Files
mempalace-toolkit/contrib
Joakim Persson 947604b25d docs: backup and recovery, plus units; move host runbook to a private repo
Adds bin/mempalace-backup and docs/backup-and-recovery.md — the mechanism a
palace actually needs, none of it site-specific.

Why a palace cannot be backed up with cp: it is chroma.sqlite3 (authoritative),
knowledge_graph.sqlite3 (usually WAL, so -wal/-shm make a plain copy a
same-instant gamble), derived HNSW segment dirs, hallways.json, the embedder
descriptor, and a HIDDEN .mempalace/origin.json. Both SQLite files are therefore
copied through the online-backup API. Two bugs are documented because both
produce a backup that looks fine: "$PALACE"/*/ silently skips the hidden dir, and
per-directory rsync collides the identically named data_level0.bin in every HNSW
segment. Treating the palace as one tree fixes both and makes a backup a faithful
palace IMAGE, so restore is a copy rather than a procedure.

Two modes: hot (default, zero downtime, ~4 s, index may lag but SQLite is
authoritative and repair --mode from-sqlite rebuilds) and cold (--cold, ~5 s
downtime, byte-consistent, restart trapped so a failed run still brings the
server back). Verification runs on the COPY — quick_check plus row counts — and
the backup is committed by mv only after it passes, with retention pruned only
after a verified commit, so a broken new backup cannot delete the last good one.

Documented because they are easy to get wrong: the sqlite3 CLI is often absent
where the Python module is present; mempalace_embedder.json must be restored with
the drawers or search silently degrades; a tested restore means running status
AND search against the restored copy, since search is what actually exercises the
index; mempalace-serve is a USER unit, so root systemctl reports "not found";
Persistent=true is what makes a missed window run after boot; and installing
against the system Python couples the palace's availability to distribution
upgrades, with the uv-managed-interpreter fix plus the two PATH traps that bite
scripted upgrades.

Also moves docs/synlig-primary-runbook.md out to a private fleet repository,
leaving a stub that explains the split, since a host inventory is operator data
for one deployment rather than part of a public toolkit. The path stays valid so
existing links do not break. Remaining host references in the README, RFCs and
ARCHITECTURE are left alone deliberately: they are load-bearing prose, contain no
secrets, and are best generalised as they are next edited rather than in one
churn-heavy pass.
2026-08-17 00:50:16 +02:00
..

contrib/ — automation recipes for mempalace-session and mempalace-pi-session

Manual invocation of the session-mining wrappers is fine on a machine you actively drive. For long-running devboxes, a weekly automated mine keeps the palace fresh without thinking about it. This directory ships ready-to-use templates for two common scheduling mechanisms, for each wrapper — plus systemd/mempalace-serve.service, which is not a mining job at all but the shared-palace server (its own section below).

pi machines: check whether you need this at all. If the pi bridge extension (extensions/pi/mempalace.ts) is installed and is ≥ 29e660e (2026-08-12), it already feeds the palace by itself on session_shutdown and a debounced agent_settled — see extensions/pi/README.md § Automatic transcript feeding. ⚠️ "Installed" is not enough — the pre-29e660e extension has no feed path at all, and a container image baked before that date ships exactly that copy. Check the deployed file, not the repo clone: grep -c MEMPALACE_FEED "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)" — zero means the templates below are not optional on that machine. As of 2026-08-14 the entire pi-devbox fleet returns zero. The templates below were written when scheduling was the only path for both harnesses; that's still true for opencode (no such extension exists), but for pi they're now a fallback — useful for a bare pi install without the bridge, a host-level catch-up job, or belt-and-braces coverage of a hard container kill (the extension's triggers don't fire on SIGKILL).

Before using either: confirm the toolkit is installed and the wrapper works — mempalace-session --dry-run (and/or mempalace-pi-session --dry-run) should list qualifying sessions. If that errors, fix the install before scheduling.

Pick one scheduler (systemd or launchd or cron). The opencode and pi jobs can be installed side by side and staggered — templates ship with Mon 03:00 for opencode, Tue 03:00 for pi to avoid racing the post-mine HNSW repair.

Templates at a glance

File What it schedules When
systemd/mempalace-session.{service,timer} opencode → palace Mon 03:00
systemd/mempalace-pi-session.{service,timer} pi → palace Tue 03:00
systemd/mempalace-session-devbox.{service,timer} opencode (inside a devbox container) → palace Mon 03:00
launchd/se.jordbo.mempalace-session.plist opencode → palace (macOS) Mon 03:00
launchd/se.jordbo.mempalace-pi-session.plist pi → palace (macOS) Tue 03:00
cron/mempalace-session.cron opencode → palace Mon 03:00
cron/mempalace-pi-session.cron pi → palace Tue 03:00
cron/mempalace-session-devbox.cron opencode (devbox) → palace Mon 03:00
systemd/mempalace-serve.service not a mining job — runs the shared palace server always-on
systemd/mempalace-backup.{service,timer} not a mining job — hot palace backup (zero downtime) daily 04:00
systemd/mempalace-backup-cold.{service,timer} not a mining job — cold palace backup (~5 s downtime, byte-consistent) Sun 04:30

The pi variants are drop-in copies of the opencode variants with script name and schedule updated; the install recipes below apply equally — just swap mempalace-session for mempalace-pi-session and the schedule day.


mempalace-serve.service — the shared palace server

The odd one out in this directory: every other template feeds a palace on a schedule, this one serves a palace over HTTP so several machines can share it (RFC-001). This unit currently runs the fleet primary on synlig — serving since 2026-08-12, seeded 2026-08-14, and the palace behind it is the only copy. Read docs/rfc-001-global-palace.md and docs/phase-1-exposure-runbook.md before installing a second one.

It is a user unit (systemctl --user), so it dies with your login session unless lingering is enabled — that is the one sudo this recipe needs:

# Install
mkdir -p ~/.config/systemd/user
cp contrib/systemd/mempalace-serve.service ~/.config/systemd/user/

sudo loginctl enable-linger "$USER"      # else the server stops when you log out
systemctl --user daemon-reload
systemctl --user enable --now mempalace-serve

# Verify — both lines matter
curl -s 172.17.0.1:8765/healthz          # -> ok
curl -s localhost:8765/healthz           # -> connection refused (exit 7), and that is CORRECT

# Logs
journalctl --user -u mempalace-serve -f

The unit refuses to start if ~/.mempalace/palace does not exist (ConditionPathExists) — better a clear failure than a server quietly creating an empty palace somewhere unexpected.

Why it binds 172.17.0.1 and not loopback, since this looks backwards and is the single most load-bearing line in the file: mempalace pins the HTTP Host header to loopback literals only on a loopback bind, so a 127.0.0.1 server behind a reverse proxy 403s every proxied request — and token auto-minting is gated on the bind being non-loopback, so the "safe-looking" loopback bind starts with no authentication at all and no warning. 172.17.0.1 is the docker0 gateway: non-loopback (so the pin relaxes and a token is minted), reachable from this host and its containers, not from the LAN. Point the tunnel/proxy at http://172.17.0.1:8765 — not https:// — because there is no --tls-cert here; TLS belongs at the proxy. Targeting https:// yields 502 from outside while local curl still says ok.

The token. There is deliberately no --token in the unit (units are world-readable). serve mints or reuses a 0600 token at ~/.mempalace/server/<sha256-prefix-of-palace-path>/token, stable across restarts. Read it from there to configure clients.

⚠️ Uninstall is where this unit differs from every other template here. Stopping it is safe; deleting its data is not.

systemctl --user disable --now mempalace-serve
rm ~/.config/systemd/user/mempalace-serve.service && systemctl --user daemon-reload

Do not rm -rf ~/.mempalace on a host that has served the fleet. That tree holds the shared palace and the only copy of the bearer token every client authenticates with. For the same reason, never rsync --delete into ~/.mempalace — the token lives inside the tree you would be syncing. And note palace directories cannot simply be moved: the directory name is a sha256 prefix of its own path.

Clients fail closed when this unit is down — they lose their palace tools entirely rather than falling back to a local palace — so stopping it is visible, reversible, and loses no data.

Operational note: the server serializes every request behind one lock, so a wedged process is a fleet-wide outage. Hence Restart=on-failure with TimeoutStopSec=30 — fail fast and let systemd recover it.

Why: runs without the user logged in (with loginctl enable-linger), survives reboots, logs to journalctl, Persistent=true catches missed runs after the machine was off. No root required — it's a user unit.

Install:

mkdir -p ~/.config/systemd/user
cp contrib/systemd/mempalace-session.service ~/.config/systemd/user/
cp contrib/systemd/mempalace-session.timer   ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now mempalace-session.timer

# Optional: keep the timer running when you log out (needed on headless servers)
sudo loginctl enable-linger "$USER"

Verify:

# Is the timer active and when will it next fire?
systemctl --user list-timers mempalace-session.timer

# Last run status + log tail
systemctl --user status mempalace-session.service

# Full run log (since today)
journalctl --user -u mempalace-session --since today

# Force a run right now (outside the schedule), for testing
systemctl --user start mempalace-session.service

Uninstall:

systemctl --user disable --now mempalace-session.timer
rm ~/.config/systemd/user/mempalace-session.{service,timer}
systemctl --user daemon-reload

What the service does

  • Type=oneshot — runs to completion, not a long-lived daemon.
  • ConditionPathExists=%h/.local/share/opencode/opencode.db — skips silently on machines that haven't used opencode (no wasted boot-time runs).
  • ConditionPathExists=!%t/mempalace-session.lock + ExecStartPre/ExecStopPost — soft mutual exclusion between overlapping runs.
  • Nice=10 + IOSchedulingClass=idle — background priority; won't interfere with interactive work.
  • TimeoutStartSec=7200 — 2 hour ceiling. The reference 60-session mine takes ~21 min; this is headroom for large corpora + slow disks.

What the timer does

  • OnCalendar=Mon 03:00 — weekly, Monday 03:00 local time. Edit to taste (see man systemd.time for syntax).
  • Persistent=true — if the machine was off at the scheduled time, run on next boot.
  • RandomizedDelaySec=30m — jitters up to 30 minutes to avoid thundering-herd across a fleet.

launchd user agent (macOS)

Why: the macOS-native equivalent of a systemd user timer. Runs without a Terminal window open, logs to ~/Library/Logs/, single-instance guarantees baked in, background-priority scheduling via ProcessType=Background. No Homebrew or third-party scheduler required.

Caveats vs. systemd:

Systemd feature launchd equivalent Notes
Persistent=true catches missed runs Partial — StartCalendarInterval fires on next system-awake time If the Mac is fully off at scheduled time, the run is skipped. Sleep-at-schedule → fires on wake.
RandomizedDelaySec=30m None native Single-user machines rarely need jitter; add a sleep $((RANDOM % 1800)) wrapper if you do.
ConditionPathExists None native mempalace-session exits cleanly when the opencode DB is missing, so no guard is strictly needed.
Lock file for overlap prevention Automatic launchd refuses to start a second instance of the same Label while one is running.

Install:

# Substitute your username into the template
sed "s|USER|$USER|g" contrib/launchd/se.jordbo.mempalace-session.plist \
  > ~/Library/LaunchAgents/se.jordbo.mempalace-session.plist

# Ensure the log directory exists
mkdir -p ~/Library/Logs

# Modern load (macOS 10.11+). "gui/$(id -u)" targets your login session.
launchctl bootstrap "gui/$(id -u)" ~/Library/LaunchAgents/se.jordbo.mempalace-session.plist
launchctl enable "gui/$(id -u)/se.jordbo.mempalace-session"

On older macOS or if you hit permissions errors with bootstrap, fall back to the legacy form: launchctl load -w ~/Library/LaunchAgents/se.jordbo.mempalace-session.plist

Verify:

# Quick check — is the job registered?
launchctl list | grep mempalace-session

# Detailed state (next run time, last exit code, throttling)
launchctl print "gui/$(id -u)/se.jordbo.mempalace-session"

# Run log tail
tail -f ~/Library/Logs/mempalace-session.log
tail -f ~/Library/Logs/mempalace-session.err.log

# Force a run right now (outside the schedule), for testing
launchctl kickstart -p "gui/$(id -u)/se.jordbo.mempalace-session"

Uninstall:

launchctl bootout "gui/$(id -u)" ~/Library/LaunchAgents/se.jordbo.mempalace-session.plist
rm ~/Library/LaunchAgents/se.jordbo.mempalace-session.plist
# Optional — keep the old logs for post-mortem, or delete them:
# rm ~/Library/Logs/mempalace-session{,.err}.log

What the plist does

  • Label=se.jordbo.mempalace-session — reverse-DNS label; shows up in launchctl list and Console.app. Change the prefix if you're forking this for a different org.
  • ProgramArguments — absolute path to mempalace-session. Template uses /Users/USER/.local/bin/mempalace-session; the install sed substitutes your actual username.
  • EnvironmentVariables.PATH — covers ~/.local/bin, Apple Silicon Homebrew (/opt/homebrew/bin), Intel Homebrew (/usr/local/bin), and system defaults. launchd agents get a minimal PATH by default, and mempalace-session needs to find mempalace + python3.
  • StartCalendarIntervalWeekday=1, Hour=3, Minute=0 = Monday 03:00. Omit any key to match "any" (e.g. drop Weekday for daily).
  • RunAtLoad=false — don't run on load/reboot, only on schedule. Flip to true if you want a run at every boot.
  • ProcessType=Background + LowPriorityIO=true + Nice=10 — macOS throttles this job's CPU and I/O so it yields to interactive work.
  • ExitTimeOut=7200 — 2h ceiling, matches the systemd unit.
  • StandardOut/ErrorPath~/Library/Logs/ is the macOS convention; Console.app picks these up automatically.

cron

Why: simpler, ubiquitous, works on any UNIX. No loginctl enable-linger dance, no user-units awareness required.

Caveats: no "persistent" semantics (a missed run while the machine was off stays missed); default cron output goes to mail or is silently dropped if no MTA. Also: cron is not installed by default on minimal Debian/Ubuntu hosts (nor on most container images). Check with command -v crontab — if absent, sudo apt install cron (or equivalent) first, or use the systemd timer instead.

Install:

# Edit the template first — replace USER with your actual username
sed "s|USER|$USER|g" contrib/cron/mempalace-session.cron > /tmp/mempalace-session.cron

# Append to your existing crontab (preserves any entries you already have)
(crontab -l 2>/dev/null; cat /tmp/mempalace-session.cron) | crontab -
rm /tmp/mempalace-session.cron

# Verify
crontab -l | grep mempalace

Ensure ~/.cache/mempalace-logs/ exists so the log file can be written:

This is the log directory only. The staging dir — the transcripts the palace keys its source_file dedup on — lives beside the palace (<palace-root>/opencode-stage/), not in ~/.cache, precisely so it cannot be cleaned away while the palace survives. Logs here are disposable.

mkdir -p ~/.cache/mempalace-logs

Verify a run is happening:

# Tail the log the cron entry writes to
tail -f ~/.cache/mempalace-logs/cron.log

# Or force a run manually to prove the command is well-formed
mempalace-session

Uninstall:

crontab -e   # remove the mempalace-session line by hand

Which should I pick?

Situation Pick
Desktop / laptop, mempalace installed directly on the host, modern systemd-based Linux systemd/mempalace-session.{service,timer}
macOS (any recent version), mempalace on the host launchd/se.jordbo.mempalace-session.plist
Long-running Linux devbox or server, mempalace on the host systemd/mempalace-session.{service,timer}
opencode-devbox (or similar) container with mempalace inside it systemd/mempalace-session-devbox.{service,timer} (preferred) or cron/mempalace-session-devbox.cron (simpler)
BSD, Alpine, or Linux distro without systemd cron/mempalace-session.cron
You already have a cron-based job scheduler on the box any cron/*.cron template
You want logs in journalctl (Linux) or Console.app (macOS) rather than a file systemd user timer / launchd

If you're not sure: systemd on Linux, launchd on macOS, cron only when neither is available. Use the -devbox variant when mempalace lives inside a long-running container rather than on the host. All templates wrap the same mempalace-session command — the difference is purely in where the scheduler lives and whether it needs to docker exec into a container to reach the tool.


Running inside a container (devbox)

If you run opencode inside a long-lived container like opencode-devbox, neither systemd nor cron is running inside that container — they're host-level services. The correct pattern is host-side scheduling that docker execs into the running container. The *-devbox templates in contrib/systemd/ and contrib/cron/ implement exactly this.

Preconditions:

  • Long-lived container. docker compose up -d with restart: unless-stopped (or equivalent). If the container is ephemeral (per-invocation), this pattern doesn't apply.
  • mempalace-session is already installed inside the container. opencode-devbox bakes it in from v1.14.30b onward (see the INSTALL_MEMPALACE_TOOLKIT arg). For earlier images or custom containers without the toolkit installed, run ./install.sh --yes from a mempalace-toolkit checkout inside the container first — the installer is idempotent and safe to re-run. Note that a manual install lives in the container's ephemeral layer and is lost on docker compose up --force-recreate, so the bake-in approach is the durable solution.
  • Host user can talk to docker. Member of the docker group on Linux, or Docker Desktop running under the current login session on macOS.
  • Canonical container name is opencode-devbox. If you renamed it via container_name: or docker-compose project naming, adjust CONTAINER in the template.

Both devbox templates guard against "container currently stopped" — they no-op silently if docker ps shows no running container with the expected name. That makes the timer safe to leave enabled across dev cycles where you tear the container down and bring it back up.

systemd user timer (host-side, devbox variant)

# Install
mkdir -p ~/.config/systemd/user
cp contrib/systemd/mempalace-session-devbox.service ~/.config/systemd/user/
cp contrib/systemd/mempalace-session-devbox.timer   ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now mempalace-session-devbox.timer

# Keep the timer running across logout (typical on dev hosts)
sudo loginctl enable-linger "$USER"

Customize container name / user (if you don't use the defaults):

systemctl --user edit mempalace-session-devbox.service
# In the override that opens, set:
#   [Service]
#   Environment=CONTAINER=my-devbox-name
#   Environment=CONTAINER_USER=my-user

Or edit the shipped service file in place before copying.

Verify:

systemctl --user list-timers mempalace-session-devbox.timer
systemctl --user status mempalace-session-devbox.service
journalctl --user -u mempalace-session-devbox --since today

# Force a run right now (while the container is up)
systemctl --user start mempalace-session-devbox.service

Behaviour when the container is down: ExecCondition fails, service is marked "condition failed" (considered a successful no-op by systemd), no alert noise. Bring the container back up and the next scheduled fire will run normally.

Uninstall:

systemctl --user disable --now mempalace-session-devbox.timer
rm ~/.config/systemd/user/mempalace-session-devbox.{service,timer}
systemctl --user daemon-reload

cron (host-side, devbox variant)

# Read the template — it has CONTAINER / CONTAINER_USER at the top.
# Adjust if your setup differs from opencode-devbox defaults.
cat contrib/cron/mempalace-session-devbox.cron

# Install (preserves existing crontab entries)
(crontab -l 2>/dev/null; cat contrib/cron/mempalace-session-devbox.cron) | crontab -

# Ensure the log directory exists
mkdir -p ~/.cache/mempalace-logs

# Verify
crontab -l | grep mempalace-session-devbox
tail -f ~/.cache/mempalace-logs/cron-devbox.log

Uninstall:

crontab -e   # remove the mempalace-session-devbox line by hand

When to pick devbox variants vs direct ones

Setup Templates to use
mempalace installed directly on the host (no devbox) mempalace-session.{service,timer}, mempalace-session.cron, se.jordbo.mempalace-session.plist
opencode-devbox container is where opencode + mempalace live mempalace-session-devbox.{service,timer}, mempalace-session-devbox.cron
Both (rare — mempalace on host AND inside a separate devbox) Install both; they write to separate palaces.

The in-container mempalace sees only the container's opencode.db and palace (via named volumes). The host's mempalace, if installed, sees only the host's. Two parallel palaces don't cross-pollinate — decide where you want the source of truth to live and schedule accordingly.

Alternative not documented here

"Run systemd inside the container" is technically viable (systemd-in-docker images exist) but adds non-trivial complexity for the sake of a once-weekly batch job. The host-scheduled approach above is equivalent in outcome and much simpler. Skip systemd-in-container unless you already have other reasons to need it.


Tuning

Frequency. Weekly is the default because:

  • New sessions you care about are typically a handful per week per user.
  • Dedup is free on unchanged sessions, so there's no cost to running daily other than the ~5 min post-mine repair.
  • Weekly keeps the palace fresh enough that searches almost always return current context.

Scheduling cadence: edit OnCalendar= or the cron DOW field. Post-mine repair is now opt-in (--repair) and should NOT be added to unattended schedules — the in-place HNSW rebuild has wiped live palaces on past runs. Run mempalace repair manually from a quiet interactive session if you ever need it.

Monthly: probably too infrequent. You'll search for "that thing we discussed last Tuesday" and miss it.


See also

  • ../../ARCHITECTURE.md §5 — operational routine (triggers, cadence) in full context.
  • ../../SKILL.md — the agent-side Operational Routine section for when an AI agent should suggest running this.