# pi ↔ MemPalace MCP bridge The canonical source of `~/.pi/agent/extensions/mempalace.ts` — the TypeScript extension that wires [MemPalace](https://github.com/MemPalace/mempalace)'s MCP server into the [pi coding-agent](https://github.com/earendil-works/pi) harness. Installs wake-up context injection, per-tool schema passthrough, and a `/mempalace-diary` slash-command. This directory **only** holds the bridge. Pi's own base config (keybindings, environment loader, settings template) lives in the sibling [`pi-toolkit`](https://gitea.jordbo.se/joakimp/pi-toolkit) repo — split out 2026-05-05 so [`opencode-devbox`](https://gitea.jordbo.se/joakimp/opencode-devbox) can build slim containers that include pi without dragging in mempalace's dependencies (~300 MB). **Jump to:** - [What it does](#what-it-does) - [Transport: local vs external](#transport-local-vs-external) - [Automatic transcript feeding](#automatic-transcript-feeding) - [The `Type.Unsafe` gotcha](#the-typeunsafe-gotcha) - [Deploying pi with mempalace on a new machine](#deploying-pi-with-mempalace-on-a-new-machine) - [Fail-soft, identity, debugging](#fail-soft) --- ## What it does 1. **Connects to MemPalace** and does the MCP handshake (`initialize` + `notifications/initialized` + `tools/list`). By default it **spawns `mempalace-mcp`** as a local stdio subprocess (`StdioMcpClient`); if `$MEMPALACE_REMOTE_URL` is set it instead talks to a shared MemPalace over HTTP (`RemoteMcpClient`) and spawns no local process — see [Transport](#transport-local-vs-external). 2. **Registers each MCP tool** as a pi tool with its real `inputSchema` passed through via `Type.Unsafe(...)` (see gotcha below). 3. **Wake-up auto-injection** (`before_agent_start`, one-shot per fresh session): calls `mempalace_status` + `mempalace_diary_read` and injects the result as a `mempalace-wakeup` system message so the agent orients itself the way `~/.agents/skills/mempalace/SKILL.md` describes. Skipped on resume/fork (context is already in the thread). 4. **Automatic transcript feeding** (`session_shutdown`, and a debounced `agent_settled`): stages + mines this pi installation's own session transcripts into the palace with no user action needed — **as of mempalace-toolkit `29e660e` (2026-08-12); see the version gate below, because "the extension is installed" does not imply "this copy can feed"**. Unlike the diary below, this needs no LLM turn — it's a subprocess + a tool call — so it *can* run on `session_shutdown` where the diary cannot. See [Automatic transcript feeding](#automatic-transcript-feeding). 5. **Manual wind-down** via a `/mempalace-diary [topic]` slash command: sends a prompt asking the LLM to call `mempalace_diary_write` with an AAAK-formatted entry summarizing the session. This one stays manual because it needs the LLM to compose the entry, and `session_shutdown` fires too late to drive another LLM turn — a constraint that applies to the diary specifically, not to feeding (see above). 6. **Stamps device provenance on every write** — `@` as the writer on `add_drawer`/`checkpoint`/`mine`/`event_append`/`artifact_put`, and a `HOST:|` prefix on diary entries — as of mempalace-toolkit `553d8657` (2026-08-25). See [Identity](#identity). ## Automatic transcript feeding > **⚠️ Version gate — requires mempalace-toolkit ≥ `29e660e` (2026-08-12), and "installed" is not > the same question as "capable".** Feeding was added to this extension on 2026-08-12. A copy baked > into a container image built before that date has *no* feed path at all — its entire > `session_shutdown` handler is `client.stop()` — and it fails the only way a memory system must not: > silently, looking exactly like a healthy run with nothing to do. > > **Check the deployed artifact, never the repo.** `/opt/*` in an image is baked at build time and can > be days behind a bind-mounted clone, and `~/.pi/agent/extensions/mempalace.ts` is usually a symlink > *into* that baked copy: > > ```sh > grep -c MEMPALACE_FEED "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)" # 0 = cannot feed > ``` > > Zero hits means this machine needs the fallback recipes in > [`contrib/`](../../contrib/README.md) until it is rebuilt, regardless of what the toolkit repo's HEAD > looks like. Date the deployed copy with `stat` plus that content probe — not `git log`, which fails > with *"detected dubious ownership"* inside a root-owned `/opt` tree. As of 2026-08-14 the whole > pi-devbox fleet fails this check. The bridge feeds this pi installation's own session transcripts into the palace by itself — no scheduler, no cron, no manual invocation. It fires on `session_shutdown` (covers quit, `/new`, `/resume`, `/fork`) and on a debounced `agent_settled` (covers a long session that later crashes, since a hard kill runs no shutdown handler at all). The work is split across two processes, and the reason is a hard constraint, not a style choice: **the palace is single-writer.** A live pi session always holds it through this extension's own `mempalace-mcp` subprocess, so an unattended `mempalace mine` from anywhere else fails outright with `palace ... is held by PID `. The bridge therefore: 1. Runs `mempalace-pi-session --prepare --reason --wing ` as a subprocess. This does every palace-free step — parse pi's JSONL, apply the quality threshold, stage the export, and (remote mode only) `rsync` it to the palace host — and prints one line, `MINE_SOURCE=`, without ever touching the palace. 2. Calls the `mempalace_mine` MCP tool **through this extension's own client** on that path. Going through the client that already holds the lock is the only way to write during a live session, and it automatically targets whichever palace the bridge is pointed at — local stdio or a shared remote one. `mempalace-pi-session` (in this repo's `bin/`) is the actual exporter and owns the quality gate, the remote transport, and every flag — see its `--help` for the full reference; this section only covers the extension's side of the wiring. **Env knobs (extension side):** | Var | Default | Effect | |---|---|---| | `MEMPALACE_FEED` | `1` | Set `0` to disable automatic feeding entirely. | | `MEMPALACE_FEED_BIN` | `mempalace-pi-session` | Helper to run. | | `MEMPALACE_FEED_WING` | `wing_conversations` | Target wing — passed to both the exporter and the `mempalace_mine` call. | | `MEMPALACE_FEED_DEBOUNCE_MS` | `600000` (10 min) | Minimum gap between mid-session (`agent_settled`) feeds. Bounds crash loss to one window instead of a whole session. | | `MEMPALACE_FEED_PREPARE_TIMEOUT_MS` | `120000` | Kills a wedged `--prepare` subprocess. | | `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `300000` (5 min) | Bounds how long the extension *waits* for `mempalace_mine`, so a stalled palace can't hang session exit. It does **not** cancel the mine — see [Debugging](#debugging). Raised from `30000` in 2026-09: the mine is the slowest call this extension makes (30–60s in normal operation), so the old deadline fired routinely and reported healthy behaviour as an error. | **Remote palace:** if `$MEMPALACE_REMOTE_URL` is set (see [Transport](#transport-local-vs-external)), `mempalace_mine`'s source path is expanded on the *server*, which cannot see this machine's transcripts — that's exactly why step 1 above rsyncs first in that mode. Configure the inbox with `MEMPALACE_PI_SSH_TARGET` (required for remote feeding — feeding is silently skipped without it), `MEMPALACE_PI_SSH_CONFIG`, and `MEMPALACE_PI_REMOTE_PATH`; see `mempalace-pi-session --help`. **Concurrency:** overlapping triggers coalesce — a `session_shutdown` landing while a debounced tick is still running joins that in-flight feed instead of racing it. `mempalace-pi-session` itself also takes a non-blocking `flock`, so even two independent invocations (e.g. this extension and the container-start catch-up some devbox images run) never race each other; losing that race is harmless because the next trigger re-exports from scratch. ## Transport: local vs external The bridge speaks the same MCP protocol over two interchangeable transports, chosen at load time: - **Local (default)** — spawns `mempalace-mcp` as a stdio subprocess; the palace lives wherever that process opens it (default `~/.mempalace`). This is the hardened path with per-request timeouts and respawn/self-heal (below). - **External** — set `MEMPALACE_REMOTE_URL` to a MemPalace HTTP endpoint (e.g. `https://mempalace.jordbo.se/mcp`, the live fleet primary — full path including `/mcp`, no trailing slash) and the bridge connects over HTTP instead, spawning no local process. Use this to share **one** palace across several harnesses/containers (pi + opencode + native). `MEMPALACE_REMOTE_TOKEN`, if set, is sent as `Authorization: Bearer `. Use `https://` for anything crossing a network — the plaintext `http://` example that stood here until 2026-08-14 predated the reverse proxy. The two transports are **either/or**, decided once at load time: with the URL set, writes go **only** to the remote palace. There is no dual-write, no local mirror, and no local `mempalace-mcp` process at all. ⚠️ **Consequence: once `MEMPALACE_REMOTE_URL` is set, the `mempalace` CLI on that machine is no longer a valid way to inspect or feed the palace the agent is using.** The CLI has no remote support whatsoever — its only selector is `--palace ` — so it reads and writes the LOCAL on-disk archive. After a flip that archive is frozen, yet `mempalace status` / `mempalace search` still report a plausible drawer count and look exactly like success: a false-positive machine. Memories filed with the CLI post-flip land in the dead archive, not in the shared palace. Use the agent's own palace tools (which go over HTTP), and mine backfills **on the palace host**. **Setting these for a *native* pi install — there is no `.env` to edit.** Worth stating plainly, because the obvious guess is wrong: **pi loads no dotenv file and has no `env` block in `settings.json`**, and this extension reads `process.env` and nothing else (`createClient()` → `process.env.MEMPALACE_REMOTE_URL`). A native install therefore inherits whatever **the shell that launches `pi`** exports — that is the only hook. So export them from your shell rc, or from a file it explicitly sources: ```sh # ~/.zshrc (or ~/.bashrc) — if you keep secrets in ~/.config/pi/.env, # nothing sources it for you; do it yourself: set -a; [ -f ~/.config/pi/.env ] && . ~/.config/pi/.env; set +a ``` Two traps. A pi launched from a **GUI** (Spotlight, dock, an editor's terminal that spawns a non-login shell) does not necessarily read that rc, so it can silently stay on the local palace. And exporting the variable in the shell where you *edited* the rc does not affect an already-running pi — the transport is chosen once at extension load. Confirm the result the same way as a container flip: ask the agent for `mempalace_status` and check the reported palace path is the **remote** host's, not your own `$HOME/.mempalace/palace`. That one check is the whole verification: a half-flipped client reports a local path while looking healthy. Serve such an endpoint with `mempalace serve --host 172.17.0.1 --port 8765` (the `pi-devbox` / `opencode-devbox` repos ship a `docker-compose.mempalace.yml` for exactly this). **The HTTP transport is authenticated as of mempalace 3.6.0** — earlier docs here said otherwise, from the v1.3.0 era. `serve` mints a bearer token, keeps it 0600, passes it via the environment (never argv), compares it with `hmac.compare_digest`, and **refuses to bind a non-loopback host without one** unless `--allow-insecure`. It also pins `Host` and allowlists `Origin` (anti-DNS-rebinding), and can terminate TLS itself. Two binds to avoid. `0.0.0.0` publishes the palace to the whole LAN. And `127.0.0.1` is the trap that looks safe: the Host pin is enforced *only* on loopback binds, so behind a tunnel every proxied request 403s — and token auto-minting is gated on the bind being non-loopback, so it starts with **no authentication at all**, no warning. Bind the docker0 gateway (`172.17.0.1`): reachable from the host and its containers, not from the LAN. Binding an interface is not the same as exposing a palace — an MCP endpoint also pins `Host`/`Origin`, so a request arriving under the wrong hostname is refused even when the port is open. Implementation note: the HTTP client (`RemoteMcpClient`) is **vendored** from [`pi-extensions`](https://gitea.jordbo.se/joakimp/pi-extensions)' `mcp-loader.ts`. A `MCP-STREAMABLE-HTTP-CLIENT-SYNC` token keeps the two copies from drifting — [`scripts/check-mcp-client-sync.sh`](../../scripts/check-mcp-client-sync.sh) fails if they diverge (it skips gracefully when the `pi-extensions` checkout isn't present). ## Fail-soft If `mempalace-mcp` can't be spawned (PATH missing, binary crashes at startup, …) the extension logs to stderr and returns early. pi keeps working without palace tools rather than refusing to start. **In remote mode the triggers differ but the outcome is identical.** An unreachable server, a DNS failure, or an HTTP 401 from a wrong/expired token all end the same way: after bounded retries the extension prints `mempalace-mcp unavailable after retries; continuing without palace tools` and **does not register the palace tools**. It is **fail-closed, not fail-local**: it does *not* quietly fall back to the local palace, so a remote outage can never scatter memories into a local copy nobody will look at again. The practical corollary, worth knowing before you debug the wrong layer: **"the agent has no `mempalace_*` tools" is the expected symptom of a server, token, or DNS fault**, not of a broken install. Diagnose it with a direct `curl` to `MEMPALACE_REMOTE_URL`, and confirm the flip with `mempalace_status` — the reported palace path must be the remote host's. The design rationale for de-registering rather than degrading is in [`docs/rfc-001-global-palace.md`](../../docs/rfc-001-global-palace.md) §2 and §4.1. ## Identity `agent_name` for diary calls comes from `$MEMPALACE_AGENT_NAME`, defaulting to `"pi"`. First diary write against that identity creates `wing_` in the palace. Set the env var if you want to run pi under a distinct identity on a given machine (e.g. `pi-laptop` vs `pi-server`). ### Device provenance is stamped here, at the edge As of mempalace-toolkit `553d8657` (2026-08-25) the bridge fills in *who wrote this* so the agent never has to: | Write | What the bridge sets | |---|---| | `add_drawer`, `checkpoint`, `mine` | `added_by = "@"` when the caller left it unset | | `event_append`, `artifact_put` | `from_agent` / `created_by` likewise | | `diary_write` | prefixes the entry text with `HOST:\|` | `` is `$MEMPALACE_PI_DEVICE`; `` is `pi`. **Both gates must hold:** `MEMPALACE_PI_DEVICE` set *and* `MEMPALACE_REMOTE_URL` pointing at a shared palace. A solitary palace stamps nothing, because there is no second machine to disambiguate from and the annotation would be pure noise. Two design points worth not re-litigating: - **The diary marker is in the entry text, not metadata.** `diary_read` returns content only, so metadata is invisible to the agent that later reads the entry — an attribution nobody can see is not an attribution. Search results are built from a fixed key list with the same consequence. - **RFC 001 §7.3.2 ranks "the agent stamps it via a skill instruction" as the worst available option**, and it was: the agent that wrote that instruction into the consumer skill then filed its own provenance drawer as `added_by=checkpoint`. Per-call boilerplate gets forgotten. Hence the edge. Callers keep two responsibilities the bridge cannot infer: pass `source_drawer_id` on `kg_add` (triples have no provenance field at all), and pass an explicit writer **only** when deliberately filing on behalf of another device. Because the extension is baked into an image, *a container older than the stamping commit satisfies both gates and still stamps nothing* — the env vars are set and the code is simply absent. The one-line check: ```bash grep -c MEMPALACE_PI_DEVICE "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)" ``` ## Agent coordination over the logstream The palace also carries an append-only coordination log (RFC 003: `mempalace_event_*`, `mempalace_artifact_*`) used for cross-machine delegation, review, patch handoff and retraction. Relevant to this extension in three ways: **1. The bridge makes directed addressing possible.** Where every machine is a thin MCP client of one shared palace, all clients report the *same* `origin_replica`, so the log cannot tell two machines apart by transport identity — `from_agent` / `to_agent` carry the entire distinction. Measured 2026-08-26 from the tor-ms22 client: `mempalace_mesh_peers` returned `peers: []` with a single replica id authoring every event from every machine. So the stamping above is what makes `to_agent="pi@tor-ms22"` mean anything, and an *unstamped* client addressed as bare `pi` is unreachable. **2. The bridge now READS the log too — auto-delivered mailbox.** It derives what this device owes and injects it, so an event addressed to this machine no longer waits for the agent to think of asking. Two delivery points, both fail-silent: - **Session start** — one more section in the existing `before_agent_start` wake-up injection, alongside `mempalace_status` and `diary_read`. - **Mid-session** — an `agent_settled` poll, floored at `MEMPALACE_MAILBOX_POLL_MS` (default 300000, i.e. 5 min), delivering via `pi.sendMessage(..., { deliverAs: "steer" })`. Note this **queues, it does not interrupt**: at `agent_settled` the agent is idle, so the message lands at the start of the next turn and spends no LLM call. There is deliberately no `triggerTurn` — waking the model on inbound fleet traffic is a much larger behavioural change than auto-delivery. **Which means a human is the trigger, and the mailbox now says so.** Measured 2026-08-26 on two devices: a delivery lands, the agent is idle, nothing happens, and the operator asks *"do I have to nudge you for you to read this?"*. Yes — because between the poll and the next turn no inference is running. The old text explained how to *close* an ask and never said when it would be *seen*, so the only reader who needed that fact was the one not told. Two additions, neither of which touches the no-`triggerTurn` decision: - **A delivery note in the message itself** — states that this is a queued message, that nothing woke the agent, and that any message starts the turn that handles it. Free, and aimed at the human reading the window. - **A notification at poll time**, because the note only helps someone who is already looking, and the case that loses an ask is nobody looking: | `MEMPALACE_MAILBOX_NOTIFY` | Behaviour | |---|---| | *unset* (default) | in-TUI `ctx.ui.notify`, the same surface `session_start` already uses | | `desktop` | additionally a terminal-native notification — Kitty `OSC 99`, else `OSC 777` (iTerm2, WezTerm, Ghostty, rxvt-unicode) | | `0` / `off` | silent; mailbox still delivers | The `desktop` path is how a notification escapes a container without `notify-send`, DBus or any host access: the escape sequence is written to stdout and interpreted by the terminal emulator on the human's own machine. It is opt-in because writing raw escapes is a behaviour change on a shared machine, not because it is unreliable. Title and body are stripped of `;` and control bytes, so a payload can neither forge an OSC field nor end the sequence early. **⚠️ Not yet observed firing through tmux (2026-08-26).** If pi runs inside a multiplexer — e.g. `kitty → tmux on the host → docker exec → pi` — tmux drops OSC sequences it does not implement, so the ping can vanish silently between the container and the human. Reaching the outer terminal needs tmux's DCS passthrough plus `allow-passthrough on`, which is not implemented here yet. Note the client can detect *neither* layer: `KITTY_WINDOW_ID` is not forwarded by `docker exec`, and `TMUX` is unset because tmux runs one level further out. Verify with a hand-written sequence in your own stack before trusting it. It fires **only when something is due** — the same condition as the delivery itself. A ping on an empty poll would train its reader to ignore it, which is the failure this whole feature exists to reverse. Owed-ness is **derived, never read off a field**, because `event_ack` appends and `status` is written once: a directed `open` event matches the mailbox query *forever*, answered or not. Two calls (`to_agent= status=open`, and `from_agent=`), then a candidate counts as answered only when one of this device's own events is **strictly later**, joins via `metadata.ack_of` or a shared `correlation_id`, *and* carries a terminal status (`applied`/`superseded`/`failed`/`blocked`). The ordering test is load-bearing: without it one terminal reply suppresses every later ask on that correlation forever. "Strictly later" means `hlc` when both events carry one — a hybrid logical clock rendered fixed-width, so a string comparison is a causal comparison across replicas — falling back to `seq` only when either side lacks an `hlc`. `seq` is this database's arrival `rowid`, so on a mesh the same event has a different `seq` per replica and a reply can arrive before its ask. Never `created_at`: it is second-precision, and a tie there can suppress an *unanswered* ask, which is the one failure this derivation exists to prevent. `*` broadcasts are excluded even though `to_agent=` matches them, because the protocol says a broadcast owes nobody a reply — which also means broadcasting an ask demonstrably reaches no owed set, giving that documented anti-pattern teeth. **Dormant asks — open, owed, and deliberately not announced.** Some asks are planted *unanswerable on purpose*: a first-boot acceptance describes work that becomes possible only when the device is next recreated, and it must stay owed until then, because closing it early to tidy the mailbox is exactly how the work gets lost. Derivation alone cannot tell that apart from neglected work, so such an ask was announced at every session start and every poll for as long as it was correctly waiting — measured at three days running on `emb-7kj4vr4g` (2026-09-28 → 2026-10-01), and across three consecutive releases before that. The ask was right; announcing it was wrong, and the cost landed on the human reading the window, the one reader who cannot filter it. An ask may therefore declare, in its own `metadata`, the condition under which it is merely waiting: ```json "dormant_unless": [ { "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json", "field": "release_tag", "baseline": "v1.9.4" }, { "kind": "file_mtime", "path": "/etc/hostname", "baseline": "2026-09-22T18:12:49Z" } ] ``` It is dormant while **every** condition still matches its baseline, and goes live the moment **any** of them differs — the trigger such asks already stated in prose ("act when EITHER differs"), now in a form the bridge can check. Two kinds, both local-file-only: `file_mtime` (compared at whole seconds in UTC, because a filesystem mtime carries sub-second residue that a reported ISO baseline never will — comparing raw milliseconds would mark every such predicate permanently "changed" and silently disable the feature) and `json_field` (dotted paths allowed, compared as strings so a manifest holding `3` matches a baseline of `"3"`). There is no expression language, no shell and no network: a general evaluator in the path that decides whether work is *visible* is a far worse trade than a clumsy schema. Dormancy is **withheld from the announcement, never from the mailbox**: the wake-up injection lists dormant asks once per session with their ids, and a mid-session poll mentions only a count, and only when the window is already open for something else. They are not added to the resurface map, so one becomes announceable the instant its baseline moves. **Dormancy must be proven, never assumed** — every unevaluable predicate shows the ask as owed. A missing file, an unreadable one, a baseline that will not parse, an unknown `kind`, a vanished field, a relative path, more than eight conditions: each announces. The dangerous failure here is not a spurious nag but work that disappears because a predicate could not be evaluated, which would be indistinguishable from the ask being lost and would not surface until a release needed it. An ask with no `dormant_unless` behaves exactly as it did before the feature existed. `scripts/test-dormancy.sh` exercises all of it, positive and negative arms both. Gated on `MEMPALACE_PI_DEVICE` **and** `MEMPALACE_REMOTE_URL` (the same pair as the stamper, since an unstamped client has no address to be reached at), and disabled outright with `MEMPALACE_MAILBOX=0` (notifications alone with `MEMPALACE_MAILBOX_NOTIFY=0`). Inert on a solitary palace: no calls, no injection. Delivery is the mechanism; the *norms* — what a reply owes, and that only a terminal event closes a thread — remain normative in the *consumer* skill (`~/.agents/skills/mempalace/SKILL.md`). **This file documents the mechanism; the skill is normative for behaviour.** **3. Live push is a deployment question, not a code one.** The palace implements an SSE endpoint (`GET /logstream/stream`, `text/event-stream` in `mempalace/mcp_server.py`), but a deployment may expose only the MCP endpoint through its reverse proxy — verified 2026-08-26 against `https://mempalace.jordbo.se`, where `/logstream/stream` and `/sync/peers` return 404 while `/mcp` serves normally. Where that is the case, polling through the existing MCP client is the only available path — which is what the mailbox in §2 does — and enabling SSE means a proxy route plus an auth decision, not an extension change. > ⚠️ **Corrected 2026-08-26.** An earlier revision of this paragraph listed > `/logstream/events` alongside those two as proxy-blocked. That route **does not > exist in the server at all** — the complete GET table in mempalace 3.8.0 is > `/healthz`, `/statusz`, `/logstream/stream` and `/sync/{version_vector,ops,artifact,peers}`, > so `/logstream/events` would 404 against a directly-reachable server too. The > proxy inference was right for the other two and wrong for that one; see > [RFC 003 §7.8](../../docs/rfc-003-coordination-log.md). Distinguish *route absent* > from *route blocked* before blaming infrastructure. As with stamping, all of this is inert unless `MEMPALACE_REMOTE_URL` points at a shared palace. On a solitary palace the event tools work fine and the log contains only this machine's own events. ## Stall protection (per-request timeout) Every JSON-RPC request to `mempalace-mcp` carries a timeout. Without it, a wedged server (classically: an OrbStack/virtiofs cold-open of a large `chroma.sqlite3` or an HNSW load) leaves the awaiting promise pending *forever*, which freezes the pi TUI — ESC cancels the LLM stream, not a pending tool `execute()`. On timeout the extension rejects the request **and** kills the stalled child (SIGTERM→SIGKILL), so pi gets a clear error instead of hanging. This is a per-REQUEST timeout, not a process-lifetime one — the long-lived server is only killed when a request genuinely stalls. - `MEMPALACE_MCP_TIMEOUT_MS` — tool-call/request timeout. Default `60000`. Kept short on purpose: a *query* taking this long is genuinely wedged. The one call that is not a query — the feed's `mempalace_mine` — passes its own deadline (`MEMPALACE_FEED_MINE_TIMEOUT_MS`) down to the transport per call, so this default does not apply to it (since 2026-09-18; see below). - `MEMPALACE_MCP_INIT_TIMEOUT_MS` — `initialize` + `tools/list` handshake timeout. Default `300000`. Deliberately generous: a genuine first cold-open over virtiofs can legitimately take minutes, and killing a still-progressing init only to respawn and re-pay the same cold cost is strictly worse than waiting. - Set either to `0` to disable (legacy unbounded behavior). ### Self-heal (respawn instead of a permanent latch) A stall-kill (or any crash) used to be a **permanent** latch: `available` flipped off and stayed off until you restarted pi. It is now self-healing — the next tool call transparently respawns `mempalace-mcp` and retries. - Respawns use **capped exponential backoff** so a persistently-broken server can't hot-loop: `MEMPALACE_MCP_MAX_RESPAWNS` attempts (default `2`; set `0` to disable self-heal and keep the old fail-fast latch), with `MEMPALACE_MCP_RESPAWN_BACKOFF_MS` (default `1000`) doubled per attempt. - The budget **resets on any successful JSON-RPC response** — proof the server is actually live — so a server that recovers regains full patience, while one that keeps dying hits the cap and stays down (then restart pi). - Why the long init timeout and bounded respawn compose rather than overlap: once a server has opened the palace once, the OS page cache is warm, so respawn cold-opens are fast. The long init timeout prevents killing a healthy *first* cold-open; the respawn handles a genuinely dead server cheaply afterwards. (Note the HNSW deserialize is CPU work that isn't page-cacheable across spawns, which is exactly why we can't rely on respawn-warming alone and keep the generous init budget.) - The initial startup is tolerant too: if the very first `start()` fails, the extension runs the same bounded respawn before falling back to fail-soft (pi keeps working without palace tools). ## Debugging - `MEMPALACE_EXT_DEBUG=1` — surface `mempalace-mcp` stderr into pi's stderr. Without this, stderr is drained silently so a misbehaving server doesn't flood the TUI. - If a tool call fails with a generic "Internal tool error", spawn `mempalace-mcp` manually with raw JSON-RPC on stdin to read the server-side error — much faster than guessing. ### `feed (tick) failed: mine timed out after …ms` **Nothing has been lost.** The deadline bounds only how long the extension *waits*; it cannot cancel the mine, which continues on the server. The transcript is already staged before the mine is invoked, and `mempalace mine --mode convos` dedups by `source_file` and is idempotent, so the work either completed after the deadline or is redone by the next tick. **Do not "fix" it by retrying harder from the client.** The palace is a single writer; a blind retry is what turns one slow mine into a queue of them. Before 2026-09 this message appeared many times per session, which made it look like a persistent failure. That was a real defect, now fixed: `lastFeedAt` was recorded only after a *successful* wait, so a timeout left the debounce clock stale and every following settled turn started another mine on top of the one still running. Two changes — recording the attempt before the wait, and raising the deadline to sit far above the slowest honest completion — mean a healthy fleet should now never see it. If you *do* still see it, it is now informative rather than noise: a mine exceeded five minutes. Check palace size and whether another writer (a host-side feeder, a scheduled `mine`) is holding the write lock, rather than raising the timeout again. ### `feed (tick) failed: mempalace remote request 'tools/call' failed: timed out after 60000ms` Same event, different deadline — and the same reassurance: **nothing has been lost**, the mine continues server-side. This is what the previous message turned into after the 2026-09 change, and it exposed that the change was incomplete. The feed's five-minute deadline was only *raced* against the call; `callTool()` had no way to carry it, so the transport's generic per-request timeout (`MEMPALACE_MCP_TIMEOUT_MS`, 60 s) fired first on every honest 60 s+ mine, and the 300 s was unreachable. On the stdio transport this was worse than noise: a per-request timeout there kills the server child, so the mine really was aborted at 60 s. Fixed 2026-09-18: `callTool(name, args, { timeoutMs })` passes a per-call deadline to both transports and the feed uses it for the mine. `scripts/test-mcp-call-timeout.sh` pins the contract (a plain call still honours the short default; the override is honoured and is itself a deadline). Seeing this message on a fixed build means a mine exceeded *five* minutes — treat it as the previous section says. ## The `Type.Unsafe` gotcha Earlier versions of this extension registered every MCP tool with `parameters: Type.Object({}, { additionalProperties: true })`, which discarded each tool's real `inputSchema`. The LLM then saw no parameter names and had to guess, leading to bugs like `mempalace_diary_read` being called with `agent=` instead of the required `agent_name=` and crashing the Python server with `TypeError: missing 1 required positional argument`. The fix (≈ lines 160-170) is to wrap the incoming JSON Schema with `Type.Unsafe<...>(tool.inputSchema)`. TypeBox schemas are plain JSON Schema at runtime plus a `Symbol` marker, so wrapping an externally-sourced schema with `Unsafe` is sufficient — no conversion to a full TypeBox tree is needed, and the LLM now sees every tool's real parameter names. If you ever need to re-loosen the schema for debugging, fall back to the `Type.Object({}, { additionalProperties: true })` default only for that specific tool, not globally. --- ## Deploying pi with mempalace on a new machine This is the "pi + memory" recipe. For pi without mempalace, see [`pi-toolkit`'s README](https://gitea.jordbo.se/joakimp/pi-toolkit/src/branch/main/README.md#deploying-pi-on-a-new-machine). ### 0. Prerequisites - Shell: zsh + oh-my-zsh recommended (both toolkits install loaders into `~/.oh-my-zsh/custom/`; bash works too, installers print the manual `source` snippet). - `git`, `node` ≥ 20, `uv`, `tmux` ≥ 3.2, pi installed upstream. - AWS credentials reachable via `AWS_PROFILE` — only if using `amazon-bedrock` as pi's provider. ### 1. Dotfiles (if you keep one) Brings `~/.config/pi/.env` (AWS creds, git-crypt encrypted), tmux CSI-u extended keys, and other machine state: ```bash git clone ~/src/dotfiles cd ~/src/dotfiles git-crypt unlock ./provision.sh --profile # or your equivalent tool ``` ### 2. Install pi upstream ```bash brew install pi-coding-agent # macOS # or see https://github.com/earendil-works/pi for Linux pi --help # creates ~/.pi/agent/ ``` ### 3. Install pi-toolkit (base pi config) ```bash git clone ssh://git@gitea.jordbo.se:2222/joakimp/pi-toolkit.git ~/pi-toolkit cd ~/pi-toolkit && ./install.sh ``` Symlinks `keybindings.json`, copies `pi-env.zsh` into `~/.oh-my-zsh/custom/`, and prints the `settings.json` bootstrap command. ### 4. Bootstrap pi settings ```bash cp ~/pi-toolkit/settings.example.json ~/.pi/agent/settings.json $EDITOR ~/.pi/agent/settings.json # eu./us./anthropic: prefix ``` ### 5. Install mempalace CLI + this toolkit ```bash uv tool install mempalace git clone ssh://git@gitea.jordbo.se:2222/joakimp/mempalace-toolkit.git ~/mempalace-toolkit cd ~/mempalace-toolkit && ./install.sh ``` Detects pi, symlinks `mempalace.ts` into `~/.pi/agent/extensions/`. Also detects pi-toolkit artifacts and prints a green check (or a warning telling you to install pi-toolkit first if you skipped step 3). ### 6. Register mempalace MCP with opencode (if applicable) Skip if this box is pi-only. Otherwise: - Install [`opencode-toolkit`](https://gitea.jordbo.se/joakimp/opencode-toolkit) so `~/.config/opencode/.env` is sourced into every shell (GitHub / Gitea / other MCP server tokens). - Register the mempalace MCP server in `~/.config/opencode/opencode.json` — see [root README § Registering mempalace with opencode](../../README.md#registering-mempalace-with-opencode-or-other-mcp-clients). ### 7. First run ```bash exec zsh pi # should start with defaults; wake-up injection shows palace status ``` If the wake-up doesn't print, run `MEMPALACE_EXT_DEBUG=1 pi` to surface `mempalace-mcp` stderr. ### Verification checklist ```bash # MCP bridge in place ls -la ~/.pi/agent/extensions/mempalace.ts # → this repo # pi-toolkit artifacts also in place ls -la ~/.pi/agent/keybindings.json # → pi-toolkit ls -la ~/.oh-my-zsh/custom/pi-env.zsh # cp from pi-toolkit # Env loaded zsh -ic 'echo $AWS_PROFILE $AWS_REGION' # Palace reachable mempalace status ``` ### Uninstall ```bash cd ~/mempalace-toolkit && ./install.sh --uninstall --yes # bridge only cd ~/pi-toolkit && ./install.sh --uninstall --yes # pi base config # Leaves pi itself, mempalace CLI, and ~/.config/pi/.env alone. ``` --- ## File layout ``` mempalace-toolkit/ └── extensions/ └── pi/ ├── README.md ← this file └── mempalace.ts ← symlinked into ~/.pi/agent/extensions/ ``` Pi base config (keybindings, env loader, settings template) lives in [`pi-toolkit`](https://gitea.jordbo.se/joakimp/pi-toolkit). `install.sh` detects pi via `~/.pi/agent/extensions/` and runs a `check_pi_toolkit` probe that warns if pi-toolkit's artifacts are missing.