From e68ee2071ca3ad396e164201181afcac65091992 Mon Sep 17 00:00:00 2001 From: Joakim Persson Date: Thu, 10 Sep 2026 20:58:12 +0200 Subject: [PATCH] docs(pi-ext): document what the mine deadline does, and what the message means Two corrections to the operator-facing docs, both exposed by 309980b. The env table listed the default as 30000, which is now wrong, and described the var as capping "the mempalace_mine call". It never did: it bounds how long the extension WAITS. The mine keeps running on the server. That exact misreading is what made a 30s deadline look safe on a call measured at 30-60s. Added a Debugging entry for "feed (tick) failed: mine timed out after ...ms", because every operator on this fleet has seen it and it was documented nowhere. It states the three things a reader needs: nothing was lost (the transcript is staged before the mine, and mine --mode convos dedups by source_file and is idempotent); do NOT retry harder from the client, because the palace is a single writer and a blind retry turns one slow mine into a queue; and after 309980b the message should not appear on a healthy fleet, so if it does it now MEANS something -- a mine exceeding five minutes, i.e. look at palace size or another writer holding the lock rather than raising the timeout again. --- extensions/pi/README.md | 26 +++++++++++++++++++++++++- 1 file changed, 25 insertions(+), 1 deletion(-) diff --git a/extensions/pi/README.md b/extensions/pi/README.md index a54c3cb..6a2eca7 100644 --- a/extensions/pi/README.md +++ b/extensions/pi/README.md @@ -117,7 +117,7 @@ side of the wiring. | `MEMPALACE_FEED_WING` | `wing_conversations` | Target wing — passed to both the exporter and the `mempalace_mine` call. | | `MEMPALACE_FEED_DEBOUNCE_MS` | `600000` (10 min) | Minimum gap between mid-session (`agent_settled`) feeds. Bounds crash loss to one window instead of a whole session. | | `MEMPALACE_FEED_PREPARE_TIMEOUT_MS` | `120000` | Kills a wedged `--prepare` subprocess. | -| `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `30000` | Caps the `mempalace_mine` call so a stalled palace can't hang session exit. | +| `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `300000` (5 min) | Bounds how long the extension *waits* for `mempalace_mine`, so a stalled palace can't hang session exit. It does **not** cancel the mine — see [Debugging](#debugging). Raised from `30000` in 2026-09: the mine is the slowest call this extension makes (30–60s in normal operation), so the old deadline fired routinely and reported healthy behaviour as an error. | **Remote palace:** if `$MEMPALACE_REMOTE_URL` is set (see [Transport](#transport-local-vs-external)), `mempalace_mine`'s source path is @@ -465,6 +465,30 @@ the next tool call transparently respawns `mempalace-mcp` and retries. `mempalace-mcp` manually with raw JSON-RPC on stdin to read the server-side error — much faster than guessing. +### `feed (tick) failed: mine timed out after …ms` + +**Nothing has been lost.** The deadline bounds only how long the extension +*waits*; it cannot cancel the mine, which continues on the server. The +transcript is already staged before the mine is invoked, and +`mempalace mine --mode convos` dedups by `source_file` and is idempotent, so the +work either completed after the deadline or is redone by the next tick. + +**Do not "fix" it by retrying harder from the client.** The palace is a single +writer; a blind retry is what turns one slow mine into a queue of them. + +Before 2026-09 this message appeared many times per session, which made it look +like a persistent failure. That was a real defect, now fixed: `lastFeedAt` was +recorded only after a *successful* wait, so a timeout left the debounce clock +stale and every following settled turn started another mine on top of the one +still running. Two changes — recording the attempt before the wait, and raising +the deadline to sit far above the slowest honest completion — mean a healthy +fleet should now never see it. + +If you *do* still see it, it is now informative rather than noise: a mine +exceeded five minutes. Check palace size and whether another writer (a +host-side feeder, a scheduled `mine`) is holding the write lock, rather than +raising the timeout again. + ## The `Type.Unsafe` gotcha Earlier versions of this extension registered every MCP tool with