docs(pi-ext): document what the mine deadline does, and what the message means

Two corrections to the operator-facing docs, both exposed by 309980b.

The env table listed the default as 30000, which is now wrong, and described the
var as capping "the mempalace_mine call". It never did: it bounds how long the
extension WAITS. The mine keeps running on the server. That exact misreading is
what made a 30s deadline look safe on a call measured at 30-60s.

Added a Debugging entry for "feed (tick) failed: mine timed out after ...ms",
because every operator on this fleet has seen it and it was documented nowhere.
It states the three things a reader needs: nothing was lost (the transcript is
staged before the mine, and mine --mode convos dedups by source_file and is
idempotent); do NOT retry harder from the client, because the palace is a single
writer and a blind retry turns one slow mine into a queue; and after 309980b the
message should not appear on a healthy fleet, so if it does it now MEANS
something -- a mine exceeding five minutes, i.e. look at palace size or another
writer holding the lock rather than raising the timeout again.
This commit is contained in:
Joakim Persson
2026-09-10 20:58:12 +02:00
parent 309980b62c
commit e68ee2071c
+25 -1
View File
@@ -117,7 +117,7 @@ side of the wiring.
| `MEMPALACE_FEED_WING` | `wing_conversations` | Target wing — passed to both the exporter and the `mempalace_mine` call. |
| `MEMPALACE_FEED_DEBOUNCE_MS` | `600000` (10 min) | Minimum gap between mid-session (`agent_settled`) feeds. Bounds crash loss to one window instead of a whole session. |
| `MEMPALACE_FEED_PREPARE_TIMEOUT_MS` | `120000` | Kills a wedged `--prepare` subprocess. |
| `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `30000` | Caps the `mempalace_mine` call so a stalled palace can't hang session exit. |
| `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `300000` (5 min) | Bounds how long the extension *waits* for `mempalace_mine`, so a stalled palace can't hang session exit. It does **not** cancel the mine — see [Debugging](#debugging). Raised from `30000` in 2026-09: the mine is the slowest call this extension makes (30–60s in normal operation), so the old deadline fired routinely and reported healthy behaviour as an error. |
**Remote palace:** if `$MEMPALACE_REMOTE_URL` is set (see
[Transport](#transport-local-vs-external)), `mempalace_mine`'s source path is
@@ -465,6 +465,30 @@ the next tool call transparently respawns `mempalace-mcp` and retries.
`mempalace-mcp` manually with raw JSON-RPC on stdin to read the
server-side error — much faster than guessing.
### `feed (tick) failed: mine timed out after …ms`
**Nothing has been lost.** The deadline bounds only how long the extension
*waits*; it cannot cancel the mine, which continues on the server. The
transcript is already staged before the mine is invoked, and
`mempalace mine --mode convos` dedups by `source_file` and is idempotent, so the
work either completed after the deadline or is redone by the next tick.
**Do not "fix" it by retrying harder from the client.** The palace is a single
writer; a blind retry is what turns one slow mine into a queue of them.
Before 2026-09 this message appeared many times per session, which made it look
like a persistent failure. That was a real defect, now fixed: `lastFeedAt` was
recorded only after a *successful* wait, so a timeout left the debounce clock
stale and every following settled turn started another mine on top of the one
still running. Two changes — recording the attempt before the wait, and raising
the deadline to sit far above the slowest honest completion — mean a healthy
fleet should now never see it.
If you *do* still see it, it is now informative rather than noise: a mine
exceeded five minutes. Check palace size and whether another writer (a
host-side feeder, a scheduled `mine`) is holding the write lock, rather than
raising the timeout again.
## The `Type.Unsafe` gotcha
Earlier versions of this extension registered every MCP tool with