Compare commits

..

6 Commits

Author SHA1 Message Date
Joakim Persson 975ab92943 feat(pi): let an ask declare dormancy, so waiting work stops nagging
A first-boot acceptance ask is planted deliberately unanswerable: it describes
work that becomes possible only when the device is next recreated, and it must
STAY owed until then, because closing it early to tidy the mailbox is exactly how
that work gets lost. deriveOwed could only see "directed, open, not answered, not
withdrawn", so such an ask was announced at every session start and every poll
for as long as it was correctly waiting -- measured at three days running on
emb-7kj4vr4g (2026-09-28 -> 2026-10-01), and across three consecutive releases
before that. The ask was right; announcing it was wrong, and the cost landed on
the human reading the window, who is the one reader that cannot filter it.

An ask may now declare the condition under which it is merely waiting:

  "dormant_unless": [
    { "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json",
      "field": "release_tag", "baseline": "v1.9.4" },
    { "kind": "file_mtime", "path": "/etc/hostname",
      "baseline": "2026-09-22T18:12:49Z" }
  ]

Dormant while EVERY condition still matches its baseline; live the moment ANY
differs -- which is the trigger those asks already stated in prose ("act when
EITHER differs"), now in a form the bridge can check. Two kinds, local files
only, no expression language, no shell, no network: a general evaluator in the
path that decides whether work is VISIBLE is a far worse trade than a clumsy
schema.

DORMANCY IS PROVEN, NEVER ASSUMED. Every unevaluable predicate announces the ask
instead of hiding it: missing file, unreadable file, unparseable baseline,
unknown kind, vanished field, relative path, more than eight conditions. The
dangerous failure here is not a spurious nag but work that disappears because a
predicate could not be evaluated -- indistinguishable from the ask being lost,
and not surfacing until a release needed it. An ask with no dormant_unless
behaves exactly as before, so this is backward compatible by construction.

Withheld from the ANNOUNCEMENT, never from the mailbox: deriveOwed now returns a
partition {owed, dormant} rather than a flat list, the wake-up injection lists
dormant asks once per session with ids, and a mid-session poll adds only a count
and only when the window is already open for something else. Dormant asks are
deliberately NOT added to the `surfaced` map, so one becomes announceable the
instant its baseline moves.

file_mtime compares WHOLE SECONDS in UTC. A filesystem mtime carries sub-second
residue (measured: /etc/hostname at .773761009) that a reported ISO baseline
never will, so comparing raw milliseconds would mark every such predicate
permanently "changed" -- silently disabling the feature while appearing to work.
The test records the residue for that reason.

scripts/test-dormancy.sh is this repo's first test: 22 cases, positive and
negative arms both, because a predicate that never fires makes the feature inert
and one that fires too eagerly hides real work. It copies the extension into a
temp tree with pi's typebox symlinked beside it (the copy is made per-run, so it
cannot drift like a vendored duplicate), owns its own fixtures rather than
reading /etc paths -- the first draft passed only on a pi-devbox container and
would have silently flipped to "not dormant" anywhere else -- and exits 2 for
INCONCLUSIVE rather than 0, since a test that skips quietly is the failure mode
it exists to catch. shellcheck clean.
2026-10-01 23:59:36 +02:00
joakimp 2167a1b033 fix(mailbox): pass order: "desc" on every cursor-less event_list call
Three of the five `mempalace_event_list` calls in the pi extension relied on
the server's default ordering: deriveOwed's `status: "open"` candidate query
(limit 50) and both windows in deriveClosed (from_agent limit 100, to_agent
limit 50). deriveOwed's other two calls already say `order: "desc"` and carry
the comment explaining why: on mempalace <= 3.9.0 the default is `asc`, so a
cursor-less `limit: N` returns the OLDEST N events, and once a device passes N
authored events its newest asks and the replies that close them fall outside
the join window. deriveClosed had exactly that latent truncation.

Why now: mempalace 3.10.0 (2026-09-15) flips the cursor-less default to
newest-first ("Logstream listings default to the newest events when no cursor
is given"). That change is server-side -- the extension talks to the fleet
hub over MEMPALACE_REMOTE_URL, so it lands when the hub upgrades, not when a
client does -- and it would have silently FIXED deriveClosed on 3.10.0 while
leaving it broken against any 3.9.0 hub. A mailbox verdict that depends on
which server version answers is the wrong shape. Saying `order` explicitly
makes both derive* functions read the same window on either version.

The selection in deriveClosed (isStrictlyAfter per correlation) is
order-independent, so this changes WHICH events are in the window, not how
the winner is picked. No behaviour change on a device with < 50 inbound and
< 100 authored events; the fleet hub is at 176 events total today, so the
window was about to matter.

Verified: file compiles under pi's own loader (jiti 2.7.0 from the installed
pi-coding-agent) before and after; `order:"desc"` literal count 3 -> 7, which
is the 3 added arguments plus 1 comment mention. `node --check` and a bare
`tsc` were tried first and both fail identically on HEAD (inline `type`
imports / no @types/node) -- checker faults, not this change.
2026-09-22 14:15:22 +02:00
joakimp 817b3a82b7 fix(feed): the mine's deadline never reached the transport; 60 s cut it off
The 2026-09 change raised MEMPALACE_FEED_MINE_TIMEOUT_MS to 300 000 and
raced it against client.callTool("mempalace_mine"). But callTool() had no
way to carry a deadline, so every call went out under the transport's
generic per-request timeout (MEMPALACE_MCP_TIMEOUT_MS, 60 000), which fired
first on every honest 60 s+ mine. Operators saw

    feed (tick) failed: mempalace remote request 'tools/call' failed:
    timed out after 60000ms

instead of the message the change had aimed at, and the 300 s was
unreachable. On stdio it was worse than noise: that transport kills the
server child on timeout, so the mine was actually aborted at 60 s.

callTool(name, args, { timeoutMs }) now passes a per-call deadline to both
transports; feedPalace() uses it for the mine. Plain calls keep the short
default — a query taking 60 s is still wedged. The Promise.race stays as the
liveness guard for a transport with its timeout disabled (0).

scripts/test-mcp-call-timeout.sh cuts RemoteMcpClient out of the shipped
file (as test-owed-withdrawal.sh does for the mailbox predicates), drives it
against a local JSON-RPC server that delays tools/call, and asserts: plain
call rejects at the generic deadline; the override outlives it; the override
is itself a deadline. Fails on the previous commit (2 of 6), passes here.
Vendored-copy delta noted in the header; protocol untouched, sync token
unchanged (check-mcp-client-sync.sh passes against pi-extensions).
2026-09-18 17:04:59 +02:00
joakimp dab989b068 fix(mailbox-tests): the owed-set suite's own gate refused to let it run on node 24
scripts/test-owed-withdrawal.sh has not executed a single assertion since the
image moved to node 24. Its precondition line

    node --experimental-strip-types --check "$SRC"

exits 2 on an unmodified extensions/pi/mempalace.ts, and the script is
`set -euo pipefail` with `|| exit 2`, so all 17 assertions and both regression
guards were skipped. Measured on pi-devbox v1.9.1, node v24.21.0.

CAUSE, MEASURED AND NARROWER THAN IT LOOKS. The reported error names the inline
type-import at mempalace.ts:92, which invites the reading "that import is
unusual". It is not the import: `node --check` does not type-strip AT ALL. A file
whose entire content is `const x: number = 1;` fails identically, with or without
--experimental-strip-types, while `node --experimental-strip-types <same file>`
executes it fine. --check has never been type-aware; execution is what gained
stripping. So this gate could never validate TypeScript, on any node.

WHY IT USED TO PASS -- INFERENCE, NOT MEASUREMENT. e2b060a's message records this
suite running on 2026-09-09 with a control pass and four mutation kills, when the
image shipped node v22.23.2; v1.9.1 ships v24.21.0. No node 22 exists on the box
where this was diagnosed, so the counterfactual was not executed. Treat "22
stripped for --check and 24 stopped" as a hypothesis consistent with the record,
not as a measured cause. What IS measured is the present-tense behaviour above.

WHY THIS MATTERS MORE THAN A RED TEST. This suite exists because isWithdrawn is
the one rule in the extension that can go wrong SILENTLY -- a wrong rule does not
throw and does not log, it makes a real unanswered ask vanish from a mailbox
forever. The rule shipped in e2b060a and has been baked since v1.9.1; the thing
that makes its failure mode visible has been dark for the same period. The
mitigating half: it failed CLOSED (exit 2, loud), never vacuously green. A gate
that cannot run must not pass, and it did not.

THE FIX. Strip first, then syntax-check the emitted JS: version-stable, and it
still refuses malformed input. `mode: "strip"` blanks type syntax without moving
anything, so offsets and line numbers survive and a reported error line still
points at the right line of the original .ts (69157 B in, 69157 B out).

THE EXIT CODES ARE NOW SPLIT, AND THAT IS THE POINT. 3 = the gate itself cannot
run (no module.stripTypeScriptTypes, i.e. node < 22.13). 2 = the source does not
parse. Collapsing the two is how this defect disguised itself: it printed
"mempalace.ts does not parse" while mempalace.ts was fine, sending a reader to
inspect the wrong file. Note the stripper is itself a parser, so a genuine syntax
error surfaces as an exception from the strip call rather than from --check; that
path is caught and reported as 2, not 3. Both remain failures. Neither passes.

VERIFIED IN SIX DIRECTIONS on this image, each expectation written down first:
  unmutated source          -> rc=0, PASSED (17/17)
  malformed TypeScript      -> rc=2, "does not parse: Expression expected"
  stripper made unavailable -> rc=3, "cannot strip", and NOT "does not parse"
                               (simulated by doctoring the runtime through
                               NODE_OPTIONS, so the shipped line ran as shipped)
  M1 marker requirement removed -> FAILED: 3   (no-marker, other-thread, prose)
  M2 third-party guard removed  -> FAILED: 1
  M3 to_agent guard removed     -> FAILED: 2   (broadcast and wrong-device share it)
  M4 ordering guard removed     -> FAILED: 1
The four kill counts are the same ones e2b060a recorded, so sensitivity is
restored rather than merely asserted.

WORTH KNOWING FOR THE NEXT PERSON WHO MUTATES THIS: the extractor carries its own
guard that rejects an isWithdrawn which no longer mentions "withdraws", so the
obvious M1 (delete the marker check outright) is refused before any assertion
runs -- correctly, but it looks like a crash. Mutate to `... || true` instead,
which removes the requirement while keeping the key mentioned.

SC2016 is disabled on the node invocation with a stated reason: the single quotes
are deliberate, the payload is JavaScript and `${process.version}` must reach
node rather than the shell. Checked that the file is shellcheck-clean at default
severity, as it was before this change, and at the -S error severity the
pi-devbox gate uses.

NOT FIXED HERE. Nothing in CI runs this suite -- it is a script an operator
invokes, which is precisely why a gate that fails loudly still went unnoticed
across a node bump. The harness's own runner two hundred lines below still passes
--experimental-strip-types, deliberately: it is a no-op on 24 and required on 22,
so it keeps the suite runnable on both. And the real reason this was found at all
is unrelated to CI: isWithdrawn was being exercised end-to-end on a released
image against the live logstream for the first time (project/pi-devbox seq 140),
and the suite was reached for as corroboration.
2026-09-14 16:27:29 +02:00
Joakim Persson e68ee2071c docs(pi-ext): document what the mine deadline does, and what the message means
Two corrections to the operator-facing docs, both exposed by 309980b.

The env table listed the default as 30000, which is now wrong, and described the
var as capping "the mempalace_mine call". It never did: it bounds how long the
extension WAITS. The mine keeps running on the server. That exact misreading is
what made a 30s deadline look safe on a call measured at 30-60s.

Added a Debugging entry for "feed (tick) failed: mine timed out after ...ms",
because every operator on this fleet has seen it and it was documented nowhere.
It states the three things a reader needs: nothing was lost (the transcript is
staged before the mine, and mine --mode convos dedups by source_file and is
idempotent); do NOT retry harder from the client, because the palace is a single
writer and a blind retry turns one slow mine into a queue; and after 309980b the
message should not appear on a healthy fleet, so if it does it now MEANS
something -- a mine exceeding five minutes, i.e. look at palace size or another
writer holding the lock rather than raising the timeout again.
2026-09-10 20:58:12 +02:00
Joakim Persson 309980b62c fix(pi-ext): stop the feed tick from launching overlapping mines
"[mempalace ext] feed (tick) failed: mine timed out after 30000ms" was parked as
cosmetic on 2026-08-27. It is not cosmetic: the tight deadline was the trigger,
but the defect is a positive feedback loop that puts multiple writers on a
single-writer palace.

lastFeedAt was assigned only AFTER a successful await. The Promise.race abandons
our WAIT and cannot cancel the server's work, so on a mine that takes longer
than the deadline -- measured at 30-60s in normal operation, against a 30s
deadline -- the catch ran with lastFeedAt UNCHANGED. That left both guards in
the agent_settled handler open at once: the debounce test
(`Date.now() - lastFeedAt < feedDebounceMs`) passed because lastFeedAt was still
stale, and feedInFlight was already null because `run` had settled. Every
subsequent settled turn therefore launched another mine on top of the one still
running, each making the next slower and the next timeout likelier -- which is
why operators saw the message many times per session instead of at most once per
10-minute debounce window.

Fix, three lines:
- move `lastFeedAt = Date.now()` to before the await, so a timeout still starts
  the debounce clock. A timeout is not a "did not happen": the mine is running
  server-side and `mine --mode convos` dedups by source_file and is idempotent.
- raise MEMPALACE_FEED_MINE_TIMEOUT_MS from 30_000 to 300_000. The mine is the
  slowest thing this extension does yet carried the tightest deadline: 4x
  tighter than the prepare step before it (120_000) and 10x tighter than the
  init handshake (300_000), a fast call. All three were introduced together in
  29e660e and this one was never revisited. 300_000 matches the init timeout
  because liveness is the only job left for this deadline -- it cannot cancel
  the server's work, so it must sit far above the slowest honest completion.

Simulated both guards over 10 minutes of settled turns at 20s intervals with a
60s mine: BEFORE 16 mines launched, 15 of them overlapping an already-running
mine; AFTER 2 launched, 0 overlapping. With a mine that exceeds even the new
deadline (400s): BEFORE 16/15, AFTER still 2/0 -- the lastFeedAt move is what
actually fixes it, and it holds even when the timeout still fires. The raise
stops the spurious message; the move stops the pile-up.

NOT fixed here: the message still goes to process.stderr, which pi renders into
the TUI input field. That needs a pi-side channel or a log file, and is tracked
separately. After this change the message should be rare, and when it does
appear it means something real: a mine exceeding five minutes.
2026-09-10 20:48:33 +02:00
5 changed files with 730 additions and 33 deletions
+99 -1
View File
@@ -117,7 +117,7 @@ side of the wiring.
| `MEMPALACE_FEED_WING` | `wing_conversations` | Target wing — passed to both the exporter and the `mempalace_mine` call. |
| `MEMPALACE_FEED_DEBOUNCE_MS` | `600000` (10 min) | Minimum gap between mid-session (`agent_settled`) feeds. Bounds crash loss to one window instead of a whole session. |
| `MEMPALACE_FEED_PREPARE_TIMEOUT_MS` | `120000` | Kills a wedged `--prepare` subprocess. |
| `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `30000` | Caps the `mempalace_mine` call so a stalled palace can't hang session exit. |
| `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `300000` (5 min) | Bounds how long the extension *waits* for `mempalace_mine`, so a stalled palace can't hang session exit. It does **not** cancel the mine — see [Debugging](#debugging). Raised from `30000` in 2026-09: the mine is the slowest call this extension makes (30–60s in normal operation), so the old deadline fired routinely and reported healthy behaviour as an error. |
**Remote palace:** if `$MEMPALACE_REMOTE_URL` is set (see
[Transport](#transport-local-vs-external)), `mempalace_mine`'s source path is
@@ -378,6 +378,57 @@ the one failure this derivation exists to prevent.
protocol says a broadcast owes nobody a reply — which also means broadcasting an
ask demonstrably reaches no owed set, giving that documented anti-pattern teeth.
**Dormant asks — open, owed, and deliberately not announced.** Some asks are
planted *unanswerable on purpose*: a first-boot acceptance describes work that
becomes possible only when the device is next recreated, and it must stay owed
until then, because closing it early to tidy the mailbox is exactly how the work
gets lost. Derivation alone cannot tell that apart from neglected work, so such
an ask was announced at every session start and every poll for as long as it was
correctly waiting — measured at three days running on `emb-7kj4vr4g`
(2026-09-28 → 2026-10-01), and across three consecutive releases before that.
The ask was right; announcing it was wrong, and the cost landed on the human
reading the window, the one reader who cannot filter it.
An ask may therefore declare, in its own `metadata`, the condition under which it
is merely waiting:
```json
"dormant_unless": [
{ "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json",
"field": "release_tag", "baseline": "v1.9.4" },
{ "kind": "file_mtime", "path": "/etc/hostname",
"baseline": "2026-09-22T18:12:49Z" }
]
```
It is dormant while **every** condition still matches its baseline, and goes live
the moment **any** of them differs — the trigger such asks already stated in
prose ("act when EITHER differs"), now in a form the bridge can check. Two kinds,
both local-file-only: `file_mtime` (compared at whole seconds in UTC, because a
filesystem mtime carries sub-second residue that a reported ISO baseline never
will — comparing raw milliseconds would mark every such predicate permanently
"changed" and silently disable the feature) and `json_field` (dotted paths
allowed, compared as strings so a manifest holding `3` matches a baseline of
`"3"`). There is no expression language, no shell and no network: a general
evaluator in the path that decides whether work is *visible* is a far worse trade
than a clumsy schema.
Dormancy is **withheld from the announcement, never from the mailbox**: the
wake-up injection lists dormant asks once per session with their ids, and a
mid-session poll mentions only a count, and only when the window is already open
for something else. They are not added to the resurface map, so one becomes
announceable the instant its baseline moves.
**Dormancy must be proven, never assumed** — every unevaluable predicate shows
the ask as owed. A missing file, an unreadable one, a baseline that will not
parse, an unknown `kind`, a vanished field, a relative path, more than eight
conditions: each announces. The dangerous failure here is not a spurious nag but
work that disappears because a predicate could not be evaluated, which would be
indistinguishable from the ask being lost and would not surface until a release
needed it. An ask with no `dormant_unless` behaves exactly as it did before the
feature existed. `scripts/test-dormancy.sh` exercises all of it, positive and
negative arms both.
Gated on `MEMPALACE_PI_DEVICE` **and** `MEMPALACE_REMOTE_URL` (the same pair as
the stamper, since an unstamped client has no address to be reached at), and
disabled outright with `MEMPALACE_MAILBOX=0` (notifications alone with
@@ -423,6 +474,9 @@ one — the long-lived server is only killed when a request genuinely stalls.
- `MEMPALACE_MCP_TIMEOUT_MS` — tool-call/request timeout. Default `60000`.
Kept short on purpose: a *query* taking this long is genuinely wedged.
The one call that is not a query — the feed's `mempalace_mine` — passes its
own deadline (`MEMPALACE_FEED_MINE_TIMEOUT_MS`) down to the transport per
call, so this default does not apply to it (since 2026-09-18; see below).
- `MEMPALACE_MCP_INIT_TIMEOUT_MS` — `initialize` + `tools/list` handshake
timeout. Default `300000`. Deliberately generous: a genuine first
cold-open over virtiofs can legitimately take minutes, and killing a
@@ -465,6 +519,50 @@ the next tool call transparently respawns `mempalace-mcp` and retries.
`mempalace-mcp` manually with raw JSON-RPC on stdin to read the
server-side error — much faster than guessing.
### `feed (tick) failed: mine timed out after …ms`
**Nothing has been lost.** The deadline bounds only how long the extension
*waits*; it cannot cancel the mine, which continues on the server. The
transcript is already staged before the mine is invoked, and
`mempalace mine --mode convos` dedups by `source_file` and is idempotent, so the
work either completed after the deadline or is redone by the next tick.
**Do not "fix" it by retrying harder from the client.** The palace is a single
writer; a blind retry is what turns one slow mine into a queue of them.
Before 2026-09 this message appeared many times per session, which made it look
like a persistent failure. That was a real defect, now fixed: `lastFeedAt` was
recorded only after a *successful* wait, so a timeout left the debounce clock
stale and every following settled turn started another mine on top of the one
still running. Two changes — recording the attempt before the wait, and raising
the deadline to sit far above the slowest honest completion — mean a healthy
fleet should now never see it.
If you *do* still see it, it is now informative rather than noise: a mine
exceeded five minutes. Check palace size and whether another writer (a
host-side feeder, a scheduled `mine`) is holding the write lock, rather than
raising the timeout again.
### `feed (tick) failed: mempalace remote request 'tools/call' failed: timed out after 60000ms`
Same event, different deadline — and the same reassurance: **nothing has been
lost**, the mine continues server-side.
This is what the previous message turned into after the 2026-09 change, and
it exposed that the change was incomplete. The feed's five-minute deadline was
only *raced* against the call; `callTool()` had no way to carry it, so the
transport's generic per-request timeout (`MEMPALACE_MCP_TIMEOUT_MS`, 60 s)
fired first on every honest 60 s+ mine, and the 300 s was unreachable. On the
stdio transport this was worse than noise: a per-request timeout there kills the
server child, so the mine really was aborted at 60 s.
Fixed 2026-09-18: `callTool(name, args, { timeoutMs })` passes a per-call
deadline to both transports and the feed uses it for the mine.
`scripts/test-mcp-call-timeout.sh` pins the contract (a plain call still honours
the short default; the override is honoured and is itself a deadline). Seeing
this message on a fixed build means a mine exceeded *five* minutes — treat it as
the previous section says.
## The `Type.Unsafe` gotcha
Earlier versions of this extension registered every MCP tool with
+297 -29
View File
@@ -67,7 +67,9 @@
* child, so pi gets an error instead of hanging and later calls fail fast.
* This is a per-REQUEST timeout, not a process-lifetime one — the
* long-lived server is only killed when a request genuinely stalls.
* - MEMPALACE_MCP_TIMEOUT_MS tool-call/request timeout (default 60000)
* - MEMPALACE_MCP_TIMEOUT_MS tool-call/request timeout (default 60000);
* the feed's mine carries its own, longer
* deadline (MEMPALACE_FEED_MINE_TIMEOUT_MS)
* - MEMPALACE_MCP_INIT_TIMEOUT_MS initialize+tools/list timeout (default 300000)
* Set either to 0 to disable (legacy unbounded behavior).
*
@@ -90,6 +92,7 @@
*/
import { type ChildProcessWithoutNullStreams, spawn } from "node:child_process";
import { readFileSync, statSync } from "node:fs";
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { Type } from "typebox";
@@ -116,7 +119,13 @@ interface IMcpClient {
readonly alive: boolean;
onExit: (() => void) | null;
start(): Promise<void>;
callTool(name: string, args: Record<string, unknown>): Promise<any>;
/**
* `opts.timeoutMs` overrides the transport's generic per-request deadline
* for THIS call only. Callers that knowingly invoke a long server-side job
* (the feed's `mempalace_mine`) pass their own deadline here; everything
* else keeps the short default, which is the wedged-query guard.
*/
callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any>;
ensureAlive(): Promise<boolean>;
stop(): void | Promise<void>;
}
@@ -149,6 +158,165 @@ const sleep = (ms: number): Promise<void> =>
if (typeof t.unref === "function") t.unref();
});
// ── Dormancy ──────────────────────────────────────────────────────────────────
/**
* Is a standing ask DORMANT — correct to leave open, wrong to announce?
*
* The problem this solves, measured rather than imagined. A first-boot
* acceptance ask is planted deliberately unanswerable: it describes work that
* becomes possible only when the device is next recreated, and it must STAY
* owed until then, because closing it early to tidy the mailbox is exactly how
* that work gets lost. But deriveOwed could only see "directed, open, not
* answered, not withdrawn", so such an ask is announced at every session start
* and every poll for as long as it is correctly waiting — three days running on
* emb-7kj4vr4g (2026-09-28 → 2026-10-01), and three consecutive releases before
* that. The ask was right; announcing it was wrong. The cost lands on the human
* reading the window, who is the one reader that cannot filter it.
*
* So an ask may now declare, in its own metadata, the condition under which it
* is merely waiting:
*
* "dormant_unless": [
* { "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json",
* "field": "release_tag", "baseline": "v1.9.4" },
* { "kind": "file_mtime", "path": "/etc/hostname",
* "baseline": "2026-09-22T18:12:49Z" }
* ]
*
* It is dormant while EVERY condition still matches its baseline, and goes live
* the moment ANY of them differs — which is the trigger those asks already
* state in prose ("act when EITHER differs"), now in a form a machine can
* check. Dormant asks are withheld from the ANNOUNCED owed set, never from the
* mailbox: the wake-up injection still lists them once per session, so they
* stay discoverable and cannot be silently dropped.
*
* THE SAFETY RULE, and the reason every branch below reports `unchanged: false`
* on doubt: dormancy must be PROVEN, never assumed. A missing file, an
* unreadable one, a baseline that will not parse, an unknown `kind`, a field
* that has vanished — each is a reason to SHOW the ask, not to hide it. The
* dangerous failure here is not a spurious nag; it is work that vanishes
* because a predicate could not be evaluated, which would be indistinguishable
* from the ask being lost and would not surface until a release needed it.
*
* No eval, no shell, no network: the predicate language is deliberately two
* fixed checks against local files. A general expression evaluator sitting in
* the path that decides whether work is VISIBLE is a far worse trade than a
* slightly clumsy schema.
*/
export type DormancyVerdict = { dormant: boolean; reason: string };
/** Split of the derived owed-set: what to announce, and what is merely waiting. */
export type OwedSplit = { owed: LogEvent[]; dormant: LogEvent[] };
// Bounds, so a malformed or hostile event cannot turn a mailbox read into real
// work. Both are deliberately small: these predicates describe boot state, and
// anything needing more than a handful of local checks is not a dormancy rule.
const MAX_DORMANCY_CONDITIONS = 8;
const MAX_DORMANCY_JSON_BYTES = 256 * 1024;
const dormancyCondition = (c: unknown): { unchanged: boolean; reason: string } => {
if (typeof c !== "object" || c === null || Array.isArray(c))
return { unchanged: false, reason: "condition is not an object" };
const { kind, path, baseline, field } = c as Record<string, unknown>;
// Absolute paths only: a relative one would resolve against whatever cwd the
// pi process happens to hold, which is not a property of the device whose
// state the baseline describes.
if (typeof path !== "string" || !path.startsWith("/") || path.includes("\0"))
return {
unchanged: false,
reason: `path must be an absolute string, got ${JSON.stringify(path)}`,
};
if (typeof baseline !== "string" && typeof baseline !== "number")
return { unchanged: false, reason: "baseline must be a string or a number" };
if (kind === "file_mtime") {
// Compared at WHOLE SECONDS in UTC, because the two routes that produce
// these baselines disagree below that: a filesystem mtime carries
// sub-second residue (measured: /etc/hostname at .773761009) that a
// hand-written or reported ISO baseline never will. Comparing raw
// milliseconds would make every such predicate permanently "changed",
// i.e. would silently disable the feature while appearing to work.
const expected = Date.parse(String(baseline));
if (!Number.isFinite(expected))
return {
unchanged: false,
reason: `baseline is not a parseable date: ${JSON.stringify(baseline)}`,
};
let mtimeMs: number;
try {
mtimeMs = statSync(path).mtimeMs;
} catch {
return { unchanged: false, reason: `cannot stat ${path}` };
}
return Math.floor(mtimeMs / 1000) === Math.floor(expected / 1000)
? { unchanged: true, reason: `${path} mtime still at baseline` }
: { unchanged: false, reason: `${path} mtime moved from baseline` };
}
if (kind === "json_field") {
if (typeof field !== "string" || field.length === 0)
return { unchanged: false, reason: "json_field needs a non-empty field name" };
let text: string;
try {
const buf = readFileSync(path);
if (buf.byteLength > MAX_DORMANCY_JSON_BYTES)
return {
unchanged: false,
reason: `${path} exceeds the ${MAX_DORMANCY_JSON_BYTES}-byte cap`,
};
text = buf.toString("utf8");
} catch {
return { unchanged: false, reason: `cannot read ${path}` };
}
let doc: unknown;
try {
doc = JSON.parse(text);
} catch {
return { unchanged: false, reason: `${path} is not valid JSON` };
}
let cur: unknown = doc;
for (const seg of field.split(".")) {
if (
typeof cur !== "object" ||
cur === null ||
!Object.prototype.hasOwnProperty.call(cur, seg)
)
return { unchanged: false, reason: `${path} has no field ${field}` };
cur = (cur as Record<string, unknown>)[seg];
}
if (cur !== null && typeof cur === "object")
return { unchanged: false, reason: `field ${field} is not a scalar` };
// Compared as strings on purpose: the baseline arrives as JSON metadata, so
// a manifest holding 3 and a baseline saying "3" are the same fact.
return String(cur) === String(baseline)
? { unchanged: true, reason: `${field} still at baseline` }
: { unchanged: false, reason: `${field} moved from baseline` };
}
return { unchanged: false, reason: `unknown dormancy kind ${JSON.stringify(kind)}` };
};
export function evaluateDormancy(
metadata: Record<string, unknown> | null | undefined,
): DormancyVerdict {
const raw = metadata?.dormant_unless;
// Every one of these returns NOT dormant, i.e. "announce it". An ask without
// a predicate behaves exactly as it did before this feature existed.
if (raw === undefined || raw === null) return { dormant: false, reason: "no dormant_unless" };
if (!Array.isArray(raw)) return { dormant: false, reason: "dormant_unless is not an array" };
if (raw.length === 0) return { dormant: false, reason: "dormant_unless is empty" };
if (raw.length > MAX_DORMANCY_CONDITIONS)
return {
dormant: false,
reason: `too many conditions (${raw.length} > ${MAX_DORMANCY_CONDITIONS})`,
};
for (const c of raw) {
const v = dormancyCondition(c);
if (!v.unchanged) return { dormant: false, reason: v.reason };
}
return { dormant: true, reason: `all ${raw.length} condition(s) still at baseline` };
}
class StdioMcpClient implements IMcpClient {
private proc: ChildProcessWithoutNullStreams | null = null;
private nextId = 1;
@@ -372,8 +540,12 @@ class StdioMcpClient implements IMcpClient {
});
}
async callTool(name: string, args: Record<string, unknown>): Promise<any> {
return this.request("tools/call", { name, arguments: args });
async callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any> {
// The per-call override matters MORE here than for the HTTP client: on
// timeout this transport kills the server child, so a generic deadline
// that undercuts a long mine does not merely abandon the wait — it
// aborts the mine.
return this.request("tools/call", { name, arguments: args }, opts?.timeoutMs ?? this.requestTimeoutMs);
}
/** SIGTERM then SIGKILL grace, for stall recovery. */
@@ -421,6 +593,10 @@ class StdioMcpClient implements IMcpClient {
// • per-request AbortController timeout honouring MEMPALACE_MCP_TIMEOUT_MS /
// MEMPALACE_MCP_INIT_TIMEOUT_MS, mirroring StdioMcpClient's timeout ethos.
// • alive / ensureAlive / onExit to satisfy IMcpClient.
// • callTool() takes an optional per-call `{ timeoutMs }` (IMcpClient
// contract, see there) so the feed's long-running mine is not cut off by
// the generic per-request deadline. Not a protocol change; sync token
// unchanged.
//
// NOTE: mempalace-mcp --transport http is a SESSIONLESS, stateless JSON-RPC
// server (no Mcp-Session-Id, always application/json, Connection: close), so
@@ -466,8 +642,8 @@ class RemoteMcpClient implements IMcpClient {
this.healthy = true;
}
async callTool(name: string, args: Record<string, unknown>): Promise<any> {
return this.request("tools/call", { name, arguments: args });
async callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any> {
return this.request("tools/call", { name, arguments: args }, { timeoutMs: opts?.timeoutMs });
}
/**
@@ -855,7 +1031,22 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
const feedWing = process.env.MEMPALACE_FEED_WING ?? "wing_conversations";
const feedDebounceMs = num(process.env.MEMPALACE_FEED_DEBOUNCE_MS, 600_000);
const feedPrepareTimeoutMs = num(process.env.MEMPALACE_FEED_PREPARE_TIMEOUT_MS, 120_000);
const feedMineTimeoutMs = num(process.env.MEMPALACE_FEED_MINE_TIMEOUT_MS, 30_000);
// 300_000, raised from 30_000 on 2026-09-10. The mine is the SLOWEST thing
// this extension does — it offers every qualifying session transcript to a
// SINGLE-WRITER palace, measured at 30–60s in normal operation and growing
// with the corpus — yet it carried by far the TIGHTEST deadline: 4x tighter
// than the prepare step that precedes it (120_000) and 10x tighter than the
// init handshake (300_000), which is a fast call. All three were introduced
// together in 29e660e (2026-08-12) and this one was never revisited, so the
// deadline fired during entirely normal operation and the resulting message
// read as an error when nothing had gone wrong.
//
// Matched to the init timeout because liveness is the ONLY legitimate job
// left for this deadline: the Promise.race below abandons our WAIT, it cannot
// cancel the server's work, so the deadline buys nothing except an escape
// from a permanently hung call. It must therefore sit far above the slowest
// honest completion, not near it.
const feedMineTimeoutMs = num(process.env.MEMPALACE_FEED_MINE_TIMEOUT_MS, 300_000);
let lastFeedAt = 0; // 0 => the first settled turn also acts as a catch-up
let feedInFlight: Promise<void> | null = null;
@@ -909,17 +1100,46 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
try {
const source = await prepareFeed(reason);
if (!source) return;
// Record the attempt HERE, before awaiting — not after a successful
// wait. The race below abandons only our WAIT; the mine keeps running
// server-side, and `mine --mode convos` dedups by source_file and is
// idempotent, so a timeout is emphatically not a "did not happen".
//
// Leaving lastFeedAt stale on the timeout path defeated the debounce
// guard in the agent_settled handler below (`Date.now() - lastFeedAt <
// feedDebounceMs`): with lastFeedAt unchanged that guard passed on EVERY
// settled turn, and because `run` had already settled, feedInFlight was
// null too — so BOTH guards stood open. Each settled turn then launched
// another mine while the previous one was still running: overlapping
// writers queueing on a single-writer palace, each making the next one
// slower and the next timeout likelier. That positive feedback loop,
// not the tight deadline by itself, is why operators saw "mine timed out
// after 30000ms" many times per session rather than at most once per
// debounce window.
lastFeedAt = Date.now();
// The deadline is passed DOWN to the transport as well as raced
// here. Until 2026-09-18 it was only raced: callTool() had no way
// to carry it, so the transport's generic per-request timeout
// (MEMPALACE_MCP_TIMEOUT_MS, 60 000) fired first on every honest
// 60 s+ mine — "remote request 'tools/call' failed: timed out
// after 60000ms" — and the 300 000 below was unreachable. The
// race stays as the liveness guard for a transport whose timeout
// is disabled (0).
await Promise.race([
client.callTool("mempalace_mine", {
source,
mode: "convos",
wing: feedWing,
// Internal call: it does not pass through the registered tool's
// execute(), so it stamps itself. These ARE this harness's own
// transcripts from this device, so the harness segment is the
// agent (not `miner`) even though the tool is `mine`.
agent: stampProvenance ? `${agentName}@${device}` : agentName,
}),
client.callTool(
"mempalace_mine",
{
source,
mode: "convos",
wing: feedWing,
// Internal call: it does not pass through the registered tool's
// execute(), so it stamps itself. These ARE this harness's own
// transcripts from this device, so the harness segment is the
// agent (not `miner`) even though the tool is `mine`.
agent: stampProvenance ? `${agentName}@${device}` : agentName,
},
{ timeoutMs: feedMineTimeoutMs },
),
new Promise((_resolve, reject) =>
setTimeout(
() => reject(new Error(`mine timed out after ${feedMineTimeoutMs}ms`)),
@@ -927,7 +1147,6 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
),
),
]);
lastFeedAt = Date.now();
} catch (err) {
process.stderr.write(
`[mempalace ext] feed (${reason}) failed: ${(err as Error).message}\n`,
@@ -1146,14 +1365,20 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
* Directed asks addressed to this device with no terminal reply from it, and
* not explicitly withdrawn by whoever sent them (see isWithdrawn).
*/
async function deriveOwed(): Promise<LogEvent[]> {
if (!mailboxEnabled || !available) return [];
async function deriveOwed(): Promise<OwedSplit> {
if (!mailboxEnabled || !available) return { owed: [], dormant: [] };
try {
const [candidatesRaw, mineRaw, inboundRaw] = await Promise.all([
// Every cursor-less event_list call in this file passes `order` explicitly.
// The server's default is not ours to lean on: mempalace <= 3.9.0 defaults
// to `asc`, 3.10.0 flips a cursor-less listing to newest-first. Both
// derive* joins want the NEWEST window (see the comment below), so say so
// and the verdict stops depending on which server version answers.
client.callTool("mempalace_event_list", {
to_agent: mailboxAddress,
status: "open",
limit: 50,
order: "desc",
}),
// `order: "desc"` is a fix, not a flourish. event_list defaults to `asc`
// (append order), so this asked for the OLDEST 100 events this device
@@ -1189,13 +1414,21 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
// You cannot owe yourself.
return c.from_agent !== mailboxAddress;
});
if (candidates.length === 0) return [];
if (candidates.length === 0) return { owed: [], dormant: [] };
const mine = parseEvents(mineRaw);
const inbound = parseEvents(inboundRaw);
return candidates.filter((c) => !isAnswered(c, mine) && !isWithdrawn(c, inbound));
const live = candidates.filter((c) => !isAnswered(c, mine) && !isWithdrawn(c, inbound));
// PARTITION, not filter. A dormant ask is still owed in the protocol
// sense — it has no terminal reply and must not be closed — so it stays in
// the return value where the mailbox can still account for it, and only
// the ANNOUNCING paths below treat the two halves differently.
const owed: LogEvent[] = [];
const dormant: LogEvent[] = [];
for (const c of live) (evaluateDormancy(c.metadata).dormant ? dormant : owed).push(c);
return { owed, dormant };
} catch {
// Fail silent and open: a mailbox read must never break a session.
return [];
return { owed: [], dormant: [] };
}
}
@@ -1218,10 +1451,15 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
async function deriveClosed(): Promise<LogEvent[]> {
if (!mailboxEnabled || !available) return [];
try {
// `order: "desc"` for the same reason as deriveOwed: without it a 3.9.0
// server hands back the OLDEST window, so past 100 authored events this
// device's newest asks (and past 50 inbound, the replies that close them)
// fall outside the join. The selection below is order-independent
// (isStrictlyAfter), so this only fixes WHICH events are in the window.
const [mineRaw, inboundRaw] = await Promise.all([
client.callTool("mempalace_event_list", { from_agent: mailboxAddress, limit: 100 }),
client.callTool("mempalace_event_list", { from_agent: mailboxAddress, limit: 100, order: "desc" }),
// NO status filter, deliberately: that omission is the entire fix.
client.callTool("mempalace_event_list", { to_agent: mailboxAddress, limit: 50 }),
client.callTool("mempalace_event_list", { to_agent: mailboxAddress, limit: 50, order: "desc" }),
]);
// Correlations this device opened as a DIRECTED ask. A broadcast owes
// nobody a reply, so it cannot be closed by one either.
@@ -1286,6 +1524,15 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
})
.join("\n");
const DORMANT_NOTE =
"Each of these declares a `dormant_unless` predicate in its metadata that STILL " +
"matches this device's current state, so the work it describes is not yet " +
"possible. They are NOT closed and NOT answered — leave them open. They return to " +
"the announced owed-set by themselves the moment a baseline stops matching. If a " +
"predicate cannot be evaluated at all (missing file, unparseable baseline, unknown " +
"kind) the ask is announced as owed instead, deliberately: dormancy has to be " +
"proven, never assumed.";
const OWED_HOWTO =
"To close one, append an event on the SAME correlation_id with a terminal status " +
"(applied/superseded/failed/blocked). An ack alone does NOT clear it, and neither " +
@@ -1393,7 +1640,8 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
lastMailboxPollAt = Date.now();
void (async () => {
try {
const [owed, closed] = await Promise.all([deriveOwed(), deriveClosed()]);
const [owedSplit, closed] = await Promise.all([deriveOwed(), deriveClosed()]);
const { owed, dormant } = owedSplit;
const now = Date.now();
const unseen = (list: LogEvent[]): LogEvent[] =>
list.filter((e) => {
@@ -1421,6 +1669,13 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
`${CLOSED_NOTE}\n\n${formatOwed(newsRaw)}`,
);
}
// A dormant ask NEVER opens this window on its own — the early return
// above is what actually stops the nag. But once the window is open for
// something else, one line of count is nearly free and keeps a waiting
// ask from feeling lost between session starts.
if (dormant.length > 0) {
blocks.push(`${dormant.length} dormant ask(s), not listed here. ${DORMANT_NOTE}`);
}
pi.sendMessage(
{
customType: "mempalace-mailbox",
@@ -1485,10 +1740,11 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
sections.push(`## mempalace_diary_read\n\n(error: ${(err as Error).message})`);
}
// Tier 1 mailbox: what this device owes a reply to, derived rather than read
// off a filter. deriveOwed() swallows its own failures and returns [], so a
// broken palace costs a missing section, never a broken wake-up.
// off a filter. deriveOwed() swallows its own failures and returns an empty
// split, so a broken palace costs a missing section, never a broken wake-up.
if (mailboxEnabled) {
const [owed, closed] = await Promise.all([deriveOwed(), deriveClosed()]);
const [owedSplit, closed] = await Promise.all([deriveOwed(), deriveClosed()]);
const { owed, dormant } = owedSplit;
if (owed.length > 0) {
// Record what this injection showed, so the first mid-session poll does
// not re-announce the identical list minutes later. Without this the
@@ -1502,6 +1758,18 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
`this device. ${OWED_HOWTO}\n\n${formatOwed(owed)}`,
);
}
// Once per session, WITH ids. The wake-up injection is the one place a
// dormant ask should be visible: "what is this device still carrying?" is a
// session-start question, and answering it here is what keeps the mid-session
// silence honest rather than concealing. Deliberately NOT added to
// `surfaced`: that map exists to suppress repeats of ANNOUNCED work, and a
// dormant ask must stay announceable the instant its baseline moves.
if (dormant.length > 0) {
sections.push(
`## logstream mailbox (${dormant.length} dormant — waiting, nothing owed yet)\n\n` +
`${DORMANT_NOTE}\n\n${formatOwed(dormant)}`,
);
}
const news = closed.filter((e) => e.id && !surfaced.has(e.id));
if (news.length > 0) {
const now = Date.now();
+156
View File
@@ -0,0 +1,156 @@
#!/usr/bin/env bash
# test-dormancy.sh — exercise the mailbox dormancy predicate (dormant_unless).
#
# Why this exists as a script and not a note: `evaluateDormancy` decides whether
# an ask is SHOWN to the agent, so a silent regression there does not look like a
# bug — it looks like an empty mailbox. The positive arms matter as much as the
# negative ones: a predicate that never fires makes the feature a no-op, and a
# predicate that fires too eagerly hides real work. Both arms run here.
#
# It cannot import extensions/pi/mempalace.ts in place, because that file imports
# `typebox`, which pi provides at RUNTIME and this repo has no node_modules for.
# So it copies the file into a temp tree with the resolved deps symlinked beside
# it. The copy is made BY this script on every run, so it cannot drift from the
# source the way a vendored duplicate would.
#
# Fixtures are created here rather than read from the host: the first draft used
# /etc/hostname and /etc/pi-devbox/build-manifest.json, which made the positive
# arms pass only on a pi-devbox container and silently flip to "not dormant"
# anywhere else — a device-dependent test that reports success by doing nothing.
#
# Exit codes, deliberately distinct:
# 0 every case behaved as specified
# 1 at least one case FAILED (a real regression)
# 2 INCONCLUSIVE — deps could not be resolved, so nothing was proven
# (never 0: a test that skips quietly is the failure mode it should catch)
set -uo pipefail
# ── Args ──────────────────────────────────────────────────────────────────────
if [[ "${1:-}" == "--help" || "${1:-}" == "-h" ]]; then
sed -n '2,25p' "$0" | sed 's/^# \?//'
exit 0
fi
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
SRC="$REPO_ROOT/extensions/pi/mempalace.ts"
[[ -r "$SRC" ]] || {
echo "INCONCLUSIVE: cannot read $SRC" >&2
exit 2
}
# ── Resolve pi's runtime deps (discovered, never hardcoded) ───────────────────
GLOBAL_ROOT="$(npm root -g 2>/dev/null || true)"
PI_PKG=""
for cand in \
"${GLOBAL_ROOT:+$GLOBAL_ROOT/@earendil-works/pi-coding-agent}" \
"/usr/lib/node_modules/@earendil-works/pi-coding-agent" \
"/usr/local/lib/node_modules/@earendil-works/pi-coding-agent"; do
[[ -n "$cand" && -d "$cand" ]] && {
PI_PKG="$cand"
break
}
done
[[ -n "$PI_PKG" && -d "$PI_PKG/node_modules/typebox" ]] || {
echo "INCONCLUSIVE: pi-coding-agent / typebox not resolvable (looked under 'npm root -g')." >&2
echo " Nothing was proven. Install pi, or run this where pi is installed." >&2
exit 2
}
# Node must be able to strip types from a .ts entry point (Node >= 22.6).
node -e 'process.exit(0)' 2>/dev/null || {
echo "INCONCLUSIVE: no usable node" >&2
exit 2
}
# ── Build the throwaway tree ──────────────────────────────────────────────────
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
mkdir -p "$WORK/extensions/pi" "$WORK/node_modules/@earendil-works"
cp "$SRC" "$WORK/extensions/pi/mempalace.ts"
ln -s "$PI_PKG/node_modules/typebox" "$WORK/node_modules/typebox"
ln -s "$PI_PKG" "$WORK/node_modules/@earendil-works/pi-coding-agent"
printf '{"type":"module"}\n' > "$WORK/package.json"
# Prove the copy is the source, so a PASS cannot be about a stale file.
if ! cmp -s "$SRC" "$WORK/extensions/pi/mempalace.ts"; then
echo "INCONCLUSIVE: the copy differs from the source" >&2
exit 2
fi
# ── Fixtures, owned by this test ──────────────────────────────────────────────
printf '{"release_tag":"v1.9.4","nested":{"deep":{"leaf":"found"}},"count":3}\n' > "$WORK/manifest.json"
printf 'not json at all\n' > "$WORK/notjson.txt"
: > "$WORK/stamp"
# ── The truth table ───────────────────────────────────────────────────────────
cat > "$WORK/run.mjs" <<'EOF'
import { statSync } from "node:fs";
import { evaluateDormancy } from "./extensions/pi/mempalace.ts";
const W = process.env.WORK;
const STAMP = `${W}/stamp`;
const MAN = `${W}/manifest.json`;
// Floor to whole seconds: that is the contract, and the residue below proves
// why the implementation must do the same.
const atBaseline = new Date(Math.floor(statSync(STAMP).mtimeMs / 1000) * 1000).toISOString();
const mt = (baseline, path = STAMP) => ({ kind: "file_mtime", path, baseline });
const jf = (field, baseline, path = MAN) => ({ kind: "json_field", path, field, baseline });
const cases = [
// No predicate => behave exactly as before the feature existed.
["no metadata", undefined, false],
["metadata without the key", { foo: "bar" }, false],
["dormant_unless is a string", { dormant_unless: "release != x" }, false],
["dormant_unless is empty", { dormant_unless: [] }, false],
["more than 8 conditions", { dormant_unless: Array(9).fill(mt(atBaseline)) }, false],
// POSITIVE arms. If these ever read false the feature is inert.
["file_mtime at baseline", { dormant_unless: [mt(atBaseline)] }, true],
["json_field at baseline", { dormant_unless: [jf("release_tag", "v1.9.4")] }, true],
["both at baseline", { dormant_unless: [jf("release_tag", "v1.9.4"), mt(atBaseline)] }, true],
["dotted field path", { dormant_unless: [jf("nested.deep.leaf", "found")] }, true],
["number vs string baseline", { dormant_unless: [jf("count", "3")] }, true],
// The trigger firing: ANY condition differing wakes the ask.
["file_mtime moved", { dormant_unless: [mt("2020-01-01T00:00:00Z")] }, false],
["json_field moved", { dormant_unless: [jf("release_tag", "v1.9.5")] }, false],
["one same, one moved", { dormant_unless: [mt(atBaseline), jf("release_tag", "v1.9.5")] }, false],
// FAIL-VISIBLE arms: dormancy unproven => announce.
["missing file", { dormant_unless: [mt(atBaseline, "/nonexistent/path")] }, false],
["unparseable baseline", { dormant_unless: [mt("not-a-date")] }, false],
["relative path", { dormant_unless: [{ kind: "file_mtime", path: "etc/hostname", baseline: atBaseline }] }, false],
["unknown kind", { dormant_unless: [{ kind: "uptime_lt", path: "/etc/hostname", baseline: "1d" }] }, false],
["missing json field", { dormant_unless: [jf("no_such_field", "x")] }, false],
["file is not JSON", { dormant_unless: [jf("a", "b", `${W}/notjson.txt`)] }, false],
["field is not scalar", { dormant_unless: [jf("nested", "x")] }, false],
["condition is not an object", { dormant_unless: ["release_tag"] }, false],
["baseline is an object", { dormant_unless: [{ kind: "file_mtime", path: STAMP, baseline: {} }] }, false],
];
let pass = 0;
const failed = [];
for (const [name, meta, want] of cases) {
const v = evaluateDormancy(meta);
const ok = v.dormant === want;
if (ok) pass++;
else failed.push(name);
console.log(
` ${ok ? "PASS" : "FAIL"} dormant=${String(v.dormant).padEnd(5)} want=${String(want).padEnd(5)} ` +
`${name.padEnd(28)} :: ${v.reason}`,
);
}
console.log(`\n ${pass}/${cases.length} passed`);
if (failed.length) console.log(` FAILED: ${failed.join(", ")}`);
// Recorded because it is the reason file_mtime floors to seconds: a real mtime
// carries sub-second residue that an ISO baseline does not.
console.log(` (stamp mtimeMs=${statSync(STAMP).mtimeMs}, baseline=${atBaseline})`);
process.exit(failed.length === 0 ? 0 : 1);
EOF
cd "$WORK" || exit 2
export WORK
node "$WORK/run.mjs"
exit $?
+134
View File
@@ -0,0 +1,134 @@
#!/usr/bin/env bash
# test-mcp-call-timeout.sh — the per-call deadline of RemoteMcpClient in
# extensions/pi/mempalace.ts, and the override that the feed's mine relies on.
#
# WHY THIS EXISTS. The feed tick calls `mempalace_mine` through
# `client.callTool()`. Its own deadline (MEMPALACE_FEED_MINE_TIMEOUT_MS,
# 300 000) was raised in 2026-09 because the mine legitimately takes 30-60 s on
# a shared single-writer palace — but callTool() carried no way to pass that
# deadline down, so the transport's generic per-request timeout
# (MEMPALACE_MCP_TIMEOUT_MS, 60 000) fired first, and operators saw
# feed (tick) failed: mempalace remote request 'tools/call' failed: timed out after 60000ms
# instead of the message the 2026-09 change had aimed at. The outer race was
# unreachable in practice. This harness pins both halves of the contract:
# 1. a plain callTool() still honours the short per-request timeout
# (a query taking that long really is wedged — keep it short);
# 2. callTool(name, args, { timeoutMs }) honours the override, so a long
# server-side job can be given its own, longer, deadline.
# Before the fix, (2) fails: the third argument was silently ignored.
#
# Like test-owed-withdrawal.sh, it runs the SHIPPED text: the class is cut out
# of mempalace.ts by brace matching, type-stripped with node's own stripper,
# and driven against a local JSON-RPC server that delays tools/call.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
SRC="${1:-$REPO_ROOT/extensions/pi/mempalace.ts}"
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
[ -r "$SRC" ] || { echo "FAIL: cannot read $SRC" >&2; exit 2; }
# ---------------------------------------------------------------- extractor ---
cat >"$WORK/extract.mjs" <<'EXTRACT'
import { readFileSync, writeFileSync } from "node:fs";
import { stripTypeScriptTypes } from "node:module";
const src = readFileSync(process.argv[2], "utf8");
/** Cut a top-level `const NAME = ...;` out of the source, verbatim. */
function decl(name) {
const start = src.indexOf(`const ${name} =`);
if (start < 0) throw new Error(`declaration not found: ${name}`);
return src.slice(start, span(start));
}
/** Cut `class NAME ... { ... }` out of the source, verbatim. */
function klass(name) {
const start = src.indexOf(`class ${name} `);
if (start < 0) throw new Error(`class not found: ${name}`);
return src.slice(start, span(start, true));
}
/** End offset of the statement starting at `start` (brace-matched). */
function span(start, braceOnly = false) {
let depth = 0, inStr = null, sawBrace = false;
for (let i = start; i < src.length; i++) {
const c = src[i], prev = src[i - 1];
if (inStr) { if (c === inStr && prev !== "\\") inStr = null; continue; }
if (c === '"' || c === "'" || c === "`") { inStr = c; continue; }
if (c === "/" && src[i + 1] === "/") { i = src.indexOf("\n", i); if (i < 0) break; continue; }
if (c === "/" && src[i + 1] === "*") { i = src.indexOf("*/", i) + 1; continue; }
if (c === "{" || c === "(" || c === "[") { depth++; if (c === "{") sawBrace = true; }
else if (c === "}" || c === ")" || c === "]") { depth--; if (braceOnly && sawBrace && depth === 0) return i + 1; }
else if (!braceOnly && c === ";" && depth === 0) return i + 1;
}
throw new Error(`unterminated statement at ${start}`);
}
const parts = [decl("num"), decl("REMOTE_PROTOCOL_VERSION"), decl("REMOTE_CLIENT_INFO"), klass("RemoteMcpClient")];
for (const [i, p] of parts.entries()) if (p.length < 30) throw new Error(`extraction ${i} implausibly short: ${p}`);
if (!parts[3].includes("tools/call")) throw new Error("RemoteMcpClient does not mention tools/call");
const ts = parts.join("\n\n") + "\nexport { RemoteMcpClient };\n";
writeFileSync(process.argv[3], stripTypeScriptTypes(ts, { mode: "strip" }));
EXTRACT
node --no-warnings "$WORK/extract.mjs" "$SRC" "$WORK/client.mjs" || exit 2
node --check "$WORK/client.mjs" || { echo "FAIL: extracted client does not parse" >&2; exit 2; }
echo "[extract] pulled RemoteMcpClient from $(basename "$SRC") ($(wc -c <"$WORK/client.mjs") bytes)"
# ------------------------------------------------------------------ harness ---
cat >"$WORK/run.mjs" <<'RUN'
import { createServer } from "node:http";
import { RemoteMcpClient } from "./client.mjs";
// A sessionless JSON-RPC server like `mempalace-mcp --transport http`:
// initialize / tools/list answer at once; tools/call sleeps SLOW_MS first.
const SLOW_MS = 1500;
const server = createServer((req, res) => {
let body = "";
req.on("data", (c) => (body += c));
req.on("end", () => {
const msg = JSON.parse(body);
if (msg.id === undefined) { res.writeHead(202); res.end(); return; } // notification
const reply = (result) => {
res.writeHead(200, { "content-type": "application/json" });
res.end(JSON.stringify({ jsonrpc: "2.0", id: msg.id, result }));
};
if (msg.method === "initialize") return reply({ protocolVersion: "2024-11-05", capabilities: {}, serverInfo: { name: "fake", version: "0" } });
if (msg.method === "tools/list") return reply({ tools: [{ name: "slow", description: "", inputSchema: { type: "object" } }] });
if (msg.method === "tools/call") return void setTimeout(() => reply({ content: [{ type: "text", text: "done" }] }), SLOW_MS);
reply({});
});
});
await new Promise((r) => server.listen(0, "127.0.0.1", r));
const url = `http://127.0.0.1:${server.address().port}/mcp`;
let failures = 0;
const check = (ok, label) => { console.log(`${ok ? "ok " : "FAIL"} ${label}`); if (!ok) failures++; };
// Short per-request deadline, deliberately below SLOW_MS.
process.env.MEMPALACE_MCP_TIMEOUT_MS = "400";
const client = new RemoteMcpClient(url);
await client.start();
check(client.alive === true, "start(): initialize + tools/list complete, client alive");
// 1. A plain call honours the short deadline — this is the wedged-query guard.
let err = null;
const t0 = Date.now();
try { await client.callTool("slow", {}); } catch (e) { err = e; }
check(err !== null && /timed out after 400ms/.test(err.message), `plain callTool() rejects at the per-request deadline (${err && err.message})`);
check(Date.now() - t0 < SLOW_MS, "…and rejects BEFORE the server would have answered");
check(client.alive === false, "a timeout marks the client unhealthy (documented side effect; ensureAlive() revives)");
// 2. A call with its own deadline outlives the generic one — the feed's mine.
await client.ensureAlive();
err = null;
let result = null;
try { result = await client.callTool("slow", {}, { timeoutMs: 5000 }); } catch (e) { err = e; }
check(err === null && result && result.content?.[0]?.text === "done", `callTool(name, args, { timeoutMs: 5000 }) waits past the generic deadline and gets the result (${err ? err.message : "ok"})`);
// 3. The override is itself a deadline, not "no deadline".
err = null;
try { await client.callTool("slow", {}, { timeoutMs: 200 }); } catch (e) { err = e; }
check(err !== null && /timed out after 200ms/.test(err.message), `callTool(..., { timeoutMs: 200 }) rejects at ITS deadline (${err && err.message})`);
server.close();
console.log(failures === 0 ? "PASS: all assertions held" : `FAIL: ${failures} assertion(s) failed`);
process.exit(failures === 0 ? 0 : 1);
RUN
cd "$WORK" && node --no-warnings run.mjs
+44 -3
View File
@@ -23,7 +23,9 @@
# withdrawal had no effect, so tor-ms22 was still being told it owed a reply 41h
# later for a release it never installed.
#
# Usage: scripts/test-owed-withdrawal.sh [source.ts] (exit 0 = all rules behave)
# Usage: scripts/test-owed-withdrawal.sh [source.ts]
# exit 0 = all rules behave 1 = a rule broke
# exit 2 = source unreadable or does not parse 3 = the gate itself cannot run
#
# The optional argument exists so the suite can be pointed at a deliberately
# MUTATED copy of the source to prove it is sensitive — a suite that has never
@@ -39,8 +41,47 @@ trap 'rm -rf "$WORK"' EXIT
[ -r "$SRC" ] || { echo "FAIL: cannot read $SRC" >&2; exit 2; }
# A gate that cannot run must not pass — the standing rule in this repo.
node --experimental-strip-types --check "$SRC" \
|| { echo "FAIL: $SRC does not parse" >&2; exit 2; }
#
# NOT `node --check`: it does not type-strip, so it rejects ANY TypeScript —
# `const x: number = 1` included, not just the inline type-import at the top of
# mempalace.ts. It passed on node 22.x and stopped passing on node 24.x
# (measured on pi-devbox v1.9.1, node v24.21.0: this gate exited 2 on an
# unmodified mempalace.ts, so all 17 assertions below refused to run). Strip
# first, then syntax-check the emitted JS. `mode: "strip"` blanks type syntax
# without moving anything, so byte offsets and line numbers survive and a
# reported error line still points at the right line of the ORIGINAL .ts.
#
# The two failure modes are reported separately and on purpose. "Cannot strip"
# is a fact about the toolchain; "does not parse" is a fact about the source.
# Collapsing them is what made this very defect present itself as
# "mempalace.ts does not parse" when mempalace.ts was fine.
#
# SC2016 is disabled deliberately: the single quotes are the point. What follows
# is JavaScript, and `${process.version}` must reach node, not be expanded by the
# shell first.
# shellcheck disable=SC2016
node --no-warnings -e '
const { readFileSync, writeFileSync } = require("node:fs");
const { stripTypeScriptTypes } = require("node:module");
if (typeof stripTypeScriptTypes !== "function") {
console.error(`FAIL: node ${process.version} cannot strip TypeScript ` +
`(module.stripTypeScriptTypes needs >= 22.13) — the gate is unavailable, ` +
`which is NOT a statement about the source`);
process.exit(3);
}
let js;
try {
js = stripTypeScriptTypes(readFileSync(process.argv[1], "utf8"), { mode: "strip" });
} catch (err) {
// The stripper is itself a parser, so a syntax error lands HERE, not in the
// --check below. This is a statement about the source: exit 2, not 3.
console.error(`FAIL: ${process.argv[1]} does not parse: ${err.message}`);
process.exit(2);
}
writeFileSync(process.argv[2], js);
' "$SRC" "$WORK/stripped.mjs" || exit $?
node --check "$WORK/stripped.mjs" \
|| { echo "FAIL: $SRC does not parse (stripped output rejected)" >&2; exit 2; }
# ---------------------------------------------------------------- extractor ---
cat >"$WORK/extract.mjs" <<'EXTRACT'