Compare commits

...

19 Commits

Author SHA1 Message Date
Joakim Persson 975ab92943 feat(pi): let an ask declare dormancy, so waiting work stops nagging
A first-boot acceptance ask is planted deliberately unanswerable: it describes
work that becomes possible only when the device is next recreated, and it must
STAY owed until then, because closing it early to tidy the mailbox is exactly how
that work gets lost. deriveOwed could only see "directed, open, not answered, not
withdrawn", so such an ask was announced at every session start and every poll
for as long as it was correctly waiting -- measured at three days running on
emb-7kj4vr4g (2026-09-28 -> 2026-10-01), and across three consecutive releases
before that. The ask was right; announcing it was wrong, and the cost landed on
the human reading the window, who is the one reader that cannot filter it.

An ask may now declare the condition under which it is merely waiting:

  "dormant_unless": [
    { "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json",
      "field": "release_tag", "baseline": "v1.9.4" },
    { "kind": "file_mtime", "path": "/etc/hostname",
      "baseline": "2026-09-22T18:12:49Z" }
  ]

Dormant while EVERY condition still matches its baseline; live the moment ANY
differs -- which is the trigger those asks already stated in prose ("act when
EITHER differs"), now in a form the bridge can check. Two kinds, local files
only, no expression language, no shell, no network: a general evaluator in the
path that decides whether work is VISIBLE is a far worse trade than a clumsy
schema.

DORMANCY IS PROVEN, NEVER ASSUMED. Every unevaluable predicate announces the ask
instead of hiding it: missing file, unreadable file, unparseable baseline,
unknown kind, vanished field, relative path, more than eight conditions. The
dangerous failure here is not a spurious nag but work that disappears because a
predicate could not be evaluated -- indistinguishable from the ask being lost,
and not surfacing until a release needed it. An ask with no dormant_unless
behaves exactly as before, so this is backward compatible by construction.

Withheld from the ANNOUNCEMENT, never from the mailbox: deriveOwed now returns a
partition {owed, dormant} rather than a flat list, the wake-up injection lists
dormant asks once per session with ids, and a mid-session poll adds only a count
and only when the window is already open for something else. Dormant asks are
deliberately NOT added to the `surfaced` map, so one becomes announceable the
instant its baseline moves.

file_mtime compares WHOLE SECONDS in UTC. A filesystem mtime carries sub-second
residue (measured: /etc/hostname at .773761009) that a reported ISO baseline
never will, so comparing raw milliseconds would mark every such predicate
permanently "changed" -- silently disabling the feature while appearing to work.
The test records the residue for that reason.

scripts/test-dormancy.sh is this repo's first test: 22 cases, positive and
negative arms both, because a predicate that never fires makes the feature inert
and one that fires too eagerly hides real work. It copies the extension into a
temp tree with pi's typebox symlinked beside it (the copy is made per-run, so it
cannot drift like a vendored duplicate), owns its own fixtures rather than
reading /etc paths -- the first draft passed only on a pi-devbox container and
would have silently flipped to "not dormant" anywhere else -- and exits 2 for
INCONCLUSIVE rather than 0, since a test that skips quietly is the failure mode
it exists to catch. shellcheck clean.
2026-10-01 23:59:36 +02:00
joakimp 2167a1b033 fix(mailbox): pass order: "desc" on every cursor-less event_list call
Three of the five `mempalace_event_list` calls in the pi extension relied on
the server's default ordering: deriveOwed's `status: "open"` candidate query
(limit 50) and both windows in deriveClosed (from_agent limit 100, to_agent
limit 50). deriveOwed's other two calls already say `order: "desc"` and carry
the comment explaining why: on mempalace <= 3.9.0 the default is `asc`, so a
cursor-less `limit: N` returns the OLDEST N events, and once a device passes N
authored events its newest asks and the replies that close them fall outside
the join window. deriveClosed had exactly that latent truncation.

Why now: mempalace 3.10.0 (2026-09-15) flips the cursor-less default to
newest-first ("Logstream listings default to the newest events when no cursor
is given"). That change is server-side -- the extension talks to the fleet
hub over MEMPALACE_REMOTE_URL, so it lands when the hub upgrades, not when a
client does -- and it would have silently FIXED deriveClosed on 3.10.0 while
leaving it broken against any 3.9.0 hub. A mailbox verdict that depends on
which server version answers is the wrong shape. Saying `order` explicitly
makes both derive* functions read the same window on either version.

The selection in deriveClosed (isStrictlyAfter per correlation) is
order-independent, so this changes WHICH events are in the window, not how
the winner is picked. No behaviour change on a device with < 50 inbound and
< 100 authored events; the fleet hub is at 176 events total today, so the
window was about to matter.

Verified: file compiles under pi's own loader (jiti 2.7.0 from the installed
pi-coding-agent) before and after; `order:"desc"` literal count 3 -> 7, which
is the 3 added arguments plus 1 comment mention. `node --check` and a bare
`tsc` were tried first and both fail identically on HEAD (inline `type`
imports / no @types/node) -- checker faults, not this change.
2026-09-22 14:15:22 +02:00
joakimp 817b3a82b7 fix(feed): the mine's deadline never reached the transport; 60 s cut it off
The 2026-09 change raised MEMPALACE_FEED_MINE_TIMEOUT_MS to 300 000 and
raced it against client.callTool("mempalace_mine"). But callTool() had no
way to carry a deadline, so every call went out under the transport's
generic per-request timeout (MEMPALACE_MCP_TIMEOUT_MS, 60 000), which fired
first on every honest 60 s+ mine. Operators saw

    feed (tick) failed: mempalace remote request 'tools/call' failed:
    timed out after 60000ms

instead of the message the change had aimed at, and the 300 s was
unreachable. On stdio it was worse than noise: that transport kills the
server child on timeout, so the mine was actually aborted at 60 s.

callTool(name, args, { timeoutMs }) now passes a per-call deadline to both
transports; feedPalace() uses it for the mine. Plain calls keep the short
default — a query taking 60 s is still wedged. The Promise.race stays as the
liveness guard for a transport with its timeout disabled (0).

scripts/test-mcp-call-timeout.sh cuts RemoteMcpClient out of the shipped
file (as test-owed-withdrawal.sh does for the mailbox predicates), drives it
against a local JSON-RPC server that delays tools/call, and asserts: plain
call rejects at the generic deadline; the override outlives it; the override
is itself a deadline. Fails on the previous commit (2 of 6), passes here.
Vendored-copy delta noted in the header; protocol untouched, sync token
unchanged (check-mcp-client-sync.sh passes against pi-extensions).
2026-09-18 17:04:59 +02:00
joakimp dab989b068 fix(mailbox-tests): the owed-set suite's own gate refused to let it run on node 24
scripts/test-owed-withdrawal.sh has not executed a single assertion since the
image moved to node 24. Its precondition line

    node --experimental-strip-types --check "$SRC"

exits 2 on an unmodified extensions/pi/mempalace.ts, and the script is
`set -euo pipefail` with `|| exit 2`, so all 17 assertions and both regression
guards were skipped. Measured on pi-devbox v1.9.1, node v24.21.0.

CAUSE, MEASURED AND NARROWER THAN IT LOOKS. The reported error names the inline
type-import at mempalace.ts:92, which invites the reading "that import is
unusual". It is not the import: `node --check` does not type-strip AT ALL. A file
whose entire content is `const x: number = 1;` fails identically, with or without
--experimental-strip-types, while `node --experimental-strip-types <same file>`
executes it fine. --check has never been type-aware; execution is what gained
stripping. So this gate could never validate TypeScript, on any node.

WHY IT USED TO PASS -- INFERENCE, NOT MEASUREMENT. e2b060a's message records this
suite running on 2026-09-09 with a control pass and four mutation kills, when the
image shipped node v22.23.2; v1.9.1 ships v24.21.0. No node 22 exists on the box
where this was diagnosed, so the counterfactual was not executed. Treat "22
stripped for --check and 24 stopped" as a hypothesis consistent with the record,
not as a measured cause. What IS measured is the present-tense behaviour above.

WHY THIS MATTERS MORE THAN A RED TEST. This suite exists because isWithdrawn is
the one rule in the extension that can go wrong SILENTLY -- a wrong rule does not
throw and does not log, it makes a real unanswered ask vanish from a mailbox
forever. The rule shipped in e2b060a and has been baked since v1.9.1; the thing
that makes its failure mode visible has been dark for the same period. The
mitigating half: it failed CLOSED (exit 2, loud), never vacuously green. A gate
that cannot run must not pass, and it did not.

THE FIX. Strip first, then syntax-check the emitted JS: version-stable, and it
still refuses malformed input. `mode: "strip"` blanks type syntax without moving
anything, so offsets and line numbers survive and a reported error line still
points at the right line of the original .ts (69157 B in, 69157 B out).

THE EXIT CODES ARE NOW SPLIT, AND THAT IS THE POINT. 3 = the gate itself cannot
run (no module.stripTypeScriptTypes, i.e. node < 22.13). 2 = the source does not
parse. Collapsing the two is how this defect disguised itself: it printed
"mempalace.ts does not parse" while mempalace.ts was fine, sending a reader to
inspect the wrong file. Note the stripper is itself a parser, so a genuine syntax
error surfaces as an exception from the strip call rather than from --check; that
path is caught and reported as 2, not 3. Both remain failures. Neither passes.

VERIFIED IN SIX DIRECTIONS on this image, each expectation written down first:
  unmutated source          -> rc=0, PASSED (17/17)
  malformed TypeScript      -> rc=2, "does not parse: Expression expected"
  stripper made unavailable -> rc=3, "cannot strip", and NOT "does not parse"
                               (simulated by doctoring the runtime through
                               NODE_OPTIONS, so the shipped line ran as shipped)
  M1 marker requirement removed -> FAILED: 3   (no-marker, other-thread, prose)
  M2 third-party guard removed  -> FAILED: 1
  M3 to_agent guard removed     -> FAILED: 2   (broadcast and wrong-device share it)
  M4 ordering guard removed     -> FAILED: 1
The four kill counts are the same ones e2b060a recorded, so sensitivity is
restored rather than merely asserted.

WORTH KNOWING FOR THE NEXT PERSON WHO MUTATES THIS: the extractor carries its own
guard that rejects an isWithdrawn which no longer mentions "withdraws", so the
obvious M1 (delete the marker check outright) is refused before any assertion
runs -- correctly, but it looks like a crash. Mutate to `... || true` instead,
which removes the requirement while keeping the key mentioned.

SC2016 is disabled on the node invocation with a stated reason: the single quotes
are deliberate, the payload is JavaScript and `${process.version}` must reach
node rather than the shell. Checked that the file is shellcheck-clean at default
severity, as it was before this change, and at the -S error severity the
pi-devbox gate uses.

NOT FIXED HERE. Nothing in CI runs this suite -- it is a script an operator
invokes, which is precisely why a gate that fails loudly still went unnoticed
across a node bump. The harness's own runner two hundred lines below still passes
--experimental-strip-types, deliberately: it is a no-op on 24 and required on 22,
so it keeps the suite runnable on both. And the real reason this was found at all
is unrelated to CI: isWithdrawn was being exercised end-to-end on a released
image against the live logstream for the first time (project/pi-devbox seq 140),
and the suite was reached for as corroboration.
2026-09-14 16:27:29 +02:00
Joakim Persson e68ee2071c docs(pi-ext): document what the mine deadline does, and what the message means
Two corrections to the operator-facing docs, both exposed by 309980b.

The env table listed the default as 30000, which is now wrong, and described the
var as capping "the mempalace_mine call". It never did: it bounds how long the
extension WAITS. The mine keeps running on the server. That exact misreading is
what made a 30s deadline look safe on a call measured at 30-60s.

Added a Debugging entry for "feed (tick) failed: mine timed out after ...ms",
because every operator on this fleet has seen it and it was documented nowhere.
It states the three things a reader needs: nothing was lost (the transcript is
staged before the mine, and mine --mode convos dedups by source_file and is
idempotent); do NOT retry harder from the client, because the palace is a single
writer and a blind retry turns one slow mine into a queue; and after 309980b the
message should not appear on a healthy fleet, so if it does it now MEANS
something -- a mine exceeding five minutes, i.e. look at palace size or another
writer holding the lock rather than raising the timeout again.
2026-09-10 20:58:12 +02:00
Joakim Persson 309980b62c fix(pi-ext): stop the feed tick from launching overlapping mines
"[mempalace ext] feed (tick) failed: mine timed out after 30000ms" was parked as
cosmetic on 2026-08-27. It is not cosmetic: the tight deadline was the trigger,
but the defect is a positive feedback loop that puts multiple writers on a
single-writer palace.

lastFeedAt was assigned only AFTER a successful await. The Promise.race abandons
our WAIT and cannot cancel the server's work, so on a mine that takes longer
than the deadline -- measured at 30-60s in normal operation, against a 30s
deadline -- the catch ran with lastFeedAt UNCHANGED. That left both guards in
the agent_settled handler open at once: the debounce test
(`Date.now() - lastFeedAt < feedDebounceMs`) passed because lastFeedAt was still
stale, and feedInFlight was already null because `run` had settled. Every
subsequent settled turn therefore launched another mine on top of the one still
running, each making the next slower and the next timeout likelier -- which is
why operators saw the message many times per session instead of at most once per
10-minute debounce window.

Fix, three lines:
- move `lastFeedAt = Date.now()` to before the await, so a timeout still starts
  the debounce clock. A timeout is not a "did not happen": the mine is running
  server-side and `mine --mode convos` dedups by source_file and is idempotent.
- raise MEMPALACE_FEED_MINE_TIMEOUT_MS from 30_000 to 300_000. The mine is the
  slowest thing this extension does yet carried the tightest deadline: 4x
  tighter than the prepare step before it (120_000) and 10x tighter than the
  init handshake (300_000), a fast call. All three were introduced together in
  29e660e and this one was never revisited. 300_000 matches the init timeout
  because liveness is the only job left for this deadline -- it cannot cancel
  the server's work, so it must sit far above the slowest honest completion.

Simulated both guards over 10 minutes of settled turns at 20s intervals with a
60s mine: BEFORE 16 mines launched, 15 of them overlapping an already-running
mine; AFTER 2 launched, 0 overlapping. With a mine that exceeds even the new
deadline (400s): BEFORE 16/15, AFTER still 2/0 -- the lastFeedAt move is what
actually fixes it, and it holds even when the timeout still fires. The raise
stops the spurious message; the move stops the pile-up.

NOT fixed here: the message still goes to process.stderr, which pi renders into
the TUI input field. That needs a pi-side channel or a log file, and is tracked
separately. After this change the message should be rare, and when it does
appear it means something real: a mine exceeding five minutes.
2026-09-10 20:48:33 +02:00
joakimp e2b060a940 feat(mailbox): let a requester withdraw its own ask, with an explicit marker
RFC 003 §3.3 clause 3 clears an ask only on "no event OF YOURS", and the skill
states the consequence outright: "there is nothing anyone can do about it from
the other end". That asymmetry is deliberate and mostly right — owed-ness is a
statement about the RECIPIENT's accountability. It is wrong in exactly one case:
the requester retracting its own ask.

MEASURED COST. pi@mbp-m1-2020 withdrew a v1.8.13 rollout ask to pi@tor-ms22 at
seq 119 — terminal `superseded`, same correlation_id, metadata.closes naming the
thread, body "DO NOT SPEND A MINUTE ON v1.8.13" — and recorded it as done. It had
no effect: seq 119's from_agent is mbp, so it could never satisfy a join that
only inspects tor-ms22's own events. tor-ms22's next wake-up, 8h later and on the
first boot of the image shipping this very derivation, still listed the ask as
owed, 41h old, for a release it never installed. The failure is invisible from
the sender's side, which is why it went unnoticed for 41h.

WHY AN EXPLICIT MARKER RATHER THAN "ANY TERMINAL EVENT FROM THE REQUESTER".
This file's standing rule is that every failure stays on the noisy-but-visible
side: a resurfacing item costs one turn of human correction, a suppressed
unanswered ask is silent and permanent. Under the naive rule a requester
appending `applied` for its own bookkeeping — on the correlation, addressed to
me, before I ever replied — would silently delete a real obligation. So release
must be STATED: metadata.withdraws (canonical) or metadata.closes (already this
fleet's de-facto marker), whose value must NAME the ask — its correlation_id or
its event id. Prose does not count.

Five further guards, each pinned by a mutation test: only the original requester;
addressed to this device exactly, never '*'; terminal status; strictly after the
ask by hlc; and joined by ack_of or correlation_id.

This does NOT break the fixed point in §9.2. That section rejects letting
terminal directed events into the owed set, because then every closure mints a
fresh obligation. That is about CANDIDATES; this adds a CLEARER. A withdrawal is
terminal, so it can never be a candidate, and the asserting shape (open) and the
clearing shape (terminal) stay disjoint.

Also fixed here, because the new rule depends on the same windowing: the `mine`
query used event_list's DEFAULT `asc` order with limit 100, i.e. the OLDEST 100
events this device ever wrote. Once a device passes 100 authored events its most
RECENT replies fall out of the join window and every ask it just answered
resurfaces as owed. Latent, not theoretical — tor-ms22 was at ~20. Both windows
are now anchored at the newest end with order: "desc".

Verification: scripts/test-owed-withdrawal.sh, 17 assertions over VERBATIM
fixtures from the real log (seq 112/119/120/122). It extracts the predicates from
mempalace.ts by brace matching and runs the SHIPPED text rather than a pasted
copy — this repo has already paid for a divergent second copy. Sensitivity
proven by four mutations: removing the marker requirement flips exactly the 3
marker assertions, and disabling the third-party / broadcast / ordering guards
each flip exactly their own. Control passes.
2026-09-09 08:56:56 +02:00
joakimp e45f6b4301 feat(mailbox): surface replies that CLOSE this device's own asks
The mailbox could report what this device OWES, and structurally nothing else.
deriveOwed() queries the log with status:"open", and a reply that closes an ask is
by definition not open — so a peer answering my delegation was invisible to it at
every poll, forever, not just late.

Measured 2026-09-07: emb-7kj4vr4g closed correlation v1813-client-rollout-emb with
a task.reply at status=applied. Nothing was announced. The feature was CORRECT by
its own definition ("owed" = "you must reply", and nothing was owed) and wrong by
the operator's, who asked why no notification arrived. The single most useful thing
a fleet can say to a human is "the thing you asked for is done" — and that was the
one category it could not say.

Independent corroboration that this was a real gap and not a preference: emb's own
reply ends "The amd64 narrowing is filed as a drawer as well, SINCE A TERMINAL
EVENT REACHES NO MAILBOX." A peer had already diagnosed the hole and was routing
around it by hand.

deriveClosed(): my task.requests (correlation + directed) joined against inbound
events with NO status filter, keeping the newest TERMINAL_STATUS reply per
correlation. Verified by computing the exact predicate over the two real events:
the new path yields evt_20260907T183143 (applied); the old status:"open" query
yields 0. Announced once each (never resurfaced — a finished ask is not a nag),
and the copy warns that a peer who did the work is the likeliest party to have
found your premise wrong. Here both premises were wrong: mine that emb was amd64,
and emb's that tor-ms22 was therefore the last candidate.

Deliberate asymmetry, documented at the call site: a broadcast is excluded from the
owed set (it owes nobody) but allowed to CLOSE, since a peer answering on my
correlation_id is news however widely it was addressed.
2026-09-07 21:42:58 +02:00
joakimp 21023e7aa0 rfc-003: an event addressed to a nobody is write-only
§7.13 — the owed set keys on to_agent, and from_agent is whatever
string the writer puts there, so nothing stops authoring under an
identity no live session runs as. When the reply comes back addressed
to that string, it is stored and delivered to no one. Distinct from
§7.12: the defect is the address, not the status.

Measured 2026-08-27: two directed task.requests planted under a
synthetic from_agent (a provenance label, not a run identity) drew two
correct replies, addressed back to the label as the protocol requires.
Neither reply was ever delivered. One carried a live-credential
exposure finding; it sat unread for ~2h20m and was found only because
a human asked whether mail had arrived.

General rule stated, not just the instance: this fleet has already
produced the same failure shape three ways (synthetic sender above; an
event addressed to a decommissioned device; a device-identity case
mismatch). Permitted exception carried over from existing fleet
practice: a synthetic sender is fine for a controlled experiment, but
the body must then name the real reply-to identity.

Proposed (not implemented): warn at event_append when from_agent
differs from the writer's own session identity, using a DISTINCT
from_agent lookup to say whether anyone has ever authored under it —
stated honestly as a heuristic that catches 'never authored, certainly
unread' but not a decommissioned device that once did.

Decision 9 in the open-decisions list gets a numbered sibling, 10,
pointing at this section — related to decision 7 (agent-name
collision between two devices) but a different failure: not two
devices sharing a name, but one writer using a name nobody runs.
2026-08-27 23:34:20 +02:00
joakimp a361b71c40 ship: don't trust the mtime the exporter deliberately backdates
rsync --update skips a file whose mtime is not strictly newer than the
receiver's. The stage file's mtime IS the source transcript's mtime
(os.utime() at :903, "preserve session mtime for dedup stability"), so
re-exporting a session that has not been appended to since its last
ship produces a mtime that is not newer than what's already at the
receiver — exactly the case a redactor upgrade needs to ship, because
content differs while mtime does not. --update reports success and
sends nothing.

Reported and patched by pi@mbp-m1-2020 (evt_20260827T211925_9674a31da0b4,
artifact art_20260827T211839_fa52af563105, sha256 f217e47e…), measured
live: a scrubbed re-export of a dormant session (pi_01a03022-…f542)
sat unshipped in the palace host's inbox while every local signal
reported a clean stage, saved only because a host-side sweep happened
to rewrite the remote copy independently that same day.

--checksum compares content and ignores size/mtime entirely. Dropping
--update outright was considered and rejected: rsync's default quick
check already transfers on a SIZE difference alone, which is why the
observed case (33-byte placeholder vs a 43-byte token) would have been
masked as "fixed" by a change that only works until a redaction whose
placeholder happens to match the secret's length. os.utime() at :903
is untouched — its backdating is a separate, load-bearing design call
for dedup stability, out of scope for this fix.

Added scripts/test-rsync-ship-idempotency.sh: ships a file, rewrites
its content to an EQUAL-LENGTH string while restoring the original
mtime (what os.utime() does), ships again, asserts the receiver's
sha256 changed. Equal length is deliberate, not cosmetic — mismatched
lengths would pass via the quick check alone and prove nothing about
--checksum specifically; this is the same reasoning that ruled out
dropping --update. Verified the test discriminates: passes against
today's --checksum, fails against --update (checked by temporarily
substituting the flag in a copy, not committed).

Runs offline — a local rsync destination path exercises the same
size/mtime/checksum comparison as the ssh transfer, no palace or
network needed.
2026-08-27 23:29:30 +02:00
joakimp b2b50afcc1 rfc-003: propose a news surface that needs no cursor, and say why widening the owed set cannot work
Decision 9 (news vs obligations) gets a proposed direction in a new §9.2, following
how §9.1 promoted the retention decision.

The load-bearing part is the negative result. The tempting fix — let terminal
directed events into the owed set so a report addressed to a device reaches it —
breaks the derivation's fixed point. Owed-ness is "directed at me, not mine, and
not joined by a later terminal event of mine", so the asserting shape (open) and
the clearing shape (terminal) have to be disjoint. Make terminal events owed and a
reply becomes owed by its requester, whose closure is itself directed + terminal
and therefore owed by the original author: every closure mints a fresh obligation
and the loop never terminates. status="open" is not editorial taste about tone, it
is what makes owed-ness terminate. Worth writing down before someone "fixes" it.

The proposal itself avoids the cost that sinks the naive version. "Since your last
session" implies a per-device read cursor, and container-local state is exactly
what --force-recreate erases. No new state is needed: the device's own last
authored event is already a cursor, it lives in the shared log, and it is
comparable across replicas for the same reason the owed-set join now uses hlc
(§7.3). Named the non-obvious constraint too — the anchor must be computed PER
STREAM, because a global one lets a chatty stream advance past unread news in a
quiet one, and that failure is silent.

Also tightened §7.12: candidacy requires exactly `open`, not merely "non-terminal",
so a `claimed` announcement is as undelivered as a finished report. Claiming still
earns its keep on handoff-prone work (it is what distinguishes "nobody started"
from "someone started and the container died"), but its audience is a log reader,
not the requester — and it does not quiet the claimer's own mailbox either, which
is what the state machine's "open --> claimed does NOT clear" already implies.
2026-08-27 17:28:21 +02:00
Joakim Persson de59571966 docs: unclip fleet-memory's diagrams, and measure the host instead of guessing
Seven labels in fleet-memory.md were losing their last line for readers, in the
document held up as the model for user-facing docs. Found by running a new
mermaid checker against a repo it was not written for, which is the strongest
evidence available that this is a systemic defect class and not one author's slip.

Measured on the real Gitea render rather than a simulation. Gitea serves each
mermaid block from a sandboxed same-origin <iframe srcdoc>, so its
.markup{line-height:1.5!important} never reaches the labels and the page is
directly measurable. On the published version: 31 labels, 6 of them rendering
MORE lines than were authored, worst overflow +0.6px past the clip edge. Mermaid
measures a label with its own metrics, commits to a box, then clips whatever does
not fit; per-line rounding accumulates, so every cut label observed anywhere in
this fleet had four or five rendered lines.

Fixes, all preserving the information rather than deleting it:

- The five store nodes drop their third line; the content shapes they carried
  ("verbatim text, embedded", "typed facts, time-bounded", "addressed events and
  their artifacts", "links between rooms", "read by recency") move into the prose
  under the diagram, which is where a qualifier belongs anyway.
- "ask by ADDRESS + ORDER" soft-wrapped its own first line because a caps line
  nearly as wide as the box overflows first; shortened to "ask by ADDRESS", with
  "read in append order" added to the prose so the ordering property survives.
  The table already lists "address, correlation, append order".
- The decision diamond loses "machine or agent" to the sentence above it, which
  now asks whether "a specific machine or agent" must act — same words, no clip.
- "to_agent = the specific agent" and the BOTH node shortened; both are spelled
  out in full in the worked-examples table directly below.

Nothing was cut for brevity's sake: every phrase removed from a node reappears in
the prose or was already in the table beneath it.
2026-08-27 16:15:44 +02:00
joakimp f0bffd1b93 feed: chase the symlink, or fail-closed takes the whole fleet's feed down
The image installs /usr/local/bin/mempalace-pi-session as a symlink into
/opt/mempalace-toolkit/bin, and ${BASH_SOURCE[0]} reports the path the script was
INVOKED as, not the resolved target. So the sibling-module lookup pointed at
/usr/local/bin, the redactor was not there, and the fail-closed import did exactly
what it was told: refused to stage.

MEASURED, not theorised: /usr/local/bin/mempalace-pi-session --dry-run exited
"[FATAL] secret scrubber unavailable ... refusing to stage" while
/opt/mempalace-toolkit/bin/mempalace-pi-session --dry-run scrubbed 40 findings on
the same input. The symlink is how every device invokes it, so at the next image
bake every feeder tick on every machine would have stopped staging — a silent,
fleet-wide memory outage, which is a worse outcome than the leak the scrubber
exists to prevent. My own tests missed it by calling bin/... directly from the
checkout, i.e. the one invocation path the fleet never uses.

FIX: chase the symlink chain in portable shell and offer fallbacks instead of
betting on a single answer. MEMPALACE_REDACT_DIR is now colon-separated —
resolved dir, invoked dir, then the image's known install path /opt/... — and the
Python side inserts the first existing candidate. readlink -f is deliberately NOT
used: it is GNU/newer-BSD only and this script also runs directly on macOS hosts,
so the chase is a plain while [ -L ] loop handling relative link targets.

Verified on all three invocation shapes: via the /usr/local/bin symlink, via the
direct /opt path, and via a second-hop symlink in an unrelated directory. All
three now report the same 40 redactions.

LESSON worth keeping: fail-closed is correct for a secret scrubber, but it
converts "module not found" into an outage, so the module lookup becomes
load-bearing infrastructure and must be tested through the real invocation path,
not the convenient one.
2026-08-27 14:18:02 +02:00
joakimp 836e35b320 redact: name the tiers where they are used, and stop the docstring lying about tier 3
Review caught that the tier vocabulary was used in the report, the commit message
and the docs without being defined anywhere the reader would land, and inspecting
that turned up two real defects rather than just a wording gap.

1. STALE DOCSTRING. The module still described tier 3 as if it redacts, which
   stopped being true when the 403-hit measurement demoted it to report-only. It
   also credited tier 3 with resolving the 40-hex-PAT-vs-commit-sha collision —
   false by default, since a reporting rule resolves nothing. Corrected, with the
   consequence stated plainly: in the default configuration a sha-shaped PAT is
   caught if and only if it belongs to THIS machine, because only a known value
   (tier 1) or a naming key (tier 3, reporting) can separate it from a commit
   sha. That is an accepted gap; the alternative is redacting every sha in the
   palace.

2. THE VOCABULARY NEVER REACHED THE OUTPUT. The tool prints rule names
   (github-pat, env-value, url-credentials) and nothing printed a tier, so the
   docs' tier language was unconnected to what an operator actually sees. Added
   RULE_TIERS as the authoritative rule -> tier mapping, tier_of(), and
   Finding.tier; the feeder now prints "T2:github-pat=6" so what matched and how
   much to trust it are both visible on one line. A self-test asserts every rule
   that can appear in a Finding maps to a tier, so adding a rule without
   classifying it fails the tests instead of printing "T?".

Tiers, for the record, are three kinds of EVIDENCE (not three severities):
T1 known value from this process's env — near-certain, zero FP by construction;
T2 known vendor shape — strong, the prefix is meaningful; T3 key name says
secret — candidate only, measured FP-heavy, reported. T0 is reserved for
suspicions(), which is a measured NON-detection.

Docs gain worked one-line examples per tier and a "which tier fired?" section
showing real output. 46 self-test cases pass.
2026-08-27 13:45:08 +02:00
joakimp 3d47937d06 feed: scrub secrets before staging a transcript, name-anchored not entropy-anchored
A palace is mined from transcripts, and transcripts contain whatever the terminal
printed. Measured on this fleet: one leaked bearer token had reached 3 drawers,
13 feeder inbox files across all three devices, and 10 local files spanning 10
days — from an agent inspecting an env var while debugging. That frequency is the
premise: this is a pipeline problem, not a discipline problem.

WHERE. bin/mempalace_redact.py, called from mempalace-pi-session at the point the
staged transcript is written. That single hook covers both transports, because
local mode mines the staged file and remote mode rsyncs that same file
byte-for-byte. Scrubbing operates on the parsed objects rather than the
serialized text, so string values are rewritten while keys, ids and structure are
untouched — scanning raw JSONL instead invents keys like "tapiKey" out of the \t
escape preceding a field name (observed, not theorised).

WHY NOT ENTROPY. The obvious "redact long random-looking strings" is actively
destructive here: drawer ids, chunk ids, event ids, replica ids, HLCs and commit
SHAs are all high-entropy, are the majority of random-looking text in a palace,
and redacting them is silent and permanent. Detection is anchored on meaning
instead: tier 1 literal values from this process's env whose NAME says secret
(zero false positives by construction, and the only tier that can tell a 40-hex
Gitea PAT from a git commit sha); tier 2 vendor-prefixed shapes (ghp_, glpat-,
xox*-, sk-, AKIA, AIza, hf_, JWT, PEM blocks, URL credentials, Authorization
headers); tier 3 name=value assignments.

THE MEASUREMENT THAT CHANGED THE DESIGN. Tier 3 was going to redact. Against
52 MB of real fleet transcripts it produced 403 hits, and inspection with values
masked showed most were ${VAR} interpolation in compose files, TypeScript
identifiers, a TYPE ANNOTATION (credentials: Credentials), an IPA attribute
holding a date (krbPasswordExpiration), AAAK diary shorthand, and terminal output
following a "Password:" prompt. Redacting those corrupts code and docs held as
memory to catch what tier 1 already catches by value. After adding guards for
interpolation, code context, non-secret key suffixes and all-digit values, the
enforced count fell 403 -> 29 on the same corpus. So tier 3 REPORTS and does not
rewrite unless MEMPALACE_REDACT_STRICT=1.

HONESTY ABOUT MISSES. Known false negatives are documented rather than papered
over: novel formats in bare prose, another machine's secrets, base64-of-a-secret,
line-split secrets. Every run prints a count including "0 redaction(s)", because
silence is indistinguishable from a scrubber that never ran, and suspicions()
reports high-entropy strings it did NOT redact as (length, fingerprint) so the
miss rate is measurable. Findings never carry the value — rule, label, length and
sha256[:8], enough to recognise a recurrence, not enough to recover the secret.

FAIL CLOSED: no scrubber, no staging (exit 3), overridable with
MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 for a machine older than this file.

Tested: 42-case corpus in --self-test, including every palace id shape as a
must-not-redact case, idempotency, and a compound case the corpus caught where a
vendor placeholder was then re-matched by the URL rule (nested placeholder — the
secret was hidden either way, so only an exact-output assertion catches it).
End-to-end on 45 real sessions: 29 enforced, fail-closed verified at rc=3,
override verified loud. Server-side layer specified in docs/secret-hygiene.md §5
but NOT implemented — it is tier 2 only there, since the hub cannot see a
client's environment.
2026-08-27 13:30:47 +02:00
joakimp ecc2a9c574 mailbox notify: record that the terminal path is unverified through tmux
Measured on tor-ms22: the layering here is kitty -> tmux (on the host) ->
docker exec -> pi (in the container), with two clients attached to one tmux
session. Test OSC sequences written straight to pi's tty produced no
notification on the remote client; the local client is still unobserved. Most
likely tmux is dropping OSC types it does not implement — reaching the outer
terminal needs tmux's DCS passthrough (ESC P tmux; ... ESC \, inner ESC doubled)
plus allow-passthrough on, which is not the default and is not implemented here.

The point worth keeping is structural, not incidental: the container can see
NEITHER layer. KITTY_WINDOW_ID is absent because docker exec does not forward it,
and TMUX is absent because tmux runs one level further out on the host. So a
containerised client cannot detect the terminal it is speaking to or the
multiplexer it must speak through, and autodetection is not merely unreliable
here — it is blind. Explicit configuration is the only route, which is why the
previous commit added forced kitty/osc777 modes.

No behaviour or defaults changed: in-TUI notify remains the default and the
terminal path stays opt-in. Also flagged the unsettled routing question — a
pane's output reaches every attached client, so a work laptop and a home machine
would both ping from one arriving ask.
2026-08-27 12:54:52 +02:00
joakimp e91766286e mailbox notify: name the protocol, because docker exec hides the terminal
Auto-detection is useless in the deployment this ships in, measured on tor-ms22.
Terminal identity lives in env vars set by the emulator (KITTY_WINDOW_ID,
TERM_PROGRAM) and docker exec does not forward them — pi inside the container
sees only TERM=xterm-256color no matter what is rendering it. So the =desktop
detection can never see Kitty from in there and always falls through to OSC 777,
which Kitty does not implement. Result: on a containerised client the ping
silently does nothing, which is the worst available failure for a feature whose
entire purpose is to break a silence.

Adds MEMPALACE_MAILBOX_NOTIFY=kitty and =osc777 to force one protocol and skip
detection. =desktop keeps the notify.ts-style autodetect for the non-container
case, unset still means in-TUI only, 0/off still silent. The Windows toast branch
is dropped from the comment rather than the code path it never had here: it needs
powershell.exe on PATH, which is a WSL fact, not a Linux-container one.
2026-08-27 12:39:16 +02:00
joakimp a92c75d070 mailbox: say it is queued, and ping the human who is not looking (RFC 003 §7.11)
Two changes, neither touching the no-triggerTurn decision, which stands.

B — a delivery note in the message itself. The mid-session text explained how to
CLOSE an ask and never said when it would be SEEN, so the only reader who needed
that fact — the human watching an idle session — was the one not told. It now
says: this is a queued message, nothing woke the agent, any message starts the
turn that handles it, and the agent is not ignoring the ask, it is not running.
Costs nothing and changes no behaviour; it converts "why is it ignoring me?" into
"right, I nudge it". Deliberately NOT added to the wake-up injection, where a
turn is already starting and the note would be false.

A — a notification at poll time, because B only helps someone already looking and
the case that loses an ask is nobody looking. MEMPALACE_MAILBOX_NOTIFY: unset →
in-TUI ctx.ui.notify, the surface session_start already uses; =desktop →
additionally a terminal-native notification (Kitty OSC 99, else OSC 777), reusing
the detection the fleet's own notify.ts already proves in this harness; =0/off →
silent. The desktop path is how a ping escapes a container with no notify-send,
no DBus and no host access: the escape sequence is written to stdout and
interpreted by the terminal emulator on the human's own machine. Opt-in because
writing raw escapes is a behaviour change on a shared machine, not because it is
unreliable — say the word and the default flips.

Placement and firing conditions are deliberate: the notify call sits AFTER
sendMessage so a ping can never be the only thing that happened, and it fires
only when something is due — the same condition as delivery. A notification on
an empty poll would train its reader to ignore it, which is the failure this
whole feature exists to reverse. The wording names the nudge ("send any message
to handle") because "you have mail" without "press a key" reproduces exactly the
confusion B fixes.

Tested: 12 cases. Mode parsing (unset/empty/desktop/DESKTOP-with-space/0/off/1),
escape hygiene, and the OSC invariant that matters — a hostile title or body
containing ";" or a BEL cannot forge an OSC field or terminate the sequence
early, which is worth asserting because event bodies arrive from other machines.
ctx.hasUI is checked and the notify call is wrapped, since the UI can be gone by
the time an unawaited poll resolves. Syntax checked with node --strip-types.
2026-08-26 23:56:21 +02:00
joakimp 982b001001 docs: RFC 003 §7.11/§7.12 — delivered is not read, and the mailbox carries obligations not news
Both measured 2026-08-26 by peer devices, and the second one measured on me:
the operator had to point at an event id before I read a report that had been
addressed to me for two hours. The mailbox was working correctly the whole time.

§7.12 is the structural one. Candidates are drawn with status="open", so an
event carrying a TERMINAL status is not a mailbox candidate at all, whoever it
is addressed to. A task.reply written to a named device to share a finding is
therefore delivered to nobody, ever — and neither is any event.ack. So the most
natural inter-device message, "here is something you should know", is exactly
the shape that gets no delivery. Two things that do work: address it as a
directed status="open" ask (correct when a response is actually wanted), or
accept it as pull-only and pair it with a drawer, which the peer's search will
surface. A terminal report plus an expectation of attention does not.

§7.11 is the peer's finding, and it explains the other half of why that report
sat unread: delivery is deliverAs "steer" with deliberately no triggerTurn, and
the poll fires on agent_settled — i.e. when the agent is IDLE, with no inference
running. The text is queued for the next turn, so the human is the trigger.
Measured on another device: a delivery landed at ~20:50Z and sat visibly
unreacted-to until the operator asked "do I have to nudge you?". Not a defect;
the no-triggerTurn decision was deliberate and stands. But the delivered text
explains how to CLOSE an ask and never says when it will be SEEN, so the one
reader who needs that fact — the human watching the window — is the one not
told. Recorded with the three fixes that do not wake a model, as decision 8.

§7.4 is upgraded from inferred to measured, on two devices independently, by a
two-arm control: a to_agent='*' event with status='open' survives the raw
candidate filter and therefore reaches the exclusion branch, which no previously
written broadcast (all non-open) ever did. Raw query returned it, derived owed
set did not, delivery named only the directed arm. Two arms differing only in
to_agent is what makes it evidence rather than an absence.

And it carries a correction of my own record, in place and dated: I had filed
that this branch COULD NOT be exercised, because every broadcast in the log is
status='ready' and the protocol forbids writing an open one. Wrong in an
instructive way — the protocol forbids it as PRODUCTION TRAFFIC, which does not
forbid planting one as a labelled, self-closed control. Where a rule appears to
block a measurement, check whether it blocks the use or the test.
2026-08-26 23:38:41 +02:00
12 changed files with 2459 additions and 56 deletions
+1
View File
@@ -66,6 +66,7 @@ Start here if you are deciding **what to put where**, or running MemPalace on mo
|---|---|
| [`docs/fleet-memory.md`](docs/fleet-memory.md) | **Start here.** The five stores MemPalace keeps and what each is for; what a *central* palace buys a fleet of machines; when to file a memory versus when to message another device. Diagrams, worked examples. |
| [`docs/rfc-003-coordination-log.md`](docs/rfc-003-coordination-log.md) | The inter-agent / inter-device coordination log (`logstream`): storage, append and query semantics, delivery and latency, the trust model, and its landmines. |
| [`docs/secret-hygiene.md`](docs/secret-hygiene.md) | Keeping credentials out of the palace — what the feeder scrubs before staging, why detection is name-anchored and never entropy-anchored, and the measured false-positive rate that made one tier report-only |
| [`docs/rfc-001-global-palace.md`](docs/rfc-001-global-palace.md) | Why and how a palace is centralised, what should *not* be global, and the Phase-0 landmines (including: never run `mempalace sync` against a shared palace). |
| [`docs/rfc-002-joiner.md`](docs/rfc-002-joiner.md) | Replaying a second palace into a shared primary, and what dedupes what. |
| [`docs/phase-1-exposure-runbook.md`](docs/phase-1-exposure-runbook.md) | *(moved)* The exposure record for one deployment is now private operator data. The stub names the reusable mechanism it also carried — the `Host`/`Origin` pin above all — which is not yet published elsewhere. |
+102 -2
View File
@@ -534,11 +534,70 @@ fi
# Also classifies each export as NEW/ALREADY FILED (by source_file lookup)
# so --dry-run reports the real mine-set size. Classification is advisory;
# `mempalace mine --mode convos` is still the authoritative dedup.
# The redactor is a sibling module, imported by the heredoc below. Exported
# rather than passed as argv so the argv unpack stays stable.
#
# ${BASH_SOURCE[0]} reports the path the script was INVOKED as, and the image
# installs /usr/local/bin/mempalace-pi-session as a symlink into
# /opt/mempalace-toolkit/bin. A naive dirname therefore yields /usr/local/bin,
# where the redactor does not exist — and because the import is fail-closed, that
# turns every feeder tick on every device into "refusing to stage". Measured: the
# symlinked invocation exited FATAL while the direct one worked, i.e. it would
# have stopped the whole fleet's memory feed at the next image bake. So chase the
# symlink chain, and offer fallbacks rather than betting on one answer.
#
# readlink -f is avoided deliberately: it is GNU/newer-BSD only, and this script
# also runs directly on macOS hosts.
_mp_resolve() {
local p="$1" target
while [ -L "$p" ]; do
target="$(readlink "$p")" || break
case "$target" in
/*) p="$target" ;;
*) p="$(dirname "$p")/$target" ;;
esac
done
printf '%s' "$p"
}
_mp_self="$(_mp_resolve "${BASH_SOURCE[0]}")"
MEMPALACE_REDACT_DIR="$(cd "$(dirname "$_mp_self")" && pwd):$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd):/opt/mempalace-toolkit/bin"
export MEMPALACE_REDACT_DIR
export_count=$(python3 - "$PI_SESSIONS_DIR" "$STAGE" "$SESSION_ID" "$SINCE" "$MIN_MESSAGES" "$MIN_ASSISTANT_CHARS" "$MODE" <<'PY'
import json, os, sqlite3, sys
from datetime import datetime, timezone
from pathlib import Path
# ── Secret scrubbing before anything is staged ───────────────────────────────
# This is the last point at which a transcript is a plain in-memory value: the
# local mine reads the staged file, and the remote path rsyncs that same file
# byte-for-byte, so scrubbing here covers BOTH transports with one hook.
#
# FAIL CLOSED. If the redactor cannot be imported, staging is refused rather
# than done unscrubbed — a missing module means a broken install, and the whole
# point of this step is that a secret must not reach a shared palace. Override
# deliberately with MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 if you ever need to feed a
# machine whose toolkit is older than this file.
# Colon-separated candidates: symlink-resolved dir, invoked dir, then the image's
# known install path. First one that has the module wins.
for _cand in os.environ.get("MEMPALACE_REDACT_DIR", "").split(":"):
if _cand and os.path.isdir(_cand):
sys.path.insert(0, _cand)
try:
import mempalace_redact as _redact
except Exception as _e: # noqa: BLE001 - any import failure is fatal by design
if os.environ.get("MEMPALACE_FEED_ALLOW_UNSCRUBBED", "").strip() in {"1", "true", "yes"}:
_redact = None
print(" [WARN] secret scrubber unavailable, staging UNSCRUBBED by request "
f"({_e})", file=sys.stderr)
else:
print(f" [FATAL] secret scrubber unavailable ({_e}); refusing to stage. "
"Set MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 to override.", file=sys.stderr)
raise SystemExit(3)
_known = _redact.env_secrets() if _redact else []
_redactions = 0
_reported = 0
sessions_dir, stage, session_filter, since, min_messages, min_assistant_chars, mode = sys.argv[1:8]
min_messages = int(min_messages)
min_assistant_chars = int(min_assistant_chars)
@@ -811,6 +870,29 @@ for path in paths:
)
continue
# Scrub the parsed objects, not the serialized text: string VALUES get
# rewritten while keys, ids and structure are left exactly as they are.
# (Scanning raw JSONL instead would also match escape artifacts like the
# "\\t" before a field name, inventing keys such as "tapiKey".)
if _redact is not None:
_found: list = []
out_lines = [_redact.scrub_obj(obj, _known, _found)[0] for obj in out_lines]
_hard = [x for x in _found if not x.rule.endswith("-reported")]
_soft = [x for x in _found if x.rule.endswith("-reported")]
_redactions += len(_hard)
_reported += len(_soft)
if _hard:
_by = {}
for x in _hard:
# "T2:github-pat" — rule says what matched, tier says how much to
# trust it, which is what the reader of this line actually needs.
_k = f"T{x.tier}:{x.rule}"
_by[_k] = _by.get(_k, 0) + 1
print(f" [REDACTED] {path.name} "
+ ", ".join(f"{k}={v}" for k, v in sorted(_by.items()))
+ " fp=" + ",".join(sorted({x.fingerprint for x in _hard})),
file=sys.stderr)
out_path = stage / f"pi_{session_uuid}.jsonl"
with out_path.open("w", encoding="utf-8") as f:
for obj in out_lines:
@@ -835,6 +917,13 @@ for path in paths:
print(f"EXPORTED {exported}")
print(f"ALREADY_FILED {-1 if already_filed is None else skipped_already_filed}")
# Report the scrub outcome even when it is zero: "0 redactions" is a measurement,
# whereas printing nothing is indistinguishable from a scrubber that never ran.
if _redact is not None:
print(f" [scrub] {_redactions} redaction(s) applied, "
f"{_reported} name-anchored candidate(s) reported only "
f"(set MEMPALACE_REDACT_STRICT=1 to redact those too)", file=sys.stderr)
if skipped_short:
print(f"SKIPPED_SHORT {skipped_short}", file=sys.stderr)
if skipped_quiet:
@@ -885,15 +974,26 @@ fi
# ── Ship to the palace host (remote mode only) ───────────────────────
# mempalace_mine expands its source path in the SERVER process, so in remote
# mode the exports have to physically exist over there. rsync --update is the
# mode the exports have to physically exist over there. --checksum is the
# idempotent half; the mine is the other half.
#
# NOT --update: the stage file's mtime is deliberately the SOURCE transcript's
# mtime (see the os.utime() in the exporter, "preserve session mtime for dedup
# stability"), so a re-export of a session that has not been appended to since
# the last ship carries an mtime that is NOT newer than the receiver copy. With
# --update rsync then SKIPS it silently -- which is exactly the case that must
# ship after a redactor change, because the content differs while the mtime does
# not. Measured on mbp-m1-2020 2026-08-27: a scrubbed re-export of a dormant
# session was skipped and unscrubbed bytes stayed in the palace host's inbox,
# while every local signal reported a clean stage. --checksum compares content
# and keeps the ship idempotent without trusting timestamps.
MINE_SOURCE="$STAGE"
if [[ "$MODE" == "remote" ]]; then
ssh_cmd="ssh"
[[ -n "$SSH_CONFIG" ]] && ssh_cmd="ssh -F $SSH_CONFIG"
echo ""
echo "Shipping stage to ${SSH_TARGET%/}/$DEVICE/ ..."
if ! rsync -a --update --no-owner --no-group \
if ! rsync -a --checksum --no-owner --no-group \
-e "$ssh_cmd" \
--include='*.jsonl' --exclude='*' \
"$STAGE/" "${SSH_TARGET%/}/$DEVICE/"; then
+604
View File
@@ -0,0 +1,604 @@
#!/usr/bin/env python3
"""Scrub secret-shaped values out of palace-bound text.
WHY THIS IS NAME-ANCHORED AND NOT ENTROPY-ANCHORED
==================================================
The obvious design — "redact long random-looking strings" — is actively wrong
for this corpus, and the reason is worth stating before anyone tries to
"improve" it. A palace is *full* of high-entropy strings that are its own
primary keys:
drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b drawer id
..._chunk_000000 chunk id
evt_20260826T213924_397b1f710c2d event id
rep_d344e349ba276d6fc11997cd552f6937 replica id
1787780364792-000000-rep_d344e349ba276d6f… hybrid logical clock
ecc2a9c574f4156403cf4e90aea8b0d088c75700 git commit sha
An entropy or "looks like base64/hex of length N" detector fires on every one
of those, i.e. on most identifiers in the corpus, and the redaction damages the
memory it was meant to protect. Worse, the damage is silent and unrecoverable:
a mined transcript whose ids have been replaced with <redacted> is no longer
traceable to anything.
So detection here is anchored on one of three things that carry *meaning*:
Tier 1 KNOWN VALUES — literal values taken from this process's own
environment, for variables whose NAME says secret.
Zero false positives by construction: the value
*is* the secret. Catches any presentation (env
dump, JSON, error message, URL, prose), which is
exactly the incident this exists for.
Tier 2 KNOWN SHAPES — vendor-prefixed credentials (ghp_…, glpat-…,
xox[abprs]-…, AKIA…, sk-ant-…, JWTs, PEM private
keys, credentials embedded in URLs, Authorization
headers). Very low false-positive rate because the
prefix is meaningful, not merely random.
Tier 3 NAME=VALUE — an assignment whose *key* says secret
(…TOKEN, …SECRET, …PASSWORD, …API_KEY, …).
REPORT-ONLY BY DEFAULT. Measured on 52 MB of real
fleet transcripts it fired 403 times, of which the
large majority were ${VAR} interpolation in compose
files, TypeScript identifiers, a type annotation
(`credentials: Credentials`), an IPA attribute
holding a date (krbPasswordExpiration) and terminal
output following a "Password:" prompt. Rewriting
those corrupts code and docs held as memory, to
catch what Tier 1 already catches by value.
Set MEMPALACE_REDACT_STRICT=1 to make it enforce.
SHAPE COLLISIONS, AND WHY TIER 1 IS THE LOAD-BEARING ONE
--------------------------------------------------------
A 40-hex Gitea PAT is byte-indistinguishable from a git commit sha, so no shape
rule can separate them. Only two things can: the value being known (Tier 1,
which redacts) or a key naming it (Tier 3, which by default only reports). So
in the default configuration a sha-shaped PAT is caught if and only if it
belongs to this machine. That is an accepted, documented gap — the alternative
is redacting every commit sha in the palace.
WHICH TIER DID THAT? — mapping printed rule names back to tiers
--------------------------------------------------------------
The output names the *rule* (`github-pat`, `env-value`), because that says what
matched. RULE_TIERS below maps rule -> tier, and `tier_of()` is what the feeder
uses to print `T2:github-pat=6`, because the tier is what tells a reader how
much to trust the hit. A self-test asserts every rule that can appear in a
Finding has a tier, so adding a rule without classifying it fails the tests.
KNOWN FALSE NEGATIVES — stated, not hidden
------------------------------------------
This will not catch: a novel credential format pasted bare into prose with no
name nearby; a secret from a machine whose env this process cannot see; a
base64-of-a-secret; a secret split across lines. Tier 1 covers the local
machine's own secrets completely, which is where the measured incidents came
from — but "the scrubber ran" must never be read as "there are no secrets in
here". `suspicions()` exists for exactly that reason: it reports high-entropy
candidates it did NOT redact, as fingerprints rather than values, so the false
negative rate can be *measured* over time instead of assumed to be zero.
Run `python3 mempalace_redact.py --self-test` for the test corpus, which
includes every benign id shape above as a must-not-redact case.
"""
from __future__ import annotations
import hashlib
import math
import os
import re
from typing import Any, Callable, Iterable
__all__ = ["scrub", "scrub_obj", "suspicions", "Finding"]
PLACEHOLDER = "<redacted:{label}>"
# ── Tier 1: names whose VALUES are secrets ────────────────────────────────────
# Anchored on the variable name, then applied as an exact literal match.
_SECRET_NAME = re.compile(
r"(TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|API_?KEY|APIKEY|ACCESS_?KEY"
r"|PRIVATE_?KEY|CREDENTIAL|BEARER|SESSION_?KEY|COOKIE|_PAT|^PAT$)",
re.IGNORECASE,
)
# …unless the name says it holds a location or a knob rather than a secret.
# SSH_KEY_PATH is the live example: it matches KEY, and its value is a path.
_NAME_NOT_SECRET = re.compile(
r"(_PATH|_FILE|_DIR|_NAME|_ID|_URL|_URI|_HOST|_PORT|_USER|_ENABLED"
r"|_TIMEOUT|_MS|_SECS|_SECONDS|_COUNT|_LIMIT|_MODE)$",
re.IGNORECASE,
)
# Keys that contain a secret word but describe a POLICY, COUNT or TYPE rather
# than holding a credential. Measured on real transcripts: krbPasswordExpiration
# (an IPA attribute whose value is a date), observationTokens (a token count),
# targetTokens, tokensOverTarget.
_KEY_NOT_A_SECRET = re.compile(
r"(expiration|expiry|_age|count|length|limit|usage|interval|policy|type|class"
r"|provider|manager|error|target|overtarget|file|path|dir|name|_id|url|uri"
r"|host|port|user|enabled|timeout|_ms|_secs)$",
re.IGNORECASE,
)
# Left of the match: a declaration keyword or a member access means this is CODE
# (`const token = x`, `obj.apiKey = y`), and the "value" is an expression.
_CODE_CONTEXT = re.compile(
r"(?:\b(?:const|let|var|def|func|fn|public|private|readonly|export|return|if|assert)\s+$)"
r"|[.\w]\s*$"
)
_TRIVIAL_VALUES = {
"", "0", "1", "true", "false", "none", "null", "nil", "unset", "changeme",
"change-me", "xxx", "***", "redacted", "your-token-here", "your_token_here",
"placeholder", "example", "dummy", "test", "secret",
}
# ── Tier 2: shapes that mean credential because the PREFIX means credential ──
_SHAPE_RULES: list[tuple[str, re.Pattern[str], Callable[[re.Match[str]], str]]] = [
# ORDER MATTERS, and the self-test is what proved it. The two structural
# rules (credentials inside a URL, Authorization header) run FIRST, before
# the vendor-prefix rules: a vendor rule rewriting `https://ghp_xxx@host` to
# `https://<redacted:github-token>@host` leaves a colon INSIDE the
# placeholder, which the URL rule then reads as user:password and redacts a
# second time, producing a nested placeholder. Ordering plus the (?!<redacted)
# guards below make the result independent of pass order.
("url-credentials",
re.compile(r"(?P<pre>[a-zA-Z][a-zA-Z0-9+.\-]*://(?!<redacted)[^/\s:@]+:)(?P<pw>(?!<redacted)[^/\s@]{4,})(?P<at>@)"),
lambda m: f"{m.group('pre')}{PLACEHOLDER.format(label='url-password')}{m.group('at')}"),
("authorization-header",
re.compile(r"(?P<pre>[Aa]uthorization:\s*(?:Bearer|Basic|token)\s+)(?P<val>(?!<redacted)[A-Za-z0-9._\-+/=]{8,})"),
lambda m: f"{m.group('pre')}{PLACEHOLDER.format(label='auth-header')}"),
("github-pat", re.compile(r"\bgh[pousr]_[A-Za-z0-9]{16,}\b"),
lambda m: PLACEHOLDER.format(label="github-token")),
("github-fine-grained", re.compile(r"\bgithub_pat_[A-Za-z0-9_]{22,}\b"),
lambda m: PLACEHOLDER.format(label="github-token")),
("gitlab-pat", re.compile(r"\bglpat-[A-Za-z0-9_\-]{16,}\b"),
lambda m: PLACEHOLDER.format(label="gitlab-token")),
("slack", re.compile(r"\bxox[abprs]-[A-Za-z0-9\-]{10,}\b"),
lambda m: PLACEHOLDER.format(label="slack-token")),
("openai-anthropic", re.compile(r"\bsk-(?:ant-)?[A-Za-z0-9_\-]{20,}\b"),
lambda m: PLACEHOLDER.format(label="api-key")),
("aws-access-key", re.compile(r"\b(?:AKIA|ASIA)[0-9A-Z]{16}\b"),
lambda m: PLACEHOLDER.format(label="aws-access-key-id")),
("google-api-key", re.compile(r"\bAIza[0-9A-Za-z_\-]{35}\b"),
lambda m: PLACEHOLDER.format(label="google-api-key")),
("huggingface", re.compile(r"\bhf_[A-Za-z0-9]{30,}\b"),
lambda m: PLACEHOLDER.format(label="hf-token")),
("npm", re.compile(r"\bnpm_[A-Za-z0-9]{30,}\b"),
lambda m: PLACEHOLDER.format(label="npm-token")),
("digitalocean", re.compile(r"\bdop_v1_[a-f0-9]{60,}\b"),
lambda m: PLACEHOLDER.format(label="do-token")),
# JWT: three base64url segments. Requires the "eyJ" header so a random
# dotted string cannot trip it.
("jwt", re.compile(r"\beyJ[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}\b"),
lambda m: PLACEHOLDER.format(label="jwt")),
# Whole PEM block, non-greedy so several in one file stay separate.
("pem-private-key",
re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----.*?-----END [A-Z ]*PRIVATE KEY-----", re.DOTALL),
lambda m: PLACEHOLDER.format(label="private-key-block")),
]
# ── Tier 3: assignment whose KEY says secret ─────────────────────────────────
# The key prefix is LAZY AND OPTIONAL, not "one char then anything": a mandatory
# leading character cannot match a key that BEGINS with the secret word, so
# `"api_key"` was missed while `"client_secret"` passed — the secret word sits
# mid-name in one and at position 0 in the other. Note plain `KEY` is absent from
# the alternation on purpose: it is too generic (SSH_KEY_PATH), and tier 1 covers
# the real ones by value anyway.
# The optional quote after the key is what makes JSON work: `"api_key": "v"`
# closes the key BEFORE the colon, so a pattern jumping straight from key to
# separator silently misses every JSON payload — i.e. most of a transcript.
_ASSIGNMENT = re.compile(
r"""(?P<key>\b[A-Za-z0-9_.\-]*?
(?:token|secret|password|passwd|passphrase|api[_-]?key|apikey
|access[_-]?key|private[_-]?key|credential)
[A-Za-z0-9_.\-]*)
(?P<kq>["']?)
(?P<sep>\s*(?:=>|[=:])\s*)
(?P<q>["']?)
(?P<val>[^\s"'<>,;)\]}]{12,})
(?P=q)""",
re.IGNORECASE | re.VERBOSE,
)
# A value that is obviously not a live secret: a reference, a path, a
# placeholder, or something we already redacted.
_NOT_A_SECRET_VALUE = re.compile(
r"""^(?:
\$\{[^}\s]*\}? # ${VAR} ${VAR:-default} ${VAR:?err}
# trailing } optional: the value
# charset stops before it, so the
# captured text is "${VAR" — which
# still means "a reference", and
# missing that is what made a
# compose file look like a secret
| \$[A-Za-z_][A-Za-z0-9_]* # $VAR
| %[A-Za-z_][A-Za-z0-9_]*%? # %VAR%
| \{\{[^}\s]*\}?\}? # {{ template }}
| <[^>]*> # <redacted:…> / <your-token>
| (?:/|~/|\./|\.\./)[^\s]* # a path
| [A-Za-z]:\\ # a windows path
| \*+ | x+ | \.+ # ***, xxx
| (?:true|false|none|null|nil|unset|changeme|placeholder|example)
)$""",
re.IGNORECASE | re.VERBOSE,
)
class Finding:
"""One redaction, recorded without ever carrying the secret itself."""
__slots__ = ("rule", "label", "length", "fingerprint")
def __init__(self, rule: str, label: str, value: str) -> None:
self.rule = rule
self.label = label
self.length = len(value)
# Stable across sessions, so a recurring leak is recognisable as the
# same one, while the value stays unrecoverable from the report.
self.fingerprint = hashlib.sha256(value.encode("utf-8")).hexdigest()[:8]
@property
def tier(self) -> int:
"""Which detection tier produced this. See RULE_TIERS for the mapping."""
return tier_of(self.rule)
def __repr__(self) -> str: # pragma: no cover - diagnostics only
return (f"<Finding T{self.tier}:{self.rule}:{self.label} "
f"len={self.length} fp={self.fingerprint}>")
def env_secrets(environ: dict[str, str] | None = None) -> list[tuple[str, str]]:
"""Tier 1 material: (name, value) pairs from the environment worth scrubbing.
Deliberately conservative: a secret-sounding NAME is not enough, the value
must also be long enough to be a credential and must not look like a path
or a knob. Longest values first so that if one secret contains another as a
substring, the longer one is replaced before the shorter can fragment it.
"""
env = os.environ if environ is None else environ
out: list[tuple[str, str]] = []
for name, value in env.items():
if not value or len(value) < 12:
continue
if not _SECRET_NAME.search(name) or _NAME_NOT_SECRET.search(name):
continue
if value.strip().lower() in _TRIVIAL_VALUES:
continue
if _NOT_A_SECRET_VALUE.match(value.strip()):
continue
if "/" in value and (value.startswith("/") or value.startswith("~")):
continue # a path that happened to be named …_KEY
if any(c.isspace() for c in value):
continue # credentials do not contain whitespace; prose does
out.append((name, value))
out.sort(key=lambda kv: len(kv[1]), reverse=True)
return out
# Rule name -> tier. Printed output names the rule (WHAT matched); the tier says
# HOW MUCH TO TRUST IT: 1 = near-certain (the value is known to be a secret),
# 2 = strong (a vendor prefix means what it says), 3 = candidate, needs eyes.
RULE_TIERS: dict[str, int] = {
"env-value": 1,
**{name: 2 for name, _pat, _lbl in _SHAPE_RULES},
"named-assignment": 3,
"named-assignment-reported": 3,
# Tier 0 is not a detection: it is a measured NON-detection, reported so the
# false-negative rate is a number instead of an assumption.
"suspicion": 0,
}
def tier_of(rule: str) -> int:
"""Tier for a rule name, or -1 if the rule was never classified."""
return RULE_TIERS.get(rule, -1)
def scrub(
text: str,
known: Iterable[tuple[str, str]] | None = None,
findings: list[Finding] | None = None,
strict: bool | None = None,
) -> tuple[str, list[Finding]]:
"""Redact secrets from one string. Returns (scrubbed, findings).
Idempotent: the replacement text matches none of the rules, so running this
twice changes nothing and cannot nest placeholders.
"""
found: list[Finding] = findings if findings is not None else []
if strict is None:
strict = os.environ.get("MEMPALACE_REDACT_STRICT", "").strip() in {"1", "true", "yes"}
if not text:
return text, found
# Tier 1 first: an exact known value should be labelled with its variable
# name (the most useful thing a reader can be told) rather than by shape.
for name, value in (env_secrets() if known is None else known):
if value and value in text:
count = text.count(value)
text = text.replace(value, PLACEHOLDER.format(label=name))
for _ in range(count):
found.append(Finding("env-value", name, value))
# Tier 2.
for rule, pattern, repl in _SHAPE_RULES:
def _sub(m: re.Match[str], _rule: str = rule) -> str:
found.append(Finding(_rule, _rule, m.group(0)))
return repl(m)
text = pattern.sub(_sub, text)
# Tier 3 — REPORT-ONLY unless strict. See STRICT note in the module docstring:
# measured on 52 MB of real fleet transcripts this rule produced 403 hits of
# which the overwhelming majority were ${VAR} interpolation, TypeScript
# identifiers, YAML env lists and documentation prose. Redacting those would
# corrupt code and docs stored as memory, to catch secrets that tier 1
# already catches by value. So by default a tier-3 hit is REPORTED (so the
# miss can be measured and the rule calibrated) and NOT rewritten.
def _sub_assignment(m: re.Match[str]) -> str:
val, key = m.group("val"), m.group("key")
pre = m.string[max(0, m.start() - 24):m.start()]
if (_NOT_A_SECRET_VALUE.match(val)
or val.strip().lower() in _TRIVIAL_VALUES
or _KEY_NOT_A_SECRET.search(key)
or _CODE_CONTEXT.search(pre)
or val.isdigit()):
return m.group(0)
found.append(Finding("named-assignment" if strict else "named-assignment-reported",
key, val))
if not strict:
return m.group(0)
return f"{m.group('key')}{m.group('kq')}{m.group('sep')}{m.group('q')}" \
f"{PLACEHOLDER.format(label=m.group('key'))}{m.group('q')}"
text = _ASSIGNMENT.sub(_sub_assignment, text)
return text, found
def scrub_obj(obj: Any, known: Iterable[tuple[str, str]] | None = None,
findings: list[Finding] | None = None,
strict: bool | None = None) -> tuple[Any, list[Finding]]:
"""Recursively scrub every string VALUE in a JSON-ish structure.
Keys are left alone: a key named "token" is metadata, not a credential, and
rewriting keys would corrupt the transcript schema.
"""
found: list[Finding] = findings if findings is not None else []
known = list(env_secrets() if known is None else known)
if isinstance(obj, str):
return scrub(obj, known, found, strict)[0], found
if isinstance(obj, dict):
return {k: scrub_obj(v, known, found, strict)[0] for k, v in obj.items()}, found
if isinstance(obj, list):
return [scrub_obj(v, known, found, strict)[0] for v in obj], found
return obj, found
# ── Measuring what we MISS, without leaking it ───────────────────────────────
_BENIGN_ID = re.compile(
r"""^(?:
drawer_[\w.\-]+ # drawer / chunk ids
| (?:evt|art)_\d{8}T\d{6}_[0-9a-f]{12} # event / artifact ids
| rep_(?:[0-9a-f]{12}|[0-9a-f]{32}) # replica ids
| \d{13}-[0-9a-f]{6}-rep_[0-9a-f]+ # hybrid logical clocks
| [0-9a-f]{7,12} # short git sha
| [0-9a-f]{40} # git sha / sha1
| [0-9a-f]{64} # sha256
| [0-9a-f]{8}(?:-[0-9a-f]{4}){3}-[0-9a-f]{12} # uuid
| sha256:[0-9a-f]{64}
| \d+
)$""",
re.VERBOSE | re.IGNORECASE,
)
_CANDIDATE = re.compile(r"[A-Za-z0-9+/_\-=]{24,}")
def _entropy(s: str) -> float:
if not s:
return 0.0
counts: dict[str, int] = {}
for ch in s:
counts[ch] = counts.get(ch, 0) + 1
n = len(s)
return -sum((c / n) * math.log2(c / n) for c in counts.values())
def suspicions(text: str, min_entropy: float = 3.6) -> list[Finding]:
"""High-entropy strings this module did NOT redact, as fingerprints.
This is the false-negative meter. It deliberately does not redact anything:
every one of these is far more likely to be one of the palace's own ids than
a credential, and redacting them would be the damage described at the top of
this file. Reporting them as (length, fingerprint) makes the miss rate
observable — and a fingerprint that keeps recurring is worth a human look.
"""
out: list[Finding] = []
for m in _CANDIDATE.finditer(text):
val = m.group(0)
if _BENIGN_ID.match(val) or val.startswith("<redacted:"):
continue
if _entropy(val) < min_entropy:
continue
out.append(Finding("suspicion", "high-entropy", val))
return out
# ── Self-test ────────────────────────────────────────────────────────────────
def _self_test() -> int:
fake_token = "pFrDBfakVKQB0SLN1hTzAoaRJeQ4J_go7xmJezojAwI" # shape-alike, not real
env = {
"MEMPALACE_REMOTE_TOKEN": fake_token,
"GITEA_TOKEN": "0123456789abcdef0123456789abcdef01234567", # 40 hex: sha-shaped
"SSH_KEY_PATH": "/Users/someone/.ssh", # name matches KEY, value is a path
"MEMPALACE_MAILBOX_POLL_MS": "300000", # knob, not a secret
"MEMPALACE_REMOTE_URL": "https://palace.example.com/mcp",
"PASSWORD": "changeme", # trivial value
"API_KEY": "${FROM_VAULT}", # a reference
}
known = env_secrets(env)
must_redact = [
("env dump line", f"MEMPALACE_REMOTE_TOKEN={fake_token}"),
("bare value in prose", f"the token is {fake_token} apparently"),
("value inside json", f'{{"env": {{"MEMPALACE_REMOTE_TOKEN": "{fake_token}"}}}}'),
("sha-shaped pat, name-anchored",
"GITEA_TOKEN=0123456789abcdef0123456789abcdef01234567"),
("github pat", "remote add origin https://ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345@git"),
("gitlab pat", "GLPAT is glpat-AbCdEfGhIjKlMnOpQrSt"),
("slack", "xoxb-1234567890-ABCdefGHIjklMNO"),
("anthropic-ish", "sk-ant-api03-AbCdEfGhIjKlMnOpQrStUvWx"),
("aws", "AKIAIOSFODNN7EXAMPLE"),
("jwt", "eyJhbGciOiJIUzI1NiJ9.eyJzdWIiOiIxMjM0NTY3ODkwIn0.dozjgNryP4J3jVmNHl0w5N_XgL0n3I9PlFUP0THsR8U"),
("pem", "-----BEGIN OPENSSH PRIVATE KEY-----\nb3BlbnNzaA\n-----END OPENSSH PRIVATE KEY-----"),
("url creds", "git clone https://joakim:hunter2hunter2@git.example.com/x.git"),
("auth header", "Authorization: Bearer abcdefghijklmnop"),
("json api_key", '"api_key": "AbCdEfGhIjKlMnOpQrSt"'),
("json nested quotes", '{"auth": {"client_secret": "AbCdEfGhIjKlMnOpQrStUv"}}'),
("yaml style", " gitea_token: 0123456789abcdefghij"),
("lowercase password kv", "db_password = s3cr3t-p4ssw0rd-xyz"),
]
must_not_redact = [
("drawer id", "drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b"),
("chunk id", "drawer_pi-devbox_landmines_ca8929fea82cf80e7770ea0d_chunk_000000"),
("event id", "evt_20260826T213924_397b1f710c2d"),
("replica id", "rep_d344e349ba276d6fc11997cd552f6937"),
("hlc", "1787780364792-000000-rep_d344e349ba276d6fc11997cd552f6937"),
("git sha", "ecc2a9c574f4156403cf4e90aea8b0d088c75700"),
("short sha", "ecc2a9c"),
("sha256", "a" * 64),
("uuid session file", "pi_01a03fd8-0dc6-73c2-a9c3-6b62ccf0318f.jsonl"),
("knob assignment", "MEMPALACE_MAILBOX_POLL_MS=300000"),
("path named key", "SSH_KEY_PATH=/Users/someone/.ssh"),
("empty token", "MEMPALACE_REMOTE_TOKEN="),
("already redacted", "MEMPALACE_REMOTE_TOKEN=<redacted:MEMPALACE_REMOTE_TOKEN>"),
("var reference", "MEMPALACE_REMOTE_TOKEN=${VAULT_TOKEN}"),
("placeholder", "MEMPALACE_REMOTE_TOKEN=<your-token-here>"),
("trivial", "PASSWORD=changeme"),
("url no creds", "MEMPALACE_REMOTE_URL=https://palace.example.com/mcp"),
("prose", "The mailbox is an obligation channel, not a news channel."),
("base64 in a diff line", "+ data = 'QUJDREVGR0hJSktMTU5PUFFSU1RVVldYWVo='"),
]
failures = 0
# Tier 3 is report-only by default, so the corpus below is checked in STRICT
# mode (where tier 3 rewrites) and the default mode is asserted separately.
print("MUST REDACT (strict: tier 3 enforcing)")
for name, sample in must_redact:
out, found = scrub(sample, known, strict=True)
ok = bool(found) and all(
secret not in out for secret in (fake_token, "hunter2hunter2",
"0123456789abcdef0123456789abcdef01234567")
)
failures += not ok
rules = ",".join(sorted({f.rule for f in found})) or "NONE"
print(f" {'pass' if ok else 'FAIL'} {name:32s} [{rules}]")
print("MUST NOT REDACT (strict: the hard cases)")
for name, sample in must_not_redact:
out, found = scrub(sample, known, strict=True)
ok = out == sample and not found
failures += not ok
detail = "" if ok else f" -> {out!r} {found}"
print(f" {'pass' if ok else 'FAIL'} {name:32s}{detail}")
print("DEFAULT MODE (tier 3 reports, does not rewrite)")
for name, sample, expect_rewrite in [
("env value still redacted", f"MEMPALACE_REMOTE_TOKEN={fake_token}", True),
("vendor shape still redacted", "ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345", True),
("name-anchored only reported", "db_password = s3cr3t-p4ssw0rd-xyz", False),
("compose interpolation untouched", "GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}", False),
("typescript identifier untouched", "const tokens = countTokensForModel", False),
("ipa attribute untouched", "--setattr=krbPasswordExpiration=20260529090000", False),
]:
out, found = scrub(sample, known)
rewritten = out != sample
ok = rewritten == expect_rewrite
if name == "name-anchored only reported":
ok = ok and any(f.rule == "named-assignment-reported" for f in found)
if name.endswith("untouched"):
ok = ok and not any(f.rule.startswith("named-assignment") for f in found)
failures += not ok
print(f" {'pass' if ok else 'FAIL'} {name:34s}"
+ ("" if ok else f" -> rewritten={rewritten} {found}"))
print("TIER MAPPING")
_rules = ({"env-value", "named-assignment", "named-assignment-reported", "suspicion"}
| {n for n, _p, _l in _SHAPE_RULES})
_unclassified = sorted(r for r in _rules if tier_of(r) < 0)
failures += bool(_unclassified)
print(f" {'pass' if not _unclassified else 'FAIL'} every rule maps to a tier"
+ (f" -> unclassified: {_unclassified}" if _unclassified
else f" ({len(_rules)} rules)"))
for label, sample, want in [
("env value is tier 1", f"MEMPALACE_REMOTE_TOKEN={fake_token}", 1),
("vendor shape is tier 2", "ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345", 2),
("name-anchored is tier 3", "db_password = s3cr3t-p4ssw0rd-xyz", 3),
]:
got = scrub(sample, known)[1]
ok = bool(got) and all(f.tier == want for f in got)
failures += not ok
print(f" {'pass' if ok else 'FAIL'} {label:26s}"
+ ("" if ok else f" -> {[(f.rule, f.tier) for f in got]}"))
print("PROPERTIES")
once, _ = scrub(f"MEMPALACE_REMOTE_TOKEN={fake_token}", known, strict=True)
twice, second = scrub(once, known, strict=True)
ok = once == twice and not second
failures += not ok
print(f" {'pass' if ok else 'FAIL'} idempotent (no nested placeholders)")
# The compound case the corpus caught: assert on EXACT output, not merely
# that the secret is gone — a nested placeholder also hides the secret, and
# would have shipped unnoticed.
compound, _ = scrub("https://ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345@git.example.com/x", known, strict=True)
ok = compound == "https://<redacted:github-token>@git.example.com/x"
failures += not ok
print(f" {'pass' if ok else 'FAIL'} vendor placeholder not re-redacted by url rule"
+ ("" if ok else f" -> {compound!r}"))
already, found_a = scrub("https://joakim:<redacted:url-password>@git.example.com", known, strict=True)
ok = already == "https://joakim:<redacted:url-password>@git.example.com" and not found_a
failures += not ok
print(f" {'pass' if ok else 'FAIL'} already-redacted url left alone")
obj = {"role": "user", "content": [{"type": "text", "text": f"export TOK={fake_token}"}],
"id": "evt_20260826T213924_397b1f710c2d"}
scrubbed, found = scrub_obj(obj, known, strict=True)
ok = (fake_token not in repr(scrubbed)
and scrubbed["id"] == obj["id"]
and scrubbed["role"] == "user"
and len(found) >= 1)
failures += not ok
print(f" {'pass' if ok else 'FAIL'} nested structure scrubbed, ids and keys preserved")
ok = all(fake_token not in repr(f) for f in scrub(f"x {fake_token}", known, strict=True)[1])
failures += not ok
print(f" {'pass' if ok else 'FAIL'} findings never carry the secret value")
sus = suspicions("drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b "
"rep_d344e349ba276d6fc11997cd552f6937 "
"1787780364792-000000-rep_d344e349ba276d6fc11997cd552f6937 "
"ecc2a9c574f4156403cf4e90aea8b0d088c75700")
ok = not sus
failures += not ok
print(f" {'pass' if ok else 'FAIL'} suspicion meter ignores our own id shapes ({len(sus)} hits)")
sus2 = suspicions("blob=Zm9vYmFyYmF6cXV1eHF1dXhmb29iYXJiYXpxdXV4Zm9vYmFy")
ok = len(sus2) == 1 and all("Zm9v" not in repr(s) for s in sus2)
failures += not ok
print(f" {'pass' if ok else 'FAIL'} suspicion meter flags unknown blobs as fingerprints only")
print(f"\n{'ALL PASS' if not failures else f'{failures} FAILED'}")
return 1 if failures else 0
if __name__ == "__main__":
import sys
if "--self-test" in sys.argv:
raise SystemExit(_self_test())
# Filter mode: scrub stdin -> stdout, report to stderr. Useful for spot
# checks and for scrubbing a file by hand without writing throwaway code.
data = sys.stdin.read()
out, found = scrub(data)
sys.stdout.write(out)
if found:
by_rule: dict[str, int] = {}
for f in found:
by_rule[f.rule] = by_rule.get(f.rule, 0) + 1
print(f"[redact] {len(found)} redaction(s): "
+ ", ".join(f"{k}={v}" for k, v in sorted(by_rule.items())), file=sys.stderr)
if "--suspicions" in sys.argv:
for s in suspicions(out):
print(f"[suspicion] len={s.length} fp={s.fingerprint}", file=sys.stderr)
+11 -11
View File
@@ -13,14 +13,14 @@ This document covers: what the stores are, what a central palace changes when se
```mermaid
flowchart LR
Q["ask by MEANING<br/>'what do we know about X?'"] --> D["<b>Drawers</b><br/>chroma.sqlite3<br/>verbatim text, embedded"]
R["ask by ENTITY + TIME<br/>'what was true in June?'"] --> G["<b>Knowledge graph</b><br/>knowledge_graph.sqlite3<br/>typed facts, time-bounded"]
S["ask by ADDRESS + ORDER<br/>'what is waiting for me?'"] --> L["<b>Coordination log</b><br/>logstream.sqlite3<br/>addressed events + artifacts"]
T["ask by ASSOCIATION<br/>'what else touches this?'"] --> P["<b>Palace graph</b><br/>tunnels.json + hallways.json<br/>links between rooms"]
D --> DI["<b>Diaries</b> live here too:<br/>drawers with room=diary,<br/>read by recency, not similarity"]
Q["ask by MEANING<br/>'what do we know about X?'"] --> D["<b>Drawers</b><br/>chroma.sqlite3"]
R["ask by ENTITY + TIME<br/>'what was true in June?'"] --> G["<b>Knowledge graph</b><br/>knowledge_graph.sqlite3"]
S["ask by ADDRESS<br/>'what is waiting for me?'"] --> L["<b>Coordination log</b><br/>logstream.sqlite3"]
T["ask by ASSOCIATION<br/>'what else touches this?'"] --> P["<b>Palace graph</b><br/>tunnels.json + hallways.json"]
D --> DI["<b>Diaries</b> live here too:<br/>drawers with room=diary"]
```
All four stores are files inside **one palace directory**, so "the palace" is a directory you can back up in one go.
All four stores are files inside **one palace directory**, so "the palace" is a directory you can back up in one go. Each holds a different *shape* of thing: drawers hold verbatim text, embedded for similarity; the knowledge graph holds typed facts that are time-bounded; the coordination log holds addressed events and their artifacts, read in append order; the palace graph holds links between rooms. Diaries are drawers, but they are read by recency rather than by similarity.
| Store | You get things back by | Typical use | Wrong use |
|---|---|---|---|
@@ -110,19 +110,19 @@ sequenceDiagram
## 3. Deciding where something goes
The question that matters is not "is this important?" but **"how will I want this back, and does anyone need to act?"**
The question that matters is not "is this important?" but **"how will I want this back, and does a specific machine or agent need to act?"**
```mermaid
flowchart TD
Start["I have something worth keeping"] --> Act{"Must a specific<br/>machine or agent<br/>DO something?"}
Start["I have something worth keeping"] --> Act{"Must someone specific<br/>DO something?"}
Act -->|no| Change{"Is it a fact that<br/>changes over time?"}
Act -->|yes| Know{"Do they also need<br/>to KNOW it later?"}
Change -->|yes| KG["Knowledge graph<br/>kg_add / kg_supersede"]
Change -->|no| Mine{"Is it about MY session<br/>— what I tried, felt, learned?"}
Change -->|no| Mine{"Is it about MY session<br/>— tried, felt, learned?"}
Mine -->|yes| Diary["Diary entry"]
Mine -->|no| Drawer["Drawer<br/>wing + room, verbatim"]
Know -->|yes| Both["BOTH:<br/>drawer for the knowledge,<br/>event pointing at it"]
Know -->|no| Event["Coordination event<br/>to_agent = the specific agent"]
Know -->|yes| Both["BOTH:<br/>drawer + event pointing at it"]
Know -->|no| Event["Coordination event<br/>to_agent = that agent"]
```
Worked examples:
+138 -1
View File
@@ -171,10 +171,76 @@ An event is **still owed** when all of these hold:
1. it is directed at you (`to_agent = <you>`, and `to_agent = '*'` is excluded — §7.4);
2. it was not written by you (you cannot owe yourself);
3. **no** event of yours has *all* of: a strictly higher `seq`, a join to it (`metadata.ack_of = <id>` or a shared `correlation_id`), **and** a terminal status (`applied` / `superseded` / `failed` / `blocked`).
3. **no** event of yours has *all* of: a strictly higher `seq`, a join to it (`metadata.ack_of = <id>` or a shared `correlation_id`), **and** a terminal status (`applied` / `superseded` / `failed` / `blocked`);
4. **the requester has not explicitly withdrawn it** — see §3.3.1.
The strictly-higher-`seq` test is load-bearing: without it, one terminal reply on a `correlation_id` suppresses every *later* ask on that same correlation, permanently. And because `seq` is local arrival order, this test is only sound on a single replica — §7.3.
#### 3.3.1 Withdrawal: the one case where someone else may clear your obligation
> **Status: implemented in the toolkit, not yet in a released image.** Clause 4 is
> live in `extensions/pi/mempalace.ts` (`isWithdrawn`) with rule tests in
> `scripts/test-owed-withdrawal.sh`. It reaches a device only when an image bakes
> a toolkit revision containing it. Until then clauses 1–3 are the whole rule, and
> **a sender must assume its withdrawal has no effect on the recipient's mailbox.**
Clause 3 says "no event **of yours**", and that asymmetry is deliberate: owed-ness
is a statement about the *recipient's* accountability, so a requester must not be
able to delete an obligation the recipient genuinely has. It is wrong in exactly
one case — the requester retracting its **own** ask.
**Measured, 2026-09-08/09.** `pi@mbp-m1-2020` withdrew a v1.8.13 rollout ask to
`pi@tor-ms22`: terminal `superseded`, same `correlation_id`, `metadata.closes`
naming the thread, body "DO NOT SPEND A MINUTE ON v1.8.13". It recorded the
withdrawal as done. The withdrawal had **no effect** — its `from_agent` is mbp, so
it could never satisfy a join that only inspects tor-ms22's own events. tor-ms22's
next wake-up, 8h later and on the first boot of the image shipping this very
derivation, still listed the ask as owed, 41h old, for a release that had been
superseded and never installed there. Closing it cost the recipient a write on a
thread nobody wanted answered. **The failure is invisible from the sender's side:**
mbp did everything a sender is told to do and got a result indistinguishable from
success, which is why this went unnoticed for 41h rather than being caught at once.
A withdrawal is honoured only when **all** of these hold. Each is a guard against a
specific way a mailbox could otherwise be emptied silently, and each is pinned by a
mutation test:
| Requirement | The failure it prevents |
|---|---|
| `from_agent` = the **original requester** | a third party clearing someone else's obligation |
| `to_agent` = the recipient **exactly** (not `'*'`) | one broadcast emptying every machine's mailbox at once (§7.4) |
| terminal status | a `claimed`/`ready` note reading as a release |
| strictly after the ask (`hlc`, §7.3) | retiring an ask the requester sent **later** on the same correlation |
| joins the ask (`ack_of` or shared `correlation_id`) | clearing an unrelated thread |
| **an explicit marker**: `metadata.withdraws` **or** `metadata.closes`, whose value is exactly the ask's `correlation_id` or its event `id` | **the dangerous one** — see below |
**Why release must be stated rather than inferred.** The tempting rule is "any
terminal event from the requester clears it". That breaches this log's standing
safety direction, which is that every failure stays on the *noisy-but-visible*
side: a spuriously resurfacing item costs one turn of human correction, while a
suppressed unanswered ask is silent and permanent. Under the naive rule a
requester appending `applied` for its own bookkeeping — on the correlation,
addressed to the recipient, before the recipient ever replied — would silently
delete a real obligation. `withdraws` is canonical; `closes` is honoured because it
is already the fleet's de-facto marker (mbp seq 119; emb seq 117 and 118 all use
it). Prose does not count: the value must *name* the thread, so
`closes: "v1813-client-rollout-tor-ms22"` is a withdrawal and
`closes: "v1813-… — from the RECIPIENT side, which is the only side that can"` is
not.
**This does not break the fixed point in §9.2, and a reader will reach for §9.2 to
object.** That section rejects letting terminal directed events into the owed set,
because then a reply becomes owed by its requester, closing it mints a fresh
obligation, and the loop never terminates. That argument is about **candidates**.
Clause 4 adds a **clearer**. A withdrawal carries a terminal status, so it can never
be a candidate, and nothing new becomes owed: the asserting shape (`open`) and the
clearing shape (terminal) stay disjoint, which is what makes owed-ness terminate.
**Recipients keep the last word.** A withdrawal suppresses; it does not rewrite
history. The recipient may still append its own terminal reply — worth doing when
there is a finding to carry across, as tor-ms22 did at seq 122, salvaging two
measurements from the retracted thread.
### 3.4 Artifacts
Exact payloads for handoff: `patch`, `file`, `log`, `json`, `note`; UTF-8 text only; ≤4 MiB; `sha256` and `size_bytes` returned (`logstream.py:744-810`).
@@ -279,6 +345,10 @@ Measured on real data: the positive-control pair carries `seq` 26/27 and `hlc` `
Two consequences: **broadcasting an ask reaches no owed set at all** (the "don't broadcast an ask" anti-pattern is mechanically enforced, not merely advised), and to reach a whole fleet with something actionable you must write **one directed event per device**, sharing a `correlation_id` so the thread stays joinable. Separately, an event written with **no** `to_agent` matches no `to_agent=` query ever and is addressed to nobody.
**✅ Measured 2026-08-26, on two devices independently.** This was inferred from source until a two-arm control settled it. A `to_agent='*'` event with `status='open'` — the anti-pattern, planted deliberately — survives the raw candidate filter and so reaches the exclusion branch, which every previously written broadcast (all non-open) never did. Result: the raw query returned it, the derived owed set did not, and the delivery named only the directed arm of the pair. Confirmed on a second machine that had merely *received* the broadcast rather than planted it. Two arms differing only in `to_agent`, same poll and same fake sender, is what makes it evidence rather than an absence: "delivered exactly one of two" cannot be explained by a mailbox that delivers everything or nothing.
> ⚠️ **Correction to an earlier record.** A note filed the same day claimed this branch *could not* be exercised, reasoning that every broadcast in the log carries `status='ready'` and the protocol forbids writing an open broadcast. The reasoning was wrong in a specific and instructive way: the protocol forbids it *as production traffic*, which does not forbid planting one as a labelled, self-closed control. Where a rule blocks a measurement, check whether it blocks the *use* or the *test* before recording the branch as untestable.
### 7.5 `since_event_id` raises on an unknown id
An unresolvable cursor raises `ValueError` rather than returning empty (`logstream.py:966-970`) — stronger than the tool description promises, and benign on one replica. On a second replica, a resuming watcher whose cursor has not yet replicated will **hard-fail instead of waiting**.
@@ -309,6 +379,49 @@ The complete GET route table is `/healthz`, `/statusz`, `/logstream/stream`, `/s
WAL. **Measured 2026-08-26** on the primary: `logstream.sqlite3` mtime three days old while `logstream.sqlite3-wal` was 865 KiB and seconds old. An operator checking whether coordination is live must look at the `-wal` file, or query.
### 7.11 Delivered is not read: the human is the trigger
The mid-session poll delivers with `deliverAs: "steer"` and **deliberately no `triggerTurn`** — waking a model on inbound fleet traffic is a much larger behavioural change than auto-delivery, and was not approved. The poll fires on `agent_settled`, which means the agent is *idle*: nothing is running to react. So the delivered text is queued and read at the top of the **next turn**, whenever a human happens to start one.
**Measured 2026-08-26:** a delivery landed in a session window at ~20:50Z and sat there, visibly unreacted-to, until the operator asked "do I have to nudge you for you to read it?" — which is precisely the question this design produces. The delay is not the agent choosing to ignore its mailbox; between the poll and the next turn there is no inference at all.
The consequence is a **UX gap, not a defect**: from the outside, an inbound ask plus an idle agent reads as neglect. The delivered text explained how to *close* an ask but said nothing about *when* it would be seen — so the one reader who needed that fact, the human watching the window, was the one not told.
**✅ Addressed 2026-08-26 (toolkit `extensions/pi/mempalace.ts`), without touching the no-`triggerTurn` decision.** Two changes, in the order their value is realised:
1. **A delivery note in the message itself** — it says this is a queued message, that nothing woke the agent, and that any message starts the turn that will handle it. Costs nothing, changes no behaviour, and converts "why is it ignoring me?" into "right, I nudge it."
2. **A notification at poll time**, because the note only reaches someone already looking, and the case that actually loses an ask is nobody looking. `MEMPALACE_MAILBOX_NOTIFY` unset → in-TUI `ctx.ui.notify` (the surface `session_start` already uses); `=desktop` → additionally a terminal-native notification (Kitty `OSC 99`, else `OSC 777`), which is how a ping escapes a container with no `notify-send`, DBus or host access — the escape sequence is interpreted by the terminal emulator on the human's machine; `=0` → silent. Title and body are stripped of `;` and control bytes so a payload cannot forge an OSC field or terminate the sequence early. It fires only when something is due, because a ping on an empty poll trains its reader to ignore it.
**⚠️ Unverified through a multiplexer — measured 2026-08-26, tor-ms22.** The terminal-native path is written and syntax-checked but has **not** been observed to fire. On this client the layering is `kitty → tmux (on the host) → docker exec → pi (in the container)`, with two clients attached to the same tmux session (one local, one remote over SSH). Test sequences written directly to pi's tty produced **no notification on the remote client**; the local client is unobserved so far. The likely mechanism is that **tmux drops OSC sequences it does not itself implement** — reaching the outer terminal requires wrapping them in tmux's DCS passthrough (`ESC P tmux; … ESC \`, with inner `ESC` bytes doubled) *and* `allow-passthrough on`, which is not the default.
This compounds §7.11's detection problem rather than repeating it: **the container cannot see either layer**. `KITTY_WINDOW_ID` is absent because `docker exec` does not forward it, and `TMUX` is absent because tmux is running one level further out, on the host — so the client cannot detect the terminal it is speaking to *or* the multiplexer it must speak through. Explicit configuration is the only reliable route; autodetection is structurally blind here.
Open, and deliberately not settled by guessing: whether the local client shows it (pending); whether to add DCS-wrapping modes; and with two clients attached, **which** client should be notified — a pane's output reaches every attached client, so a work laptop and a home machine would both ping.
Still deliberately **not** done: `MEMPALACE_MAILBOX_TRIGGER=1` (opt-in, default off) to start a turn on arrival. Worth revisiting once notification is proven in daily use — but a machine that auto-turns on inbound events writes to a shared log with nobody watching, and §7.1 (no idempotency guard) is the reason to be slow about it.
### 7.12 The mailbox is an obligation channel, so a report addressed to you is never delivered
Candidates are drawn with `status="open"` (§3.3), so an event carrying a **terminal** status is not a mailbox candidate at all — no matter who it is addressed to. A `task.reply` sent to a named device to share a finding is therefore delivered to nobody, ever, and neither is any `event.ack`. It sits in the log until someone reads the log.
Read the filter precisely: candidacy requires **exactly `open`**, not merely "non-terminal". So a `claimed` announcement is equally undelivered — announcing that you have started work reaches the requester's mailbox no more than your finished report does. Claiming is still worth doing on long or handoff-prone work, because it puts *pickup* in the log for whoever later asks "did anyone start this before that container died?", but its audience is a log reader, not the requester. Note also that claiming does not quiet **your own** mailbox: the original ask stays owed until a terminal event of yours joins it (§3.3), which is exactly what the state machine means by `open --> claimed` *does NOT clear*.
**Measured 2026-08-26:** a device wrote a detailed report addressed to `pi@<peer>` with `status="applied"`, on a thread whose ask belonged to a third party. The recipient never saw it and only read it when the operator pointed at it by id — with the mailbox working exactly as specified throughout.
This is the shape of the channel, and it is worth stating because the natural inter-device message — *"here is something you should know"* — is exactly the shape that gets no delivery. Two ways to make news reach a peer: address it as a **directed `status="open"` ask** so it enters the owed set and gets closed when acted on (correct when a response is genuinely wanted), or accept that it is **pull-only** and pair it with a palace drawer, which the peer's search *will* surface later. What does not work is a terminal-status report and an expectation of attention.
### 7.13 An event addressed to an identity no session runs as is write-only
The owed set (§3.3) is the log's only push channel — it is what a client injects at wake-up. It keys on `to_agent`. A `from_agent` is whatever string the writer puts there (§6: "authenticates the fleet, not the agent"), so nothing stops a writer from authoring under an identity that no live session has ever run as. When it later receives a reply, that reply is addressed right back to the string the writer chose — and if nothing runs as that string, the reply is stored, queryable, and delivered to no one. Unlike §7.12, this is not about status; a **directed, `status="open"`** reply is undeliverable here, because the *address* rather than the *shape* is the defect.
**Measured 2026-08-27.** A device planted two directed `task.request`s under a synthetic `from_agent` (a release-provenance label, not an identity any session runs as), intending only to record where the ask came from. Both recipients replied correctly, addressed to that synthetic sender as the protocol requires — one an `event.ack` (`status="claimed"`), one a `task.reply` (`status="applied"`). Neither reply was ever delivered to anyone. One contained a live-credential exposure finding. It sat unread for roughly two and a half hours and was found only because a human asked whether mail had arrived.
**The general rule, not the instance:** before addressing an event, the writer is responsible for the address being an identity a live session actually runs as; and a writer that authors under an identity other than its own session identity has made every reply to that event undeliverable to itself. This fleet has already produced the failure in three unrelated shapes — a synthetic sender (above); an event addressed to a device since decommissioned; and a device-identity case mismatch (an uppercase form written where the fleet's convention is lowercase) — which is why this is recorded as a property of the channel rather than a mistake to remember not to repeat.
**The permitted exception, and its obligation.** Authoring under a synthetic identity is legitimate for controlled experiments — this fleet has run positive/negative mailbox controls this way deliberately — and this RFC does not forbid it. But a synthetic sender then **must name the real identity to reply to** in the body; the address line is not a place to record provenance if a reply is ever wanted.
**Proposed, not implemented:** on `event_append`, if `from_agent` differs from the writer's own stamped session identity, warn that replies to this event will be addressed to that other identity, and say whether any event in the log has ever been authored under it — a plain `DISTINCT from_agent` lookup. State the heuristic's limit honestly: an identity that has never authored is certainly unread by anyone; one that *has* authored may still belong to a device that no longer exists, which this check cannot see.
---
## 8. Phasing
@@ -338,6 +451,9 @@ Replication exists in code — `version_vector()`, `list_ops()`, `apply_remote_e
5. **Should `GET /logstream/events` exist?** A read-only HTTP tail would let non-MCP consumers (dashboards, CI) follow a stream without an MCP client. Today they must hold SSE or speak MCP.
6. **Undocumented limits.** The 64 KiB `metadata` cap and the lowercase-`type` regex are enforced server-side but absent from the tool schemas (§3.1). Document, or relax.
7. **Agent-name registry.** Nothing prevents two devices sharing one `from_agent` (§7.7). A warning at append time would be cheap.
8. ~~**Does an arriving ask deserve a notification?**~~ **Done 2026-08-26** — delivery note plus `MEMPALACE_MAILBOX_NOTIFY` (§7.11). What stays open is the narrower question: should `desktop` be the *default* rather than opt-in, and does an arriving ask ever deserve a turn (`MEMPALACE_MAILBOX_TRIGGER`)?
9. **Is there a delivery path for news rather than obligations?** (§7.12) Today the only delivered shape is a directed open ask. Options: leave it pull-only and rely on the paired drawer, or give informational events a distinct low-priority surface at wake-up ("3 reports since your last session") separate from the owed set. **Proposed direction in §9.2** — including why the obvious fix (widening the owed set) is not available.
10. **Should a write warn when `from_agent` is not the writer's own identity?** (§7.13) Related to decision 7 but distinct: decision 7 is about two devices *sharing* one identity by accident; this is about a writer deliberately authoring under an identity nobody runs, which makes every reply to that event undeliverable to itself. A `DISTINCT from_agent` lookup at append time is cheap and catches the unread-for-certain case; it cannot catch a decommissioned device that once wrote, which is the harder half of the same problem.
### 9.1 Retention — decided direction (2026-08-26)
@@ -356,6 +472,27 @@ Three constraints any implementation has to respect:
Open sub-questions: the window length (a quarter is the obvious first guess); whether cold storage is a sibling `logstream-archive-<period>.sqlite3` or plain files on disk; and whether archives are queryable through the same tools behind an explicit opt-in flag, or simply left as files for a human to open when a question reaches back that far.
### 9.2 A news surface, without new state — proposed direction (2026-08-27)
Decision 9 above is **proposed in direction**: news should get its own low-priority surface at wake-up, and the owed set should not be touched.
**Why the obvious fix is not available.** The tempting change is to let terminal directed events into the owed set, so a report addressed to a device reaches it. That breaks the derivation's fixed point. Owed-ness is defined (§3.3) as *directed at me, not written by me, and not joined by a later terminal event of mine* — so the asserting shape (`open`) and the clearing shape (terminal) must be disjoint. Make terminal events owed, and: a reply becomes owed by its requester; the requester clears it by writing a terminal event joined to it; that event is directed, terminal, and therefore owed by the original author; and every closure mints a fresh obligation. The loop never terminates. You could exempt `task.reply` and `event.ack` by `type`, but that is the same filter re-entered through a different door, with more surface to get wrong. **`status="open"` is not an arbitrary editorial choice about tone; it is what makes owed-ness terminate.**
**The proposal.** A second, clearly-separated section at wake-up — "N reports since your last session" — listing directed non-`open` events newer than the reader's own last activity, count-capped, bodies not inlined.
The cost that sinks the naive version is state: "since your last session" implies a per-device read cursor, and a cursor in container-local storage is precisely what a `--force-recreate` erases (the palace exists because container state does not survive). No new state is needed, because **the device's own last authored event is already a cursor**: surface directed events whose `hlc` is greater than the newest `hlc` among events written by this device. That anchor lives in the shared log, survives recreate, and is comparable across replicas for the same reason the owed-set join now uses `hlc` (§7.3).
Properties worth stating before anyone implements it:
- **Per stream, not global.** Take the anchor as the newest `hlc` of this device's events *in that stream*. A global anchor lets a chatty stream advance past unread news in a quiet one — the failure is silent, so this is not a refinement to leave for later.
- **A device that has never written sees everything.** Bounded by the count cap, and arguably correct on first boot; it is the same shape as a new joiner reading recent history.
- **A busy device gets a narrow window.** Correct by construction: you saw the log when you last wrote to it.
- **There is no read receipt, so repeats are possible** until the anchor advances. Acceptable only because this surface is a summary, not an interruption: count and one-line subjects, never bodies.
- **It must not resurface.** Owed items re-announce (`MEMPALACE_MAILBOX_RESURFACE_MS`); news must not, or it becomes nagging without obligation.
- **Opt-in first**, following the notification precedent (§7.11): inert unless explicitly enabled, and never counted in the owed total.
**What this does not fix.** It is still not push (§7.12 and the 404 on `/logstream/stream` for this deployment), still queued into the next turn rather than waking anyone (§7.11), and still useless to a device that never starts a session. For anything that must survive indefinitely, the paired **drawer** remains the durable channel; this surface only shortens the delay before a peer notices something already written.
---
## 10. Evidence index
+185
View File
@@ -0,0 +1,185 @@
# Secret hygiene: keeping credentials out of the palace
A palace is mined from **transcripts**, and transcripts contain whatever the
terminal printed. Print an environment, `cat` a `.env`, paste a `curl -H
"Authorization: …"`, and the secret becomes a drawer — searchable by every agent
on the fleet, on every machine, indefinitely. This document describes the
scrubbing that exists, what it deliberately does **not** try to do, and the two
layers it belongs in.
Measured on this fleet, 2026-08-27: a single leaked bearer token had reached
**3 drawers, 13 feeder inbox files across all three devices, and 10 local files
dating back 10 days**. Nobody was careless; an agent inspected an environment
variable while debugging. That frequency is the design premise — this is a
routine event to be handled by the pipeline, not an incident to be handled by
discipline.
## 1. Where the hook goes, and why there are two of them
| Layer | Implemented in | Sees | Catches |
|---|---|---|---|
| **Client, pre-staging** | `bin/mempalace_redact.py`, called from `bin/mempalace-pi-session` at the point the staged transcript is written | the machine's own environment **and** the transcript | the real-world case: a secret this machine holds, printed into a session, before it is uploaded anywhere |
| **Server, pre-persist** | *not yet implemented* — see §5 | only the text it is handed | secrets from clients that predate the scrubber, from non-pi clients, and secrets an agent types straight into `add_drawer` |
The client hook is the load-bearing one, because it is the only layer that can
compare text against **the actual secret values it holds** (see tier 1 below) and
because it stops the leak *before transmission*. The server hook is defence in
depth: pattern-only, but it covers write paths the feeder never sees.
Client placement matters for one specific reason: the pi feeder writes the staged
transcript from an in-memory structure, and the remote path then `rsync`s **that
same file** byte-for-byte. Scrubbing at the write point therefore covers local
mining and remote upload with a single hook — there is no second serialization to
forget.
## 2. Detection is anchored on meaning, never on entropy
The tempting design — "redact long random-looking strings" — is actively
destructive here, because a palace is *full* of high-entropy strings that are its
own primary keys:
```
drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b drawer id
evt_20260826T213924_397b1f710c2d event id
rep_d344e349ba276d6fc11997cd552f6937 replica id
1787780364792-000000-rep_d344e349ba276d6f… hybrid logical clock
ecc2a9c574f4156403cf4e90aea8b0d088c75700 git commit sha
```
An entropy detector fires on every one of those, and the resulting redaction is
silent, permanent, and destroys traceability. So detection uses three anchors
that carry meaning instead:
Each tier is a different *kind of evidence* that a string is a credential. The
tier is not a severity ranking of the secret — it is how much to trust the
detection. Rule names appear in the output; the tier tells you how to read them.
| Tier | Anchor | Default | False-positive risk |
|---|---|---|---|
| **1 — known values** | literal values from this process's env, for variables whose *name* says secret (`…TOKEN`, `…SECRET`, `…PASSWORD`, `…API_KEY`) | **redact** | none by construction: the value *is* the secret |
| **2 — known shapes** | vendor-prefixed credentials: `ghp_…`, `github_pat_…`, `glpat-…`, `xox[abprs]-…`, `sk-…`, `AKIA…`, `AIza…`, `hf_…`, JWTs, PEM private-key blocks, credentials inside URLs, `Authorization:` headers | **redact** | very low: the prefix is meaningful, not random |
| **3 — name=value** | an assignment whose *key* says secret | **report only** | measured **high** — see §3 |
### Worked examples, one line each
```
MEMPALACE_REMOTE_TOKEN=pfrDBfak… tier 1 — value matches this env's secret
the token is pfrDBfak… apparently tier 1 — same value, bare in prose, still caught
git clone https://joakim:hunter2@git/x tier 2 — credentials in a URL
Authorization: Bearer abcdefghijklmnop tier 2 — header shape
ghp_AbCdEf… / glpat-… / AKIA… / sk-ant-… tier 2 — vendor prefix
db_password = s3cr3t-p4ssw0rd-xyz tier 3 — only the KEY suggests it (reported)
GITEA_TOKEN=0123456789abcdef… (40 hex) tier 3 — indistinguishable from a commit sha
```
### Which tier fired? Read it off the output
The tool prints **rule names**, not tier numbers, because the rule says *what*
matched. The feeder prefixes them with the tier so both are visible:
```
[REDACTED] 2026-06-27T23-13-49.jsonl T1:env-value=1, T2:github-fine-grained=1 fp=ad78c7d4,fcd95ab5
[scrub] 29 redaction(s) applied, 141 name-anchored candidate(s) reported only
```
`RULE_TIERS` in `mempalace_redact.py` is the authoritative mapping, `tier_of()`
reads it, and a self-test fails if any rule is left unclassified — so the two
vocabularies cannot drift apart silently.
Tier 1 catches any presentation of a secret — env dump, JSON, error message,
URL, prose — because it matches the value itself. It is also, **in the default
configuration, the only tier that resolves a shape collision**: a 40-hex Gitea
PAT is byte-identical to a git commit sha, so only a known value (tier 1,
redacts) or a naming key (tier 3, reports) can tell them apart. A sha-shaped PAT
is therefore caught if and only if it belongs to this machine — an accepted gap,
since the alternative is redacting every commit sha in the palace.
## 3. The false-positive measurement, which changed the design
Tier 3 was originally going to redact. Measured against **52 MB of real fleet
transcripts (45 sessions)** it produced **403 hits**, and inspection of them (with
values masked) showed the overwhelming majority were not secrets:
- `GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}` — docker-compose interpolation
- `const tokens = countTokensForModel` — TypeScript source stored as memory
- `refreshToken(credentials: Credentials …)` — a *type annotation*
- `--setattr=krbPasswordExpiration=20260529…` — an IPA attribute holding a date
- `|1.FORGE.TOKENS:…` — AAAK diary shorthand
- `(root@host) Password:` followed by unrelated terminal output
Redacting those would corrupt code, configuration and documentation held as
memory, in order to catch secrets tier 1 already catches by value. After adding
guards for interpolation (`${VAR}`, `$VAR`, `%VAR%`, `{{tpl}}`), code context,
non-secret key suffixes (`…Expiration`, `…Count`, `…Type`) and all-digit values,
the enforced count on the same corpus fell from **403 to 29** — and those 29 are
tier-1 and tier-2 hits, i.e. real credential shapes.
So tier 3 **reports and does not rewrite** by default. Set
`MEMPALACE_REDACT_STRICT=1` to make it enforce, e.g. on a machine whose
transcripts are configuration-heavy rather than code-heavy.
## 4. Known false negatives — stated, not hidden
The scrubber will not catch a novel credential format pasted bare into prose
with no name nearby, a secret belonging to a machine whose environment this
process cannot see, a base64-of-a-secret, or a secret split across lines.
**"The scrubber ran" must never be read as "there are no secrets in here."** To
keep that measurable rather than assumed, two things are reported:
- every scrub prints a count — including `0 redaction(s)`, because printing
nothing is indistinguishable from a scrubber that never ran;
- `suspicions()` reports high-entropy strings it did **not** redact, as
`(length, fingerprint)` pairs rather than values, so the miss rate can be
tracked over time and a recurring fingerprint can be investigated by hand.
Findings never carry the secret. A `Finding` holds the rule, the label, the
length and `sha256(value)[:8]` — enough to recognise the same leak recurring,
not enough to recover it.
**Fail closed.** If the redactor cannot be imported, the feeder refuses to stage
rather than staging unscrubbed (`exit 3`). Override deliberately with
`MEMPALACE_FEED_ALLOW_UNSCRUBBED=1`.
## 5. The server-side layer (not yet implemented)
Three call sites, because the server has three near-duplicate validators rather
than one:
| Path | Function | Covers |
|---|---|---|
| drawers | `config.py` → `sanitize_content()` | `add_drawer`, `update_drawer`, `diary_write`, `checkpoint` — all four route through it |
| events | `logstream.py` → `_sanitize_body()` | `event_append`, and `event_ack` transitively |
| artifacts | `logstream.py` → `put_artifact()` | inlines its own checks; needs its own edit |
Server-side scrubbing is **tier 2 only** (plus optional tier-3 reporting): the
hub cannot see a client's environment, so tier 1 is structurally unavailable
there. That asymmetry is the reason the client hook is not redundant.
## 6. Cleaning up a leak that already landed
1. **Never re-echo the value while hunting it.** Pass it via stdin, never argv
(visible in `ps` on a shared host). Note `get_drawer` is a trap: reading a
drawer in order to redact it prints the secret back into the live transcript.
2. **Redact, don't delete.** Replacing the value in place keeps the mined
transcript's memory value; a stale embedding vector is a cheap price.
3. **Scrub the feeder inbox too**, or the next mine re-files it.
4. **A live session file needs an equal-length in-place overwrite** — the agent
holds it open, so temp-file-plus-rename loses everything appended afterwards.
5. **Expect residue.** SQLite keeps old page content in freed pages until
`VACUUM`, so a raw byte scan still matches after a successful `UPDATE`. Decide
explicitly whether that matters: if the credential store on the same host is
plaintext anyway, it usually does not.
## 7. Usage
```bash
# unit tests, including the must-not-redact corpus of palace id shapes
python3 bin/mempalace_redact.py --self-test
# scrub anything on stdin; report goes to stderr
some-command | python3 bin/mempalace_redact.py > clean.txt
# also list high-entropy strings that were NOT redacted, as fingerprints
python3 bin/mempalace_redact.py --suspicions < transcript.jsonl > clean.jsonl
```
+141 -2
View File
@@ -117,7 +117,7 @@ side of the wiring.
| `MEMPALACE_FEED_WING` | `wing_conversations` | Target wing — passed to both the exporter and the `mempalace_mine` call. |
| `MEMPALACE_FEED_DEBOUNCE_MS` | `600000` (10 min) | Minimum gap between mid-session (`agent_settled`) feeds. Bounds crash loss to one window instead of a whole session. |
| `MEMPALACE_FEED_PREPARE_TIMEOUT_MS` | `120000` | Kills a wedged `--prepare` subprocess. |
| `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `30000` | Caps the `mempalace_mine` call so a stalled palace can't hang session exit. |
| `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `300000` (5 min) | Bounds how long the extension *waits* for `mempalace_mine`, so a stalled palace can't hang session exit. It does **not** cancel the mine — see [Debugging](#debugging). Raised from `30000` in 2026-09: the mine is the slowest call this extension makes (30–60s in normal operation), so the old deadline fired routinely and reported healthy behaviour as an error. |
**Remote palace:** if `$MEMPALACE_REMOTE_URL` is set (see
[Transport](#transport-local-vs-external)), `mempalace_mine`'s source path is
@@ -316,6 +316,46 @@ waits for the agent to think of asking. Two delivery points, both fail-silent:
`triggerTurn` — waking the model on inbound fleet traffic is a much larger
behavioural change than auto-delivery.
**Which means a human is the trigger, and the mailbox now says so.** Measured
2026-08-26 on two devices: a delivery lands, the agent is idle, nothing happens,
and the operator asks *"do I have to nudge you for you to read this?"*. Yes —
because between the poll and the next turn no inference is running. The old text
explained how to *close* an ask and never said when it would be *seen*, so the
only reader who needed that fact was the one not told. Two additions, neither of
which touches the no-`triggerTurn` decision:
- **A delivery note in the message itself** — states that this is a queued
message, that nothing woke the agent, and that any message starts the turn
that handles it. Free, and aimed at the human reading the window.
- **A notification at poll time**, because the note only helps someone who is
already looking, and the case that loses an ask is nobody looking:
| `MEMPALACE_MAILBOX_NOTIFY` | Behaviour |
|---|---|
| *unset* (default) | in-TUI `ctx.ui.notify`, the same surface `session_start` already uses |
| `desktop` | additionally a terminal-native notification — Kitty `OSC 99`, else `OSC 777` (iTerm2, WezTerm, Ghostty, rxvt-unicode) |
| `0` / `off` | silent; mailbox still delivers |
The `desktop` path is how a notification escapes a container without
`notify-send`, DBus or any host access: the escape sequence is written to stdout
and interpreted by the terminal emulator on the human's own machine. It is
opt-in because writing raw escapes is a behaviour change on a shared machine,
not because it is unreliable. Title and body are stripped of `;` and control
bytes, so a payload can neither forge an OSC field nor end the sequence early.
**⚠️ Not yet observed firing through tmux (2026-08-26).** If pi runs inside a
multiplexer — e.g. `kitty → tmux on the host → docker exec → pi` — tmux drops
OSC sequences it does not implement, so the ping can vanish silently between
the container and the human. Reaching the outer terminal needs tmux's DCS
passthrough plus `allow-passthrough on`, which is not implemented here yet.
Note the client can detect *neither* layer: `KITTY_WINDOW_ID` is not forwarded
by `docker exec`, and `TMUX` is unset because tmux runs one level further out.
Verify with a hand-written sequence in your own stack before trusting it.
It fires **only when something is due** — the same condition as the delivery
itself. A ping on an empty poll would train its reader to ignore it, which is
the failure this whole feature exists to reverse.
Owed-ness is **derived, never read off a field**, because `event_ack` appends and
`status` is written once: a directed `open` event matches the mailbox query
*forever*, answered or not. Two calls (`to_agent=<me> status=open`, and
@@ -338,9 +378,61 @@ the one failure this derivation exists to prevent.
protocol says a broadcast owes nobody a reply — which also means broadcasting an
ask demonstrably reaches no owed set, giving that documented anti-pattern teeth.
**Dormant asks — open, owed, and deliberately not announced.** Some asks are
planted *unanswerable on purpose*: a first-boot acceptance describes work that
becomes possible only when the device is next recreated, and it must stay owed
until then, because closing it early to tidy the mailbox is exactly how the work
gets lost. Derivation alone cannot tell that apart from neglected work, so such
an ask was announced at every session start and every poll for as long as it was
correctly waiting — measured at three days running on `emb-7kj4vr4g`
(2026-09-28 → 2026-10-01), and across three consecutive releases before that.
The ask was right; announcing it was wrong, and the cost landed on the human
reading the window, the one reader who cannot filter it.
An ask may therefore declare, in its own `metadata`, the condition under which it
is merely waiting:
```json
"dormant_unless": [
{ "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json",
"field": "release_tag", "baseline": "v1.9.4" },
{ "kind": "file_mtime", "path": "/etc/hostname",
"baseline": "2026-09-22T18:12:49Z" }
]
```
It is dormant while **every** condition still matches its baseline, and goes live
the moment **any** of them differs — the trigger such asks already stated in
prose ("act when EITHER differs"), now in a form the bridge can check. Two kinds,
both local-file-only: `file_mtime` (compared at whole seconds in UTC, because a
filesystem mtime carries sub-second residue that a reported ISO baseline never
will — comparing raw milliseconds would mark every such predicate permanently
"changed" and silently disable the feature) and `json_field` (dotted paths
allowed, compared as strings so a manifest holding `3` matches a baseline of
`"3"`). There is no expression language, no shell and no network: a general
evaluator in the path that decides whether work is *visible* is a far worse trade
than a clumsy schema.
Dormancy is **withheld from the announcement, never from the mailbox**: the
wake-up injection lists dormant asks once per session with their ids, and a
mid-session poll mentions only a count, and only when the window is already open
for something else. They are not added to the resurface map, so one becomes
announceable the instant its baseline moves.
**Dormancy must be proven, never assumed** — every unevaluable predicate shows
the ask as owed. A missing file, an unreadable one, a baseline that will not
parse, an unknown `kind`, a vanished field, a relative path, more than eight
conditions: each announces. The dangerous failure here is not a spurious nag but
work that disappears because a predicate could not be evaluated, which would be
indistinguishable from the ask being lost and would not surface until a release
needed it. An ask with no `dormant_unless` behaves exactly as it did before the
feature existed. `scripts/test-dormancy.sh` exercises all of it, positive and
negative arms both.
Gated on `MEMPALACE_PI_DEVICE` **and** `MEMPALACE_REMOTE_URL` (the same pair as
the stamper, since an unstamped client has no address to be reached at), and
disabled outright with `MEMPALACE_MAILBOX=0`. Inert on a solitary palace: no
disabled outright with `MEMPALACE_MAILBOX=0` (notifications alone with
`MEMPALACE_MAILBOX_NOTIFY=0`). Inert on a solitary palace: no
calls, no injection. Delivery is the mechanism; the *norms* — what a reply owes,
and that only a terminal event closes a thread — remain normative in the
*consumer* skill (`~/.agents/skills/mempalace/SKILL.md`). **This file documents
@@ -382,6 +474,9 @@ one — the long-lived server is only killed when a request genuinely stalls.
- `MEMPALACE_MCP_TIMEOUT_MS` — tool-call/request timeout. Default `60000`.
Kept short on purpose: a *query* taking this long is genuinely wedged.
The one call that is not a query — the feed's `mempalace_mine` — passes its
own deadline (`MEMPALACE_FEED_MINE_TIMEOUT_MS`) down to the transport per
call, so this default does not apply to it (since 2026-09-18; see below).
- `MEMPALACE_MCP_INIT_TIMEOUT_MS` — `initialize` + `tools/list` handshake
timeout. Default `300000`. Deliberately generous: a genuine first
cold-open over virtiofs can legitimately take minutes, and killing a
@@ -424,6 +519,50 @@ the next tool call transparently respawns `mempalace-mcp` and retries.
`mempalace-mcp` manually with raw JSON-RPC on stdin to read the
server-side error — much faster than guessing.
### `feed (tick) failed: mine timed out after …ms`
**Nothing has been lost.** The deadline bounds only how long the extension
*waits*; it cannot cancel the mine, which continues on the server. The
transcript is already staged before the mine is invoked, and
`mempalace mine --mode convos` dedups by `source_file` and is idempotent, so the
work either completed after the deadline or is redone by the next tick.
**Do not "fix" it by retrying harder from the client.** The palace is a single
writer; a blind retry is what turns one slow mine into a queue of them.
Before 2026-09 this message appeared many times per session, which made it look
like a persistent failure. That was a real defect, now fixed: `lastFeedAt` was
recorded only after a *successful* wait, so a timeout left the debounce clock
stale and every following settled turn started another mine on top of the one
still running. Two changes — recording the attempt before the wait, and raising
the deadline to sit far above the slowest honest completion — mean a healthy
fleet should now never see it.
If you *do* still see it, it is now informative rather than noise: a mine
exceeded five minutes. Check palace size and whether another writer (a
host-side feeder, a scheduled `mine`) is holding the write lock, rather than
raising the timeout again.
### `feed (tick) failed: mempalace remote request 'tools/call' failed: timed out after 60000ms`
Same event, different deadline — and the same reassurance: **nothing has been
lost**, the mine continues server-side.
This is what the previous message turned into after the 2026-09 change, and
it exposed that the change was incomplete. The feed's five-minute deadline was
only *raced* against the call; `callTool()` had no way to carry it, so the
transport's generic per-request timeout (`MEMPALACE_MCP_TIMEOUT_MS`, 60 s)
fired first on every honest 60 s+ mine, and the 300 s was unreachable. On the
stdio transport this was worse than noise: a per-request timeout there kills the
server child, so the mine really was aborted at 60 s.
Fixed 2026-09-18: `callTool(name, args, { timeoutMs })` passes a per-call
deadline to both transports and the feed uses it for the mine.
`scripts/test-mcp-call-timeout.sh` pins the contract (a plain call still honours
the short default; the override is honoured and is itself a deadline). Seeing
this message on a fixed build means a mine exceeded *five* minutes — treat it as
the previous section says.
## The `Type.Unsafe` gotcha
Earlier versions of this extension registered every MCP tool with
+587 -40
View File
@@ -67,7 +67,9 @@
* child, so pi gets an error instead of hanging and later calls fail fast.
* This is a per-REQUEST timeout, not a process-lifetime one — the
* long-lived server is only killed when a request genuinely stalls.
* - MEMPALACE_MCP_TIMEOUT_MS tool-call/request timeout (default 60000)
* - MEMPALACE_MCP_TIMEOUT_MS tool-call/request timeout (default 60000);
* the feed's mine carries its own, longer
* deadline (MEMPALACE_FEED_MINE_TIMEOUT_MS)
* - MEMPALACE_MCP_INIT_TIMEOUT_MS initialize+tools/list timeout (default 300000)
* Set either to 0 to disable (legacy unbounded behavior).
*
@@ -90,6 +92,7 @@
*/
import { type ChildProcessWithoutNullStreams, spawn } from "node:child_process";
import { readFileSync, statSync } from "node:fs";
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { Type } from "typebox";
@@ -116,7 +119,13 @@ interface IMcpClient {
readonly alive: boolean;
onExit: (() => void) | null;
start(): Promise<void>;
callTool(name: string, args: Record<string, unknown>): Promise<any>;
/**
* `opts.timeoutMs` overrides the transport's generic per-request deadline
* for THIS call only. Callers that knowingly invoke a long server-side job
* (the feed's `mempalace_mine`) pass their own deadline here; everything
* else keeps the short default, which is the wedged-query guard.
*/
callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any>;
ensureAlive(): Promise<boolean>;
stop(): void | Promise<void>;
}
@@ -149,6 +158,165 @@ const sleep = (ms: number): Promise<void> =>
if (typeof t.unref === "function") t.unref();
});
// ── Dormancy ──────────────────────────────────────────────────────────────────
/**
* Is a standing ask DORMANT — correct to leave open, wrong to announce?
*
* The problem this solves, measured rather than imagined. A first-boot
* acceptance ask is planted deliberately unanswerable: it describes work that
* becomes possible only when the device is next recreated, and it must STAY
* owed until then, because closing it early to tidy the mailbox is exactly how
* that work gets lost. But deriveOwed could only see "directed, open, not
* answered, not withdrawn", so such an ask is announced at every session start
* and every poll for as long as it is correctly waiting — three days running on
* emb-7kj4vr4g (2026-09-28 → 2026-10-01), and three consecutive releases before
* that. The ask was right; announcing it was wrong. The cost lands on the human
* reading the window, who is the one reader that cannot filter it.
*
* So an ask may now declare, in its own metadata, the condition under which it
* is merely waiting:
*
* "dormant_unless": [
* { "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json",
* "field": "release_tag", "baseline": "v1.9.4" },
* { "kind": "file_mtime", "path": "/etc/hostname",
* "baseline": "2026-09-22T18:12:49Z" }
* ]
*
* It is dormant while EVERY condition still matches its baseline, and goes live
* the moment ANY of them differs — which is the trigger those asks already
* state in prose ("act when EITHER differs"), now in a form a machine can
* check. Dormant asks are withheld from the ANNOUNCED owed set, never from the
* mailbox: the wake-up injection still lists them once per session, so they
* stay discoverable and cannot be silently dropped.
*
* THE SAFETY RULE, and the reason every branch below reports `unchanged: false`
* on doubt: dormancy must be PROVEN, never assumed. A missing file, an
* unreadable one, a baseline that will not parse, an unknown `kind`, a field
* that has vanished — each is a reason to SHOW the ask, not to hide it. The
* dangerous failure here is not a spurious nag; it is work that vanishes
* because a predicate could not be evaluated, which would be indistinguishable
* from the ask being lost and would not surface until a release needed it.
*
* No eval, no shell, no network: the predicate language is deliberately two
* fixed checks against local files. A general expression evaluator sitting in
* the path that decides whether work is VISIBLE is a far worse trade than a
* slightly clumsy schema.
*/
export type DormancyVerdict = { dormant: boolean; reason: string };
/** Split of the derived owed-set: what to announce, and what is merely waiting. */
export type OwedSplit = { owed: LogEvent[]; dormant: LogEvent[] };
// Bounds, so a malformed or hostile event cannot turn a mailbox read into real
// work. Both are deliberately small: these predicates describe boot state, and
// anything needing more than a handful of local checks is not a dormancy rule.
const MAX_DORMANCY_CONDITIONS = 8;
const MAX_DORMANCY_JSON_BYTES = 256 * 1024;
const dormancyCondition = (c: unknown): { unchanged: boolean; reason: string } => {
if (typeof c !== "object" || c === null || Array.isArray(c))
return { unchanged: false, reason: "condition is not an object" };
const { kind, path, baseline, field } = c as Record<string, unknown>;
// Absolute paths only: a relative one would resolve against whatever cwd the
// pi process happens to hold, which is not a property of the device whose
// state the baseline describes.
if (typeof path !== "string" || !path.startsWith("/") || path.includes("\0"))
return {
unchanged: false,
reason: `path must be an absolute string, got ${JSON.stringify(path)}`,
};
if (typeof baseline !== "string" && typeof baseline !== "number")
return { unchanged: false, reason: "baseline must be a string or a number" };
if (kind === "file_mtime") {
// Compared at WHOLE SECONDS in UTC, because the two routes that produce
// these baselines disagree below that: a filesystem mtime carries
// sub-second residue (measured: /etc/hostname at .773761009) that a
// hand-written or reported ISO baseline never will. Comparing raw
// milliseconds would make every such predicate permanently "changed",
// i.e. would silently disable the feature while appearing to work.
const expected = Date.parse(String(baseline));
if (!Number.isFinite(expected))
return {
unchanged: false,
reason: `baseline is not a parseable date: ${JSON.stringify(baseline)}`,
};
let mtimeMs: number;
try {
mtimeMs = statSync(path).mtimeMs;
} catch {
return { unchanged: false, reason: `cannot stat ${path}` };
}
return Math.floor(mtimeMs / 1000) === Math.floor(expected / 1000)
? { unchanged: true, reason: `${path} mtime still at baseline` }
: { unchanged: false, reason: `${path} mtime moved from baseline` };
}
if (kind === "json_field") {
if (typeof field !== "string" || field.length === 0)
return { unchanged: false, reason: "json_field needs a non-empty field name" };
let text: string;
try {
const buf = readFileSync(path);
if (buf.byteLength > MAX_DORMANCY_JSON_BYTES)
return {
unchanged: false,
reason: `${path} exceeds the ${MAX_DORMANCY_JSON_BYTES}-byte cap`,
};
text = buf.toString("utf8");
} catch {
return { unchanged: false, reason: `cannot read ${path}` };
}
let doc: unknown;
try {
doc = JSON.parse(text);
} catch {
return { unchanged: false, reason: `${path} is not valid JSON` };
}
let cur: unknown = doc;
for (const seg of field.split(".")) {
if (
typeof cur !== "object" ||
cur === null ||
!Object.prototype.hasOwnProperty.call(cur, seg)
)
return { unchanged: false, reason: `${path} has no field ${field}` };
cur = (cur as Record<string, unknown>)[seg];
}
if (cur !== null && typeof cur === "object")
return { unchanged: false, reason: `field ${field} is not a scalar` };
// Compared as strings on purpose: the baseline arrives as JSON metadata, so
// a manifest holding 3 and a baseline saying "3" are the same fact.
return String(cur) === String(baseline)
? { unchanged: true, reason: `${field} still at baseline` }
: { unchanged: false, reason: `${field} moved from baseline` };
}
return { unchanged: false, reason: `unknown dormancy kind ${JSON.stringify(kind)}` };
};
export function evaluateDormancy(
metadata: Record<string, unknown> | null | undefined,
): DormancyVerdict {
const raw = metadata?.dormant_unless;
// Every one of these returns NOT dormant, i.e. "announce it". An ask without
// a predicate behaves exactly as it did before this feature existed.
if (raw === undefined || raw === null) return { dormant: false, reason: "no dormant_unless" };
if (!Array.isArray(raw)) return { dormant: false, reason: "dormant_unless is not an array" };
if (raw.length === 0) return { dormant: false, reason: "dormant_unless is empty" };
if (raw.length > MAX_DORMANCY_CONDITIONS)
return {
dormant: false,
reason: `too many conditions (${raw.length} > ${MAX_DORMANCY_CONDITIONS})`,
};
for (const c of raw) {
const v = dormancyCondition(c);
if (!v.unchanged) return { dormant: false, reason: v.reason };
}
return { dormant: true, reason: `all ${raw.length} condition(s) still at baseline` };
}
class StdioMcpClient implements IMcpClient {
private proc: ChildProcessWithoutNullStreams | null = null;
private nextId = 1;
@@ -372,8 +540,12 @@ class StdioMcpClient implements IMcpClient {
});
}
async callTool(name: string, args: Record<string, unknown>): Promise<any> {
return this.request("tools/call", { name, arguments: args });
async callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any> {
// The per-call override matters MORE here than for the HTTP client: on
// timeout this transport kills the server child, so a generic deadline
// that undercuts a long mine does not merely abandon the wait — it
// aborts the mine.
return this.request("tools/call", { name, arguments: args }, opts?.timeoutMs ?? this.requestTimeoutMs);
}
/** SIGTERM then SIGKILL grace, for stall recovery. */
@@ -421,6 +593,10 @@ class StdioMcpClient implements IMcpClient {
// • per-request AbortController timeout honouring MEMPALACE_MCP_TIMEOUT_MS /
// MEMPALACE_MCP_INIT_TIMEOUT_MS, mirroring StdioMcpClient's timeout ethos.
// • alive / ensureAlive / onExit to satisfy IMcpClient.
// • callTool() takes an optional per-call `{ timeoutMs }` (IMcpClient
// contract, see there) so the feed's long-running mine is not cut off by
// the generic per-request deadline. Not a protocol change; sync token
// unchanged.
//
// NOTE: mempalace-mcp --transport http is a SESSIONLESS, stateless JSON-RPC
// server (no Mcp-Session-Id, always application/json, Connection: close), so
@@ -466,8 +642,8 @@ class RemoteMcpClient implements IMcpClient {
this.healthy = true;
}
async callTool(name: string, args: Record<string, unknown>): Promise<any> {
return this.request("tools/call", { name, arguments: args });
async callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any> {
return this.request("tools/call", { name, arguments: args }, { timeoutMs: opts?.timeoutMs });
}
/**
@@ -855,7 +1031,22 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
const feedWing = process.env.MEMPALACE_FEED_WING ?? "wing_conversations";
const feedDebounceMs = num(process.env.MEMPALACE_FEED_DEBOUNCE_MS, 600_000);
const feedPrepareTimeoutMs = num(process.env.MEMPALACE_FEED_PREPARE_TIMEOUT_MS, 120_000);
const feedMineTimeoutMs = num(process.env.MEMPALACE_FEED_MINE_TIMEOUT_MS, 30_000);
// 300_000, raised from 30_000 on 2026-09-10. The mine is the SLOWEST thing
// this extension does — it offers every qualifying session transcript to a
// SINGLE-WRITER palace, measured at 30–60s in normal operation and growing
// with the corpus — yet it carried by far the TIGHTEST deadline: 4x tighter
// than the prepare step that precedes it (120_000) and 10x tighter than the
// init handshake (300_000), which is a fast call. All three were introduced
// together in 29e660e (2026-08-12) and this one was never revisited, so the
// deadline fired during entirely normal operation and the resulting message
// read as an error when nothing had gone wrong.
//
// Matched to the init timeout because liveness is the ONLY legitimate job
// left for this deadline: the Promise.race below abandons our WAIT, it cannot
// cancel the server's work, so the deadline buys nothing except an escape
// from a permanently hung call. It must therefore sit far above the slowest
// honest completion, not near it.
const feedMineTimeoutMs = num(process.env.MEMPALACE_FEED_MINE_TIMEOUT_MS, 300_000);
let lastFeedAt = 0; // 0 => the first settled turn also acts as a catch-up
let feedInFlight: Promise<void> | null = null;
@@ -909,17 +1100,46 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
try {
const source = await prepareFeed(reason);
if (!source) return;
// Record the attempt HERE, before awaiting — not after a successful
// wait. The race below abandons only our WAIT; the mine keeps running
// server-side, and `mine --mode convos` dedups by source_file and is
// idempotent, so a timeout is emphatically not a "did not happen".
//
// Leaving lastFeedAt stale on the timeout path defeated the debounce
// guard in the agent_settled handler below (`Date.now() - lastFeedAt <
// feedDebounceMs`): with lastFeedAt unchanged that guard passed on EVERY
// settled turn, and because `run` had already settled, feedInFlight was
// null too — so BOTH guards stood open. Each settled turn then launched
// another mine while the previous one was still running: overlapping
// writers queueing on a single-writer palace, each making the next one
// slower and the next timeout likelier. That positive feedback loop,
// not the tight deadline by itself, is why operators saw "mine timed out
// after 30000ms" many times per session rather than at most once per
// debounce window.
lastFeedAt = Date.now();
// The deadline is passed DOWN to the transport as well as raced
// here. Until 2026-09-18 it was only raced: callTool() had no way
// to carry it, so the transport's generic per-request timeout
// (MEMPALACE_MCP_TIMEOUT_MS, 60 000) fired first on every honest
// 60 s+ mine — "remote request 'tools/call' failed: timed out
// after 60000ms" — and the 300 000 below was unreachable. The
// race stays as the liveness guard for a transport whose timeout
// is disabled (0).
await Promise.race([
client.callTool("mempalace_mine", {
source,
mode: "convos",
wing: feedWing,
// Internal call: it does not pass through the registered tool's
// execute(), so it stamps itself. These ARE this harness's own
// transcripts from this device, so the harness segment is the
// agent (not `miner`) even though the tool is `mine`.
agent: stampProvenance ? `${agentName}@${device}` : agentName,
}),
client.callTool(
"mempalace_mine",
{
source,
mode: "convos",
wing: feedWing,
// Internal call: it does not pass through the registered tool's
// execute(), so it stamps itself. These ARE this harness's own
// transcripts from this device, so the harness segment is the
// agent (not `miner`) even though the tool is `mine`.
agent: stampProvenance ? `${agentName}@${device}` : agentName,
},
{ timeoutMs: feedMineTimeoutMs },
),
new Promise((_resolve, reject) =>
setTimeout(
() => reject(new Error(`mine timed out after ${feedMineTimeoutMs}ms`)),
@@ -927,7 +1147,6 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
),
),
]);
lastFeedAt = Date.now();
} catch (err) {
process.stderr.write(
`[mempalace ext] feed (${reason}) failed: ${(err as Error).message}\n`,
@@ -1068,17 +1287,120 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
return Boolean(m.correlation_id && m.correlation_id === candidate.correlation_id);
});
/** Directed asks addressed to this device with no terminal reply from it. */
async function deriveOwed(): Promise<LogEvent[]> {
if (!mailboxEnabled || !available) return [];
/**
* Has the ORIGINAL REQUESTER explicitly withdrawn this candidate?
*
* RFC 003 §3.3 clause 3 clears an ask only on "no event **of yours**", and the
* skill states the consequence outright: "there is nothing anyone can do about
* it from the other end". That asymmetry is deliberate and mostly right — owed
* ness is a statement about the RECIPIENT's accountability, and a requester
* must not be able to delete an obligation the recipient genuinely has.
*
* It is wrong in exactly one case: the requester withdrawing its OWN ask.
*
* MEASURED COST, 2026-09-08/09. pi@mbp-m1-2020 withdrew a v1.8.13 rollout ask
* to pi@tor-ms22 (evt seq 119: terminal `superseded`, same correlation_id,
* `metadata.closes` naming it, body "DO NOT SPEND A MINUTE ON v1.8.13") and
* recorded the withdrawal as done. It had no effect: seq 119's `from_agent` is
* mbp, so it could never satisfy a join that only looks at tor-ms22's own
* events. tor-ms22's next wake-up — 8h later, on the first boot of the image
* shipping this very derivation — still listed the ask as owed, 41h old, for a
* release that had been superseded and never installed there. It had to spend a
* write closing an ask nobody wanted answered. The asymmetry is also INVISIBLE
* from the sender's side: mbp did everything a sender is told to do and got a
* result indistinguishable from success.
*
* WHY AN EXPLICIT MARKER AND NOT "ANY TERMINAL EVENT FROM THE REQUESTER".
* This file's standing rule is that every failure mode stays on the
* noisy-but-visible side, because a spuriously resurfacing item is one turn of
* human correction whereas a suppressed unanswered ask is silent and permanent.
* Inferring release from any terminal event would breach that: a requester
* appending `applied` for its own bookkeeping — on the correlation, addressed
* to me, before I ever replied — would silently delete a real obligation.
* So release must be STATED, not inferred. `withdraws` is the canonical key;
* `closes` is honoured because it is already this fleet's de-facto marker (mbp
* seq 119, emb seq 117/118 all use it), and either must name this exact ask —
* its `correlation_id` or its event `id`. Prose does not count.
*
* WHY THIS DOES NOT BREAK THE FIXED POINT IN RFC 003 §9.2. §9.2 rejects letting
* terminal directed events into the owed set, because then every closure mints a
* fresh obligation and the loop never terminates. That argument is about
* CANDIDATES. This adds a CLEARER. A withdrawal carries a terminal status, so it
* can never be a candidate, and nothing new becomes owed. The asserting shape
* (`open`) and the clearing shape (terminal) stay disjoint. A reader who reaches
* for §9.2 to object is answering a different question.
*/
const isWithdrawn = (candidate: LogEvent, inbound: LogEvent[]): boolean => {
const requester = candidate.from_agent;
if (!requester) return false;
// Names THIS ask specifically: its correlation or its id. A marker naming
// something else, or carrying prose, is not a withdrawal of this ask.
const names = (v: unknown): boolean =>
typeof v === "string" &&
v.length > 0 &&
((Boolean(candidate.correlation_id) && v === candidate.correlation_id) ||
(Boolean(candidate.id) && v === candidate.id));
return inbound.some((e) => {
// Only the party that ASKED may retract. A third party writing a terminal
// event on a shared correlation must never clear my obligation.
if (e.from_agent !== requester) return false;
// Addressed to the party being released. `to_agent=<me>` also matches '*'
// at the SQL level (RFC 003 §7.4), and a broadcast must not be able to
// quietly empty every machine's mailbox at once.
if (e.to_agent !== mailboxAddress) return false;
if (!TERMINAL_STATUS.has((e.status ?? "").toLowerCase())) return false;
// Same ordering guard, same reason as isAnswered: a withdrawal must not
// retire an ask the requester sent LATER on the same correlation.
if (!isStrictlyAfter(e, candidate)) return false;
const ackOf = e.metadata?.ack_of;
const joins =
(typeof ackOf === "string" && Boolean(candidate.id) && ackOf === candidate.id) ||
Boolean(e.correlation_id && e.correlation_id === candidate.correlation_id);
if (!joins) return false;
return names(e.metadata?.withdraws) || names(e.metadata?.closes);
});
};
/**
* Directed asks addressed to this device with no terminal reply from it, and
* not explicitly withdrawn by whoever sent them (see isWithdrawn).
*/
async function deriveOwed(): Promise<OwedSplit> {
if (!mailboxEnabled || !available) return { owed: [], dormant: [] };
try {
const [candidatesRaw, mineRaw] = await Promise.all([
const [candidatesRaw, mineRaw, inboundRaw] = await Promise.all([
// Every cursor-less event_list call in this file passes `order` explicitly.
// The server's default is not ours to lean on: mempalace <= 3.9.0 defaults
// to `asc`, 3.10.0 flips a cursor-less listing to newest-first. Both
// derive* joins want the NEWEST window (see the comment below), so say so
// and the verdict stops depending on which server version answers.
client.callTool("mempalace_event_list", {
to_agent: mailboxAddress,
status: "open",
limit: 50,
order: "desc",
}),
// `order: "desc"` is a fix, not a flourish. event_list defaults to `asc`
// (append order), so this asked for the OLDEST 100 events this device
// ever wrote — meaning that once a device passes 100 authored events its
// most RECENT replies fall out of the join window and every ask it just
// answered resurfaces as owed. Latent rather than theoretical: tor-ms22
// was at ~20 authored events when this was written. The window has to be
// anchored at the newest end for the same reason the join uses `hlc`.
client.callTool("mempalace_event_list", {
from_agent: mailboxAddress,
limit: 100,
order: "desc",
}),
// Inbound with NO status filter, deliberately: a withdrawal carries a
// TERMINAL status and is therefore structurally invisible to the
// `status: "open"` candidate query above — the same omission deriveClosed
// is built on. No amount of polling the open set could ever see one.
client.callTool("mempalace_event_list", {
to_agent: mailboxAddress,
limit: 100,
order: "desc",
}),
client.callTool("mempalace_event_list", { from_agent: mailboxAddress, limit: 100 }),
]);
const candidates = parseEvents(candidatesRaw).filter((c) => {
// `to_agent: <me>` ALSO matches `*` broadcasts, per the tool contract.
@@ -1092,11 +1414,80 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
// You cannot owe yourself.
return c.from_agent !== mailboxAddress;
});
if (candidates.length === 0) return [];
if (candidates.length === 0) return { owed: [], dormant: [] };
const mine = parseEvents(mineRaw);
return candidates.filter((c) => !isAnswered(c, mine));
const inbound = parseEvents(inboundRaw);
const live = candidates.filter((c) => !isAnswered(c, mine) && !isWithdrawn(c, inbound));
// PARTITION, not filter. A dormant ask is still owed in the protocol
// sense — it has no terminal reply and must not be closed — so it stays in
// the return value where the mailbox can still account for it, and only
// the ANNOUNCING paths below treat the two halves differently.
const owed: LogEvent[] = [];
const dormant: LogEvent[] = [];
for (const c of live) (evaluateDormancy(c.metadata).dormant ? dormant : owed).push(c);
return { owed, dormant };
} catch {
// Fail silent and open: a mailbox read must never break a session.
return { owed: [], dormant: [] };
}
}
/**
* Terminal replies to asks THIS device sent. NOT work — news.
*
* Why this is a second query rather than a widened deriveOwed(): deriveOwed()
* asks the log for `status: "open"`, and a reply that CLOSES an ask is by
* definition not open, so it is structurally invisible to that filter — no
* amount of polling or waiting could ever have surfaced it.
*
* Measured 2026-09-07: emb-7kj4vr4g closed a v1.8.13 rollout ask with a
* task.reply at status=applied, and the operator reasonably expected to be
* told. The mailbox stayed silent and was CORRECT to — "owed" means "you must
* reply", and nothing was owed. But the single most useful thing a fleet can
* tell a human is "the thing you asked for is done" (here: done, AND your
* premise was wrong), and that was the one category the feature could not
* report. Silence was right by its own definition and wrong by the user's.
*/
async function deriveClosed(): Promise<LogEvent[]> {
if (!mailboxEnabled || !available) return [];
try {
// `order: "desc"` for the same reason as deriveOwed: without it a 3.9.0
// server hands back the OLDEST window, so past 100 authored events this
// device's newest asks (and past 50 inbound, the replies that close them)
// fall outside the join. The selection below is order-independent
// (isStrictlyAfter), so this only fixes WHICH events are in the window.
const [mineRaw, inboundRaw] = await Promise.all([
client.callTool("mempalace_event_list", { from_agent: mailboxAddress, limit: 100, order: "desc" }),
// NO status filter, deliberately: that omission is the entire fix.
client.callTool("mempalace_event_list", { to_agent: mailboxAddress, limit: 50, order: "desc" }),
]);
// Correlations this device opened as a DIRECTED ask. A broadcast owes
// nobody a reply, so it cannot be closed by one either.
const myAsks = new Map<string, LogEvent>();
for (const e of parseEvents(mineRaw)) {
if (!e.correlation_id || !e.to_agent) continue;
if (e.to_agent === "*" || e.to_agent === mailboxAddress) continue;
if (e.type !== "task.request") continue;
myAsks.set(e.correlation_id, e);
}
if (myAsks.size === 0) return [];
// Newest terminal reply per correlation only. A peer that appends
// applied-then-superseded should cost one line, not a wall of them.
const best = new Map<string, LogEvent>();
for (const e of parseEvents(inboundRaw)) {
if (e.from_agent === mailboxAddress) continue; // cannot inform myself
const cid = e.correlation_id;
if (!cid || !myAsks.has(cid)) continue;
// NOTE the deliberate asymmetry with deriveOwed: a broadcast is excluded
// there (it owes nobody) but allowed here, because a peer answering on MY
// correlation_id is news to me regardless of how widely it was addressed.
if (!TERMINAL_STATUS.has((e.status ?? "").toLowerCase())) continue;
const prev = best.get(cid);
if (!prev || isStrictlyAfter(e, prev)) best.set(cid, e);
}
return [...best.values()];
} catch {
// Same contract as deriveOwed: a mailbox read never breaks a session.
return [];
}
}
@@ -1133,11 +1524,41 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
})
.join("\n");
const DORMANT_NOTE =
"Each of these declares a `dormant_unless` predicate in its metadata that STILL " +
"matches this device's current state, so the work it describes is not yet " +
"possible. They are NOT closed and NOT answered — leave them open. They return to " +
"the announced owed-set by themselves the moment a baseline stops matching. If a " +
"predicate cannot be evaluated at all (missing file, unparseable baseline, unknown " +
"kind) the ask is announced as owed instead, deliberately: dormancy has to be " +
"proven, never assumed.";
const OWED_HOWTO =
"To close one, append an event on the SAME correlation_id with a terminal status " +
"(applied/superseded/failed/blocked). An ack alone does NOT clear it, and neither " +
"does claimed/ready — those deliberately keep it visible as taken-but-unfinished.";
// Written for the ONE reader the rest of this text ignores: the human watching
// the window. Measured 2026-08-26 on two devices — a delivery lands, the agent
// is idle, nothing happens, and the operator asks "do I have to nudge you for
// you to read this?". Yes, and it is by design (see the deliverAs comment
// below): at agent_settled no inference is running, so the text is queued for
// the next turn. The message explained how to CLOSE an ask but never when it
// would be SEEN, so the person who needed that fact was the only one not told.
// Cheapest possible fix, and it changes no behaviour: say so in the text.
const QUEUED_NOTE =
"DELIVERY NOTE — this is a QUEUED message, not an action: it arrived while this " +
"agent was idle, so no model call was made and nothing woke it. It is read at the " +
"start of the next turn. If you are the human watching and want it handled now, " +
"send any message to start a turn; the agent is not ignoring the ask, it is not running.";
const CLOSED_NOTE =
"NO ACTION IS OWED on these — they are replies to asks THIS device sent, shown " +
"once each because 'the thing you asked for is done' is news you wanted and the " +
"owed-set could never carry it. Read the reply before assuming your original ask " +
"was right: a peer that did the work is the most likely party to have found your " +
"premise wrong.";
// Mid-session mailbox. Tier 1 (the wake-up injection below) owns the first
// look; this exists because arrivals are bursty and correlate with our own
// activity — measured 2026-08-26, 11 of 22 events on this log landed inside a
@@ -1149,7 +1570,68 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
// the owed set holds something not already surfaced. Re-announcing the same
// ask every poll is precisely how a channel teaches its reader to ignore it,
// which is the failure this whole feature exists to reverse.
pi.on("agent_settled", async () => {
// --- A: tell the human at poll time (RFC 003 §7.11) -----------------------
//
// B (the QUEUED_NOTE above) fixes the confusion of someone who IS looking at
// the window. It does nothing for the case that actually loses an ask: nobody
// is looking. The delivered text is queued for a turn that only a human starts,
// so an ask can wait indefinitely on an idle session while its recipient makes
// coffee. A notification at poll time is the only part of this feature that
// reaches a person who is not watching.
//
// Two surfaces, deliberately split by risk:
// default in-TUI ctx.ui.notify, exactly as session_start already does.
// Zero new output channels; visible only if you are looking.
// =desktop additionally emit a terminal-native notification, reusing
// the detection the fleet's own notify.ts already proves in
// this harness (Kitty OSC 99 / Windows toast / OSC 777). This
// escapes the container without notify-send, DBus or any host
// access: the escape sequence is interpreted by the terminal
// emulator on the human's machine. Opt-in because writing raw
// escapes to stdout is a behaviour change on shared machines,
// not because it is unreliable.
// =kitty =osc777 force one protocol, skipping detection entirely.
// =0 silent, for anyone who wants the mailbox without pings.
//
// WHY FORCING EXISTS — measured on tor-ms22 2026-08-26, and it invalidates
// detection in exactly the deployment this ships in. Terminal identity lives in
// env vars set by the emulator (KITTY_WINDOW_ID, TERM_PROGRAM) and `docker exec`
// does NOT forward them: inside the container pi sees only TERM=xterm-256color
// no matter what is rendering it. So detection can never see Kitty from in here
// and always falls through to OSC 777, which Kitty does not implement — on a
// containerised client the notification would silently do nothing, the worst
// possible failure for a feature whose whole job is to break a silence. Naming
// the protocol ends the guessing.
//
// Fired only when something is actually due, i.e. the same condition as the
// delivery itself: a notification that fires on an empty poll would teach its
// reader to ignore it, which is the failure this whole feature exists to undo.
const mailboxNotifyMode = (process.env.MEMPALACE_MAILBOX_NOTIFY ?? "").trim().toLowerCase();
const mailboxNotifyEnabled = mailboxNotifyMode !== "0" && mailboxNotifyMode !== "off";
const mailboxNotifyTerminal =
mailboxNotifyMode === "desktop" ||
mailboxNotifyMode === "kitty" ||
mailboxNotifyMode === "osc777";
/** Terminal-native notification. Best-effort; never throws into a handler. */
const notifyTerminal = (title: string, body: string): void => {
try {
const esc = (s: string): string => s.replace(/[\x00-\x1f\x7f;]/g, " ");
if (
mailboxNotifyMode === "kitty" ||
(mailboxNotifyMode === "desktop" && !!process.env.KITTY_WINDOW_ID)
) {
process.stdout.write(`\x1b]99;i=1:d=0;${esc(title)}\x1b\\`);
process.stdout.write(`\x1b]99;i=1:p=body;${esc(body)}\x1b\\`);
return;
}
process.stdout.write(`\x1b]777;notify;${esc(title)};${esc(body)}\x07`);
} catch {
/* a notification is never worth an exception */
}
};
pi.on("agent_settled", async (_event, ctx) => {
if (!mailboxEnabled || !available) return;
if (!wokeUp) return; // the wake-up injection has not run yet; it does look #1
if (Date.now() - lastMailboxPollAt < mailboxPollMs) return;
@@ -1158,22 +1640,48 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
lastMailboxPollAt = Date.now();
void (async () => {
try {
const owed = await deriveOwed();
const [owedSplit, closed] = await Promise.all([deriveOwed(), deriveClosed()]);
const { owed, dormant } = owedSplit;
const now = Date.now();
const due = owed.filter((e) => {
if (!e.id) return false;
const last = surfaced.get(e.id);
return last === undefined || now - last >= mailboxResurfaceMs;
});
if (due.length === 0) return; // silence is the correct output here
for (const e of due) if (e.id) surfaced.set(e.id, now);
const unseen = (list: LogEvent[]): LogEvent[] =>
list.filter((e) => {
if (!e.id) return false;
const last = surfaced.get(e.id);
return last === undefined || now - last >= mailboxResurfaceMs;
});
const due = unseen(owed);
// Closing replies are announced ONCE and never resurface: an ask you
// already know is finished is not a nag, and re-announcing it hourly is
// the train-the-reader-to-ignore-it failure this window exists to stop.
const newsRaw = closed.filter((e) => e.id && !surfaced.has(e.id));
if (due.length === 0 && newsRaw.length === 0) return; // silence is correct
for (const e of [...due, ...newsRaw]) if (e.id) surfaced.set(e.id, now);
const blocks: string[] = [];
if (due.length > 0) {
blocks.push(
`${due.length} directed ask(s) addressed to "${mailboxAddress}" with no ` +
`terminal reply from this device. ${OWED_HOWTO}\n\n${formatOwed(due)}`,
);
}
if (newsRaw.length > 0) {
blocks.push(
`${newsRaw.length} repl(y|ies) CLOSING an ask this device sent. ` +
`${CLOSED_NOTE}\n\n${formatOwed(newsRaw)}`,
);
}
// A dormant ask NEVER opens this window on its own — the early return
// above is what actually stops the nag. But once the window is open for
// something else, one line of count is nearly free and keeps a waiting
// ask from feeling lost between session starts.
if (dormant.length > 0) {
blocks.push(`${dormant.length} dormant ask(s), not listed here. ${DORMANT_NOTE}`);
}
pi.sendMessage(
{
customType: "mempalace-mailbox",
content:
`MemPalace logstream mailbox: ${due.length} directed ask(s) addressed to ` +
`"${mailboxAddress}" with no terminal reply from this device. ${OWED_HOWTO}\n\n` +
formatOwed(due),
`MemPalace logstream mailbox.\n\n${QUEUED_NOTE}\n\n` +
blocks.join("\n\n---\n\n"),
display: true,
},
// "steer" is queue-on-arrival, NOT an interrupt: at agent_settled the
@@ -1185,6 +1693,23 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
// than auto-delivery and is not what was approved.
{ deliverAs: "steer" },
);
// AFTER the queueing call, so a notification can never be the only
// thing that happened: the message is in the thread first, then we
// point at it. Wording names the nudge explicitly, because "you have
// mail" without "press a key" reproduces the exact confusion B fixes.
if (mailboxNotifyEnabled) {
const parts = [
due.length > 0 ? `${due.length} ask(s) owed` : "",
newsRaw.length > 0 ? `${newsRaw.length} closed` : "",
].filter(Boolean);
const summary = `${parts.join(" + ")} for ${mailboxAddress} — send any message to handle`;
try {
if (ctx?.hasUI) ctx.ui.notify(`MemPalace mailbox: ${summary}`, "info");
} catch {
/* ignore: UI may be gone by the time the poll resolves */
}
if (mailboxNotifyTerminal) notifyTerminal("MemPalace mailbox", summary);
}
} catch {
/* best-effort delivery: never break a turn over a mailbox read */
}
@@ -1215,10 +1740,11 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
sections.push(`## mempalace_diary_read\n\n(error: ${(err as Error).message})`);
}
// Tier 1 mailbox: what this device owes a reply to, derived rather than read
// off a filter. deriveOwed() swallows its own failures and returns [], so a
// broken palace costs a missing section, never a broken wake-up.
// off a filter. deriveOwed() swallows its own failures and returns an empty
// split, so a broken palace costs a missing section, never a broken wake-up.
if (mailboxEnabled) {
const owed = await deriveOwed();
const [owedSplit, closed] = await Promise.all([deriveOwed(), deriveClosed()]);
const { owed, dormant } = owedSplit;
if (owed.length > 0) {
// Record what this injection showed, so the first mid-session poll does
// not re-announce the identical list minutes later. Without this the
@@ -1232,6 +1758,27 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
`this device. ${OWED_HOWTO}\n\n${formatOwed(owed)}`,
);
}
// Once per session, WITH ids. The wake-up injection is the one place a
// dormant ask should be visible: "what is this device still carrying?" is a
// session-start question, and answering it here is what keeps the mid-session
// silence honest rather than concealing. Deliberately NOT added to
// `surfaced`: that map exists to suppress repeats of ANNOUNCED work, and a
// dormant ask must stay announceable the instant its baseline moves.
if (dormant.length > 0) {
sections.push(
`## logstream mailbox (${dormant.length} dormant — waiting, nothing owed yet)\n\n` +
`${DORMANT_NOTE}\n\n${formatOwed(dormant)}`,
);
}
const news = closed.filter((e) => e.id && !surfaced.has(e.id));
if (news.length > 0) {
const now = Date.now();
for (const e of news) if (e.id) surfaced.set(e.id, now);
sections.push(
`## logstream mailbox (${news.length} closed — no action owed)\n\n` +
`${CLOSED_NOTE}\n\n${formatOwed(news)}`,
);
}
}
if (sections.length === 0) return;
+156
View File
@@ -0,0 +1,156 @@
#!/usr/bin/env bash
# test-dormancy.sh — exercise the mailbox dormancy predicate (dormant_unless).
#
# Why this exists as a script and not a note: `evaluateDormancy` decides whether
# an ask is SHOWN to the agent, so a silent regression there does not look like a
# bug — it looks like an empty mailbox. The positive arms matter as much as the
# negative ones: a predicate that never fires makes the feature a no-op, and a
# predicate that fires too eagerly hides real work. Both arms run here.
#
# It cannot import extensions/pi/mempalace.ts in place, because that file imports
# `typebox`, which pi provides at RUNTIME and this repo has no node_modules for.
# So it copies the file into a temp tree with the resolved deps symlinked beside
# it. The copy is made BY this script on every run, so it cannot drift from the
# source the way a vendored duplicate would.
#
# Fixtures are created here rather than read from the host: the first draft used
# /etc/hostname and /etc/pi-devbox/build-manifest.json, which made the positive
# arms pass only on a pi-devbox container and silently flip to "not dormant"
# anywhere else — a device-dependent test that reports success by doing nothing.
#
# Exit codes, deliberately distinct:
# 0 every case behaved as specified
# 1 at least one case FAILED (a real regression)
# 2 INCONCLUSIVE — deps could not be resolved, so nothing was proven
# (never 0: a test that skips quietly is the failure mode it should catch)
set -uo pipefail
# ── Args ──────────────────────────────────────────────────────────────────────
if [[ "${1:-}" == "--help" || "${1:-}" == "-h" ]]; then
sed -n '2,25p' "$0" | sed 's/^# \?//'
exit 0
fi
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
SRC="$REPO_ROOT/extensions/pi/mempalace.ts"
[[ -r "$SRC" ]] || {
echo "INCONCLUSIVE: cannot read $SRC" >&2
exit 2
}
# ── Resolve pi's runtime deps (discovered, never hardcoded) ───────────────────
GLOBAL_ROOT="$(npm root -g 2>/dev/null || true)"
PI_PKG=""
for cand in \
"${GLOBAL_ROOT:+$GLOBAL_ROOT/@earendil-works/pi-coding-agent}" \
"/usr/lib/node_modules/@earendil-works/pi-coding-agent" \
"/usr/local/lib/node_modules/@earendil-works/pi-coding-agent"; do
[[ -n "$cand" && -d "$cand" ]] && {
PI_PKG="$cand"
break
}
done
[[ -n "$PI_PKG" && -d "$PI_PKG/node_modules/typebox" ]] || {
echo "INCONCLUSIVE: pi-coding-agent / typebox not resolvable (looked under 'npm root -g')." >&2
echo " Nothing was proven. Install pi, or run this where pi is installed." >&2
exit 2
}
# Node must be able to strip types from a .ts entry point (Node >= 22.6).
node -e 'process.exit(0)' 2>/dev/null || {
echo "INCONCLUSIVE: no usable node" >&2
exit 2
}
# ── Build the throwaway tree ──────────────────────────────────────────────────
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
mkdir -p "$WORK/extensions/pi" "$WORK/node_modules/@earendil-works"
cp "$SRC" "$WORK/extensions/pi/mempalace.ts"
ln -s "$PI_PKG/node_modules/typebox" "$WORK/node_modules/typebox"
ln -s "$PI_PKG" "$WORK/node_modules/@earendil-works/pi-coding-agent"
printf '{"type":"module"}\n' > "$WORK/package.json"
# Prove the copy is the source, so a PASS cannot be about a stale file.
if ! cmp -s "$SRC" "$WORK/extensions/pi/mempalace.ts"; then
echo "INCONCLUSIVE: the copy differs from the source" >&2
exit 2
fi
# ── Fixtures, owned by this test ──────────────────────────────────────────────
printf '{"release_tag":"v1.9.4","nested":{"deep":{"leaf":"found"}},"count":3}\n' > "$WORK/manifest.json"
printf 'not json at all\n' > "$WORK/notjson.txt"
: > "$WORK/stamp"
# ── The truth table ───────────────────────────────────────────────────────────
cat > "$WORK/run.mjs" <<'EOF'
import { statSync } from "node:fs";
import { evaluateDormancy } from "./extensions/pi/mempalace.ts";
const W = process.env.WORK;
const STAMP = `${W}/stamp`;
const MAN = `${W}/manifest.json`;
// Floor to whole seconds: that is the contract, and the residue below proves
// why the implementation must do the same.
const atBaseline = new Date(Math.floor(statSync(STAMP).mtimeMs / 1000) * 1000).toISOString();
const mt = (baseline, path = STAMP) => ({ kind: "file_mtime", path, baseline });
const jf = (field, baseline, path = MAN) => ({ kind: "json_field", path, field, baseline });
const cases = [
// No predicate => behave exactly as before the feature existed.
["no metadata", undefined, false],
["metadata without the key", { foo: "bar" }, false],
["dormant_unless is a string", { dormant_unless: "release != x" }, false],
["dormant_unless is empty", { dormant_unless: [] }, false],
["more than 8 conditions", { dormant_unless: Array(9).fill(mt(atBaseline)) }, false],
// POSITIVE arms. If these ever read false the feature is inert.
["file_mtime at baseline", { dormant_unless: [mt(atBaseline)] }, true],
["json_field at baseline", { dormant_unless: [jf("release_tag", "v1.9.4")] }, true],
["both at baseline", { dormant_unless: [jf("release_tag", "v1.9.4"), mt(atBaseline)] }, true],
["dotted field path", { dormant_unless: [jf("nested.deep.leaf", "found")] }, true],
["number vs string baseline", { dormant_unless: [jf("count", "3")] }, true],
// The trigger firing: ANY condition differing wakes the ask.
["file_mtime moved", { dormant_unless: [mt("2020-01-01T00:00:00Z")] }, false],
["json_field moved", { dormant_unless: [jf("release_tag", "v1.9.5")] }, false],
["one same, one moved", { dormant_unless: [mt(atBaseline), jf("release_tag", "v1.9.5")] }, false],
// FAIL-VISIBLE arms: dormancy unproven => announce.
["missing file", { dormant_unless: [mt(atBaseline, "/nonexistent/path")] }, false],
["unparseable baseline", { dormant_unless: [mt("not-a-date")] }, false],
["relative path", { dormant_unless: [{ kind: "file_mtime", path: "etc/hostname", baseline: atBaseline }] }, false],
["unknown kind", { dormant_unless: [{ kind: "uptime_lt", path: "/etc/hostname", baseline: "1d" }] }, false],
["missing json field", { dormant_unless: [jf("no_such_field", "x")] }, false],
["file is not JSON", { dormant_unless: [jf("a", "b", `${W}/notjson.txt`)] }, false],
["field is not scalar", { dormant_unless: [jf("nested", "x")] }, false],
["condition is not an object", { dormant_unless: ["release_tag"] }, false],
["baseline is an object", { dormant_unless: [{ kind: "file_mtime", path: STAMP, baseline: {} }] }, false],
];
let pass = 0;
const failed = [];
for (const [name, meta, want] of cases) {
const v = evaluateDormancy(meta);
const ok = v.dormant === want;
if (ok) pass++;
else failed.push(name);
console.log(
` ${ok ? "PASS" : "FAIL"} dormant=${String(v.dormant).padEnd(5)} want=${String(want).padEnd(5)} ` +
`${name.padEnd(28)} :: ${v.reason}`,
);
}
console.log(`\n ${pass}/${cases.length} passed`);
if (failed.length) console.log(` FAILED: ${failed.join(", ")}`);
// Recorded because it is the reason file_mtime floors to seconds: a real mtime
// carries sub-second residue that an ISO baseline does not.
console.log(` (stamp mtimeMs=${statSync(STAMP).mtimeMs}, baseline=${atBaseline})`);
process.exit(failed.length === 0 ? 0 : 1);
EOF
cd "$WORK" || exit 2
export WORK
node "$WORK/run.mjs"
exit $?
+134
View File
@@ -0,0 +1,134 @@
#!/usr/bin/env bash
# test-mcp-call-timeout.sh — the per-call deadline of RemoteMcpClient in
# extensions/pi/mempalace.ts, and the override that the feed's mine relies on.
#
# WHY THIS EXISTS. The feed tick calls `mempalace_mine` through
# `client.callTool()`. Its own deadline (MEMPALACE_FEED_MINE_TIMEOUT_MS,
# 300 000) was raised in 2026-09 because the mine legitimately takes 30-60 s on
# a shared single-writer palace — but callTool() carried no way to pass that
# deadline down, so the transport's generic per-request timeout
# (MEMPALACE_MCP_TIMEOUT_MS, 60 000) fired first, and operators saw
# feed (tick) failed: mempalace remote request 'tools/call' failed: timed out after 60000ms
# instead of the message the 2026-09 change had aimed at. The outer race was
# unreachable in practice. This harness pins both halves of the contract:
# 1. a plain callTool() still honours the short per-request timeout
# (a query taking that long really is wedged — keep it short);
# 2. callTool(name, args, { timeoutMs }) honours the override, so a long
# server-side job can be given its own, longer, deadline.
# Before the fix, (2) fails: the third argument was silently ignored.
#
# Like test-owed-withdrawal.sh, it runs the SHIPPED text: the class is cut out
# of mempalace.ts by brace matching, type-stripped with node's own stripper,
# and driven against a local JSON-RPC server that delays tools/call.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
SRC="${1:-$REPO_ROOT/extensions/pi/mempalace.ts}"
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
[ -r "$SRC" ] || { echo "FAIL: cannot read $SRC" >&2; exit 2; }
# ---------------------------------------------------------------- extractor ---
cat >"$WORK/extract.mjs" <<'EXTRACT'
import { readFileSync, writeFileSync } from "node:fs";
import { stripTypeScriptTypes } from "node:module";
const src = readFileSync(process.argv[2], "utf8");
/** Cut a top-level `const NAME = ...;` out of the source, verbatim. */
function decl(name) {
const start = src.indexOf(`const ${name} =`);
if (start < 0) throw new Error(`declaration not found: ${name}`);
return src.slice(start, span(start));
}
/** Cut `class NAME ... { ... }` out of the source, verbatim. */
function klass(name) {
const start = src.indexOf(`class ${name} `);
if (start < 0) throw new Error(`class not found: ${name}`);
return src.slice(start, span(start, true));
}
/** End offset of the statement starting at `start` (brace-matched). */
function span(start, braceOnly = false) {
let depth = 0, inStr = null, sawBrace = false;
for (let i = start; i < src.length; i++) {
const c = src[i], prev = src[i - 1];
if (inStr) { if (c === inStr && prev !== "\\") inStr = null; continue; }
if (c === '"' || c === "'" || c === "`") { inStr = c; continue; }
if (c === "/" && src[i + 1] === "/") { i = src.indexOf("\n", i); if (i < 0) break; continue; }
if (c === "/" && src[i + 1] === "*") { i = src.indexOf("*/", i) + 1; continue; }
if (c === "{" || c === "(" || c === "[") { depth++; if (c === "{") sawBrace = true; }
else if (c === "}" || c === ")" || c === "]") { depth--; if (braceOnly && sawBrace && depth === 0) return i + 1; }
else if (!braceOnly && c === ";" && depth === 0) return i + 1;
}
throw new Error(`unterminated statement at ${start}`);
}
const parts = [decl("num"), decl("REMOTE_PROTOCOL_VERSION"), decl("REMOTE_CLIENT_INFO"), klass("RemoteMcpClient")];
for (const [i, p] of parts.entries()) if (p.length < 30) throw new Error(`extraction ${i} implausibly short: ${p}`);
if (!parts[3].includes("tools/call")) throw new Error("RemoteMcpClient does not mention tools/call");
const ts = parts.join("\n\n") + "\nexport { RemoteMcpClient };\n";
writeFileSync(process.argv[3], stripTypeScriptTypes(ts, { mode: "strip" }));
EXTRACT
node --no-warnings "$WORK/extract.mjs" "$SRC" "$WORK/client.mjs" || exit 2
node --check "$WORK/client.mjs" || { echo "FAIL: extracted client does not parse" >&2; exit 2; }
echo "[extract] pulled RemoteMcpClient from $(basename "$SRC") ($(wc -c <"$WORK/client.mjs") bytes)"
# ------------------------------------------------------------------ harness ---
cat >"$WORK/run.mjs" <<'RUN'
import { createServer } from "node:http";
import { RemoteMcpClient } from "./client.mjs";
// A sessionless JSON-RPC server like `mempalace-mcp --transport http`:
// initialize / tools/list answer at once; tools/call sleeps SLOW_MS first.
const SLOW_MS = 1500;
const server = createServer((req, res) => {
let body = "";
req.on("data", (c) => (body += c));
req.on("end", () => {
const msg = JSON.parse(body);
if (msg.id === undefined) { res.writeHead(202); res.end(); return; } // notification
const reply = (result) => {
res.writeHead(200, { "content-type": "application/json" });
res.end(JSON.stringify({ jsonrpc: "2.0", id: msg.id, result }));
};
if (msg.method === "initialize") return reply({ protocolVersion: "2024-11-05", capabilities: {}, serverInfo: { name: "fake", version: "0" } });
if (msg.method === "tools/list") return reply({ tools: [{ name: "slow", description: "", inputSchema: { type: "object" } }] });
if (msg.method === "tools/call") return void setTimeout(() => reply({ content: [{ type: "text", text: "done" }] }), SLOW_MS);
reply({});
});
});
await new Promise((r) => server.listen(0, "127.0.0.1", r));
const url = `http://127.0.0.1:${server.address().port}/mcp`;
let failures = 0;
const check = (ok, label) => { console.log(`${ok ? "ok " : "FAIL"} ${label}`); if (!ok) failures++; };
// Short per-request deadline, deliberately below SLOW_MS.
process.env.MEMPALACE_MCP_TIMEOUT_MS = "400";
const client = new RemoteMcpClient(url);
await client.start();
check(client.alive === true, "start(): initialize + tools/list complete, client alive");
// 1. A plain call honours the short deadline — this is the wedged-query guard.
let err = null;
const t0 = Date.now();
try { await client.callTool("slow", {}); } catch (e) { err = e; }
check(err !== null && /timed out after 400ms/.test(err.message), `plain callTool() rejects at the per-request deadline (${err && err.message})`);
check(Date.now() - t0 < SLOW_MS, "…and rejects BEFORE the server would have answered");
check(client.alive === false, "a timeout marks the client unhealthy (documented side effect; ensureAlive() revives)");
// 2. A call with its own deadline outlives the generic one — the feed's mine.
await client.ensureAlive();
err = null;
let result = null;
try { result = await client.callTool("slow", {}, { timeoutMs: 5000 }); } catch (e) { err = e; }
check(err === null && result && result.content?.[0]?.text === "done", `callTool(name, args, { timeoutMs: 5000 }) waits past the generic deadline and gets the result (${err ? err.message : "ok"})`);
// 3. The override is itself a deadline, not "no deadline".
err = null;
try { await client.callTool("slow", {}, { timeoutMs: 200 }); } catch (e) { err = e; }
check(err !== null && /timed out after 200ms/.test(err.message), `callTool(..., { timeoutMs: 200 }) rejects at ITS deadline (${err && err.message})`);
server.close();
console.log(failures === 0 ? "PASS: all assertions held" : `FAIL: ${failures} assertion(s) failed`);
process.exit(failures === 0 ? 0 : 1);
RUN
cd "$WORK" && node --no-warnings run.mjs
+312
View File
@@ -0,0 +1,312 @@
#!/usr/bin/env bash
# test-owed-withdrawal.sh — rule tests for the owed-set derivation in
# extensions/pi/mempalace.ts: isStrictlyAfter, isAnswered, isWithdrawn.
#
# WHY THIS EXISTS
# isWithdrawn changes who is allowed to clear an obligation, which is the one
# thing in this extension that can go wrong SILENTLY. A wrong rule here does
# not throw and does not show up in a log — it makes a real unanswered ask
# vanish from a mailbox forever. So the rules get pinned down before they ship.
#
# WHY IT EXTRACTS THE PREDICATES FROM THE SOURCE INSTEAD OF RESTATING THEM
# These are closures inside createExtension(), so they cannot be imported. The
# tempting shortcut is to paste a copy of the logic into the test — which tests
# the copy. This repo has already paid for a divergent second copy (the
# pi-extensions skill mirror, 9579 B behind for weeks; the shell-lint logic
# extracted to one file for the same reason). So the harness cuts the actual
# declaration text out of mempalace.ts by brace matching and runs THAT. Change
# the source and this test follows; paraphrase the source and it cannot.
#
# FIXTURES ARE VERBATIM REAL EVENTS from the fleet log (project/pi-devbox), not
# invented shapes — including the exact incident that motivated isWithdrawn:
# mbp-m1-2020 withdrew a v1.8.13 rollout ask to tor-ms22 at seq 119 and the
# withdrawal had no effect, so tor-ms22 was still being told it owed a reply 41h
# later for a release it never installed.
#
# Usage: scripts/test-owed-withdrawal.sh [source.ts]
# exit 0 = all rules behave 1 = a rule broke
# exit 2 = source unreadable or does not parse 3 = the gate itself cannot run
#
# The optional argument exists so the suite can be pointed at a deliberately
# MUTATED copy of the source to prove it is sensitive — a suite that has never
# been observed to fail is not evidence. See the mutation check in the CHANGELOG
# entry that introduced isWithdrawn.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
SRC="${1:-$REPO_ROOT/extensions/pi/mempalace.ts}"
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
[ -r "$SRC" ] || { echo "FAIL: cannot read $SRC" >&2; exit 2; }
# A gate that cannot run must not pass — the standing rule in this repo.
#
# NOT `node --check`: it does not type-strip, so it rejects ANY TypeScript —
# `const x: number = 1` included, not just the inline type-import at the top of
# mempalace.ts. It passed on node 22.x and stopped passing on node 24.x
# (measured on pi-devbox v1.9.1, node v24.21.0: this gate exited 2 on an
# unmodified mempalace.ts, so all 17 assertions below refused to run). Strip
# first, then syntax-check the emitted JS. `mode: "strip"` blanks type syntax
# without moving anything, so byte offsets and line numbers survive and a
# reported error line still points at the right line of the ORIGINAL .ts.
#
# The two failure modes are reported separately and on purpose. "Cannot strip"
# is a fact about the toolchain; "does not parse" is a fact about the source.
# Collapsing them is what made this very defect present itself as
# "mempalace.ts does not parse" when mempalace.ts was fine.
#
# SC2016 is disabled deliberately: the single quotes are the point. What follows
# is JavaScript, and `${process.version}` must reach node, not be expanded by the
# shell first.
# shellcheck disable=SC2016
node --no-warnings -e '
const { readFileSync, writeFileSync } = require("node:fs");
const { stripTypeScriptTypes } = require("node:module");
if (typeof stripTypeScriptTypes !== "function") {
console.error(`FAIL: node ${process.version} cannot strip TypeScript ` +
`(module.stripTypeScriptTypes needs >= 22.13) — the gate is unavailable, ` +
`which is NOT a statement about the source`);
process.exit(3);
}
let js;
try {
js = stripTypeScriptTypes(readFileSync(process.argv[1], "utf8"), { mode: "strip" });
} catch (err) {
// The stripper is itself a parser, so a syntax error lands HERE, not in the
// --check below. This is a statement about the source: exit 2, not 3.
console.error(`FAIL: ${process.argv[1]} does not parse: ${err.message}`);
process.exit(2);
}
writeFileSync(process.argv[2], js);
' "$SRC" "$WORK/stripped.mjs" || exit $?
node --check "$WORK/stripped.mjs" \
|| { echo "FAIL: $SRC does not parse (stripped output rejected)" >&2; exit 2; }
# ---------------------------------------------------------------- extractor ---
cat >"$WORK/extract.mjs" <<'EXTRACT'
import { readFileSync, writeFileSync } from "node:fs";
const src = readFileSync(process.argv[2], "utf8");
/**
* Cut `const <name> = ...;` out of the source by matching delimiters, so the
* test runs the shipped text. Returns the declaration verbatim.
*/
function decl(name) {
const start = src.indexOf(`const ${name} =`);
if (start < 0) throw new Error(`declaration not found in source: ${name}`);
let depth = 0;
let inStr = null;
for (let i = start; i < src.length; i++) {
const c = src[i];
const prev = src[i - 1];
if (inStr) {
if (c === inStr && prev !== "\\") inStr = null;
continue;
}
if (c === '"' || c === "'" || c === "`") { inStr = c; continue; }
if (c === "/" && src[i + 1] === "/") { i = src.indexOf("\n", i); if (i < 0) break; continue; }
if (c === "{" || c === "(" || c === "[") depth++;
else if (c === "}" || c === ")" || c === "]") depth--;
else if (c === ";" && depth === 0) return src.slice(start, i + 1);
}
throw new Error(`unterminated declaration: ${name}`);
}
const parts = ["TERMINAL_STATUS", "isStrictlyAfter", "isAnswered", "isWithdrawn"].map(decl);
// Sanity: the extractor must have found real bodies, not empty matches. A
// silently-empty extraction would make every assertion below pass vacuously.
for (const [i, p] of parts.entries()) {
if (p.length < 40) throw new Error(`extracted declaration ${i} is implausibly short: ${p}`);
}
if (!parts[3].includes("withdraws")) throw new Error("isWithdrawn does not mention its marker key");
writeFileSync(process.argv[3], parts.join("\n\n"));
EXTRACT
node "$WORK/extract.mjs" "$SRC" "$WORK/extracted.ts"
echo "[extract] pulled 4 declarations from $(basename "$SRC") ($(wc -c <"$WORK/extracted.ts") bytes)"
# ------------------------------------------------------------------ harness ---
{
cat <<'HEAD'
type LogEvent = {
id?: string;
seq?: number;
hlc?: string;
type?: string;
status?: string;
from_agent?: string;
to_agent?: string;
correlation_id?: string | null;
created_at?: string;
body?: string;
metadata?: Record<string, unknown> | null;
};
const mailboxAddress = "pi@tor-ms22";
HEAD
cat "$WORK/extracted.ts"
cat <<'TAIL'
// ---- VERBATIM fixtures from the real fleet log, stream project/pi-devbox ----
const REP = "rep_d344e349ba276d6fc11997cd552f6937";
/** seq 112 — mbp's v1.8.13 rollout ask to tor-ms22. The obligation. */
const ask112: LogEvent = {
id: "evt_20260907T125110_2bef6f9b07a8",
seq: 112,
hlc: `1788785470242-000000-${REP}`,
type: "task.request",
status: "open",
from_agent: "pi@mbp-m1-2020",
to_agent: "pi@tor-ms22",
correlation_id: "v1813-client-rollout-tor-ms22",
};
/** seq 119 — mbp's withdrawal of its OWN ask. Terminal, directed, marker present. */
const withdraw119: LogEvent = {
id: "evt_20260908T223949_8f5b9bec7c92",
seq: 119,
hlc: `1788907189286-000000-${REP}`,
type: "task.reply",
status: "superseded",
from_agent: "pi@mbp-m1-2020",
to_agent: "pi@tor-ms22",
correlation_id: "v1813-client-rollout-tor-ms22",
metadata: {
closes: "v1813-client-rollout-tor-ms22",
nothing_owed: "no reply required to this event",
replacement: "v1814-client-rollout-tor-ms22",
},
};
/** seq 120 — the LIVE v1.8.14 ask. Must survive the v1813 withdrawal. */
const ask120: LogEvent = {
id: "evt_20260908T224118_808de43d6982",
seq: 120,
hlc: `1788907278075-000000-${REP}`,
type: "task.request",
status: "open",
from_agent: "pi@mbp-m1-2020",
to_agent: "pi@tor-ms22",
correlation_id: "v1814-client-rollout-tor-ms22",
};
/** seq 122 — tor-ms22's OWN terminal reply, which is what actually cleared 112. */
const myReply122: LogEvent = {
id: "evt_20260909T061851_cf9bd18800db",
seq: 122,
hlc: `1788934731146-000000-${REP}`,
type: "task.reply",
status: "superseded",
from_agent: "pi@tor-ms22",
to_agent: "pi@mbp-m1-2020",
correlation_id: "v1813-client-rollout-tor-ms22",
metadata: { closes: "v1813-client-rollout-tor-ms22 — from the RECIPIENT side, which is the only side that can" },
};
const clone = (e: LogEvent, over: Partial<LogEvent>): LogEvent => ({ ...e, ...over });
let failed = 0;
const check = (name: string, expected: boolean, actual: boolean): void => {
const ok = expected === actual;
if (!ok) failed++;
console.log(`${ok ? " ok " : " FAIL"} ${name} (expected ${expected}, got ${actual})`);
};
console.log("\nisWithdrawn — the requester retracting its own ask");
// THE INCIDENT. Before this rule existed the answer was false and tor-ms22 was
// told it owed a reply for a release that no longer existed.
check("real seq 119 withdraws real seq 112", true, isWithdrawn(ask112, [withdraw119]));
check("withdrawal does NOT touch the live v1814 ask", false, isWithdrawn(ask120, [withdraw119]));
console.log("\nisWithdrawn — the ways a mailbox must NOT be silently emptied");
check(
"no marker: a bare terminal event from the requester is not a withdrawal",
false,
isWithdrawn(ask112, [clone(withdraw119, { metadata: { nothing_owed: "no reply required" } })]),
);
check(
"marker naming a DIFFERENT thread does not clear this one",
false,
isWithdrawn(ask112, [clone(withdraw119, { metadata: { closes: "v1814-client-rollout-tor-ms22" } })]),
);
check(
"marker carrying PROSE does not count (tor-ms22's own seq 122 shape)",
false,
isWithdrawn(ask112, [
clone(withdraw119, {
metadata: { closes: "v1813-client-rollout-tor-ms22 — from the RECIPIENT side, which is the only side that can" },
}),
]),
);
check(
"a THIRD PARTY cannot retract someone else's ask",
false,
isWithdrawn(ask112, [clone(withdraw119, { from_agent: "pi@emb-7kj4vr4g" })]),
);
check(
"a BROADCAST cannot empty every machine's mailbox at once",
false,
isWithdrawn(ask112, [clone(withdraw119, { to_agent: "*" })]),
);
check(
"a NON-TERMINAL status is not a withdrawal",
false,
isWithdrawn(ask112, [clone(withdraw119, { status: "open" })]),
);
check(
"an event addressed to a DIFFERENT device does not clear my ask",
false,
isWithdrawn(ask112, [clone(withdraw119, { to_agent: "pi@emb-7kj4vr4g" })]),
);
check(
"ORDERING: a withdrawal cannot retire an ask the requester sent LATER",
false,
isWithdrawn(clone(ask112, { seq: 121, hlc: `1788907300000-000000-${REP}` }), [withdraw119]),
);
console.log("\nisWithdrawn — accepted marker spellings");
check(
"canonical `withdraws` naming the correlation",
true,
isWithdrawn(ask112, [clone(withdraw119, { metadata: { withdraws: "v1813-client-rollout-tor-ms22" } })]),
);
check(
"marker naming the ask's EVENT ID instead of its correlation",
true,
isWithdrawn(ask112, [clone(withdraw119, { metadata: { withdraws: ask112.id } })]),
);
console.log("\nisAnswered — regression guard, unchanged behaviour");
check("my own terminal reply still closes my own ask", true, isAnswered(ask112, [myReply122]));
check("my reply on one thread does not close another", false, isAnswered(ask120, [myReply122]));
check(
"ORDERING still load-bearing: my reply cannot pre-close a LATER ask",
false,
isAnswered(clone(ask112, { seq: 130, hlc: `1788999999999-000000-${REP}` }), [myReply122]),
);
console.log("\nEND TO END — the derived owed set for tor-ms22 on 2026-09-09");
// The behaviour change, stated as the derivation states it. `mine` deliberately
// EXCLUDES seq 122 so this measures the new rule rather than the reply that
// happened to be written first.
const candidates = [ask112, ask120];
const mine: LogEvent[] = [];
const inbound = [withdraw119];
const owed = candidates.filter((c) => !isAnswered(c, mine) && !isWithdrawn(c, inbound));
const owedIds = owed.map((e) => e.correlation_id).join(", ");
check("exactly one ask remains owed", true, owed.length === 1);
check("and it is the v1814 one, not the withdrawn v1813", true, owedIds === "v1814-client-rollout-tor-ms22");
console.log(` derived owed set: [${owedIds}]`);
console.log(failed === 0 ? "\n=== PASSED ===" : `\n=== FAILED: ${failed} ===`);
process.exit(failed === 0 ? 0 : 1);
TAIL
} >"$WORK/harness.ts"
node --experimental-strip-types "$WORK/harness.ts"
+88
View File
@@ -0,0 +1,88 @@
#!/usr/bin/env bash
# test-rsync-ship-idempotency.sh — regression test for the ship-step bug fixed
# 2026-08-27 (bin/mempalace-pi-session: rsync --update -> --checksum).
#
# THE BUG: the stage file's mtime is deliberately the SOURCE transcript's mtime
# (os.utime() at :903, "preserve session mtime for dedup stability"), so a
# re-export of a session that has not been appended to since the last ship
# carries a mtime that is NOT newer than the receiver copy. rsync --update
# skips a file whose mtime is not strictly newer than the destination's, so a
# scrubbed re-export of a DORMANT session was silently dropped: content
# differed, mtime did not, "success" was reported, and unscrubbed bytes stayed
# in the palace host's inbox indefinitely. Measured on mbp-m1-2020 2026-08-27.
#
# THE ASSERTION mirrors that exactly and needs no ssh, no palace, no secret:
# stage a file, ship it, rewrite the content while RESTORING the original
# mtime (the same thing os.utime() does), ship again, assert the receiver's
# sha256 changed. On the pre-fix flag (--update) this fails; with --checksum
# (content comparison, mtime-independent) it passes.
#
# Deliberately does NOT invoke bin/mempalace-pi-session itself: that needs a
# live ssh target, a palace, MEMPALACE_* env — disproportionate scaffolding for
# what is, at its core, one rsync flag. Pulls the flag list out of the script
# instead of hand-copying it, so a future change to the ship command either
# updates this test's expectation or fails it loudly rather than drifting
# silently out of sync with what actually ships.
#
# Usage: scripts/test-rsync-ship-idempotency.sh (exit 0 = fix still holds)
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
SHIP_SCRIPT="$REPO_ROOT/bin/mempalace-pi-session"
SRC="$(mktemp -d)"; DST="$(mktemp -d)"
trap 'rm -rf "$SRC" "$DST"' EXIT
EXPECT='rsync -a --checksum --no-owner --no-group'
if ! grep -qF "$EXPECT" "$SHIP_SCRIPT"; then
echo "FAIL: $SHIP_SCRIPT no longer ships with '$EXPECT'." >&2
echo " Either the fix regressed, or the flags changed -- update this" >&2
echo " test's EXPECT string to match, don't just delete the test." >&2
exit 1
fi
ship() {
rsync -a --checksum --no-owner --no-group \
--include='*.jsonl' --exclude='*' \
"$SRC/" "$DST/"
}
# Same BYTE LENGTH on purpose, mirroring mbp-m1-2020's own reasoning for why
# dropping --update outright would not be enough: rsync's default quick check
# (no --update, no --checksum) already transfers when SIZE differs, so a test
# with mismatched lengths would pass by accident and prove nothing about
# --checksum specifically. A real redaction whose placeholder happens to match
# the secret's length is exactly the case that stays silently undetected unless
# content itself, not size or mtime, is compared.
UNSCRUBBED='live: MEMPALACE_REMOTE_TOKEN=AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA'
SCRUBBED='live: MEMPALACE_REMOTE_TOKEN=<redacted:MEMPALACE_REMOTE_TOKEN---------->'
if [[ ${#UNSCRUBBED} -ne ${#SCRUBBED} ]]; then
echo "FAIL: test fixture bug -- the two fixture strings must be equal length (${#UNSCRUBBED} vs ${#SCRUBBED}); this test would not exercise --checksum at all otherwise." >&2
exit 1
fi
# 1. Baseline: ship a session, receiver now matches sender.
printf '%s\n' "$UNSCRUBBED" > "$SRC/session.jsonl"
touch -d '2026-08-23T23:53:18' "$SRC/session.jsonl"
ship
before="$(sha256sum "$DST/session.jsonl" | cut -d' ' -f1)"
# 2. The bug's exact shape: rewrite content (same length!) as a scrubbed
# re-export would, then restore the ORIGINAL mtime, as the exporter's
# os.utime() does.
printf '%s\n' "$SCRUBBED" > "$SRC/session.jsonl"
touch -d '2026-08-23T23:53:18' "$SRC/session.jsonl"
ship
after="$(sha256sum "$DST/session.jsonl" | cut -d' ' -f1)"
if [[ "$before" == "$after" ]]; then
echo "FAIL: receiver sha256 unchanged after a content-only re-export at a held-constant mtime." >&2
echo " This is the exact defect --checksum was added to fix." >&2
exit 1
fi
if ! grep -q '<redacted:MEMPALACE_REMOTE_TOKEN' "$DST/session.jsonl"; then
echo "FAIL: receiver does not contain the scrubbed content." >&2
exit 1
fi
echo "PASS: ship step transfers changed content even when mtime is deliberately held constant (before=$before after=$after)"