Compare commits

...

23 Commits

Author SHA1 Message Date
Joakim Persson 975ab92943 feat(pi): let an ask declare dormancy, so waiting work stops nagging
A first-boot acceptance ask is planted deliberately unanswerable: it describes
work that becomes possible only when the device is next recreated, and it must
STAY owed until then, because closing it early to tidy the mailbox is exactly how
that work gets lost. deriveOwed could only see "directed, open, not answered, not
withdrawn", so such an ask was announced at every session start and every poll
for as long as it was correctly waiting -- measured at three days running on
emb-7kj4vr4g (2026-09-28 -> 2026-10-01), and across three consecutive releases
before that. The ask was right; announcing it was wrong, and the cost landed on
the human reading the window, who is the one reader that cannot filter it.

An ask may now declare the condition under which it is merely waiting:

  "dormant_unless": [
    { "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json",
      "field": "release_tag", "baseline": "v1.9.4" },
    { "kind": "file_mtime", "path": "/etc/hostname",
      "baseline": "2026-09-22T18:12:49Z" }
  ]

Dormant while EVERY condition still matches its baseline; live the moment ANY
differs -- which is the trigger those asks already stated in prose ("act when
EITHER differs"), now in a form the bridge can check. Two kinds, local files
only, no expression language, no shell, no network: a general evaluator in the
path that decides whether work is VISIBLE is a far worse trade than a clumsy
schema.

DORMANCY IS PROVEN, NEVER ASSUMED. Every unevaluable predicate announces the ask
instead of hiding it: missing file, unreadable file, unparseable baseline,
unknown kind, vanished field, relative path, more than eight conditions. The
dangerous failure here is not a spurious nag but work that disappears because a
predicate could not be evaluated -- indistinguishable from the ask being lost,
and not surfacing until a release needed it. An ask with no dormant_unless
behaves exactly as before, so this is backward compatible by construction.

Withheld from the ANNOUNCEMENT, never from the mailbox: deriveOwed now returns a
partition {owed, dormant} rather than a flat list, the wake-up injection lists
dormant asks once per session with ids, and a mid-session poll adds only a count
and only when the window is already open for something else. Dormant asks are
deliberately NOT added to the `surfaced` map, so one becomes announceable the
instant its baseline moves.

file_mtime compares WHOLE SECONDS in UTC. A filesystem mtime carries sub-second
residue (measured: /etc/hostname at .773761009) that a reported ISO baseline
never will, so comparing raw milliseconds would mark every such predicate
permanently "changed" -- silently disabling the feature while appearing to work.
The test records the residue for that reason.

scripts/test-dormancy.sh is this repo's first test: 22 cases, positive and
negative arms both, because a predicate that never fires makes the feature inert
and one that fires too eagerly hides real work. It copies the extension into a
temp tree with pi's typebox symlinked beside it (the copy is made per-run, so it
cannot drift like a vendored duplicate), owns its own fixtures rather than
reading /etc paths -- the first draft passed only on a pi-devbox container and
would have silently flipped to "not dormant" anywhere else -- and exits 2 for
INCONCLUSIVE rather than 0, since a test that skips quietly is the failure mode
it exists to catch. shellcheck clean.
2026-10-01 23:59:36 +02:00
joakimp 2167a1b033 fix(mailbox): pass order: "desc" on every cursor-less event_list call
Three of the five `mempalace_event_list` calls in the pi extension relied on
the server's default ordering: deriveOwed's `status: "open"` candidate query
(limit 50) and both windows in deriveClosed (from_agent limit 100, to_agent
limit 50). deriveOwed's other two calls already say `order: "desc"` and carry
the comment explaining why: on mempalace <= 3.9.0 the default is `asc`, so a
cursor-less `limit: N` returns the OLDEST N events, and once a device passes N
authored events its newest asks and the replies that close them fall outside
the join window. deriveClosed had exactly that latent truncation.

Why now: mempalace 3.10.0 (2026-09-15) flips the cursor-less default to
newest-first ("Logstream listings default to the newest events when no cursor
is given"). That change is server-side -- the extension talks to the fleet
hub over MEMPALACE_REMOTE_URL, so it lands when the hub upgrades, not when a
client does -- and it would have silently FIXED deriveClosed on 3.10.0 while
leaving it broken against any 3.9.0 hub. A mailbox verdict that depends on
which server version answers is the wrong shape. Saying `order` explicitly
makes both derive* functions read the same window on either version.

The selection in deriveClosed (isStrictlyAfter per correlation) is
order-independent, so this changes WHICH events are in the window, not how
the winner is picked. No behaviour change on a device with < 50 inbound and
< 100 authored events; the fleet hub is at 176 events total today, so the
window was about to matter.

Verified: file compiles under pi's own loader (jiti 2.7.0 from the installed
pi-coding-agent) before and after; `order:"desc"` literal count 3 -> 7, which
is the 3 added arguments plus 1 comment mention. `node --check` and a bare
`tsc` were tried first and both fail identically on HEAD (inline `type`
imports / no @types/node) -- checker faults, not this change.
2026-09-22 14:15:22 +02:00
joakimp 817b3a82b7 fix(feed): the mine's deadline never reached the transport; 60 s cut it off
The 2026-09 change raised MEMPALACE_FEED_MINE_TIMEOUT_MS to 300 000 and
raced it against client.callTool("mempalace_mine"). But callTool() had no
way to carry a deadline, so every call went out under the transport's
generic per-request timeout (MEMPALACE_MCP_TIMEOUT_MS, 60 000), which fired
first on every honest 60 s+ mine. Operators saw

    feed (tick) failed: mempalace remote request 'tools/call' failed:
    timed out after 60000ms

instead of the message the change had aimed at, and the 300 s was
unreachable. On stdio it was worse than noise: that transport kills the
server child on timeout, so the mine was actually aborted at 60 s.

callTool(name, args, { timeoutMs }) now passes a per-call deadline to both
transports; feedPalace() uses it for the mine. Plain calls keep the short
default — a query taking 60 s is still wedged. The Promise.race stays as the
liveness guard for a transport with its timeout disabled (0).

scripts/test-mcp-call-timeout.sh cuts RemoteMcpClient out of the shipped
file (as test-owed-withdrawal.sh does for the mailbox predicates), drives it
against a local JSON-RPC server that delays tools/call, and asserts: plain
call rejects at the generic deadline; the override outlives it; the override
is itself a deadline. Fails on the previous commit (2 of 6), passes here.
Vendored-copy delta noted in the header; protocol untouched, sync token
unchanged (check-mcp-client-sync.sh passes against pi-extensions).
2026-09-18 17:04:59 +02:00
joakimp dab989b068 fix(mailbox-tests): the owed-set suite's own gate refused to let it run on node 24
scripts/test-owed-withdrawal.sh has not executed a single assertion since the
image moved to node 24. Its precondition line

    node --experimental-strip-types --check "$SRC"

exits 2 on an unmodified extensions/pi/mempalace.ts, and the script is
`set -euo pipefail` with `|| exit 2`, so all 17 assertions and both regression
guards were skipped. Measured on pi-devbox v1.9.1, node v24.21.0.

CAUSE, MEASURED AND NARROWER THAN IT LOOKS. The reported error names the inline
type-import at mempalace.ts:92, which invites the reading "that import is
unusual". It is not the import: `node --check` does not type-strip AT ALL. A file
whose entire content is `const x: number = 1;` fails identically, with or without
--experimental-strip-types, while `node --experimental-strip-types <same file>`
executes it fine. --check has never been type-aware; execution is what gained
stripping. So this gate could never validate TypeScript, on any node.

WHY IT USED TO PASS -- INFERENCE, NOT MEASUREMENT. e2b060a's message records this
suite running on 2026-09-09 with a control pass and four mutation kills, when the
image shipped node v22.23.2; v1.9.1 ships v24.21.0. No node 22 exists on the box
where this was diagnosed, so the counterfactual was not executed. Treat "22
stripped for --check and 24 stopped" as a hypothesis consistent with the record,
not as a measured cause. What IS measured is the present-tense behaviour above.

WHY THIS MATTERS MORE THAN A RED TEST. This suite exists because isWithdrawn is
the one rule in the extension that can go wrong SILENTLY -- a wrong rule does not
throw and does not log, it makes a real unanswered ask vanish from a mailbox
forever. The rule shipped in e2b060a and has been baked since v1.9.1; the thing
that makes its failure mode visible has been dark for the same period. The
mitigating half: it failed CLOSED (exit 2, loud), never vacuously green. A gate
that cannot run must not pass, and it did not.

THE FIX. Strip first, then syntax-check the emitted JS: version-stable, and it
still refuses malformed input. `mode: "strip"` blanks type syntax without moving
anything, so offsets and line numbers survive and a reported error line still
points at the right line of the original .ts (69157 B in, 69157 B out).

THE EXIT CODES ARE NOW SPLIT, AND THAT IS THE POINT. 3 = the gate itself cannot
run (no module.stripTypeScriptTypes, i.e. node < 22.13). 2 = the source does not
parse. Collapsing the two is how this defect disguised itself: it printed
"mempalace.ts does not parse" while mempalace.ts was fine, sending a reader to
inspect the wrong file. Note the stripper is itself a parser, so a genuine syntax
error surfaces as an exception from the strip call rather than from --check; that
path is caught and reported as 2, not 3. Both remain failures. Neither passes.

VERIFIED IN SIX DIRECTIONS on this image, each expectation written down first:
  unmutated source          -> rc=0, PASSED (17/17)
  malformed TypeScript      -> rc=2, "does not parse: Expression expected"
  stripper made unavailable -> rc=3, "cannot strip", and NOT "does not parse"
                               (simulated by doctoring the runtime through
                               NODE_OPTIONS, so the shipped line ran as shipped)
  M1 marker requirement removed -> FAILED: 3   (no-marker, other-thread, prose)
  M2 third-party guard removed  -> FAILED: 1
  M3 to_agent guard removed     -> FAILED: 2   (broadcast and wrong-device share it)
  M4 ordering guard removed     -> FAILED: 1
The four kill counts are the same ones e2b060a recorded, so sensitivity is
restored rather than merely asserted.

WORTH KNOWING FOR THE NEXT PERSON WHO MUTATES THIS: the extractor carries its own
guard that rejects an isWithdrawn which no longer mentions "withdraws", so the
obvious M1 (delete the marker check outright) is refused before any assertion
runs -- correctly, but it looks like a crash. Mutate to `... || true` instead,
which removes the requirement while keeping the key mentioned.

SC2016 is disabled on the node invocation with a stated reason: the single quotes
are deliberate, the payload is JavaScript and `${process.version}` must reach
node rather than the shell. Checked that the file is shellcheck-clean at default
severity, as it was before this change, and at the -S error severity the
pi-devbox gate uses.

NOT FIXED HERE. Nothing in CI runs this suite -- it is a script an operator
invokes, which is precisely why a gate that fails loudly still went unnoticed
across a node bump. The harness's own runner two hundred lines below still passes
--experimental-strip-types, deliberately: it is a no-op on 24 and required on 22,
so it keeps the suite runnable on both. And the real reason this was found at all
is unrelated to CI: isWithdrawn was being exercised end-to-end on a released
image against the live logstream for the first time (project/pi-devbox seq 140),
and the suite was reached for as corroboration.
2026-09-14 16:27:29 +02:00
Joakim Persson e68ee2071c docs(pi-ext): document what the mine deadline does, and what the message means
Two corrections to the operator-facing docs, both exposed by 309980b.

The env table listed the default as 30000, which is now wrong, and described the
var as capping "the mempalace_mine call". It never did: it bounds how long the
extension WAITS. The mine keeps running on the server. That exact misreading is
what made a 30s deadline look safe on a call measured at 30-60s.

Added a Debugging entry for "feed (tick) failed: mine timed out after ...ms",
because every operator on this fleet has seen it and it was documented nowhere.
It states the three things a reader needs: nothing was lost (the transcript is
staged before the mine, and mine --mode convos dedups by source_file and is
idempotent); do NOT retry harder from the client, because the palace is a single
writer and a blind retry turns one slow mine into a queue; and after 309980b the
message should not appear on a healthy fleet, so if it does it now MEANS
something -- a mine exceeding five minutes, i.e. look at palace size or another
writer holding the lock rather than raising the timeout again.
2026-09-10 20:58:12 +02:00
Joakim Persson 309980b62c fix(pi-ext): stop the feed tick from launching overlapping mines
"[mempalace ext] feed (tick) failed: mine timed out after 30000ms" was parked as
cosmetic on 2026-08-27. It is not cosmetic: the tight deadline was the trigger,
but the defect is a positive feedback loop that puts multiple writers on a
single-writer palace.

lastFeedAt was assigned only AFTER a successful await. The Promise.race abandons
our WAIT and cannot cancel the server's work, so on a mine that takes longer
than the deadline -- measured at 30-60s in normal operation, against a 30s
deadline -- the catch ran with lastFeedAt UNCHANGED. That left both guards in
the agent_settled handler open at once: the debounce test
(`Date.now() - lastFeedAt < feedDebounceMs`) passed because lastFeedAt was still
stale, and feedInFlight was already null because `run` had settled. Every
subsequent settled turn therefore launched another mine on top of the one still
running, each making the next slower and the next timeout likelier -- which is
why operators saw the message many times per session instead of at most once per
10-minute debounce window.

Fix, three lines:
- move `lastFeedAt = Date.now()` to before the await, so a timeout still starts
  the debounce clock. A timeout is not a "did not happen": the mine is running
  server-side and `mine --mode convos` dedups by source_file and is idempotent.
- raise MEMPALACE_FEED_MINE_TIMEOUT_MS from 30_000 to 300_000. The mine is the
  slowest thing this extension does yet carried the tightest deadline: 4x
  tighter than the prepare step before it (120_000) and 10x tighter than the
  init handshake (300_000), a fast call. All three were introduced together in
  29e660e and this one was never revisited. 300_000 matches the init timeout
  because liveness is the only job left for this deadline -- it cannot cancel
  the server's work, so it must sit far above the slowest honest completion.

Simulated both guards over 10 minutes of settled turns at 20s intervals with a
60s mine: BEFORE 16 mines launched, 15 of them overlapping an already-running
mine; AFTER 2 launched, 0 overlapping. With a mine that exceeds even the new
deadline (400s): BEFORE 16/15, AFTER still 2/0 -- the lastFeedAt move is what
actually fixes it, and it holds even when the timeout still fires. The raise
stops the spurious message; the move stops the pile-up.

NOT fixed here: the message still goes to process.stderr, which pi renders into
the TUI input field. That needs a pi-side channel or a log file, and is tracked
separately. After this change the message should be rare, and when it does
appear it means something real: a mine exceeding five minutes.
2026-09-10 20:48:33 +02:00
joakimp e2b060a940 feat(mailbox): let a requester withdraw its own ask, with an explicit marker
RFC 003 §3.3 clause 3 clears an ask only on "no event OF YOURS", and the skill
states the consequence outright: "there is nothing anyone can do about it from
the other end". That asymmetry is deliberate and mostly right — owed-ness is a
statement about the RECIPIENT's accountability. It is wrong in exactly one case:
the requester retracting its own ask.

MEASURED COST. pi@mbp-m1-2020 withdrew a v1.8.13 rollout ask to pi@tor-ms22 at
seq 119 — terminal `superseded`, same correlation_id, metadata.closes naming the
thread, body "DO NOT SPEND A MINUTE ON v1.8.13" — and recorded it as done. It had
no effect: seq 119's from_agent is mbp, so it could never satisfy a join that
only inspects tor-ms22's own events. tor-ms22's next wake-up, 8h later and on the
first boot of the image shipping this very derivation, still listed the ask as
owed, 41h old, for a release it never installed. The failure is invisible from
the sender's side, which is why it went unnoticed for 41h.

WHY AN EXPLICIT MARKER RATHER THAN "ANY TERMINAL EVENT FROM THE REQUESTER".
This file's standing rule is that every failure stays on the noisy-but-visible
side: a resurfacing item costs one turn of human correction, a suppressed
unanswered ask is silent and permanent. Under the naive rule a requester
appending `applied` for its own bookkeeping — on the correlation, addressed to
me, before I ever replied — would silently delete a real obligation. So release
must be STATED: metadata.withdraws (canonical) or metadata.closes (already this
fleet's de-facto marker), whose value must NAME the ask — its correlation_id or
its event id. Prose does not count.

Five further guards, each pinned by a mutation test: only the original requester;
addressed to this device exactly, never '*'; terminal status; strictly after the
ask by hlc; and joined by ack_of or correlation_id.

This does NOT break the fixed point in §9.2. That section rejects letting
terminal directed events into the owed set, because then every closure mints a
fresh obligation. That is about CANDIDATES; this adds a CLEARER. A withdrawal is
terminal, so it can never be a candidate, and the asserting shape (open) and the
clearing shape (terminal) stay disjoint.

Also fixed here, because the new rule depends on the same windowing: the `mine`
query used event_list's DEFAULT `asc` order with limit 100, i.e. the OLDEST 100
events this device ever wrote. Once a device passes 100 authored events its most
RECENT replies fall out of the join window and every ask it just answered
resurfaces as owed. Latent, not theoretical — tor-ms22 was at ~20. Both windows
are now anchored at the newest end with order: "desc".

Verification: scripts/test-owed-withdrawal.sh, 17 assertions over VERBATIM
fixtures from the real log (seq 112/119/120/122). It extracts the predicates from
mempalace.ts by brace matching and runs the SHIPPED text rather than a pasted
copy — this repo has already paid for a divergent second copy. Sensitivity
proven by four mutations: removing the marker requirement flips exactly the 3
marker assertions, and disabling the third-party / broadcast / ordering guards
each flip exactly their own. Control passes.
2026-09-09 08:56:56 +02:00
joakimp e45f6b4301 feat(mailbox): surface replies that CLOSE this device's own asks
The mailbox could report what this device OWES, and structurally nothing else.
deriveOwed() queries the log with status:"open", and a reply that closes an ask is
by definition not open — so a peer answering my delegation was invisible to it at
every poll, forever, not just late.

Measured 2026-09-07: emb-7kj4vr4g closed correlation v1813-client-rollout-emb with
a task.reply at status=applied. Nothing was announced. The feature was CORRECT by
its own definition ("owed" = "you must reply", and nothing was owed) and wrong by
the operator's, who asked why no notification arrived. The single most useful thing
a fleet can say to a human is "the thing you asked for is done" — and that was the
one category it could not say.

Independent corroboration that this was a real gap and not a preference: emb's own
reply ends "The amd64 narrowing is filed as a drawer as well, SINCE A TERMINAL
EVENT REACHES NO MAILBOX." A peer had already diagnosed the hole and was routing
around it by hand.

deriveClosed(): my task.requests (correlation + directed) joined against inbound
events with NO status filter, keeping the newest TERMINAL_STATUS reply per
correlation. Verified by computing the exact predicate over the two real events:
the new path yields evt_20260907T183143 (applied); the old status:"open" query
yields 0. Announced once each (never resurfaced — a finished ask is not a nag),
and the copy warns that a peer who did the work is the likeliest party to have
found your premise wrong. Here both premises were wrong: mine that emb was amd64,
and emb's that tor-ms22 was therefore the last candidate.

Deliberate asymmetry, documented at the call site: a broadcast is excluded from the
owed set (it owes nobody) but allowed to CLOSE, since a peer answering on my
correlation_id is news however widely it was addressed.
2026-09-07 21:42:58 +02:00
joakimp 21023e7aa0 rfc-003: an event addressed to a nobody is write-only
§7.13 — the owed set keys on to_agent, and from_agent is whatever
string the writer puts there, so nothing stops authoring under an
identity no live session runs as. When the reply comes back addressed
to that string, it is stored and delivered to no one. Distinct from
§7.12: the defect is the address, not the status.

Measured 2026-08-27: two directed task.requests planted under a
synthetic from_agent (a provenance label, not a run identity) drew two
correct replies, addressed back to the label as the protocol requires.
Neither reply was ever delivered. One carried a live-credential
exposure finding; it sat unread for ~2h20m and was found only because
a human asked whether mail had arrived.

General rule stated, not just the instance: this fleet has already
produced the same failure shape three ways (synthetic sender above; an
event addressed to a decommissioned device; a device-identity case
mismatch). Permitted exception carried over from existing fleet
practice: a synthetic sender is fine for a controlled experiment, but
the body must then name the real reply-to identity.

Proposed (not implemented): warn at event_append when from_agent
differs from the writer's own session identity, using a DISTINCT
from_agent lookup to say whether anyone has ever authored under it —
stated honestly as a heuristic that catches 'never authored, certainly
unread' but not a decommissioned device that once did.

Decision 9 in the open-decisions list gets a numbered sibling, 10,
pointing at this section — related to decision 7 (agent-name
collision between two devices) but a different failure: not two
devices sharing a name, but one writer using a name nobody runs.
2026-08-27 23:34:20 +02:00
joakimp a361b71c40 ship: don't trust the mtime the exporter deliberately backdates
rsync --update skips a file whose mtime is not strictly newer than the
receiver's. The stage file's mtime IS the source transcript's mtime
(os.utime() at :903, "preserve session mtime for dedup stability"), so
re-exporting a session that has not been appended to since its last
ship produces a mtime that is not newer than what's already at the
receiver — exactly the case a redactor upgrade needs to ship, because
content differs while mtime does not. --update reports success and
sends nothing.

Reported and patched by pi@mbp-m1-2020 (evt_20260827T211925_9674a31da0b4,
artifact art_20260827T211839_fa52af563105, sha256 f217e47e…), measured
live: a scrubbed re-export of a dormant session (pi_01a03022-…f542)
sat unshipped in the palace host's inbox while every local signal
reported a clean stage, saved only because a host-side sweep happened
to rewrite the remote copy independently that same day.

--checksum compares content and ignores size/mtime entirely. Dropping
--update outright was considered and rejected: rsync's default quick
check already transfers on a SIZE difference alone, which is why the
observed case (33-byte placeholder vs a 43-byte token) would have been
masked as "fixed" by a change that only works until a redaction whose
placeholder happens to match the secret's length. os.utime() at :903
is untouched — its backdating is a separate, load-bearing design call
for dedup stability, out of scope for this fix.

Added scripts/test-rsync-ship-idempotency.sh: ships a file, rewrites
its content to an EQUAL-LENGTH string while restoring the original
mtime (what os.utime() does), ships again, asserts the receiver's
sha256 changed. Equal length is deliberate, not cosmetic — mismatched
lengths would pass via the quick check alone and prove nothing about
--checksum specifically; this is the same reasoning that ruled out
dropping --update. Verified the test discriminates: passes against
today's --checksum, fails against --update (checked by temporarily
substituting the flag in a copy, not committed).

Runs offline — a local rsync destination path exercises the same
size/mtime/checksum comparison as the ssh transfer, no palace or
network needed.
2026-08-27 23:29:30 +02:00
joakimp b2b50afcc1 rfc-003: propose a news surface that needs no cursor, and say why widening the owed set cannot work
Decision 9 (news vs obligations) gets a proposed direction in a new §9.2, following
how §9.1 promoted the retention decision.

The load-bearing part is the negative result. The tempting fix — let terminal
directed events into the owed set so a report addressed to a device reaches it —
breaks the derivation's fixed point. Owed-ness is "directed at me, not mine, and
not joined by a later terminal event of mine", so the asserting shape (open) and
the clearing shape (terminal) have to be disjoint. Make terminal events owed and a
reply becomes owed by its requester, whose closure is itself directed + terminal
and therefore owed by the original author: every closure mints a fresh obligation
and the loop never terminates. status="open" is not editorial taste about tone, it
is what makes owed-ness terminate. Worth writing down before someone "fixes" it.

The proposal itself avoids the cost that sinks the naive version. "Since your last
session" implies a per-device read cursor, and container-local state is exactly
what --force-recreate erases. No new state is needed: the device's own last
authored event is already a cursor, it lives in the shared log, and it is
comparable across replicas for the same reason the owed-set join now uses hlc
(§7.3). Named the non-obvious constraint too — the anchor must be computed PER
STREAM, because a global one lets a chatty stream advance past unread news in a
quiet one, and that failure is silent.

Also tightened §7.12: candidacy requires exactly `open`, not merely "non-terminal",
so a `claimed` announcement is as undelivered as a finished report. Claiming still
earns its keep on handoff-prone work (it is what distinguishes "nobody started"
from "someone started and the container died"), but its audience is a log reader,
not the requester — and it does not quiet the claimer's own mailbox either, which
is what the state machine's "open --> claimed does NOT clear" already implies.
2026-08-27 17:28:21 +02:00
Joakim Persson de59571966 docs: unclip fleet-memory's diagrams, and measure the host instead of guessing
Seven labels in fleet-memory.md were losing their last line for readers, in the
document held up as the model for user-facing docs. Found by running a new
mermaid checker against a repo it was not written for, which is the strongest
evidence available that this is a systemic defect class and not one author's slip.

Measured on the real Gitea render rather than a simulation. Gitea serves each
mermaid block from a sandboxed same-origin <iframe srcdoc>, so its
.markup{line-height:1.5!important} never reaches the labels and the page is
directly measurable. On the published version: 31 labels, 6 of them rendering
MORE lines than were authored, worst overflow +0.6px past the clip edge. Mermaid
measures a label with its own metrics, commits to a box, then clips whatever does
not fit; per-line rounding accumulates, so every cut label observed anywhere in
this fleet had four or five rendered lines.

Fixes, all preserving the information rather than deleting it:

- The five store nodes drop their third line; the content shapes they carried
  ("verbatim text, embedded", "typed facts, time-bounded", "addressed events and
  their artifacts", "links between rooms", "read by recency") move into the prose
  under the diagram, which is where a qualifier belongs anyway.
- "ask by ADDRESS + ORDER" soft-wrapped its own first line because a caps line
  nearly as wide as the box overflows first; shortened to "ask by ADDRESS", with
  "read in append order" added to the prose so the ordering property survives.
  The table already lists "address, correlation, append order".
- The decision diamond loses "machine or agent" to the sentence above it, which
  now asks whether "a specific machine or agent" must act — same words, no clip.
- "to_agent = the specific agent" and the BOTH node shortened; both are spelled
  out in full in the worked-examples table directly below.

Nothing was cut for brevity's sake: every phrase removed from a node reappears in
the prose or was already in the table beneath it.
2026-08-27 16:15:44 +02:00
joakimp f0bffd1b93 feed: chase the symlink, or fail-closed takes the whole fleet's feed down
The image installs /usr/local/bin/mempalace-pi-session as a symlink into
/opt/mempalace-toolkit/bin, and ${BASH_SOURCE[0]} reports the path the script was
INVOKED as, not the resolved target. So the sibling-module lookup pointed at
/usr/local/bin, the redactor was not there, and the fail-closed import did exactly
what it was told: refused to stage.

MEASURED, not theorised: /usr/local/bin/mempalace-pi-session --dry-run exited
"[FATAL] secret scrubber unavailable ... refusing to stage" while
/opt/mempalace-toolkit/bin/mempalace-pi-session --dry-run scrubbed 40 findings on
the same input. The symlink is how every device invokes it, so at the next image
bake every feeder tick on every machine would have stopped staging — a silent,
fleet-wide memory outage, which is a worse outcome than the leak the scrubber
exists to prevent. My own tests missed it by calling bin/... directly from the
checkout, i.e. the one invocation path the fleet never uses.

FIX: chase the symlink chain in portable shell and offer fallbacks instead of
betting on a single answer. MEMPALACE_REDACT_DIR is now colon-separated —
resolved dir, invoked dir, then the image's known install path /opt/... — and the
Python side inserts the first existing candidate. readlink -f is deliberately NOT
used: it is GNU/newer-BSD only and this script also runs directly on macOS hosts,
so the chase is a plain while [ -L ] loop handling relative link targets.

Verified on all three invocation shapes: via the /usr/local/bin symlink, via the
direct /opt path, and via a second-hop symlink in an unrelated directory. All
three now report the same 40 redactions.

LESSON worth keeping: fail-closed is correct for a secret scrubber, but it
converts "module not found" into an outage, so the module lookup becomes
load-bearing infrastructure and must be tested through the real invocation path,
not the convenient one.
2026-08-27 14:18:02 +02:00
joakimp 836e35b320 redact: name the tiers where they are used, and stop the docstring lying about tier 3
Review caught that the tier vocabulary was used in the report, the commit message
and the docs without being defined anywhere the reader would land, and inspecting
that turned up two real defects rather than just a wording gap.

1. STALE DOCSTRING. The module still described tier 3 as if it redacts, which
   stopped being true when the 403-hit measurement demoted it to report-only. It
   also credited tier 3 with resolving the 40-hex-PAT-vs-commit-sha collision —
   false by default, since a reporting rule resolves nothing. Corrected, with the
   consequence stated plainly: in the default configuration a sha-shaped PAT is
   caught if and only if it belongs to THIS machine, because only a known value
   (tier 1) or a naming key (tier 3, reporting) can separate it from a commit
   sha. That is an accepted gap; the alternative is redacting every sha in the
   palace.

2. THE VOCABULARY NEVER REACHED THE OUTPUT. The tool prints rule names
   (github-pat, env-value, url-credentials) and nothing printed a tier, so the
   docs' tier language was unconnected to what an operator actually sees. Added
   RULE_TIERS as the authoritative rule -> tier mapping, tier_of(), and
   Finding.tier; the feeder now prints "T2:github-pat=6" so what matched and how
   much to trust it are both visible on one line. A self-test asserts every rule
   that can appear in a Finding maps to a tier, so adding a rule without
   classifying it fails the tests instead of printing "T?".

Tiers, for the record, are three kinds of EVIDENCE (not three severities):
T1 known value from this process's env — near-certain, zero FP by construction;
T2 known vendor shape — strong, the prefix is meaningful; T3 key name says
secret — candidate only, measured FP-heavy, reported. T0 is reserved for
suspicions(), which is a measured NON-detection.

Docs gain worked one-line examples per tier and a "which tier fired?" section
showing real output. 46 self-test cases pass.
2026-08-27 13:45:08 +02:00
joakimp 3d47937d06 feed: scrub secrets before staging a transcript, name-anchored not entropy-anchored
A palace is mined from transcripts, and transcripts contain whatever the terminal
printed. Measured on this fleet: one leaked bearer token had reached 3 drawers,
13 feeder inbox files across all three devices, and 10 local files spanning 10
days — from an agent inspecting an env var while debugging. That frequency is the
premise: this is a pipeline problem, not a discipline problem.

WHERE. bin/mempalace_redact.py, called from mempalace-pi-session at the point the
staged transcript is written. That single hook covers both transports, because
local mode mines the staged file and remote mode rsyncs that same file
byte-for-byte. Scrubbing operates on the parsed objects rather than the
serialized text, so string values are rewritten while keys, ids and structure are
untouched — scanning raw JSONL instead invents keys like "tapiKey" out of the \t
escape preceding a field name (observed, not theorised).

WHY NOT ENTROPY. The obvious "redact long random-looking strings" is actively
destructive here: drawer ids, chunk ids, event ids, replica ids, HLCs and commit
SHAs are all high-entropy, are the majority of random-looking text in a palace,
and redacting them is silent and permanent. Detection is anchored on meaning
instead: tier 1 literal values from this process's env whose NAME says secret
(zero false positives by construction, and the only tier that can tell a 40-hex
Gitea PAT from a git commit sha); tier 2 vendor-prefixed shapes (ghp_, glpat-,
xox*-, sk-, AKIA, AIza, hf_, JWT, PEM blocks, URL credentials, Authorization
headers); tier 3 name=value assignments.

THE MEASUREMENT THAT CHANGED THE DESIGN. Tier 3 was going to redact. Against
52 MB of real fleet transcripts it produced 403 hits, and inspection with values
masked showed most were ${VAR} interpolation in compose files, TypeScript
identifiers, a TYPE ANNOTATION (credentials: Credentials), an IPA attribute
holding a date (krbPasswordExpiration), AAAK diary shorthand, and terminal output
following a "Password:" prompt. Redacting those corrupts code and docs held as
memory to catch what tier 1 already catches by value. After adding guards for
interpolation, code context, non-secret key suffixes and all-digit values, the
enforced count fell 403 -> 29 on the same corpus. So tier 3 REPORTS and does not
rewrite unless MEMPALACE_REDACT_STRICT=1.

HONESTY ABOUT MISSES. Known false negatives are documented rather than papered
over: novel formats in bare prose, another machine's secrets, base64-of-a-secret,
line-split secrets. Every run prints a count including "0 redaction(s)", because
silence is indistinguishable from a scrubber that never ran, and suspicions()
reports high-entropy strings it did NOT redact as (length, fingerprint) so the
miss rate is measurable. Findings never carry the value — rule, label, length and
sha256[:8], enough to recognise a recurrence, not enough to recover the secret.

FAIL CLOSED: no scrubber, no staging (exit 3), overridable with
MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 for a machine older than this file.

Tested: 42-case corpus in --self-test, including every palace id shape as a
must-not-redact case, idempotency, and a compound case the corpus caught where a
vendor placeholder was then re-matched by the URL rule (nested placeholder — the
secret was hidden either way, so only an exact-output assertion catches it).
End-to-end on 45 real sessions: 29 enforced, fail-closed verified at rc=3,
override verified loud. Server-side layer specified in docs/secret-hygiene.md §5
but NOT implemented — it is tier 2 only there, since the hub cannot see a
client's environment.
2026-08-27 13:30:47 +02:00
joakimp ecc2a9c574 mailbox notify: record that the terminal path is unverified through tmux
Measured on tor-ms22: the layering here is kitty -> tmux (on the host) ->
docker exec -> pi (in the container), with two clients attached to one tmux
session. Test OSC sequences written straight to pi's tty produced no
notification on the remote client; the local client is still unobserved. Most
likely tmux is dropping OSC types it does not implement — reaching the outer
terminal needs tmux's DCS passthrough (ESC P tmux; ... ESC \, inner ESC doubled)
plus allow-passthrough on, which is not the default and is not implemented here.

The point worth keeping is structural, not incidental: the container can see
NEITHER layer. KITTY_WINDOW_ID is absent because docker exec does not forward it,
and TMUX is absent because tmux runs one level further out on the host. So a
containerised client cannot detect the terminal it is speaking to or the
multiplexer it must speak through, and autodetection is not merely unreliable
here — it is blind. Explicit configuration is the only route, which is why the
previous commit added forced kitty/osc777 modes.

No behaviour or defaults changed: in-TUI notify remains the default and the
terminal path stays opt-in. Also flagged the unsettled routing question — a
pane's output reaches every attached client, so a work laptop and a home machine
would both ping from one arriving ask.
2026-08-27 12:54:52 +02:00
joakimp e91766286e mailbox notify: name the protocol, because docker exec hides the terminal
Auto-detection is useless in the deployment this ships in, measured on tor-ms22.
Terminal identity lives in env vars set by the emulator (KITTY_WINDOW_ID,
TERM_PROGRAM) and docker exec does not forward them — pi inside the container
sees only TERM=xterm-256color no matter what is rendering it. So the =desktop
detection can never see Kitty from in there and always falls through to OSC 777,
which Kitty does not implement. Result: on a containerised client the ping
silently does nothing, which is the worst available failure for a feature whose
entire purpose is to break a silence.

Adds MEMPALACE_MAILBOX_NOTIFY=kitty and =osc777 to force one protocol and skip
detection. =desktop keeps the notify.ts-style autodetect for the non-container
case, unset still means in-TUI only, 0/off still silent. The Windows toast branch
is dropped from the comment rather than the code path it never had here: it needs
powershell.exe on PATH, which is a WSL fact, not a Linux-container one.
2026-08-27 12:39:16 +02:00
joakimp a92c75d070 mailbox: say it is queued, and ping the human who is not looking (RFC 003 §7.11)
Two changes, neither touching the no-triggerTurn decision, which stands.

B — a delivery note in the message itself. The mid-session text explained how to
CLOSE an ask and never said when it would be SEEN, so the only reader who needed
that fact — the human watching an idle session — was the one not told. It now
says: this is a queued message, nothing woke the agent, any message starts the
turn that handles it, and the agent is not ignoring the ask, it is not running.
Costs nothing and changes no behaviour; it converts "why is it ignoring me?" into
"right, I nudge it". Deliberately NOT added to the wake-up injection, where a
turn is already starting and the note would be false.

A — a notification at poll time, because B only helps someone already looking and
the case that loses an ask is nobody looking. MEMPALACE_MAILBOX_NOTIFY: unset →
in-TUI ctx.ui.notify, the surface session_start already uses; =desktop →
additionally a terminal-native notification (Kitty OSC 99, else OSC 777), reusing
the detection the fleet's own notify.ts already proves in this harness; =0/off →
silent. The desktop path is how a ping escapes a container with no notify-send,
no DBus and no host access: the escape sequence is written to stdout and
interpreted by the terminal emulator on the human's own machine. Opt-in because
writing raw escapes is a behaviour change on a shared machine, not because it is
unreliable — say the word and the default flips.

Placement and firing conditions are deliberate: the notify call sits AFTER
sendMessage so a ping can never be the only thing that happened, and it fires
only when something is due — the same condition as delivery. A notification on
an empty poll would train its reader to ignore it, which is the failure this
whole feature exists to reverse. The wording names the nudge ("send any message
to handle") because "you have mail" without "press a key" reproduces exactly the
confusion B fixes.

Tested: 12 cases. Mode parsing (unset/empty/desktop/DESKTOP-with-space/0/off/1),
escape hygiene, and the OSC invariant that matters — a hostile title or body
containing ";" or a BEL cannot forge an OSC field or terminate the sequence
early, which is worth asserting because event bodies arrive from other machines.
ctx.hasUI is checked and the notify call is wrapped, since the UI can be gone by
the time an unawaited poll resolves. Syntax checked with node --strip-types.
2026-08-26 23:56:21 +02:00
joakimp 982b001001 docs: RFC 003 §7.11/§7.12 — delivered is not read, and the mailbox carries obligations not news
Both measured 2026-08-26 by peer devices, and the second one measured on me:
the operator had to point at an event id before I read a report that had been
addressed to me for two hours. The mailbox was working correctly the whole time.

§7.12 is the structural one. Candidates are drawn with status="open", so an
event carrying a TERMINAL status is not a mailbox candidate at all, whoever it
is addressed to. A task.reply written to a named device to share a finding is
therefore delivered to nobody, ever — and neither is any event.ack. So the most
natural inter-device message, "here is something you should know", is exactly
the shape that gets no delivery. Two things that do work: address it as a
directed status="open" ask (correct when a response is actually wanted), or
accept it as pull-only and pair it with a drawer, which the peer's search will
surface. A terminal report plus an expectation of attention does not.

§7.11 is the peer's finding, and it explains the other half of why that report
sat unread: delivery is deliverAs "steer" with deliberately no triggerTurn, and
the poll fires on agent_settled — i.e. when the agent is IDLE, with no inference
running. The text is queued for the next turn, so the human is the trigger.
Measured on another device: a delivery landed at ~20:50Z and sat visibly
unreacted-to until the operator asked "do I have to nudge you?". Not a defect;
the no-triggerTurn decision was deliberate and stands. But the delivered text
explains how to CLOSE an ask and never says when it will be SEEN, so the one
reader who needs that fact — the human watching the window — is the one not
told. Recorded with the three fixes that do not wake a model, as decision 8.

§7.4 is upgraded from inferred to measured, on two devices independently, by a
two-arm control: a to_agent='*' event with status='open' survives the raw
candidate filter and therefore reaches the exclusion branch, which no previously
written broadcast (all non-open) ever did. Raw query returned it, derived owed
set did not, delivery named only the directed arm. Two arms differing only in
to_agent is what makes it evidence rather than an absence.

And it carries a correction of my own record, in place and dated: I had filed
that this branch COULD NOT be exercised, because every broadcast in the log is
status='ready' and the protocol forbids writing an open one. Wrong in an
instructive way — the protocol forbids it as PRODUCTION TRAFFIC, which does not
forbid planting one as a labelled, self-closed control. Where a rule appears to
block a measurement, check whether it blocks the use or the test.
2026-08-26 23:38:41 +02:00
joakimp bfe9c5cd4f mailbox: join the owed set on hlc, not seq, before a second replica exists
deriveOwed decides "is this ask still owed?" by asking whether one of my own
terminal replies is LATER than the ask. It compared `seq` — this database's
arrival rowid. On a single hub that is global order, so it was correct; the
comment above it already said hlc was the durable key "once mesh_peers reports
actual peers". Making the switch now, while one replica means the two orderings
agree, costs nothing; making it later means changing the rule while two machines
already disagree about order.

The bug being pre-empted is specific: with a second replica the same event gets
a different `seq` in each database, because arrival order is not authorship
order. A reply authored after its ask can arrive first and take the lower seq;
the join then concludes "no later reply exists" and an already-answered ask
reappears as owed — permanently, on that machine.

isStrictlyAfter() prefers `hlc` when both events carry one and falls back to
`seq` otherwise (a server predating the field, or an un-backfilled row). hlc is
rendered fixed-width, <unix_ms:13 digits>-<counter:6 hex>-<replica_id>, so a
plain string comparison IS the causal comparison, with the replica id as final
tiebreak. Verified before writing the code that the field is actually on the
wire — event_list returns it per event (logstream.py:593) — because a fallback
that never fires would have made this a no-op dressed as a fix.

`created_at` stays rejected, and the reason is now written down where the
decision is: it is server-generated at second precision, so ties are routine,
and a tie can suppress an UNANSWERED ask. Both remaining failure modes are on
the noisy-but-visible side — an answered item resurfacing is annoying, an
unanswered ask going silent defeats the mailbox.

Tested: 18 cases against the extracted comparator — hlc later/earlier/equal,
same-ms counter ties in hex (0x10 vs 0x9, which is where a non-padded format
would break), cross-replica tiebreak, both mesh reorderings, every seq-fallback
path, and degenerate input (null seq, numeric hlc, empty strings, no keys) which
must never claim "answered". All pass. Real-data check on the positive-control
pair: seq 26/27 carry hlc ...3071857/...3574085, so the orderings agree today
and the switch is a no-op now and correct later.

The cursor keeps the opposite ordering ON PURPOSE — since_event_id is
local-arrival ordered so a tail consumer still sees late-arriving remote ops
whose hlc is older (hlc.py:19-21). RFC 003 §9.3 now says so explicitly, because
that asymmetry looks like a bug worth "fixing" and is not.
2026-08-26 23:30:09 +02:00
joakimp d2764bf78e docs: take the Phase 1 exposure record private, leave a moved-note
463 lines with 64 mentions of specific hosts, in a public repo: the primary and
tunnel hosts by name, the registrar/DNS step, the tunnel resource wiring, shared
token custody, per-machine flip dates, and a palace lineage naming three work
machines. Now in the private fleet repo (fleet-ops 093fb65); this file becomes a
moved-note in the shape docs/synlig-primary-runbook.md already established.

Git history keeps the old text, so this limits future exposure rather than
undoing it.

Unlike the primary-host runbook, this file was MIXED — and the stub says so
instead of quietly implying the toolkit still documents HTTP exposure. §1.1-§1.3
(why one shared fleet token rather than per-device proxy users, and the one place
per-device identity does exist), §2 (the bind trap and the Host/Origin pin) and
§3.7 (a client is flipped by three env vars that travel as a set) are reusable
mechanism now published nowhere else. Named in the stub so the extraction is
tracked debt rather than a silent loss, and named in the private copy too so
whoever extracts it can delete the duplicate.

Inbound references fixed rather than left pointing at content that moved: two
§3.8 pointers in extensions/pi/README.md are replaced by the instruction they
were pointing at (the palace path reported by mempalace_status must be the
remote host's — a half-flipped client looks healthy while reporting a local
path), and the bind-trap reference is replaced by the Host/Origin sentence
itself, so the extension README no longer depends on the moved file. contrib
also loses a hostname and a seeding date it did not need to make its point.
2026-08-26 23:30:09 +02:00
joakimp d4d8bb6109 docs: retention direction for the coordination log, and drop the operator's name from RFC 002 §5
RFC 003 §9.1 — retention is settled in direction (rotate old traffic out of
the way, logrotate-style, moved aside rather than destroyed) and the sketch
records the parts that are not obvious:

- Tier artifacts before events. Events are a few KB; artifacts are capped at
  4 MiB and stored in-row, so moving artifact CONTENT cold while keeping the
  kind/sha256/size/created_by stub reclaims nearly all the space and keeps the
  audit trail ("what was handed over, by whom, verified how") intact.
- If events rotate at all, the unit is the correlation_id THREAD whose latest
  event is terminal — never the row. Archiving an ask while leaving its reply
  (or the reverse) breaks the owed-set join, and both failure modes are bad:
  an ask that can never be cleared resurfaces as owed forever, or a reply is
  orphaned from what it answered.
- Never rotate an event that is still owed. Owed-ness is DERIVED at read time,
  so an unanswered ask is indistinguishable from a stale one except by that
  derivation — a purely time-based sweep would discard the live obligations of
  a machine that has merely been offline for a month, which is precisely the
  case this log exists to serve.
- Rotation invalidates held cursors: since_event_id RAISES on an unknown id
  (§7.5), so archiving an event a watcher holds as its resume point turns its
  next poll into an error. Either announce rotation ahead of live cursors, or
  teach the anchor lookup to fall back to created_at/hlc.
- Archive → verify (row counts, artifact sha256) → only then DELETE + VACUUM,
  with a --dry-run that reports in thread units.

RFC 002 §5: "Open decisions for ALC" → "Open decisions". RFC 001 never names
the operator anywhere; impersonal is the mature precedent and ALC was never a
real identifier in the first place (it is the AAAK spec's illustrative code for
"Alice", copy-forwarded into ~700 diary entries without verification).

Same edit also removes two device names and a hostname from §5.3, which is
host inventory and belongs in the private fleet repo, not a public one. The
mechanism it teaches — a palace in a Docker named volume dies on the next
container recreate, so census before flipping — is unchanged and is the part
that mattered.
2026-08-26 23:07:40 +02:00
joakimp e1cc7592a1 docs: write the RFC 003 that the code has been citing all along, plus an operator-facing fleet-memory guide
`logstream.py`'s module docstring is headed "Agent coordination event log for
MemPalace (RFC 003)" and enumerates five "Design constraints (RFC 003)".
Comments cite "RFC 003 phase 5", "RFC 003 suggested defaults" and "the first
RFC 003 dogfood". Every event/artifact tool description cites RFC 003. The
document has never existed — confirmed by searching this repo and the primary
host. So this is a retrospective spec: it transcribes what the implementation
already believes, and records what it does NOT do.

docs/rfc-003-coordination-log.md, verified line-by-line against mempalace
3.8.0 (every claim cites file.py:LINE, indexed in §10 for re-verification).
The parts that are not visible from the tool descriptions:

- No idempotency guard on event_append or put_artifact (§7.1). The replication
  path checks `id OR (origin_replica, origin_seq)` before applying; the client
  path checks nothing. So peer replay is safe and CLIENT RETRY IS NOT — a
  retried append forks a coordination thread into two ids. This makes the
  fleet's "a timeout is not a failure, verify before retrying" rule
  load-bearing rather than advisory.
- Coordination traffic is deliberately exempt from BOTH palace locks
  (§4) — _HTTP_LOCK_FREE_TOOLS and _PEER_WRITER_EXEMPT_TOOLS, each with its
  own rationale in-source. "One large mine blocks every client" is true of
  drawer writes and false of coordination writes.
- `mempalace sync` never touches the log, and no DELETE FROM events exists
  anywhere (§2, §9.1) — answering for the logstream a question RFC 001 §7.2
  left open, and making the log permanent and unbounded.
- from_agent is shape-validated and never authenticated; there is no read
  scoping at all (§6). The log authenticates the fleet, not the agent. That is
  simultaneously the security limitation and the only way to positive-control
  the mailbox.
- Three distinct orderings — seq (local arrival rowid), origin_seq (author's
  counter), hlc (fleet-wide, lexicographically sortable) — and origin_replica
  identifies the PALACE, not the writer, which is why from_agent/to_agent carry
  the whole distinction between machines (§3.2).
- GET /logstream/events does not exist and never did (§7.8). An earlier
  measurement saw it 404 and blamed the reverse proxy; that inference was right
  for /sync/* and /logstream/stream and wrong for this one. Corrected in
  extensions/pi/README.md §3 too, in place, dated.

docs/fleet-memory.md is the operator-facing companion the repo lacked
entirely: the front-door README had zero mentions of coordination, so the
channel was undiscoverable unless a message happened to arrive. It covers the
five stores and what each is for (drawers/wings/rooms, diaries, KG, palace
graph, coordination log), what a central palace buys a fleet — awareness,
non-repetition of expensive work, and retractions that travel — and a decision
flow for drawer vs KG vs event. Six mermaid diagrams, all rendered and
inspected as images, not merely parsed: the first pass produced a truncated
state label from an HTML entity and a self-loop that drew a meaningless dotted
lasso. Validation says "no syntax error"; only looking says "correct".

Also: a Documentation table in the top-level README so all of the above is
reachable from the front door.

Deliberately host-agnostic, per synlig-primary-runbook.md's precedent — no
device names, hostnames or operator names in either new document.
2026-08-26 22:55:55 +02:00
16 changed files with 3227 additions and 539 deletions
+28 -3
View File
@@ -58,6 +58,28 @@ Who owns what:
---
## Documentation
Start here if you are deciding **what to put where**, or running MemPalace on more than one machine:
| Document | What it answers |
|---|---|
| [`docs/fleet-memory.md`](docs/fleet-memory.md) | **Start here.** The five stores MemPalace keeps and what each is for; what a *central* palace buys a fleet of machines; when to file a memory versus when to message another device. Diagrams, worked examples. |
| [`docs/rfc-003-coordination-log.md`](docs/rfc-003-coordination-log.md) | The inter-agent / inter-device coordination log (`logstream`): storage, append and query semantics, delivery and latency, the trust model, and its landmines. |
| [`docs/secret-hygiene.md`](docs/secret-hygiene.md) | Keeping credentials out of the palace — what the feeder scrubs before staging, why detection is name-anchored and never entropy-anchored, and the measured false-positive rate that made one tier report-only |
| [`docs/rfc-001-global-palace.md`](docs/rfc-001-global-palace.md) | Why and how a palace is centralised, what should *not* be global, and the Phase-0 landmines (including: never run `mempalace sync` against a shared palace). |
| [`docs/rfc-002-joiner.md`](docs/rfc-002-joiner.md) | Replaying a second palace into a shared primary, and what dedupes what. |
| [`docs/phase-1-exposure-runbook.md`](docs/phase-1-exposure-runbook.md) | *(moved)* The exposure record for one deployment is now private operator data. The stub names the reusable mechanism it also carried — the `Host`/`Origin` pin above all — which is not yet published elsewhere. |
| [`docs/backup-and-recovery.md`](docs/backup-and-recovery.md) | Why a palace needs its own backup procedure, and how to restore one. |
| [`extensions/pi/README.md`](extensions/pi/README.md) | The pi-side client: provenance stamping at the edge, and the auto-delivered mailbox. |
Behaviour for the *agents* is normative in the `mempalace` skill (`SKILL.md` here,
installed to `~/.agents/skills/mempalace/`), not in these documents — the split is
deliberate: the skill says what an agent should do, the docs say what the machinery
actually does.
---
## Why this exists
MemPalace is the agent memory layer. Its stock CLI has two gaps that bite on a machine running opencode with a docs-first palace policy:
@@ -575,9 +597,12 @@ the palace host and asks the server to mine its own local copy. Requires
`MEMPALACE_PI_REMOTE_PATH`, and `MEMPALACE_PI_DEVICE` in `--help` for the
rest. `MEMPALACE_PI_DEVICE` also defaults `--agent` to `pi@<device>` so each
drawer's `added_by` records which harness and which machine produced it —
mempalace itself stores neither. Deploying that primary — newt, DNS, and why the
auth is a shared bearer token rather than per-device proxy users — is
[`docs/phase-1-exposure-runbook.md`](docs/phase-1-exposure-runbook.md).
mempalace itself stores neither. Deploying a primary — tunnel, DNS, token
custody — is deployment-specific operator data and lives in a private fleet
repository; [`docs/phase-1-exposure-runbook.md`](docs/phase-1-exposure-runbook.md)
records what moved and why. The *consequence* of choosing one shared fleet token
over per-device proxy users is public, in
[`docs/rfc-003-coordination-log.md`](docs/rfc-003-coordination-log.md) §6.
---
+102 -2
View File
@@ -534,11 +534,70 @@ fi
# Also classifies each export as NEW/ALREADY FILED (by source_file lookup)
# so --dry-run reports the real mine-set size. Classification is advisory;
# `mempalace mine --mode convos` is still the authoritative dedup.
# The redactor is a sibling module, imported by the heredoc below. Exported
# rather than passed as argv so the argv unpack stays stable.
#
# ${BASH_SOURCE[0]} reports the path the script was INVOKED as, and the image
# installs /usr/local/bin/mempalace-pi-session as a symlink into
# /opt/mempalace-toolkit/bin. A naive dirname therefore yields /usr/local/bin,
# where the redactor does not exist — and because the import is fail-closed, that
# turns every feeder tick on every device into "refusing to stage". Measured: the
# symlinked invocation exited FATAL while the direct one worked, i.e. it would
# have stopped the whole fleet's memory feed at the next image bake. So chase the
# symlink chain, and offer fallbacks rather than betting on one answer.
#
# readlink -f is avoided deliberately: it is GNU/newer-BSD only, and this script
# also runs directly on macOS hosts.
_mp_resolve() {
local p="$1" target
while [ -L "$p" ]; do
target="$(readlink "$p")" || break
case "$target" in
/*) p="$target" ;;
*) p="$(dirname "$p")/$target" ;;
esac
done
printf '%s' "$p"
}
_mp_self="$(_mp_resolve "${BASH_SOURCE[0]}")"
MEMPALACE_REDACT_DIR="$(cd "$(dirname "$_mp_self")" && pwd):$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd):/opt/mempalace-toolkit/bin"
export MEMPALACE_REDACT_DIR
export_count=$(python3 - "$PI_SESSIONS_DIR" "$STAGE" "$SESSION_ID" "$SINCE" "$MIN_MESSAGES" "$MIN_ASSISTANT_CHARS" "$MODE" <<'PY'
import json, os, sqlite3, sys
from datetime import datetime, timezone
from pathlib import Path
# ── Secret scrubbing before anything is staged ───────────────────────────────
# This is the last point at which a transcript is a plain in-memory value: the
# local mine reads the staged file, and the remote path rsyncs that same file
# byte-for-byte, so scrubbing here covers BOTH transports with one hook.
#
# FAIL CLOSED. If the redactor cannot be imported, staging is refused rather
# than done unscrubbed — a missing module means a broken install, and the whole
# point of this step is that a secret must not reach a shared palace. Override
# deliberately with MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 if you ever need to feed a
# machine whose toolkit is older than this file.
# Colon-separated candidates: symlink-resolved dir, invoked dir, then the image's
# known install path. First one that has the module wins.
for _cand in os.environ.get("MEMPALACE_REDACT_DIR", "").split(":"):
if _cand and os.path.isdir(_cand):
sys.path.insert(0, _cand)
try:
import mempalace_redact as _redact
except Exception as _e: # noqa: BLE001 - any import failure is fatal by design
if os.environ.get("MEMPALACE_FEED_ALLOW_UNSCRUBBED", "").strip() in {"1", "true", "yes"}:
_redact = None
print(" [WARN] secret scrubber unavailable, staging UNSCRUBBED by request "
f"({_e})", file=sys.stderr)
else:
print(f" [FATAL] secret scrubber unavailable ({_e}); refusing to stage. "
"Set MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 to override.", file=sys.stderr)
raise SystemExit(3)
_known = _redact.env_secrets() if _redact else []
_redactions = 0
_reported = 0
sessions_dir, stage, session_filter, since, min_messages, min_assistant_chars, mode = sys.argv[1:8]
min_messages = int(min_messages)
min_assistant_chars = int(min_assistant_chars)
@@ -811,6 +870,29 @@ for path in paths:
)
continue
# Scrub the parsed objects, not the serialized text: string VALUES get
# rewritten while keys, ids and structure are left exactly as they are.
# (Scanning raw JSONL instead would also match escape artifacts like the
# "\\t" before a field name, inventing keys such as "tapiKey".)
if _redact is not None:
_found: list = []
out_lines = [_redact.scrub_obj(obj, _known, _found)[0] for obj in out_lines]
_hard = [x for x in _found if not x.rule.endswith("-reported")]
_soft = [x for x in _found if x.rule.endswith("-reported")]
_redactions += len(_hard)
_reported += len(_soft)
if _hard:
_by = {}
for x in _hard:
# "T2:github-pat" — rule says what matched, tier says how much to
# trust it, which is what the reader of this line actually needs.
_k = f"T{x.tier}:{x.rule}"
_by[_k] = _by.get(_k, 0) + 1
print(f" [REDACTED] {path.name} "
+ ", ".join(f"{k}={v}" for k, v in sorted(_by.items()))
+ " fp=" + ",".join(sorted({x.fingerprint for x in _hard})),
file=sys.stderr)
out_path = stage / f"pi_{session_uuid}.jsonl"
with out_path.open("w", encoding="utf-8") as f:
for obj in out_lines:
@@ -835,6 +917,13 @@ for path in paths:
print(f"EXPORTED {exported}")
print(f"ALREADY_FILED {-1 if already_filed is None else skipped_already_filed}")
# Report the scrub outcome even when it is zero: "0 redactions" is a measurement,
# whereas printing nothing is indistinguishable from a scrubber that never ran.
if _redact is not None:
print(f" [scrub] {_redactions} redaction(s) applied, "
f"{_reported} name-anchored candidate(s) reported only "
f"(set MEMPALACE_REDACT_STRICT=1 to redact those too)", file=sys.stderr)
if skipped_short:
print(f"SKIPPED_SHORT {skipped_short}", file=sys.stderr)
if skipped_quiet:
@@ -885,15 +974,26 @@ fi
# ── Ship to the palace host (remote mode only) ───────────────────────
# mempalace_mine expands its source path in the SERVER process, so in remote
# mode the exports have to physically exist over there. rsync --update is the
# mode the exports have to physically exist over there. --checksum is the
# idempotent half; the mine is the other half.
#
# NOT --update: the stage file's mtime is deliberately the SOURCE transcript's
# mtime (see the os.utime() in the exporter, "preserve session mtime for dedup
# stability"), so a re-export of a session that has not been appended to since
# the last ship carries an mtime that is NOT newer than the receiver copy. With
# --update rsync then SKIPS it silently -- which is exactly the case that must
# ship after a redactor change, because the content differs while the mtime does
# not. Measured on mbp-m1-2020 2026-08-27: a scrubbed re-export of a dormant
# session was skipped and unscrubbed bytes stayed in the palace host's inbox,
# while every local signal reported a clean stage. --checksum compares content
# and keeps the ship idempotent without trusting timestamps.
MINE_SOURCE="$STAGE"
if [[ "$MODE" == "remote" ]]; then
ssh_cmd="ssh"
[[ -n "$SSH_CONFIG" ]] && ssh_cmd="ssh -F $SSH_CONFIG"
echo ""
echo "Shipping stage to ${SSH_TARGET%/}/$DEVICE/ ..."
if ! rsync -a --update --no-owner --no-group \
if ! rsync -a --checksum --no-owner --no-group \
-e "$ssh_cmd" \
--include='*.jsonl' --exclude='*' \
"$STAGE/" "${SSH_TARGET%/}/$DEVICE/"; then
+604
View File
@@ -0,0 +1,604 @@
#!/usr/bin/env python3
"""Scrub secret-shaped values out of palace-bound text.
WHY THIS IS NAME-ANCHORED AND NOT ENTROPY-ANCHORED
==================================================
The obvious design — "redact long random-looking strings" — is actively wrong
for this corpus, and the reason is worth stating before anyone tries to
"improve" it. A palace is *full* of high-entropy strings that are its own
primary keys:
drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b drawer id
..._chunk_000000 chunk id
evt_20260826T213924_397b1f710c2d event id
rep_d344e349ba276d6fc11997cd552f6937 replica id
1787780364792-000000-rep_d344e349ba276d6f… hybrid logical clock
ecc2a9c574f4156403cf4e90aea8b0d088c75700 git commit sha
An entropy or "looks like base64/hex of length N" detector fires on every one
of those, i.e. on most identifiers in the corpus, and the redaction damages the
memory it was meant to protect. Worse, the damage is silent and unrecoverable:
a mined transcript whose ids have been replaced with <redacted> is no longer
traceable to anything.
So detection here is anchored on one of three things that carry *meaning*:
Tier 1 KNOWN VALUES — literal values taken from this process's own
environment, for variables whose NAME says secret.
Zero false positives by construction: the value
*is* the secret. Catches any presentation (env
dump, JSON, error message, URL, prose), which is
exactly the incident this exists for.
Tier 2 KNOWN SHAPES — vendor-prefixed credentials (ghp_…, glpat-…,
xox[abprs]-…, AKIA…, sk-ant-…, JWTs, PEM private
keys, credentials embedded in URLs, Authorization
headers). Very low false-positive rate because the
prefix is meaningful, not merely random.
Tier 3 NAME=VALUE — an assignment whose *key* says secret
(…TOKEN, …SECRET, …PASSWORD, …API_KEY, …).
REPORT-ONLY BY DEFAULT. Measured on 52 MB of real
fleet transcripts it fired 403 times, of which the
large majority were ${VAR} interpolation in compose
files, TypeScript identifiers, a type annotation
(`credentials: Credentials`), an IPA attribute
holding a date (krbPasswordExpiration) and terminal
output following a "Password:" prompt. Rewriting
those corrupts code and docs held as memory, to
catch what Tier 1 already catches by value.
Set MEMPALACE_REDACT_STRICT=1 to make it enforce.
SHAPE COLLISIONS, AND WHY TIER 1 IS THE LOAD-BEARING ONE
--------------------------------------------------------
A 40-hex Gitea PAT is byte-indistinguishable from a git commit sha, so no shape
rule can separate them. Only two things can: the value being known (Tier 1,
which redacts) or a key naming it (Tier 3, which by default only reports). So
in the default configuration a sha-shaped PAT is caught if and only if it
belongs to this machine. That is an accepted, documented gap — the alternative
is redacting every commit sha in the palace.
WHICH TIER DID THAT? — mapping printed rule names back to tiers
--------------------------------------------------------------
The output names the *rule* (`github-pat`, `env-value`), because that says what
matched. RULE_TIERS below maps rule -> tier, and `tier_of()` is what the feeder
uses to print `T2:github-pat=6`, because the tier is what tells a reader how
much to trust the hit. A self-test asserts every rule that can appear in a
Finding has a tier, so adding a rule without classifying it fails the tests.
KNOWN FALSE NEGATIVES — stated, not hidden
------------------------------------------
This will not catch: a novel credential format pasted bare into prose with no
name nearby; a secret from a machine whose env this process cannot see; a
base64-of-a-secret; a secret split across lines. Tier 1 covers the local
machine's own secrets completely, which is where the measured incidents came
from — but "the scrubber ran" must never be read as "there are no secrets in
here". `suspicions()` exists for exactly that reason: it reports high-entropy
candidates it did NOT redact, as fingerprints rather than values, so the false
negative rate can be *measured* over time instead of assumed to be zero.
Run `python3 mempalace_redact.py --self-test` for the test corpus, which
includes every benign id shape above as a must-not-redact case.
"""
from __future__ import annotations
import hashlib
import math
import os
import re
from typing import Any, Callable, Iterable
__all__ = ["scrub", "scrub_obj", "suspicions", "Finding"]
PLACEHOLDER = "<redacted:{label}>"
# ── Tier 1: names whose VALUES are secrets ────────────────────────────────────
# Anchored on the variable name, then applied as an exact literal match.
_SECRET_NAME = re.compile(
r"(TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|API_?KEY|APIKEY|ACCESS_?KEY"
r"|PRIVATE_?KEY|CREDENTIAL|BEARER|SESSION_?KEY|COOKIE|_PAT|^PAT$)",
re.IGNORECASE,
)
# …unless the name says it holds a location or a knob rather than a secret.
# SSH_KEY_PATH is the live example: it matches KEY, and its value is a path.
_NAME_NOT_SECRET = re.compile(
r"(_PATH|_FILE|_DIR|_NAME|_ID|_URL|_URI|_HOST|_PORT|_USER|_ENABLED"
r"|_TIMEOUT|_MS|_SECS|_SECONDS|_COUNT|_LIMIT|_MODE)$",
re.IGNORECASE,
)
# Keys that contain a secret word but describe a POLICY, COUNT or TYPE rather
# than holding a credential. Measured on real transcripts: krbPasswordExpiration
# (an IPA attribute whose value is a date), observationTokens (a token count),
# targetTokens, tokensOverTarget.
_KEY_NOT_A_SECRET = re.compile(
r"(expiration|expiry|_age|count|length|limit|usage|interval|policy|type|class"
r"|provider|manager|error|target|overtarget|file|path|dir|name|_id|url|uri"
r"|host|port|user|enabled|timeout|_ms|_secs)$",
re.IGNORECASE,
)
# Left of the match: a declaration keyword or a member access means this is CODE
# (`const token = x`, `obj.apiKey = y`), and the "value" is an expression.
_CODE_CONTEXT = re.compile(
r"(?:\b(?:const|let|var|def|func|fn|public|private|readonly|export|return|if|assert)\s+$)"
r"|[.\w]\s*$"
)
_TRIVIAL_VALUES = {
"", "0", "1", "true", "false", "none", "null", "nil", "unset", "changeme",
"change-me", "xxx", "***", "redacted", "your-token-here", "your_token_here",
"placeholder", "example", "dummy", "test", "secret",
}
# ── Tier 2: shapes that mean credential because the PREFIX means credential ──
_SHAPE_RULES: list[tuple[str, re.Pattern[str], Callable[[re.Match[str]], str]]] = [
# ORDER MATTERS, and the self-test is what proved it. The two structural
# rules (credentials inside a URL, Authorization header) run FIRST, before
# the vendor-prefix rules: a vendor rule rewriting `https://ghp_xxx@host` to
# `https://<redacted:github-token>@host` leaves a colon INSIDE the
# placeholder, which the URL rule then reads as user:password and redacts a
# second time, producing a nested placeholder. Ordering plus the (?!<redacted)
# guards below make the result independent of pass order.
("url-credentials",
re.compile(r"(?P<pre>[a-zA-Z][a-zA-Z0-9+.\-]*://(?!<redacted)[^/\s:@]+:)(?P<pw>(?!<redacted)[^/\s@]{4,})(?P<at>@)"),
lambda m: f"{m.group('pre')}{PLACEHOLDER.format(label='url-password')}{m.group('at')}"),
("authorization-header",
re.compile(r"(?P<pre>[Aa]uthorization:\s*(?:Bearer|Basic|token)\s+)(?P<val>(?!<redacted)[A-Za-z0-9._\-+/=]{8,})"),
lambda m: f"{m.group('pre')}{PLACEHOLDER.format(label='auth-header')}"),
("github-pat", re.compile(r"\bgh[pousr]_[A-Za-z0-9]{16,}\b"),
lambda m: PLACEHOLDER.format(label="github-token")),
("github-fine-grained", re.compile(r"\bgithub_pat_[A-Za-z0-9_]{22,}\b"),
lambda m: PLACEHOLDER.format(label="github-token")),
("gitlab-pat", re.compile(r"\bglpat-[A-Za-z0-9_\-]{16,}\b"),
lambda m: PLACEHOLDER.format(label="gitlab-token")),
("slack", re.compile(r"\bxox[abprs]-[A-Za-z0-9\-]{10,}\b"),
lambda m: PLACEHOLDER.format(label="slack-token")),
("openai-anthropic", re.compile(r"\bsk-(?:ant-)?[A-Za-z0-9_\-]{20,}\b"),
lambda m: PLACEHOLDER.format(label="api-key")),
("aws-access-key", re.compile(r"\b(?:AKIA|ASIA)[0-9A-Z]{16}\b"),
lambda m: PLACEHOLDER.format(label="aws-access-key-id")),
("google-api-key", re.compile(r"\bAIza[0-9A-Za-z_\-]{35}\b"),
lambda m: PLACEHOLDER.format(label="google-api-key")),
("huggingface", re.compile(r"\bhf_[A-Za-z0-9]{30,}\b"),
lambda m: PLACEHOLDER.format(label="hf-token")),
("npm", re.compile(r"\bnpm_[A-Za-z0-9]{30,}\b"),
lambda m: PLACEHOLDER.format(label="npm-token")),
("digitalocean", re.compile(r"\bdop_v1_[a-f0-9]{60,}\b"),
lambda m: PLACEHOLDER.format(label="do-token")),
# JWT: three base64url segments. Requires the "eyJ" header so a random
# dotted string cannot trip it.
("jwt", re.compile(r"\beyJ[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}\b"),
lambda m: PLACEHOLDER.format(label="jwt")),
# Whole PEM block, non-greedy so several in one file stay separate.
("pem-private-key",
re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----.*?-----END [A-Z ]*PRIVATE KEY-----", re.DOTALL),
lambda m: PLACEHOLDER.format(label="private-key-block")),
]
# ── Tier 3: assignment whose KEY says secret ─────────────────────────────────
# The key prefix is LAZY AND OPTIONAL, not "one char then anything": a mandatory
# leading character cannot match a key that BEGINS with the secret word, so
# `"api_key"` was missed while `"client_secret"` passed — the secret word sits
# mid-name in one and at position 0 in the other. Note plain `KEY` is absent from
# the alternation on purpose: it is too generic (SSH_KEY_PATH), and tier 1 covers
# the real ones by value anyway.
# The optional quote after the key is what makes JSON work: `"api_key": "v"`
# closes the key BEFORE the colon, so a pattern jumping straight from key to
# separator silently misses every JSON payload — i.e. most of a transcript.
_ASSIGNMENT = re.compile(
r"""(?P<key>\b[A-Za-z0-9_.\-]*?
(?:token|secret|password|passwd|passphrase|api[_-]?key|apikey
|access[_-]?key|private[_-]?key|credential)
[A-Za-z0-9_.\-]*)
(?P<kq>["']?)
(?P<sep>\s*(?:=>|[=:])\s*)
(?P<q>["']?)
(?P<val>[^\s"'<>,;)\]}]{12,})
(?P=q)""",
re.IGNORECASE | re.VERBOSE,
)
# A value that is obviously not a live secret: a reference, a path, a
# placeholder, or something we already redacted.
_NOT_A_SECRET_VALUE = re.compile(
r"""^(?:
\$\{[^}\s]*\}? # ${VAR} ${VAR:-default} ${VAR:?err}
# trailing } optional: the value
# charset stops before it, so the
# captured text is "${VAR" — which
# still means "a reference", and
# missing that is what made a
# compose file look like a secret
| \$[A-Za-z_][A-Za-z0-9_]* # $VAR
| %[A-Za-z_][A-Za-z0-9_]*%? # %VAR%
| \{\{[^}\s]*\}?\}? # {{ template }}
| <[^>]*> # <redacted:…> / <your-token>
| (?:/|~/|\./|\.\./)[^\s]* # a path
| [A-Za-z]:\\ # a windows path
| \*+ | x+ | \.+ # ***, xxx
| (?:true|false|none|null|nil|unset|changeme|placeholder|example)
)$""",
re.IGNORECASE | re.VERBOSE,
)
class Finding:
"""One redaction, recorded without ever carrying the secret itself."""
__slots__ = ("rule", "label", "length", "fingerprint")
def __init__(self, rule: str, label: str, value: str) -> None:
self.rule = rule
self.label = label
self.length = len(value)
# Stable across sessions, so a recurring leak is recognisable as the
# same one, while the value stays unrecoverable from the report.
self.fingerprint = hashlib.sha256(value.encode("utf-8")).hexdigest()[:8]
@property
def tier(self) -> int:
"""Which detection tier produced this. See RULE_TIERS for the mapping."""
return tier_of(self.rule)
def __repr__(self) -> str: # pragma: no cover - diagnostics only
return (f"<Finding T{self.tier}:{self.rule}:{self.label} "
f"len={self.length} fp={self.fingerprint}>")
def env_secrets(environ: dict[str, str] | None = None) -> list[tuple[str, str]]:
"""Tier 1 material: (name, value) pairs from the environment worth scrubbing.
Deliberately conservative: a secret-sounding NAME is not enough, the value
must also be long enough to be a credential and must not look like a path
or a knob. Longest values first so that if one secret contains another as a
substring, the longer one is replaced before the shorter can fragment it.
"""
env = os.environ if environ is None else environ
out: list[tuple[str, str]] = []
for name, value in env.items():
if not value or len(value) < 12:
continue
if not _SECRET_NAME.search(name) or _NAME_NOT_SECRET.search(name):
continue
if value.strip().lower() in _TRIVIAL_VALUES:
continue
if _NOT_A_SECRET_VALUE.match(value.strip()):
continue
if "/" in value and (value.startswith("/") or value.startswith("~")):
continue # a path that happened to be named …_KEY
if any(c.isspace() for c in value):
continue # credentials do not contain whitespace; prose does
out.append((name, value))
out.sort(key=lambda kv: len(kv[1]), reverse=True)
return out
# Rule name -> tier. Printed output names the rule (WHAT matched); the tier says
# HOW MUCH TO TRUST IT: 1 = near-certain (the value is known to be a secret),
# 2 = strong (a vendor prefix means what it says), 3 = candidate, needs eyes.
RULE_TIERS: dict[str, int] = {
"env-value": 1,
**{name: 2 for name, _pat, _lbl in _SHAPE_RULES},
"named-assignment": 3,
"named-assignment-reported": 3,
# Tier 0 is not a detection: it is a measured NON-detection, reported so the
# false-negative rate is a number instead of an assumption.
"suspicion": 0,
}
def tier_of(rule: str) -> int:
"""Tier for a rule name, or -1 if the rule was never classified."""
return RULE_TIERS.get(rule, -1)
def scrub(
text: str,
known: Iterable[tuple[str, str]] | None = None,
findings: list[Finding] | None = None,
strict: bool | None = None,
) -> tuple[str, list[Finding]]:
"""Redact secrets from one string. Returns (scrubbed, findings).
Idempotent: the replacement text matches none of the rules, so running this
twice changes nothing and cannot nest placeholders.
"""
found: list[Finding] = findings if findings is not None else []
if strict is None:
strict = os.environ.get("MEMPALACE_REDACT_STRICT", "").strip() in {"1", "true", "yes"}
if not text:
return text, found
# Tier 1 first: an exact known value should be labelled with its variable
# name (the most useful thing a reader can be told) rather than by shape.
for name, value in (env_secrets() if known is None else known):
if value and value in text:
count = text.count(value)
text = text.replace(value, PLACEHOLDER.format(label=name))
for _ in range(count):
found.append(Finding("env-value", name, value))
# Tier 2.
for rule, pattern, repl in _SHAPE_RULES:
def _sub(m: re.Match[str], _rule: str = rule) -> str:
found.append(Finding(_rule, _rule, m.group(0)))
return repl(m)
text = pattern.sub(_sub, text)
# Tier 3 — REPORT-ONLY unless strict. See STRICT note in the module docstring:
# measured on 52 MB of real fleet transcripts this rule produced 403 hits of
# which the overwhelming majority were ${VAR} interpolation, TypeScript
# identifiers, YAML env lists and documentation prose. Redacting those would
# corrupt code and docs stored as memory, to catch secrets that tier 1
# already catches by value. So by default a tier-3 hit is REPORTED (so the
# miss can be measured and the rule calibrated) and NOT rewritten.
def _sub_assignment(m: re.Match[str]) -> str:
val, key = m.group("val"), m.group("key")
pre = m.string[max(0, m.start() - 24):m.start()]
if (_NOT_A_SECRET_VALUE.match(val)
or val.strip().lower() in _TRIVIAL_VALUES
or _KEY_NOT_A_SECRET.search(key)
or _CODE_CONTEXT.search(pre)
or val.isdigit()):
return m.group(0)
found.append(Finding("named-assignment" if strict else "named-assignment-reported",
key, val))
if not strict:
return m.group(0)
return f"{m.group('key')}{m.group('kq')}{m.group('sep')}{m.group('q')}" \
f"{PLACEHOLDER.format(label=m.group('key'))}{m.group('q')}"
text = _ASSIGNMENT.sub(_sub_assignment, text)
return text, found
def scrub_obj(obj: Any, known: Iterable[tuple[str, str]] | None = None,
findings: list[Finding] | None = None,
strict: bool | None = None) -> tuple[Any, list[Finding]]:
"""Recursively scrub every string VALUE in a JSON-ish structure.
Keys are left alone: a key named "token" is metadata, not a credential, and
rewriting keys would corrupt the transcript schema.
"""
found: list[Finding] = findings if findings is not None else []
known = list(env_secrets() if known is None else known)
if isinstance(obj, str):
return scrub(obj, known, found, strict)[0], found
if isinstance(obj, dict):
return {k: scrub_obj(v, known, found, strict)[0] for k, v in obj.items()}, found
if isinstance(obj, list):
return [scrub_obj(v, known, found, strict)[0] for v in obj], found
return obj, found
# ── Measuring what we MISS, without leaking it ───────────────────────────────
_BENIGN_ID = re.compile(
r"""^(?:
drawer_[\w.\-]+ # drawer / chunk ids
| (?:evt|art)_\d{8}T\d{6}_[0-9a-f]{12} # event / artifact ids
| rep_(?:[0-9a-f]{12}|[0-9a-f]{32}) # replica ids
| \d{13}-[0-9a-f]{6}-rep_[0-9a-f]+ # hybrid logical clocks
| [0-9a-f]{7,12} # short git sha
| [0-9a-f]{40} # git sha / sha1
| [0-9a-f]{64} # sha256
| [0-9a-f]{8}(?:-[0-9a-f]{4}){3}-[0-9a-f]{12} # uuid
| sha256:[0-9a-f]{64}
| \d+
)$""",
re.VERBOSE | re.IGNORECASE,
)
_CANDIDATE = re.compile(r"[A-Za-z0-9+/_\-=]{24,}")
def _entropy(s: str) -> float:
if not s:
return 0.0
counts: dict[str, int] = {}
for ch in s:
counts[ch] = counts.get(ch, 0) + 1
n = len(s)
return -sum((c / n) * math.log2(c / n) for c in counts.values())
def suspicions(text: str, min_entropy: float = 3.6) -> list[Finding]:
"""High-entropy strings this module did NOT redact, as fingerprints.
This is the false-negative meter. It deliberately does not redact anything:
every one of these is far more likely to be one of the palace's own ids than
a credential, and redacting them would be the damage described at the top of
this file. Reporting them as (length, fingerprint) makes the miss rate
observable — and a fingerprint that keeps recurring is worth a human look.
"""
out: list[Finding] = []
for m in _CANDIDATE.finditer(text):
val = m.group(0)
if _BENIGN_ID.match(val) or val.startswith("<redacted:"):
continue
if _entropy(val) < min_entropy:
continue
out.append(Finding("suspicion", "high-entropy", val))
return out
# ── Self-test ────────────────────────────────────────────────────────────────
def _self_test() -> int:
fake_token = "pFrDBfakVKQB0SLN1hTzAoaRJeQ4J_go7xmJezojAwI" # shape-alike, not real
env = {
"MEMPALACE_REMOTE_TOKEN": fake_token,
"GITEA_TOKEN": "0123456789abcdef0123456789abcdef01234567", # 40 hex: sha-shaped
"SSH_KEY_PATH": "/Users/someone/.ssh", # name matches KEY, value is a path
"MEMPALACE_MAILBOX_POLL_MS": "300000", # knob, not a secret
"MEMPALACE_REMOTE_URL": "https://palace.example.com/mcp",
"PASSWORD": "changeme", # trivial value
"API_KEY": "${FROM_VAULT}", # a reference
}
known = env_secrets(env)
must_redact = [
("env dump line", f"MEMPALACE_REMOTE_TOKEN={fake_token}"),
("bare value in prose", f"the token is {fake_token} apparently"),
("value inside json", f'{{"env": {{"MEMPALACE_REMOTE_TOKEN": "{fake_token}"}}}}'),
("sha-shaped pat, name-anchored",
"GITEA_TOKEN=0123456789abcdef0123456789abcdef01234567"),
("github pat", "remote add origin https://ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345@git"),
("gitlab pat", "GLPAT is glpat-AbCdEfGhIjKlMnOpQrSt"),
("slack", "xoxb-1234567890-ABCdefGHIjklMNO"),
("anthropic-ish", "sk-ant-api03-AbCdEfGhIjKlMnOpQrStUvWx"),
("aws", "AKIAIOSFODNN7EXAMPLE"),
("jwt", "eyJhbGciOiJIUzI1NiJ9.eyJzdWIiOiIxMjM0NTY3ODkwIn0.dozjgNryP4J3jVmNHl0w5N_XgL0n3I9PlFUP0THsR8U"),
("pem", "-----BEGIN OPENSSH PRIVATE KEY-----\nb3BlbnNzaA\n-----END OPENSSH PRIVATE KEY-----"),
("url creds", "git clone https://joakim:hunter2hunter2@git.example.com/x.git"),
("auth header", "Authorization: Bearer abcdefghijklmnop"),
("json api_key", '"api_key": "AbCdEfGhIjKlMnOpQrSt"'),
("json nested quotes", '{"auth": {"client_secret": "AbCdEfGhIjKlMnOpQrStUv"}}'),
("yaml style", " gitea_token: 0123456789abcdefghij"),
("lowercase password kv", "db_password = s3cr3t-p4ssw0rd-xyz"),
]
must_not_redact = [
("drawer id", "drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b"),
("chunk id", "drawer_pi-devbox_landmines_ca8929fea82cf80e7770ea0d_chunk_000000"),
("event id", "evt_20260826T213924_397b1f710c2d"),
("replica id", "rep_d344e349ba276d6fc11997cd552f6937"),
("hlc", "1787780364792-000000-rep_d344e349ba276d6fc11997cd552f6937"),
("git sha", "ecc2a9c574f4156403cf4e90aea8b0d088c75700"),
("short sha", "ecc2a9c"),
("sha256", "a" * 64),
("uuid session file", "pi_01a03fd8-0dc6-73c2-a9c3-6b62ccf0318f.jsonl"),
("knob assignment", "MEMPALACE_MAILBOX_POLL_MS=300000"),
("path named key", "SSH_KEY_PATH=/Users/someone/.ssh"),
("empty token", "MEMPALACE_REMOTE_TOKEN="),
("already redacted", "MEMPALACE_REMOTE_TOKEN=<redacted:MEMPALACE_REMOTE_TOKEN>"),
("var reference", "MEMPALACE_REMOTE_TOKEN=${VAULT_TOKEN}"),
("placeholder", "MEMPALACE_REMOTE_TOKEN=<your-token-here>"),
("trivial", "PASSWORD=changeme"),
("url no creds", "MEMPALACE_REMOTE_URL=https://palace.example.com/mcp"),
("prose", "The mailbox is an obligation channel, not a news channel."),
("base64 in a diff line", "+ data = 'QUJDREVGR0hJSktMTU5PUFFSU1RVVldYWVo='"),
]
failures = 0
# Tier 3 is report-only by default, so the corpus below is checked in STRICT
# mode (where tier 3 rewrites) and the default mode is asserted separately.
print("MUST REDACT (strict: tier 3 enforcing)")
for name, sample in must_redact:
out, found = scrub(sample, known, strict=True)
ok = bool(found) and all(
secret not in out for secret in (fake_token, "hunter2hunter2",
"0123456789abcdef0123456789abcdef01234567")
)
failures += not ok
rules = ",".join(sorted({f.rule for f in found})) or "NONE"
print(f" {'pass' if ok else 'FAIL'} {name:32s} [{rules}]")
print("MUST NOT REDACT (strict: the hard cases)")
for name, sample in must_not_redact:
out, found = scrub(sample, known, strict=True)
ok = out == sample and not found
failures += not ok
detail = "" if ok else f" -> {out!r} {found}"
print(f" {'pass' if ok else 'FAIL'} {name:32s}{detail}")
print("DEFAULT MODE (tier 3 reports, does not rewrite)")
for name, sample, expect_rewrite in [
("env value still redacted", f"MEMPALACE_REMOTE_TOKEN={fake_token}", True),
("vendor shape still redacted", "ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345", True),
("name-anchored only reported", "db_password = s3cr3t-p4ssw0rd-xyz", False),
("compose interpolation untouched", "GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}", False),
("typescript identifier untouched", "const tokens = countTokensForModel", False),
("ipa attribute untouched", "--setattr=krbPasswordExpiration=20260529090000", False),
]:
out, found = scrub(sample, known)
rewritten = out != sample
ok = rewritten == expect_rewrite
if name == "name-anchored only reported":
ok = ok and any(f.rule == "named-assignment-reported" for f in found)
if name.endswith("untouched"):
ok = ok and not any(f.rule.startswith("named-assignment") for f in found)
failures += not ok
print(f" {'pass' if ok else 'FAIL'} {name:34s}"
+ ("" if ok else f" -> rewritten={rewritten} {found}"))
print("TIER MAPPING")
_rules = ({"env-value", "named-assignment", "named-assignment-reported", "suspicion"}
| {n for n, _p, _l in _SHAPE_RULES})
_unclassified = sorted(r for r in _rules if tier_of(r) < 0)
failures += bool(_unclassified)
print(f" {'pass' if not _unclassified else 'FAIL'} every rule maps to a tier"
+ (f" -> unclassified: {_unclassified}" if _unclassified
else f" ({len(_rules)} rules)"))
for label, sample, want in [
("env value is tier 1", f"MEMPALACE_REMOTE_TOKEN={fake_token}", 1),
("vendor shape is tier 2", "ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345", 2),
("name-anchored is tier 3", "db_password = s3cr3t-p4ssw0rd-xyz", 3),
]:
got = scrub(sample, known)[1]
ok = bool(got) and all(f.tier == want for f in got)
failures += not ok
print(f" {'pass' if ok else 'FAIL'} {label:26s}"
+ ("" if ok else f" -> {[(f.rule, f.tier) for f in got]}"))
print("PROPERTIES")
once, _ = scrub(f"MEMPALACE_REMOTE_TOKEN={fake_token}", known, strict=True)
twice, second = scrub(once, known, strict=True)
ok = once == twice and not second
failures += not ok
print(f" {'pass' if ok else 'FAIL'} idempotent (no nested placeholders)")
# The compound case the corpus caught: assert on EXACT output, not merely
# that the secret is gone — a nested placeholder also hides the secret, and
# would have shipped unnoticed.
compound, _ = scrub("https://ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345@git.example.com/x", known, strict=True)
ok = compound == "https://<redacted:github-token>@git.example.com/x"
failures += not ok
print(f" {'pass' if ok else 'FAIL'} vendor placeholder not re-redacted by url rule"
+ ("" if ok else f" -> {compound!r}"))
already, found_a = scrub("https://joakim:<redacted:url-password>@git.example.com", known, strict=True)
ok = already == "https://joakim:<redacted:url-password>@git.example.com" and not found_a
failures += not ok
print(f" {'pass' if ok else 'FAIL'} already-redacted url left alone")
obj = {"role": "user", "content": [{"type": "text", "text": f"export TOK={fake_token}"}],
"id": "evt_20260826T213924_397b1f710c2d"}
scrubbed, found = scrub_obj(obj, known, strict=True)
ok = (fake_token not in repr(scrubbed)
and scrubbed["id"] == obj["id"]
and scrubbed["role"] == "user"
and len(found) >= 1)
failures += not ok
print(f" {'pass' if ok else 'FAIL'} nested structure scrubbed, ids and keys preserved")
ok = all(fake_token not in repr(f) for f in scrub(f"x {fake_token}", known, strict=True)[1])
failures += not ok
print(f" {'pass' if ok else 'FAIL'} findings never carry the secret value")
sus = suspicions("drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b "
"rep_d344e349ba276d6fc11997cd552f6937 "
"1787780364792-000000-rep_d344e349ba276d6fc11997cd552f6937 "
"ecc2a9c574f4156403cf4e90aea8b0d088c75700")
ok = not sus
failures += not ok
print(f" {'pass' if ok else 'FAIL'} suspicion meter ignores our own id shapes ({len(sus)} hits)")
sus2 = suspicions("blob=Zm9vYmFyYmF6cXV1eHF1dXhmb29iYXJiYXpxdXV4Zm9vYmFy")
ok = len(sus2) == 1 and all("Zm9v" not in repr(s) for s in sus2)
failures += not ok
print(f" {'pass' if ok else 'FAIL'} suspicion meter flags unknown blobs as fingerprints only")
print(f"\n{'ALL PASS' if not failures else f'{failures} FAILED'}")
return 1 if failures else 0
if __name__ == "__main__":
import sys
if "--self-test" in sys.argv:
raise SystemExit(_self_test())
# Filter mode: scrub stdin -> stdout, report to stderr. Useful for spot
# checks and for scrubbing a file by hand without writing throwaway code.
data = sys.stdin.read()
out, found = scrub(data)
sys.stdout.write(out)
if found:
by_rule: dict[str, int] = {}
for f in found:
by_rule[f.rule] = by_rule.get(f.rule, 0) + 1
print(f"[redact] {len(found)} redaction(s): "
+ ", ".join(f"{k}={v}" for k, v in sorted(by_rule.items())), file=sys.stderr)
if "--suspicions" in sys.argv:
for s in suspicions(out):
print(f"[suspicion] len={s.length} fp={s.fingerprint}", file=sys.stderr)
+2 -3
View File
@@ -48,9 +48,8 @@ The pi variants are drop-in copies of the opencode variants with script name and
The odd one out in this directory: every other template *feeds* a palace on a schedule, this one
**serves** a palace over HTTP so several machines can share it (RFC-001). **This unit currently runs
the fleet primary on `synlig`** — serving since 2026-08-12, seeded 2026-08-14, and the palace behind
it is the only copy. Read [`docs/rfc-001-global-palace.md`](../docs/rfc-001-global-palace.md) and
[`docs/phase-1-exposure-runbook.md`](../docs/phase-1-exposure-runbook.md) before installing a second one.
a fleet primary** — whose palace may be the only copy. Read [`docs/rfc-001-global-palace.md`](../docs/rfc-001-global-palace.md)
before installing a second one.
It is a **user** unit (`systemctl --user`), so it dies with your login session unless lingering is
enabled — that is the one `sudo` this recipe needs:
+218
View File
@@ -0,0 +1,218 @@
# Fleet memory — what MemPalace stores, and what to put where
**Audience:** the person running MemPalace on more than one machine, or thinking about it.
**Companion documents:** `docs/rfc-003-coordination-log.md` for the coordination log's mechanism, `docs/rfc-001-global-palace.md` for the centralisation design, and `~/.agents/skills/mempalace/SKILL.md` for what the *agents* are told to do.
MemPalace is usually described as "memory for agents", which is true and not very actionable. It is really **five stores with different retrieval models**, and most of the value — especially across a fleet — comes from putting each kind of thing in the store whose retrieval model matches how you will want it back.
This document covers: what the stores are, what a central palace changes when several machines share one, and how to decide between filing a memory and sending a message.
---
## 1. The stores at a glance
```mermaid
flowchart LR
Q["ask by MEANING<br/>'what do we know about X?'"] --> D["<b>Drawers</b><br/>chroma.sqlite3"]
R["ask by ENTITY + TIME<br/>'what was true in June?'"] --> G["<b>Knowledge graph</b><br/>knowledge_graph.sqlite3"]
S["ask by ADDRESS<br/>'what is waiting for me?'"] --> L["<b>Coordination log</b><br/>logstream.sqlite3"]
T["ask by ASSOCIATION<br/>'what else touches this?'"] --> P["<b>Palace graph</b><br/>tunnels.json + hallways.json"]
D --> DI["<b>Diaries</b> live here too:<br/>drawers with room=diary"]
```
All four stores are files inside **one palace directory**, so "the palace" is a directory you can back up in one go. Each holds a different *shape* of thing: drawers hold verbatim text, embedded for similarity; the knowledge graph holds typed facts that are time-bounded; the coordination log holds addressed events and their artifacts, read in append order; the palace graph holds links between rooms. Diaries are drawers, but they are read by recency rather than by similarity.
| Store | You get things back by | Typical use | Wrong use |
|---|---|---|---|
| **Drawers** (wings → rooms) | semantic similarity | a verbatim finding, a decision and its reasoning, a runbook, a transcript excerpt | anything a specific machine must *act* on; anything whose value is its exact byte content |
| **Diaries** (drawers with `room="diary"`) | agent + recency | "what did I do last session, and what did it feel like" — the continuity thread across sessions | facts other agents need to find by searching; a diary is read by *its author*, chronologically |
| **Knowledge graph** | entity, relationship, point in time | facts that *change*: versions, employers, who owns what, an injury that heals | prose, reasoning, anything you'd want to read rather than query |
| **Coordination log** | address, correlation, append order | "device B must review this patch"; "this claim is retracted, stop building on it" | durable knowledge — an event is invisible to semantic search |
| **Palace graph** | traversal from a room | discovering that an API design in one project touches a schema in another | primary storage — it links drawers, it does not hold content |
Two smaller files exist and are implementation detail, not user surface: `sqlite_exact.sqlite3` (an exact-match index over drawer metadata) and, if the daemon runs, `queue.sqlite3` (its job queue).
### 1.1 Drawers: wings and rooms
A **wing** is a project or domain; a **room** is an aspect within it. `wing="pi-devbox", room="landmines"` is a good pair; `wing="misc", room="stuff"` is how a palace becomes a landfill. Content is stored **verbatim and chunked** — never summarised — and retrieved by embedding similarity, so a drawer is found by someone who *doesn't already know it exists*. That is the property to optimise for: write the drawer that the next person's search will match.
The one counter-intuitive consequence: **fresh drawers rank worst.** A drawer filed an hour ago has no advantage in a similarity search, and a well-worn older drawer will outrank it. For anything recent, enumerate by date (`list_drawers(since=…)`) instead of searching.
### 1.2 Diaries
A diary entry is a drawer with `room="diary"`, filed by default into `wing_<agent>`, tagged with the writing agent. It is the *first-person* record: what I did, what surprised me, what I would do differently. Agents are told to write one before a session ends, and to read the last few at session start.
In a fleet this is the highest-signal store per byte, for a reason that is easy to miss: a diary entry is the only place that records **what did not work**. A drawer tends to record the conclusion; the diary records the three hours that produced it.
### 1.3 Knowledge graph
Triples — subject, predicate, object — with `valid_from` / `valid_to`, so a fact can *stop* being true without being deleted. `supersede` replaces a single-valued fact at a shared boundary, so a point-in-time query at that instant returns exactly one value.
Use it for anything you will later want to ask "what was true at time T?" about: which version was released when, who owned a service, what model an assistant was using. Do not use it for prose — a triple whose object is a paragraph is a drawer wearing a costume.
### 1.4 The coordination log
Addressed, ordered, exact. `mempalace_event_*` carries messages between agents; `mempalace_artifact_*` carries byte-exact payloads (patches, logs, files) that events can reference. This is the only store where one machine can *reach* another.
Its full mechanism, limits and landmines are in `docs/rfc-003-coordination-log.md`. §4 below covers what an operator needs to decide.
---
## 2. What changes when a fleet shares one palace
A single machine's palace is a notebook. A shared palace is something different in kind: **the fleet stops being a set of independent agents that each learn the same lessons separately.**
```mermaid
flowchart LR
A["laptop<br/>agent session"] -->|MCP over HTTPS| H(("central palace<br/>one server"))
B["workstation<br/>agent session"] -->|MCP over HTTPS| H
C["build box<br/>agent session"] -->|MCP over HTTPS| H
H --> D["drawers + diaries"]
H --> G["knowledge graph"]
H --> L["coordination log"]
```
Note the topology: in the common deployment the machines are **thin clients of one server**, not peer replicas. Everything one machine writes is immediately visible to the others — there is no sync delay to reason about, and equally no local copy to fall back on when the server is unreachable. (`docs/rfc-001-global-palace.md` §4 designs an edge proxy with a local palace and a durable outbox for deployments that need to keep working offline; the plain thin-client shape above does not.)
### 2.1 The three things this actually buys
**Awareness.** "What has anyone been doing?" becomes answerable. Each machine's diary is readable by every other machine, so an agent starting work on a shared project can see that another machine spent yesterday on it, and how far it got.
**Non-repetition of expensive work.** This is the biggest measurable win. Anything that cost real time to obtain — a scraped API surface, a spec read end to end, a bisect, a benchmark, a long investigation into why a build fails on one platform — is filed once and searchable everywhere. The second machine's cost drops from hours to one search.
**Mistakes and retractions travel.** The subtle one, and the reason a shared palace is worth more than a shared wiki. When a machine discovers that a belief was *wrong*, it can file the retraction where every other machine will hit it. Without that, each machine independently rediscovers the same dead end — and worse, a machine can spend a day rebuilding something another machine already proved doesn't work.
```mermaid
sequenceDiagram
participant W as workstation
participant P as central palace
participant L as laptop, asleep 9 days
W->>W: spends 3h finding why the build breaks
W->>P: drawer — the finding, verbatim, with evidence
W->>P: KG fact — "component X requires flag Y"
W->>P: diary — what was tried and failed
Note over L: ...9 days pass, laptop asleep...
L->>P: wake-up: read diaries + search before starting
P-->>L: the finding, the failed attempts, the fact
Note over L: cost: one search instead of 3 hours
```
### 2.2 The two costs, stated plainly
**Everything is visible to everyone.** One shared token, no per-agent read scoping. Anything filed into a shared palace should be considered readable by every machine and every agent on it. Do not put secrets in drawers.
**Provenance stops being obvious.** On a single machine, every drawer is yours and every path exists. On a shared palace, most drawers came from other machines and most `source_file` paths **do not exist locally**. Two consequences worth internalising:
- A file path in a drawer is evidence about *some* machine, not necessarily this one.
- ⚠️ **Never run `mempalace sync` against a shared palace.** It prunes drawers whose source files look gitignored, deleted or moved — which on a shared palace describes most of the content, including every other machine's. See `docs/rfc-001-global-palace.md` §7.2. The coordination log is *not* affected by this (RFC 003 §2), but drawers very much are.
---
## 3. Deciding where something goes
The question that matters is not "is this important?" but **"how will I want this back, and does a specific machine or agent need to act?"**
```mermaid
flowchart TD
Start["I have something worth keeping"] --> Act{"Must someone specific<br/>DO something?"}
Act -->|no| Change{"Is it a fact that<br/>changes over time?"}
Act -->|yes| Know{"Do they also need<br/>to KNOW it later?"}
Change -->|yes| KG["Knowledge graph<br/>kg_add / kg_supersede"]
Change -->|no| Mine{"Is it about MY session<br/>— tried, felt, learned?"}
Mine -->|yes| Diary["Diary entry"]
Mine -->|no| Drawer["Drawer<br/>wing + room, verbatim"]
Know -->|yes| Both["BOTH:<br/>drawer + event pointing at it"]
Know -->|no| Event["Coordination event<br/>to_agent = that agent"]
```
Worked examples:
| Situation | Where | Why |
|---|---|---|
| "The release takes 76 min, and 40 of those are the base image build." | Drawer | durable, nobody must act, next person finds it by searching "release timing" |
| "v1.8.9 is the released version, as of this timestamp." | KG (`supersede`) | it will change; you will want "what was released in August?" |
| "I spent two hours chasing a watcher that was already dead." | Diary | first-person, chronological, tells the next session what *not* to retry |
| "Build box: this patch is ready, please review and apply." | Event (directed) | a named machine must act; the patch itself goes in as an artifact |
| "The claim in that drawer is wrong — I measured the opposite." | Both | file the corrected finding as a drawer, then an event so the machine building on it stops |
| "Everyone should know the new toolkit is live." | Drawer + broadcast event | the drawer is what anyone will *find*; the broadcast is a notice, not an ask (§4.2) |
The failure mode in each direction is worth naming, because both are common:
- **A finding filed only as an event** is invisible to semantic search. Nobody will ever find it again, and the next agent will re-derive it.
- **An ask filed only as a drawer** is addressed to nobody. It will be found, if ever, by accident — long after it mattered.
---
## 4. What the coordination log can and cannot do
This is the part most likely to be mis-set expectations, so it is worth being blunt: **it is a durable log, not a chat.** Nothing is listening. Events are appended and persist; there is no delivery window; nothing is lost by being offline when one is written. A message waits indefinitely, and your reply waits just as patiently for a sender who has since gone away.
That sounds like a limitation and is actually the correct design for a fleet where few machines are awake at once and any given machine may sleep for weeks.
### 4.1 Latency, honestly
| Recipient state | When they see it | Notes |
|---|---|---|
| In a live session, mailbox-enabled client | **≈2–5 minutes** | measured ≈2–3 min on first live delivery; a poll floor of 5 min applies between checks |
| Holding an SSE connection (`GET /logstream/stream`) | sub-second | for daemons/dashboards, not interactive agents |
| Asleep, next session tomorrow | tomorrow | delivered in the session-start wake-up |
| Asleep for three weeks | in three weeks | nothing expires; the log is permanent |
| Never runs again | never | there is no re-routing and no dead-letter path |
So: appropriate for "handle this when you next wake", "here is a patch", "stop building on that claim". Not appropriate for anything with a deadline inside the hour, unless you know the recipient is awake.
One pleasant property, worth knowing because it is counter-intuitive: coordination traffic is **exempt from the palace's write lock**, so you can message another machine and it can reply *while* a long mine is running on the server (RFC 003 §4). Coordination stays alive when memory writes are blocked.
### 4.2 The one thing that does not work: broadcasting an ask
You can write a broadcast (`to_agent="*"`) and every machine that *lists* events will see it. But a broadcast **never enters any machine's mailbox** and is never auto-delivered — by design, because "everyone owes this answer" degenerates into either N duplicate replies or nobody acting.
- **Broadcast** = a notice on a wall. Fine for "v1.8.9 is out".
- **Directed event** = a message in a named mailbox. Required for anything that must be done.
To reach a whole fleet with something actionable, **fan out**: one directed event per device, sharing one `correlation_id` so the thread stays joinable. Each machine then owes its own reply.
```mermaid
flowchart LR
You["you"] -->|"to_agent=pi@laptop"| A["laptop owes a reply"]
You -->|"to_agent=pi@workstation"| B["workstation owes a reply"]
You -->|"to_agent=pi@build-box"| C["build box owes a reply"]
A --> Corr["one shared correlation_id<br/>joins the three threads"]
B --> Corr
C --> Corr
```
### 4.3 Addressing, and why the format matters
Addresses are `<harness>@<device>` — `pi@laptop`, `opencode@build-box`. Two rules follow:
1. **An unstamped client is unreachable.** If a machine's events say `from_agent: pi` with no device, nobody can address it, because "pi" is every machine.
2. **Two machines must never share one address.** Nothing prevents it, nothing warns, and the result is that each silently discards the other's asks as its own (RFC 003 §7.7).
### 4.4 A reply is owed until it is *terminally* closed
An acknowledgement does not close a thread. Neither does "claimed" or "ready". Only a terminal status — `applied`, `superseded`, `failed`, `blocked` — clears an ask from the recipient's mailbox. Until then, a mailbox-enabled client will keep resurfacing it, which is the intended behaviour: an unanswered ask should nag.
---
## 5. Habits that make a shared palace work
Small, and the whole value rests on them:
1. **Search before you answer, and enumerate before you conclude.** One empty search is not proof of silence — fresh drawers rank worst, so for anything from the last couple of days list by date and read the other machines' diaries.
2. **Write the diary entry before the session ends.** It is the store other machines learn from fastest, and the only one that records failed attempts.
3. **File retractions as loudly as findings.** "I was wrong about X, here is the measurement" is worth more than a new finding, because it stops N machines repeating a dead end.
4. **Check the mailbox at wake-up, even when you expect nothing.** An empty result costs one call. Silence is only informative once you know delivery works.
5. **Say which machine you are talking about.** On a shared palace, "the container" and "the host" are ambiguous and a path is not self-identifying.
6. **After a write times out, verify — do not blindly retry.** The palace is single-writer for memory writes, so a timeout usually means the write *completed*. For coordination events this matters twice over: there is no idempotency guard, so a retried event forks the thread into two (RFC 003 §7.1).
---
## 6. See also
- `docs/rfc-003-coordination-log.md` — the coordination log: storage, semantics, security model, landmines.
- `docs/rfc-001-global-palace.md` — how and why a palace is centralised; §5 what should not be global; §7 the landmines, including the `sync` hazard.
- `docs/phase-1-exposure-runbook.md` — (moved) the site-specific exposure record is now private; the stub names the mechanism still awaiting extraction.
- `docs/backup-and-recovery.md` — why a palace needs its own backup procedure.
- `extensions/pi/README.md` — the pi-side client: provenance stamping and the auto-delivered mailbox.
- `~/.agents/skills/mempalace/SKILL.md` — the protocol the agents themselves follow.
+47 -463
View File
@@ -1,463 +1,47 @@
# Phase 1 exposure — newt on synlig, DNS, and the client auth model
Companion to [`rfc-001-global-palace.md`](./rfc-001-global-palace.md) (design + decisions) and
[`synlig-primary-runbook.md`](./synlig-primary-runbook.md) (what is already installed on the primary).
This doc covers only the step the other two leave open: **making the primary reachable** — runbook §4
items 2 and 5.
> **Status 2026-08-14 — DONE. Exposed, seeded, and one client flipped.**
> `https://mempalace.jordbo.se/mcp` has been serving since 2026-08-12 (§3.3–§3.6 are all ✅ below).
> The palace was seeded 2026-08-14 15:07 from EMB-7KJ4VR4G (14,777 → 14,803 drawers) — whose palace was
> itself carried over from the previous work computer **EMB-X1JY06WJ** around 2026-07-06, so the
> primary's lineage is EMB-X1JY06WJ → EMB-7KJ4VR4G → synlig by two file-level copies (rfc-001 §4.4
> Deviation) — and that machine's pi-devbox container is flipped and **verified end-to-end — see §3.8**, which is the
> verification procedure that did not exist when the first flip was performed.
> Still outstanding: the transcript feeder is **inert on every deployed image** (§3.7), and §7.6 is
> still a hard blocker for the *second* machine to join (§4).
Original status, kept for the record — it contradicted this file's own ✅ section markers for two days,
which is the failure mode a file-top status block invites:
**Status 2026-08-12 — Pangolin updated on nyvaken (done, yours). newt not yet installed on synlig.
Nothing exposed. No client `.env` flipped.**
Read this before touching Pangolin: three of the four questions this step raises were **already decided**
in RFC §6.2 on 2026-08-09, and re-deciding them differently is how the fleet ends up in two states.
---
## 1. The four questions, answered
| Question | Answer | Where it was decided |
| --- | --- | --- |
| Which port? | **8765**, path **`/mcp`** (liveness: `/healthz`) | `cli.py:2141` default; runbook §2.4 |
| What does newt target? | **`172.17.0.1:8765`** (docker0), **never** `127.0.0.1` | RFC §6.2 Transport; runbook §2.4 |
| Open, or authenticated? | **Authenticated. The primary is never an open public resource.** | RFC §6.2 Network posture |
| Per-device credentials? | **No — Phase 1 ships the single shared bearer token.** Per-device tokens are Phase 4. | RFC §6.2 Authentication |
### 1.1 Why not per-device users at the proxy
The instinct — "create a Pangolin user per container, put the credentials in each `.env`, keep the
usernames distinct" — is the right *goal* (revocation, attribution) reached through the wrong *layer*,
twice over:
1. **mempalace validates exactly one token.** `hmac.compare_digest(provided, f"Bearer {srv.auth_token}")`
(`mcp_server.py:5292-5295`) — there is no user table and no second credential. Per-device HTTP identity
is not a configuration you can express today; it is Phase 4 work (a server-side
`token → {device_id, scopes}` registry). RFC §6.2 chose the shared token for Phase 1 deliberately:
*"iterate more feature rich but more complex solutions over time."*
2. **Pangolin's HTTP auth is browser-shaped; the clients are not.** SSO login, resource PIN and resource
password all assume something that can follow a redirect, render a form and hold a session cookie.
Every MemPalace client here is a headless JSON-RPC `POST` with an `Authorization` header — pi's
extension, opencode's `type:remote` MCP entry, and `mempalace-pi-session --mode remote`'s
`urllib.request.urlopen`. Point those at a user-authenticated resource and they receive a login page
where JSON should be. Enabling that protection breaks precisely the clients it is meant to protect.
So: **Pangolin terminates TLS and nothing more** (RFC §6.2 Transport, decided 2026-08-09). The bearer token
is the authentication. This is not "unprotected" — an unauthenticated request to `/mcp` gets a 401 from
mempalace itself, verified A4/A5 in runbook §2.4.
**Consequence to accept consciously** (RFC §6.2, §7.3.2): until Phase 4 the primary **cannot tell devices
apart**. `origin_device` is client-asserted and advisory — nothing load-bearing may depend on it, and
revoking one laptop means rotating the token everywhere.
### 1.2 The one place per-device identity *does* exist today
Remote mode is not only HTTP. `mempalace_mine` expands its source path in the **server** process, so a
client's staged transcripts must physically exist on the primary. The feeder therefore ships them over
SSH into a **per-device inbox** before asking the server to mine its own copy:
```sh
rsync -a --update -e "$ssh_cmd" "$STAGE/" "${SSH_TARGET%/}/$DEVICE/" # bin/mempalace-pi-session:677-680
```
That SSH key **is** per-device identity, and it is individually revocable (one line out of
`authorized_keys`) years before Phase 4 lands. It costs nothing extra, because the mining path needs SSH
regardless.
Two implications people miss:
- **`mempalace.jordbo.se` alone does not enable mining.** HTTPS covers the read/write tool surface
(`search`, `add_drawer`, `diary_write`, `kg_*`) — genuinely useful on its own, and the reason to do this
at all. But `--mode remote` also needs `MEMPALACE_PI_SSH_TARGET` reachable. Budget for both paths.
- **`DEVICE` defaults to `$(hostname)`** (`bin/mempalace-pi-session:163`). In a container that is the
container hostname: either random per recreate (inboxes proliferate; each recreate re-mines into a fresh
empty inbox) or identical across sibling devboxes (two containers writing one inbox). **Set
`MEMPALACE_PI_DEVICE` explicitly per container.** It is a label, not a secret, so put it somewhere
reviewable — a committed compose file — where duplicates are visible. That, not username hygiene in
`.env`, is the discipline this design actually asks of you.
### 1.3 "Then why Pangolin at all, if the feeder uses SSH?"
Because they are not alternatives — they carry different traffic, and neither substitutes for the other.
| | Pangolin/newt (HTTPS) | SSH + rsync |
| --- | --- | --- |
| Carries | the **MCP tool surface**: `search`, `add_drawer`, `diary_write`, `kg_*` — every live tool call | **transcript files only**, once per session or cron run |
| Used by | the pi extension, opencode `type:remote`, any MCP client | the feeder, internally (`bin/mempalace-pi-session:677-680`) |
| Needed because | clients need one stable URL, reachable from wherever they are | `mempalace_mine` expands its source path **server-side**, so the server can only mine files on its own disk |
HTTPS alone is a palace you can query but cannot feed. SSH alone is files shipped with no live query API.
The rsync is not a transport preference; it is a workaround for *where `mine` resolves paths*.
**Could SSH replace Pangolin?** Partly, and it is worth being honest about it:
`ssh -L 8765:172.17.0.1:8765 synlig` yields a working local MCP endpoint with no public HTTPS at all.
Three reasons this runbook does not do that:
1. **Direction.** synlig dials *out* through newt. That we reached for a dial-out tunnel rather than a
port-forward is itself the evidence that inbound was not available — a corporate host does not accept
connections from a phone on a foreign network.
2. **MCP clients want a durable URL**, not a per-session forwarded port. opencode `type:remote` takes a
URL; a forward that drops takes the tools down mid-session.
3. The forward must be up on **every device before every session**. Pangolin is up once.
**The weak point, stated plainly.** The rsync runs *client → synlig*, so it needs synlig's SSH reachable
**from the client**. Were that already true everywhere, no tunnel would be needed for MCP either. So the
honest expectation after Phase 1 is: **query and write from anywhere, mine only from devices that can
reach synlig's SSH** (corporate network / VPN / LAN). See §4 for the change that would remove that limit.
---
## 2. The bind trap, in full
RFC §6.2 and runbook §2.4 already say **do not bind loopback behind the tunnel**, because
`enforce_host_pin = _http_is_loopback(host)` (`mcp_server.py:5367`) makes a loopback bind reject the
proxy's forwarded `Host:` with a **403** that reads exactly like a Pangolin misconfiguration.
**Additional finding, 2026-08-12 — the same reflex also silently removes authentication.** Token
resolution in `cmd_serve` (`cli.py:1447-1450`) is:
```python
loopback = _server_is_loopback(host)
if not token and not loopback and not args.allow_insecure:
token, token_created = _load_or_create_server_token(palace_path)
```
Auto-minting is gated on the bind being **non-loopback**. A loopback bind therefore starts with **no token
at all** — no error, no warning, `--allow-insecure` not required — because the server has concluded it is
only reachable locally, while the tunnel is serving it to the internet. Bind loopback behind newt and you
get a 403 wall *and*, the moment anything relaxes the Host pin, an unauthenticated palace.
Both failure modes have the same cure, already implemented in
`contrib/systemd/mempalace-serve.service`: **bind `172.17.0.1`**. Non-loopback, so the Host pin relaxes and
the token is mandatory; docker0-only, so newt reaches it and the LAN does not.
> Belt and braces: set `MEMPALACE_MCP_HTTP_TOKEN` explicitly in the unit rather than relying on
> auto-minting. Then no future bind change can quietly drop authentication.
---
## 3. Steps
Ordered so nothing is reachable before it is authenticated.
### 3.1 Start the primary (runbook §4.3 — one `sudo`, unit already staged)
```sh
sudo loginctl enable-linger ecsjper
cd ~/.config/systemd/user && mv mempalace-serve.service.staged mempalace-serve.service
systemctl --user daemon-reload && systemctl --user enable --now mempalace-serve
curl -s 172.17.0.1:8765/healthz # expect ok
curl -s 127.0.0.1:8765/healthz # expect NOTHING — connection refused, exit 7 (see below)
ss -ltnp | grep 8765 # expect 172.17.0.1:8765 only
```
⚠ **The loopback probe returns empty, not 403** — corrected 2026-08-12 against the real run. Nothing is
listening on `127.0.0.1`, so the connection is refused at TCP level and `curl -s` prints nothing; check it
with `-w '%{http_code}'` → `000` and `$?` → `7`. The 403 belongs to a *different* configuration: server
bound **to loopback**, receiving a proxy-forwarded foreign `Host:` (§2, verified 2026-08-10). With a
docker0-only bind you cannot get 403 from loopback, because you never get far enough to send a header.
Refusal is the stronger signal of the two: it proves the loopback and LAN surface is not listening at all.
If it *hangs* instead, or `ss` shows `0.0.0.0:8765`, stop — that is not this configuration.
### 3.2 Collect the shared token
```sh
cat ~/.mempalace/server/f5d849287f6d73f0141b29d7/token
```
Directory name is `sha256(realpath(palace))[:24]` — it changes if the palace path ever moves. Store via the
`.env.age` flow, 0600 (RFC §6.2).
### 3.3 newt on synlig — ✅ done 2026-08-12 (installed, connected to Pangolin)
synlig runs Docker (Gitea Actions runner + digikam) but **no tunnel client** — runbook §4.2. Pangolin on
nyvaken cannot dial in; synlig must dial out. Add a `newt` container with the credentials Pangolin issues
for a new site.
Because newt runs in Docker on this box, the docker0 bind is already correct for it: from inside the
container the primary is `172.17.0.1:8765`. Verify from *inside* newt's network namespace, not from the
host, before touching DNS.
> synlig has 7.8 GiB shared with a CI runner (runbook §1). newt is small, but do not colocate anything
> else here casually.
**Confirm next, now that newt is up.** "Connected to Pangolin" proves newt reached *nyvaken* — a different
claim from newt reaching *the palace*, and the two fail independently:
```sh
# from inside newt's namespace, not from the host
docker exec <newt-container> wget -qO- http://172.17.0.1:8765/healthz # expect ok
```
If that hangs or refuses while the Pangolin dashboard shows the site online, the tunnel is fine and the
*target* is wrong — look at the resource's upstream address (§3.5), not at newt. Note this check needs
§3.1 done first: if `mempalace-serve` is not running yet, it fails for that reason alone.
### 3.4 DNS at the web hotel — ✅ done 2026-08-12
One CNAME: `mempalace` → **the same target your existing Pangolin resources use** (nyvaken's public
hostname). RFC §6.2 costed this as *"one DNS record per service on the web hotel is the whole setup cost."*
Done: `mempalace.jordbo.se` resolves and terminates TLS at Pangolin on nyvaken — the §6.2 estimate held.
⚠️ Not verified from here: nyvaken's public FQDN, and whether your web hotel permits a CNAME at that label
(some require an A record, or forbid CNAME where other records exist). Confirm before assuming a 5-minute job.
### 3.5 Pangolin resource
- Target: newt site → **`http://172.17.0.1:8765`** — path `/mcp` (plus `/healthz` for the external probe).
- ⚠️ **The scheme is `http`, not `https`.** The primary runs `serve --host 172.17.0.1 --port 8765` with no
cert (`contrib/systemd/mempalace-serve.service`): TLS terminates **at Pangolin**, which is the entire
point of the §6.2 decision. Point the resource at `https://172.17.0.1:8765` and Pangolin attempts a TLS
handshake against a plaintext listener — you get a 502/Bad Gateway from outside while the server itself
looks perfectly healthy on `curl 172.17.0.1:8765/healthz`. (Got this wrong on the first attempt
2026-08-12, because this line used to omit the scheme.)
- **Auth: none at the Pangolin layer** (§1.1). TLS termination only.
- ⚠️ **If resource auth is left on, the signature is a `302`, not a 401 or 403** — hit for real 2026-08-12:
```
HTTP/2 302
location: https://pangolin.jordbo.se/auth/resource/<uuid>?redirect=https%3A%2F%2Fmempalace.jordbo.se%2Fhealthz
content-length: 0
```
This is §1.1's "browser-shaped auth" arriving as a concrete symptom: Pangolin sends the login redirect
**before** proxying, so the palace never sees the request and its journal stays silent. `curl -s` shows
an empty body and an MCP client sees non-JSON. Diagnose with `-D-` or
`-w '%{http_code} %{redirect_url}'` — a `location:` pointing at `/auth/resource/…` means the fix is in
the Pangolin UI (switch the site's authentication off), not in the palace, the unit, or newt.
**A 302 is unambiguously good news:** DNS, TLS and routing all worked — only the auth layer intervened.
- Do **not** attach an `Origin`-injecting proxy or browser client: a *present* non-loopback `Origin` is a
hard 403 with no override (runbook §2.4 B3).
**The server's entire HTTP surface is two exact paths**, so path-scoped rules cover it completely
(`mcp_server.py:5299-5318`, read 2026-08-12):
| Method + path | Auth | Notes |
| --- | --- | --- |
| `GET /healthz` | none (Host/Origin gated only) | the liveness probe; works with no creds by design |
| `POST /mcp` | `Authorization: Bearer <token>`, `hmac.compare_digest` on the exact string | the whole tool surface |
| anything else | — | `send_error(404)` from the palace itself |
Three consequences worth having in writing:
- **Path-scoped Pangolin rules are not a compromise here, they are tighter than a host-wide proxy** and
lose nothing — there is no third endpoint to forget.
- **`/mcp` is matched exactly** (`if path != "/mcp"`), so a client URL with a trailing slash gets a 404
from the palace. Configure clients as `https://mempalace.jordbo.se/mcp` — no trailing slash.
- **There is no `GET /mcp`, no SSE, no session id, no `DELETE`.** This is plain JSON-RPC over POST, not MCP
streamable-HTTP. A strict client that opens with a `GET` handshake will see 404; pi's extension,
opencode `type:remote` and the feeder all POST directly and are fine.
### 3.6 Verify end-to-end before flipping any client — ✅ passed 2026-08-12
```sh
curl -s https://mempalace.jordbo.se/healthz # ok ✓ from devbox AND synlig
curl -s -o /dev/null -w '%{http_code}\n' -X POST \
https://mempalace.jordbo.se/mcp # 401 ✓ the token is doing its job
TOKEN=$(cat ~/.mempalace/server/f5d849287f6d73f0141b29d7/token) # on synlig
curl -s -X POST https://mempalace.jordbo.se/mcp \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | head -c 300 # 36 tools
```
No `initialize` and no `Accept: text/event-stream` needed — see the surface table above; a bare
`tools/list` POST is a complete request. Run the 200 and the 401 from **two different networks** (a client
box and the primary itself): passing from only one leaves split-horizon DNS untested.
The 401 check matters as much as the 200: it is the only evidence that the thing you just published to the
internet is not open. Then, and only then, Phase 1 client flip — **one machine first** (RFC §8), and
remember opencode containers need the §4.1 sidecar merge before their `.env` takes effect.
### 3.7 Flipping a client: set three variables, or none
⚠ **`MEMPALACE_REMOTE_URL` on its own does not degrade to local feeding — it stops feeding.** `auto` mode
switches to `remote` the moment the URL is set, and remote mode then refuses to run without an SSH target:
```sh
auto) if [[ -n "$REMOTE_URL" ]]; then MODE="remote"; else MODE="local"; fi ;; # :286
...
command -v rsync >/dev/null 2>&1 || { echo "error: rsync not found ..."; exit 3; } # :297
if [[ -z "$SSH_TARGET" ]]; then
echo "error: MEMPALACE_PI_SSH_TARGET unset (needed for --mode remote)" >&2; exit 1 # :298-300
fi
```
That exit happens **before anything is staged or filed**, and a cron-driven feeder will simply start
failing — the loudest symptom is silence, which is the hardest kind to notice. Two safe orders:
- **Both paths at once:** set `MEMPALACE_REMOTE_URL`, `MEMPALACE_REMOTE_TOKEN` **and**
`MEMPALACE_PI_SSH_TARGET` (plus `MEMPALACE_PI_DEVICE`, §1.2) in the same edit.
- **HTTPS first, mining later:** set the URL and token, and pin the feeder to `--mode local` until the SSH
target exists. Tools then read/write the shared palace while transcripts keep landing in the local one.
`MEMPALACE_REMOTE_URL` is the **full endpoint including `/mcp`, with no trailing slash** — the feeder POSTs
to it verbatim (`urllib.request.Request(url, data=payload, …)`, `:718`), and §3.5 shows the server matches
`/mcp` exactly. Matches the existing examples (`docker-compose.mempalace.yml:7`,
`MEMPALACE_REMOTE_URL=http://<reachable-host>:8765/mcp`). For this fleet:
`MEMPALACE_REMOTE_URL=https://mempalace.jordbo.se/mcp`
Either way, run the feeder once by hand and read its exit code before trusting the timer. This is
precisely the failure "one machine first" is meant to contain.
**The same misconfiguration has two different symptoms depending on who invokes the feeder** — checked in
the code 2026-08-12, and the quieter one is the trap:
| Caller | Behaviour with `REMOTE_URL` set and `SSH_TARGET` unset |
| --- | --- |
| direct run, session-end hook, cron | `exit 1` with `error: MEMPALACE_PI_SSH_TARGET unset` (`:298-300`) |
| **pi-devbox container start, images ≤ v1.7.0** | **silently skips — no error, no log file** (`entrypoint-user.sh:134`: `: # remote palace but no inbox configured — nothing we can ship to; skip quietly`) |
| pi-devbox container start, images after `cbd7cf5` | prints `MemPalace catch-up skipped: remote palace with no transcript inbox` to the start output *and* to `mempalace-catchup.log`, naming both variables |
The silent skip is fixed in pi-devbox (`cbd7cf5`), but the fix is in
`entrypoint-user.sh`, which is `COPY`d in `Dockerfile.base` — so **every container running an image built
before that base rebuild still skips silently.** That is the whole fleet today. Until the rebuild lands,
assume silence and check by hand.
The old skip happened *before* the subshell that writes `~/.pi/agent/mempalace-catchup.log`, so there was
not even an empty log to notice. Someone asking "why is nothing from this container in the palace?" found
no artifact at all. The skip itself is correct — there is genuinely nothing to ship to — but it was
indistinguishable from a healthy run that had nothing to do, which is the worst property a memory system
can have: **the failure looks exactly like success.**
→ On a pi-devbox container, confirm the feeder is actually alive after a flip rather than assuming:
```sh
mempalace-pi-session --reason manual-check; echo "exit=$?" # exit=0 and a filed count, not silence
cat ~/.pi/agent/mempalace-catchup.log # missing file = the entrypoint skipped
```
### 3.8 Verify the flip actually took — ✅ done 2026-08-14 on EMB-7KJ4VR4G
§3.6 verifies the *endpoint* before you flip. §3.7 tells you *how* to flip. Neither verifies the thing
you actually care about afterwards: **that the agent's palace tools are now talking to synlig.** That
gap is why the first flip was "done" for an hour before anyone could say whether it had worked.
There are **three separate claims** here and they fail independently. Check them in order; each one is
cheap and rules out a different fault.
**(a) Did the container receive the variable?** In a shell *inside* the container:
```sh
env | grep MEMPALACE_REMOTE_URL # -> https://mempalace.jordbo.se/mcp
```
⚠ Do **not** run a bare `env | grep MEMPALACE` — that prints the bearer token into your scrollback.
If the variable is absent, the cause is almost always §3.7's edit not having been applied: `env_file`
is read at container **create** time and baked into the container config, so **`docker compose up -d`
is required; `docker compose restart` silently reuses the old config.** Evidence from 2026-08-14 —
running container config-hash `3a55e09ac19e0118` vs compose-computed `8a7e0cd4a677a9c1`; a differing
hash is what makes `up -d` recreate. `--force-recreate` is not needed.
**Never `docker compose down -v`** to "pick up" a change: on pi-devbox that destroys seven named
volumes, including `devbox-pi-config` (pi's config **and every session transcript**), `devbox-uv`, and
`devbox-chroma-cache` (a large embedding-model re-download).
**(b) Is the server reachable and the token accepted?** Still inside the container — this proves the
network path and the credential, independently of any agent:
```sh
curl -s -X POST "$MEMPALACE_REMOTE_URL" \
-H "Authorization: Bearer $MEMPALACE_REMOTE_TOKEN" \
-H 'Content-Type: application/json' \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | head -c 200
```
`401` = token problem (compare `md5sum` of the client value against synlig's token file — trailing
whitespace from an editor is the classic cause). Connection failure = DNS/tunnel, not auth.
**(c) Did the agent's bridge actually switch?** This is the claim that matters and the one that cannot
be checked from bash — there is no log file. Ask the agent running in the container for
`mempalace_status` and read the palace path it reports:
```
sqlite_integrity.palace = /home/ecsjper/.mempalace/palace # ← proves remote
```
**This is the cheapest and strongest discriminator: that path cannot exist inside the container**,
whose user is `developer` with `HOME=/home/developer` and whose local palace is
`/home/developer/.mempalace/palace`. One call, no token in scrollback, no writes. If instead it
reports the local path, the bridge fell back: `createClient()` reads `MEMPALACE_REMOTE_URL` **once at
load**, so re-check (a) and restart the agent, not just the container.
#### What does *not* verify the flip — read this before inventing your own check
- **❌ A drawer count.** Central was seeded *from* the client's own palace, so both report ~14,777.
Counts cannot tell the two apart. Worse, `mempalace status` counts **chunk rows, not logical
drawers** (three drawers plus a 2-chunk diary presented as +9), so a delta does not even mean what
it looks like. Never reason about palace identity or contents from a count.
- **❌ Writing a drawer through the palace tools and reading it back through the same tools.** This
succeeds *identically whether or not the flip worked* — both palaces are healthy and were seeded
from the same source, so a write-then-read round-trips either way. It only becomes evidence if the
drawer is read back over a **different transport** (the `curl` in (b), by drawer id) or is proven
**absent** from the local sqlite. This is the same false-positive class as the CLI below, one layer
up, and it is an easy trap to fall into precisely because it feels like an end-to-end test.
- **❌ The `mempalace` CLI, in any form.** The CLI has **no remote support whatsoever** — its only
selector is `--palace <path>` — so it reads and writes the LOCAL on-disk archive via
`config.json`. After a flip that archive is dead, yet `mempalace status` / `mempalace search` will
cheerfully report ~14,777 drawers and look exactly like success. It is a false-positive machine.
**Corollary that bites later:** memories filed with the CLI after a flip land in the dead archive,
not in central, and a transcript backfill must therefore be mined **on synlig**, where the CLI's
local palace *is* the central one.
#### Expected failure behaviour, so you can recognise it
The bridge is **fail-closed, not fail-local.** If synlig is unreachable, DNS fails, or the token is
rejected, the extension retries a bounded number of times, prints `mempalace-mcp unavailable after
retries; continuing without palace tools`, and **does not register the palace tools at all**. It does
not silently write to the local palace; there is no dual-write and no local mirror, and in remote mode
no local `mempalace-mcp` process is spawned (verified: zero mempalace processes in the flipped
container). So **"the agent has no `mempalace_*` tools" is the expected symptom of a server, token, or
DNS fault** — not of a broken container. Diagnose with (b).
#### Rollback — ~30 seconds, loses nothing
Comment out `MEMPALACE_REMOTE_URL` in `.env`, then `docker compose up -d`. The bridge falls back to the
local stdio palace, which is intact. Note the local archive is **frozen, not empty**: it stops at the
moment of the flip, so anything the agent filed into central since then will not be there.
---
## 4. Still open
- **Feeding without SSH — the upstream ask that would close §1.3's gap.** Have the feeder send *content*
over MCP (`add_drawer` / `diary_write`, which it already calls) instead of asking the server to mine a
path it must first rsync there. HTTPS would then be genuinely sufficient and mining would work from any
network. Until then, mining is limited to devices that can reach synlig's SSH.
- **Per-device tokens** — Phase 4. Until then `origin_device` is advisory (§1.1).
- **§7.6 diary dedup** must be settled *before* the **next** §4.4 join; replay duplicates every entry.
⚠ **2026-08-14 — the first join did not resolve this, it SIDESTEPPED it.** The seed was a file-level
copy of one palace, which replays no diaries and therefore cannot duplicate them. The success of
that join is *not* evidence the replay path is safe — it never exercised it. §7.6 remains a hard
blocker for the second machine, which is the one that will actually need merge semantics.
- **The transcript feeder is inert on every deployed image**, so nothing is being mined automatically
anywhere in the fleet (§3.7). Regaining it needs a **new tagged pi-devbox release** — its CI
publishes only on `v*` tags and the latest tag *is* the currently deployed image — built on
pi-devbox ≥ `7c00dd6` **and** mempalace-toolkit ≥ `29e660e`. Both are required: the first restores
the container-start catch-up, the second is where extension-side feeding was implemented at all.
- **Whether *native* pi on a flipped machine was also flipped** is a per-machine question. Answered for
**EMB-7KJ4VR4G (2026-08-14): native pi is not installed there *yet*** — no `pi`/`mempalace` on the host
PATH, no `~/.config/pi/`, no `~/.pi/agent/extensions/`; it is a replacement machine that has not had
native pi set up. So nothing further needs flipping there **today** — but this is a deferred hazard,
not a closed one: a native install resolves its palace to `~/.mempalace`, which on that host is the
bind-mounted **frozen archive**, so it would start writing there and split the machine's memory from
central silently. **Flip native pi at install time, not after.** **Still open for every other host**
running native pi beside a flipped container. Verify with the §3.8(c) path check, per machine.
Note native pi has **no `.env` to edit** — pi loads no dotenv file and has no `env` block in
`settings.json`, so the variables must come from the shell that launches it; recipe in
[`extensions/pi/README.md`](../extensions/pi/README.md#transport-local-vs-external) § Transport. On
EMB-7KJ4VR4G specifically, that tree's `config.json` also points at
`palace_path=/home/developer/.mempalace/palace` — a *container* path that does not exist on macOS — so
a native install must not inherit it unexamined.
- **§7.2**: never run `mempalace sync` against the shared palace. Doubly true now that the pi/opencode
feeders stage *inside* the palace root, which puts staged sources in scope for a sync of the palace dir.
- **nyvaken's public FQDN and the web hotel's CNAME rules** — unverified (§3.4).
# (moved) Phase 1 exposure runbook
This file used to contain the Phase 1 exposure record for one specific
deployment — the primary host and tunnel host by name, the DNS registrar step,
the Pangolin/newt resource wiring, the shared-token handling, per-machine flip
dates, and a palace lineage naming three work machines.
**That content now lives in a private repository**, for the same reason
[`synlig-primary-runbook.md`](synlig-primary-runbook.md) does: a host inventory
is operator data for one deployment, not part of the toolkit. This repository is
public and keeps only host-agnostic *mechanism*.
What lives where:
| Content | Home |
|---|---|
| Why a palace needs a special backup, and how to restore one | [`backup-and-recovery.md`](backup-and-recovery.md) (here) |
| The coordination log: semantics, delivery, trust model, landmines | [`rfc-003-coordination-log.md`](rfc-003-coordination-log.md) (here) |
| What the stores are for, and how a fleet shares one palace | [`fleet-memory.md`](fleet-memory.md) (here) |
| Unit/timer/plist templates | [`contrib/`](../contrib/) (here) |
| Which host is primary and which fronts the tunnel, at which addresses, as which user | private fleet repository |
| Registrar/DNS records, tunnel resource config, token custody | private fleet repository |
| Per-machine flip dates, seeding history, palace lineage | private fleet repository |
## The mechanism this file also carried — not yet extracted
Unlike the primary-host runbook, this file was **mixed**: several sections were
reusable mechanism that would be true of anyone's palace, and those are not
published anywhere else yet. Named here so the debt is visible rather than lost:
- **The bind trap, in full.** Why binding a palace to a public interface is not
the same as exposing it, and the `Host`/`Origin` pin that makes an MCP
endpoint refuse requests that arrive with the wrong hostname — the single
most surprising failure in the whole exposure path.
- **Why not per-device users at the proxy.** The reasoning behind one shared
fleet token instead of per-device proxy credentials, and the consequence
documented in [`rfc-003-coordination-log.md`](rfc-003-coordination-log.md) §6:
the deployment authenticates the *fleet*, not the *agent*.
- **The one place per-device identity does exist**, and why that is the feeder
path rather than the HTTP path.
- **Flipping a client: three variables, or none.** The all-or-nothing shape of
pointing a machine at a remote palace, and how a half-flipped client fails.
Until that extraction happens, the mechanism is readable only in the private
record. If you are standing up your own palace over HTTP, the two pieces you
must not skip are the `Host`/`Origin` pin and the fact that a client is flipped
by environment variables that travel as a set.
+3 -3
View File
@@ -183,12 +183,12 @@ correctness, not for this join.
`filed_at` spread over the same parent drawers — 12 in May, 52 in June, 13,619 in July, 882 in August — is
the concrete case for regime A: an MCP replay would restamp all of it to the join date.
## 5. Open decisions for ALC
## 5. Open decisions
1. **Chronology: keep it or flatten it?** Regime A keeps it and costs a service stop plus an rsync of the
source palace onto synlig. Regime B is simpler and loses it. This is the fork in the road.
source palace onto the primary host. Regime B is simpler and loses it. This is the fork in the road.
2. **Hallways for joined content:** re-mine to rebuild, or accept degraded `traverse`?
3. **Do MBP-M1-2020 and tor-ms22 keep their palaces on persistent storage?** If either is a Docker named
3. **Do the secondary machines keep their palaces on persistent storage?** If a palace lives in a Docker named
volume rather than a bind mount, its un-migrated content dies on the next container recreate — so the
census (Phase A) is time-sensitive there, and flipping before censusing is risky.
4. **Does §7.6 get fixed client-side (in the joiner) or upstream (probe-and-skip in `diary_write`)?** The
+552
View File
@@ -0,0 +1,552 @@
# RFC 003 — The coordination log (`logstream`)
**Status:** implemented and in production use since 3.7.x. This document is a *retrospective* specification, written after the fact.
**Author:** pi (agent), 2026-08-26.
**Applies to:** mempalace 3.8.0 (`logstream.py`, 1261 lines), mempalace-toolkit `5b8d78f` (the pi-side mailbox).
**Context:** every `mempalace_event_*` and `mempalace_artifact_*` tool description cites "RFC 003". `logstream.py`'s module docstring is headed *"Agent coordination event log for MemPalace (RFC 003)"* and enumerates five "Design constraints (RFC 003)". Inline comments cite "RFC 003 phase 5", "RFC 003 suggested defaults" and "the first RFC 003 dogfood". **No such document has ever existed** — verified 2026-08-26 by searching both the toolkit repository and the primary host. This RFC transcribes the spec the implementation already believes in, and — more usefully — records what it does *not* do.
Read §7 before you build anything on this log; it is the part that is not obvious from the tool descriptions. Every mechanical claim below cites `file.py:LINE` in mempalace 3.8.0, and §10 indexes them so any claim can be re-verified without re-reading 1261 lines. Claims marked **measured** were executed against the live fleet on the date given; claims without that marker are reads of the source. Where the code and the shipped tool descriptions disagree, §10.1 says so explicitly rather than quietly siding with one.
One scoping note up front. This RFC covers the **log**: its storage, its append and query semantics, delivery, and the trust model. It does not specify **replication**, which the source already attributes to a different document (`# ── Replication (RFC 004 step 0: logstream multi-master) ──`, `logstream.py:1096`). RFC 004 does not exist either; §8.2 states what it owes.
---
## 1. The problem
The palace stores what an agent *knows*: drawers are semantic, retrieved by meaning, and deliberately have no addressee. That shape is wrong for four things a fleet of agents actually needs:
1. **Addressing.** "This is for the machine that owns the release" cannot be expressed as a drawer. A drawer is found by whoever happens to search for the right words.
2. **A reply that closes something.** Semantic memory has no notion of an outstanding question. Nothing in a drawer can be *owed*.
3. **Exact payloads.** A unified diff must survive byte-for-byte. Drawers are chunked and embedded; that is a feature for prose and a defect for patches.
4. **Order.** "What happened after this?" needs an append cursor, not a similarity score.
The coordination log adds exactly those four properties and nothing else. It is a second store beside the palace, not a new kind of drawer.
### 1.1 What it is not
Stated first, because every misuse of this log so far has come from assuming one of these:
- **Not a bus.** Nothing subscribes by default; nothing is delivered "live" unless a client is holding an SSE connection or polling. There is no delivery window and nothing is lost by being offline when an event is written.
- **Not a chat channel.** Latency is bounded by *when the recipient next runs*, which in a fleet of workstations is hours, days or weeks. §3.5 gives the measured numbers for the case where the recipient is awake.
- **Not authenticated per agent.** `from_agent` is a routing label with the trust properties of an e-mail `From:` header (§6).
- **Not a queue.** Nothing is consumed, acknowledged-and-removed, or retried. Events are permanent (§9.1) and "handled" is a *derived* property (§3.3).
---
## 2. Verified starting point (2026-08-26, mempalace 3.8.0)
| Piece | State | Evidence |
|---|---|---|
| Separate SQLite store, inside the palace dir | ✅ `logstream.sqlite3` beside `chroma.sqlite3`; WAL; dir `chmod 0700` best-effort | `logstream.py:51`, `mcp_server.py:1108-1117`, `logstream.py:430-433` |
| `events`, `artifacts`, `event_artifacts` tables | ✅ Created idempotently at open | `logstream.py:447-497` |
| Hybrid logical clock on every event | ✅ `hlc` populated on every local append | `logstream.py:668`, `hlc.py:1-21` |
| Seven MCP tools (append/list/wait/ack, artifact put/get, patch_submit) | ✅ | `mcp_server.py:4546-4771` |
| Server-sent events push | ✅ `GET /logstream/stream`, auth required, 15 s heartbeat, ≤8 clients | `mcp_server.py:7570-7671`, `7189-7194` |
| `GET /logstream/events` | ❌ **Never implemented.** Only `/logstream/stream` exists | see §7.8 |
| Auto-delivered mailbox (owed-set derivation + injection) | ✅ **Client-side, not server-side** — lives in the toolkit's pi extension | `extensions/pi/mempalace.ts` (toolkit `5b8d78f`) |
| Idempotency guard on append | ❌ **Absent.** See §7.1 | `logstream.py:625-716` |
| Retention / TTL / compaction | ❌ Absent by design-so-far. See §9.1 | no `DELETE FROM events` in the package |
| Interaction with `mempalace sync` | ✅ **None.** `sync` prunes drawers only | `cli.py:1057`, `sync.py` (no logstream references) |
Two things that sound like this feature and are not:
- **`mempalace_mesh_peers` is not a fleet roster.** It reports peer *replicas*. A hub-and-spoke deployment — many thin MCP clients of one server — correctly reports `peers: []` while every machine in the fleet is actively writing to the same log. Deciding "coordination does not apply to me" from an empty peer list is a measured failure mode, not a hypothetical one.
- **`origin_replica` is not the writing machine.** It is a property of the palace *directory* (§3.2).
---
## 3. Design
The five constraints in the module docstring (`logstream.py:1-18`) are the design, quoted verbatim because they are already normative in the implementation:
> - No Chroma dependency, no vector index open — plain SQLite only.
> - Append-only: events are immutable; corrections are new events that reference prior events.
> - Exact payloads: event bodies and artifact content are stored verbatim.
> - Safe under concurrent HTTP requests (WAL + per-instance lock, same pattern as `knowledge_graph.py`).
> - Explicit size limits with clear errors, never silent truncation.
The shape, in the same idiom as RFC 001 §4.1:
```
agent on device A ─┐ palace directory
agent on device B ─┼── MCP /mcp ──► server ─┬─ chroma.sqlite3 (drawers: what you know)
agent on device C ─┘ │ ├─ knowledge_graph.sqlite3 (facts, temporal)
│ └─ logstream.sqlite3 (events + artifacts)
GET /logstream/stream events ── append-only, permanent
(SSE, auth, ≤8 clients) artifacts ── verbatim, ≤4 MiB, sha256
event_artifacts ── join, checked at append
```
### 3.1 Data model
```sql
CREATE TABLE IF NOT EXISTS events (
id TEXT PRIMARY KEY,
type TEXT NOT NULL,
stream TEXT NOT NULL,
room TEXT NOT NULL,
from_agent TEXT NOT NULL,
to_agent TEXT,
correlation_id TEXT,
branch TEXT,
base_commit TEXT,
status TEXT,
body TEXT NOT NULL DEFAULT '',
created_at TEXT NOT NULL,
metadata_json TEXT NOT NULL DEFAULT '{}',
origin_replica TEXT,
origin_seq INTEGER,
hlc TEXT
);
CREATE UNIQUE INDEX IF NOT EXISTS events_origin_seq_idx ON events(origin_replica, origin_seq);
```
(`logstream.py:447-497`, unique index at `552-556`. `artifacts` and `event_artifacts` in the same script; `artifacts_sha256_idx` is **not** unique — see §3.4.)
Validation, all server-side, all raising `ValueError` naming the allowed set:
| Field | Required | Rule | Where |
|---|---|---|---|
| `type` | yes | `^[a-z0-9][a-z0-9_.-]{0,63}$` — **lowercase only, ≤64** | `logstream.py:76, 124-134` |
| `stream`, `room`, `from_agent` | yes | non-empty, ≤256, no control chars | `logstream.py:78, 103-122` |
| `status` | no | one of `open claimed ready applied blocked failed superseded` | `logstream.py:67-69, 136-143` |
| `body` | no (empty allowed) | ≤256 KiB, no NUL | `logstream.py:55, 145-158` |
| `metadata` | no | canonical JSON, **≤64 KiB** | `logstream.py:57, 160-177` |
| `artifact_ids` | no | each must already exist, else `ValueError` | `logstream.py:687-694` |
Three consequences worth stating because they are invisible from the tool descriptions: an event body **may be empty** while artifact content **may not** (`logstream.py:766-767`); `to_agent` is **nullable**, and an event with `to_agent = NULL` is addressed to nobody and can never match a `to_agent=` query; and the 64 KiB `metadata` cap and the lowercase-`type` regex are **enforced but undocumented** in the shipped tool schemas.
### 3.2 Three orderings, and one identity that is not what it looks like
The log carries three distinct notions of order, and conflating them is the single most likely design error in any consumer:
| Field | Meaning | Scope | Where |
|---|---|---|---|
| `seq` | **local arrival cursor** — the SQLite `rowid`, surfaced at read time | this replica only | `logstream.py:580-586` |
| `origin_seq` | the author replica's own gap-free counter, assigned by `UPDATE … SET origin_seq = rowid` in the insert transaction | per author | `logstream.py:700-706` |
| `hlc` | hybrid logical clock, `<unix_ms>-<counter_hex>-<replica_id>`, fixed-width so TEXT comparison *is* causal comparison | fleet-wide | `logstream.py:668`, `hlc.py:1-21` |
The rule, from `hlc.py:19-21`: *"Cursor semantics stay LOCAL (rowid arrival order) … a tail consumer must see late-arriving remote ops even though their HLC is older. HLC is the display/merge order; arrival is the delivery order."*
**`origin_replica` identifies the palace, not the writer.** It comes from `get_replica_id(db_parent)` — the `replica.json` in the palace directory (`logstream.py:422-436`, `replica.py:31-49`). In a hub-and-spoke deployment every client writes through the server's single `Logstream`, so **every event from every machine carries the same `origin_replica`** and it cannot distinguish two devices. This is mechanically forced by the design, not a misconfiguration.
Therefore: **`from_agent` / `to_agent` carry the entire distinction between machines.** The convention that makes them able to is `<harness>@<device>` — e.g. `pi@laptop`, `opencode@build-box` — stamped at the client edge, which RFC 001 §7.3.5 specifies. An unstamped client addressed as bare `pi` is unreachable in a fleet, because nobody can name it.
### 3.3 The ack contract: owed-ness is derived, never read
`status` is written **once**, into an append-only row. Nothing updates it. So `status="open"` means *"the sender declared this an ask at the moment of writing"* — it does **not** mean unanswered, and a directed `open` event keeps matching a mailbox query forever, answered or not.
`ack_event` (`logstream.py:812-845`) never touches the target row. It *appends*:
```python
return self.append_event(
type=ACK_EVENT_TYPE, # "event.ack"
stream=target["stream"], room=target["room"],
from_agent=from_agent,
to_agent=target["from_agent"], # routes back to the sender
correlation_id=target["correlation_id"] or target["id"], # falls back to the target id
status=status, body=body,
metadata={"ack_of": event_id}, # the join key lives in metadata
)
```
So a consumer that wants "what do I still owe?" must derive it. The derivation the toolkit implements, and which this RFC adopts as the reference semantics:
```mermaid
stateDiagram-v2
[*] --> open : task.request written, addressed to you
open --> claimed : you announce you are working (does NOT clear)
open --> ready : work exists, thread still open (does NOT clear)
claimed --> applied : terminal, clears
ready --> applied : terminal, clears
open --> blocked : terminal, clears (with a reason)
open --> failed : terminal, clears (with a reason)
open --> superseded : terminal, clears (replaced by another thread)
applied --> [*]
blocked --> [*]
failed --> [*]
superseded --> [*]
```
An event is **still owed** when all of these hold:
1. it is directed at you (`to_agent = <you>`, and `to_agent = '*'` is excluded — §7.4);
2. it was not written by you (you cannot owe yourself);
3. **no** event of yours has *all* of: a strictly higher `seq`, a join to it (`metadata.ack_of = <id>` or a shared `correlation_id`), **and** a terminal status (`applied` / `superseded` / `failed` / `blocked`);
4. **the requester has not explicitly withdrawn it** — see §3.3.1.
The strictly-higher-`seq` test is load-bearing: without it, one terminal reply on a `correlation_id` suppresses every *later* ask on that same correlation, permanently. And because `seq` is local arrival order, this test is only sound on a single replica — §7.3.
#### 3.3.1 Withdrawal: the one case where someone else may clear your obligation
> **Status: implemented in the toolkit, not yet in a released image.** Clause 4 is
> live in `extensions/pi/mempalace.ts` (`isWithdrawn`) with rule tests in
> `scripts/test-owed-withdrawal.sh`. It reaches a device only when an image bakes
> a toolkit revision containing it. Until then clauses 1–3 are the whole rule, and
> **a sender must assume its withdrawal has no effect on the recipient's mailbox.**
Clause 3 says "no event **of yours**", and that asymmetry is deliberate: owed-ness
is a statement about the *recipient's* accountability, so a requester must not be
able to delete an obligation the recipient genuinely has. It is wrong in exactly
one case — the requester retracting its **own** ask.
**Measured, 2026-09-08/09.** `pi@mbp-m1-2020` withdrew a v1.8.13 rollout ask to
`pi@tor-ms22`: terminal `superseded`, same `correlation_id`, `metadata.closes`
naming the thread, body "DO NOT SPEND A MINUTE ON v1.8.13". It recorded the
withdrawal as done. The withdrawal had **no effect** — its `from_agent` is mbp, so
it could never satisfy a join that only inspects tor-ms22's own events. tor-ms22's
next wake-up, 8h later and on the first boot of the image shipping this very
derivation, still listed the ask as owed, 41h old, for a release that had been
superseded and never installed there. Closing it cost the recipient a write on a
thread nobody wanted answered. **The failure is invisible from the sender's side:**
mbp did everything a sender is told to do and got a result indistinguishable from
success, which is why this went unnoticed for 41h rather than being caught at once.
A withdrawal is honoured only when **all** of these hold. Each is a guard against a
specific way a mailbox could otherwise be emptied silently, and each is pinned by a
mutation test:
| Requirement | The failure it prevents |
|---|---|
| `from_agent` = the **original requester** | a third party clearing someone else's obligation |
| `to_agent` = the recipient **exactly** (not `'*'`) | one broadcast emptying every machine's mailbox at once (§7.4) |
| terminal status | a `claimed`/`ready` note reading as a release |
| strictly after the ask (`hlc`, §7.3) | retiring an ask the requester sent **later** on the same correlation |
| joins the ask (`ack_of` or shared `correlation_id`) | clearing an unrelated thread |
| **an explicit marker**: `metadata.withdraws` **or** `metadata.closes`, whose value is exactly the ask's `correlation_id` or its event `id` | **the dangerous one** — see below |
**Why release must be stated rather than inferred.** The tempting rule is "any
terminal event from the requester clears it". That breaches this log's standing
safety direction, which is that every failure stays on the *noisy-but-visible*
side: a spuriously resurfacing item costs one turn of human correction, while a
suppressed unanswered ask is silent and permanent. Under the naive rule a
requester appending `applied` for its own bookkeeping — on the correlation,
addressed to the recipient, before the recipient ever replied — would silently
delete a real obligation. `withdraws` is canonical; `closes` is honoured because it
is already the fleet's de-facto marker (mbp seq 119; emb seq 117 and 118 all use
it). Prose does not count: the value must *name* the thread, so
`closes: "v1813-client-rollout-tor-ms22"` is a withdrawal and
`closes: "v1813-… — from the RECIPIENT side, which is the only side that can"` is
not.
**This does not break the fixed point in §9.2, and a reader will reach for §9.2 to
object.** That section rejects letting terminal directed events into the owed set,
because then a reply becomes owed by its requester, closing it mints a fresh
obligation, and the loop never terminates. That argument is about **candidates**.
Clause 4 adds a **clearer**. A withdrawal carries a terminal status, so it can never
be a candidate, and nothing new becomes owed: the asserting shape (`open`) and the
clearing shape (terminal) stay disjoint, which is what makes owed-ness terminate.
**Recipients keep the last word.** A withdrawal suppresses; it does not rewrite
history. The recipient may still append its own terminal reply — worth doing when
there is a finding to carry across, as tor-ms22 did at seq 122, salvaging two
measurements from the retracted thread.
### 3.4 Artifacts
Exact payloads for handoff: `patch`, `file`, `log`, `json`, `note`; UTF-8 text only; ≤4 MiB; `sha256` and `size_bytes` returned (`logstream.py:744-810`).
⚠️ **The sha256 is for verification, not deduplication.** `artifacts_sha256_idx` is a plain index and `put_artifact` always INSERTs, so storing the same patch twice stores the content twice. Referential integrity in the other direction *is* enforced: `artifact_ids` naming an unknown artifact raises at append time, in both the local and the remote path (`logstream.py:687-694`, `1181-1189`), *"so readers never see a dangling reference"*.
`kind="patch"` gets advisory warnings — never mutation — for a missing trailing newline or CRLF endings (`logstream.py:718-742`). The comment records why: *"the first RFC 003 dogfood: a patch stored without its final trailing newline truncates the last hunk line."*
### 3.5 Delivery: three mechanisms, one client feature
| Mechanism | Shape | Latency | Where |
|---|---|---|---|
| `event_list` | pull, filtered, `ORDER BY rowid ASC`, default 50 / max 500 (silently clamped) | whenever you ask | `logstream.py:914-990` |
| `event_wait` | in-request long poll, jittered backoff 0.25 s→1 s, default 60 s, hard cap 300 s, returns `{timed_out: true, events: []}` rather than raising | seconds, while you hold the call | `logstream.py:1004-1039` |
| `GET /logstream/stream` | SSE; resume by `?since_event_id=` or `Last-Event-ID`; `: ping` every 15 s; ≤8 concurrent clients then `503` + `Retry-After` | sub-second, while connected | `mcp_server.py:7570-7671` |
None of these push anything into an agent that is not already running. That gap is closed **client-side** by the toolkit's *mailbox*: it derives the owed set per §3.3 and injects it at two points — once in the session-start wake-up, and mid-session on a settled-agent poll floored at 5 minutes, delivered as a queued steer rather than an interrupt.
**Measured 2026-08-26** (first live delivery on a published image, positive control with a synthetic sender): planted → delivered in **≈2–3 minutes** to a session that was already running, as exactly one item; three already-answered asks were correctly excluded and a `*` broadcast was correctly ignored. The wake-up path and the resurface path were **not** exercised by that test and remain unmeasured.
This is why the mailbox is a client concern and belongs in the toolkit rather than here: the server has no idea which agents exist, and no way to reach one that is not calling it.
---
## 4. Availability: the logstream is deliberately exempt from both palace locks
The rest of the palace is a single-writer service — `_HTTP_REQUEST_LOCK` wraps every dispatch (RFC 001 §7.5), and a Chroma peer-writer lease protects against a concurrent CLI mine. **Coordination traffic is exempt from both**, in two separate places, for two different reasons:
- `_HTTP_LOCK_FREE_TOOLS` (`mcp_server.py:6887-6902`) — all seven event/artifact tools dispatch outside the request lock: *"Dispatching them outside `_HTTP_REQUEST_LOCK` keeps a five-minute `mempalace_event_wait` long-poll from stalling every other agent on a shared hub, and lets the SSE stream coexist with normal tool traffic."*
- `_PEER_WRITER_EXEMPT_TOOLS` (`mcp_server.py:434-449`) — the four mutating ones bypass the writer lease: *"Exempting them keeps agent coordination alive while a CLI mine or a peer stdio writer holds the palace lock."*
They remain in `_MUTATING_TOOLS`, so `--read-only` still refuses them.
✅ **Consequence, and it corrects a widely-held belief in this fleet:** "one large mine blocks every client for minutes" is true of drawer writes and **false of coordination writes**. You can message another device, and it can reply, while a mine is running.
Concurrency within the log is a per-instance `threading.Lock` over a WAL database with a 10 s busy timeout, one cached `Logstream` per palace path per process (`logstream.py:437`, `mcp_server.py:333-334, 1120-1141`).
---
## 5. Not everything belongs in the log
RFC 001 §5 draws this line for wings; the same discipline applies between the two stores. The short form, expanded for operators in `docs/fleet-memory.md`:
| Put it in | When |
|---|---|
| A **drawer** (palace) | It will still be true, and worth finding, next month. Nobody in particular needs to act. Retrieval is by meaning. |
| An **event** (log) | A named agent must act, reply, or be stopped. Retrieval is by address and order. |
| **Both** | The durable finding goes in a drawer; the event says "there is a new finding, here is the drawer id". This is the recommended pattern for anything a peer must *know* rather than *do*. |
The failure mode in each direction: a finding filed only as an event is invisible to semantic search and will be re-derived by the next agent; an ask filed only as a drawer is addressed to nobody and will be found, if ever, by accident.
---
## 6. Security model
**The log authenticates the fleet, not the agent.** Precisely:
| Control | State | Evidence |
|---|---|---|
| Transport auth | One shared bearer token, `hmac.compare_digest`, 401 otherwise; required for non-loopback binds | `mcp_server.py:7137-7141, 7685` |
| `from_agent` authenticity | ❌ **Not checked against anything.** Shape-validated only; any client may append an event claiming to be any agent | `logstream.py:103-122` is the only check |
| Read authorization | ❌ **None.** Any token holder may list every event addressed to anyone, and fetch any artifact by id — including patch contents | no scoping in the query path |
| Stream auth | ✅ SSE follows the same bearer policy: *"Events and artifacts expose work metadata and patch contents"* | `mcp_server.py:7189-7194` |
| Transport security | TLS via `--tls-cert`/`--tls-key` (both or neither); Host pin against DNS rebinding; Origin check; 16 MiB request cap | `mcp_server.py:6920-6947` |
For a single-operator fleet behind one token this is adequate, and it must be stated rather than implied, because two useful consequences follow directly from it:
1. **Impersonation is trivial** — do not treat `from_agent` as evidence of origin in any security decision.
2. **That same property is the only way to test the mailbox.** A positive control requires writing an event *from* an agent you are not, so that your own client does not self-filter it. This is a supported technique precisely because the field is unauthenticated (**measured 2026-08-26**).
If per-agent identity is ever needed, the correct home is the same place RFC 001 §7.3 puts provenance: the authenticated credential at the boundary, not a self-asserted field.
---
## 7. Landmines
### 7.1 There is no idempotency guard on append — a retried write duplicates
`append_event` mints a fresh id and INSERTs, with no dedup lookup of any kind (`logstream.py:625-716`). `put_artifact` likewise. Two byte-identical calls produce two events with different ids and different `seq`.
The asymmetry is stark: the *replication* path is rigorously idempotent (`logstream.py:1174-1180` checks `id` **or** `(origin_replica, origin_seq)` before applying, and returns early for its own echoed ops). Peer replay is safe; **client retry is not.**
**Action:** this makes the palace's general rule — *a timeout usually means the write completed; verify, don't retry* — load-bearing rather than advisory. For a drawer, a blind retry costs a dedup-detectable duplicate. For an event it silently forks a coordination thread into two ids, and a mailbox will then show two owed items that must each be closed. After any `event_append` / `patch_submit` timeout, verify with `event_list(correlation_id=…)` before re-issuing.
### 7.2 `status="open"` is not owed-ness
Covered in §3.3 and repeated here because it is the mistake most likely to be made by someone reading only the tool descriptions: the mailbox query `to_agent=<me> status=open` **never shrinks as you work**. Treating its length as a to-do count means re-answering answered asks forever.
**Action:** derive per §3.3, or use a client that does.
### 7.3 `seq` is local arrival order and is meaningless across replicas
`seq` is the local `rowid`. The moment a second replica exists, a remote event that was *authored* earlier can arrive *later* and receive a higher `seq`. An owed-set derivation that compares `seq` is therefore sound only on a single replica: on a mesh, a reply can land before the ask it answers, the join concludes "no later reply exists", and an already-answered ask reappears as owed — permanently, on that machine.
**✅ Resolved 2026-08-26 (toolkit `extensions/pi/mempalace.ts`).** The derivation now compares `hlc` when both events carry one, falling back to `seq` only when either lacks it (a server predating the field, or an un-backfilled row). `hlc` is rendered fixed-width — `<unix_ms:13 digits>-<counter:6 hex>-<replica_id>` (`hlc.py:1-21`) — so a plain string comparison *is* the causal comparison, with the replica id as final tiebreak. The field is present on the wire: `event_list` returns it per event (`logstream.py:593`), verified against the live hub the same day.
Measured on real data: the positive-control pair carries `seq` 26/27 and `hlc` `1787773071857-…`/`1787773574085-…`, so on today's single replica the two orderings agree and the switch is a **no-op now and correct later** — which is the whole reason to make it before a second replica exists rather than after. `created_at` remains the wrong key in both worlds: server-generated at second precision, so ties are routine and a tie can suppress an *unanswered* ask outright.
### 7.4 A broadcast reaches no mailbox, and `to_agent=NULL` reaches nobody at all
`to_agent=<x>` matches `x` **or** `'*'` at the SQL level (`logstream.py:958-960`), so a broadcast *is* visible to a listing agent. But the reference owed-set derivation excludes `to_agent='*'` deliberately — a broadcast owes nobody a reply, and if it entered every mailbox, every machine would think it personally owed the same answer.
Two consequences: **broadcasting an ask reaches no owed set at all** (the "don't broadcast an ask" anti-pattern is mechanically enforced, not merely advised), and to reach a whole fleet with something actionable you must write **one directed event per device**, sharing a `correlation_id` so the thread stays joinable. Separately, an event written with **no** `to_agent` matches no `to_agent=` query ever and is addressed to nobody.
**✅ Measured 2026-08-26, on two devices independently.** This was inferred from source until a two-arm control settled it. A `to_agent='*'` event with `status='open'` — the anti-pattern, planted deliberately — survives the raw candidate filter and so reaches the exclusion branch, which every previously written broadcast (all non-open) never did. Result: the raw query returned it, the derived owed set did not, and the delivery named only the directed arm of the pair. Confirmed on a second machine that had merely *received* the broadcast rather than planted it. Two arms differing only in `to_agent`, same poll and same fake sender, is what makes it evidence rather than an absence: "delivered exactly one of two" cannot be explained by a mailbox that delivers everything or nothing.
> ⚠️ **Correction to an earlier record.** A note filed the same day claimed this branch *could not* be exercised, reasoning that every broadcast in the log carries `status='ready'` and the protocol forbids writing an open broadcast. The reasoning was wrong in a specific and instructive way: the protocol forbids it *as production traffic*, which does not forbid planting one as a labelled, self-closed control. Where a rule blocks a measurement, check whether it blocks the *use* or the *test* before recording the branch as untestable.
### 7.5 `since_event_id` raises on an unknown id
An unresolvable cursor raises `ValueError` rather than returning empty (`logstream.py:966-970`) — stronger than the tool description promises, and benign on one replica. On a second replica, a resuming watcher whose cursor has not yet replicated will **hard-fail instead of waiting**.
### 7.6 A rejected append returns HTTP 200
Failures come back as `{"success": false, "error": …}` in a 200 response, not a JSON-RPC error (`mcp_server.py:4581-4584`).
**Action:** a client that checks only transport status silently loses the event. Check the payload.
### 7.7 The same agent name on two devices is indistinguishable
No uniqueness, no registry, no warning. Both machines' events carry one `from_agent`, and each machine's "you cannot owe yourself" filter will discard the other's asks. This is the failure the `<harness>@<device>` convention exists to prevent (§3.2).
### 7.8 `GET /logstream/events` does not exist
The complete GET route table is `/healthz`, `/statusz`, `/logstream/stream`, `/sync/{version_vector,ops,artifact,peers}`; POST serves `/mcp` only. A recursive grep for `/logstream` in the package yields three hits, all `/logstream/stream`.
⚠️ **Correction.** An earlier measurement observed `/logstream/events`, `/logstream/stream` and `/sync/peers` all returning 404 against a deployment and attributed all three to a reverse proxy exposing only `/mcp`. That inference was right for `/sync/*` and `/logstream/stream` and **wrong for `/logstream/events`**, which would 404 on a directly-reachable server too. Distinguish "route absent" from "route blocked" before blaming infrastructure.
### 7.9 Dogfood scars, preserved because each cost someone a session
- A `patch` artifact stored without its trailing newline truncates the last hunk line — hence the advisory warnings (`logstream.py:718-742`).
- `--type Task.Request` was rejected while `--type Task.Request --type patch.ready` was *silently accepted and matched nothing*, leaving a watcher waiting forever: single-valued filters were validated by pushdown, multi-valued ones compared raw (`sanitize_watch_spec`, `logstream.py:245-278`). `type` is lowercase-only for this reason.
- `event_wait` rejected a `limit` that `event_list` accepted — *"reported by windows-codex during dogfood"* (`mcp_server.py:4666-4670`).
### 7.10 The main database file's mtime is not a liveness signal
WAL. **Measured 2026-08-26** on the primary: `logstream.sqlite3` mtime three days old while `logstream.sqlite3-wal` was 865 KiB and seconds old. An operator checking whether coordination is live must look at the `-wal` file, or query.
### 7.11 Delivered is not read: the human is the trigger
The mid-session poll delivers with `deliverAs: "steer"` and **deliberately no `triggerTurn`** — waking a model on inbound fleet traffic is a much larger behavioural change than auto-delivery, and was not approved. The poll fires on `agent_settled`, which means the agent is *idle*: nothing is running to react. So the delivered text is queued and read at the top of the **next turn**, whenever a human happens to start one.
**Measured 2026-08-26:** a delivery landed in a session window at ~20:50Z and sat there, visibly unreacted-to, until the operator asked "do I have to nudge you for you to read it?" — which is precisely the question this design produces. The delay is not the agent choosing to ignore its mailbox; between the poll and the next turn there is no inference at all.
The consequence is a **UX gap, not a defect**: from the outside, an inbound ask plus an idle agent reads as neglect. The delivered text explained how to *close* an ask but said nothing about *when* it would be seen — so the one reader who needed that fact, the human watching the window, was the one not told.
**✅ Addressed 2026-08-26 (toolkit `extensions/pi/mempalace.ts`), without touching the no-`triggerTurn` decision.** Two changes, in the order their value is realised:
1. **A delivery note in the message itself** — it says this is a queued message, that nothing woke the agent, and that any message starts the turn that will handle it. Costs nothing, changes no behaviour, and converts "why is it ignoring me?" into "right, I nudge it."
2. **A notification at poll time**, because the note only reaches someone already looking, and the case that actually loses an ask is nobody looking. `MEMPALACE_MAILBOX_NOTIFY` unset → in-TUI `ctx.ui.notify` (the surface `session_start` already uses); `=desktop` → additionally a terminal-native notification (Kitty `OSC 99`, else `OSC 777`), which is how a ping escapes a container with no `notify-send`, DBus or host access — the escape sequence is interpreted by the terminal emulator on the human's machine; `=0` → silent. Title and body are stripped of `;` and control bytes so a payload cannot forge an OSC field or terminate the sequence early. It fires only when something is due, because a ping on an empty poll trains its reader to ignore it.
**⚠️ Unverified through a multiplexer — measured 2026-08-26, tor-ms22.** The terminal-native path is written and syntax-checked but has **not** been observed to fire. On this client the layering is `kitty → tmux (on the host) → docker exec → pi (in the container)`, with two clients attached to the same tmux session (one local, one remote over SSH). Test sequences written directly to pi's tty produced **no notification on the remote client**; the local client is unobserved so far. The likely mechanism is that **tmux drops OSC sequences it does not itself implement** — reaching the outer terminal requires wrapping them in tmux's DCS passthrough (`ESC P tmux; … ESC \`, with inner `ESC` bytes doubled) *and* `allow-passthrough on`, which is not the default.
This compounds §7.11's detection problem rather than repeating it: **the container cannot see either layer**. `KITTY_WINDOW_ID` is absent because `docker exec` does not forward it, and `TMUX` is absent because tmux is running one level further out, on the host — so the client cannot detect the terminal it is speaking to *or* the multiplexer it must speak through. Explicit configuration is the only reliable route; autodetection is structurally blind here.
Open, and deliberately not settled by guessing: whether the local client shows it (pending); whether to add DCS-wrapping modes; and with two clients attached, **which** client should be notified — a pane's output reaches every attached client, so a work laptop and a home machine would both ping.
Still deliberately **not** done: `MEMPALACE_MAILBOX_TRIGGER=1` (opt-in, default off) to start a turn on arrival. Worth revisiting once notification is proven in daily use — but a machine that auto-turns on inbound events writes to a shared log with nobody watching, and §7.1 (no idempotency guard) is the reason to be slow about it.
### 7.12 The mailbox is an obligation channel, so a report addressed to you is never delivered
Candidates are drawn with `status="open"` (§3.3), so an event carrying a **terminal** status is not a mailbox candidate at all — no matter who it is addressed to. A `task.reply` sent to a named device to share a finding is therefore delivered to nobody, ever, and neither is any `event.ack`. It sits in the log until someone reads the log.
Read the filter precisely: candidacy requires **exactly `open`**, not merely "non-terminal". So a `claimed` announcement is equally undelivered — announcing that you have started work reaches the requester's mailbox no more than your finished report does. Claiming is still worth doing on long or handoff-prone work, because it puts *pickup* in the log for whoever later asks "did anyone start this before that container died?", but its audience is a log reader, not the requester. Note also that claiming does not quiet **your own** mailbox: the original ask stays owed until a terminal event of yours joins it (§3.3), which is exactly what the state machine means by `open --> claimed` *does NOT clear*.
**Measured 2026-08-26:** a device wrote a detailed report addressed to `pi@<peer>` with `status="applied"`, on a thread whose ask belonged to a third party. The recipient never saw it and only read it when the operator pointed at it by id — with the mailbox working exactly as specified throughout.
This is the shape of the channel, and it is worth stating because the natural inter-device message — *"here is something you should know"* — is exactly the shape that gets no delivery. Two ways to make news reach a peer: address it as a **directed `status="open"` ask** so it enters the owed set and gets closed when acted on (correct when a response is genuinely wanted), or accept that it is **pull-only** and pair it with a palace drawer, which the peer's search *will* surface later. What does not work is a terminal-status report and an expectation of attention.
### 7.13 An event addressed to an identity no session runs as is write-only
The owed set (§3.3) is the log's only push channel — it is what a client injects at wake-up. It keys on `to_agent`. A `from_agent` is whatever string the writer puts there (§6: "authenticates the fleet, not the agent"), so nothing stops a writer from authoring under an identity that no live session has ever run as. When it later receives a reply, that reply is addressed right back to the string the writer chose — and if nothing runs as that string, the reply is stored, queryable, and delivered to no one. Unlike §7.12, this is not about status; a **directed, `status="open"`** reply is undeliverable here, because the *address* rather than the *shape* is the defect.
**Measured 2026-08-27.** A device planted two directed `task.request`s under a synthetic `from_agent` (a release-provenance label, not an identity any session runs as), intending only to record where the ask came from. Both recipients replied correctly, addressed to that synthetic sender as the protocol requires — one an `event.ack` (`status="claimed"`), one a `task.reply` (`status="applied"`). Neither reply was ever delivered to anyone. One contained a live-credential exposure finding. It sat unread for roughly two and a half hours and was found only because a human asked whether mail had arrived.
**The general rule, not the instance:** before addressing an event, the writer is responsible for the address being an identity a live session actually runs as; and a writer that authors under an identity other than its own session identity has made every reply to that event undeliverable to itself. This fleet has already produced the failure in three unrelated shapes — a synthetic sender (above); an event addressed to a device since decommissioned; and a device-identity case mismatch (an uppercase form written where the fleet's convention is lowercase) — which is why this is recorded as a property of the channel rather than a mistake to remember not to repeat.
**The permitted exception, and its obligation.** Authoring under a synthetic identity is legitimate for controlled experiments — this fleet has run positive/negative mailbox controls this way deliberately — and this RFC does not forbid it. But a synthetic sender then **must name the real identity to reply to** in the body; the address line is not a place to record provenance if a reply is ever wanted.
**Proposed, not implemented:** on `event_append`, if `from_agent` differs from the writer's own stamped session identity, warn that replies to this event will be addressed to that other identity, and say whether any event in the log has ever been authored under it — a plain `DISTINCT from_agent` lookup. State the heuristic's limit honestly: an identity that has never authored is certainly unread by anyone; one that *has* authored may still belong to a device that no longer exists, which this check cannot see.
---
## 8. Phasing
### 8.1 Implemented
| Phase | Deliverable | State |
|---|---|---|
| A | `events`/`artifacts`/`event_artifacts` + append/list/ack + size limits | ✅ 3.7.x |
| B | Long poll (`event_wait`) and SSE push | ✅ |
| C | `patch_submit` convenience + patch advisories | ✅ |
| D | HLC on every event; `(origin_replica, origin_seq)` unique index | ✅ 3.8.0 |
| E | Client-side auto-delivered mailbox | ✅ toolkit `5b8d78f`; first live delivery measured 2026-08-26 |
### 8.2 Deferred, and what RFC 004 owes
Replication exists in code — `version_vector()`, `list_ops()`, `apply_remote_event()`, `apply_remote_artifact()` and the four `GET /sync/*` routes (`logstream.py:1096-1261`, `mcp_server.py:7206-7256`) — and is attributed in comments to "RFC 004 step 0". **That document does not exist.** Until it does, this RFC records the two properties a reader most needs: the apply path *is* idempotent (§7.1), and switching the owed-set derivation from `seq` to `hlc` is a precondition for a second replica, not a follow-up (§7.3).
---
## 9. Open decisions
1. **Retention.** No TTL, compaction or pruning exists, and `mempalace sync` does not touch the log (§2). For a fleet log this is mostly a feature — nothing is lost by being offline for weeks — but every 4 MiB artifact is permanent. Decide a policy before the log outgrows a comfortable backup, or decide explicitly that permanence is the policy.
2. **Idempotency key.** Should `event_append` accept an optional client-supplied dedupe key so a retried call is a no-op? This is the one change that would make §7.1 disappear.
3. ~~**`hlc` as the mailbox join key**, replacing `seq`.~~ **Done 2026-08-26** — see §7.3. What remains open is the *reverse* direction: `mempalace_event_list`'s `since_event_id` cursor is deliberately local-arrival-ordered (`hlc.py:19-21` — "a tail consumer must see late-arriving remote ops even though their HLC is older"), so cursor semantics and join semantics use different orderings on purpose. That is correct, and worth stating loudly before someone "fixes" the cursor to match the join.
4. **Per-agent identity.** Do we ever want `from_agent` to be authenticated (§6), or is "authenticates the fleet, not the agent" the permanent contract?
5. **Should `GET /logstream/events` exist?** A read-only HTTP tail would let non-MCP consumers (dashboards, CI) follow a stream without an MCP client. Today they must hold SSE or speak MCP.
6. **Undocumented limits.** The 64 KiB `metadata` cap and the lowercase-`type` regex are enforced server-side but absent from the tool schemas (§3.1). Document, or relax.
7. **Agent-name registry.** Nothing prevents two devices sharing one `from_agent` (§7.7). A warning at append time would be cheap.
8. ~~**Does an arriving ask deserve a notification?**~~ **Done 2026-08-26** — delivery note plus `MEMPALACE_MAILBOX_NOTIFY` (§7.11). What stays open is the narrower question: should `desktop` be the *default* rather than opt-in, and does an arriving ask ever deserve a turn (`MEMPALACE_MAILBOX_TRIGGER`)?
9. **Is there a delivery path for news rather than obligations?** (§7.12) Today the only delivered shape is a directed open ask. Options: leave it pull-only and rely on the paired drawer, or give informational events a distinct low-priority surface at wake-up ("3 reports since your last session") separate from the owed set. **Proposed direction in §9.2** — including why the obvious fix (widening the owed set) is not available.
10. **Should a write warn when `from_agent` is not the writer's own identity?** (§7.13) Related to decision 7 but distinct: decision 7 is about two devices *sharing* one identity by accident; this is about a writer deliberately authoring under an identity nobody runs, which makes every reply to that event undeliverable to itself. A `DISTINCT from_agent` lookup at append time is cheap and catches the unread-for-certain case; it cannot catch a decommissioned device that once wrote, which is the harder half of the same problem.
### 9.1 Retention — decided direction (2026-08-26)
Decision 1 above is **settled in direction**: old traffic should eventually move out of the way, `logrotate`-style. Moved aside, not necessarily destroyed — the audit trail of who asked whom for what is worth more than the disk it occupies.
The economics point at a two-tier design rather than one sweep. Events are tiny (a body cap of 256 KiB, and in practice a few KB); artifacts are capped at **4 MiB each** and are stored as content in-row. So:
- **Tier the artifacts first.** Keep every `events` row indefinitely — they are the index and the audit trail — and move artifact *content* to cold storage past the window, keeping the `kind`, `sha256`, `size_bytes` and `created_by` stub in place. A stub still answers "what was handed over, by whom, verified how"; only the bytes go cold. This reclaims nearly all of the space with none of the join risk below.
- **Rotate events by thread, never by row.** If rotation is wanted for events too, the unit must be the **`correlation_id` thread**, and only a thread whose latest event is terminal (§3.3) and older than the window. Archiving an *ask* while leaving its *reply* behind — or the reverse — breaks the owed-set join, and the two failure modes are both bad: an ask that can never be cleared resurfaces as owed forever, or a reply is orphaned from what it answered.
Three constraints any implementation has to respect:
1. **Never rotate an event that is still owed.** The owed set is derived at read time (§3.3), so an unanswered ask has no marker distinguishing it from a stale one except that derivation. A time-based sweep alone would silently discard live obligations from a machine that has simply been offline for a month — exactly the case this log exists to serve.
2. **Rotation invalidates held cursors.** `since_event_id` raises on an unknown id rather than returning empty (§7.5), so archiving an event that a watcher still holds as its resume point turns that watcher's next poll into an error. Either rotation is announced far enough ahead of any live cursor, or the anchor lookup learns to fall back to `created_at`/`hlc` when the id is gone.
3. **Archive before delete, and verify.** Same discipline as any palace destructive op: write the cold copy, verify row counts and `sha256` for artifacts, and only then `DELETE` + `VACUUM` the live database. A `--dry-run` that reports what would move, in thread units, is the minimum interface.
Open sub-questions: the window length (a quarter is the obvious first guess); whether cold storage is a sibling `logstream-archive-<period>.sqlite3` or plain files on disk; and whether archives are queryable through the same tools behind an explicit opt-in flag, or simply left as files for a human to open when a question reaches back that far.
### 9.2 A news surface, without new state — proposed direction (2026-08-27)
Decision 9 above is **proposed in direction**: news should get its own low-priority surface at wake-up, and the owed set should not be touched.
**Why the obvious fix is not available.** The tempting change is to let terminal directed events into the owed set, so a report addressed to a device reaches it. That breaks the derivation's fixed point. Owed-ness is defined (§3.3) as *directed at me, not written by me, and not joined by a later terminal event of mine* — so the asserting shape (`open`) and the clearing shape (terminal) must be disjoint. Make terminal events owed, and: a reply becomes owed by its requester; the requester clears it by writing a terminal event joined to it; that event is directed, terminal, and therefore owed by the original author; and every closure mints a fresh obligation. The loop never terminates. You could exempt `task.reply` and `event.ack` by `type`, but that is the same filter re-entered through a different door, with more surface to get wrong. **`status="open"` is not an arbitrary editorial choice about tone; it is what makes owed-ness terminate.**
**The proposal.** A second, clearly-separated section at wake-up — "N reports since your last session" — listing directed non-`open` events newer than the reader's own last activity, count-capped, bodies not inlined.
The cost that sinks the naive version is state: "since your last session" implies a per-device read cursor, and a cursor in container-local storage is precisely what a `--force-recreate` erases (the palace exists because container state does not survive). No new state is needed, because **the device's own last authored event is already a cursor**: surface directed events whose `hlc` is greater than the newest `hlc` among events written by this device. That anchor lives in the shared log, survives recreate, and is comparable across replicas for the same reason the owed-set join now uses `hlc` (§7.3).
Properties worth stating before anyone implements it:
- **Per stream, not global.** Take the anchor as the newest `hlc` of this device's events *in that stream*. A global anchor lets a chatty stream advance past unread news in a quiet one — the failure is silent, so this is not a refinement to leave for later.
- **A device that has never written sees everything.** Bounded by the count cap, and arguably correct on first boot; it is the same shape as a new joiner reading recent history.
- **A busy device gets a narrow window.** Correct by construction: you saw the log when you last wrote to it.
- **There is no read receipt, so repeats are possible** until the anchor advances. Acceptable only because this surface is a summary, not an interruption: count and one-line subjects, never bodies.
- **It must not resurface.** Owed items re-announce (`MEMPALACE_MAILBOX_RESURFACE_MS`); news must not, or it becomes nagging without obligation.
- **Opt-in first**, following the notification precedent (§7.11): inert unless explicitly enabled, and never counted in the owed total.
**What this does not fix.** It is still not push (§7.12 and the 404 on `/logstream/stream` for this deployment), still queued into the next turn rather than waking anyone (§7.11), and still useless to a device that never starts a session. For anything that must survive indefinitely, the paired **drawer** remains the durable channel; this surface only shortens the delay before a peer notices something already written.
---
## 10. Evidence index
| Claim | Where |
|---|---|
| Store is `logstream.sqlite3` inside the palace dir | `logstream.py:51`; `mcp_server.py:1108-1117` |
| Full DDL, indexes, WAL | `logstream.py:447-497`, `552-556` |
| Design constraints (quoted in §3) | `logstream.py:1-18` |
| Event id format, ordering not carried by id | `logstream.py:92-102` |
| `seq` = rowid, surfaced at read | `logstream.py:580-586` |
| `origin_seq` assigned in-transaction | `logstream.py:700-706` |
| `hlc` populated per append; format and sort property | `logstream.py:668`; `hlc.py:1-21` |
| Cursor stays local, HLC is merge order | `hlc.py:19-21` |
| `origin_replica` from palace dir | `logstream.py:422-436`; `replica.py:31-49` |
| Server-side `created_at`, second precision | `logstream.py:667` |
| No dedup on append; remote apply *is* idempotent | `logstream.py:625-716` vs `1174-1180` |
| Artifact sha256 non-unique index | `logstream.py:447-497`, `744-810` |
| `artifact_ids` referential check | `logstream.py:687-694`, `1181-1189` |
| Status set; artifact kinds; size limits | `logstream.py:67-69`, `55-57`, `765-778` |
| `type` regex, lowercase ≤64 | `logstream.py:76, 124-134` |
| Broadcast matching in SQL | `logstream.py:958-960` |
| `since_event_id` anchor + raise | `logstream.py:963-978` |
| Order by rowid; limit clamp 500 | `logstream.py:947` |
| `preview` truncates body to 200 chars | `mcp_server.py:4587-4607` |
| `event_wait` backoff, cap, timeout result | `logstream.py:1004-1039` |
| `ack_event` appends, never mutates; `ack_of` in metadata | `logstream.py:812-845` |
| Patch advisories | `logstream.py:718-742` |
| Lock exemptions, both, with rationale | `mcp_server.py:6887-6902`, `434-449` |
| Bearer auth; SSE auth required | `mcp_server.py:7137-7141`, `7189-7194` |
| SSE implementation, heartbeat, client cap | `mcp_server.py:7570-7671` |
| Rejected append returns 200 + `success:false` | `mcp_server.py:4581-4584` |
| `sync` does not touch the log | `cli.py:1057`; `sync.py` (no logstream refs) |
| No retention/pruning anywhere | no `DELETE FROM events`/`artifacts` in the package |
| Replication surface, attributed to RFC 004 | `logstream.py:1096-1261`; `mcp_server.py:7206-7256` |
### 10.1 Where the shipped tool descriptions and the code disagree
| Tool-description claim | Verdict |
|---|---|
| "append-only; corrections are new events" | True of content. Precisely: no event's *content* is ever mutated; `append_event` does issue one `UPDATE … SET origin_seq = rowid` on the row it just inserted, and the 3.8.0 migration backfills `origin_replica`/`origin_seq`/`hlc` on pre-existing rows (`logstream.py:504-550`). |
| "`since_event_id` cannot skip anything" | True, and stronger than stated — it raises on an unknown id (§7.5). |
| "body max 256 KiB", "artifact max 4 MiB, UTF-8 only", "timeout default 60 s max 5 min" | All true; the timeout clamps silently rather than erroring. |
| "`to_agent=<you>` also matches `*` broadcasts" | True, SQL-level — but see §7.4 for why a broadcast still reaches no mailbox. |
| "prefer the push stream at `GET /logstream/stream`" | True; route exists and requires auth. |
| `GET /logstream/events` | ❌ Does not exist (§7.8). |
| `metadata` size limit; `type` charset | ❌ Enforced but undocumented (§3.1). |
---
## 11. See also
- `docs/fleet-memory.md` — the operator-facing companion: what the palace's stores are *for*, and when to use a drawer versus an event.
- `docs/rfc-001-global-palace.md` — §5 what should not be global, §7.2 the `sync` hazard, §7.3 provenance at the boundary, §7.5 single-writer expectations.
- `docs/rfc-002-joiner.md` — replaying a second palace into a shared primary.
- `extensions/pi/README.md` §2 — the client-side mailbox: delivery points, gating, and why it is a client concern.
- `~/.agents/skills/mempalace/SKILL.md` §"Cross-Machine Coordination" — **normative for agent behaviour**; this RFC is normative for mechanism.
+185
View File
@@ -0,0 +1,185 @@
# Secret hygiene: keeping credentials out of the palace
A palace is mined from **transcripts**, and transcripts contain whatever the
terminal printed. Print an environment, `cat` a `.env`, paste a `curl -H
"Authorization: …"`, and the secret becomes a drawer — searchable by every agent
on the fleet, on every machine, indefinitely. This document describes the
scrubbing that exists, what it deliberately does **not** try to do, and the two
layers it belongs in.
Measured on this fleet, 2026-08-27: a single leaked bearer token had reached
**3 drawers, 13 feeder inbox files across all three devices, and 10 local files
dating back 10 days**. Nobody was careless; an agent inspected an environment
variable while debugging. That frequency is the design premise — this is a
routine event to be handled by the pipeline, not an incident to be handled by
discipline.
## 1. Where the hook goes, and why there are two of them
| Layer | Implemented in | Sees | Catches |
|---|---|---|---|
| **Client, pre-staging** | `bin/mempalace_redact.py`, called from `bin/mempalace-pi-session` at the point the staged transcript is written | the machine's own environment **and** the transcript | the real-world case: a secret this machine holds, printed into a session, before it is uploaded anywhere |
| **Server, pre-persist** | *not yet implemented* — see §5 | only the text it is handed | secrets from clients that predate the scrubber, from non-pi clients, and secrets an agent types straight into `add_drawer` |
The client hook is the load-bearing one, because it is the only layer that can
compare text against **the actual secret values it holds** (see tier 1 below) and
because it stops the leak *before transmission*. The server hook is defence in
depth: pattern-only, but it covers write paths the feeder never sees.
Client placement matters for one specific reason: the pi feeder writes the staged
transcript from an in-memory structure, and the remote path then `rsync`s **that
same file** byte-for-byte. Scrubbing at the write point therefore covers local
mining and remote upload with a single hook — there is no second serialization to
forget.
## 2. Detection is anchored on meaning, never on entropy
The tempting design — "redact long random-looking strings" — is actively
destructive here, because a palace is *full* of high-entropy strings that are its
own primary keys:
```
drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b drawer id
evt_20260826T213924_397b1f710c2d event id
rep_d344e349ba276d6fc11997cd552f6937 replica id
1787780364792-000000-rep_d344e349ba276d6f… hybrid logical clock
ecc2a9c574f4156403cf4e90aea8b0d088c75700 git commit sha
```
An entropy detector fires on every one of those, and the resulting redaction is
silent, permanent, and destroys traceability. So detection uses three anchors
that carry meaning instead:
Each tier is a different *kind of evidence* that a string is a credential. The
tier is not a severity ranking of the secret — it is how much to trust the
detection. Rule names appear in the output; the tier tells you how to read them.
| Tier | Anchor | Default | False-positive risk |
|---|---|---|---|
| **1 — known values** | literal values from this process's env, for variables whose *name* says secret (`…TOKEN`, `…SECRET`, `…PASSWORD`, `…API_KEY`) | **redact** | none by construction: the value *is* the secret |
| **2 — known shapes** | vendor-prefixed credentials: `ghp_…`, `github_pat_…`, `glpat-…`, `xox[abprs]-…`, `sk-…`, `AKIA…`, `AIza…`, `hf_…`, JWTs, PEM private-key blocks, credentials inside URLs, `Authorization:` headers | **redact** | very low: the prefix is meaningful, not random |
| **3 — name=value** | an assignment whose *key* says secret | **report only** | measured **high** — see §3 |
### Worked examples, one line each
```
MEMPALACE_REMOTE_TOKEN=pfrDBfak… tier 1 — value matches this env's secret
the token is pfrDBfak… apparently tier 1 — same value, bare in prose, still caught
git clone https://joakim:hunter2@git/x tier 2 — credentials in a URL
Authorization: Bearer abcdefghijklmnop tier 2 — header shape
ghp_AbCdEf… / glpat-… / AKIA… / sk-ant-… tier 2 — vendor prefix
db_password = s3cr3t-p4ssw0rd-xyz tier 3 — only the KEY suggests it (reported)
GITEA_TOKEN=0123456789abcdef… (40 hex) tier 3 — indistinguishable from a commit sha
```
### Which tier fired? Read it off the output
The tool prints **rule names**, not tier numbers, because the rule says *what*
matched. The feeder prefixes them with the tier so both are visible:
```
[REDACTED] 2026-06-27T23-13-49.jsonl T1:env-value=1, T2:github-fine-grained=1 fp=ad78c7d4,fcd95ab5
[scrub] 29 redaction(s) applied, 141 name-anchored candidate(s) reported only
```
`RULE_TIERS` in `mempalace_redact.py` is the authoritative mapping, `tier_of()`
reads it, and a self-test fails if any rule is left unclassified — so the two
vocabularies cannot drift apart silently.
Tier 1 catches any presentation of a secret — env dump, JSON, error message,
URL, prose — because it matches the value itself. It is also, **in the default
configuration, the only tier that resolves a shape collision**: a 40-hex Gitea
PAT is byte-identical to a git commit sha, so only a known value (tier 1,
redacts) or a naming key (tier 3, reports) can tell them apart. A sha-shaped PAT
is therefore caught if and only if it belongs to this machine — an accepted gap,
since the alternative is redacting every commit sha in the palace.
## 3. The false-positive measurement, which changed the design
Tier 3 was originally going to redact. Measured against **52 MB of real fleet
transcripts (45 sessions)** it produced **403 hits**, and inspection of them (with
values masked) showed the overwhelming majority were not secrets:
- `GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}` — docker-compose interpolation
- `const tokens = countTokensForModel` — TypeScript source stored as memory
- `refreshToken(credentials: Credentials …)` — a *type annotation*
- `--setattr=krbPasswordExpiration=20260529…` — an IPA attribute holding a date
- `|1.FORGE.TOKENS:…` — AAAK diary shorthand
- `(root@host) Password:` followed by unrelated terminal output
Redacting those would corrupt code, configuration and documentation held as
memory, in order to catch secrets tier 1 already catches by value. After adding
guards for interpolation (`${VAR}`, `$VAR`, `%VAR%`, `{{tpl}}`), code context,
non-secret key suffixes (`…Expiration`, `…Count`, `…Type`) and all-digit values,
the enforced count on the same corpus fell from **403 to 29** — and those 29 are
tier-1 and tier-2 hits, i.e. real credential shapes.
So tier 3 **reports and does not rewrite** by default. Set
`MEMPALACE_REDACT_STRICT=1` to make it enforce, e.g. on a machine whose
transcripts are configuration-heavy rather than code-heavy.
## 4. Known false negatives — stated, not hidden
The scrubber will not catch a novel credential format pasted bare into prose
with no name nearby, a secret belonging to a machine whose environment this
process cannot see, a base64-of-a-secret, or a secret split across lines.
**"The scrubber ran" must never be read as "there are no secrets in here."** To
keep that measurable rather than assumed, two things are reported:
- every scrub prints a count — including `0 redaction(s)`, because printing
nothing is indistinguishable from a scrubber that never ran;
- `suspicions()` reports high-entropy strings it did **not** redact, as
`(length, fingerprint)` pairs rather than values, so the miss rate can be
tracked over time and a recurring fingerprint can be investigated by hand.
Findings never carry the secret. A `Finding` holds the rule, the label, the
length and `sha256(value)[:8]` — enough to recognise the same leak recurring,
not enough to recover it.
**Fail closed.** If the redactor cannot be imported, the feeder refuses to stage
rather than staging unscrubbed (`exit 3`). Override deliberately with
`MEMPALACE_FEED_ALLOW_UNSCRUBBED=1`.
## 5. The server-side layer (not yet implemented)
Three call sites, because the server has three near-duplicate validators rather
than one:
| Path | Function | Covers |
|---|---|---|
| drawers | `config.py` → `sanitize_content()` | `add_drawer`, `update_drawer`, `diary_write`, `checkpoint` — all four route through it |
| events | `logstream.py` → `_sanitize_body()` | `event_append`, and `event_ack` transitively |
| artifacts | `logstream.py` → `put_artifact()` | inlines its own checks; needs its own edit |
Server-side scrubbing is **tier 2 only** (plus optional tier-3 reporting): the
hub cannot see a client's environment, so tier 1 is structurally unavailable
there. That asymmetry is the reason the client hook is not redundant.
## 6. Cleaning up a leak that already landed
1. **Never re-echo the value while hunting it.** Pass it via stdin, never argv
(visible in `ps` on a shared host). Note `get_drawer` is a trap: reading a
drawer in order to redact it prints the secret back into the live transcript.
2. **Redact, don't delete.** Replacing the value in place keeps the mined
transcript's memory value; a stale embedding vector is a cheap price.
3. **Scrub the feeder inbox too**, or the next mine re-files it.
4. **A live session file needs an equal-length in-place overwrite** — the agent
holds it open, so temp-file-plus-rename loses everything appended afterwards.
5. **Expect residue.** SQLite keeps old page content in freed pages until
`VACUUM`, so a raw byte scan still matches after a successful `UPDATE`. Decide
explicitly whether that matters: if the credential store on the same host is
plaintext anyway, it usually does not.
## 7. Usage
```bash
# unit tests, including the must-not-redact corpus of palace id shapes
python3 bin/mempalace_redact.py --self-test
# scrub anything on stdin; report goes to stderr
some-command | python3 bin/mempalace_redact.py > clean.txt
# also list high-entropy strings that were NOT redacted, as fingerprints
python3 bin/mempalace_redact.py --suspicions < transcript.jsonl > clean.jsonl
```
+1 -1
View File
@@ -11,7 +11,7 @@ What lives where:
| Content | Home |
|---|---|
| How to expose a palace over HTTP, and the Host/Origin pin | `docs/phase-1-exposure-runbook.md` (here) |
| How this deployment was exposed over HTTP (tunnel, DNS, token custody) | private fleet repository — see [`phase-1-exposure-runbook.md`](phase-1-exposure-runbook.md), also moved |
| Why a palace needs a special backup, and how to restore one | `docs/backup-and-recovery.md` (here) |
| Unit/timer/plist templates | `contrib/` (here) |
| Which machine is primary, its addresses, users, tunnels, offsite target | private fleet repository |
+171 -15
View File
@@ -117,7 +117,7 @@ side of the wiring.
| `MEMPALACE_FEED_WING` | `wing_conversations` | Target wing — passed to both the exporter and the `mempalace_mine` call. |
| `MEMPALACE_FEED_DEBOUNCE_MS` | `600000` (10 min) | Minimum gap between mid-session (`agent_settled`) feeds. Bounds crash loss to one window instead of a whole session. |
| `MEMPALACE_FEED_PREPARE_TIMEOUT_MS` | `120000` | Kills a wedged `--prepare` subprocess. |
| `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `30000` | Caps the `mempalace_mine` call so a stalled palace can't hang session exit. |
| `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `300000` (5 min) | Bounds how long the extension *waits* for `mempalace_mine`, so a stalled palace can't hang session exit. It does **not** cancel the mine — see [Debugging](#debugging). Raised from `30000` in 2026-09: the mine is the slowest call this extension makes (30–60s in normal operation), so the old deadline fired routinely and reported healthy behaviour as an error. |
**Remote palace:** if `$MEMPALACE_REMOTE_URL` is set (see
[Transport](#transport-local-vs-external)), `mempalace_mine`'s source path is
@@ -187,8 +187,8 @@ chosen at load time:
transport is chosen once at extension load. Confirm the result the same way
as a container flip: ask the agent for `mempalace_status` and check the
reported palace path is the **remote** host's, not your own
`$HOME/.mempalace/palace` — see
[`docs/phase-1-exposure-runbook.md`](../../docs/phase-1-exposure-runbook.md) §3.8.
`$HOME/.mempalace/palace`. That one check is the whole verification: a
half-flipped client reports a local path while looking healthy.
Serve such an endpoint with `mempalace serve --host 172.17.0.1 --port 8765`
(the `pi-devbox` / `opencode-devbox` repos ship a
@@ -207,8 +207,9 @@ chosen at load time:
token auto-minting is gated on the bind being non-loopback, so it starts with
**no authentication at all**, no warning. Bind the docker0 gateway
(`172.17.0.1`): reachable from the host and its containers, not from the LAN.
See
[`docs/phase-1-exposure-runbook.md`](../../docs/phase-1-exposure-runbook.md).
Binding an interface is not the same as exposing a palace — an MCP endpoint
also pins `Host`/`Origin`, so a request arriving under the wrong hostname is
refused even when the port is open.
Implementation note: the HTTP client (`RemoteMcpClient`) is **vendored** from
[`pi-extensions`](https://gitea.jordbo.se/joakimp/pi-extensions)'
@@ -234,9 +235,9 @@ local palace, so a remote outage can never scatter memories into a local copy
nobody will look at again. The practical corollary, worth knowing before you
debug the wrong layer: **"the agent has no `mempalace_*` tools" is the
expected symptom of a server, token, or DNS fault**, not of a broken install.
Diagnose it with a direct `curl` to `MEMPALACE_REMOTE_URL` — see
[`docs/phase-1-exposure-runbook.md`](../../docs/phase-1-exposure-runbook.md)
§3.8. The design rationale for de-registering rather than degrading is in
Diagnose it with a direct `curl` to `MEMPALACE_REMOTE_URL`, and confirm the flip
with `mempalace_status` — the reported palace path must be the remote host's.
The design rationale for de-registering rather than degrading is in
[`docs/rfc-001-global-palace.md`](../../docs/rfc-001-global-palace.md) §2 and §4.1.
## Identity
@@ -315,24 +316,123 @@ waits for the agent to think of asking. Two delivery points, both fail-silent:
`triggerTurn` — waking the model on inbound fleet traffic is a much larger
behavioural change than auto-delivery.
**Which means a human is the trigger, and the mailbox now says so.** Measured
2026-08-26 on two devices: a delivery lands, the agent is idle, nothing happens,
and the operator asks *"do I have to nudge you for you to read this?"*. Yes —
because between the poll and the next turn no inference is running. The old text
explained how to *close* an ask and never said when it would be *seen*, so the
only reader who needed that fact was the one not told. Two additions, neither of
which touches the no-`triggerTurn` decision:
- **A delivery note in the message itself** — states that this is a queued
message, that nothing woke the agent, and that any message starts the turn
that handles it. Free, and aimed at the human reading the window.
- **A notification at poll time**, because the note only helps someone who is
already looking, and the case that loses an ask is nobody looking:
| `MEMPALACE_MAILBOX_NOTIFY` | Behaviour |
|---|---|
| *unset* (default) | in-TUI `ctx.ui.notify`, the same surface `session_start` already uses |
| `desktop` | additionally a terminal-native notification — Kitty `OSC 99`, else `OSC 777` (iTerm2, WezTerm, Ghostty, rxvt-unicode) |
| `0` / `off` | silent; mailbox still delivers |
The `desktop` path is how a notification escapes a container without
`notify-send`, DBus or any host access: the escape sequence is written to stdout
and interpreted by the terminal emulator on the human's own machine. It is
opt-in because writing raw escapes is a behaviour change on a shared machine,
not because it is unreliable. Title and body are stripped of `;` and control
bytes, so a payload can neither forge an OSC field nor end the sequence early.
**⚠️ Not yet observed firing through tmux (2026-08-26).** If pi runs inside a
multiplexer — e.g. `kitty → tmux on the host → docker exec → pi` — tmux drops
OSC sequences it does not implement, so the ping can vanish silently between
the container and the human. Reaching the outer terminal needs tmux's DCS
passthrough plus `allow-passthrough on`, which is not implemented here yet.
Note the client can detect *neither* layer: `KITTY_WINDOW_ID` is not forwarded
by `docker exec`, and `TMUX` is unset because tmux runs one level further out.
Verify with a hand-written sequence in your own stack before trusting it.
It fires **only when something is due** — the same condition as the delivery
itself. A ping on an empty poll would train its reader to ignore it, which is
the failure this whole feature exists to reverse.
Owed-ness is **derived, never read off a field**, because `event_ack` appends and
`status` is written once: a directed `open` event matches the mailbox query
*forever*, answered or not. Two calls (`to_agent=<me> status=open`, and
`from_agent=<me>`), then a candidate counts as answered only when one of this
device's own events has a **strictly higher `seq`**, joins via
device's own events is **strictly later**, joins via
`metadata.ack_of` or a shared `correlation_id`, *and* carries a terminal status
(`applied`/`superseded`/`failed`/`blocked`). The `seq` test is load-bearing:
(`applied`/`superseded`/`failed`/`blocked`). The ordering test is load-bearing:
without it one terminal reply suppresses every later ask on that correlation
forever. `seq` is replica-local, so `hlc` is the correct key once `mesh_peers`
reports actual peers.
forever.
"Strictly later" means `hlc` when both events carry one — a hybrid logical clock
rendered fixed-width, so a string comparison is a causal comparison across
replicas — falling back to `seq` only when either side lacks an `hlc`. `seq` is
this database's arrival `rowid`, so on a mesh the same event has a different
`seq` per replica and a reply can arrive before its ask. Never `created_at`: it
is second-precision, and a tie there can suppress an *unanswered* ask, which is
the one failure this derivation exists to prevent.
`*` broadcasts are excluded even though `to_agent=<me>` matches them, because the
protocol says a broadcast owes nobody a reply — which also means broadcasting an
ask demonstrably reaches no owed set, giving that documented anti-pattern teeth.
**Dormant asks — open, owed, and deliberately not announced.** Some asks are
planted *unanswerable on purpose*: a first-boot acceptance describes work that
becomes possible only when the device is next recreated, and it must stay owed
until then, because closing it early to tidy the mailbox is exactly how the work
gets lost. Derivation alone cannot tell that apart from neglected work, so such
an ask was announced at every session start and every poll for as long as it was
correctly waiting — measured at three days running on `emb-7kj4vr4g`
(2026-09-28 → 2026-10-01), and across three consecutive releases before that.
The ask was right; announcing it was wrong, and the cost landed on the human
reading the window, the one reader who cannot filter it.
An ask may therefore declare, in its own `metadata`, the condition under which it
is merely waiting:
```json
"dormant_unless": [
{ "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json",
"field": "release_tag", "baseline": "v1.9.4" },
{ "kind": "file_mtime", "path": "/etc/hostname",
"baseline": "2026-09-22T18:12:49Z" }
]
```
It is dormant while **every** condition still matches its baseline, and goes live
the moment **any** of them differs — the trigger such asks already stated in
prose ("act when EITHER differs"), now in a form the bridge can check. Two kinds,
both local-file-only: `file_mtime` (compared at whole seconds in UTC, because a
filesystem mtime carries sub-second residue that a reported ISO baseline never
will — comparing raw milliseconds would mark every such predicate permanently
"changed" and silently disable the feature) and `json_field` (dotted paths
allowed, compared as strings so a manifest holding `3` matches a baseline of
`"3"`). There is no expression language, no shell and no network: a general
evaluator in the path that decides whether work is *visible* is a far worse trade
than a clumsy schema.
Dormancy is **withheld from the announcement, never from the mailbox**: the
wake-up injection lists dormant asks once per session with their ids, and a
mid-session poll mentions only a count, and only when the window is already open
for something else. They are not added to the resurface map, so one becomes
announceable the instant its baseline moves.
**Dormancy must be proven, never assumed** — every unevaluable predicate shows
the ask as owed. A missing file, an unreadable one, a baseline that will not
parse, an unknown `kind`, a vanished field, a relative path, more than eight
conditions: each announces. The dangerous failure here is not a spurious nag but
work that disappears because a predicate could not be evaluated, which would be
indistinguishable from the ask being lost and would not surface until a release
needed it. An ask with no `dormant_unless` behaves exactly as it did before the
feature existed. `scripts/test-dormancy.sh` exercises all of it, positive and
negative arms both.
Gated on `MEMPALACE_PI_DEVICE` **and** `MEMPALACE_REMOTE_URL` (the same pair as
the stamper, since an unstamped client has no address to be reached at), and
disabled outright with `MEMPALACE_MAILBOX=0`. Inert on a solitary palace: no
disabled outright with `MEMPALACE_MAILBOX=0` (notifications alone with
`MEMPALACE_MAILBOX_NOTIFY=0`). Inert on a solitary palace: no
calls, no injection. Delivery is the mechanism; the *norms* — what a reply owes,
and that only a terminal event closes a thread — remain normative in the
*consumer* skill (`~/.agents/skills/mempalace/SKILL.md`). **This file documents
@@ -342,12 +442,21 @@ the mechanism; the skill is normative for behaviour.**
an SSE endpoint (`GET /logstream/stream`, `text/event-stream` in
`mempalace/mcp_server.py`), but a deployment may expose only the MCP endpoint
through its reverse proxy — verified 2026-08-26 against
`https://mempalace.jordbo.se`, where `/logstream/events`, `/logstream/stream`
and `/sync/peers` all return 404 while `/mcp` serves normally. Where that is the
`https://mempalace.jordbo.se`, where `/logstream/stream` and `/sync/peers`
return 404 while `/mcp` serves normally. Where that is the
case, polling through the existing MCP client is the only available path — which
is what the mailbox in §2 does — and enabling SSE means a proxy route plus an
auth decision, not an extension change.
> ⚠️ **Corrected 2026-08-26.** An earlier revision of this paragraph listed
> `/logstream/events` alongside those two as proxy-blocked. That route **does not
> exist in the server at all** — the complete GET table in mempalace 3.8.0 is
> `/healthz`, `/statusz`, `/logstream/stream` and `/sync/{version_vector,ops,artifact,peers}`,
> so `/logstream/events` would 404 against a directly-reachable server too. The
> proxy inference was right for the other two and wrong for that one; see
> [RFC 003 §7.8](../../docs/rfc-003-coordination-log.md). Distinguish *route absent*
> from *route blocked* before blaming infrastructure.
As with stamping, all of this is inert unless `MEMPALACE_REMOTE_URL` points at a
shared palace. On a solitary palace the event tools work fine and the log
contains only this machine's own events.
@@ -365,6 +474,9 @@ one — the long-lived server is only killed when a request genuinely stalls.
- `MEMPALACE_MCP_TIMEOUT_MS` — tool-call/request timeout. Default `60000`.
Kept short on purpose: a *query* taking this long is genuinely wedged.
The one call that is not a query — the feed's `mempalace_mine` — passes its
own deadline (`MEMPALACE_FEED_MINE_TIMEOUT_MS`) down to the transport per
call, so this default does not apply to it (since 2026-09-18; see below).
- `MEMPALACE_MCP_INIT_TIMEOUT_MS` — `initialize` + `tools/list` handshake
timeout. Default `300000`. Deliberately generous: a genuine first
cold-open over virtiofs can legitimately take minutes, and killing a
@@ -407,6 +519,50 @@ the next tool call transparently respawns `mempalace-mcp` and retries.
`mempalace-mcp` manually with raw JSON-RPC on stdin to read the
server-side error — much faster than guessing.
### `feed (tick) failed: mine timed out after …ms`
**Nothing has been lost.** The deadline bounds only how long the extension
*waits*; it cannot cancel the mine, which continues on the server. The
transcript is already staged before the mine is invoked, and
`mempalace mine --mode convos` dedups by `source_file` and is idempotent, so the
work either completed after the deadline or is redone by the next tick.
**Do not "fix" it by retrying harder from the client.** The palace is a single
writer; a blind retry is what turns one slow mine into a queue of them.
Before 2026-09 this message appeared many times per session, which made it look
like a persistent failure. That was a real defect, now fixed: `lastFeedAt` was
recorded only after a *successful* wait, so a timeout left the debounce clock
stale and every following settled turn started another mine on top of the one
still running. Two changes — recording the attempt before the wait, and raising
the deadline to sit far above the slowest honest completion — mean a healthy
fleet should now never see it.
If you *do* still see it, it is now informative rather than noise: a mine
exceeded five minutes. Check palace size and whether another writer (a
host-side feeder, a scheduled `mine`) is holding the write lock, rather than
raising the timeout again.
### `feed (tick) failed: mempalace remote request 'tools/call' failed: timed out after 60000ms`
Same event, different deadline — and the same reassurance: **nothing has been
lost**, the mine continues server-side.
This is what the previous message turned into after the 2026-09 change, and
it exposed that the change was incomplete. The feed's five-minute deadline was
only *raced* against the call; `callTool()` had no way to carry it, so the
transport's generic per-request timeout (`MEMPALACE_MCP_TIMEOUT_MS`, 60 s)
fired first on every honest 60 s+ mine, and the 300 s was unreachable. On the
stdio transport this was worse than noise: a per-request timeout there kills the
server child, so the mine really was aborted at 60 s.
Fixed 2026-09-18: `callTool(name, args, { timeoutMs })` passes a per-call
deadline to both transports and the feed uses it for the mine.
`scripts/test-mcp-call-timeout.sh` pins the contract (a plain call still honours
the short default; the override is honoured and is itself a deadline). Seeing
this message on a fixed build means a mine exceeded *five* minutes — treat it as
the previous section says.
## The `Type.Unsafe` gotcha
Earlier versions of this extension registered every MCP tool with
+624 -49
View File
@@ -67,7 +67,9 @@
* child, so pi gets an error instead of hanging and later calls fail fast.
* This is a per-REQUEST timeout, not a process-lifetime one — the
* long-lived server is only killed when a request genuinely stalls.
* - MEMPALACE_MCP_TIMEOUT_MS tool-call/request timeout (default 60000)
* - MEMPALACE_MCP_TIMEOUT_MS tool-call/request timeout (default 60000);
* the feed's mine carries its own, longer
* deadline (MEMPALACE_FEED_MINE_TIMEOUT_MS)
* - MEMPALACE_MCP_INIT_TIMEOUT_MS initialize+tools/list timeout (default 300000)
* Set either to 0 to disable (legacy unbounded behavior).
*
@@ -90,6 +92,7 @@
*/
import { type ChildProcessWithoutNullStreams, spawn } from "node:child_process";
import { readFileSync, statSync } from "node:fs";
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
import { Type } from "typebox";
@@ -116,7 +119,13 @@ interface IMcpClient {
readonly alive: boolean;
onExit: (() => void) | null;
start(): Promise<void>;
callTool(name: string, args: Record<string, unknown>): Promise<any>;
/**
* `opts.timeoutMs` overrides the transport's generic per-request deadline
* for THIS call only. Callers that knowingly invoke a long server-side job
* (the feed's `mempalace_mine`) pass their own deadline here; everything
* else keeps the short default, which is the wedged-query guard.
*/
callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any>;
ensureAlive(): Promise<boolean>;
stop(): void | Promise<void>;
}
@@ -132,6 +141,7 @@ const num = (envVal: string | undefined, fallback: number): number => {
type LogEvent = {
id?: string;
seq?: number;
hlc?: string;
type?: string;
status?: string;
from_agent?: string;
@@ -148,6 +158,165 @@ const sleep = (ms: number): Promise<void> =>
if (typeof t.unref === "function") t.unref();
});
// ── Dormancy ──────────────────────────────────────────────────────────────────
/**
* Is a standing ask DORMANT — correct to leave open, wrong to announce?
*
* The problem this solves, measured rather than imagined. A first-boot
* acceptance ask is planted deliberately unanswerable: it describes work that
* becomes possible only when the device is next recreated, and it must STAY
* owed until then, because closing it early to tidy the mailbox is exactly how
* that work gets lost. But deriveOwed could only see "directed, open, not
* answered, not withdrawn", so such an ask is announced at every session start
* and every poll for as long as it is correctly waiting — three days running on
* emb-7kj4vr4g (2026-09-28 → 2026-10-01), and three consecutive releases before
* that. The ask was right; announcing it was wrong. The cost lands on the human
* reading the window, who is the one reader that cannot filter it.
*
* So an ask may now declare, in its own metadata, the condition under which it
* is merely waiting:
*
* "dormant_unless": [
* { "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json",
* "field": "release_tag", "baseline": "v1.9.4" },
* { "kind": "file_mtime", "path": "/etc/hostname",
* "baseline": "2026-09-22T18:12:49Z" }
* ]
*
* It is dormant while EVERY condition still matches its baseline, and goes live
* the moment ANY of them differs — which is the trigger those asks already
* state in prose ("act when EITHER differs"), now in a form a machine can
* check. Dormant asks are withheld from the ANNOUNCED owed set, never from the
* mailbox: the wake-up injection still lists them once per session, so they
* stay discoverable and cannot be silently dropped.
*
* THE SAFETY RULE, and the reason every branch below reports `unchanged: false`
* on doubt: dormancy must be PROVEN, never assumed. A missing file, an
* unreadable one, a baseline that will not parse, an unknown `kind`, a field
* that has vanished — each is a reason to SHOW the ask, not to hide it. The
* dangerous failure here is not a spurious nag; it is work that vanishes
* because a predicate could not be evaluated, which would be indistinguishable
* from the ask being lost and would not surface until a release needed it.
*
* No eval, no shell, no network: the predicate language is deliberately two
* fixed checks against local files. A general expression evaluator sitting in
* the path that decides whether work is VISIBLE is a far worse trade than a
* slightly clumsy schema.
*/
export type DormancyVerdict = { dormant: boolean; reason: string };
/** Split of the derived owed-set: what to announce, and what is merely waiting. */
export type OwedSplit = { owed: LogEvent[]; dormant: LogEvent[] };
// Bounds, so a malformed or hostile event cannot turn a mailbox read into real
// work. Both are deliberately small: these predicates describe boot state, and
// anything needing more than a handful of local checks is not a dormancy rule.
const MAX_DORMANCY_CONDITIONS = 8;
const MAX_DORMANCY_JSON_BYTES = 256 * 1024;
const dormancyCondition = (c: unknown): { unchanged: boolean; reason: string } => {
if (typeof c !== "object" || c === null || Array.isArray(c))
return { unchanged: false, reason: "condition is not an object" };
const { kind, path, baseline, field } = c as Record<string, unknown>;
// Absolute paths only: a relative one would resolve against whatever cwd the
// pi process happens to hold, which is not a property of the device whose
// state the baseline describes.
if (typeof path !== "string" || !path.startsWith("/") || path.includes("\0"))
return {
unchanged: false,
reason: `path must be an absolute string, got ${JSON.stringify(path)}`,
};
if (typeof baseline !== "string" && typeof baseline !== "number")
return { unchanged: false, reason: "baseline must be a string or a number" };
if (kind === "file_mtime") {
// Compared at WHOLE SECONDS in UTC, because the two routes that produce
// these baselines disagree below that: a filesystem mtime carries
// sub-second residue (measured: /etc/hostname at .773761009) that a
// hand-written or reported ISO baseline never will. Comparing raw
// milliseconds would make every such predicate permanently "changed",
// i.e. would silently disable the feature while appearing to work.
const expected = Date.parse(String(baseline));
if (!Number.isFinite(expected))
return {
unchanged: false,
reason: `baseline is not a parseable date: ${JSON.stringify(baseline)}`,
};
let mtimeMs: number;
try {
mtimeMs = statSync(path).mtimeMs;
} catch {
return { unchanged: false, reason: `cannot stat ${path}` };
}
return Math.floor(mtimeMs / 1000) === Math.floor(expected / 1000)
? { unchanged: true, reason: `${path} mtime still at baseline` }
: { unchanged: false, reason: `${path} mtime moved from baseline` };
}
if (kind === "json_field") {
if (typeof field !== "string" || field.length === 0)
return { unchanged: false, reason: "json_field needs a non-empty field name" };
let text: string;
try {
const buf = readFileSync(path);
if (buf.byteLength > MAX_DORMANCY_JSON_BYTES)
return {
unchanged: false,
reason: `${path} exceeds the ${MAX_DORMANCY_JSON_BYTES}-byte cap`,
};
text = buf.toString("utf8");
} catch {
return { unchanged: false, reason: `cannot read ${path}` };
}
let doc: unknown;
try {
doc = JSON.parse(text);
} catch {
return { unchanged: false, reason: `${path} is not valid JSON` };
}
let cur: unknown = doc;
for (const seg of field.split(".")) {
if (
typeof cur !== "object" ||
cur === null ||
!Object.prototype.hasOwnProperty.call(cur, seg)
)
return { unchanged: false, reason: `${path} has no field ${field}` };
cur = (cur as Record<string, unknown>)[seg];
}
if (cur !== null && typeof cur === "object")
return { unchanged: false, reason: `field ${field} is not a scalar` };
// Compared as strings on purpose: the baseline arrives as JSON metadata, so
// a manifest holding 3 and a baseline saying "3" are the same fact.
return String(cur) === String(baseline)
? { unchanged: true, reason: `${field} still at baseline` }
: { unchanged: false, reason: `${field} moved from baseline` };
}
return { unchanged: false, reason: `unknown dormancy kind ${JSON.stringify(kind)}` };
};
export function evaluateDormancy(
metadata: Record<string, unknown> | null | undefined,
): DormancyVerdict {
const raw = metadata?.dormant_unless;
// Every one of these returns NOT dormant, i.e. "announce it". An ask without
// a predicate behaves exactly as it did before this feature existed.
if (raw === undefined || raw === null) return { dormant: false, reason: "no dormant_unless" };
if (!Array.isArray(raw)) return { dormant: false, reason: "dormant_unless is not an array" };
if (raw.length === 0) return { dormant: false, reason: "dormant_unless is empty" };
if (raw.length > MAX_DORMANCY_CONDITIONS)
return {
dormant: false,
reason: `too many conditions (${raw.length} > ${MAX_DORMANCY_CONDITIONS})`,
};
for (const c of raw) {
const v = dormancyCondition(c);
if (!v.unchanged) return { dormant: false, reason: v.reason };
}
return { dormant: true, reason: `all ${raw.length} condition(s) still at baseline` };
}
class StdioMcpClient implements IMcpClient {
private proc: ChildProcessWithoutNullStreams | null = null;
private nextId = 1;
@@ -371,8 +540,12 @@ class StdioMcpClient implements IMcpClient {
});
}
async callTool(name: string, args: Record<string, unknown>): Promise<any> {
return this.request("tools/call", { name, arguments: args });
async callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any> {
// The per-call override matters MORE here than for the HTTP client: on
// timeout this transport kills the server child, so a generic deadline
// that undercuts a long mine does not merely abandon the wait — it
// aborts the mine.
return this.request("tools/call", { name, arguments: args }, opts?.timeoutMs ?? this.requestTimeoutMs);
}
/** SIGTERM then SIGKILL grace, for stall recovery. */
@@ -420,6 +593,10 @@ class StdioMcpClient implements IMcpClient {
// • per-request AbortController timeout honouring MEMPALACE_MCP_TIMEOUT_MS /
// MEMPALACE_MCP_INIT_TIMEOUT_MS, mirroring StdioMcpClient's timeout ethos.
// • alive / ensureAlive / onExit to satisfy IMcpClient.
// • callTool() takes an optional per-call `{ timeoutMs }` (IMcpClient
// contract, see there) so the feed's long-running mine is not cut off by
// the generic per-request deadline. Not a protocol change; sync token
// unchanged.
//
// NOTE: mempalace-mcp --transport http is a SESSIONLESS, stateless JSON-RPC
// server (no Mcp-Session-Id, always application/json, Connection: close), so
@@ -465,8 +642,8 @@ class RemoteMcpClient implements IMcpClient {
this.healthy = true;
}
async callTool(name: string, args: Record<string, unknown>): Promise<any> {
return this.request("tools/call", { name, arguments: args });
async callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any> {
return this.request("tools/call", { name, arguments: args }, { timeoutMs: opts?.timeoutMs });
}
/**
@@ -854,7 +1031,22 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
const feedWing = process.env.MEMPALACE_FEED_WING ?? "wing_conversations";
const feedDebounceMs = num(process.env.MEMPALACE_FEED_DEBOUNCE_MS, 600_000);
const feedPrepareTimeoutMs = num(process.env.MEMPALACE_FEED_PREPARE_TIMEOUT_MS, 120_000);
const feedMineTimeoutMs = num(process.env.MEMPALACE_FEED_MINE_TIMEOUT_MS, 30_000);
// 300_000, raised from 30_000 on 2026-09-10. The mine is the SLOWEST thing
// this extension does — it offers every qualifying session transcript to a
// SINGLE-WRITER palace, measured at 30–60s in normal operation and growing
// with the corpus — yet it carried by far the TIGHTEST deadline: 4x tighter
// than the prepare step that precedes it (120_000) and 10x tighter than the
// init handshake (300_000), which is a fast call. All three were introduced
// together in 29e660e (2026-08-12) and this one was never revisited, so the
// deadline fired during entirely normal operation and the resulting message
// read as an error when nothing had gone wrong.
//
// Matched to the init timeout because liveness is the ONLY legitimate job
// left for this deadline: the Promise.race below abandons our WAIT, it cannot
// cancel the server's work, so the deadline buys nothing except an escape
// from a permanently hung call. It must therefore sit far above the slowest
// honest completion, not near it.
const feedMineTimeoutMs = num(process.env.MEMPALACE_FEED_MINE_TIMEOUT_MS, 300_000);
let lastFeedAt = 0; // 0 => the first settled turn also acts as a catch-up
let feedInFlight: Promise<void> | null = null;
@@ -908,17 +1100,46 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
try {
const source = await prepareFeed(reason);
if (!source) return;
// Record the attempt HERE, before awaiting — not after a successful
// wait. The race below abandons only our WAIT; the mine keeps running
// server-side, and `mine --mode convos` dedups by source_file and is
// idempotent, so a timeout is emphatically not a "did not happen".
//
// Leaving lastFeedAt stale on the timeout path defeated the debounce
// guard in the agent_settled handler below (`Date.now() - lastFeedAt <
// feedDebounceMs`): with lastFeedAt unchanged that guard passed on EVERY
// settled turn, and because `run` had already settled, feedInFlight was
// null too — so BOTH guards stood open. Each settled turn then launched
// another mine while the previous one was still running: overlapping
// writers queueing on a single-writer palace, each making the next one
// slower and the next timeout likelier. That positive feedback loop,
// not the tight deadline by itself, is why operators saw "mine timed out
// after 30000ms" many times per session rather than at most once per
// debounce window.
lastFeedAt = Date.now();
// The deadline is passed DOWN to the transport as well as raced
// here. Until 2026-09-18 it was only raced: callTool() had no way
// to carry it, so the transport's generic per-request timeout
// (MEMPALACE_MCP_TIMEOUT_MS, 60 000) fired first on every honest
// 60 s+ mine — "remote request 'tools/call' failed: timed out
// after 60000ms" — and the 300 000 below was unreachable. The
// race stays as the liveness guard for a transport whose timeout
// is disabled (0).
await Promise.race([
client.callTool("mempalace_mine", {
source,
mode: "convos",
wing: feedWing,
// Internal call: it does not pass through the registered tool's
// execute(), so it stamps itself. These ARE this harness's own
// transcripts from this device, so the harness segment is the
// agent (not `miner`) even though the tool is `mine`.
agent: stampProvenance ? `${agentName}@${device}` : agentName,
}),
client.callTool(
"mempalace_mine",
{
source,
mode: "convos",
wing: feedWing,
// Internal call: it does not pass through the registered tool's
// execute(), so it stamps itself. These ARE this harness's own
// transcripts from this device, so the harness segment is the
// agent (not `miner`) even though the tool is `mine`.
agent: stampProvenance ? `${agentName}@${device}` : agentName,
},
{ timeoutMs: feedMineTimeoutMs },
),
new Promise((_resolve, reject) =>
setTimeout(
() => reject(new Error(`mine timed out after ${feedMineTimeoutMs}ms`)),
@@ -926,7 +1147,6 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
),
),
]);
lastFeedAt = Date.now();
} catch (err) {
process.stderr.write(
`[mempalace ext] feed (${reason}) failed: ${(err as Error).message}\n`,
@@ -1007,6 +1227,40 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
}
};
/**
* Is `later` strictly after `earlier` in the fleet's ordering?
*
* Prefer `hlc` — a hybrid logical clock rendered fixed-width
* (`<unix_ms:13 digits>-<counter:6 hex>-<replica_id>`), so a plain string
* comparison IS the causal comparison, and it stays correct once a second
* replica exists. `seq` is the arrival rowid of THIS database, so on a mesh
* the same event carries a different `seq` per replica and a reply can land
* before the ask it answers.
*
* Fall back to `seq` only when either side lacks an `hlc` (a server predating
* the field, or a row whose backfill did not run). On a single replica the two
* agree, so the fallback is not a downgrade today — it is the single-replica
* case still being handled after the mesh case became primary.
*
* NEVER compare `created_at`: it is server-generated at second precision, so
* ties are routine, and a tie or a skew there can suppress an UNANSWERED ask
* outright. Every failure mode here is deliberately kept on the
* noisy-but-visible side — an answered item resurfacing is annoying, an
* unanswered ask going silent defeats the mailbox.
*/
const isStrictlyAfter = (later: LogEvent, earlier: LogEvent): boolean => {
if (
typeof later.hlc === "string" &&
typeof earlier.hlc === "string" &&
later.hlc !== "" &&
earlier.hlc !== ""
) {
return later.hlc > earlier.hlc;
}
if (typeof later.seq !== "number" || typeof earlier.seq !== "number") return false;
return later.seq > earlier.seq;
};
/**
* Has one of MY events closed this candidate?
*
@@ -1023,15 +1277,8 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
// it one terminal reply suppresses every LATER ask on the same
// correlation_id forever — silently, permanently, and worst on exactly
// the long-running threads the correlation join is for.
//
// Compare `seq`, NEVER `created_at`. `seq` is replica-local (it equals
// origin_seq only while a single replica authors for every machine); the
// durable key once `mempalace_mesh_peers` reports real peers is `hlc`,
// which is total and causally consistent. Local-seq skew can only make an
// ANSWERED item resurface (noise, visible), whereas a timestamp
// comparison can suppress an UNANSWERED ask outright.
if (typeof m.seq !== "number" || typeof candidate.seq !== "number") return false;
if (m.seq <= candidate.seq) return false;
// See isStrictlyAfter for why the key is `hlc` and not `seq`/`created_at`.
if (!isStrictlyAfter(m, candidate)) return false;
// (b) join on the exact ack (written for us by event_ack), else on a
// shared correlation_id — which is why correlation_id is mandatory on a
// directed open: without it there is no key to join a reply back to.
@@ -1040,17 +1287,120 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
return Boolean(m.correlation_id && m.correlation_id === candidate.correlation_id);
});
/** Directed asks addressed to this device with no terminal reply from it. */
async function deriveOwed(): Promise<LogEvent[]> {
if (!mailboxEnabled || !available) return [];
/**
* Has the ORIGINAL REQUESTER explicitly withdrawn this candidate?
*
* RFC 003 §3.3 clause 3 clears an ask only on "no event **of yours**", and the
* skill states the consequence outright: "there is nothing anyone can do about
* it from the other end". That asymmetry is deliberate and mostly right — owed
* ness is a statement about the RECIPIENT's accountability, and a requester
* must not be able to delete an obligation the recipient genuinely has.
*
* It is wrong in exactly one case: the requester withdrawing its OWN ask.
*
* MEASURED COST, 2026-09-08/09. pi@mbp-m1-2020 withdrew a v1.8.13 rollout ask
* to pi@tor-ms22 (evt seq 119: terminal `superseded`, same correlation_id,
* `metadata.closes` naming it, body "DO NOT SPEND A MINUTE ON v1.8.13") and
* recorded the withdrawal as done. It had no effect: seq 119's `from_agent` is
* mbp, so it could never satisfy a join that only looks at tor-ms22's own
* events. tor-ms22's next wake-up — 8h later, on the first boot of the image
* shipping this very derivation — still listed the ask as owed, 41h old, for a
* release that had been superseded and never installed there. It had to spend a
* write closing an ask nobody wanted answered. The asymmetry is also INVISIBLE
* from the sender's side: mbp did everything a sender is told to do and got a
* result indistinguishable from success.
*
* WHY AN EXPLICIT MARKER AND NOT "ANY TERMINAL EVENT FROM THE REQUESTER".
* This file's standing rule is that every failure mode stays on the
* noisy-but-visible side, because a spuriously resurfacing item is one turn of
* human correction whereas a suppressed unanswered ask is silent and permanent.
* Inferring release from any terminal event would breach that: a requester
* appending `applied` for its own bookkeeping — on the correlation, addressed
* to me, before I ever replied — would silently delete a real obligation.
* So release must be STATED, not inferred. `withdraws` is the canonical key;
* `closes` is honoured because it is already this fleet's de-facto marker (mbp
* seq 119, emb seq 117/118 all use it), and either must name this exact ask —
* its `correlation_id` or its event `id`. Prose does not count.
*
* WHY THIS DOES NOT BREAK THE FIXED POINT IN RFC 003 §9.2. §9.2 rejects letting
* terminal directed events into the owed set, because then every closure mints a
* fresh obligation and the loop never terminates. That argument is about
* CANDIDATES. This adds a CLEARER. A withdrawal carries a terminal status, so it
* can never be a candidate, and nothing new becomes owed. The asserting shape
* (`open`) and the clearing shape (terminal) stay disjoint. A reader who reaches
* for §9.2 to object is answering a different question.
*/
const isWithdrawn = (candidate: LogEvent, inbound: LogEvent[]): boolean => {
const requester = candidate.from_agent;
if (!requester) return false;
// Names THIS ask specifically: its correlation or its id. A marker naming
// something else, or carrying prose, is not a withdrawal of this ask.
const names = (v: unknown): boolean =>
typeof v === "string" &&
v.length > 0 &&
((Boolean(candidate.correlation_id) && v === candidate.correlation_id) ||
(Boolean(candidate.id) && v === candidate.id));
return inbound.some((e) => {
// Only the party that ASKED may retract. A third party writing a terminal
// event on a shared correlation must never clear my obligation.
if (e.from_agent !== requester) return false;
// Addressed to the party being released. `to_agent=<me>` also matches '*'
// at the SQL level (RFC 003 §7.4), and a broadcast must not be able to
// quietly empty every machine's mailbox at once.
if (e.to_agent !== mailboxAddress) return false;
if (!TERMINAL_STATUS.has((e.status ?? "").toLowerCase())) return false;
// Same ordering guard, same reason as isAnswered: a withdrawal must not
// retire an ask the requester sent LATER on the same correlation.
if (!isStrictlyAfter(e, candidate)) return false;
const ackOf = e.metadata?.ack_of;
const joins =
(typeof ackOf === "string" && Boolean(candidate.id) && ackOf === candidate.id) ||
Boolean(e.correlation_id && e.correlation_id === candidate.correlation_id);
if (!joins) return false;
return names(e.metadata?.withdraws) || names(e.metadata?.closes);
});
};
/**
* Directed asks addressed to this device with no terminal reply from it, and
* not explicitly withdrawn by whoever sent them (see isWithdrawn).
*/
async function deriveOwed(): Promise<OwedSplit> {
if (!mailboxEnabled || !available) return { owed: [], dormant: [] };
try {
const [candidatesRaw, mineRaw] = await Promise.all([
const [candidatesRaw, mineRaw, inboundRaw] = await Promise.all([
// Every cursor-less event_list call in this file passes `order` explicitly.
// The server's default is not ours to lean on: mempalace <= 3.9.0 defaults
// to `asc`, 3.10.0 flips a cursor-less listing to newest-first. Both
// derive* joins want the NEWEST window (see the comment below), so say so
// and the verdict stops depending on which server version answers.
client.callTool("mempalace_event_list", {
to_agent: mailboxAddress,
status: "open",
limit: 50,
order: "desc",
}),
// `order: "desc"` is a fix, not a flourish. event_list defaults to `asc`
// (append order), so this asked for the OLDEST 100 events this device
// ever wrote — meaning that once a device passes 100 authored events its
// most RECENT replies fall out of the join window and every ask it just
// answered resurfaces as owed. Latent rather than theoretical: tor-ms22
// was at ~20 authored events when this was written. The window has to be
// anchored at the newest end for the same reason the join uses `hlc`.
client.callTool("mempalace_event_list", {
from_agent: mailboxAddress,
limit: 100,
order: "desc",
}),
// Inbound with NO status filter, deliberately: a withdrawal carries a
// TERMINAL status and is therefore structurally invisible to the
// `status: "open"` candidate query above — the same omission deriveClosed
// is built on. No amount of polling the open set could ever see one.
client.callTool("mempalace_event_list", {
to_agent: mailboxAddress,
limit: 100,
order: "desc",
}),
client.callTool("mempalace_event_list", { from_agent: mailboxAddress, limit: 100 }),
]);
const candidates = parseEvents(candidatesRaw).filter((c) => {
// `to_agent: <me>` ALSO matches `*` broadcasts, per the tool contract.
@@ -1064,11 +1414,80 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
// You cannot owe yourself.
return c.from_agent !== mailboxAddress;
});
if (candidates.length === 0) return [];
if (candidates.length === 0) return { owed: [], dormant: [] };
const mine = parseEvents(mineRaw);
return candidates.filter((c) => !isAnswered(c, mine));
const inbound = parseEvents(inboundRaw);
const live = candidates.filter((c) => !isAnswered(c, mine) && !isWithdrawn(c, inbound));
// PARTITION, not filter. A dormant ask is still owed in the protocol
// sense — it has no terminal reply and must not be closed — so it stays in
// the return value where the mailbox can still account for it, and only
// the ANNOUNCING paths below treat the two halves differently.
const owed: LogEvent[] = [];
const dormant: LogEvent[] = [];
for (const c of live) (evaluateDormancy(c.metadata).dormant ? dormant : owed).push(c);
return { owed, dormant };
} catch {
// Fail silent and open: a mailbox read must never break a session.
return { owed: [], dormant: [] };
}
}
/**
* Terminal replies to asks THIS device sent. NOT work — news.
*
* Why this is a second query rather than a widened deriveOwed(): deriveOwed()
* asks the log for `status: "open"`, and a reply that CLOSES an ask is by
* definition not open, so it is structurally invisible to that filter — no
* amount of polling or waiting could ever have surfaced it.
*
* Measured 2026-09-07: emb-7kj4vr4g closed a v1.8.13 rollout ask with a
* task.reply at status=applied, and the operator reasonably expected to be
* told. The mailbox stayed silent and was CORRECT to — "owed" means "you must
* reply", and nothing was owed. But the single most useful thing a fleet can
* tell a human is "the thing you asked for is done" (here: done, AND your
* premise was wrong), and that was the one category the feature could not
* report. Silence was right by its own definition and wrong by the user's.
*/
async function deriveClosed(): Promise<LogEvent[]> {
if (!mailboxEnabled || !available) return [];
try {
// `order: "desc"` for the same reason as deriveOwed: without it a 3.9.0
// server hands back the OLDEST window, so past 100 authored events this
// device's newest asks (and past 50 inbound, the replies that close them)
// fall outside the join. The selection below is order-independent
// (isStrictlyAfter), so this only fixes WHICH events are in the window.
const [mineRaw, inboundRaw] = await Promise.all([
client.callTool("mempalace_event_list", { from_agent: mailboxAddress, limit: 100, order: "desc" }),
// NO status filter, deliberately: that omission is the entire fix.
client.callTool("mempalace_event_list", { to_agent: mailboxAddress, limit: 50, order: "desc" }),
]);
// Correlations this device opened as a DIRECTED ask. A broadcast owes
// nobody a reply, so it cannot be closed by one either.
const myAsks = new Map<string, LogEvent>();
for (const e of parseEvents(mineRaw)) {
if (!e.correlation_id || !e.to_agent) continue;
if (e.to_agent === "*" || e.to_agent === mailboxAddress) continue;
if (e.type !== "task.request") continue;
myAsks.set(e.correlation_id, e);
}
if (myAsks.size === 0) return [];
// Newest terminal reply per correlation only. A peer that appends
// applied-then-superseded should cost one line, not a wall of them.
const best = new Map<string, LogEvent>();
for (const e of parseEvents(inboundRaw)) {
if (e.from_agent === mailboxAddress) continue; // cannot inform myself
const cid = e.correlation_id;
if (!cid || !myAsks.has(cid)) continue;
// NOTE the deliberate asymmetry with deriveOwed: a broadcast is excluded
// there (it owes nobody) but allowed here, because a peer answering on MY
// correlation_id is news to me regardless of how widely it was addressed.
if (!TERMINAL_STATUS.has((e.status ?? "").toLowerCase())) continue;
const prev = best.get(cid);
if (!prev || isStrictlyAfter(e, prev)) best.set(cid, e);
}
return [...best.values()];
} catch {
// Same contract as deriveOwed: a mailbox read never breaks a session.
return [];
}
}
@@ -1105,11 +1524,41 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
})
.join("\n");
const DORMANT_NOTE =
"Each of these declares a `dormant_unless` predicate in its metadata that STILL " +
"matches this device's current state, so the work it describes is not yet " +
"possible. They are NOT closed and NOT answered — leave them open. They return to " +
"the announced owed-set by themselves the moment a baseline stops matching. If a " +
"predicate cannot be evaluated at all (missing file, unparseable baseline, unknown " +
"kind) the ask is announced as owed instead, deliberately: dormancy has to be " +
"proven, never assumed.";
const OWED_HOWTO =
"To close one, append an event on the SAME correlation_id with a terminal status " +
"(applied/superseded/failed/blocked). An ack alone does NOT clear it, and neither " +
"does claimed/ready — those deliberately keep it visible as taken-but-unfinished.";
// Written for the ONE reader the rest of this text ignores: the human watching
// the window. Measured 2026-08-26 on two devices — a delivery lands, the agent
// is idle, nothing happens, and the operator asks "do I have to nudge you for
// you to read this?". Yes, and it is by design (see the deliverAs comment
// below): at agent_settled no inference is running, so the text is queued for
// the next turn. The message explained how to CLOSE an ask but never when it
// would be SEEN, so the person who needed that fact was the only one not told.
// Cheapest possible fix, and it changes no behaviour: say so in the text.
const QUEUED_NOTE =
"DELIVERY NOTE — this is a QUEUED message, not an action: it arrived while this " +
"agent was idle, so no model call was made and nothing woke it. It is read at the " +
"start of the next turn. If you are the human watching and want it handled now, " +
"send any message to start a turn; the agent is not ignoring the ask, it is not running.";
const CLOSED_NOTE =
"NO ACTION IS OWED on these — they are replies to asks THIS device sent, shown " +
"once each because 'the thing you asked for is done' is news you wanted and the " +
"owed-set could never carry it. Read the reply before assuming your original ask " +
"was right: a peer that did the work is the most likely party to have found your " +
"premise wrong.";
// Mid-session mailbox. Tier 1 (the wake-up injection below) owns the first
// look; this exists because arrivals are bursty and correlate with our own
// activity — measured 2026-08-26, 11 of 22 events on this log landed inside a
@@ -1121,7 +1570,68 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
// the owed set holds something not already surfaced. Re-announcing the same
// ask every poll is precisely how a channel teaches its reader to ignore it,
// which is the failure this whole feature exists to reverse.
pi.on("agent_settled", async () => {
// --- A: tell the human at poll time (RFC 003 §7.11) -----------------------
//
// B (the QUEUED_NOTE above) fixes the confusion of someone who IS looking at
// the window. It does nothing for the case that actually loses an ask: nobody
// is looking. The delivered text is queued for a turn that only a human starts,
// so an ask can wait indefinitely on an idle session while its recipient makes
// coffee. A notification at poll time is the only part of this feature that
// reaches a person who is not watching.
//
// Two surfaces, deliberately split by risk:
// default in-TUI ctx.ui.notify, exactly as session_start already does.
// Zero new output channels; visible only if you are looking.
// =desktop additionally emit a terminal-native notification, reusing
// the detection the fleet's own notify.ts already proves in
// this harness (Kitty OSC 99 / Windows toast / OSC 777). This
// escapes the container without notify-send, DBus or any host
// access: the escape sequence is interpreted by the terminal
// emulator on the human's machine. Opt-in because writing raw
// escapes to stdout is a behaviour change on shared machines,
// not because it is unreliable.
// =kitty =osc777 force one protocol, skipping detection entirely.
// =0 silent, for anyone who wants the mailbox without pings.
//
// WHY FORCING EXISTS — measured on tor-ms22 2026-08-26, and it invalidates
// detection in exactly the deployment this ships in. Terminal identity lives in
// env vars set by the emulator (KITTY_WINDOW_ID, TERM_PROGRAM) and `docker exec`
// does NOT forward them: inside the container pi sees only TERM=xterm-256color
// no matter what is rendering it. So detection can never see Kitty from in here
// and always falls through to OSC 777, which Kitty does not implement — on a
// containerised client the notification would silently do nothing, the worst
// possible failure for a feature whose whole job is to break a silence. Naming
// the protocol ends the guessing.
//
// Fired only when something is actually due, i.e. the same condition as the
// delivery itself: a notification that fires on an empty poll would teach its
// reader to ignore it, which is the failure this whole feature exists to undo.
const mailboxNotifyMode = (process.env.MEMPALACE_MAILBOX_NOTIFY ?? "").trim().toLowerCase();
const mailboxNotifyEnabled = mailboxNotifyMode !== "0" && mailboxNotifyMode !== "off";
const mailboxNotifyTerminal =
mailboxNotifyMode === "desktop" ||
mailboxNotifyMode === "kitty" ||
mailboxNotifyMode === "osc777";
/** Terminal-native notification. Best-effort; never throws into a handler. */
const notifyTerminal = (title: string, body: string): void => {
try {
const esc = (s: string): string => s.replace(/[\x00-\x1f\x7f;]/g, " ");
if (
mailboxNotifyMode === "kitty" ||
(mailboxNotifyMode === "desktop" && !!process.env.KITTY_WINDOW_ID)
) {
process.stdout.write(`\x1b]99;i=1:d=0;${esc(title)}\x1b\\`);
process.stdout.write(`\x1b]99;i=1:p=body;${esc(body)}\x1b\\`);
return;
}
process.stdout.write(`\x1b]777;notify;${esc(title)};${esc(body)}\x07`);
} catch {
/* a notification is never worth an exception */
}
};
pi.on("agent_settled", async (_event, ctx) => {
if (!mailboxEnabled || !available) return;
if (!wokeUp) return; // the wake-up injection has not run yet; it does look #1
if (Date.now() - lastMailboxPollAt < mailboxPollMs) return;
@@ -1130,22 +1640,48 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
lastMailboxPollAt = Date.now();
void (async () => {
try {
const owed = await deriveOwed();
const [owedSplit, closed] = await Promise.all([deriveOwed(), deriveClosed()]);
const { owed, dormant } = owedSplit;
const now = Date.now();
const due = owed.filter((e) => {
if (!e.id) return false;
const last = surfaced.get(e.id);
return last === undefined || now - last >= mailboxResurfaceMs;
});
if (due.length === 0) return; // silence is the correct output here
for (const e of due) if (e.id) surfaced.set(e.id, now);
const unseen = (list: LogEvent[]): LogEvent[] =>
list.filter((e) => {
if (!e.id) return false;
const last = surfaced.get(e.id);
return last === undefined || now - last >= mailboxResurfaceMs;
});
const due = unseen(owed);
// Closing replies are announced ONCE and never resurface: an ask you
// already know is finished is not a nag, and re-announcing it hourly is
// the train-the-reader-to-ignore-it failure this window exists to stop.
const newsRaw = closed.filter((e) => e.id && !surfaced.has(e.id));
if (due.length === 0 && newsRaw.length === 0) return; // silence is correct
for (const e of [...due, ...newsRaw]) if (e.id) surfaced.set(e.id, now);
const blocks: string[] = [];
if (due.length > 0) {
blocks.push(
`${due.length} directed ask(s) addressed to "${mailboxAddress}" with no ` +
`terminal reply from this device. ${OWED_HOWTO}\n\n${formatOwed(due)}`,
);
}
if (newsRaw.length > 0) {
blocks.push(
`${newsRaw.length} repl(y|ies) CLOSING an ask this device sent. ` +
`${CLOSED_NOTE}\n\n${formatOwed(newsRaw)}`,
);
}
// A dormant ask NEVER opens this window on its own — the early return
// above is what actually stops the nag. But once the window is open for
// something else, one line of count is nearly free and keeps a waiting
// ask from feeling lost between session starts.
if (dormant.length > 0) {
blocks.push(`${dormant.length} dormant ask(s), not listed here. ${DORMANT_NOTE}`);
}
pi.sendMessage(
{
customType: "mempalace-mailbox",
content:
`MemPalace logstream mailbox: ${due.length} directed ask(s) addressed to ` +
`"${mailboxAddress}" with no terminal reply from this device. ${OWED_HOWTO}\n\n` +
formatOwed(due),
`MemPalace logstream mailbox.\n\n${QUEUED_NOTE}\n\n` +
blocks.join("\n\n---\n\n"),
display: true,
},
// "steer" is queue-on-arrival, NOT an interrupt: at agent_settled the
@@ -1157,6 +1693,23 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
// than auto-delivery and is not what was approved.
{ deliverAs: "steer" },
);
// AFTER the queueing call, so a notification can never be the only
// thing that happened: the message is in the thread first, then we
// point at it. Wording names the nudge explicitly, because "you have
// mail" without "press a key" reproduces the exact confusion B fixes.
if (mailboxNotifyEnabled) {
const parts = [
due.length > 0 ? `${due.length} ask(s) owed` : "",
newsRaw.length > 0 ? `${newsRaw.length} closed` : "",
].filter(Boolean);
const summary = `${parts.join(" + ")} for ${mailboxAddress} — send any message to handle`;
try {
if (ctx?.hasUI) ctx.ui.notify(`MemPalace mailbox: ${summary}`, "info");
} catch {
/* ignore: UI may be gone by the time the poll resolves */
}
if (mailboxNotifyTerminal) notifyTerminal("MemPalace mailbox", summary);
}
} catch {
/* best-effort delivery: never break a turn over a mailbox read */
}
@@ -1187,10 +1740,11 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
sections.push(`## mempalace_diary_read\n\n(error: ${(err as Error).message})`);
}
// Tier 1 mailbox: what this device owes a reply to, derived rather than read
// off a filter. deriveOwed() swallows its own failures and returns [], so a
// broken palace costs a missing section, never a broken wake-up.
// off a filter. deriveOwed() swallows its own failures and returns an empty
// split, so a broken palace costs a missing section, never a broken wake-up.
if (mailboxEnabled) {
const owed = await deriveOwed();
const [owedSplit, closed] = await Promise.all([deriveOwed(), deriveClosed()]);
const { owed, dormant } = owedSplit;
if (owed.length > 0) {
// Record what this injection showed, so the first mid-session poll does
// not re-announce the identical list minutes later. Without this the
@@ -1204,6 +1758,27 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
`this device. ${OWED_HOWTO}\n\n${formatOwed(owed)}`,
);
}
// Once per session, WITH ids. The wake-up injection is the one place a
// dormant ask should be visible: "what is this device still carrying?" is a
// session-start question, and answering it here is what keeps the mid-session
// silence honest rather than concealing. Deliberately NOT added to
// `surfaced`: that map exists to suppress repeats of ANNOUNCED work, and a
// dormant ask must stay announceable the instant its baseline moves.
if (dormant.length > 0) {
sections.push(
`## logstream mailbox (${dormant.length} dormant — waiting, nothing owed yet)\n\n` +
`${DORMANT_NOTE}\n\n${formatOwed(dormant)}`,
);
}
const news = closed.filter((e) => e.id && !surfaced.has(e.id));
if (news.length > 0) {
const now = Date.now();
for (const e of news) if (e.id) surfaced.set(e.id, now);
sections.push(
`## logstream mailbox (${news.length} closed — no action owed)\n\n` +
`${CLOSED_NOTE}\n\n${formatOwed(news)}`,
);
}
}
if (sections.length === 0) return;
+156
View File
@@ -0,0 +1,156 @@
#!/usr/bin/env bash
# test-dormancy.sh — exercise the mailbox dormancy predicate (dormant_unless).
#
# Why this exists as a script and not a note: `evaluateDormancy` decides whether
# an ask is SHOWN to the agent, so a silent regression there does not look like a
# bug — it looks like an empty mailbox. The positive arms matter as much as the
# negative ones: a predicate that never fires makes the feature a no-op, and a
# predicate that fires too eagerly hides real work. Both arms run here.
#
# It cannot import extensions/pi/mempalace.ts in place, because that file imports
# `typebox`, which pi provides at RUNTIME and this repo has no node_modules for.
# So it copies the file into a temp tree with the resolved deps symlinked beside
# it. The copy is made BY this script on every run, so it cannot drift from the
# source the way a vendored duplicate would.
#
# Fixtures are created here rather than read from the host: the first draft used
# /etc/hostname and /etc/pi-devbox/build-manifest.json, which made the positive
# arms pass only on a pi-devbox container and silently flip to "not dormant"
# anywhere else — a device-dependent test that reports success by doing nothing.
#
# Exit codes, deliberately distinct:
# 0 every case behaved as specified
# 1 at least one case FAILED (a real regression)
# 2 INCONCLUSIVE — deps could not be resolved, so nothing was proven
# (never 0: a test that skips quietly is the failure mode it should catch)
set -uo pipefail
# ── Args ──────────────────────────────────────────────────────────────────────
if [[ "${1:-}" == "--help" || "${1:-}" == "-h" ]]; then
sed -n '2,25p' "$0" | sed 's/^# \?//'
exit 0
fi
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
SRC="$REPO_ROOT/extensions/pi/mempalace.ts"
[[ -r "$SRC" ]] || {
echo "INCONCLUSIVE: cannot read $SRC" >&2
exit 2
}
# ── Resolve pi's runtime deps (discovered, never hardcoded) ───────────────────
GLOBAL_ROOT="$(npm root -g 2>/dev/null || true)"
PI_PKG=""
for cand in \
"${GLOBAL_ROOT:+$GLOBAL_ROOT/@earendil-works/pi-coding-agent}" \
"/usr/lib/node_modules/@earendil-works/pi-coding-agent" \
"/usr/local/lib/node_modules/@earendil-works/pi-coding-agent"; do
[[ -n "$cand" && -d "$cand" ]] && {
PI_PKG="$cand"
break
}
done
[[ -n "$PI_PKG" && -d "$PI_PKG/node_modules/typebox" ]] || {
echo "INCONCLUSIVE: pi-coding-agent / typebox not resolvable (looked under 'npm root -g')." >&2
echo " Nothing was proven. Install pi, or run this where pi is installed." >&2
exit 2
}
# Node must be able to strip types from a .ts entry point (Node >= 22.6).
node -e 'process.exit(0)' 2>/dev/null || {
echo "INCONCLUSIVE: no usable node" >&2
exit 2
}
# ── Build the throwaway tree ──────────────────────────────────────────────────
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
mkdir -p "$WORK/extensions/pi" "$WORK/node_modules/@earendil-works"
cp "$SRC" "$WORK/extensions/pi/mempalace.ts"
ln -s "$PI_PKG/node_modules/typebox" "$WORK/node_modules/typebox"
ln -s "$PI_PKG" "$WORK/node_modules/@earendil-works/pi-coding-agent"
printf '{"type":"module"}\n' > "$WORK/package.json"
# Prove the copy is the source, so a PASS cannot be about a stale file.
if ! cmp -s "$SRC" "$WORK/extensions/pi/mempalace.ts"; then
echo "INCONCLUSIVE: the copy differs from the source" >&2
exit 2
fi
# ── Fixtures, owned by this test ──────────────────────────────────────────────
printf '{"release_tag":"v1.9.4","nested":{"deep":{"leaf":"found"}},"count":3}\n' > "$WORK/manifest.json"
printf 'not json at all\n' > "$WORK/notjson.txt"
: > "$WORK/stamp"
# ── The truth table ───────────────────────────────────────────────────────────
cat > "$WORK/run.mjs" <<'EOF'
import { statSync } from "node:fs";
import { evaluateDormancy } from "./extensions/pi/mempalace.ts";
const W = process.env.WORK;
const STAMP = `${W}/stamp`;
const MAN = `${W}/manifest.json`;
// Floor to whole seconds: that is the contract, and the residue below proves
// why the implementation must do the same.
const atBaseline = new Date(Math.floor(statSync(STAMP).mtimeMs / 1000) * 1000).toISOString();
const mt = (baseline, path = STAMP) => ({ kind: "file_mtime", path, baseline });
const jf = (field, baseline, path = MAN) => ({ kind: "json_field", path, field, baseline });
const cases = [
// No predicate => behave exactly as before the feature existed.
["no metadata", undefined, false],
["metadata without the key", { foo: "bar" }, false],
["dormant_unless is a string", { dormant_unless: "release != x" }, false],
["dormant_unless is empty", { dormant_unless: [] }, false],
["more than 8 conditions", { dormant_unless: Array(9).fill(mt(atBaseline)) }, false],
// POSITIVE arms. If these ever read false the feature is inert.
["file_mtime at baseline", { dormant_unless: [mt(atBaseline)] }, true],
["json_field at baseline", { dormant_unless: [jf("release_tag", "v1.9.4")] }, true],
["both at baseline", { dormant_unless: [jf("release_tag", "v1.9.4"), mt(atBaseline)] }, true],
["dotted field path", { dormant_unless: [jf("nested.deep.leaf", "found")] }, true],
["number vs string baseline", { dormant_unless: [jf("count", "3")] }, true],
// The trigger firing: ANY condition differing wakes the ask.
["file_mtime moved", { dormant_unless: [mt("2020-01-01T00:00:00Z")] }, false],
["json_field moved", { dormant_unless: [jf("release_tag", "v1.9.5")] }, false],
["one same, one moved", { dormant_unless: [mt(atBaseline), jf("release_tag", "v1.9.5")] }, false],
// FAIL-VISIBLE arms: dormancy unproven => announce.
["missing file", { dormant_unless: [mt(atBaseline, "/nonexistent/path")] }, false],
["unparseable baseline", { dormant_unless: [mt("not-a-date")] }, false],
["relative path", { dormant_unless: [{ kind: "file_mtime", path: "etc/hostname", baseline: atBaseline }] }, false],
["unknown kind", { dormant_unless: [{ kind: "uptime_lt", path: "/etc/hostname", baseline: "1d" }] }, false],
["missing json field", { dormant_unless: [jf("no_such_field", "x")] }, false],
["file is not JSON", { dormant_unless: [jf("a", "b", `${W}/notjson.txt`)] }, false],
["field is not scalar", { dormant_unless: [jf("nested", "x")] }, false],
["condition is not an object", { dormant_unless: ["release_tag"] }, false],
["baseline is an object", { dormant_unless: [{ kind: "file_mtime", path: STAMP, baseline: {} }] }, false],
];
let pass = 0;
const failed = [];
for (const [name, meta, want] of cases) {
const v = evaluateDormancy(meta);
const ok = v.dormant === want;
if (ok) pass++;
else failed.push(name);
console.log(
` ${ok ? "PASS" : "FAIL"} dormant=${String(v.dormant).padEnd(5)} want=${String(want).padEnd(5)} ` +
`${name.padEnd(28)} :: ${v.reason}`,
);
}
console.log(`\n ${pass}/${cases.length} passed`);
if (failed.length) console.log(` FAILED: ${failed.join(", ")}`);
// Recorded because it is the reason file_mtime floors to seconds: a real mtime
// carries sub-second residue that an ISO baseline does not.
console.log(` (stamp mtimeMs=${statSync(STAMP).mtimeMs}, baseline=${atBaseline})`);
process.exit(failed.length === 0 ? 0 : 1);
EOF
cd "$WORK" || exit 2
export WORK
node "$WORK/run.mjs"
exit $?
+134
View File
@@ -0,0 +1,134 @@
#!/usr/bin/env bash
# test-mcp-call-timeout.sh — the per-call deadline of RemoteMcpClient in
# extensions/pi/mempalace.ts, and the override that the feed's mine relies on.
#
# WHY THIS EXISTS. The feed tick calls `mempalace_mine` through
# `client.callTool()`. Its own deadline (MEMPALACE_FEED_MINE_TIMEOUT_MS,
# 300 000) was raised in 2026-09 because the mine legitimately takes 30-60 s on
# a shared single-writer palace — but callTool() carried no way to pass that
# deadline down, so the transport's generic per-request timeout
# (MEMPALACE_MCP_TIMEOUT_MS, 60 000) fired first, and operators saw
# feed (tick) failed: mempalace remote request 'tools/call' failed: timed out after 60000ms
# instead of the message the 2026-09 change had aimed at. The outer race was
# unreachable in practice. This harness pins both halves of the contract:
# 1. a plain callTool() still honours the short per-request timeout
# (a query taking that long really is wedged — keep it short);
# 2. callTool(name, args, { timeoutMs }) honours the override, so a long
# server-side job can be given its own, longer, deadline.
# Before the fix, (2) fails: the third argument was silently ignored.
#
# Like test-owed-withdrawal.sh, it runs the SHIPPED text: the class is cut out
# of mempalace.ts by brace matching, type-stripped with node's own stripper,
# and driven against a local JSON-RPC server that delays tools/call.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
SRC="${1:-$REPO_ROOT/extensions/pi/mempalace.ts}"
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
[ -r "$SRC" ] || { echo "FAIL: cannot read $SRC" >&2; exit 2; }
# ---------------------------------------------------------------- extractor ---
cat >"$WORK/extract.mjs" <<'EXTRACT'
import { readFileSync, writeFileSync } from "node:fs";
import { stripTypeScriptTypes } from "node:module";
const src = readFileSync(process.argv[2], "utf8");
/** Cut a top-level `const NAME = ...;` out of the source, verbatim. */
function decl(name) {
const start = src.indexOf(`const ${name} =`);
if (start < 0) throw new Error(`declaration not found: ${name}`);
return src.slice(start, span(start));
}
/** Cut `class NAME ... { ... }` out of the source, verbatim. */
function klass(name) {
const start = src.indexOf(`class ${name} `);
if (start < 0) throw new Error(`class not found: ${name}`);
return src.slice(start, span(start, true));
}
/** End offset of the statement starting at `start` (brace-matched). */
function span(start, braceOnly = false) {
let depth = 0, inStr = null, sawBrace = false;
for (let i = start; i < src.length; i++) {
const c = src[i], prev = src[i - 1];
if (inStr) { if (c === inStr && prev !== "\\") inStr = null; continue; }
if (c === '"' || c === "'" || c === "`") { inStr = c; continue; }
if (c === "/" && src[i + 1] === "/") { i = src.indexOf("\n", i); if (i < 0) break; continue; }
if (c === "/" && src[i + 1] === "*") { i = src.indexOf("*/", i) + 1; continue; }
if (c === "{" || c === "(" || c === "[") { depth++; if (c === "{") sawBrace = true; }
else if (c === "}" || c === ")" || c === "]") { depth--; if (braceOnly && sawBrace && depth === 0) return i + 1; }
else if (!braceOnly && c === ";" && depth === 0) return i + 1;
}
throw new Error(`unterminated statement at ${start}`);
}
const parts = [decl("num"), decl("REMOTE_PROTOCOL_VERSION"), decl("REMOTE_CLIENT_INFO"), klass("RemoteMcpClient")];
for (const [i, p] of parts.entries()) if (p.length < 30) throw new Error(`extraction ${i} implausibly short: ${p}`);
if (!parts[3].includes("tools/call")) throw new Error("RemoteMcpClient does not mention tools/call");
const ts = parts.join("\n\n") + "\nexport { RemoteMcpClient };\n";
writeFileSync(process.argv[3], stripTypeScriptTypes(ts, { mode: "strip" }));
EXTRACT
node --no-warnings "$WORK/extract.mjs" "$SRC" "$WORK/client.mjs" || exit 2
node --check "$WORK/client.mjs" || { echo "FAIL: extracted client does not parse" >&2; exit 2; }
echo "[extract] pulled RemoteMcpClient from $(basename "$SRC") ($(wc -c <"$WORK/client.mjs") bytes)"
# ------------------------------------------------------------------ harness ---
cat >"$WORK/run.mjs" <<'RUN'
import { createServer } from "node:http";
import { RemoteMcpClient } from "./client.mjs";
// A sessionless JSON-RPC server like `mempalace-mcp --transport http`:
// initialize / tools/list answer at once; tools/call sleeps SLOW_MS first.
const SLOW_MS = 1500;
const server = createServer((req, res) => {
let body = "";
req.on("data", (c) => (body += c));
req.on("end", () => {
const msg = JSON.parse(body);
if (msg.id === undefined) { res.writeHead(202); res.end(); return; } // notification
const reply = (result) => {
res.writeHead(200, { "content-type": "application/json" });
res.end(JSON.stringify({ jsonrpc: "2.0", id: msg.id, result }));
};
if (msg.method === "initialize") return reply({ protocolVersion: "2024-11-05", capabilities: {}, serverInfo: { name: "fake", version: "0" } });
if (msg.method === "tools/list") return reply({ tools: [{ name: "slow", description: "", inputSchema: { type: "object" } }] });
if (msg.method === "tools/call") return void setTimeout(() => reply({ content: [{ type: "text", text: "done" }] }), SLOW_MS);
reply({});
});
});
await new Promise((r) => server.listen(0, "127.0.0.1", r));
const url = `http://127.0.0.1:${server.address().port}/mcp`;
let failures = 0;
const check = (ok, label) => { console.log(`${ok ? "ok " : "FAIL"} ${label}`); if (!ok) failures++; };
// Short per-request deadline, deliberately below SLOW_MS.
process.env.MEMPALACE_MCP_TIMEOUT_MS = "400";
const client = new RemoteMcpClient(url);
await client.start();
check(client.alive === true, "start(): initialize + tools/list complete, client alive");
// 1. A plain call honours the short deadline — this is the wedged-query guard.
let err = null;
const t0 = Date.now();
try { await client.callTool("slow", {}); } catch (e) { err = e; }
check(err !== null && /timed out after 400ms/.test(err.message), `plain callTool() rejects at the per-request deadline (${err && err.message})`);
check(Date.now() - t0 < SLOW_MS, "…and rejects BEFORE the server would have answered");
check(client.alive === false, "a timeout marks the client unhealthy (documented side effect; ensureAlive() revives)");
// 2. A call with its own deadline outlives the generic one — the feed's mine.
await client.ensureAlive();
err = null;
let result = null;
try { result = await client.callTool("slow", {}, { timeoutMs: 5000 }); } catch (e) { err = e; }
check(err === null && result && result.content?.[0]?.text === "done", `callTool(name, args, { timeoutMs: 5000 }) waits past the generic deadline and gets the result (${err ? err.message : "ok"})`);
// 3. The override is itself a deadline, not "no deadline".
err = null;
try { await client.callTool("slow", {}, { timeoutMs: 200 }); } catch (e) { err = e; }
check(err !== null && /timed out after 200ms/.test(err.message), `callTool(..., { timeoutMs: 200 }) rejects at ITS deadline (${err && err.message})`);
server.close();
console.log(failures === 0 ? "PASS: all assertions held" : `FAIL: ${failures} assertion(s) failed`);
process.exit(failures === 0 ? 0 : 1);
RUN
cd "$WORK" && node --no-warnings run.mjs
+312
View File
@@ -0,0 +1,312 @@
#!/usr/bin/env bash
# test-owed-withdrawal.sh — rule tests for the owed-set derivation in
# extensions/pi/mempalace.ts: isStrictlyAfter, isAnswered, isWithdrawn.
#
# WHY THIS EXISTS
# isWithdrawn changes who is allowed to clear an obligation, which is the one
# thing in this extension that can go wrong SILENTLY. A wrong rule here does
# not throw and does not show up in a log — it makes a real unanswered ask
# vanish from a mailbox forever. So the rules get pinned down before they ship.
#
# WHY IT EXTRACTS THE PREDICATES FROM THE SOURCE INSTEAD OF RESTATING THEM
# These are closures inside createExtension(), so they cannot be imported. The
# tempting shortcut is to paste a copy of the logic into the test — which tests
# the copy. This repo has already paid for a divergent second copy (the
# pi-extensions skill mirror, 9579 B behind for weeks; the shell-lint logic
# extracted to one file for the same reason). So the harness cuts the actual
# declaration text out of mempalace.ts by brace matching and runs THAT. Change
# the source and this test follows; paraphrase the source and it cannot.
#
# FIXTURES ARE VERBATIM REAL EVENTS from the fleet log (project/pi-devbox), not
# invented shapes — including the exact incident that motivated isWithdrawn:
# mbp-m1-2020 withdrew a v1.8.13 rollout ask to tor-ms22 at seq 119 and the
# withdrawal had no effect, so tor-ms22 was still being told it owed a reply 41h
# later for a release it never installed.
#
# Usage: scripts/test-owed-withdrawal.sh [source.ts]
# exit 0 = all rules behave 1 = a rule broke
# exit 2 = source unreadable or does not parse 3 = the gate itself cannot run
#
# The optional argument exists so the suite can be pointed at a deliberately
# MUTATED copy of the source to prove it is sensitive — a suite that has never
# been observed to fail is not evidence. See the mutation check in the CHANGELOG
# entry that introduced isWithdrawn.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
SRC="${1:-$REPO_ROOT/extensions/pi/mempalace.ts}"
WORK="$(mktemp -d)"
trap 'rm -rf "$WORK"' EXIT
[ -r "$SRC" ] || { echo "FAIL: cannot read $SRC" >&2; exit 2; }
# A gate that cannot run must not pass — the standing rule in this repo.
#
# NOT `node --check`: it does not type-strip, so it rejects ANY TypeScript —
# `const x: number = 1` included, not just the inline type-import at the top of
# mempalace.ts. It passed on node 22.x and stopped passing on node 24.x
# (measured on pi-devbox v1.9.1, node v24.21.0: this gate exited 2 on an
# unmodified mempalace.ts, so all 17 assertions below refused to run). Strip
# first, then syntax-check the emitted JS. `mode: "strip"` blanks type syntax
# without moving anything, so byte offsets and line numbers survive and a
# reported error line still points at the right line of the ORIGINAL .ts.
#
# The two failure modes are reported separately and on purpose. "Cannot strip"
# is a fact about the toolchain; "does not parse" is a fact about the source.
# Collapsing them is what made this very defect present itself as
# "mempalace.ts does not parse" when mempalace.ts was fine.
#
# SC2016 is disabled deliberately: the single quotes are the point. What follows
# is JavaScript, and `${process.version}` must reach node, not be expanded by the
# shell first.
# shellcheck disable=SC2016
node --no-warnings -e '
const { readFileSync, writeFileSync } = require("node:fs");
const { stripTypeScriptTypes } = require("node:module");
if (typeof stripTypeScriptTypes !== "function") {
console.error(`FAIL: node ${process.version} cannot strip TypeScript ` +
`(module.stripTypeScriptTypes needs >= 22.13) — the gate is unavailable, ` +
`which is NOT a statement about the source`);
process.exit(3);
}
let js;
try {
js = stripTypeScriptTypes(readFileSync(process.argv[1], "utf8"), { mode: "strip" });
} catch (err) {
// The stripper is itself a parser, so a syntax error lands HERE, not in the
// --check below. This is a statement about the source: exit 2, not 3.
console.error(`FAIL: ${process.argv[1]} does not parse: ${err.message}`);
process.exit(2);
}
writeFileSync(process.argv[2], js);
' "$SRC" "$WORK/stripped.mjs" || exit $?
node --check "$WORK/stripped.mjs" \
|| { echo "FAIL: $SRC does not parse (stripped output rejected)" >&2; exit 2; }
# ---------------------------------------------------------------- extractor ---
cat >"$WORK/extract.mjs" <<'EXTRACT'
import { readFileSync, writeFileSync } from "node:fs";
const src = readFileSync(process.argv[2], "utf8");
/**
* Cut `const <name> = ...;` out of the source by matching delimiters, so the
* test runs the shipped text. Returns the declaration verbatim.
*/
function decl(name) {
const start = src.indexOf(`const ${name} =`);
if (start < 0) throw new Error(`declaration not found in source: ${name}`);
let depth = 0;
let inStr = null;
for (let i = start; i < src.length; i++) {
const c = src[i];
const prev = src[i - 1];
if (inStr) {
if (c === inStr && prev !== "\\") inStr = null;
continue;
}
if (c === '"' || c === "'" || c === "`") { inStr = c; continue; }
if (c === "/" && src[i + 1] === "/") { i = src.indexOf("\n", i); if (i < 0) break; continue; }
if (c === "{" || c === "(" || c === "[") depth++;
else if (c === "}" || c === ")" || c === "]") depth--;
else if (c === ";" && depth === 0) return src.slice(start, i + 1);
}
throw new Error(`unterminated declaration: ${name}`);
}
const parts = ["TERMINAL_STATUS", "isStrictlyAfter", "isAnswered", "isWithdrawn"].map(decl);
// Sanity: the extractor must have found real bodies, not empty matches. A
// silently-empty extraction would make every assertion below pass vacuously.
for (const [i, p] of parts.entries()) {
if (p.length < 40) throw new Error(`extracted declaration ${i} is implausibly short: ${p}`);
}
if (!parts[3].includes("withdraws")) throw new Error("isWithdrawn does not mention its marker key");
writeFileSync(process.argv[3], parts.join("\n\n"));
EXTRACT
node "$WORK/extract.mjs" "$SRC" "$WORK/extracted.ts"
echo "[extract] pulled 4 declarations from $(basename "$SRC") ($(wc -c <"$WORK/extracted.ts") bytes)"
# ------------------------------------------------------------------ harness ---
{
cat <<'HEAD'
type LogEvent = {
id?: string;
seq?: number;
hlc?: string;
type?: string;
status?: string;
from_agent?: string;
to_agent?: string;
correlation_id?: string | null;
created_at?: string;
body?: string;
metadata?: Record<string, unknown> | null;
};
const mailboxAddress = "pi@tor-ms22";
HEAD
cat "$WORK/extracted.ts"
cat <<'TAIL'
// ---- VERBATIM fixtures from the real fleet log, stream project/pi-devbox ----
const REP = "rep_d344e349ba276d6fc11997cd552f6937";
/** seq 112 — mbp's v1.8.13 rollout ask to tor-ms22. The obligation. */
const ask112: LogEvent = {
id: "evt_20260907T125110_2bef6f9b07a8",
seq: 112,
hlc: `1788785470242-000000-${REP}`,
type: "task.request",
status: "open",
from_agent: "pi@mbp-m1-2020",
to_agent: "pi@tor-ms22",
correlation_id: "v1813-client-rollout-tor-ms22",
};
/** seq 119 — mbp's withdrawal of its OWN ask. Terminal, directed, marker present. */
const withdraw119: LogEvent = {
id: "evt_20260908T223949_8f5b9bec7c92",
seq: 119,
hlc: `1788907189286-000000-${REP}`,
type: "task.reply",
status: "superseded",
from_agent: "pi@mbp-m1-2020",
to_agent: "pi@tor-ms22",
correlation_id: "v1813-client-rollout-tor-ms22",
metadata: {
closes: "v1813-client-rollout-tor-ms22",
nothing_owed: "no reply required to this event",
replacement: "v1814-client-rollout-tor-ms22",
},
};
/** seq 120 — the LIVE v1.8.14 ask. Must survive the v1813 withdrawal. */
const ask120: LogEvent = {
id: "evt_20260908T224118_808de43d6982",
seq: 120,
hlc: `1788907278075-000000-${REP}`,
type: "task.request",
status: "open",
from_agent: "pi@mbp-m1-2020",
to_agent: "pi@tor-ms22",
correlation_id: "v1814-client-rollout-tor-ms22",
};
/** seq 122 — tor-ms22's OWN terminal reply, which is what actually cleared 112. */
const myReply122: LogEvent = {
id: "evt_20260909T061851_cf9bd18800db",
seq: 122,
hlc: `1788934731146-000000-${REP}`,
type: "task.reply",
status: "superseded",
from_agent: "pi@tor-ms22",
to_agent: "pi@mbp-m1-2020",
correlation_id: "v1813-client-rollout-tor-ms22",
metadata: { closes: "v1813-client-rollout-tor-ms22 — from the RECIPIENT side, which is the only side that can" },
};
const clone = (e: LogEvent, over: Partial<LogEvent>): LogEvent => ({ ...e, ...over });
let failed = 0;
const check = (name: string, expected: boolean, actual: boolean): void => {
const ok = expected === actual;
if (!ok) failed++;
console.log(`${ok ? " ok " : " FAIL"} ${name} (expected ${expected}, got ${actual})`);
};
console.log("\nisWithdrawn — the requester retracting its own ask");
// THE INCIDENT. Before this rule existed the answer was false and tor-ms22 was
// told it owed a reply for a release that no longer existed.
check("real seq 119 withdraws real seq 112", true, isWithdrawn(ask112, [withdraw119]));
check("withdrawal does NOT touch the live v1814 ask", false, isWithdrawn(ask120, [withdraw119]));
console.log("\nisWithdrawn — the ways a mailbox must NOT be silently emptied");
check(
"no marker: a bare terminal event from the requester is not a withdrawal",
false,
isWithdrawn(ask112, [clone(withdraw119, { metadata: { nothing_owed: "no reply required" } })]),
);
check(
"marker naming a DIFFERENT thread does not clear this one",
false,
isWithdrawn(ask112, [clone(withdraw119, { metadata: { closes: "v1814-client-rollout-tor-ms22" } })]),
);
check(
"marker carrying PROSE does not count (tor-ms22's own seq 122 shape)",
false,
isWithdrawn(ask112, [
clone(withdraw119, {
metadata: { closes: "v1813-client-rollout-tor-ms22 — from the RECIPIENT side, which is the only side that can" },
}),
]),
);
check(
"a THIRD PARTY cannot retract someone else's ask",
false,
isWithdrawn(ask112, [clone(withdraw119, { from_agent: "pi@emb-7kj4vr4g" })]),
);
check(
"a BROADCAST cannot empty every machine's mailbox at once",
false,
isWithdrawn(ask112, [clone(withdraw119, { to_agent: "*" })]),
);
check(
"a NON-TERMINAL status is not a withdrawal",
false,
isWithdrawn(ask112, [clone(withdraw119, { status: "open" })]),
);
check(
"an event addressed to a DIFFERENT device does not clear my ask",
false,
isWithdrawn(ask112, [clone(withdraw119, { to_agent: "pi@emb-7kj4vr4g" })]),
);
check(
"ORDERING: a withdrawal cannot retire an ask the requester sent LATER",
false,
isWithdrawn(clone(ask112, { seq: 121, hlc: `1788907300000-000000-${REP}` }), [withdraw119]),
);
console.log("\nisWithdrawn — accepted marker spellings");
check(
"canonical `withdraws` naming the correlation",
true,
isWithdrawn(ask112, [clone(withdraw119, { metadata: { withdraws: "v1813-client-rollout-tor-ms22" } })]),
);
check(
"marker naming the ask's EVENT ID instead of its correlation",
true,
isWithdrawn(ask112, [clone(withdraw119, { metadata: { withdraws: ask112.id } })]),
);
console.log("\nisAnswered — regression guard, unchanged behaviour");
check("my own terminal reply still closes my own ask", true, isAnswered(ask112, [myReply122]));
check("my reply on one thread does not close another", false, isAnswered(ask120, [myReply122]));
check(
"ORDERING still load-bearing: my reply cannot pre-close a LATER ask",
false,
isAnswered(clone(ask112, { seq: 130, hlc: `1788999999999-000000-${REP}` }), [myReply122]),
);
console.log("\nEND TO END — the derived owed set for tor-ms22 on 2026-09-09");
// The behaviour change, stated as the derivation states it. `mine` deliberately
// EXCLUDES seq 122 so this measures the new rule rather than the reply that
// happened to be written first.
const candidates = [ask112, ask120];
const mine: LogEvent[] = [];
const inbound = [withdraw119];
const owed = candidates.filter((c) => !isAnswered(c, mine) && !isWithdrawn(c, inbound));
const owedIds = owed.map((e) => e.correlation_id).join(", ");
check("exactly one ask remains owed", true, owed.length === 1);
check("and it is the v1814 one, not the withdrawn v1813", true, owedIds === "v1814-client-rollout-tor-ms22");
console.log(` derived owed set: [${owedIds}]`);
console.log(failed === 0 ? "\n=== PASSED ===" : `\n=== FAILED: ${failed} ===`);
process.exit(failed === 0 ? 0 : 1);
TAIL
} >"$WORK/harness.ts"
node --experimental-strip-types "$WORK/harness.ts"
+88
View File
@@ -0,0 +1,88 @@
#!/usr/bin/env bash
# test-rsync-ship-idempotency.sh — regression test for the ship-step bug fixed
# 2026-08-27 (bin/mempalace-pi-session: rsync --update -> --checksum).
#
# THE BUG: the stage file's mtime is deliberately the SOURCE transcript's mtime
# (os.utime() at :903, "preserve session mtime for dedup stability"), so a
# re-export of a session that has not been appended to since the last ship
# carries a mtime that is NOT newer than the receiver copy. rsync --update
# skips a file whose mtime is not strictly newer than the destination's, so a
# scrubbed re-export of a DORMANT session was silently dropped: content
# differed, mtime did not, "success" was reported, and unscrubbed bytes stayed
# in the palace host's inbox indefinitely. Measured on mbp-m1-2020 2026-08-27.
#
# THE ASSERTION mirrors that exactly and needs no ssh, no palace, no secret:
# stage a file, ship it, rewrite the content while RESTORING the original
# mtime (the same thing os.utime() does), ship again, assert the receiver's
# sha256 changed. On the pre-fix flag (--update) this fails; with --checksum
# (content comparison, mtime-independent) it passes.
#
# Deliberately does NOT invoke bin/mempalace-pi-session itself: that needs a
# live ssh target, a palace, MEMPALACE_* env — disproportionate scaffolding for
# what is, at its core, one rsync flag. Pulls the flag list out of the script
# instead of hand-copying it, so a future change to the ship command either
# updates this test's expectation or fails it loudly rather than drifting
# silently out of sync with what actually ships.
#
# Usage: scripts/test-rsync-ship-idempotency.sh (exit 0 = fix still holds)
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
SHIP_SCRIPT="$REPO_ROOT/bin/mempalace-pi-session"
SRC="$(mktemp -d)"; DST="$(mktemp -d)"
trap 'rm -rf "$SRC" "$DST"' EXIT
EXPECT='rsync -a --checksum --no-owner --no-group'
if ! grep -qF "$EXPECT" "$SHIP_SCRIPT"; then
echo "FAIL: $SHIP_SCRIPT no longer ships with '$EXPECT'." >&2
echo " Either the fix regressed, or the flags changed -- update this" >&2
echo " test's EXPECT string to match, don't just delete the test." >&2
exit 1
fi
ship() {
rsync -a --checksum --no-owner --no-group \
--include='*.jsonl' --exclude='*' \
"$SRC/" "$DST/"
}
# Same BYTE LENGTH on purpose, mirroring mbp-m1-2020's own reasoning for why
# dropping --update outright would not be enough: rsync's default quick check
# (no --update, no --checksum) already transfers when SIZE differs, so a test
# with mismatched lengths would pass by accident and prove nothing about
# --checksum specifically. A real redaction whose placeholder happens to match
# the secret's length is exactly the case that stays silently undetected unless
# content itself, not size or mtime, is compared.
UNSCRUBBED='live: MEMPALACE_REMOTE_TOKEN=AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA'
SCRUBBED='live: MEMPALACE_REMOTE_TOKEN=<redacted:MEMPALACE_REMOTE_TOKEN---------->'
if [[ ${#UNSCRUBBED} -ne ${#SCRUBBED} ]]; then
echo "FAIL: test fixture bug -- the two fixture strings must be equal length (${#UNSCRUBBED} vs ${#SCRUBBED}); this test would not exercise --checksum at all otherwise." >&2
exit 1
fi
# 1. Baseline: ship a session, receiver now matches sender.
printf '%s\n' "$UNSCRUBBED" > "$SRC/session.jsonl"
touch -d '2026-08-23T23:53:18' "$SRC/session.jsonl"
ship
before="$(sha256sum "$DST/session.jsonl" | cut -d' ' -f1)"
# 2. The bug's exact shape: rewrite content (same length!) as a scrubbed
# re-export would, then restore the ORIGINAL mtime, as the exporter's
# os.utime() does.
printf '%s\n' "$SCRUBBED" > "$SRC/session.jsonl"
touch -d '2026-08-23T23:53:18' "$SRC/session.jsonl"
ship
after="$(sha256sum "$DST/session.jsonl" | cut -d' ' -f1)"
if [[ "$before" == "$after" ]]; then
echo "FAIL: receiver sha256 unchanged after a content-only re-export at a held-constant mtime." >&2
echo " This is the exact defect --checksum was added to fix." >&2
exit 1
fi
if ! grep -q '<redacted:MEMPALACE_REMOTE_TOKEN' "$DST/session.jsonl"; then
echo "FAIL: receiver does not contain the scrubbed content." >&2
exit 1
fi
echo "PASS: ship step transfers changed content even when mtime is deliberately held constant (before=$before after=$after)"