2167a1b0332400f0d616729c11e6785b8f0ea6cd
6 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
817b3a82b7 |
fix(feed): the mine's deadline never reached the transport; 60 s cut it off
The 2026-09 change raised MEMPALACE_FEED_MINE_TIMEOUT_MS to 300 000 and
raced it against client.callTool("mempalace_mine"). But callTool() had no
way to carry a deadline, so every call went out under the transport's
generic per-request timeout (MEMPALACE_MCP_TIMEOUT_MS, 60 000), which fired
first on every honest 60 s+ mine. Operators saw
feed (tick) failed: mempalace remote request 'tools/call' failed:
timed out after 60000ms
instead of the message the change had aimed at, and the 300 s was
unreachable. On stdio it was worse than noise: that transport kills the
server child on timeout, so the mine was actually aborted at 60 s.
callTool(name, args, { timeoutMs }) now passes a per-call deadline to both
transports; feedPalace() uses it for the mine. Plain calls keep the short
default — a query taking 60 s is still wedged. The Promise.race stays as the
liveness guard for a transport with its timeout disabled (0).
scripts/test-mcp-call-timeout.sh cuts RemoteMcpClient out of the shipped
file (as test-owed-withdrawal.sh does for the mailbox predicates), drives it
against a local JSON-RPC server that delays tools/call, and asserts: plain
call rejects at the generic deadline; the override outlives it; the override
is itself a deadline. Fails on the previous commit (2 of 6), passes here.
Vendored-copy delta noted in the header; protocol untouched, sync token
unchanged (check-mcp-client-sync.sh passes against pi-extensions).
|
||
|
|
dab989b068 |
fix(mailbox-tests): the owed-set suite's own gate refused to let it run on node 24
scripts/test-owed-withdrawal.sh has not executed a single assertion since the
image moved to node 24. Its precondition line
node --experimental-strip-types --check "$SRC"
exits 2 on an unmodified extensions/pi/mempalace.ts, and the script is
`set -euo pipefail` with `|| exit 2`, so all 17 assertions and both regression
guards were skipped. Measured on pi-devbox v1.9.1, node v24.21.0.
CAUSE, MEASURED AND NARROWER THAN IT LOOKS. The reported error names the inline
type-import at mempalace.ts:92, which invites the reading "that import is
unusual". It is not the import: `node --check` does not type-strip AT ALL. A file
whose entire content is `const x: number = 1;` fails identically, with or without
--experimental-strip-types, while `node --experimental-strip-types <same file>`
executes it fine. --check has never been type-aware; execution is what gained
stripping. So this gate could never validate TypeScript, on any node.
WHY IT USED TO PASS -- INFERENCE, NOT MEASUREMENT. e2b060a's message records this
suite running on 2026-09-09 with a control pass and four mutation kills, when the
image shipped node v22.23.2; v1.9.1 ships v24.21.0. No node 22 exists on the box
where this was diagnosed, so the counterfactual was not executed. Treat "22
stripped for --check and 24 stopped" as a hypothesis consistent with the record,
not as a measured cause. What IS measured is the present-tense behaviour above.
WHY THIS MATTERS MORE THAN A RED TEST. This suite exists because isWithdrawn is
the one rule in the extension that can go wrong SILENTLY -- a wrong rule does not
throw and does not log, it makes a real unanswered ask vanish from a mailbox
forever. The rule shipped in
|
||
|
|
e2b060a940 |
feat(mailbox): let a requester withdraw its own ask, with an explicit marker
RFC 003 §3.3 clause 3 clears an ask only on "no event OF YOURS", and the skill states the consequence outright: "there is nothing anyone can do about it from the other end". That asymmetry is deliberate and mostly right — owed-ness is a statement about the RECIPIENT's accountability. It is wrong in exactly one case: the requester retracting its own ask. MEASURED COST. pi@mbp-m1-2020 withdrew a v1.8.13 rollout ask to pi@tor-ms22 at seq 119 — terminal `superseded`, same correlation_id, metadata.closes naming the thread, body "DO NOT SPEND A MINUTE ON v1.8.13" — and recorded it as done. It had no effect: seq 119's from_agent is mbp, so it could never satisfy a join that only inspects tor-ms22's own events. tor-ms22's next wake-up, 8h later and on the first boot of the image shipping this very derivation, still listed the ask as owed, 41h old, for a release it never installed. The failure is invisible from the sender's side, which is why it went unnoticed for 41h. WHY AN EXPLICIT MARKER RATHER THAN "ANY TERMINAL EVENT FROM THE REQUESTER". This file's standing rule is that every failure stays on the noisy-but-visible side: a resurfacing item costs one turn of human correction, a suppressed unanswered ask is silent and permanent. Under the naive rule a requester appending `applied` for its own bookkeeping — on the correlation, addressed to me, before I ever replied — would silently delete a real obligation. So release must be STATED: metadata.withdraws (canonical) or metadata.closes (already this fleet's de-facto marker), whose value must NAME the ask — its correlation_id or its event id. Prose does not count. Five further guards, each pinned by a mutation test: only the original requester; addressed to this device exactly, never '*'; terminal status; strictly after the ask by hlc; and joined by ack_of or correlation_id. This does NOT break the fixed point in §9.2. That section rejects letting terminal directed events into the owed set, because then every closure mints a fresh obligation. That is about CANDIDATES; this adds a CLEARER. A withdrawal is terminal, so it can never be a candidate, and the asserting shape (open) and the clearing shape (terminal) stay disjoint. Also fixed here, because the new rule depends on the same windowing: the `mine` query used event_list's DEFAULT `asc` order with limit 100, i.e. the OLDEST 100 events this device ever wrote. Once a device passes 100 authored events its most RECENT replies fall out of the join window and every ask it just answered resurfaces as owed. Latent, not theoretical — tor-ms22 was at ~20. Both windows are now anchored at the newest end with order: "desc". Verification: scripts/test-owed-withdrawal.sh, 17 assertions over VERBATIM fixtures from the real log (seq 112/119/120/122). It extracts the predicates from mempalace.ts by brace matching and runs the SHIPPED text rather than a pasted copy — this repo has already paid for a divergent second copy. Sensitivity proven by four mutations: removing the marker requirement flips exactly the 3 marker assertions, and disabling the third-party / broadcast / ordering guards each flip exactly their own. Control passes. |
||
|
|
a361b71c40 |
ship: don't trust the mtime the exporter deliberately backdates
rsync --update skips a file whose mtime is not strictly newer than the receiver's. The stage file's mtime IS the source transcript's mtime (os.utime() at :903, "preserve session mtime for dedup stability"), so re-exporting a session that has not been appended to since its last ship produces a mtime that is not newer than what's already at the receiver — exactly the case a redactor upgrade needs to ship, because content differs while mtime does not. --update reports success and sends nothing. Reported and patched by pi@mbp-m1-2020 (evt_20260827T211925_9674a31da0b4, artifact art_20260827T211839_fa52af563105, sha256 f217e47e…), measured live: a scrubbed re-export of a dormant session (pi_01a03022-…f542) sat unshipped in the palace host's inbox while every local signal reported a clean stage, saved only because a host-side sweep happened to rewrite the remote copy independently that same day. --checksum compares content and ignores size/mtime entirely. Dropping --update outright was considered and rejected: rsync's default quick check already transfers on a SIZE difference alone, which is why the observed case (33-byte placeholder vs a 43-byte token) would have been masked as "fixed" by a change that only works until a redaction whose placeholder happens to match the secret's length. os.utime() at :903 is untouched — its backdating is a separate, load-bearing design call for dedup stability, out of scope for this fix. Added scripts/test-rsync-ship-idempotency.sh: ships a file, rewrites its content to an EQUAL-LENGTH string while restoring the original mtime (what os.utime() does), ships again, asserts the receiver's sha256 changed. Equal length is deliberate, not cosmetic — mismatched lengths would pass via the quick check alone and prove nothing about --checksum specifically; this is the same reasoning that ruled out dropping --update. Verified the test discriminates: passes against today's --checksum, fails against --update (checked by temporarily substituting the flag in a copy, not committed). Runs offline — a local rsync destination path exercises the same size/mtime/checksum comparison as the ssh transfer, no palace or network needed. |
||
|
|
553d86570c |
provenance: stamp device+harness at the edge, not in the agent's head
RFC 001 §7.3.2 ranks "agent stamps provenance via a skill instruction" as the ❌ worst possible place — per-call boilerplate, forgettable, improvisable. It was right, and we had shipped exactly that: the mempalace skill told the agent to pass added_by="<harness>@<device>" by hand. Measured on the shared palace, 199 rows had reached it unresolvable, 10 of them filed by the very agent that wrote the instruction, in a drawer about host provenance. The trigger was a cross-host misattribution: a session on tor-ms22 read its own diary, could not tell that the entries were written on EMB-7KJ4VR4G, and reported another machine's verification as this one's. Move the same convention into the ⚠️ edge row, where it is uniform and unforgettable (§7.3.5): * extensions/pi/mempalace.ts defaults the writer field on every tool that has one — added_by (add_drawer, checkpoint), agent (mine), from_agent (event_append), created_by (artifact_put) — from $MEMPALACE_PI_DEVICE. An explicit value always wins, so filing for another device stays possible. The allowlist is per tool, never blanket: 3.8.0's dispatcher hard-rejects undeclared args with -32602, so injecting added_by into diary_write or kg_add (which have no such property) would break the call outright. * mine gets miner@<device> when the caller invokes it, but <harness>@<device> for the bridge's own transcript feed — bulk extraction is not agent-authored memory, and that keeps the pi/opencode/miner taxonomy honest. * diary_write has no metadata slot at all, and the device must never go in agent_name (wing = f"wing_{agent_name}" would splinter the diary per host). So the entry TEXT carries an AAAK field, HOST:<device>|SESSION:… — which is also the only channel a READER sees: search projects a fixed key set and diary_read returns content, so no metadata fix, not even a server-authoritative one, would have prevented the misattribution. * The wake-up block now states the device and warns that diary_read interleaves every machine's diary. * R1: doubly gated on MEMPALACE_PI_DEVICE and MEMPALACE_REMOTE_URL, so a solitary devbox stamps nothing and behaves exactly as before — which is also the correct semantics per §7.3.3. Version the reconciler that was living only on synlig (bin/ + contrib/systemd/), add --dry-run, and teach it two new rules: diary_host_marker reads the HOST: field, and sibling_chunk propagates a resolved origin across a drawer's chunks (a text marker lands in chunk 0 only, so a 5-chunk diary entry would otherwise stamp 1 and leave 4 blank). --dry-run against the real palace before deploying earned its keep twice, and scripts/test-device-stamp.sh pins both findings with the strings it found: HOST: was ALREADY in use with a composite grammar (HOST:emb-7kj4vr4g.f1d3c3f89e3e.v1.8.3.pi0.84.2) and for bare container ids, so an unvalidated rule invented devices like "f1d3c3f89e3e.pi0.84.2"; and HOST: also carries a different SENSE elsewhere (HOST:exec.via.ssh-controlmaster->…, meaning where I was executing). Validating against the known-device set both refuses those and recovers the composite entries correctly. A marker convention inherits every prior meaning of its own name. Deployed and verified on synlig: device 14,217 → 14,317, integrity ok, idempotent on immediate re-run, no invented device values. RFC updates: §7.3.1 corrected (the arg whitelist is a hard -32602 in 3.8.0, not a silent drop; get_drawer DOES return metadata, search structurally cannot; triples and logstream live in separate databases the stamper cannot reach), §7.3.5 added (what is deployed, including the divergence from §7.3.4's opaque origin_device — tor-ms22 vs tor-ms22-native is that cost already visible), and Phase 4 now carries per-device tokens motivated FIRST by revocation, with the finding that tokens are the cheap half: core holds one scalar auth_token and has zero device concept, so authoritative stamping needs a component we own. |
||
|
|
96699f2a17 |
feat(pi-bridge): external MemPalace transport via MEMPALACE_REMOTE_URL
Let the pi<->mempalace bridge connect to a shared MemPalace over HTTP instead of always spawning a local mempalace-mcp: - Extract IMcpClient; rename McpClient -> StdioMcpClient (ctor command, arg-less start()). - Add RemoteMcpClient (vendored from pi-extensions/mcp-loader.ts): streamable-HTTP with AbortController timeouts, protocolVersion pinned 2024-11-05, alive/ensureAlive. mempalace-mcp --transport http is sessionless JSON-RPC today; session/SSE/404 branches retained for a future streamable-HTTP server. - createClient() selects transport from MEMPALACE_REMOTE_URL; MEMPALACE_REMOTE_TOKEN -> Authorization: Bearer. Lifecycle automation (wake-up, /mempalace-diary) unchanged. - scripts/check-mcp-client-sync.sh: drift guard vs canonical mcp-loader.ts. - README: document local-vs-external transport. Typechecks clean (strict); both transports smoke-tested against live mempalace-mcp. |