9 Commits

Author SHA1 Message Date
joakimp e2b060a940 feat(mailbox): let a requester withdraw its own ask, with an explicit marker
RFC 003 §3.3 clause 3 clears an ask only on "no event OF YOURS", and the skill
states the consequence outright: "there is nothing anyone can do about it from
the other end". That asymmetry is deliberate and mostly right — owed-ness is a
statement about the RECIPIENT's accountability. It is wrong in exactly one case:
the requester retracting its own ask.

MEASURED COST. pi@mbp-m1-2020 withdrew a v1.8.13 rollout ask to pi@tor-ms22 at
seq 119 — terminal `superseded`, same correlation_id, metadata.closes naming the
thread, body "DO NOT SPEND A MINUTE ON v1.8.13" — and recorded it as done. It had
no effect: seq 119's from_agent is mbp, so it could never satisfy a join that
only inspects tor-ms22's own events. tor-ms22's next wake-up, 8h later and on the
first boot of the image shipping this very derivation, still listed the ask as
owed, 41h old, for a release it never installed. The failure is invisible from
the sender's side, which is why it went unnoticed for 41h.

WHY AN EXPLICIT MARKER RATHER THAN "ANY TERMINAL EVENT FROM THE REQUESTER".
This file's standing rule is that every failure stays on the noisy-but-visible
side: a resurfacing item costs one turn of human correction, a suppressed
unanswered ask is silent and permanent. Under the naive rule a requester
appending `applied` for its own bookkeeping — on the correlation, addressed to
me, before I ever replied — would silently delete a real obligation. So release
must be STATED: metadata.withdraws (canonical) or metadata.closes (already this
fleet's de-facto marker), whose value must NAME the ask — its correlation_id or
its event id. Prose does not count.

Five further guards, each pinned by a mutation test: only the original requester;
addressed to this device exactly, never '*'; terminal status; strictly after the
ask by hlc; and joined by ack_of or correlation_id.

This does NOT break the fixed point in §9.2. That section rejects letting
terminal directed events into the owed set, because then every closure mints a
fresh obligation. That is about CANDIDATES; this adds a CLEARER. A withdrawal is
terminal, so it can never be a candidate, and the asserting shape (open) and the
clearing shape (terminal) stay disjoint.

Also fixed here, because the new rule depends on the same windowing: the `mine`
query used event_list's DEFAULT `asc` order with limit 100, i.e. the OLDEST 100
events this device ever wrote. Once a device passes 100 authored events its most
RECENT replies fall out of the join window and every ask it just answered
resurfaces as owed. Latent, not theoretical — tor-ms22 was at ~20. Both windows
are now anchored at the newest end with order: "desc".

Verification: scripts/test-owed-withdrawal.sh, 17 assertions over VERBATIM
fixtures from the real log (seq 112/119/120/122). It extracts the predicates from
mempalace.ts by brace matching and runs the SHIPPED text rather than a pasted
copy — this repo has already paid for a divergent second copy. Sensitivity
proven by four mutations: removing the marker requirement flips exactly the 3
marker assertions, and disabling the third-party / broadcast / ordering guards
each flip exactly their own. Control passes.
2026-09-09 08:56:56 +02:00
joakimp 21023e7aa0 rfc-003: an event addressed to a nobody is write-only
§7.13 — the owed set keys on to_agent, and from_agent is whatever
string the writer puts there, so nothing stops authoring under an
identity no live session runs as. When the reply comes back addressed
to that string, it is stored and delivered to no one. Distinct from
§7.12: the defect is the address, not the status.

Measured 2026-08-27: two directed task.requests planted under a
synthetic from_agent (a provenance label, not a run identity) drew two
correct replies, addressed back to the label as the protocol requires.
Neither reply was ever delivered. One carried a live-credential
exposure finding; it sat unread for ~2h20m and was found only because
a human asked whether mail had arrived.

General rule stated, not just the instance: this fleet has already
produced the same failure shape three ways (synthetic sender above; an
event addressed to a decommissioned device; a device-identity case
mismatch). Permitted exception carried over from existing fleet
practice: a synthetic sender is fine for a controlled experiment, but
the body must then name the real reply-to identity.

Proposed (not implemented): warn at event_append when from_agent
differs from the writer's own session identity, using a DISTINCT
from_agent lookup to say whether anyone has ever authored under it —
stated honestly as a heuristic that catches 'never authored, certainly
unread' but not a decommissioned device that once did.

Decision 9 in the open-decisions list gets a numbered sibling, 10,
pointing at this section — related to decision 7 (agent-name
collision between two devices) but a different failure: not two
devices sharing a name, but one writer using a name nobody runs.
2026-08-27 23:34:20 +02:00
joakimp b2b50afcc1 rfc-003: propose a news surface that needs no cursor, and say why widening the owed set cannot work
Decision 9 (news vs obligations) gets a proposed direction in a new §9.2, following
how §9.1 promoted the retention decision.

The load-bearing part is the negative result. The tempting fix — let terminal
directed events into the owed set so a report addressed to a device reaches it —
breaks the derivation's fixed point. Owed-ness is "directed at me, not mine, and
not joined by a later terminal event of mine", so the asserting shape (open) and
the clearing shape (terminal) have to be disjoint. Make terminal events owed and a
reply becomes owed by its requester, whose closure is itself directed + terminal
and therefore owed by the original author: every closure mints a fresh obligation
and the loop never terminates. status="open" is not editorial taste about tone, it
is what makes owed-ness terminate. Worth writing down before someone "fixes" it.

The proposal itself avoids the cost that sinks the naive version. "Since your last
session" implies a per-device read cursor, and container-local state is exactly
what --force-recreate erases. No new state is needed: the device's own last
authored event is already a cursor, it lives in the shared log, and it is
comparable across replicas for the same reason the owed-set join now uses hlc
(§7.3). Named the non-obvious constraint too — the anchor must be computed PER
STREAM, because a global one lets a chatty stream advance past unread news in a
quiet one, and that failure is silent.

Also tightened §7.12: candidacy requires exactly `open`, not merely "non-terminal",
so a `claimed` announcement is as undelivered as a finished report. Claiming still
earns its keep on handoff-prone work (it is what distinguishes "nobody started"
from "someone started and the container died"), but its audience is a log reader,
not the requester — and it does not quiet the claimer's own mailbox either, which
is what the state machine's "open --> claimed does NOT clear" already implies.
2026-08-27 17:28:21 +02:00
joakimp ecc2a9c574 mailbox notify: record that the terminal path is unverified through tmux
Measured on tor-ms22: the layering here is kitty -> tmux (on the host) ->
docker exec -> pi (in the container), with two clients attached to one tmux
session. Test OSC sequences written straight to pi's tty produced no
notification on the remote client; the local client is still unobserved. Most
likely tmux is dropping OSC types it does not implement — reaching the outer
terminal needs tmux's DCS passthrough (ESC P tmux; ... ESC \, inner ESC doubled)
plus allow-passthrough on, which is not the default and is not implemented here.

The point worth keeping is structural, not incidental: the container can see
NEITHER layer. KITTY_WINDOW_ID is absent because docker exec does not forward it,
and TMUX is absent because tmux runs one level further out on the host. So a
containerised client cannot detect the terminal it is speaking to or the
multiplexer it must speak through, and autodetection is not merely unreliable
here — it is blind. Explicit configuration is the only route, which is why the
previous commit added forced kitty/osc777 modes.

No behaviour or defaults changed: in-TUI notify remains the default and the
terminal path stays opt-in. Also flagged the unsettled routing question — a
pane's output reaches every attached client, so a work laptop and a home machine
would both ping from one arriving ask.
2026-08-27 12:54:52 +02:00
joakimp a92c75d070 mailbox: say it is queued, and ping the human who is not looking (RFC 003 §7.11)
Two changes, neither touching the no-triggerTurn decision, which stands.

B — a delivery note in the message itself. The mid-session text explained how to
CLOSE an ask and never said when it would be SEEN, so the only reader who needed
that fact — the human watching an idle session — was the one not told. It now
says: this is a queued message, nothing woke the agent, any message starts the
turn that handles it, and the agent is not ignoring the ask, it is not running.
Costs nothing and changes no behaviour; it converts "why is it ignoring me?" into
"right, I nudge it". Deliberately NOT added to the wake-up injection, where a
turn is already starting and the note would be false.

A — a notification at poll time, because B only helps someone already looking and
the case that loses an ask is nobody looking. MEMPALACE_MAILBOX_NOTIFY: unset →
in-TUI ctx.ui.notify, the surface session_start already uses; =desktop →
additionally a terminal-native notification (Kitty OSC 99, else OSC 777), reusing
the detection the fleet's own notify.ts already proves in this harness; =0/off →
silent. The desktop path is how a ping escapes a container with no notify-send,
no DBus and no host access: the escape sequence is written to stdout and
interpreted by the terminal emulator on the human's own machine. Opt-in because
writing raw escapes is a behaviour change on a shared machine, not because it is
unreliable — say the word and the default flips.

Placement and firing conditions are deliberate: the notify call sits AFTER
sendMessage so a ping can never be the only thing that happened, and it fires
only when something is due — the same condition as delivery. A notification on
an empty poll would train its reader to ignore it, which is the failure this
whole feature exists to reverse. The wording names the nudge ("send any message
to handle") because "you have mail" without "press a key" reproduces exactly the
confusion B fixes.

Tested: 12 cases. Mode parsing (unset/empty/desktop/DESKTOP-with-space/0/off/1),
escape hygiene, and the OSC invariant that matters — a hostile title or body
containing ";" or a BEL cannot forge an OSC field or terminate the sequence
early, which is worth asserting because event bodies arrive from other machines.
ctx.hasUI is checked and the notify call is wrapped, since the UI can be gone by
the time an unawaited poll resolves. Syntax checked with node --strip-types.
2026-08-26 23:56:21 +02:00
joakimp 982b001001 docs: RFC 003 §7.11/§7.12 — delivered is not read, and the mailbox carries obligations not news
Both measured 2026-08-26 by peer devices, and the second one measured on me:
the operator had to point at an event id before I read a report that had been
addressed to me for two hours. The mailbox was working correctly the whole time.

§7.12 is the structural one. Candidates are drawn with status="open", so an
event carrying a TERMINAL status is not a mailbox candidate at all, whoever it
is addressed to. A task.reply written to a named device to share a finding is
therefore delivered to nobody, ever — and neither is any event.ack. So the most
natural inter-device message, "here is something you should know", is exactly
the shape that gets no delivery. Two things that do work: address it as a
directed status="open" ask (correct when a response is actually wanted), or
accept it as pull-only and pair it with a drawer, which the peer's search will
surface. A terminal report plus an expectation of attention does not.

§7.11 is the peer's finding, and it explains the other half of why that report
sat unread: delivery is deliverAs "steer" with deliberately no triggerTurn, and
the poll fires on agent_settled — i.e. when the agent is IDLE, with no inference
running. The text is queued for the next turn, so the human is the trigger.
Measured on another device: a delivery landed at ~20:50Z and sat visibly
unreacted-to until the operator asked "do I have to nudge you?". Not a defect;
the no-triggerTurn decision was deliberate and stands. But the delivered text
explains how to CLOSE an ask and never says when it will be SEEN, so the one
reader who needs that fact — the human watching the window — is the one not
told. Recorded with the three fixes that do not wake a model, as decision 8.

§7.4 is upgraded from inferred to measured, on two devices independently, by a
two-arm control: a to_agent='*' event with status='open' survives the raw
candidate filter and therefore reaches the exclusion branch, which no previously
written broadcast (all non-open) ever did. Raw query returned it, derived owed
set did not, delivery named only the directed arm. Two arms differing only in
to_agent is what makes it evidence rather than an absence.

And it carries a correction of my own record, in place and dated: I had filed
that this branch COULD NOT be exercised, because every broadcast in the log is
status='ready' and the protocol forbids writing an open one. Wrong in an
instructive way — the protocol forbids it as PRODUCTION TRAFFIC, which does not
forbid planting one as a labelled, self-closed control. Where a rule appears to
block a measurement, check whether it blocks the use or the test.
2026-08-26 23:38:41 +02:00
joakimp bfe9c5cd4f mailbox: join the owed set on hlc, not seq, before a second replica exists
deriveOwed decides "is this ask still owed?" by asking whether one of my own
terminal replies is LATER than the ask. It compared `seq` — this database's
arrival rowid. On a single hub that is global order, so it was correct; the
comment above it already said hlc was the durable key "once mesh_peers reports
actual peers". Making the switch now, while one replica means the two orderings
agree, costs nothing; making it later means changing the rule while two machines
already disagree about order.

The bug being pre-empted is specific: with a second replica the same event gets
a different `seq` in each database, because arrival order is not authorship
order. A reply authored after its ask can arrive first and take the lower seq;
the join then concludes "no later reply exists" and an already-answered ask
reappears as owed — permanently, on that machine.

isStrictlyAfter() prefers `hlc` when both events carry one and falls back to
`seq` otherwise (a server predating the field, or an un-backfilled row). hlc is
rendered fixed-width, <unix_ms:13 digits>-<counter:6 hex>-<replica_id>, so a
plain string comparison IS the causal comparison, with the replica id as final
tiebreak. Verified before writing the code that the field is actually on the
wire — event_list returns it per event (logstream.py:593) — because a fallback
that never fires would have made this a no-op dressed as a fix.

`created_at` stays rejected, and the reason is now written down where the
decision is: it is server-generated at second precision, so ties are routine,
and a tie can suppress an UNANSWERED ask. Both remaining failure modes are on
the noisy-but-visible side — an answered item resurfacing is annoying, an
unanswered ask going silent defeats the mailbox.

Tested: 18 cases against the extracted comparator — hlc later/earlier/equal,
same-ms counter ties in hex (0x10 vs 0x9, which is where a non-padded format
would break), cross-replica tiebreak, both mesh reorderings, every seq-fallback
path, and degenerate input (null seq, numeric hlc, empty strings, no keys) which
must never claim "answered". All pass. Real-data check on the positive-control
pair: seq 26/27 carry hlc ...3071857/...3574085, so the orderings agree today
and the switch is a no-op now and correct later.

The cursor keeps the opposite ordering ON PURPOSE — since_event_id is
local-arrival ordered so a tail consumer still sees late-arriving remote ops
whose hlc is older (hlc.py:19-21). RFC 003 §9.3 now says so explicitly, because
that asymmetry looks like a bug worth "fixing" and is not.
2026-08-26 23:30:09 +02:00
joakimp d4d8bb6109 docs: retention direction for the coordination log, and drop the operator's name from RFC 002 §5
RFC 003 §9.1 — retention is settled in direction (rotate old traffic out of
the way, logrotate-style, moved aside rather than destroyed) and the sketch
records the parts that are not obvious:

- Tier artifacts before events. Events are a few KB; artifacts are capped at
  4 MiB and stored in-row, so moving artifact CONTENT cold while keeping the
  kind/sha256/size/created_by stub reclaims nearly all the space and keeps the
  audit trail ("what was handed over, by whom, verified how") intact.
- If events rotate at all, the unit is the correlation_id THREAD whose latest
  event is terminal — never the row. Archiving an ask while leaving its reply
  (or the reverse) breaks the owed-set join, and both failure modes are bad:
  an ask that can never be cleared resurfaces as owed forever, or a reply is
  orphaned from what it answered.
- Never rotate an event that is still owed. Owed-ness is DERIVED at read time,
  so an unanswered ask is indistinguishable from a stale one except by that
  derivation — a purely time-based sweep would discard the live obligations of
  a machine that has merely been offline for a month, which is precisely the
  case this log exists to serve.
- Rotation invalidates held cursors: since_event_id RAISES on an unknown id
  (§7.5), so archiving an event a watcher holds as its resume point turns its
  next poll into an error. Either announce rotation ahead of live cursors, or
  teach the anchor lookup to fall back to created_at/hlc.
- Archive → verify (row counts, artifact sha256) → only then DELETE + VACUUM,
  with a --dry-run that reports in thread units.

RFC 002 §5: "Open decisions for ALC" → "Open decisions". RFC 001 never names
the operator anywhere; impersonal is the mature precedent and ALC was never a
real identifier in the first place (it is the AAAK spec's illustrative code for
"Alice", copy-forwarded into ~700 diary entries without verification).

Same edit also removes two device names and a hostname from §5.3, which is
host inventory and belongs in the private fleet repo, not a public one. The
mechanism it teaches — a palace in a Docker named volume dies on the next
container recreate, so census before flipping — is unchanged and is the part
that mattered.
2026-08-26 23:07:40 +02:00
joakimp e1cc7592a1 docs: write the RFC 003 that the code has been citing all along, plus an operator-facing fleet-memory guide
`logstream.py`'s module docstring is headed "Agent coordination event log for
MemPalace (RFC 003)" and enumerates five "Design constraints (RFC 003)".
Comments cite "RFC 003 phase 5", "RFC 003 suggested defaults" and "the first
RFC 003 dogfood". Every event/artifact tool description cites RFC 003. The
document has never existed — confirmed by searching this repo and the primary
host. So this is a retrospective spec: it transcribes what the implementation
already believes, and records what it does NOT do.

docs/rfc-003-coordination-log.md, verified line-by-line against mempalace
3.8.0 (every claim cites file.py:LINE, indexed in §10 for re-verification).
The parts that are not visible from the tool descriptions:

- No idempotency guard on event_append or put_artifact (§7.1). The replication
  path checks `id OR (origin_replica, origin_seq)` before applying; the client
  path checks nothing. So peer replay is safe and CLIENT RETRY IS NOT — a
  retried append forks a coordination thread into two ids. This makes the
  fleet's "a timeout is not a failure, verify before retrying" rule
  load-bearing rather than advisory.
- Coordination traffic is deliberately exempt from BOTH palace locks
  (§4) — _HTTP_LOCK_FREE_TOOLS and _PEER_WRITER_EXEMPT_TOOLS, each with its
  own rationale in-source. "One large mine blocks every client" is true of
  drawer writes and false of coordination writes.
- `mempalace sync` never touches the log, and no DELETE FROM events exists
  anywhere (§2, §9.1) — answering for the logstream a question RFC 001 §7.2
  left open, and making the log permanent and unbounded.
- from_agent is shape-validated and never authenticated; there is no read
  scoping at all (§6). The log authenticates the fleet, not the agent. That is
  simultaneously the security limitation and the only way to positive-control
  the mailbox.
- Three distinct orderings — seq (local arrival rowid), origin_seq (author's
  counter), hlc (fleet-wide, lexicographically sortable) — and origin_replica
  identifies the PALACE, not the writer, which is why from_agent/to_agent carry
  the whole distinction between machines (§3.2).
- GET /logstream/events does not exist and never did (§7.8). An earlier
  measurement saw it 404 and blamed the reverse proxy; that inference was right
  for /sync/* and /logstream/stream and wrong for this one. Corrected in
  extensions/pi/README.md §3 too, in place, dated.

docs/fleet-memory.md is the operator-facing companion the repo lacked
entirely: the front-door README had zero mentions of coordination, so the
channel was undiscoverable unless a message happened to arrive. It covers the
five stores and what each is for (drawers/wings/rooms, diaries, KG, palace
graph, coordination log), what a central palace buys a fleet — awareness,
non-repetition of expensive work, and retractions that travel — and a decision
flow for drawer vs KG vs event. Six mermaid diagrams, all rendered and
inspected as images, not merely parsed: the first pass produced a truncated
state label from an HTML entity and a self-loop that drew a meaningless dotted
lasso. Validation says "no syntax error"; only looking says "correct".

Also: a Documentation table in the top-level README so all of the above is
reachable from the front door.

Deliberately host-agnostic, per synlig-primary-runbook.md's precedent — no
device names, hostnames or operator names in either new document.
2026-08-26 22:55:55 +02:00