Commit Graph

11 Commits

Author SHA1 Message Date
joakimp e91766286e mailbox notify: name the protocol, because docker exec hides the terminal
Auto-detection is useless in the deployment this ships in, measured on tor-ms22.
Terminal identity lives in env vars set by the emulator (KITTY_WINDOW_ID,
TERM_PROGRAM) and docker exec does not forward them — pi inside the container
sees only TERM=xterm-256color no matter what is rendering it. So the =desktop
detection can never see Kitty from in there and always falls through to OSC 777,
which Kitty does not implement. Result: on a containerised client the ping
silently does nothing, which is the worst available failure for a feature whose
entire purpose is to break a silence.

Adds MEMPALACE_MAILBOX_NOTIFY=kitty and =osc777 to force one protocol and skip
detection. =desktop keeps the notify.ts-style autodetect for the non-container
case, unset still means in-TUI only, 0/off still silent. The Windows toast branch
is dropped from the comment rather than the code path it never had here: it needs
powershell.exe on PATH, which is a WSL fact, not a Linux-container one.
2026-08-27 12:39:16 +02:00
joakimp a92c75d070 mailbox: say it is queued, and ping the human who is not looking (RFC 003 §7.11)
Two changes, neither touching the no-triggerTurn decision, which stands.

B — a delivery note in the message itself. The mid-session text explained how to
CLOSE an ask and never said when it would be SEEN, so the only reader who needed
that fact — the human watching an idle session — was the one not told. It now
says: this is a queued message, nothing woke the agent, any message starts the
turn that handles it, and the agent is not ignoring the ask, it is not running.
Costs nothing and changes no behaviour; it converts "why is it ignoring me?" into
"right, I nudge it". Deliberately NOT added to the wake-up injection, where a
turn is already starting and the note would be false.

A — a notification at poll time, because B only helps someone already looking and
the case that loses an ask is nobody looking. MEMPALACE_MAILBOX_NOTIFY: unset →
in-TUI ctx.ui.notify, the surface session_start already uses; =desktop →
additionally a terminal-native notification (Kitty OSC 99, else OSC 777), reusing
the detection the fleet's own notify.ts already proves in this harness; =0/off →
silent. The desktop path is how a ping escapes a container with no notify-send,
no DBus and no host access: the escape sequence is written to stdout and
interpreted by the terminal emulator on the human's own machine. Opt-in because
writing raw escapes is a behaviour change on a shared machine, not because it is
unreliable — say the word and the default flips.

Placement and firing conditions are deliberate: the notify call sits AFTER
sendMessage so a ping can never be the only thing that happened, and it fires
only when something is due — the same condition as delivery. A notification on
an empty poll would train its reader to ignore it, which is the failure this
whole feature exists to reverse. The wording names the nudge ("send any message
to handle") because "you have mail" without "press a key" reproduces exactly the
confusion B fixes.

Tested: 12 cases. Mode parsing (unset/empty/desktop/DESKTOP-with-space/0/off/1),
escape hygiene, and the OSC invariant that matters — a hostile title or body
containing ";" or a BEL cannot forge an OSC field or terminate the sequence
early, which is worth asserting because event bodies arrive from other machines.
ctx.hasUI is checked and the notify call is wrapped, since the UI can be gone by
the time an unawaited poll resolves. Syntax checked with node --strip-types.
2026-08-26 23:56:21 +02:00
joakimp bfe9c5cd4f mailbox: join the owed set on hlc, not seq, before a second replica exists
deriveOwed decides "is this ask still owed?" by asking whether one of my own
terminal replies is LATER than the ask. It compared `seq` — this database's
arrival rowid. On a single hub that is global order, so it was correct; the
comment above it already said hlc was the durable key "once mesh_peers reports
actual peers". Making the switch now, while one replica means the two orderings
agree, costs nothing; making it later means changing the rule while two machines
already disagree about order.

The bug being pre-empted is specific: with a second replica the same event gets
a different `seq` in each database, because arrival order is not authorship
order. A reply authored after its ask can arrive first and take the lower seq;
the join then concludes "no later reply exists" and an already-answered ask
reappears as owed — permanently, on that machine.

isStrictlyAfter() prefers `hlc` when both events carry one and falls back to
`seq` otherwise (a server predating the field, or an un-backfilled row). hlc is
rendered fixed-width, <unix_ms:13 digits>-<counter:6 hex>-<replica_id>, so a
plain string comparison IS the causal comparison, with the replica id as final
tiebreak. Verified before writing the code that the field is actually on the
wire — event_list returns it per event (logstream.py:593) — because a fallback
that never fires would have made this a no-op dressed as a fix.

`created_at` stays rejected, and the reason is now written down where the
decision is: it is server-generated at second precision, so ties are routine,
and a tie can suppress an UNANSWERED ask. Both remaining failure modes are on
the noisy-but-visible side — an answered item resurfacing is annoying, an
unanswered ask going silent defeats the mailbox.

Tested: 18 cases against the extracted comparator — hlc later/earlier/equal,
same-ms counter ties in hex (0x10 vs 0x9, which is where a non-padded format
would break), cross-replica tiebreak, both mesh reorderings, every seq-fallback
path, and degenerate input (null seq, numeric hlc, empty strings, no keys) which
must never claim "answered". All pass. Real-data check on the positive-control
pair: seq 26/27 carry hlc ...3071857/...3574085, so the orderings agree today
and the switch is a no-op now and correct later.

The cursor keeps the opposite ordering ON PURPOSE — since_event_id is
local-arrival ordered so a tail consumer still sees late-arriving remote ops
whose hlc is older (hlc.py:19-21). RFC 003 §9.3 now says so explicitly, because
that asymmetry looks like a bug worth "fixing" and is not.
2026-08-26 23:30:09 +02:00
joakimp 5b8d78f946 pi bridge: the log gets read, not just written
Adds the auto-delivered mailbox. Until now the bridge stamped events on the way
out and never read the log, so a directed ask reached an agent only if that agent
happened to run event_list itself — which in practice meant ALC telling it to.
The channel had real cross-machine traffic since 2026-08-18 and no reader.

DELIVERY, two points, both fail-silent and both additive (223 insertions, 0
deletions; the feed's agent_settled handler is byte-identical):
- session start: one more sections.push() in the existing before_agent_start
  wake-up injection, beside mempalace_status and diary_read.
- mid-session: a second agent_settled handler, poll floored at
  MEMPALACE_MAILBOX_POLL_MS (default 300000 = 5 min), delivered with
  pi.sendMessage(deliverAs: "steer").

The cadence is chosen from measured arrival, not taste: 22 events since
2026-08-18, of which ELEVEN landed inside one 5h38m window today. Arrival is
bursty and correlates with the agent's own activity, because events arrive when
another machine is working the same thread — so agent_settled (activity-coupled)
is the right trigger and a wall-clock timer is the wrong one. Tightest observed
gap was 2m12s, so a 5-minute floor coalesces a burst into one message instead of
delivering five.

POLLING IS DECOUPLED FROM DELIVERY, which is the part that keeps this from
becoming noise: polling is cheap and frequent, but an item is only announced if
it has not been surfaced this session, or was surfaced more than
MEMPALACE_MAILBOX_RESURFACE_MS ago (default 1 h). Re-announcing the same ask
every five minutes would train the reader to ignore it — the exact failure the
status filter was introduced to prevent. The dedup map is in memory on purpose:
after a restart it may re-show something already seen, and that is the SAFE
failure direction (a resurfacing item is visible noise; a suppressed unanswered
ask is silent and permanent).

Owed-ness is DERIVED, never read off a field. event_ack appends and status is
written once, so a directed `open` matches the mailbox query forever, answered or
not — measured on this device, where the raw filter returned 3 asks of which 2
were already answered. A candidate is answered only when one of this device's own
events has a strictly higher seq, joins via metadata.ack_of or a shared
correlation_id, and carries a terminal status. The seq test is load-bearing:
without it one terminal reply suppresses every later ask on that correlation
forever, verified against the live thread where a seq-16 reply precedes the
seq-17 request it cannot have answered.

`*` broadcasts are excluded even though to_agent=<me> matches them, because the
protocol says a broadcast owes nobody a reply. Leaving them in would have made
this code contradict the skill documenting it, and would have made every machine
think it personally owed the same answer. It also gives "don't broadcast an ask"
teeth: broadcasting one now demonstrably reaches no owed set.

GATE: on when MEMPALACE_PI_DEVICE and MEMPALACE_REMOTE_URL are both set (the same
pair as the stamper — an unstamped client has no address to be reached at), off
with MEMPALACE_MAILBOX=0. Default-on is deliberate and ALC's call: the problem
being fixed is that nobody reads the inbox, and an opt-in fix for a nobody-does-it
problem only relocates the forgetting. Inert on a solitary palace.

README §2 rewritten in the same commit — it asserted "the bridge is write-only
today: there is no mailbox, no poll, no delivery", which this commit falsifies.
Shipping the code without the doc edit would have left a record asserting
something untrue in the very file documenting the fix for that class of defect.

VERIFIED: tsc 5.9.3 --strict, 0 errors, against the real pi types, with the
harness mutation-tested first (an injected error on a new line was caught, then
restored clean) because this repo has no package.json, no tsconfig and no tsc on
PATH — nothing in-repo will re-run this. Owed-set logic extracted verbatim from
the implementation and run against the live fixture: 3 candidates -> owed
[seq 21] only; demoting the seq-19 reply to seq 5 makes seq 17 owed again
(proves the ordering guard is live, not dead code); an open `*` broadcast and a
self-authored open ask are both excluded; null correlation_id does NOT join
itself (a plain === would have had null === null clear every uncorrelated ask).
NOT verified: no live palace call from the implementation, and the queue-vs-
interrupt semantics of "steer" are read from docs/extensions.md, not observed.
2026-08-26 17:23:24 +02:00
pi 553d86570c provenance: stamp device+harness at the edge, not in the agent's head
RFC 001 §7.3.2 ranks "agent stamps provenance via a skill instruction" as the
❌ worst possible place — per-call boilerplate, forgettable, improvisable. It
was right, and we had shipped exactly that: the mempalace skill told the agent
to pass added_by="<harness>@<device>" by hand. Measured on the shared palace,
199 rows had reached it unresolvable, 10 of them filed by the very agent that
wrote the instruction, in a drawer about host provenance. The trigger was a
cross-host misattribution: a session on tor-ms22 read its own diary, could not
tell that the entries were written on EMB-7KJ4VR4G, and reported another
machine's verification as this one's.

Move the same convention into the ⚠️ edge row, where it is uniform and
unforgettable (§7.3.5):

* extensions/pi/mempalace.ts defaults the writer field on every tool that has
  one — added_by (add_drawer, checkpoint), agent (mine), from_agent
  (event_append), created_by (artifact_put) — from $MEMPALACE_PI_DEVICE. An
  explicit value always wins, so filing for another device stays possible. The
  allowlist is per tool, never blanket: 3.8.0's dispatcher hard-rejects
  undeclared args with -32602, so injecting added_by into diary_write or kg_add
  (which have no such property) would break the call outright.
* mine gets miner@<device> when the caller invokes it, but <harness>@<device>
  for the bridge's own transcript feed — bulk extraction is not agent-authored
  memory, and that keeps the pi/opencode/miner taxonomy honest.
* diary_write has no metadata slot at all, and the device must never go in
  agent_name (wing = f"wing_{agent_name}" would splinter the diary per host).
  So the entry TEXT carries an AAAK field, HOST:<device>|SESSION:… — which is
  also the only channel a READER sees: search projects a fixed key set and
  diary_read returns content, so no metadata fix, not even a
  server-authoritative one, would have prevented the misattribution.
* The wake-up block now states the device and warns that diary_read interleaves
  every machine's diary.
* R1: doubly gated on MEMPALACE_PI_DEVICE and MEMPALACE_REMOTE_URL, so a
  solitary devbox stamps nothing and behaves exactly as before — which is also
  the correct semantics per §7.3.3.

Version the reconciler that was living only on synlig (bin/ + contrib/systemd/),
add --dry-run, and teach it two new rules: diary_host_marker reads the HOST:
field, and sibling_chunk propagates a resolved origin across a drawer's chunks
(a text marker lands in chunk 0 only, so a 5-chunk diary entry would otherwise
stamp 1 and leave 4 blank).

--dry-run against the real palace before deploying earned its keep twice, and
scripts/test-device-stamp.sh pins both findings with the strings it found:
HOST: was ALREADY in use with a composite grammar
(HOST:emb-7kj4vr4g.f1d3c3f89e3e.v1.8.3.pi0.84.2) and for bare container ids, so
an unvalidated rule invented devices like "f1d3c3f89e3e.pi0.84.2"; and HOST:
also carries a different SENSE elsewhere (HOST:exec.via.ssh-controlmaster->…,
meaning where I was executing). Validating against the known-device set both
refuses those and recovers the composite entries correctly. A marker convention
inherits every prior meaning of its own name.

Deployed and verified on synlig: device 14,217 → 14,317, integrity ok,
idempotent on immediate re-run, no invented device values.

RFC updates: §7.3.1 corrected (the arg whitelist is a hard -32602 in 3.8.0, not
a silent drop; get_drawer DOES return metadata, search structurally cannot;
triples and logstream live in separate databases the stamper cannot reach),
§7.3.5 added (what is deployed, including the divergence from §7.3.4's opaque
origin_device — tor-ms22 vs tor-ms22-native is that cost already visible), and
Phase 4 now carries per-device tokens motivated FIRST by revocation, with the
finding that tokens are the cheap half: core holds one scalar auth_token and has
zero device concept, so authoritative stamping needs a component we own.
2026-08-25 22:26:46 +02:00
Joakim Persson 29e660e18f feeders: stage beside the palace, not in ~/.cache; document Phase 1 exposure
Staging default moves out of ~/.cache to <palace-root>/pi-stage (pi) and
<palace-root>/opencode-stage (opencode), resolved with mempalace's own
palace-path precedence ($MEMPALACE_PALACE_PATH -> $MEMPAL_PALACE_PATH ->
~/.mempalace/config.json -> ~/.mempalace/palace), then dirname.

Why: the convos miner keys dedup on the *staged* path, so a wiped stage plus a
sync scoped to include it prunes the drawers mined from those sources --
deleting memories, not a cache. Under ~/.cache that state was reachable by
anything treating a cache as disposable. Staging inside the palace makes the
coupling structural: the stage cannot be wiped without touching the palace
itself. Overrides ($MEMPALACE_PI_STAGE / $MEMPALACE_SESSION_STAGE, --stage) are
unchanged. Note the old default had never been created on any host, so this
closed a latent hazard, not a live one.

Measured, and the docs now claim only this much: sync prunes only within the
scope it is given -- wing-only, 1299 scanned / 1299 out of scope / 0 removed;
scoped at the palace root, 651 kept / 648 out of scope. The previous blanket
"sync prunes every drawer" wording overstated it, which is a liability: the next
reader disproves the overstatement and discards the real constraint with it.

Also in this change:
- cron log dir ~/.cache/mempalace-session -> ~/.cache/mempalace-logs. The stage
  left that namespace, so the old name now read as "the stage".
- AGENTS.md: the convos miner *does* check mtime (verified against upstream
  convo_miner.py); the previous "no mtime check" claim was wrong.
- smoke-test assertions use `mktemp -d` for --sessions-dir. One pointed at /tmp,
  which still held earlier synthetic transcripts, so a --dry-run exported a fake
  session into the real stage: --dry-run skips the mine, not the export.

docs/phase-1-exposure-runbook.md -- the newt/DNS/auth step that RFC 001 and the
synlig runbook leave open (runbook section 4, items 2 and 5). Port 8765 at /mcp,
newt targets 172.17.0.1, and the authentication is the single shared bearer
token (RFC 6.2, decided 2026-08-09) rather than per-device proxy users. The
latter cannot work today: mempalace validates exactly one token, and Pangolin's
SSO/PIN/password are browser-shaped while every client here is a headless
JSON-RPC POST -- enabling that protection breaks the clients it protects. The
per-device axis that *does* exist is the feeder's SSH key + per-device inbox.

New finding recorded there: a loopback bind does not merely 403 behind a tunnel
(already known, runbook 2.4) -- it also silently starts the server with no token
at all, because auto-minting is gated on the bind being non-loopback.

extensions/pi/README.md: the HTTP transport IS authenticated as of mempalace
3.6.0; the "sessionless and unauthenticated" note dated from the v1.3.0 era.
Closes the RFC section 8 Phase-0 hygiene item.
2026-08-12 17:04:01 +02:00
pi 96699f2a17 feat(pi-bridge): external MemPalace transport via MEMPALACE_REMOTE_URL
Let the pi<->mempalace bridge connect to a shared MemPalace over HTTP instead
of always spawning a local mempalace-mcp:
- Extract IMcpClient; rename McpClient -> StdioMcpClient (ctor command, arg-less start()).
- Add RemoteMcpClient (vendored from pi-extensions/mcp-loader.ts): streamable-HTTP
  with AbortController timeouts, protocolVersion pinned 2024-11-05, alive/ensureAlive.
  mempalace-mcp --transport http is sessionless JSON-RPC today; session/SSE/404
  branches retained for a future streamable-HTTP server.
- createClient() selects transport from MEMPALACE_REMOTE_URL; MEMPALACE_REMOTE_TOKEN
  -> Authorization: Bearer. Lifecycle automation (wake-up, /mempalace-diary) unchanged.
- scripts/check-mcp-client-sync.sh: drift guard vs canonical mcp-loader.ts.
- README: document local-vs-external transport.

Typechecks clean (strict); both transports smoke-tested against live mempalace-mcp.
2026-07-02 13:09:29 +02:00
pi e12b624cf7 feat(pi-ext): self-healing respawn + scoped init timeout for mempalace-mcp
A stall-kill (or any crash) of mempalace-mcp was a permanent latch:
available flipped off and stayed off until pi restart. Now the next tool
call transparently respawns the server and retries.

- ensureAlive(): bounded respawn with capped exponential backoff
  (MEMPALACE_MCP_MAX_RESPAWNS, default 2; MEMPALACE_MCP_RESPAWN_BACKOFF_MS,
  default 1000). Respawn budget resets on any successful JSON-RPC response,
  so a recovered server regains full patience while a persistently-broken
  one hits the cap and stays down (no hot-loop).
- Init timeout default raised 120000 -> 300000 (scoped to init only): a
  genuine virtiofs cold-open shouldn't be killed mid-progress only to
  respawn and re-pay the same cost. Per-call timeout stays 60000.
- Concurrency hardening: generation counter so a late exit from a killed
  old process can't tear down a fresh respawn; explicit healthy flag
  replaces racy proc!=null liveness check.
- README: document self-heal, new env vars, and why generous-init +
  bounded-respawn compose rather than overlap.
2026-06-26 00:22:21 +02:00
joakimp a3b8829991 feat(pi-ext): per-request timeout + stall-kill for mempalace-mcp
A wedged mempalace-mcp (classically an OrbStack virtiofs cold-open of a
large chroma.sqlite3 / HNSW load) left the awaiting JSON-RPC promise
pending forever, freezing the pi TUI uninterruptibly: ESC cancels the
LLM stream, not a pending tool execute().

The JSON-RPC client now arms a per-request timer. On expiry it rejects
the request AND kills the stalled child (SIGTERM->SIGKILL), so pi gets
an error instead of hanging; the extension flips available=false so
later calls fail fast (restart pi to retry). Per-REQUEST, not
per-process: the long-lived server only dies on a genuine stall.

Knobs: MEMPALACE_MCP_TIMEOUT_MS (default 60000),
MEMPALACE_MCP_INIT_TIMEOUT_MS (default 120000), 0 = disable.

This supersedes the planned standalone stdio-watchdog shim: the
extension already owns request/response correlation, so a separate
framing-reparsing shim is unnecessary.
2026-06-13 23:48:33 +02:00
joakimp ce09d25c97 Rename to @earendil-works/pi-coding-agent + earendil-works/pi URL
Pi moved to its new home at earendil-works on 2026-05-07
(https://pi.dev/news/2026/5/7/pi-has-a-new-home).

Sweep:
- extensions/pi/mempalace.ts: 'import type { ExtensionAPI } from
  "@mariozechner/pi-coding-agent"' -> @earendil-works/pi-coding-agent.
- README and extensions/pi/README: github.com/mariozechner/pi-coding-agent
  URL refs -> github.com/earendil-works/pi.
- install.sh: same URL substitution in the user-facing pointer line.

Brew install references (`brew install pi-coding-agent`) left as-is:
formula still works at 0.73.1, tap update tracked upstream at
earendil-works/pi#2755.
2026-05-09 17:56:46 +02:00
joakimp ef1d022fbc feat(extensions): version-control pi mempalace extension + install.sh symlink
The pi coding-agent extension at ~/.pi/agent/extensions/mempalace.ts was
living only on tor-ms22, including hand-edited fixes (Type.Unsafe
schema-passthrough for MCP tool parameters). One disk wipe away from
losing it, and no way to reproduce the install on a new machine.

- extensions/pi/mempalace.ts: canonical copy (matches tor-ms22 byte-for-byte)
- extensions/pi/README.md: what it does, the schema-passthrough gotcha,
  debugging knobs
- install.sh: new install_pi_extension step — gated on ~/.pi/agent/extensions/
  existing, backs up any real file in the way, idempotent re-runs, mirror
  block in uninstall. Works on macOS and Linux (plain ln -s, readlink -f).
- README.md: mention extensions/pi/ in the repo-contents list and in the
  Setup section

Verified on tor-ms22: install (backs up existing real file) → uninstall
(removes symlink) → reinstall (clean symlink). Re-runs are no-ops.
2026-05-05 13:42:47 +02:00