Commit Graph

29 Commits

Author SHA1 Message Date
Joakim Persson 975ab92943 feat(pi): let an ask declare dormancy, so waiting work stops nagging
A first-boot acceptance ask is planted deliberately unanswerable: it describes
work that becomes possible only when the device is next recreated, and it must
STAY owed until then, because closing it early to tidy the mailbox is exactly how
that work gets lost. deriveOwed could only see "directed, open, not answered, not
withdrawn", so such an ask was announced at every session start and every poll
for as long as it was correctly waiting -- measured at three days running on
emb-7kj4vr4g (2026-09-28 -> 2026-10-01), and across three consecutive releases
before that. The ask was right; announcing it was wrong, and the cost landed on
the human reading the window, who is the one reader that cannot filter it.

An ask may now declare the condition under which it is merely waiting:

  "dormant_unless": [
    { "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json",
      "field": "release_tag", "baseline": "v1.9.4" },
    { "kind": "file_mtime", "path": "/etc/hostname",
      "baseline": "2026-09-22T18:12:49Z" }
  ]

Dormant while EVERY condition still matches its baseline; live the moment ANY
differs -- which is the trigger those asks already stated in prose ("act when
EITHER differs"), now in a form the bridge can check. Two kinds, local files
only, no expression language, no shell, no network: a general evaluator in the
path that decides whether work is VISIBLE is a far worse trade than a clumsy
schema.

DORMANCY IS PROVEN, NEVER ASSUMED. Every unevaluable predicate announces the ask
instead of hiding it: missing file, unreadable file, unparseable baseline,
unknown kind, vanished field, relative path, more than eight conditions. The
dangerous failure here is not a spurious nag but work that disappears because a
predicate could not be evaluated -- indistinguishable from the ask being lost,
and not surfacing until a release needed it. An ask with no dormant_unless
behaves exactly as before, so this is backward compatible by construction.

Withheld from the ANNOUNCEMENT, never from the mailbox: deriveOwed now returns a
partition {owed, dormant} rather than a flat list, the wake-up injection lists
dormant asks once per session with ids, and a mid-session poll adds only a count
and only when the window is already open for something else. Dormant asks are
deliberately NOT added to the `surfaced` map, so one becomes announceable the
instant its baseline moves.

file_mtime compares WHOLE SECONDS in UTC. A filesystem mtime carries sub-second
residue (measured: /etc/hostname at .773761009) that a reported ISO baseline
never will, so comparing raw milliseconds would mark every such predicate
permanently "changed" -- silently disabling the feature while appearing to work.
The test records the residue for that reason.

scripts/test-dormancy.sh is this repo's first test: 22 cases, positive and
negative arms both, because a predicate that never fires makes the feature inert
and one that fires too eagerly hides real work. It copies the extension into a
temp tree with pi's typebox symlinked beside it (the copy is made per-run, so it
cannot drift like a vendored duplicate), owns its own fixtures rather than
reading /etc paths -- the first draft passed only on a pi-devbox container and
would have silently flipped to "not dormant" anywhere else -- and exits 2 for
INCONCLUSIVE rather than 0, since a test that skips quietly is the failure mode
it exists to catch. shellcheck clean.
2026-10-01 23:59:36 +02:00
joakimp 2167a1b033 fix(mailbox): pass order: "desc" on every cursor-less event_list call
Three of the five `mempalace_event_list` calls in the pi extension relied on
the server's default ordering: deriveOwed's `status: "open"` candidate query
(limit 50) and both windows in deriveClosed (from_agent limit 100, to_agent
limit 50). deriveOwed's other two calls already say `order: "desc"` and carry
the comment explaining why: on mempalace <= 3.9.0 the default is `asc`, so a
cursor-less `limit: N` returns the OLDEST N events, and once a device passes N
authored events its newest asks and the replies that close them fall outside
the join window. deriveClosed had exactly that latent truncation.

Why now: mempalace 3.10.0 (2026-09-15) flips the cursor-less default to
newest-first ("Logstream listings default to the newest events when no cursor
is given"). That change is server-side -- the extension talks to the fleet
hub over MEMPALACE_REMOTE_URL, so it lands when the hub upgrades, not when a
client does -- and it would have silently FIXED deriveClosed on 3.10.0 while
leaving it broken against any 3.9.0 hub. A mailbox verdict that depends on
which server version answers is the wrong shape. Saying `order` explicitly
makes both derive* functions read the same window on either version.

The selection in deriveClosed (isStrictlyAfter per correlation) is
order-independent, so this changes WHICH events are in the window, not how
the winner is picked. No behaviour change on a device with < 50 inbound and
< 100 authored events; the fleet hub is at 176 events total today, so the
window was about to matter.

Verified: file compiles under pi's own loader (jiti 2.7.0 from the installed
pi-coding-agent) before and after; `order:"desc"` literal count 3 -> 7, which
is the 3 added arguments plus 1 comment mention. `node --check` and a bare
`tsc` were tried first and both fail identically on HEAD (inline `type`
imports / no @types/node) -- checker faults, not this change.
2026-09-22 14:15:22 +02:00
joakimp 817b3a82b7 fix(feed): the mine's deadline never reached the transport; 60 s cut it off
The 2026-09 change raised MEMPALACE_FEED_MINE_TIMEOUT_MS to 300 000 and
raced it against client.callTool("mempalace_mine"). But callTool() had no
way to carry a deadline, so every call went out under the transport's
generic per-request timeout (MEMPALACE_MCP_TIMEOUT_MS, 60 000), which fired
first on every honest 60 s+ mine. Operators saw

    feed (tick) failed: mempalace remote request 'tools/call' failed:
    timed out after 60000ms

instead of the message the change had aimed at, and the 300 s was
unreachable. On stdio it was worse than noise: that transport kills the
server child on timeout, so the mine was actually aborted at 60 s.

callTool(name, args, { timeoutMs }) now passes a per-call deadline to both
transports; feedPalace() uses it for the mine. Plain calls keep the short
default — a query taking 60 s is still wedged. The Promise.race stays as the
liveness guard for a transport with its timeout disabled (0).

scripts/test-mcp-call-timeout.sh cuts RemoteMcpClient out of the shipped
file (as test-owed-withdrawal.sh does for the mailbox predicates), drives it
against a local JSON-RPC server that delays tools/call, and asserts: plain
call rejects at the generic deadline; the override outlives it; the override
is itself a deadline. Fails on the previous commit (2 of 6), passes here.
Vendored-copy delta noted in the header; protocol untouched, sync token
unchanged (check-mcp-client-sync.sh passes against pi-extensions).
2026-09-18 17:04:59 +02:00
Joakim Persson e68ee2071c docs(pi-ext): document what the mine deadline does, and what the message means
Two corrections to the operator-facing docs, both exposed by 309980b.

The env table listed the default as 30000, which is now wrong, and described the
var as capping "the mempalace_mine call". It never did: it bounds how long the
extension WAITS. The mine keeps running on the server. That exact misreading is
what made a 30s deadline look safe on a call measured at 30-60s.

Added a Debugging entry for "feed (tick) failed: mine timed out after ...ms",
because every operator on this fleet has seen it and it was documented nowhere.
It states the three things a reader needs: nothing was lost (the transcript is
staged before the mine, and mine --mode convos dedups by source_file and is
idempotent); do NOT retry harder from the client, because the palace is a single
writer and a blind retry turns one slow mine into a queue; and after 309980b the
message should not appear on a healthy fleet, so if it does it now MEANS
something -- a mine exceeding five minutes, i.e. look at palace size or another
writer holding the lock rather than raising the timeout again.
2026-09-10 20:58:12 +02:00
Joakim Persson 309980b62c fix(pi-ext): stop the feed tick from launching overlapping mines
"[mempalace ext] feed (tick) failed: mine timed out after 30000ms" was parked as
cosmetic on 2026-08-27. It is not cosmetic: the tight deadline was the trigger,
but the defect is a positive feedback loop that puts multiple writers on a
single-writer palace.

lastFeedAt was assigned only AFTER a successful await. The Promise.race abandons
our WAIT and cannot cancel the server's work, so on a mine that takes longer
than the deadline -- measured at 30-60s in normal operation, against a 30s
deadline -- the catch ran with lastFeedAt UNCHANGED. That left both guards in
the agent_settled handler open at once: the debounce test
(`Date.now() - lastFeedAt < feedDebounceMs`) passed because lastFeedAt was still
stale, and feedInFlight was already null because `run` had settled. Every
subsequent settled turn therefore launched another mine on top of the one still
running, each making the next slower and the next timeout likelier -- which is
why operators saw the message many times per session instead of at most once per
10-minute debounce window.

Fix, three lines:
- move `lastFeedAt = Date.now()` to before the await, so a timeout still starts
  the debounce clock. A timeout is not a "did not happen": the mine is running
  server-side and `mine --mode convos` dedups by source_file and is idempotent.
- raise MEMPALACE_FEED_MINE_TIMEOUT_MS from 30_000 to 300_000. The mine is the
  slowest thing this extension does yet carried the tightest deadline: 4x
  tighter than the prepare step before it (120_000) and 10x tighter than the
  init handshake (300_000), a fast call. All three were introduced together in
  29e660e and this one was never revisited. 300_000 matches the init timeout
  because liveness is the only job left for this deadline -- it cannot cancel
  the server's work, so it must sit far above the slowest honest completion.

Simulated both guards over 10 minutes of settled turns at 20s intervals with a
60s mine: BEFORE 16 mines launched, 15 of them overlapping an already-running
mine; AFTER 2 launched, 0 overlapping. With a mine that exceeds even the new
deadline (400s): BEFORE 16/15, AFTER still 2/0 -- the lastFeedAt move is what
actually fixes it, and it holds even when the timeout still fires. The raise
stops the spurious message; the move stops the pile-up.

NOT fixed here: the message still goes to process.stderr, which pi renders into
the TUI input field. That needs a pi-side channel or a log file, and is tracked
separately. After this change the message should be rare, and when it does
appear it means something real: a mine exceeding five minutes.
2026-09-10 20:48:33 +02:00
joakimp e2b060a940 feat(mailbox): let a requester withdraw its own ask, with an explicit marker
RFC 003 §3.3 clause 3 clears an ask only on "no event OF YOURS", and the skill
states the consequence outright: "there is nothing anyone can do about it from
the other end". That asymmetry is deliberate and mostly right — owed-ness is a
statement about the RECIPIENT's accountability. It is wrong in exactly one case:
the requester retracting its own ask.

MEASURED COST. pi@mbp-m1-2020 withdrew a v1.8.13 rollout ask to pi@tor-ms22 at
seq 119 — terminal `superseded`, same correlation_id, metadata.closes naming the
thread, body "DO NOT SPEND A MINUTE ON v1.8.13" — and recorded it as done. It had
no effect: seq 119's from_agent is mbp, so it could never satisfy a join that
only inspects tor-ms22's own events. tor-ms22's next wake-up, 8h later and on the
first boot of the image shipping this very derivation, still listed the ask as
owed, 41h old, for a release it never installed. The failure is invisible from
the sender's side, which is why it went unnoticed for 41h.

WHY AN EXPLICIT MARKER RATHER THAN "ANY TERMINAL EVENT FROM THE REQUESTER".
This file's standing rule is that every failure stays on the noisy-but-visible
side: a resurfacing item costs one turn of human correction, a suppressed
unanswered ask is silent and permanent. Under the naive rule a requester
appending `applied` for its own bookkeeping — on the correlation, addressed to
me, before I ever replied — would silently delete a real obligation. So release
must be STATED: metadata.withdraws (canonical) or metadata.closes (already this
fleet's de-facto marker), whose value must NAME the ask — its correlation_id or
its event id. Prose does not count.

Five further guards, each pinned by a mutation test: only the original requester;
addressed to this device exactly, never '*'; terminal status; strictly after the
ask by hlc; and joined by ack_of or correlation_id.

This does NOT break the fixed point in §9.2. That section rejects letting
terminal directed events into the owed set, because then every closure mints a
fresh obligation. That is about CANDIDATES; this adds a CLEARER. A withdrawal is
terminal, so it can never be a candidate, and the asserting shape (open) and the
clearing shape (terminal) stay disjoint.

Also fixed here, because the new rule depends on the same windowing: the `mine`
query used event_list's DEFAULT `asc` order with limit 100, i.e. the OLDEST 100
events this device ever wrote. Once a device passes 100 authored events its most
RECENT replies fall out of the join window and every ask it just answered
resurfaces as owed. Latent, not theoretical — tor-ms22 was at ~20. Both windows
are now anchored at the newest end with order: "desc".

Verification: scripts/test-owed-withdrawal.sh, 17 assertions over VERBATIM
fixtures from the real log (seq 112/119/120/122). It extracts the predicates from
mempalace.ts by brace matching and runs the SHIPPED text rather than a pasted
copy — this repo has already paid for a divergent second copy. Sensitivity
proven by four mutations: removing the marker requirement flips exactly the 3
marker assertions, and disabling the third-party / broadcast / ordering guards
each flip exactly their own. Control passes.
2026-09-09 08:56:56 +02:00
joakimp e45f6b4301 feat(mailbox): surface replies that CLOSE this device's own asks
The mailbox could report what this device OWES, and structurally nothing else.
deriveOwed() queries the log with status:"open", and a reply that closes an ask is
by definition not open — so a peer answering my delegation was invisible to it at
every poll, forever, not just late.

Measured 2026-09-07: emb-7kj4vr4g closed correlation v1813-client-rollout-emb with
a task.reply at status=applied. Nothing was announced. The feature was CORRECT by
its own definition ("owed" = "you must reply", and nothing was owed) and wrong by
the operator's, who asked why no notification arrived. The single most useful thing
a fleet can say to a human is "the thing you asked for is done" — and that was the
one category it could not say.

Independent corroboration that this was a real gap and not a preference: emb's own
reply ends "The amd64 narrowing is filed as a drawer as well, SINCE A TERMINAL
EVENT REACHES NO MAILBOX." A peer had already diagnosed the hole and was routing
around it by hand.

deriveClosed(): my task.requests (correlation + directed) joined against inbound
events with NO status filter, keeping the newest TERMINAL_STATUS reply per
correlation. Verified by computing the exact predicate over the two real events:
the new path yields evt_20260907T183143 (applied); the old status:"open" query
yields 0. Announced once each (never resurfaced — a finished ask is not a nag),
and the copy warns that a peer who did the work is the likeliest party to have
found your premise wrong. Here both premises were wrong: mine that emb was amd64,
and emb's that tor-ms22 was therefore the last candidate.

Deliberate asymmetry, documented at the call site: a broadcast is excluded from the
owed set (it owes nobody) but allowed to CLOSE, since a peer answering on my
correlation_id is news however widely it was addressed.
2026-09-07 21:42:58 +02:00
joakimp ecc2a9c574 mailbox notify: record that the terminal path is unverified through tmux
Measured on tor-ms22: the layering here is kitty -> tmux (on the host) ->
docker exec -> pi (in the container), with two clients attached to one tmux
session. Test OSC sequences written straight to pi's tty produced no
notification on the remote client; the local client is still unobserved. Most
likely tmux is dropping OSC types it does not implement — reaching the outer
terminal needs tmux's DCS passthrough (ESC P tmux; ... ESC \, inner ESC doubled)
plus allow-passthrough on, which is not the default and is not implemented here.

The point worth keeping is structural, not incidental: the container can see
NEITHER layer. KITTY_WINDOW_ID is absent because docker exec does not forward it,
and TMUX is absent because tmux runs one level further out on the host. So a
containerised client cannot detect the terminal it is speaking to or the
multiplexer it must speak through, and autodetection is not merely unreliable
here — it is blind. Explicit configuration is the only route, which is why the
previous commit added forced kitty/osc777 modes.

No behaviour or defaults changed: in-TUI notify remains the default and the
terminal path stays opt-in. Also flagged the unsettled routing question — a
pane's output reaches every attached client, so a work laptop and a home machine
would both ping from one arriving ask.
2026-08-27 12:54:52 +02:00
joakimp e91766286e mailbox notify: name the protocol, because docker exec hides the terminal
Auto-detection is useless in the deployment this ships in, measured on tor-ms22.
Terminal identity lives in env vars set by the emulator (KITTY_WINDOW_ID,
TERM_PROGRAM) and docker exec does not forward them — pi inside the container
sees only TERM=xterm-256color no matter what is rendering it. So the =desktop
detection can never see Kitty from in there and always falls through to OSC 777,
which Kitty does not implement. Result: on a containerised client the ping
silently does nothing, which is the worst available failure for a feature whose
entire purpose is to break a silence.

Adds MEMPALACE_MAILBOX_NOTIFY=kitty and =osc777 to force one protocol and skip
detection. =desktop keeps the notify.ts-style autodetect for the non-container
case, unset still means in-TUI only, 0/off still silent. The Windows toast branch
is dropped from the comment rather than the code path it never had here: it needs
powershell.exe on PATH, which is a WSL fact, not a Linux-container one.
2026-08-27 12:39:16 +02:00
joakimp a92c75d070 mailbox: say it is queued, and ping the human who is not looking (RFC 003 §7.11)
Two changes, neither touching the no-triggerTurn decision, which stands.

B — a delivery note in the message itself. The mid-session text explained how to
CLOSE an ask and never said when it would be SEEN, so the only reader who needed
that fact — the human watching an idle session — was the one not told. It now
says: this is a queued message, nothing woke the agent, any message starts the
turn that handles it, and the agent is not ignoring the ask, it is not running.
Costs nothing and changes no behaviour; it converts "why is it ignoring me?" into
"right, I nudge it". Deliberately NOT added to the wake-up injection, where a
turn is already starting and the note would be false.

A — a notification at poll time, because B only helps someone already looking and
the case that loses an ask is nobody looking. MEMPALACE_MAILBOX_NOTIFY: unset →
in-TUI ctx.ui.notify, the surface session_start already uses; =desktop →
additionally a terminal-native notification (Kitty OSC 99, else OSC 777), reusing
the detection the fleet's own notify.ts already proves in this harness; =0/off →
silent. The desktop path is how a ping escapes a container with no notify-send,
no DBus and no host access: the escape sequence is written to stdout and
interpreted by the terminal emulator on the human's own machine. Opt-in because
writing raw escapes is a behaviour change on a shared machine, not because it is
unreliable — say the word and the default flips.

Placement and firing conditions are deliberate: the notify call sits AFTER
sendMessage so a ping can never be the only thing that happened, and it fires
only when something is due — the same condition as delivery. A notification on
an empty poll would train its reader to ignore it, which is the failure this
whole feature exists to reverse. The wording names the nudge ("send any message
to handle") because "you have mail" without "press a key" reproduces exactly the
confusion B fixes.

Tested: 12 cases. Mode parsing (unset/empty/desktop/DESKTOP-with-space/0/off/1),
escape hygiene, and the OSC invariant that matters — a hostile title or body
containing ";" or a BEL cannot forge an OSC field or terminate the sequence
early, which is worth asserting because event bodies arrive from other machines.
ctx.hasUI is checked and the notify call is wrapped, since the UI can be gone by
the time an unawaited poll resolves. Syntax checked with node --strip-types.
2026-08-26 23:56:21 +02:00
joakimp bfe9c5cd4f mailbox: join the owed set on hlc, not seq, before a second replica exists
deriveOwed decides "is this ask still owed?" by asking whether one of my own
terminal replies is LATER than the ask. It compared `seq` — this database's
arrival rowid. On a single hub that is global order, so it was correct; the
comment above it already said hlc was the durable key "once mesh_peers reports
actual peers". Making the switch now, while one replica means the two orderings
agree, costs nothing; making it later means changing the rule while two machines
already disagree about order.

The bug being pre-empted is specific: with a second replica the same event gets
a different `seq` in each database, because arrival order is not authorship
order. A reply authored after its ask can arrive first and take the lower seq;
the join then concludes "no later reply exists" and an already-answered ask
reappears as owed — permanently, on that machine.

isStrictlyAfter() prefers `hlc` when both events carry one and falls back to
`seq` otherwise (a server predating the field, or an un-backfilled row). hlc is
rendered fixed-width, <unix_ms:13 digits>-<counter:6 hex>-<replica_id>, so a
plain string comparison IS the causal comparison, with the replica id as final
tiebreak. Verified before writing the code that the field is actually on the
wire — event_list returns it per event (logstream.py:593) — because a fallback
that never fires would have made this a no-op dressed as a fix.

`created_at` stays rejected, and the reason is now written down where the
decision is: it is server-generated at second precision, so ties are routine,
and a tie can suppress an UNANSWERED ask. Both remaining failure modes are on
the noisy-but-visible side — an answered item resurfacing is annoying, an
unanswered ask going silent defeats the mailbox.

Tested: 18 cases against the extracted comparator — hlc later/earlier/equal,
same-ms counter ties in hex (0x10 vs 0x9, which is where a non-padded format
would break), cross-replica tiebreak, both mesh reorderings, every seq-fallback
path, and degenerate input (null seq, numeric hlc, empty strings, no keys) which
must never claim "answered". All pass. Real-data check on the positive-control
pair: seq 26/27 carry hlc ...3071857/...3574085, so the orderings agree today
and the switch is a no-op now and correct later.

The cursor keeps the opposite ordering ON PURPOSE — since_event_id is
local-arrival ordered so a tail consumer still sees late-arriving remote ops
whose hlc is older (hlc.py:19-21). RFC 003 §9.3 now says so explicitly, because
that asymmetry looks like a bug worth "fixing" and is not.
2026-08-26 23:30:09 +02:00
joakimp d2764bf78e docs: take the Phase 1 exposure record private, leave a moved-note
463 lines with 64 mentions of specific hosts, in a public repo: the primary and
tunnel hosts by name, the registrar/DNS step, the tunnel resource wiring, shared
token custody, per-machine flip dates, and a palace lineage naming three work
machines. Now in the private fleet repo (fleet-ops 093fb65); this file becomes a
moved-note in the shape docs/synlig-primary-runbook.md already established.

Git history keeps the old text, so this limits future exposure rather than
undoing it.

Unlike the primary-host runbook, this file was MIXED — and the stub says so
instead of quietly implying the toolkit still documents HTTP exposure. §1.1-§1.3
(why one shared fleet token rather than per-device proxy users, and the one place
per-device identity does exist), §2 (the bind trap and the Host/Origin pin) and
§3.7 (a client is flipped by three env vars that travel as a set) are reusable
mechanism now published nowhere else. Named in the stub so the extraction is
tracked debt rather than a silent loss, and named in the private copy too so
whoever extracts it can delete the duplicate.

Inbound references fixed rather than left pointing at content that moved: two
§3.8 pointers in extensions/pi/README.md are replaced by the instruction they
were pointing at (the palace path reported by mempalace_status must be the
remote host's — a half-flipped client looks healthy while reporting a local
path), and the bind-trap reference is replaced by the Host/Origin sentence
itself, so the extension README no longer depends on the moved file. contrib
also loses a hostname and a seeding date it did not need to make its point.
2026-08-26 23:30:09 +02:00
joakimp e1cc7592a1 docs: write the RFC 003 that the code has been citing all along, plus an operator-facing fleet-memory guide
`logstream.py`'s module docstring is headed "Agent coordination event log for
MemPalace (RFC 003)" and enumerates five "Design constraints (RFC 003)".
Comments cite "RFC 003 phase 5", "RFC 003 suggested defaults" and "the first
RFC 003 dogfood". Every event/artifact tool description cites RFC 003. The
document has never existed — confirmed by searching this repo and the primary
host. So this is a retrospective spec: it transcribes what the implementation
already believes, and records what it does NOT do.

docs/rfc-003-coordination-log.md, verified line-by-line against mempalace
3.8.0 (every claim cites file.py:LINE, indexed in §10 for re-verification).
The parts that are not visible from the tool descriptions:

- No idempotency guard on event_append or put_artifact (§7.1). The replication
  path checks `id OR (origin_replica, origin_seq)` before applying; the client
  path checks nothing. So peer replay is safe and CLIENT RETRY IS NOT — a
  retried append forks a coordination thread into two ids. This makes the
  fleet's "a timeout is not a failure, verify before retrying" rule
  load-bearing rather than advisory.
- Coordination traffic is deliberately exempt from BOTH palace locks
  (§4) — _HTTP_LOCK_FREE_TOOLS and _PEER_WRITER_EXEMPT_TOOLS, each with its
  own rationale in-source. "One large mine blocks every client" is true of
  drawer writes and false of coordination writes.
- `mempalace sync` never touches the log, and no DELETE FROM events exists
  anywhere (§2, §9.1) — answering for the logstream a question RFC 001 §7.2
  left open, and making the log permanent and unbounded.
- from_agent is shape-validated and never authenticated; there is no read
  scoping at all (§6). The log authenticates the fleet, not the agent. That is
  simultaneously the security limitation and the only way to positive-control
  the mailbox.
- Three distinct orderings — seq (local arrival rowid), origin_seq (author's
  counter), hlc (fleet-wide, lexicographically sortable) — and origin_replica
  identifies the PALACE, not the writer, which is why from_agent/to_agent carry
  the whole distinction between machines (§3.2).
- GET /logstream/events does not exist and never did (§7.8). An earlier
  measurement saw it 404 and blamed the reverse proxy; that inference was right
  for /sync/* and /logstream/stream and wrong for this one. Corrected in
  extensions/pi/README.md §3 too, in place, dated.

docs/fleet-memory.md is the operator-facing companion the repo lacked
entirely: the front-door README had zero mentions of coordination, so the
channel was undiscoverable unless a message happened to arrive. It covers the
five stores and what each is for (drawers/wings/rooms, diaries, KG, palace
graph, coordination log), what a central palace buys a fleet — awareness,
non-repetition of expensive work, and retractions that travel — and a decision
flow for drawer vs KG vs event. Six mermaid diagrams, all rendered and
inspected as images, not merely parsed: the first pass produced a truncated
state label from an HTML entity and a self-loop that drew a meaningless dotted
lasso. Validation says "no syntax error"; only looking says "correct".

Also: a Documentation table in the top-level README so all of the above is
reachable from the front door.

Deliberately host-agnostic, per synlig-primary-runbook.md's precedent — no
device names, hostnames or operator names in either new document.
2026-08-26 22:55:55 +02:00
joakimp 5b8d78f946 pi bridge: the log gets read, not just written
Adds the auto-delivered mailbox. Until now the bridge stamped events on the way
out and never read the log, so a directed ask reached an agent only if that agent
happened to run event_list itself — which in practice meant ALC telling it to.
The channel had real cross-machine traffic since 2026-08-18 and no reader.

DELIVERY, two points, both fail-silent and both additive (223 insertions, 0
deletions; the feed's agent_settled handler is byte-identical):
- session start: one more sections.push() in the existing before_agent_start
  wake-up injection, beside mempalace_status and diary_read.
- mid-session: a second agent_settled handler, poll floored at
  MEMPALACE_MAILBOX_POLL_MS (default 300000 = 5 min), delivered with
  pi.sendMessage(deliverAs: "steer").

The cadence is chosen from measured arrival, not taste: 22 events since
2026-08-18, of which ELEVEN landed inside one 5h38m window today. Arrival is
bursty and correlates with the agent's own activity, because events arrive when
another machine is working the same thread — so agent_settled (activity-coupled)
is the right trigger and a wall-clock timer is the wrong one. Tightest observed
gap was 2m12s, so a 5-minute floor coalesces a burst into one message instead of
delivering five.

POLLING IS DECOUPLED FROM DELIVERY, which is the part that keeps this from
becoming noise: polling is cheap and frequent, but an item is only announced if
it has not been surfaced this session, or was surfaced more than
MEMPALACE_MAILBOX_RESURFACE_MS ago (default 1 h). Re-announcing the same ask
every five minutes would train the reader to ignore it — the exact failure the
status filter was introduced to prevent. The dedup map is in memory on purpose:
after a restart it may re-show something already seen, and that is the SAFE
failure direction (a resurfacing item is visible noise; a suppressed unanswered
ask is silent and permanent).

Owed-ness is DERIVED, never read off a field. event_ack appends and status is
written once, so a directed `open` matches the mailbox query forever, answered or
not — measured on this device, where the raw filter returned 3 asks of which 2
were already answered. A candidate is answered only when one of this device's own
events has a strictly higher seq, joins via metadata.ack_of or a shared
correlation_id, and carries a terminal status. The seq test is load-bearing:
without it one terminal reply suppresses every later ask on that correlation
forever, verified against the live thread where a seq-16 reply precedes the
seq-17 request it cannot have answered.

`*` broadcasts are excluded even though to_agent=<me> matches them, because the
protocol says a broadcast owes nobody a reply. Leaving them in would have made
this code contradict the skill documenting it, and would have made every machine
think it personally owed the same answer. It also gives "don't broadcast an ask"
teeth: broadcasting one now demonstrably reaches no owed set.

GATE: on when MEMPALACE_PI_DEVICE and MEMPALACE_REMOTE_URL are both set (the same
pair as the stamper — an unstamped client has no address to be reached at), off
with MEMPALACE_MAILBOX=0. Default-on is deliberate and ALC's call: the problem
being fixed is that nobody reads the inbox, and an opt-in fix for a nobody-does-it
problem only relocates the forgetting. Inert on a solitary palace.

README §2 rewritten in the same commit — it asserted "the bridge is write-only
today: there is no mailbox, no poll, no delivery", which this commit falsifies.
Shipping the code without the doc edit would have left a record asserting
something untrue in the very file documenting the fix for that class of defect.

VERIFIED: tsc 5.9.3 --strict, 0 errors, against the real pi types, with the
harness mutation-tested first (an injected error on a new line was caught, then
restored clean) because this repo has no package.json, no tsconfig and no tsc on
PATH — nothing in-repo will re-run this. Owed-set logic extracted verbatim from
the implementation and run against the live fixture: 3 candidates -> owed
[seq 21] only; demoting the seq-19 reply to seq 5 makes seq 17 owed again
(proves the ordering guard is live, not dead code); an open `*` broadcast and a
self-authored open ask are both excluded; null correlation_id does NOT join
itself (a plain === would have had null === null clear every uncorrelated ask).
NOT verified: no live palace call from the implementation, and the queue-vs-
interrupt semantics of "steer" are read from docs/extensions.md, not observed.
2026-08-26 17:23:24 +02:00
joakimp e70bef2b5d docs: the edge stamper and the coordination log the fleet actually uses
Two gaps, both found by using the thing rather than reading it.

PROVENANCE WAS SHIPPED UNDOCUMENTED. 553d8657 moved device attribution to the
edge — writer stamped on add_drawer/checkpoint/mine/event_append/artifact_put,
HOST:<device>| prefixed on diary entries — and this README, the file that
documents the extension, never mentioned it. So the only description of the
behaviour lived in the consumer skill, i.e. in the place the code does NOT live.
Now recorded next to Identity, with the two design points that keep getting
re-litigated: the diary marker is in the entry TEXT because diary_read returns
content only (an attribution nobody can see is not an attribution), and RFC 001
§7.3.2 ranks agent-side stamping worst — demonstrated when the agent that wrote
that skill instruction filed its own provenance drawer as added_by=checkpoint.
Includes the version gate that matters in practice: an older image satisfies both
env gates and still stamps nothing, because the extension is baked.

COORDINATION WAS UNDOCUMENTED ANYWHERE. The fleet has used the RFC 003 logstream
for real work since 2026-08-18 (patch handoff, review, a v1->v2 supersede) and no
file in this repo said so. Three facts belong here because they are mechanism:

- The stamper is what makes directed addressing possible. Where every machine is
  a thin MCP client of one shared palace, all clients report the SAME
  origin_replica, so from_agent/to_agent carry the entire distinction between
  machines. Measured 2026-08-26: mesh_peers returns peers: [] with a single
  replica id authoring every event from every machine. An unstamped client
  addressed as bare "pi" is unreachable.
- The bridge is WRITE-ONLY today: it stamps events going out and never reads the
  log. An event addressed to this machine by name reaches the agent only if the
  agent queries for it. Stated plainly because it is the current weak point, and
  it is exactly how a retraction addressed to pi@tor-ms22 sat unread while that
  agent rebuilt the thing it warned about.
- Live push is a DEPLOYMENT question. The palace implements SSE
  (GET /logstream/stream, text/event-stream in mcp_server.py) but a deployment
  may expose only /mcp: verified against mempalace.jordbo.se, where
  /logstream/events, /logstream/stream and /sync/peers all 404 while /mcp serves.
  Enabling it is a proxy route plus an auth decision, not an extension change.

Division of labour made explicit rather than implied: this file documents the
MECHANISM, the consumer skill is NORMATIVE for behaviour. Duplicating the ack
contract here would guarantee two copies that disagree.
2026-08-26 12:49:19 +02:00
pi 553d86570c provenance: stamp device+harness at the edge, not in the agent's head
RFC 001 §7.3.2 ranks "agent stamps provenance via a skill instruction" as the
❌ worst possible place — per-call boilerplate, forgettable, improvisable. It
was right, and we had shipped exactly that: the mempalace skill told the agent
to pass added_by="<harness>@<device>" by hand. Measured on the shared palace,
199 rows had reached it unresolvable, 10 of them filed by the very agent that
wrote the instruction, in a drawer about host provenance. The trigger was a
cross-host misattribution: a session on tor-ms22 read its own diary, could not
tell that the entries were written on EMB-7KJ4VR4G, and reported another
machine's verification as this one's.

Move the same convention into the ⚠️ edge row, where it is uniform and
unforgettable (§7.3.5):

* extensions/pi/mempalace.ts defaults the writer field on every tool that has
  one — added_by (add_drawer, checkpoint), agent (mine), from_agent
  (event_append), created_by (artifact_put) — from $MEMPALACE_PI_DEVICE. An
  explicit value always wins, so filing for another device stays possible. The
  allowlist is per tool, never blanket: 3.8.0's dispatcher hard-rejects
  undeclared args with -32602, so injecting added_by into diary_write or kg_add
  (which have no such property) would break the call outright.
* mine gets miner@<device> when the caller invokes it, but <harness>@<device>
  for the bridge's own transcript feed — bulk extraction is not agent-authored
  memory, and that keeps the pi/opencode/miner taxonomy honest.
* diary_write has no metadata slot at all, and the device must never go in
  agent_name (wing = f"wing_{agent_name}" would splinter the diary per host).
  So the entry TEXT carries an AAAK field, HOST:<device>|SESSION:… — which is
  also the only channel a READER sees: search projects a fixed key set and
  diary_read returns content, so no metadata fix, not even a
  server-authoritative one, would have prevented the misattribution.
* The wake-up block now states the device and warns that diary_read interleaves
  every machine's diary.
* R1: doubly gated on MEMPALACE_PI_DEVICE and MEMPALACE_REMOTE_URL, so a
  solitary devbox stamps nothing and behaves exactly as before — which is also
  the correct semantics per §7.3.3.

Version the reconciler that was living only on synlig (bin/ + contrib/systemd/),
add --dry-run, and teach it two new rules: diary_host_marker reads the HOST:
field, and sibling_chunk propagates a resolved origin across a drawer's chunks
(a text marker lands in chunk 0 only, so a 5-chunk diary entry would otherwise
stamp 1 and leave 4 blank).

--dry-run against the real palace before deploying earned its keep twice, and
scripts/test-device-stamp.sh pins both findings with the strings it found:
HOST: was ALREADY in use with a composite grammar
(HOST:emb-7kj4vr4g.f1d3c3f89e3e.v1.8.3.pi0.84.2) and for bare container ids, so
an unvalidated rule invented devices like "f1d3c3f89e3e.pi0.84.2"; and HOST:
also carries a different SENSE elsewhere (HOST:exec.via.ssh-controlmaster->…,
meaning where I was executing). Validating against the known-device set both
refuses those and recovers the composite entries correctly. A marker convention
inherits every prior meaning of its own name.

Deployed and verified on synlig: device 14,217 → 14,317, integrity ok,
idempotent on immediate re-run, no invented device values.

RFC updates: §7.3.1 corrected (the arg whitelist is a hard -32602 in 3.8.0, not
a silent drop; get_drawer DOES return metadata, search structurally cannot;
triples and logstream live in separate databases the stamper cannot reach),
§7.3.5 added (what is deployed, including the divergence from §7.3.4's opaque
origin_device — tor-ms22 vs tor-ms22-native is that cost already visible), and
Phase 4 now carries per-device tokens motivated FIRST by revocation, with the
finding that tokens are the cheap half: core holds one scalar auth_token and has
zero device concept, so authoritative stamping needs a component we own.
2026-08-25 22:26:46 +02:00
Joakim Persson a94eb7fdd0 docs: native pi has no .env to flip, and EMB-7KJ4VR4G has no native pi
Follow-up to ec436ed, closing the one item that commit left as "unverified".

Verified on the EMB-7KJ4VR4G host: native pi is not installed at all -- no `pi`
or `mempalace` on PATH, no ~/.config/pi/, no ~/.pi/agent/extensions/. That
host's ~/.mempalace exists solely to back the devbox container through the bind
mount (.devbox-owner holds 1000:1000). So the 2026-08-14 flip covers every pi
on that machine and there is no split-brain to fix. The general hazard stays
documented, because it is a per-machine question.

Also documents how a native install *would* be flipped, since the obvious guess
is wrong: pi loads no dotenv file and has no `env` block in settings.json, and
the extension reads process.env only. The launching shell is the sole hook, so
the vars must be exported from a shell rc (a GUI-launched pi may not read one),
and ~/.config/pi/.env is not sourced by anything automatically.
2026-08-14 23:00:46 +02:00
Joakim Persson ec436ed3ad docs: reconcile the RFC-001 docs with what is actually deployed
Audit of every doc touching the global-palace rollout against the running
fleet. Each correction below was verified against the filesystem or the host,
not against another doc:

- synlig-primary-runbook: the decommission `rm -rf ~/.mempalace` now carries a
  STOP block. That tree holds the fleet palace *and* the only copy of the
  bearer token every client authenticates with; the old "empty today" comment
  stopped being true when the palace was seeded on 2026-08-14. Adds an ordered
  safe decommission, and drops count-based join verification.

- phase-1-exposure-runbook: new S3.8, how to verify a flip actually took --
  the procedure that until now existed only in an untracked handover file.
  Three claims that fail independently (env var / curl / the palace-path
  discriminator) plus an explicit list of checks that produce FALSE POSITIVES:
  drawer counts (both sides were seeded from the same palace, and `status`
  counts chunks not drawers), write-then-read through the same transport, and
  the `mempalace` CLI -- which has no remote support at all, so post-flip it
  reads the dead local archive and reports success.

- rfc-001: status Draft -> Phases 0-1 implemented. Records that the join was a
  file-level copy, which SIDESTEPPED the S7.6 diary-dedup question rather than
  answering it -- so S7.6 remains a hard blocker for the second machine, which
  is the one that will actually exercise merge semantics.

- ARCHITECTURE, SKILL, contrib/README, extensions/pi/README all claimed pi
  feeds the palace automatically, unconditionally. That is gated on
  mempalace-toolkit >= 29e660e and every deployed image predates it, so the
  claim is currently false fleet-wide. Each site now states the gate plus a
  check that inspects the *deployed* file rather than repo HEAD.

- extensions/pi/README: plaintext http://mempalace.lan example -> https
  endpoint; the two transports are either/or (no dual-write, no local mirror);
  the bridge fails CLOSED, so "the agent has no mempalace_* tools" is the
  expected symptom of a server/token/DNS fault, not of a broken install.

- contrib/README: documents mempalace-serve.service, which this directory has
  shipped since day one without explaining it (linger, the load-bearing
  172.17.0.1 bind and why loopback is the unsafe-looking-safe option, the
  token path, and an uninstall warning).

- Fixes a pre-existing stray ```sh fence that was swallowing S3.2's heading and
  the token command into a code block.

Docs only; no behaviour change.
2026-08-14 22:57:22 +02:00
Joakim Persson 29e660e18f feeders: stage beside the palace, not in ~/.cache; document Phase 1 exposure
Staging default moves out of ~/.cache to <palace-root>/pi-stage (pi) and
<palace-root>/opencode-stage (opencode), resolved with mempalace's own
palace-path precedence ($MEMPALACE_PALACE_PATH -> $MEMPAL_PALACE_PATH ->
~/.mempalace/config.json -> ~/.mempalace/palace), then dirname.

Why: the convos miner keys dedup on the *staged* path, so a wiped stage plus a
sync scoped to include it prunes the drawers mined from those sources --
deleting memories, not a cache. Under ~/.cache that state was reachable by
anything treating a cache as disposable. Staging inside the palace makes the
coupling structural: the stage cannot be wiped without touching the palace
itself. Overrides ($MEMPALACE_PI_STAGE / $MEMPALACE_SESSION_STAGE, --stage) are
unchanged. Note the old default had never been created on any host, so this
closed a latent hazard, not a live one.

Measured, and the docs now claim only this much: sync prunes only within the
scope it is given -- wing-only, 1299 scanned / 1299 out of scope / 0 removed;
scoped at the palace root, 651 kept / 648 out of scope. The previous blanket
"sync prunes every drawer" wording overstated it, which is a liability: the next
reader disproves the overstatement and discards the real constraint with it.

Also in this change:
- cron log dir ~/.cache/mempalace-session -> ~/.cache/mempalace-logs. The stage
  left that namespace, so the old name now read as "the stage".
- AGENTS.md: the convos miner *does* check mtime (verified against upstream
  convo_miner.py); the previous "no mtime check" claim was wrong.
- smoke-test assertions use `mktemp -d` for --sessions-dir. One pointed at /tmp,
  which still held earlier synthetic transcripts, so a --dry-run exported a fake
  session into the real stage: --dry-run skips the mine, not the export.

docs/phase-1-exposure-runbook.md -- the newt/DNS/auth step that RFC 001 and the
synlig runbook leave open (runbook section 4, items 2 and 5). Port 8765 at /mcp,
newt targets 172.17.0.1, and the authentication is the single shared bearer
token (RFC 6.2, decided 2026-08-09) rather than per-device proxy users. The
latter cannot work today: mempalace validates exactly one token, and Pangolin's
SSO/PIN/password are browser-shaped while every client here is a headless
JSON-RPC POST -- enabling that protection breaks the clients it protects. The
per-device axis that *does* exist is the feeder's SSH key + per-device inbox.

New finding recorded there: a loopback bind does not merely 403 behind a tunnel
(already known, runbook 2.4) -- it also silently starts the server with no token
at all, because auto-minting is gated on the bind being non-loopback.

extensions/pi/README.md: the HTTP transport IS authenticated as of mempalace
3.6.0; the "sessionless and unauthenticated" note dated from the v1.3.0 era.
Closes the RFC section 8 Phase-0 hygiene item.
2026-08-12 17:04:01 +02:00
pi 96699f2a17 feat(pi-bridge): external MemPalace transport via MEMPALACE_REMOTE_URL
Let the pi<->mempalace bridge connect to a shared MemPalace over HTTP instead
of always spawning a local mempalace-mcp:
- Extract IMcpClient; rename McpClient -> StdioMcpClient (ctor command, arg-less start()).
- Add RemoteMcpClient (vendored from pi-extensions/mcp-loader.ts): streamable-HTTP
  with AbortController timeouts, protocolVersion pinned 2024-11-05, alive/ensureAlive.
  mempalace-mcp --transport http is sessionless JSON-RPC today; session/SSE/404
  branches retained for a future streamable-HTTP server.
- createClient() selects transport from MEMPALACE_REMOTE_URL; MEMPALACE_REMOTE_TOKEN
  -> Authorization: Bearer. Lifecycle automation (wake-up, /mempalace-diary) unchanged.
- scripts/check-mcp-client-sync.sh: drift guard vs canonical mcp-loader.ts.
- README: document local-vs-external transport.

Typechecks clean (strict); both transports smoke-tested against live mempalace-mcp.
2026-07-02 13:09:29 +02:00
pi e12b624cf7 feat(pi-ext): self-healing respawn + scoped init timeout for mempalace-mcp
A stall-kill (or any crash) of mempalace-mcp was a permanent latch:
available flipped off and stayed off until pi restart. Now the next tool
call transparently respawns the server and retries.

- ensureAlive(): bounded respawn with capped exponential backoff
  (MEMPALACE_MCP_MAX_RESPAWNS, default 2; MEMPALACE_MCP_RESPAWN_BACKOFF_MS,
  default 1000). Respawn budget resets on any successful JSON-RPC response,
  so a recovered server regains full patience while a persistently-broken
  one hits the cap and stays down (no hot-loop).
- Init timeout default raised 120000 -> 300000 (scoped to init only): a
  genuine virtiofs cold-open shouldn't be killed mid-progress only to
  respawn and re-pay the same cost. Per-call timeout stays 60000.
- Concurrency hardening: generation counter so a late exit from a killed
  old process can't tear down a fresh respawn; explicit healthy flag
  replaces racy proc!=null liveness check.
- README: document self-heal, new env vars, and why generous-init +
  bounded-respawn compose rather than overlap.
2026-06-26 00:22:21 +02:00
joakimp a3b8829991 feat(pi-ext): per-request timeout + stall-kill for mempalace-mcp
A wedged mempalace-mcp (classically an OrbStack virtiofs cold-open of a
large chroma.sqlite3 / HNSW load) left the awaiting JSON-RPC promise
pending forever, freezing the pi TUI uninterruptibly: ESC cancels the
LLM stream, not a pending tool execute().

The JSON-RPC client now arms a per-request timer. On expiry it rejects
the request AND kills the stalled child (SIGTERM->SIGKILL), so pi gets
an error instead of hanging; the extension flips available=false so
later calls fail fast (restart pi to retry). Per-REQUEST, not
per-process: the long-lived server only dies on a genuine stall.

Knobs: MEMPALACE_MCP_TIMEOUT_MS (default 60000),
MEMPALACE_MCP_INIT_TIMEOUT_MS (default 120000), 0 = disable.

This supersedes the planned standalone stdio-watchdog shim: the
extension already owns request/response correlation, so a separate
framing-reparsing shim is unnecessary.
2026-06-13 23:48:33 +02:00
joakimp ce09d25c97 Rename to @earendil-works/pi-coding-agent + earendil-works/pi URL
Pi moved to its new home at earendil-works on 2026-05-07
(https://pi.dev/news/2026/5/7/pi-has-a-new-home).

Sweep:
- extensions/pi/mempalace.ts: 'import type { ExtensionAPI } from
  "@mariozechner/pi-coding-agent"' -> @earendil-works/pi-coding-agent.
- README and extensions/pi/README: github.com/mariozechner/pi-coding-agent
  URL refs -> github.com/earendil-works/pi.
- install.sh: same URL substitution in the user-facing pointer line.

Brew install references (`brew install pi-coding-agent`) left as-is:
formula still works at 0.73.1, tap update tracked upstream at
earendil-works/pi#2755.
2026-05-09 17:56:46 +02:00
joakimp 16915f0e55 refactor: split pi-generic config into pi-toolkit repo
Parallel to the opencode-toolkit split earlier today. Pi's own config
(keybindings, shell env loader, settings template) moves to a new
sibling repo so opencode-devbox's mempalace opt-out can build slim
containers that include pi without dragging in chromadb + embedding
models (~300 MB).

What moved to pi-toolkit (https://gitea.jordbo.se/joakimp/pi-toolkit):
- extensions/pi/keybindings.json          (mosh/tmux newline fix)
- extensions/pi/pi-env.zsh                (sources ~/.config/pi/.env)
- extensions/pi/settings.example.json     (Bedrock template)
- install.sh::install_pi_keybindings      (symlink step)
- install.sh::install_pi_env_loader       (cp step + bash fallback)
- install.sh::check_pi_settings           (probe)
- install.sh::check_aws_env               (probe)

What stays here (this is the pi\u2194mempalace bridge, mempalace-side):
- extensions/pi/mempalace.ts              (the MCP extension)
- install.sh::install_pi_extension        (symlink step)
- NEW: install.sh::check_pi_toolkit       (probe: warns if pi is
                                           installed but pi-toolkit's
                                           artifacts are missing, with
                                           git-clone pointer)

install.sh shrank from 520 to 403 lines. Uninstall mirror correctly
does NOT touch pi-toolkit-owned files (explicit comment).

Docs updated:
- extensions/pi/README.md: rewritten as 'pi\u2194MemPalace MCP bridge',
  recipe becomes 'Deploying pi with mempalace' (pi-toolkit step 3,
  this repo step 5).
- AGENTS.md: Structure block + 'What install.sh does' section reflect
  the narrower scope and list the four things that moved out.
- README.md: repo-contents line + Setup section's deploy summary.

Verified on tor-ms22: full install\u2192uninstall\u2192reinstall lifecycle clean.
After mempalace-toolkit uninstall, pi-toolkit artifacts
(~/.pi/agent/keybindings.json, ~/.oh-my-zsh/custom/pi-env.zsh) remain
intact \u2014 correctly untouched. check_pi_toolkit probe fires green when
both exist.
2026-05-05 17:26:53 +02:00
joakimp 3d3a0fb125 docs(pi): add opencode-toolkit pointer in deploy recipe step 6
Cross-reference the newly-extracted opencode-toolkit repo, which owns
~/.config/opencode/.env loading. The recipe now distinguishes between
registering mempalace in opencode.json (still this repo's concern) and
ensuring opencode's own env loader is in place (opencode-toolkit's).
2026-05-05 17:14:32 +02:00
joakimp 118bd20fec feat(extensions/pi): ship pi-env.zsh shell loader
The loader that sources ~/.config/pi/.env into every shell was only
living in the myconfigs tor-ms22 backup \u2014 a fresh machine had nowhere
to get it from except copying by hand. Now canonical here.

- extensions/pi/pi-env.zsh: 20-line POSIX-compatible loader
  (set -a; source ~/.config/pi/.env; set +a). Works in bash and zsh.
- install.sh install_pi_env_loader:
  * oh-my-zsh detected (~/.oh-my-zsh/custom/ exists)
    \u2192 cp into that dir (NOT symlink \u2014 that dir is typically part of
      a dotfiles backup, and a symlink to mempalace-toolkit would
      break when restored on another host).
    \u2192 Idempotent: if target content matches repo, says 'already
      installed'. If it differs, leaves user edits alone and points
      at diff for manual reconcile.
  * No oh-my-zsh \u2192 prints source-this-line snippet for ~/.zshrc or
    ~/.bashrc (derived from $SHELL). Does NOT auto-edit rc files.
- install.sh uninstall: only removes the copy if content still matches
  repo. Local edits preserved.
- Docs:
  * extensions/pi/README.md Environment setup section rewritten with
    both install paths, step 5 of deploy recipe updated.
  * AGENTS.md Structure block lists pi-env.zsh.
  * Root README repo-contents line mentions it.

Verified on tor-ms22: install fresh \u2192 uninstall (content match \u2192 remove)
\u2192 reinstall \u2192 zsh -ic loads AWS vars correctly. Also tested bash fallback
path via HOME=/tmp/fake-home SHELL=/bin/bash \u2014 prints right .bashrc snippet.
2026-05-05 16:58:56 +02:00
joakimp 71c335148a docs(pi): full 'new machine' deploy recipe in extensions/pi/README
Consolidates the step-by-step recipe that's been living in diary entries
and session chat into the canonical pi bring-up doc. Covers:

  0. Prerequisites (zsh+oh-my-zsh, uv, tmux 3.2+, AWS creds)
  1. Dotfiles: myconfigs provision (tmux CSI-u, ~/.config/pi/.env, zsh loader)
  2. pi install (upstream brew/npm)
  3. mempalace CLI (uv tool install) + mempalace-toolkit install.sh
  4. pi settings bootstrap (start without --model, region prefix table)
  5. AWS env verification (git-crypt unlock gotcha)
  6. Opencode MCP registration pointer (if applicable)
  7. First run + wake-up injection smoke test
  + Verification checklist + uninstall

Root README.md adds a short summary box in the Setup section pointing at
the full recipe, so readers coming in from the front door find the pi
path immediately but the details stay with the files they install.

Covers: macOS + Linux. Works for homelab / work-macos / any myconfigs
profile that ships .config/pi/ + pi-env.zsh.
2026-05-05 15:20:47 +02:00
joakimp 854ae41f65 feat(extensions/pi): keybindings symlink + settings template + AWS/pi probes
Round out the pi bring-up story so a fresh machine can reach a working
pi+mempalace install with just `git clone && ./install.sh`:

- extensions/pi/keybindings.json: generic mosh/tmux newline fix
  (shift+enter, ctrl+j, alt+j). Safe on any machine — not
  region/account-specific. Symlinked into ~/.pi/agent/.
- extensions/pi/settings.example.json: template for `settings.json`
  so pi can start without --provider/--model. NOT symlinked — pi
  rewrites settings.json at runtime (lastChangelogVersion bumps),
  which would dirty the repo. Installer prints the cp + edit hint.
- install.sh: new install_pi_keybindings + uninstall mirror; new
  check_pi_settings probe (warns if settings.json missing); new
  check_aws_env probe (warns if AWS_PROFILE/AWS_REGION unset and
  settings.json selects amazon-bedrock). All new steps gated on
  pi being installed (~/.pi/agent/extensions/ exists).
- extensions/pi/README.md: documents keybindings rationale,
  settings bootstrap, and the recommended ~/.config/pi/.env +
  ~/.oh-my-zsh/custom/pi-env.zsh env layout (paired with the
  myconfigs commit 884e329 that split AWS vars out of
  ~/.config/opencode/.env).

Verified on tor-ms22: full install → uninstall → reinstall cycle,
new shell loads AWS_PROFILE/AWS_REGION from the new pi-env.zsh hook.

Works on macOS and Linux (plain ln -s, POSIX bash).
2026-05-05 13:59:20 +02:00
joakimp ef1d022fbc feat(extensions): version-control pi mempalace extension + install.sh symlink
The pi coding-agent extension at ~/.pi/agent/extensions/mempalace.ts was
living only on tor-ms22, including hand-edited fixes (Type.Unsafe
schema-passthrough for MCP tool parameters). One disk wipe away from
losing it, and no way to reproduce the install on a new machine.

- extensions/pi/mempalace.ts: canonical copy (matches tor-ms22 byte-for-byte)
- extensions/pi/README.md: what it does, the schema-passthrough gotcha,
  debugging knobs
- install.sh: new install_pi_extension step — gated on ~/.pi/agent/extensions/
  existing, backs up any real file in the way, idempotent re-runs, mirror
  block in uninstall. Works on macOS and Linux (plain ln -s, readlink -f).
- README.md: mention extensions/pi/ in the repo-contents list and in the
  Setup section

Verified on tor-ms22: install (backs up existing real file) → uninstall
(removes symlink) → reinstall (clean symlink). Re-runs are no-ops.
2026-05-05 13:42:47 +02:00