Commit Graph

66 Commits

Author SHA1 Message Date
joakimp d2764bf78e docs: take the Phase 1 exposure record private, leave a moved-note
463 lines with 64 mentions of specific hosts, in a public repo: the primary and
tunnel hosts by name, the registrar/DNS step, the tunnel resource wiring, shared
token custody, per-machine flip dates, and a palace lineage naming three work
machines. Now in the private fleet repo (fleet-ops 093fb65); this file becomes a
moved-note in the shape docs/synlig-primary-runbook.md already established.

Git history keeps the old text, so this limits future exposure rather than
undoing it.

Unlike the primary-host runbook, this file was MIXED — and the stub says so
instead of quietly implying the toolkit still documents HTTP exposure. §1.1-§1.3
(why one shared fleet token rather than per-device proxy users, and the one place
per-device identity does exist), §2 (the bind trap and the Host/Origin pin) and
§3.7 (a client is flipped by three env vars that travel as a set) are reusable
mechanism now published nowhere else. Named in the stub so the extraction is
tracked debt rather than a silent loss, and named in the private copy too so
whoever extracts it can delete the duplicate.

Inbound references fixed rather than left pointing at content that moved: two
§3.8 pointers in extensions/pi/README.md are replaced by the instruction they
were pointing at (the palace path reported by mempalace_status must be the
remote host's — a half-flipped client looks healthy while reporting a local
path), and the bind-trap reference is replaced by the Host/Origin sentence
itself, so the extension README no longer depends on the moved file. contrib
also loses a hostname and a seeding date it did not need to make its point.
2026-08-26 23:30:09 +02:00
joakimp d4d8bb6109 docs: retention direction for the coordination log, and drop the operator's name from RFC 002 §5
RFC 003 §9.1 — retention is settled in direction (rotate old traffic out of
the way, logrotate-style, moved aside rather than destroyed) and the sketch
records the parts that are not obvious:

- Tier artifacts before events. Events are a few KB; artifacts are capped at
  4 MiB and stored in-row, so moving artifact CONTENT cold while keeping the
  kind/sha256/size/created_by stub reclaims nearly all the space and keeps the
  audit trail ("what was handed over, by whom, verified how") intact.
- If events rotate at all, the unit is the correlation_id THREAD whose latest
  event is terminal — never the row. Archiving an ask while leaving its reply
  (or the reverse) breaks the owed-set join, and both failure modes are bad:
  an ask that can never be cleared resurfaces as owed forever, or a reply is
  orphaned from what it answered.
- Never rotate an event that is still owed. Owed-ness is DERIVED at read time,
  so an unanswered ask is indistinguishable from a stale one except by that
  derivation — a purely time-based sweep would discard the live obligations of
  a machine that has merely been offline for a month, which is precisely the
  case this log exists to serve.
- Rotation invalidates held cursors: since_event_id RAISES on an unknown id
  (§7.5), so archiving an event a watcher holds as its resume point turns its
  next poll into an error. Either announce rotation ahead of live cursors, or
  teach the anchor lookup to fall back to created_at/hlc.
- Archive → verify (row counts, artifact sha256) → only then DELETE + VACUUM,
  with a --dry-run that reports in thread units.

RFC 002 §5: "Open decisions for ALC" → "Open decisions". RFC 001 never names
the operator anywhere; impersonal is the mature precedent and ALC was never a
real identifier in the first place (it is the AAAK spec's illustrative code for
"Alice", copy-forwarded into ~700 diary entries without verification).

Same edit also removes two device names and a hostname from §5.3, which is
host inventory and belongs in the private fleet repo, not a public one. The
mechanism it teaches — a palace in a Docker named volume dies on the next
container recreate, so census before flipping — is unchanged and is the part
that mattered.
2026-08-26 23:07:40 +02:00
joakimp e1cc7592a1 docs: write the RFC 003 that the code has been citing all along, plus an operator-facing fleet-memory guide
`logstream.py`'s module docstring is headed "Agent coordination event log for
MemPalace (RFC 003)" and enumerates five "Design constraints (RFC 003)".
Comments cite "RFC 003 phase 5", "RFC 003 suggested defaults" and "the first
RFC 003 dogfood". Every event/artifact tool description cites RFC 003. The
document has never existed — confirmed by searching this repo and the primary
host. So this is a retrospective spec: it transcribes what the implementation
already believes, and records what it does NOT do.

docs/rfc-003-coordination-log.md, verified line-by-line against mempalace
3.8.0 (every claim cites file.py:LINE, indexed in §10 for re-verification).
The parts that are not visible from the tool descriptions:

- No idempotency guard on event_append or put_artifact (§7.1). The replication
  path checks `id OR (origin_replica, origin_seq)` before applying; the client
  path checks nothing. So peer replay is safe and CLIENT RETRY IS NOT — a
  retried append forks a coordination thread into two ids. This makes the
  fleet's "a timeout is not a failure, verify before retrying" rule
  load-bearing rather than advisory.
- Coordination traffic is deliberately exempt from BOTH palace locks
  (§4) — _HTTP_LOCK_FREE_TOOLS and _PEER_WRITER_EXEMPT_TOOLS, each with its
  own rationale in-source. "One large mine blocks every client" is true of
  drawer writes and false of coordination writes.
- `mempalace sync` never touches the log, and no DELETE FROM events exists
  anywhere (§2, §9.1) — answering for the logstream a question RFC 001 §7.2
  left open, and making the log permanent and unbounded.
- from_agent is shape-validated and never authenticated; there is no read
  scoping at all (§6). The log authenticates the fleet, not the agent. That is
  simultaneously the security limitation and the only way to positive-control
  the mailbox.
- Three distinct orderings — seq (local arrival rowid), origin_seq (author's
  counter), hlc (fleet-wide, lexicographically sortable) — and origin_replica
  identifies the PALACE, not the writer, which is why from_agent/to_agent carry
  the whole distinction between machines (§3.2).
- GET /logstream/events does not exist and never did (§7.8). An earlier
  measurement saw it 404 and blamed the reverse proxy; that inference was right
  for /sync/* and /logstream/stream and wrong for this one. Corrected in
  extensions/pi/README.md §3 too, in place, dated.

docs/fleet-memory.md is the operator-facing companion the repo lacked
entirely: the front-door README had zero mentions of coordination, so the
channel was undiscoverable unless a message happened to arrive. It covers the
five stores and what each is for (drawers/wings/rooms, diaries, KG, palace
graph, coordination log), what a central palace buys a fleet — awareness,
non-repetition of expensive work, and retractions that travel — and a decision
flow for drawer vs KG vs event. Six mermaid diagrams, all rendered and
inspected as images, not merely parsed: the first pass produced a truncated
state label from an HTML entity and a self-loop that drew a meaningless dotted
lasso. Validation says "no syntax error"; only looking says "correct".

Also: a Documentation table in the top-level README so all of the above is
reachable from the front door.

Deliberately host-agnostic, per synlig-primary-runbook.md's precedent — no
device names, hostnames or operator names in either new document.
2026-08-26 22:55:55 +02:00
joakimp 5b8d78f946 pi bridge: the log gets read, not just written
Adds the auto-delivered mailbox. Until now the bridge stamped events on the way
out and never read the log, so a directed ask reached an agent only if that agent
happened to run event_list itself — which in practice meant ALC telling it to.
The channel had real cross-machine traffic since 2026-08-18 and no reader.

DELIVERY, two points, both fail-silent and both additive (223 insertions, 0
deletions; the feed's agent_settled handler is byte-identical):
- session start: one more sections.push() in the existing before_agent_start
  wake-up injection, beside mempalace_status and diary_read.
- mid-session: a second agent_settled handler, poll floored at
  MEMPALACE_MAILBOX_POLL_MS (default 300000 = 5 min), delivered with
  pi.sendMessage(deliverAs: "steer").

The cadence is chosen from measured arrival, not taste: 22 events since
2026-08-18, of which ELEVEN landed inside one 5h38m window today. Arrival is
bursty and correlates with the agent's own activity, because events arrive when
another machine is working the same thread — so agent_settled (activity-coupled)
is the right trigger and a wall-clock timer is the wrong one. Tightest observed
gap was 2m12s, so a 5-minute floor coalesces a burst into one message instead of
delivering five.

POLLING IS DECOUPLED FROM DELIVERY, which is the part that keeps this from
becoming noise: polling is cheap and frequent, but an item is only announced if
it has not been surfaced this session, or was surfaced more than
MEMPALACE_MAILBOX_RESURFACE_MS ago (default 1 h). Re-announcing the same ask
every five minutes would train the reader to ignore it — the exact failure the
status filter was introduced to prevent. The dedup map is in memory on purpose:
after a restart it may re-show something already seen, and that is the SAFE
failure direction (a resurfacing item is visible noise; a suppressed unanswered
ask is silent and permanent).

Owed-ness is DERIVED, never read off a field. event_ack appends and status is
written once, so a directed `open` matches the mailbox query forever, answered or
not — measured on this device, where the raw filter returned 3 asks of which 2
were already answered. A candidate is answered only when one of this device's own
events has a strictly higher seq, joins via metadata.ack_of or a shared
correlation_id, and carries a terminal status. The seq test is load-bearing:
without it one terminal reply suppresses every later ask on that correlation
forever, verified against the live thread where a seq-16 reply precedes the
seq-17 request it cannot have answered.

`*` broadcasts are excluded even though to_agent=<me> matches them, because the
protocol says a broadcast owes nobody a reply. Leaving them in would have made
this code contradict the skill documenting it, and would have made every machine
think it personally owed the same answer. It also gives "don't broadcast an ask"
teeth: broadcasting one now demonstrably reaches no owed set.

GATE: on when MEMPALACE_PI_DEVICE and MEMPALACE_REMOTE_URL are both set (the same
pair as the stamper — an unstamped client has no address to be reached at), off
with MEMPALACE_MAILBOX=0. Default-on is deliberate and ALC's call: the problem
being fixed is that nobody reads the inbox, and an opt-in fix for a nobody-does-it
problem only relocates the forgetting. Inert on a solitary palace.

README §2 rewritten in the same commit — it asserted "the bridge is write-only
today: there is no mailbox, no poll, no delivery", which this commit falsifies.
Shipping the code without the doc edit would have left a record asserting
something untrue in the very file documenting the fix for that class of defect.

VERIFIED: tsc 5.9.3 --strict, 0 errors, against the real pi types, with the
harness mutation-tested first (an injected error on a new line was caught, then
restored clean) because this repo has no package.json, no tsconfig and no tsc on
PATH — nothing in-repo will re-run this. Owed-set logic extracted verbatim from
the implementation and run against the live fixture: 3 candidates -> owed
[seq 21] only; demoting the seq-19 reply to seq 5 makes seq 17 owed again
(proves the ordering guard is live, not dead code); an open `*` broadcast and a
self-authored open ask are both excluded; null correlation_id does NOT join
itself (a plain === would have had null === null clear every uncorrelated ask).
NOT verified: no live palace call from the implementation, and the queue-vs-
interrupt semantics of "steer" are read from docs/extensions.md, not observed.
2026-08-26 17:23:24 +02:00
joakimp e70bef2b5d docs: the edge stamper and the coordination log the fleet actually uses
Two gaps, both found by using the thing rather than reading it.

PROVENANCE WAS SHIPPED UNDOCUMENTED. 553d8657 moved device attribution to the
edge — writer stamped on add_drawer/checkpoint/mine/event_append/artifact_put,
HOST:<device>| prefixed on diary entries — and this README, the file that
documents the extension, never mentioned it. So the only description of the
behaviour lived in the consumer skill, i.e. in the place the code does NOT live.
Now recorded next to Identity, with the two design points that keep getting
re-litigated: the diary marker is in the entry TEXT because diary_read returns
content only (an attribution nobody can see is not an attribution), and RFC 001
§7.3.2 ranks agent-side stamping worst — demonstrated when the agent that wrote
that skill instruction filed its own provenance drawer as added_by=checkpoint.
Includes the version gate that matters in practice: an older image satisfies both
env gates and still stamps nothing, because the extension is baked.

COORDINATION WAS UNDOCUMENTED ANYWHERE. The fleet has used the RFC 003 logstream
for real work since 2026-08-18 (patch handoff, review, a v1->v2 supersede) and no
file in this repo said so. Three facts belong here because they are mechanism:

- The stamper is what makes directed addressing possible. Where every machine is
  a thin MCP client of one shared palace, all clients report the SAME
  origin_replica, so from_agent/to_agent carry the entire distinction between
  machines. Measured 2026-08-26: mesh_peers returns peers: [] with a single
  replica id authoring every event from every machine. An unstamped client
  addressed as bare "pi" is unreachable.
- The bridge is WRITE-ONLY today: it stamps events going out and never reads the
  log. An event addressed to this machine by name reaches the agent only if the
  agent queries for it. Stated plainly because it is the current weak point, and
  it is exactly how a retraction addressed to pi@tor-ms22 sat unread while that
  agent rebuilt the thing it warned about.
- Live push is a DEPLOYMENT question. The palace implements SSE
  (GET /logstream/stream, text/event-stream in mcp_server.py) but a deployment
  may expose only /mcp: verified against mempalace.jordbo.se, where
  /logstream/events, /logstream/stream and /sync/peers all 404 while /mcp serves.
  Enabling it is a proxy route plus an auth decision, not an extension change.

Division of labour made explicit rather than implied: this file documents the
MECHANISM, the consumer skill is NORMATIVE for behaviour. Duplicating the ack
contract here would guarantee two copies that disagree.
2026-08-26 12:49:19 +02:00
pi 553d86570c provenance: stamp device+harness at the edge, not in the agent's head
RFC 001 §7.3.2 ranks "agent stamps provenance via a skill instruction" as the
❌ worst possible place — per-call boilerplate, forgettable, improvisable. It
was right, and we had shipped exactly that: the mempalace skill told the agent
to pass added_by="<harness>@<device>" by hand. Measured on the shared palace,
199 rows had reached it unresolvable, 10 of them filed by the very agent that
wrote the instruction, in a drawer about host provenance. The trigger was a
cross-host misattribution: a session on tor-ms22 read its own diary, could not
tell that the entries were written on EMB-7KJ4VR4G, and reported another
machine's verification as this one's.

Move the same convention into the ⚠️ edge row, where it is uniform and
unforgettable (§7.3.5):

* extensions/pi/mempalace.ts defaults the writer field on every tool that has
  one — added_by (add_drawer, checkpoint), agent (mine), from_agent
  (event_append), created_by (artifact_put) — from $MEMPALACE_PI_DEVICE. An
  explicit value always wins, so filing for another device stays possible. The
  allowlist is per tool, never blanket: 3.8.0's dispatcher hard-rejects
  undeclared args with -32602, so injecting added_by into diary_write or kg_add
  (which have no such property) would break the call outright.
* mine gets miner@<device> when the caller invokes it, but <harness>@<device>
  for the bridge's own transcript feed — bulk extraction is not agent-authored
  memory, and that keeps the pi/opencode/miner taxonomy honest.
* diary_write has no metadata slot at all, and the device must never go in
  agent_name (wing = f"wing_{agent_name}" would splinter the diary per host).
  So the entry TEXT carries an AAAK field, HOST:<device>|SESSION:… — which is
  also the only channel a READER sees: search projects a fixed key set and
  diary_read returns content, so no metadata fix, not even a
  server-authoritative one, would have prevented the misattribution.
* The wake-up block now states the device and warns that diary_read interleaves
  every machine's diary.
* R1: doubly gated on MEMPALACE_PI_DEVICE and MEMPALACE_REMOTE_URL, so a
  solitary devbox stamps nothing and behaves exactly as before — which is also
  the correct semantics per §7.3.3.

Version the reconciler that was living only on synlig (bin/ + contrib/systemd/),
add --dry-run, and teach it two new rules: diary_host_marker reads the HOST:
field, and sibling_chunk propagates a resolved origin across a drawer's chunks
(a text marker lands in chunk 0 only, so a 5-chunk diary entry would otherwise
stamp 1 and leave 4 blank).

--dry-run against the real palace before deploying earned its keep twice, and
scripts/test-device-stamp.sh pins both findings with the strings it found:
HOST: was ALREADY in use with a composite grammar
(HOST:emb-7kj4vr4g.f1d3c3f89e3e.v1.8.3.pi0.84.2) and for bare container ids, so
an unvalidated rule invented devices like "f1d3c3f89e3e.pi0.84.2"; and HOST:
also carries a different SENSE elsewhere (HOST:exec.via.ssh-controlmaster->…,
meaning where I was executing). Validating against the known-device set both
refuses those and recovers the composite entries correctly. A marker convention
inherits every prior meaning of its own name.

Deployed and verified on synlig: device 14,217 → 14,317, integrity ok,
idempotent on immediate re-run, no invented device values.

RFC updates: §7.3.1 corrected (the arg whitelist is a hard -32602 in 3.8.0, not
a silent drop; get_drawer DOES return metadata, search structurally cannot;
triples and logstream live in separate databases the stamper cannot reach),
§7.3.5 added (what is deployed, including the divergence from §7.3.4's opaque
origin_device — tor-ms22 vs tor-ms22-native is that cost already visible), and
Phase 4 now carries per-device tokens motivated FIRST by revocation, with the
finding that tokens are the cheap half: core holds one scalar auth_token and has
zero device concept, so authoritative stamping needs a component we own.
2026-08-25 22:26:46 +02:00
joakimp 0fe64c480e docs: MEMPALACE_PI_DEVICE now also attributes drawers
Follow-up to c64ffa1, which changed the --agent default but left --help
claiming $USER. Fixes that text and states in README/ARCHITECTURE that the
device label reaches the palace as added_by, since mempalace stores neither the
machine nor the harness on a write.
2026-08-23 13:29:51 +02:00
joakimp c64ffa1d93 feeder: default agent to pi@<device> so palace writes carry provenance
mempalace core 3.7.1 records neither which machine nor which harness produced a
drawer, and the single shared bearer token means the server cannot distinguish
clients. Today both facts survive only incidentally -- device in the per-device
inbox path, harness in the pi_*.jsonl filename -- so attribution for anything
filed outside the feeder has to be inferred after the fact.

Defaulting --agent to pi@$MEMPALACE_PI_DEVICE records both explicitly in a field
that already flows through to drawer metadata (added_by), for every enrolled
device, with no core change. Falls back to $USER, which is what pre-existing
drawers carry (added_by=joakim on fed transcripts, hence the ambiguity).

Not pushed: lands on machines at the next image build.
2026-08-23 13:08:40 +02:00
joakimp fd8b15f570 docs: the convos miner does check mtime — finish a correction that stopped half-way
ARCHITECTURE.md and README.md still carried the claim from 954c3f2 (initial
commit) that the convos miner "keys on source_file path alone (convos miner
doesn't check mtime)", and told the operator to delete the staging dir to force a
re-mine. 29e660e corrected exactly that claim in AGENTS.md and SKILL.md — and
missed these two files, so the repo has been documenting both behaviours at once
ever since. Two files said mtime is checked, two said it is not.

Ground truth, read off the deployed mempalace 3.7.1 rather than inferred:
convo_miner.py:657 calls file_already_mined(..., check_mtime=True) inside
mine_lock(source_file), and palace.py:1430 re-mines when no drawers exist for the
source_file, when the stored normalize_version predates the current schema, or
when the mtime differs. On a mismatch the file's existing drawers are purged
(_source_file_delete_ids -> collection.delete) before refiling, so a changed
transcript is replaced rather than doubled. The docstring names the case outright:
transcripts are not assumed immutable, since a session keeps appending to its own
file while active and /compact or /clear can rewrite one in place.

The stale advice was not merely out of date, it was expensive. "Delete the staging
dir to force a re-mine" is the one gesture that re-keys dedup: the staged path IS
the key, so a stage that is wiped or recreated elsewhere makes the palace refile a
whole wing as duplicates instead of replacing it. The docs now say so, name `touch`
as the non-destructive way to force a single session, and record why staged copies
must carry the source's mtime — with the corollary that an old mtime in a stage or
a remote inbox says nothing about when the file was shipped, so a ship is judged by
the feeder's log instead.

Sample output blocks quoting "(dedup by source_file)" are deliberately left
verbatim: that is what bin/mempalace-session:426 and bin/mempalace-pi-session:857
actually print. Tightening the wrappers' own wording is a separate change, because
the samples have to move with it.
2026-08-18 09:42:40 +02:00
Joakim Persson 947604b25d docs: backup and recovery, plus units; move host runbook to a private repo
Adds bin/mempalace-backup and docs/backup-and-recovery.md — the mechanism a
palace actually needs, none of it site-specific.

Why a palace cannot be backed up with cp: it is chroma.sqlite3 (authoritative),
knowledge_graph.sqlite3 (usually WAL, so -wal/-shm make a plain copy a
same-instant gamble), derived HNSW segment dirs, hallways.json, the embedder
descriptor, and a HIDDEN .mempalace/origin.json. Both SQLite files are therefore
copied through the online-backup API. Two bugs are documented because both
produce a backup that looks fine: "$PALACE"/*/ silently skips the hidden dir, and
per-directory rsync collides the identically named data_level0.bin in every HNSW
segment. Treating the palace as one tree fixes both and makes a backup a faithful
palace IMAGE, so restore is a copy rather than a procedure.

Two modes: hot (default, zero downtime, ~4 s, index may lag but SQLite is
authoritative and repair --mode from-sqlite rebuilds) and cold (--cold, ~5 s
downtime, byte-consistent, restart trapped so a failed run still brings the
server back). Verification runs on the COPY — quick_check plus row counts — and
the backup is committed by mv only after it passes, with retention pruned only
after a verified commit, so a broken new backup cannot delete the last good one.

Documented because they are easy to get wrong: the sqlite3 CLI is often absent
where the Python module is present; mempalace_embedder.json must be restored with
the drawers or search silently degrades; a tested restore means running status
AND search against the restored copy, since search is what actually exercises the
index; mempalace-serve is a USER unit, so root systemctl reports "not found";
Persistent=true is what makes a missed window run after boot; and installing
against the system Python couples the palace's availability to distribution
upgrades, with the uv-managed-interpreter fix plus the two PATH traps that bite
scripted upgrades.

Also moves docs/synlig-primary-runbook.md out to a private fleet repository,
leaving a stub that explains the split, since a host inventory is operator data
for one deployment rather than part of a public toolkit. The path stays valid so
existing links do not break. Remaining host references in the README, RFCs and
ARCHITECTURE are left alone deliberately: they are load-bearing prose, contain no
secrets, and are best generalised as they are next edited rather than in one
churn-heavy pass.
2026-08-17 00:50:16 +02:00
Joakim Persson b609cf5a69 docs(synlig): the transcript inbox is the primary's third moving part; and survive an unset HOME
Runbook gaps found while fixing the 2026-08-15 feed failure:

- §2.5 still titled "written but not installed" and still asserting "Not
  installed, not enabled" — false since 2026-08-12. A reader landing there got
  a flat contradiction of §4 item 3. Retitled, with the verified-2026-08-16
  process line, and it now states the fact §2.6 depends on: the server is a
  NATIVE process (no mempalace container on synlig), so it sees host paths.
- New §2.6 documents ~/mempalace-feed/<device>/: why transcripts cannot travel
  over the HTTPS leg at all, the three client variables, and why
  MEMPALACE_PI_REMOTE_PATH is the trap (its /data/feed default assumes a
  containerized server; here it must equal the ssh-target path, and a mismatch
  fails with rsync succeeding and only the mine failing). Plus the operational
  notes that cost time: dedup keys on the absolute path so the inbox path is
  load-bearing, grown sessions are purged+refiled by mtime, and a client-side
  MCP timeout is NOT a failed mine.
- Header status: counts refreshed with an explicit "treat counts as timestamps".

Also, bin/mempalace-pi-session: default HOME from the passwd database when it is
unset. `docker run --entrypoint="" <image>` inherits no HOME when the image
config declares none, and every default is HOME-anchored under `set -u`, so the
script — including the palace-free --self-test — died with "HOME: unbound
variable" in exactly the environment pi-devbox's smoke suite uses. pi-devbox
v1.8.0 lost a release to the same assumption from the other side.
2026-08-16 00:52:27 +02:00
Joakim Persson 6e1f4f30fc fix(pi-session): a failed remote mine reported success — MCP escapes the payload
The remote-mine leg decided success with `'"error"' in body`. MCP answers a
hard tool failure with HTTP 200 and a JSON-RPC *result* whose content[].text
carries the tool's own JSON as an ESCAPED string, so those bytes are
\"error\" and the substring never matches. On 2026-08-15 (EMB-7KJ4VR4G, first
boot of the fresh pi-devbox image) a mine that failed with

  {"success": false, "error": "source directory not found: '/data/feed/emb-7kj4vr4g'"}

printed "Done. Wing 'wing_conversations' updated." directly under that error
and exited 0. Transcripts had been rsynced for the whole session and filed
nowhere; the only artifact anyone would check said it worked.

- classify(): parse the envelope instead of grepping it. Catches JSON-RPC
  errors, MCP isError, and inner success=false/error, and distinguishes
  "verified ok" from "unverified: no JSON tool payload" rather than assuming.
- --self-test: six recorded MCP responses (fixture 1 is the real 2026-08-15
  body) plus a regression guard asserting the old substring check is blind to
  it. Needs no palace, no network, no sessions dir.
- Preflight warning when the rsync destination path and
  MEMPALACE_PI_REMOTE_PATH disagree. The /data/feed default assumes a
  CONTAINERIZED palace server; a native one (systemd unit / uv tool) sees host
  paths, and then the two must match. Warned in preflight so --dry-run and
  --prepare surface it too.
- Remote mode no longer previews NEW/SKIP from the LOCAL palace: dedup happens
  on the palace host keyed on the remote inbox path, so this machine cannot
  answer it. Tags become [?] and the summary says who decides. It had been
  reporting "6 already filed" about a palace it was not feeding.
2026-08-16 00:28:50 +02:00
Joakim Persson f60cf9c732 feat(census): RFC 002 Phase A — read-only join census, and three RFC corrections it found
bin/mempalace-census classifies a palace on disk into the RFC 002 §2 classes
(mined / diary / agent-authored) and emits a human report or a --json manifest
that feeds Phases B/C. Read-only: every sqlite handle is opened mode=ro, no
-wal/-shm is created, safe against a live mempalace-serve. Reads LOCAL DISK only
and warns if MEMPALACE_REMOTE_URL is set, so a local census can't be mistaken
for a remote one.

The design point is that it SELF-VERIFIES instead of trusting ids.py's
docstrings: for every replayable drawer it reassembles content from chunks,
recomputes the upstream id and compares to the stored id. That single check
covers the id recipe, the chunk reassembly order and the classifier at once --
176/176 accounted for on the reference palace -- and it falsified three things
the RFC previously asserted from a docs-only reading (now RFC 002 §2.1):

  (a) The hash input is LENGTH-PREFIXED, not "|"-joined. ids.py:31 defines
      _DELIM = "|" and the make_* docstrings describe f"{wing}|{room}|{content}",
      but _DELIM is dead code and _delimited_sha256 builds
      "".join(f"{len(part)}:{part}"). Measured: length-prefixed reproduces real
      ids 5/5, pipe-joined 0/5. Diary ids differ again -- a PLAIN sha256.

  (b) id_recipe is NOT a mined-only marker. It looked like a clean
      discriminator (same 14,586 count as source_file) but the server stamps
      'v3' on content ids too, so classifying on `source_file OR id_recipe`
      swallowed all 60 agent-authored drawers into MINED -- the dangerous
      direction, since Phase C would try to re-mine drawers that have no source
      file and silently drop them. Discriminator is a TRUTHY source_file (the
      writer stores "" rather than omitting the key), cross-checked against the
      miner-only keys source_mtime / normalize_version; disagreement is now a
      first-class warning.

  (c) Content ids DRIFT: 9 of 60 agent-authored drawers no longer reproduce
      their own id, because update_drawer preserves the id while rewriting and
      re-chunking. So "recompute the content id and skip if present" -- the
      strategy this RFC specified for regime A -- misses every drifted drawer
      and duplicates it. Phase C must key on the STORED id. Flagged as
      edited_since_filing in the manifest so Phase C can be tested on them.

Also corrects §4.1's headline number: 15,949 mined / 98.9% was reconstructible
exactly as 16,338 (all embeddings rows) - 192 (diary rows) - 197 (agent rows),
i.e. it counted rows rather than parent drawers AND spanned both collections,
absorbing all 1,560 non-joinable mempalace_closets rows into the mined total.
Correct figures: 14,389 mined / 116 diary / 60 agent-authored = 14,565 parents,
replay surface 176 (1.2%). Two rules now enforced in the tool: always filter by
collection (one sqlite file holds both), and always say whether a count is rows
or parent drawers -- a chunked drawer contributes N rows and no parent row.

One implementation trap worth recording: Chroma splits metadata across
string_value and int_value, so reading only string_value nulls every numeric key
(source_mtime, chunk_index, line_start) -- which made the miner-marker
cross-check report 100% conflict until the loader coalesced the two columns.
2026-08-15 09:06:40 +02:00
Joakim Persson 2f9170428c docs(rfc-002): run the census for real — the replay surface is 176 records, and closets were missing
Ran Phase A's classifier against the frozen EMB-7KJ4VR4G archive. Since the
fleet shares one devbox image this is a reasonable prior for tor-ms22 and
MBP-M1-2020:

  mined (source_file set)   15949   98.9%   re-mine, never replay
  diary entries               116    0.7%   replay + §7.6 suffix skip
  agent-authored drawers       60    0.4%   replay, idempotent
  KG open facts                34      --   server guard dedupes
  KG closed facts               0      --   nothing to do

Two consequences that shrink this project: the replay-only surface is 176
records, not thousands, so the writer is a small job and the §7.6 diary guard
treated as the blocker governs 116 records; and there are ZERO closed KG facts,
so the unguarded-closed-fact gap is real in the code but empty in the data.

filed_at spread in the same archive -- 12 in May, 52 in June, 15174 in July, 887
in August -- is the concrete argument for the history-preserving regime.

Two corrections to the first draft:

1. CLOSETS were omitted entirely. mempalace_closets is a second Chroma
   collection (1560 rows, ~10% of the palace) and there is NO MCP tool that
   writes one -- mcp_server.py exposes only _purge_source_closets, and
   closet_llm.py says outright that regex closets are always created by the
   miner. So closets cannot be replayed even in principle; they return only by
   re-mining. Same category as hallways/known_entities/palace-graph, and it
   resolves itself for the ~99% that gets re-mined anyway.

2. Census gotcha: chroma.sqlite3 holds BOTH collections, so an embeddings-wide
   query over-counts by ~10%. Must join segments->collections and keep
   mempalace_drawers. Doing so reconciles exactly: 14,778 = the 14,777 seeded to
   the primary + the one known post-snapshot chunk.

Also drops the earlier "438 of 14,829" figure, which conflated central's
no_source count (which includes its own later agent writes) with the archive's
actual replay surface.
2026-08-15 00:29:58 +02:00
Joakim Persson a25e22922b docs: scope the joiner (RFC 002) — and the chronology loss nobody had costed
RFC 001 §4.4 designs a join as "idempotent replay of local history", but no
replay tool exists and the two joins done so far were whole-palace file copies
that cannot merge. tor-ms22 and MBP-M1-2020 each hold a palace that needs to
reach the primary, so this scopes the tool.

The finding that drives the design: MCP replay CANNOT preserve filed_at.
add_drawer stamps it server-side (mcp_server.py:2580) with no override, and
diary_write builds its own now()-based id. A pure MCP replay would therefore
collapse months of history into the join instant -- on a palace whose value IS
its chronology, and where list_drawers filters on filed_at. RFC 001 does not
mention this. kg_add is the exception: it takes valid_from/valid_to, so fact
windows survive.

Hence two regimes, and a recommendation to build the history-preserving one
first: direct disk write on synlig (preserves ids + filed_at, bypasses the
server guards, needs the service stopped) vs MCP replay (guards work, timestamps
flatten). migrate.py is already a working model for the direct path -- it reads
drawers straight from the palace sqlite and re-adds them preserving ids,
documents and metadata.

Also concretised: every dedup key verified against mempalace 3.6.0 source rather
than assumed --
  - agent-authored drawers: sha256(wing|room|content)[:24], fully deterministic
  - mined drawers: sha256(source_file|chunk_index)[:24], PATH-dependent, so
    replay duplicates instead of deduping -> re-mine on synlig, do not replay
  - diaries: full ids never repeat (wall-clock component); match the
    sha256(entry)[:12] suffix only -- this is §7.6
  - KG closed facts: no server guard at all (it is scoped to valid_to IS NULL)
  - hallways/known_entities/palace-graph are built at MINE time, so replayed
    drawers arrive with no co-occurrence edges and traverse under-reports

Scope is phased so the useful half lands first: Phase A is a read-only census
that sizes the job and cannot break anything, and is the piece to run the moment
tor-ms22 is reachable. The replay-only surface is small -- only 438 of 14,829
drawers on the primary have no source_file.

Method note recorded in the doc: an attempt to gather these mechanisms via a
delegated subagent returned confident, fabricated code for a package path that
does not exist on this machine. Everything here is cited to file:line and was
re-read directly.
2026-08-15 00:27:36 +02:00
Joakim Persson c349d007e1 docs(§7.2): measure the shared-palace sync blast radius — it is 0 today, and the feeder is what changes that
Unscoped `mempalace_sync` dry-run against the live primary (14,829 drawers):
out_of_scope 14391, no_source 438, missing 0, kept 0. Zero drawers deletable.

The mechanism is the finding: an entirely-absent source root yields
`out_of_scope`, NOT `missing`, and only missing/gitignored drawers get removed.
So a wipe needs the root to EXIST while files under it do not -- not merely a
host that lacks the repos. §7.2's "from a laptop that lacks the repos, it is a
fleet-wide wipe" therefore overstates today's risk.

It also understates tomorrow's, which is the part worth acting on. Two guards
currently prevent the laptop scenario and neither was designed to: the CLI has
no remote support, so a client's `mempalace sync` cannot reach central at all;
and the MCP tool runs server-side on synlig, where clients' /workspace roots do
not exist. That is safety by coincidence of layout.

The feeder's remote mode dissolves it: it rsyncs staged transcripts into
per-device inboxes ON synlig, so those sources begin existing on the palace host
and become in-scope for the first time. A later stage rotation or inbox cleanup
then marks that device's conversation drawers `missing` -- prunable, and
prunable from a different device. mempalace-pi-session already documents the
single-machine form of this; a shared palace makes it cross-device.

So stage retention on synlig plus the sync guard belong to enabling the feeder
fleet-wide, not to a later cleanup pass.
2026-08-14 23:15:01 +02:00
Joakim Persson 6e8172d93a docs: record the primary's real lineage — EMB-X1JY06WJ -> EMB-7KJ4VR4G -> synlig
Three docs said the primary was "seeded from EMB-7KJ4VR4G's palace", which is
true but stops one hop short. That palace was itself carried over from
EMB-X1JY06WJ, the previous work computer, when it was replaced around
2026-07-06.

Evidence, read read-only out of the frozen archive on EMB-7KJ4VR4G: 76 drawers
carry `source_machine=EMB-X1JY06WJ` and they are the oldest in the store --
earliest filed_at 2026-05-04, two months before that machine's stack existed --
with no other source_machine value present. The remaining ~16k were filed
locally afterwards (15,189 in July, 1,073 in August). Last write is
2026-08-14T15:07:06, the seed instant.

Two consequences, recorded in rfc-001 S4.4:

- It was the *second* whole-palace file copy, not the first, and both were
  single-source clones. So the method has two successes behind it and still
  zero exercises of merge semantics -- S7.6 is *less* tested than "we've done
  this twice" would suggest, not more.
- The primary now holds records from a machine that no longer exists, and
  `source_machine` is the only thing marking them. Not noise; don't prune it.

Also corrects a94eb7f, which inferred the archive's origin from the
`.devbox-owner` marker. That marker only records that the devbox adopted the
directory and says nothing about provenance -- the archive predates the stack.
And native pi on EMB-7KJ4VR4G is "not installed *yet*", a deferred hazard rather
than a closed one: a native install there would resolve its palace to
~/.mempalace, i.e. the frozen archive, and split that machine's memory from
central silently. Flip at install time.
2026-08-14 23:04:35 +02:00
Joakim Persson a94eb7fdd0 docs: native pi has no .env to flip, and EMB-7KJ4VR4G has no native pi
Follow-up to ec436ed, closing the one item that commit left as "unverified".

Verified on the EMB-7KJ4VR4G host: native pi is not installed at all -- no `pi`
or `mempalace` on PATH, no ~/.config/pi/, no ~/.pi/agent/extensions/. That
host's ~/.mempalace exists solely to back the devbox container through the bind
mount (.devbox-owner holds 1000:1000). So the 2026-08-14 flip covers every pi
on that machine and there is no split-brain to fix. The general hazard stays
documented, because it is a per-machine question.

Also documents how a native install *would* be flipped, since the obvious guess
is wrong: pi loads no dotenv file and has no `env` block in settings.json, and
the extension reads process.env only. The launching shell is the sole hook, so
the vars must be exported from a shell rc (a GUI-launched pi may not read one),
and ~/.config/pi/.env is not sourced by anything automatically.
2026-08-14 23:00:46 +02:00
Joakim Persson ec436ed3ad docs: reconcile the RFC-001 docs with what is actually deployed
Audit of every doc touching the global-palace rollout against the running
fleet. Each correction below was verified against the filesystem or the host,
not against another doc:

- synlig-primary-runbook: the decommission `rm -rf ~/.mempalace` now carries a
  STOP block. That tree holds the fleet palace *and* the only copy of the
  bearer token every client authenticates with; the old "empty today" comment
  stopped being true when the palace was seeded on 2026-08-14. Adds an ordered
  safe decommission, and drops count-based join verification.

- phase-1-exposure-runbook: new S3.8, how to verify a flip actually took --
  the procedure that until now existed only in an untracked handover file.
  Three claims that fail independently (env var / curl / the palace-path
  discriminator) plus an explicit list of checks that produce FALSE POSITIVES:
  drawer counts (both sides were seeded from the same palace, and `status`
  counts chunks not drawers), write-then-read through the same transport, and
  the `mempalace` CLI -- which has no remote support at all, so post-flip it
  reads the dead local archive and reports success.

- rfc-001: status Draft -> Phases 0-1 implemented. Records that the join was a
  file-level copy, which SIDESTEPPED the S7.6 diary-dedup question rather than
  answering it -- so S7.6 remains a hard blocker for the second machine, which
  is the one that will actually exercise merge semantics.

- ARCHITECTURE, SKILL, contrib/README, extensions/pi/README all claimed pi
  feeds the palace automatically, unconditionally. That is gated on
  mempalace-toolkit >= 29e660e and every deployed image predates it, so the
  claim is currently false fleet-wide. Each site now states the gate plus a
  check that inspects the *deployed* file rather than repo HEAD.

- extensions/pi/README: plaintext http://mempalace.lan example -> https
  endpoint; the two transports are either/or (no dual-write, no local mirror);
  the bridge fails CLOSED, so "the agent has no mempalace_* tools" is the
  expected symptom of a server/token/DNS fault, not of a broken install.

- contrib/README: documents mempalace-serve.service, which this directory has
  shipped since day one without explaining it (linger, the load-bearing
  172.17.0.1 bind and why loopback is the unsafe-looking-safe option, the
  token path, and an uninstall warning).

- Fixes a pre-existing stray ```sh fence that was swallowing S3.2's heading and
  the token command into a code block.

Docs only; no behaviour change.
2026-08-14 22:57:22 +02:00
Joakim Persson 2293f1c89b docs: the silent-skip fix is base-image-gated, so today's fleet still skips silently
pi-devbox cbd7cf5 makes the remote-palace-without-inbox skip announce itself, but
entrypoint-user.sh is COPY'd in Dockerfile.base, so the fix only reaches a
container after a base rebuild. Every image running today still skips silently --
recording both behaviours with the cut-off, so the table stays true for whichever
image a container is actually on, rather than describing a fix that has not
shipped yet.
2026-08-13 16:31:18 +02:00
Joakim Persson 2e73a9eae8 docs: REMOTE_URL includes /mcp, and the entrypoint skips silently where the feeder exits 1
Two details needed before the first client flip.

1. MEMPALACE_REMOTE_URL is the full endpoint including /mcp with no trailing
   slash. The feeder POSTs to it verbatim (bin/mempalace-pi-session:718), and the
   server matches `path != "/mcp"` exactly, so a base URL or a trailing slash
   both 404 -- and a 404 here looks like a routing/proxy fault, not a config
   typo, which is a bad hour to spend. Matches the existing pi-devbox examples.
   For this fleet: https://mempalace.jordbo.se/mcp

2. The 3.7 trap has TWO symptoms, not one, and I had only documented the loud
   one. Direct runs / session-end / cron exit 1 with a clear error (:298-300).
   But pi-devbox's entrypoint-user.sh:134 checks the same condition and skips
   *quietly* -- and the skip happens before the subshell that writes
   ~/.pi/agent/mempalace-catchup.log, so there is not even an empty log to find.
   A container flipped with REMOTE_URL but no SSH_TARGET therefore contributes
   nothing to the palace and leaves no artifact explaining why. The skip is
   correct in itself (there is genuinely no inbox to ship to) but it is
   indistinguishable from a healthy run with nothing to do. Documented with two
   commands that tell those apart after a flip.
2026-08-13 16:22:42 +02:00
Joakim Persson 4cb70ce3e1 docs: record the 302 auth-redirect fingerprint and the server's exact HTTP surface
Phase 1 exposure now verified end to end (2026-08-12): /healthz -> ok and an
unauthenticated POST /mcp -> 401, both from a client container and from the
primary itself. 3.4 and 3.6 marked passed.

Two findings from the failure in between, worth more than a checkbox:

1. Leaving Pangolin's resource authentication on does NOT surface as 401 or 403.
   It is a 302 with `location: https://pangolin.jordbo.se/auth/resource/<uuid>
   ?redirect=...`, sent *before* the request is proxied -- so the palace never
   sees it and its journal stays silent, `curl -s` prints an empty body, and an
   MCP client gets non-JSON. That is 1.1's "Pangolin's HTTP auth is
   browser-shaped; the clients are not" arriving as a concrete symptom rather
   than an argument. Now recorded with the exact header shape and the two curl
   flags that reveal it (-D-, -w '%{redirect_url}'), plus the reading: a 302 is
   good news, because DNS, TLS and routing all worked and only auth intervened.

2. Read the server's routing to settle whether path-scoped proxy rules are
   sufficient. The entire HTTP surface is two exact paths (mcp_server.py:5299-
   5318): GET /healthz (no token, Host/Origin gated) and POST /mcp (Bearer,
   compare_digest on the exact string). Everything else is a 404 from the palace
   itself. So path-scoped rules are tighter than a host-wide proxy and lose
   nothing. Two client-facing consequences: /mcp is matched exactly, so a
   trailing slash 404s -- configure clients without one; and there is no GET
   /mcp, no SSE, no session id, no DELETE, so this is plain JSON-RPC over POST,
   not MCP streamable-HTTP. A strict client opening with a GET handshake sees
   404. Also means the verify step needs no initialize and no Accept:
   text/event-stream, which the old snippet left ambiguous.

Also: run the 200 and 401 from two different networks, not one -- passing from
only the primary leaves split-horizon DNS untested.
2026-08-13 16:21:32 +02:00
Joakim Persson 08e344b047 docs: proxy target is http:// not https://; loopback probe refuses, it does not 403
Both corrections come from the first real Phase 1 start on synlig (2026-08-12).

1. The Pangolin resource target was documented as bare `172.17.0.1:8765` with no
   scheme, and the obvious guess from that is `https://` -- which cannot work.
   `contrib/systemd/mempalace-serve.service` runs `serve --host 172.17.0.1
   --port 8765` with no --tls-cert, so the primary speaks plaintext HTTP; TLS
   terminates at Pangolin, which is the whole point of the RFC 6.2 decision.
   Point a proxy at https:// and it attempts a TLS handshake against a plaintext
   listener: 502 from outside, while `curl 172.17.0.1:8765/healthz` on the box
   still says ok -- a confusing pair of symptoms. Now spelled `http://` with the
   failure mode named, in the runbook and in the unit's comments.

2. `curl -s 127.0.0.1:8765/healthz` was documented as "expect 403". Wrong: the
   real run returns empty. With a docker0-only bind nothing is listening on
   loopback, so the connection is refused at TCP level before any header is sent
   (%{http_code} -> 000, exit 7). The 403 is the *loopback-bind* case verified
   2026-08-10 -- server on 127.0.0.1 answering a proxy-forwarded foreign Host.
   Two distinct behaviours had been collapsed into one expectation in three
   places (both runbooks and the unit).

   Worth stating why the correction matters rather than just fixing it: refusal
   is the *stronger* signal. A 403 proves only that a request was rejected; a
   refused connection proves the loopback and LAN surface is not listening at
   all. Someone who expected 403, saw silence, and "fixed" it by rebinding to
   0.0.0.0 would have converted a correct configuration into an exposed one.
   The docs now also say what to do if it hangs, or if ss shows 0.0.0.0:8765.
2026-08-13 16:01:02 +02:00
Joakim Persson 00a95d1a2f docs: why the tunnel and the feeder's SSH path are not redundant; newt done
Asked "why do we need Pangolin if you proposed rsync/ssh?", and the runbook did
not actually answer it -- it stated both were needed without saying why neither
substitutes. New section 1.3:

  - Pangolin/HTTPS carries the MCP tool surface (search, add_drawer,
    diary_write, kg_*) -- every live tool call, from any MCP client.
  - SSH/rsync carries transcript *files* only, because mempalace_mine expands
    its source path server-side, so the server can only mine its own disk.

HTTPS alone is a palace you can query but cannot feed; SSH alone is files with
no query API. The rsync is not a transport preference, it is a workaround for
where `mine` resolves paths.

Records honestly that `ssh -L 8765:172.17.0.1:8765 synlig` *would* replace the
tunnel for MCP, and why we don't: synlig dials out (reaching for a dial-out
tunnel is itself the evidence inbound was unavailable), MCP clients want a
durable URL rather than a per-session forward, and the forward must be up on
every device before every session.

And the design's weak point, stated instead of glossed: the rsync runs
client -> synlig, so mining needs synlig's SSH reachable *from the client*. Were
that true everywhere, no tunnel would be needed for MCP either. Honest
expectation after Phase 1 is therefore: query/write from anywhere, mine only
from devices that can reach synlig's SSH. Section 4 now carries the upstream ask
that would close the gap -- have the feeder send content over MCP (add_drawer /
diary_write, which it already calls) instead of asking the server to mine a path
it must first rsync there.

New section 3.7, a live trap for the imminent client flip:
MEMPALACE_REMOTE_URL on its own does not degrade to local feeding, it *stops*
feeding. auto mode switches to remote as soon as the URL is set (:286) and
remote mode then exits 1 without MEMPALACE_PI_SSH_TARGET (:298-300), before
anything is staged or filed -- so a cron feeder just starts failing, and the
loudest symptom is silence. Two safe orders given: set all three variables in
one edit, or set URL+token and pin --mode local until the SSH target exists.

newt is installed on synlig and connected to Pangolin (done 2026-08-12), marked
here and in the synlig runbook's item 2; the blocker is now item 3, the one
sudo. Added the follow-up that "connected to Pangolin" only proves newt reached
nyvaken -- reaching the *palace* is a separate claim that fails independently,
so probe 172.17.0.1:8765/healthz from inside newt's namespace.

All seven code citations verified against the source at commit time.
2026-08-13 00:15:57 +02:00
Joakim Persson 29e660e18f feeders: stage beside the palace, not in ~/.cache; document Phase 1 exposure
Staging default moves out of ~/.cache to <palace-root>/pi-stage (pi) and
<palace-root>/opencode-stage (opencode), resolved with mempalace's own
palace-path precedence ($MEMPALACE_PALACE_PATH -> $MEMPAL_PALACE_PATH ->
~/.mempalace/config.json -> ~/.mempalace/palace), then dirname.

Why: the convos miner keys dedup on the *staged* path, so a wiped stage plus a
sync scoped to include it prunes the drawers mined from those sources --
deleting memories, not a cache. Under ~/.cache that state was reachable by
anything treating a cache as disposable. Staging inside the palace makes the
coupling structural: the stage cannot be wiped without touching the palace
itself. Overrides ($MEMPALACE_PI_STAGE / $MEMPALACE_SESSION_STAGE, --stage) are
unchanged. Note the old default had never been created on any host, so this
closed a latent hazard, not a live one.

Measured, and the docs now claim only this much: sync prunes only within the
scope it is given -- wing-only, 1299 scanned / 1299 out of scope / 0 removed;
scoped at the palace root, 651 kept / 648 out of scope. The previous blanket
"sync prunes every drawer" wording overstated it, which is a liability: the next
reader disproves the overstatement and discards the real constraint with it.

Also in this change:
- cron log dir ~/.cache/mempalace-session -> ~/.cache/mempalace-logs. The stage
  left that namespace, so the old name now read as "the stage".
- AGENTS.md: the convos miner *does* check mtime (verified against upstream
  convo_miner.py); the previous "no mtime check" claim was wrong.
- smoke-test assertions use `mktemp -d` for --sessions-dir. One pointed at /tmp,
  which still held earlier synthetic transcripts, so a --dry-run exported a fake
  session into the real stage: --dry-run skips the mine, not the export.

docs/phase-1-exposure-runbook.md -- the newt/DNS/auth step that RFC 001 and the
synlig runbook leave open (runbook section 4, items 2 and 5). Port 8765 at /mcp,
newt targets 172.17.0.1, and the authentication is the single shared bearer
token (RFC 6.2, decided 2026-08-09) rather than per-device proxy users. The
latter cannot work today: mempalace validates exactly one token, and Pangolin's
SSO/PIN/password are browser-shaped while every client here is a headless
JSON-RPC POST -- enabling that protection breaks the clients it protects. The
per-device axis that *does* exist is the feeder's SSH key + per-device inbox.

New finding recorded there: a loopback bind does not merely 403 behind a tunnel
(already known, runbook 2.4) -- it also silently starts the server with no token
at all, because auto-minting is gated on the bind being non-loopback.

extensions/pi/README.md: the HTTP transport IS authenticated as of mempalace
3.6.0; the "sessionless and unauthenticated" note dated from the v1.3.0 era.
Closes the RFC section 8 Phase-0 hygiene item.
2026-08-12 17:04:01 +02:00
Joakim Persson 3626946013 Phase 0 on synlig: install, palace layout, verified Host/Origin policy
synlig is greenfield — no mempalace, no ~/.mempalace at all — so Phase 0
became "provision correctly from birth" rather than "migrate carefully".
Nothing is serving; no client config was touched.

Done:
- mempalace 3.6.0 installed via uv, pinned to the fleet version (the id
  recipes and idempotency probes this RFC leans on are version-specific).
- Embedder pre-warmed. This was the real unknown: the first embed pulls a
  79.3 MB ONNX model from the chroma CDN, and an egress-filtered work VM
  would have failed at the worst moment — the first client write. Pulled
  at ~20 MB/s, no proxy interference. Done in a throwaway palace so the
  real one never saw it.
- Palace at the stock default ~/.mempalace/palace, so no config file and
  no MEMPALACE_PALACE_PATH is needed on synlig at all.

Two corrections to the RFC, both from provisioning rather than reading:

- §7.1 was understated. DEFAULT_PALACE_PATH (~/.mempalace/palace) and
  DEFAULT_KG_PATH (~/.mempalace/knowledge_graph.sqlite3) already differ
  with stock defaults, so the KG split is out-of-the-box behaviour, not a
  consequence of a custom --palace, and it is permanent rather than
  one-time: serve always passes --palace, any CLI call without it uses
  the HOME path. A one-time mv does not fix that, it only picks which of
  the two files gets populated. Fixed instead by converging both rules on
  one inode via relative symlinks, and verified the load-bearing
  assumption: a dangling symlink is created on connect, -wal/-shm land
  next to the target (so the palace dir stays a self-contained backup
  unit, which is the part that mattered), cross-path read works, same
  inode. hallways.json deliberately left alone — already palace-derived,
  HOME path is a warning-only probe.
- §6.2 upgraded from "test early" to verified: 11/11 as predicted. The
  headline is that the safe-sounding reflex is the failure mode — loopback
  bind + proxy forwarding a public Host is 403, non-loopback is 200. Also
  confirmed Origin is never relaxed on either bind, and /healthz is
  Host/Origin-gated but token-free, so it works as the tunnel probe.
  Recommends binding the docker0 gateway over 0.0.0.0: non-loopback so
  the pin relaxes, but reachable only from the host and its containers.

New docs/synlig-primary-runbook.md carries the discovered facts about the
box, the evidence tables, an explicit "deliberately not done" list, and
tomorrow's ordered steps. New contrib/systemd/mempalace-serve.service
carries the bind rationale inline so nobody "fixes" it back to loopback;
staged on synlig with a .staged suffix so systemd cannot pick it up by
accident.

Flagged for tomorrow: synlig has no newt/tunnel client (docker ps shows
only the Gitea runner and digikam), so Pangolin on nyvaken cannot reach
it until one is added — easy to miss, because Pangolin will look healthy
from its own side.
2026-08-10 00:14:06 +02:00
Joakim Persson 7e51055c96 rfc-001: join protocol, diary non-idempotency, deployment decisions
Second planning round. Three findings from verification, and the
decisions that were blocking phasing.

Verified in mempalace 3.6.0 and written up:

- NEW §7.6 — diary_write has no idempotency guard at all. The id embeds
  datetime.now() at microsecond resolution and the write is a bare
  col.add (mcp_server.py:3546) with no col.get probe, in direct contrast
  to add_drawer's probe 900 lines earlier (:2593-2600). So §3's "log the
  intent, replay the intent" is false for diaries — replay duplicates.
  This lands on the critical path because diaries are `replicated` and
  are exactly what a join replays.
- NEW §4.4 — join/bootstrap protocol, which the RFC simply lacked. Not
  "seed the primary from one palace": every container joins the same way,
  repeatedly, so a join is idempotent replay and the only question per
  record type is what dedupes it. Drawers and open KG facts need zero
  client bookkeeping; closed KG facts and diaries need client-side keys.
  Join state belongs in the palace dir, not the container (~/.mempalace
  is not preserved by default), which also makes two containers sharing
  one host's palace the easy case rather than a double-upload hazard.
- §3.2 — the add_triple guard is scoped to `valid_to IS NULL`, so closed
  historical facts are unguarded. §3 read as unconditionally idempotent.
- §7.3 — drawer/diary metadata is exhaustive: no session, PID or
  conversation field. Two concurrent pi sessions are indistinguishable,
  and pi-vs-opencode is only accidentally distinguishable because the pi
  extension never sets added_by (it sets identity for diaries only,
  mempalace.ts:758). Corrects the §7.3.3 bullet that claimed
  multi-harness attribution was already solved — the field is the right
  home, but nothing populates it. Design rule stated: device + agent,
  never session; provenance granularity equals token granularity.
- §6.2 — Host-pinning is coupled to the bind address
  (enforce_host_pin = _http_is_loopback(host), :5367). The trap is the
  RFC's own "bind loopback" reflex: behind a proxy that forwards a public
  Host, a loopback bind 403s, while a non-loopback bind deliberately
  relaxes the pin. The Origin check is never relaxed. This replaces the
  warning I first wrote, which had it backwards.
- §9.3 — largely resolved: the last-chunk-only probe is deliberate
  (batched upsert is all-or-nothing, :2586-2592). Only a crash mid-upsert
  remains untested.
- §4.1 — fix shape for opencode's stale-config asymmetry: make the
  mcp.mempalace subtree env-authoritative, with a fingerprint so
  hand-edits still win. pi-devbox/entrypoint-user.sh:131-162 already has
  the jq deep-merge pattern to port; no OPENCODE_CONFIG* layer exists
  upstream, so the real file must be written. Now Phase 1.5.

Decisions (new §8.1): primary = synlig; TLS at Pangolin; single shared
token for now (so origin_device stays advisory); diaries replicated.
Work/personal is per-wing, not per-device — two primaries split along
machine lines is rejected, because pi-devbox/opencode-devbox are
simultaneously work and home projects and the device where work happened
cannot classify the project. MEMPALACE_REMOTE_URL stays scalar so
multi-store remains additive later (new §9.7 keeps the placement
question open).
2026-08-09 23:48:57 +02:00
Joakim Persson fdcd5871de rfc-001: resolve open questions 1 and 6 — both permissive
Recon before planning the implementation, and both blockers dissolved.

Q1, does opencode support remote MCP: yes. Its published schema defines
McpRemoteConfig (type/url/headers/oauth) as a sibling of McpLocalConfig, with
headers as a free string→string map. The RFC had been treating "every sampled
config in myconfigs is type:local" as evidence about the schema when it was
only evidence about deployments. Consequence: the edge proxy is not the only
option for opencode, so nothing in the phasing hangs on this. What opencode
still lacks is offline/local-first, which is the honest Phase 2 argument.

Q6, how opencode-devbox learns the opt-in: it already does. generate-config.py,
run from entrypoint-user.sh:117, registers mempalace as remote+bearer when
MEMPALACE_REMOTE_URL is set and local stdio otherwise, and its comment says it
deliberately mirrors the mempalace.ts env contract. The image also ships
mempalace by default, so R5's premise was wrong and is corrected. Phase 2 there
is a third branch in an existing script, not a new mechanism.

One new constraint found while confirming this, and it is the sharpest edge in
the whole opt-in story: generate-config.py never overwrites an existing config
and ~/.config/opencode is a named volume, so flipping the .env on a container
that already has a config is a no-op — it only writes an opencode.jsonc.proposed
sidecar. pi re-reads env every start; opencode does not. Documented in §4.1 and
§9.6, and it applies symmetrically to R6 reversibility.
2026-08-09 22:42:00 +02:00
Joakim Persson 052dbb8038 docs(rfc-001): provenance belongs to the sync boundary, not the agent
Reverses the previous §7.3 on review. It said "stamp added_by everywhere,
now, because it cannot be backfilled" and was about to become a skill
instruction telling agents to do it. Both halves were wrong.

Wrong on ownership: provenance answers "which device asserted this?", so
only a party that can verify the answer should write it. An agent must shell
out to read env, can forget, and will improvise when the values are absent —
the worst possible stamper, and its claim is unverifiable by anyone. Under
the §4 design every write reaches the primary over an authenticated channel,
including offline ones at outbox-flush time, so the primary can stamp the
complete set with no client cooperation. Provenance is a property of the sync
channel, not of the record's author. §7.3.2 adds the trust ladder; edge-side
stamping is demoted to an advisory interim, because 3.6.0's serve takes a
single shared bearer token (mcp_server.py:5291-5293) and the package has zero
device/origin concept, so authoritative stamping needs the per-device
credentials of §6 — Phase 4, not Phase 0.

Wrong on backfill: a solitary devbox is a single-origin store by definition,
so origin is a property of the whole palace and can be assigned wholesale at
import (one --origin-device flag) at the moment it stops being solitary. Bulk
attribution is strictly more reliable than per-record stamping since it
cannot be partially applied. Per-record provenance is only needed where
origins interleave, which is only the primary. Solitary containers therefore
stamp nothing and lose nothing — more R1-compliant than the previous draft,
which quietly asked users who had opted out to carry metadata for the
feature. Multi-harness-on-one-host stays solved by added_by = agent name.

Keeps the verified mechanics (fixed metadata schema, argument whitelisting at
mcp_server.py:4777 silently dropping unknown fields, the diary agent_name →
wing_pi@host trap, kg_add having no provenance slot, added_by absent from
search results) and the identity findings (a container cannot discover its
host's identity; hostnames are neither unique nor stable; rename splits one
device's history in two). §7.3.5 keeps the fail-closed rule for whichever
component does stamp.

Phase 0 drops its provenance item accordingly, and §6.2 now specifies
server-side stamping from the authenticated credential.

Also flags in-document that these notes are a poisoning vector: a future
agent reading them out of the palace must not conclude it should hand-stamp.
2026-08-09 15:40:25 +02:00
joakimp 661ee20b39 docs(rfc-001): require solitary-first operation, opt-in centralization
Adds §1.1 (R1–R6) and §1.2 as a hard constraint on the design rather than a
preference. A shared palace is valuable only for the multi-machine /
multi-container pattern; for most users of the published pi-devbox and
opencode-devbox images it is useless overhead. Evidence: all three sampled
opencode-devbox deployments contain zero mempalace references.

Requirements: solitary operation stays the default and stays byte-identical
(no extra process, no outbox, no network calls); opt-in via docker-compose.yml
+ .env only, never an image rebuild; credentials only in .env, never in
compose, the image, a command line, or a log; no new required services (the
primary stays a separate standalone compose project); degrade-not-fail where
mempalace is absent; fully reversible.

§1.2 extends the convention pi-devbox already ships (.env.example:12-23,
"local by default" with commented MEMPALACE_REMOTE_URL/_TOKEN) into a
three-state ladder — local stdio (default, unchanged) / direct remote
(unchanged) / edge (new, MEMPALACE_EDGE=1) — instead of inventing a new
mechanism. Notes that the devbox-palace volume coupling reverses under edge
mode: the local palace holds the outbox, so persisting it becomes required
rather than irrelevant.

Consequence recorded in §4.1: mempalace-edge must be selected at registration
time, not left always-in-path to decide by env at runtime, since that would
insert a process and a failure mode into every solitary user's setup. pi
branches in createClient(); opencode needs its static MCP JSON templated at
container start, which is new open question §9.6.

Also: adds an R1 acceptance test, marks pi-devbox/.env.example:21 as stale
(advertises mempalace-mcp --transport http rather than mempalace serve
--token/--tls-cert), and extends the evidence index.
2026-08-09 15:02:51 +02:00
joakimp 35b1e3d81d docs: add RFC 001 — global palace with local fallback (mempalace-edge)
Design for moving from one MemPalace per machine per harness to a single
primary palace with per-machine offline fallback, plus the source-verified
archaeology behind it.

Key decisions recorded:

- Sync operations (MCP tool calls), not databases. Embeddings are computed
  client-side and are not portable across architectures; the KG `triples`
  table has no UNIQUE(subject,predicate,object,valid_from) and its ids embed
  datetime.now(), so row copies duplicate facts. Replaying tool calls is
  idempotent where it matters.
- Implement as `mempalace-edge`, a local stdio MCP proxy (child mempalace-mcp
  + HTTPS to the primary + outbox.sqlite), not as per-harness patches. Needs
  zero mempalace internals, so it serves pi, opencode and the CLI alike and
  survives mempalace upgrades.
- Merged reads (query both, re-sort, dedupe by drawer id) give fleet-wide
  recall without replication — which is why Phase 3 (pull replication) is
  deferred: "own writes plus whatever it can reach" is good enough.
- Per-wing replication policy: curated/content-addressed wings replicate,
  mined code/docs stay local (derived, re-mineable, path-dependent ids).

Also documents two silently destructive footguns to avoid during rollout:
`--palace` vs MEMPALACE_PALACE_PATH (the KG follows the flag only, so
`mempalace serve` can start with a silently empty knowledge graph), and
`mempalace sync`, which is gitignore-aware drawer deletion rather than
replication and would wipe fleet memory when run from a host lacking the
repos. Notes that mempalace 3.6.0's `serve` already ships token auth + TLS,
making the "unauthenticated, front it with a proxy" notes elsewhere stale.

Includes an evidence index mapping each claim to file:line in mempalace
3.6.0, and four upstream candidates (origin_host provenance, per-wing ACL,
mempalace_kg_supersede missing from service.py WRITE_TOOLS, and a
sync --refuse-shared guard).
2026-08-09 00:30:00 +02:00
pi 96699f2a17 feat(pi-bridge): external MemPalace transport via MEMPALACE_REMOTE_URL
Let the pi<->mempalace bridge connect to a shared MemPalace over HTTP instead
of always spawning a local mempalace-mcp:
- Extract IMcpClient; rename McpClient -> StdioMcpClient (ctor command, arg-less start()).
- Add RemoteMcpClient (vendored from pi-extensions/mcp-loader.ts): streamable-HTTP
  with AbortController timeouts, protocolVersion pinned 2024-11-05, alive/ensureAlive.
  mempalace-mcp --transport http is sessionless JSON-RPC today; session/SSE/404
  branches retained for a future streamable-HTTP server.
- createClient() selects transport from MEMPALACE_REMOTE_URL; MEMPALACE_REMOTE_TOKEN
  -> Authorization: Bearer. Lifecycle automation (wake-up, /mempalace-diary) unchanged.
- scripts/check-mcp-client-sync.sh: drift guard vs canonical mcp-loader.ts.
- README: document local-vs-external transport.

Typechecks clean (strict); both transports smoke-tested against live mempalace-mcp.
2026-07-02 13:09:29 +02:00
pi e12b624cf7 feat(pi-ext): self-healing respawn + scoped init timeout for mempalace-mcp
A stall-kill (or any crash) of mempalace-mcp was a permanent latch:
available flipped off and stayed off until pi restart. Now the next tool
call transparently respawns the server and retries.

- ensureAlive(): bounded respawn with capped exponential backoff
  (MEMPALACE_MCP_MAX_RESPAWNS, default 2; MEMPALACE_MCP_RESPAWN_BACKOFF_MS,
  default 1000). Respawn budget resets on any successful JSON-RPC response,
  so a recovered server regains full patience while a persistently-broken
  one hits the cap and stays down (no hot-loop).
- Init timeout default raised 120000 -> 300000 (scoped to init only): a
  genuine virtiofs cold-open shouldn't be killed mid-progress only to
  respawn and re-pay the same cost. Per-call timeout stays 60000.
- Concurrency hardening: generation counter so a late exit from a killed
  old process can't tear down a fresh respawn; explicit healthy flag
  replaces racy proc!=null liveness check.
- README: document self-heal, new env vars, and why generous-init +
  bounded-respawn compose rather than overlap.
2026-06-26 00:22:21 +02:00
joakimp a3b8829991 feat(pi-ext): per-request timeout + stall-kill for mempalace-mcp
A wedged mempalace-mcp (classically an OrbStack virtiofs cold-open of a
large chroma.sqlite3 / HNSW load) left the awaiting JSON-RPC promise
pending forever, freezing the pi TUI uninterruptibly: ESC cancels the
LLM stream, not a pending tool execute().

The JSON-RPC client now arms a per-request timer. On expiry it rejects
the request AND kills the stalled child (SIGTERM->SIGKILL), so pi gets
an error instead of hanging; the extension flips available=false so
later calls fail fast (restart pi to retry). Per-REQUEST, not
per-process: the long-lived server only dies on a genuine stall.

Knobs: MEMPALACE_MCP_TIMEOUT_MS (default 60000),
MEMPALACE_MCP_INIT_TIMEOUT_MS (default 120000), 0 = disable.

This supersedes the planned standalone stdio-watchdog shim: the
extension already owns request/response correlation, so a separate
framing-reparsing shim is unnecessary.
2026-06-13 23:48:33 +02:00
joakimp ce09d25c97 Rename to @earendil-works/pi-coding-agent + earendil-works/pi URL
Pi moved to its new home at earendil-works on 2026-05-07
(https://pi.dev/news/2026/5/7/pi-has-a-new-home).

Sweep:
- extensions/pi/mempalace.ts: 'import type { ExtensionAPI } from
  "@mariozechner/pi-coding-agent"' -> @earendil-works/pi-coding-agent.
- README and extensions/pi/README: github.com/mariozechner/pi-coding-agent
  URL refs -> github.com/earendil-works/pi.
- install.sh: same URL substitution in the user-facing pointer line.

Brew install references (`brew install pi-coding-agent`) left as-is:
formula still works at 0.73.1, tap update tracked upstream at
earendil-works/pi#2755.
2026-05-09 17:56:46 +02:00
joakimp 90e70fff61 docs: add Ecosystem diagram to README + update harness-extension conventions post-split
Two small doc updates consolidating the session's architectural arc:

1. README.md gets a new 'Ecosystem' section (after the contents list,
   before 'Why this exists') showing the five-repo composition:
     myconfigs -> opencode-toolkit + pi-toolkit -> mempalace-toolkit
   Plus an ownership table clarifying which scope lives where. The
   diagram makes the opt-out pattern visible \u2014 opencode-devbox's slim
   container path skips mempalace-toolkit and still gets a functional
   stack.

2. AGENTS.md 'Adding a new harness extension' section was still written
   for the pre-split extensions/pi/ which had keybindings.json +
   settings.example.json + pi-env.zsh. Rewrote to reflect:
   - Bridge-only scope (harness-generic config goes in <harness>-toolkit).
   - 'Probe for the sibling toolkit' step replaces the old symlink-keybindings
     and template-settings steps.
   - Worked example now points at install_pi_extension + check_pi_toolkit
     rather than the four functions that moved out.
   - Explicitly names the pattern: opencode-toolkit + pi-toolkit as
     the two existing examples of the sibling-toolkit convention.
2026-05-05 17:50:13 +02:00
joakimp 16915f0e55 refactor: split pi-generic config into pi-toolkit repo
Parallel to the opencode-toolkit split earlier today. Pi's own config
(keybindings, shell env loader, settings template) moves to a new
sibling repo so opencode-devbox's mempalace opt-out can build slim
containers that include pi without dragging in chromadb + embedding
models (~300 MB).

What moved to pi-toolkit (https://gitea.jordbo.se/joakimp/pi-toolkit):
- extensions/pi/keybindings.json          (mosh/tmux newline fix)
- extensions/pi/pi-env.zsh                (sources ~/.config/pi/.env)
- extensions/pi/settings.example.json     (Bedrock template)
- install.sh::install_pi_keybindings      (symlink step)
- install.sh::install_pi_env_loader       (cp step + bash fallback)
- install.sh::check_pi_settings           (probe)
- install.sh::check_aws_env               (probe)

What stays here (this is the pi\u2194mempalace bridge, mempalace-side):
- extensions/pi/mempalace.ts              (the MCP extension)
- install.sh::install_pi_extension        (symlink step)
- NEW: install.sh::check_pi_toolkit       (probe: warns if pi is
                                           installed but pi-toolkit's
                                           artifacts are missing, with
                                           git-clone pointer)

install.sh shrank from 520 to 403 lines. Uninstall mirror correctly
does NOT touch pi-toolkit-owned files (explicit comment).

Docs updated:
- extensions/pi/README.md: rewritten as 'pi\u2194MemPalace MCP bridge',
  recipe becomes 'Deploying pi with mempalace' (pi-toolkit step 3,
  this repo step 5).
- AGENTS.md: Structure block + 'What install.sh does' section reflect
  the narrower scope and list the four things that moved out.
- README.md: repo-contents line + Setup section's deploy summary.

Verified on tor-ms22: full install\u2192uninstall\u2192reinstall lifecycle clean.
After mempalace-toolkit uninstall, pi-toolkit artifacts
(~/.pi/agent/keybindings.json, ~/.oh-my-zsh/custom/pi-env.zsh) remain
intact \u2014 correctly untouched. check_pi_toolkit probe fires green when
both exist.
2026-05-05 17:26:53 +02:00
joakimp 3d3a0fb125 docs(pi): add opencode-toolkit pointer in deploy recipe step 6
Cross-reference the newly-extracted opencode-toolkit repo, which owns
~/.config/opencode/.env loading. The recipe now distinguishes between
registering mempalace in opencode.json (still this repo's concern) and
ensuring opencode's own env loader is in place (opencode-toolkit's).
2026-05-05 17:14:32 +02:00
joakimp 118bd20fec feat(extensions/pi): ship pi-env.zsh shell loader
The loader that sources ~/.config/pi/.env into every shell was only
living in the myconfigs tor-ms22 backup \u2014 a fresh machine had nowhere
to get it from except copying by hand. Now canonical here.

- extensions/pi/pi-env.zsh: 20-line POSIX-compatible loader
  (set -a; source ~/.config/pi/.env; set +a). Works in bash and zsh.
- install.sh install_pi_env_loader:
  * oh-my-zsh detected (~/.oh-my-zsh/custom/ exists)
    \u2192 cp into that dir (NOT symlink \u2014 that dir is typically part of
      a dotfiles backup, and a symlink to mempalace-toolkit would
      break when restored on another host).
    \u2192 Idempotent: if target content matches repo, says 'already
      installed'. If it differs, leaves user edits alone and points
      at diff for manual reconcile.
  * No oh-my-zsh \u2192 prints source-this-line snippet for ~/.zshrc or
    ~/.bashrc (derived from $SHELL). Does NOT auto-edit rc files.
- install.sh uninstall: only removes the copy if content still matches
  repo. Local edits preserved.
- Docs:
  * extensions/pi/README.md Environment setup section rewritten with
    both install paths, step 5 of deploy recipe updated.
  * AGENTS.md Structure block lists pi-env.zsh.
  * Root README repo-contents line mentions it.

Verified on tor-ms22: install fresh \u2192 uninstall (content match \u2192 remove)
\u2192 reinstall \u2192 zsh -ic loads AWS vars correctly. Also tested bash fallback
path via HOME=/tmp/fake-home SHELL=/bin/bash \u2014 prints right .bashrc snippet.
2026-05-05 16:58:56 +02:00
joakimp 5d8f523cb3 docs(AGENTS): unstale wrapper count + new 'Adding a new harness extension' section
Two stale lines from the 2-wrapper era (we have three now), plus a
missing convention doc for future harness extensions:

- 'What this is': two wrappers \u2192 three + extensions/ tree
- 'Adding a new wrapper': drop the 'third wrapper triggers helper lib'
  claim since three coexist fine as standalone scripts; keep the
  threshold at four, re-evaluate then.
- New 'Adding a new harness extension' section codifies the extensions/
  pattern so a future claude-code or kiro bridge follows the same shape
  (per-harness dir, symlink-vs-template rules, gated install steps,
  probe-don't-halt, uninstall mirror, README + Structure updates).
2026-05-05 15:29:15 +02:00
joakimp 71c335148a docs(pi): full 'new machine' deploy recipe in extensions/pi/README
Consolidates the step-by-step recipe that's been living in diary entries
and session chat into the canonical pi bring-up doc. Covers:

  0. Prerequisites (zsh+oh-my-zsh, uv, tmux 3.2+, AWS creds)
  1. Dotfiles: myconfigs provision (tmux CSI-u, ~/.config/pi/.env, zsh loader)
  2. pi install (upstream brew/npm)
  3. mempalace CLI (uv tool install) + mempalace-toolkit install.sh
  4. pi settings bootstrap (start without --model, region prefix table)
  5. AWS env verification (git-crypt unlock gotcha)
  6. Opencode MCP registration pointer (if applicable)
  7. First run + wake-up injection smoke test
  + Verification checklist + uninstall

Root README.md adds a short summary box in the Setup section pointing at
the full recipe, so readers coming in from the front door find the pi
path immediately but the details stay with the files they install.

Covers: macOS + Linux. Works for homelab / work-macos / any myconfigs
profile that ships .config/pi/ + pi-env.zsh.
2026-05-05 15:20:47 +02:00
joakimp 79e0692dac docs: update AGENTS.md structure + ARCHITECTURE.md see-also for pi bring-up
AGENTS.md Structure block was stale (listed only 2 bin/ wrappers, no
extensions/, no contrib/). Added full tree + a new 'What install.sh does'
section enumerating all steps, gates, and probes so maintainers see at a
glance what the installer touches.

ARCHITECTURE.md is scoped to the producer side (feeding the palace);
pi extension is consumer side, so out of scope for the main body. Added
a pointer in the See also section so readers can find extensions/pi/README.md.
2026-05-05 14:47:06 +02:00
joakimp 75876e5c41 fix(install): silence AWS probe when pi settings.json absent
check_aws_env was warning about missing AWS_PROFILE/AWS_REGION even on
fresh machines with no settings.json yet \u2014 but at that point we don't
know which provider the user will pick, so the warning is noise.
check_pi_settings already tells the user to bootstrap settings.json;
the AWS probe now stays quiet until it has evidence (amazon-bedrock in
settings.json) that AWS creds are actually needed.
2026-05-05 14:44:49 +02:00
joakimp 854ae41f65 feat(extensions/pi): keybindings symlink + settings template + AWS/pi probes
Round out the pi bring-up story so a fresh machine can reach a working
pi+mempalace install with just `git clone && ./install.sh`:

- extensions/pi/keybindings.json: generic mosh/tmux newline fix
  (shift+enter, ctrl+j, alt+j). Safe on any machine — not
  region/account-specific. Symlinked into ~/.pi/agent/.
- extensions/pi/settings.example.json: template for `settings.json`
  so pi can start without --provider/--model. NOT symlinked — pi
  rewrites settings.json at runtime (lastChangelogVersion bumps),
  which would dirty the repo. Installer prints the cp + edit hint.
- install.sh: new install_pi_keybindings + uninstall mirror; new
  check_pi_settings probe (warns if settings.json missing); new
  check_aws_env probe (warns if AWS_PROFILE/AWS_REGION unset and
  settings.json selects amazon-bedrock). All new steps gated on
  pi being installed (~/.pi/agent/extensions/ exists).
- extensions/pi/README.md: documents keybindings rationale,
  settings bootstrap, and the recommended ~/.config/pi/.env +
  ~/.oh-my-zsh/custom/pi-env.zsh env layout (paired with the
  myconfigs commit 884e329 that split AWS vars out of
  ~/.config/opencode/.env).

Verified on tor-ms22: full install → uninstall → reinstall cycle,
new shell loads AWS_PROFILE/AWS_REGION from the new pi-env.zsh hook.

Works on macOS and Linux (plain ln -s, POSIX bash).
2026-05-05 13:59:20 +02:00
joakimp ef1d022fbc feat(extensions): version-control pi mempalace extension + install.sh symlink
The pi coding-agent extension at ~/.pi/agent/extensions/mempalace.ts was
living only on tor-ms22, including hand-edited fixes (Type.Unsafe
schema-passthrough for MCP tool parameters). One disk wipe away from
losing it, and no way to reproduce the install on a new machine.

- extensions/pi/mempalace.ts: canonical copy (matches tor-ms22 byte-for-byte)
- extensions/pi/README.md: what it does, the schema-passthrough gotcha,
  debugging knobs
- install.sh: new install_pi_extension step — gated on ~/.pi/agent/extensions/
  existing, backs up any real file in the way, idempotent re-runs, mirror
  block in uninstall. Works on macOS and Linux (plain ln -s, readlink -f).
- README.md: mention extensions/pi/ in the repo-contents list and in the
  Setup section

Verified on tor-ms22: install (backs up existing real file) → uninstall
(removes symlink) → reinstall (clean symlink). Re-runs are no-ops.
2026-05-05 13:42:47 +02:00
joakimp 6352373a1f fix(feeders): make post-mine repair opt-in, not default
The three feeder wrappers (mempalace-docs, mempalace-pi-session,
mempalace-session) unconditionally ran 'mempalace repair --yes' after
mining, controllable only via --no-repair opt-out. The contrib launchd
and systemd templates did not pass --no-repair, so every scheduled tick
invoked the destructive in-place HNSW rebuild.

This has bitten us twice:
  - 2026-05-04 09:08: a kickstart triggered repair while an MCP
    subprocess held the DB open; the live collection was wiped (0
    drawers) and had to be restored from the palace.backup snapshot.
  - 2026-05-05 10:00: post-mine repair crashed mid-rebuild with
    'NotFoundError: Collection [<uuid>] does not exist' - chromadb's
    rebuild recreated the collection under a new UUID while the code
    still held the old handle. Live DB survived only by luck (crash
    hit before the swap).

Fix: flip the default.
  - New flag: --repair (opt-in). Prints a warning and sleeps 3s before
    invoking 'mempalace repair --yes'.
  - --no-repair is retained as a deprecated no-op alias for backward
    compatibility with any scripts/units still passing it.
  - Default behavior: no repair. Routine ChromaDB add() keeps HNSW
    consistent; repair is a recovery op, not a maintenance tick.

Docs updated to match: README, SKILL, ARCHITECTURE, AGENTS,
contrib/README. Scheduling guidance now explicitly warns against
enabling --repair on cron/launchd/systemd-timer runs.
2026-05-05 12:35:04 +02:00
joakimp 53d96adc65 docs(contrib): scheduling templates for mempalace-pi-session
Drop-in equivalents of the opencode templates for each scheduler
mechanism:

  systemd/mempalace-pi-session.{service,timer}
  launchd/se.jordbo.mempalace-pi-session.plist
  cron/mempalace-pi-session.cron

Schedule is staggered from the opencode jobs (Mon 03:00 -> Tue 03:00)
so machines running both don't race each other on the post-mine HNSW
repair step. Service unit uses ConditionPathExists=%h/.pi/agent/sessions
to no-op silently on machines that haven't used pi, matching the
opencode template's guard on ~/.local/share/opencode/opencode.db.

contrib/README.md grows a 'Templates at a glance' table so the set is
discoverable without reading the whole doc.
2026-05-05 08:48:33 +02:00
joakimp 14d253f929 feat(session): tag opencode staging headers with '| source: opencode'
Complement to the mempalace-pi-session feeder: now that a second source
mines into wing_conversations, every session's synthetic header carries
an explicit source tag so the LLM can discriminate at read time when
searches return first-chunk content:

  [session: <title> | <directory> | <YYYY-MM-DD> | source: opencode]

The primary disambiguator in search results remains source_file basename
(opencode: '<slug>_<id>.jsonl', pi: 'pi_<uuid>.jsonl'), which is present
in every chunk's metadata regardless of where the search hit landed in
the session. This header tag is a cosmetic second signal on first-chunk
hits.

Caveat: existing drawers keep their old header — mempalace mine dedups
by source_file path, which didn't change, so old opencode sessions are
not re-mined. They are implicitly opencode (the only pre-pi source).
2026-05-05 08:48:27 +02:00
joakimp 9450a45194 feat(pi-session): add mempalace-pi-session feeder for pi coding-agent sessions
Parallel to mempalace-session, this wrapper walks ~/.pi/agent/sessions/
JSONL files and mines qualifying sessions into wing_conversations via
'mempalace mine --mode convos'.

Design choices mirror mempalace-session:
- Export-stage-mine idiom with deterministic per-session staging paths
  under ~/.cache/mempalace-pi-session/<wing>/, so 'mempalace mine' dedup
  on source_file makes re-runs idempotent.
- --dry-run classifies each export as [NEW] or [SKIP] by matching staging
  path against the palace's already-filed source_files.
- --min-messages filter skips throwaway single-prompt sessions.

Pi-specific parsing:
- Pi JSONL is a typed tree (id/parentId) per docs/session-format.md;
  this walks in file order, which is correct for the overwhelmingly
  linear case and harmlessly duplicative on branched sessions (palace
  semantic dedup handles it).
- Roles mapped to Claude Code JSONL shape:
    user      -> {type:user, content:text}
    assistant -> {type:assistant, content:[text, tool_use]}
    toolResult-> {type:human, content:[tool_result]} (folded back by normalizer)
    bashExecution/custom(display)/branchSummary/compactionSummary
              -> rendered as text annotations
- thinking blocks and image blocks dropped (noise / palace is text-only).

Source labelling:
- Staging filenames prefixed 'pi_<uuid>.jsonl' so every drawer's
  source_file metadata (visible in search results) unambiguously
  identifies the harness. Opencode's convention ('<slug>_<id>.jsonl')
  is preserved to keep the existing 19k+ drawers deduped.
- Inline synthetic header on first chunk:
    [session: <title> | <cwd> | <date> | source: pi]
  as a secondary signal.
2026-05-05 08:48:20 +02:00
Joakim Persson 98baabe7a0 contrib: flag cron-not-installed as a common caveat
Minimal Debian/Ubuntu hosts (and most base container images) don't
ship cron by default. `crontab: command not found` is the first
thing a user hits if they try the cron path without installing it.
Previous caveats block covered semantics (no Persistent, mail-drop
stderr) but silently assumed cron was present. Add an explicit
"check command -v crontab, apt install cron, or pick systemd"
preflight to the caveats so the error is surfaced before the
user runs into it.

Caught during 2026-04-30 Phase 4 runtime validation on a Debian
trixie host: `crontab -T` lint failed because cron wasn't
installed, even though the underlying docker-exec shell command
(the actual workload) ran fine.
2026-04-30 21:00:53 +00:00