463 lines with 64 mentions of specific hosts, in a public repo: the primary and tunnel hosts by name, the registrar/DNS step, the tunnel resource wiring, shared token custody, per-machine flip dates, and a palace lineage naming three work machines. Now in the private fleet repo (fleet-ops 093fb65); this file becomes a moved-note in the shape docs/synlig-primary-runbook.md already established. Git history keeps the old text, so this limits future exposure rather than undoing it. Unlike the primary-host runbook, this file was MIXED — and the stub says so instead of quietly implying the toolkit still documents HTTP exposure. §1.1-§1.3 (why one shared fleet token rather than per-device proxy users, and the one place per-device identity does exist), §2 (the bind trap and the Host/Origin pin) and §3.7 (a client is flipped by three env vars that travel as a set) are reusable mechanism now published nowhere else. Named in the stub so the extraction is tracked debt rather than a silent loss, and named in the private copy too so whoever extracts it can delete the duplicate. Inbound references fixed rather than left pointing at content that moved: two §3.8 pointers in extensions/pi/README.md are replaced by the instruction they were pointing at (the palace path reported by mempalace_status must be the remote host's — a half-flipped client looks healthy while reporting a local path), and the bind-trap reference is replaced by the Host/Origin sentence itself, so the extension README no longer depends on the moved file. contrib also loses a hostname and a seeding date it did not need to make its point.
16 KiB
Fleet memory — what MemPalace stores, and what to put where
Audience: the person running MemPalace on more than one machine, or thinking about it.
Companion documents: docs/rfc-003-coordination-log.md for the coordination log's mechanism, docs/rfc-001-global-palace.md for the centralisation design, and ~/.agents/skills/mempalace/SKILL.md for what the agents are told to do.
MemPalace is usually described as "memory for agents", which is true and not very actionable. It is really five stores with different retrieval models, and most of the value — especially across a fleet — comes from putting each kind of thing in the store whose retrieval model matches how you will want it back.
This document covers: what the stores are, what a central palace changes when several machines share one, and how to decide between filing a memory and sending a message.
1. The stores at a glance
flowchart LR
Q["ask by MEANING<br/>'what do we know about X?'"] --> D["<b>Drawers</b><br/>chroma.sqlite3<br/>verbatim text, embedded"]
R["ask by ENTITY + TIME<br/>'what was true in June?'"] --> G["<b>Knowledge graph</b><br/>knowledge_graph.sqlite3<br/>typed facts, time-bounded"]
S["ask by ADDRESS + ORDER<br/>'what is waiting for me?'"] --> L["<b>Coordination log</b><br/>logstream.sqlite3<br/>addressed events + artifacts"]
T["ask by ASSOCIATION<br/>'what else touches this?'"] --> P["<b>Palace graph</b><br/>tunnels.json + hallways.json<br/>links between rooms"]
D --> DI["<b>Diaries</b> live here too:<br/>drawers with room=diary,<br/>read by recency, not similarity"]
All four stores are files inside one palace directory, so "the palace" is a directory you can back up in one go.
| Store | You get things back by | Typical use | Wrong use |
|---|---|---|---|
| Drawers (wings → rooms) | semantic similarity | a verbatim finding, a decision and its reasoning, a runbook, a transcript excerpt | anything a specific machine must act on; anything whose value is its exact byte content |
Diaries (drawers with room="diary") |
agent + recency | "what did I do last session, and what did it feel like" — the continuity thread across sessions | facts other agents need to find by searching; a diary is read by its author, chronologically |
| Knowledge graph | entity, relationship, point in time | facts that change: versions, employers, who owns what, an injury that heals | prose, reasoning, anything you'd want to read rather than query |
| Coordination log | address, correlation, append order | "device B must review this patch"; "this claim is retracted, stop building on it" | durable knowledge — an event is invisible to semantic search |
| Palace graph | traversal from a room | discovering that an API design in one project touches a schema in another | primary storage — it links drawers, it does not hold content |
Two smaller files exist and are implementation detail, not user surface: sqlite_exact.sqlite3 (an exact-match index over drawer metadata) and, if the daemon runs, queue.sqlite3 (its job queue).
1.1 Drawers: wings and rooms
A wing is a project or domain; a room is an aspect within it. wing="pi-devbox", room="landmines" is a good pair; wing="misc", room="stuff" is how a palace becomes a landfill. Content is stored verbatim and chunked — never summarised — and retrieved by embedding similarity, so a drawer is found by someone who doesn't already know it exists. That is the property to optimise for: write the drawer that the next person's search will match.
The one counter-intuitive consequence: fresh drawers rank worst. A drawer filed an hour ago has no advantage in a similarity search, and a well-worn older drawer will outrank it. For anything recent, enumerate by date (list_drawers(since=…)) instead of searching.
1.2 Diaries
A diary entry is a drawer with room="diary", filed by default into wing_<agent>, tagged with the writing agent. It is the first-person record: what I did, what surprised me, what I would do differently. Agents are told to write one before a session ends, and to read the last few at session start.
In a fleet this is the highest-signal store per byte, for a reason that is easy to miss: a diary entry is the only place that records what did not work. A drawer tends to record the conclusion; the diary records the three hours that produced it.
1.3 Knowledge graph
Triples — subject, predicate, object — with valid_from / valid_to, so a fact can stop being true without being deleted. supersede replaces a single-valued fact at a shared boundary, so a point-in-time query at that instant returns exactly one value.
Use it for anything you will later want to ask "what was true at time T?" about: which version was released when, who owned a service, what model an assistant was using. Do not use it for prose — a triple whose object is a paragraph is a drawer wearing a costume.
1.4 The coordination log
Addressed, ordered, exact. mempalace_event_* carries messages between agents; mempalace_artifact_* carries byte-exact payloads (patches, logs, files) that events can reference. This is the only store where one machine can reach another.
Its full mechanism, limits and landmines are in docs/rfc-003-coordination-log.md. §4 below covers what an operator needs to decide.
2. What changes when a fleet shares one palace
A single machine's palace is a notebook. A shared palace is something different in kind: the fleet stops being a set of independent agents that each learn the same lessons separately.
flowchart LR
A["laptop<br/>agent session"] -->|MCP over HTTPS| H(("central palace<br/>one server"))
B["workstation<br/>agent session"] -->|MCP over HTTPS| H
C["build box<br/>agent session"] -->|MCP over HTTPS| H
H --> D["drawers + diaries"]
H --> G["knowledge graph"]
H --> L["coordination log"]
Note the topology: in the common deployment the machines are thin clients of one server, not peer replicas. Everything one machine writes is immediately visible to the others — there is no sync delay to reason about, and equally no local copy to fall back on when the server is unreachable. (docs/rfc-001-global-palace.md §4 designs an edge proxy with a local palace and a durable outbox for deployments that need to keep working offline; the plain thin-client shape above does not.)
2.1 The three things this actually buys
Awareness. "What has anyone been doing?" becomes answerable. Each machine's diary is readable by every other machine, so an agent starting work on a shared project can see that another machine spent yesterday on it, and how far it got.
Non-repetition of expensive work. This is the biggest measurable win. Anything that cost real time to obtain — a scraped API surface, a spec read end to end, a bisect, a benchmark, a long investigation into why a build fails on one platform — is filed once and searchable everywhere. The second machine's cost drops from hours to one search.
Mistakes and retractions travel. The subtle one, and the reason a shared palace is worth more than a shared wiki. When a machine discovers that a belief was wrong, it can file the retraction where every other machine will hit it. Without that, each machine independently rediscovers the same dead end — and worse, a machine can spend a day rebuilding something another machine already proved doesn't work.
sequenceDiagram
participant W as workstation
participant P as central palace
participant L as laptop, asleep 9 days
W->>W: spends 3h finding why the build breaks
W->>P: drawer — the finding, verbatim, with evidence
W->>P: KG fact — "component X requires flag Y"
W->>P: diary — what was tried and failed
Note over L: ...9 days pass, laptop asleep...
L->>P: wake-up: read diaries + search before starting
P-->>L: the finding, the failed attempts, the fact
Note over L: cost: one search instead of 3 hours
2.2 The two costs, stated plainly
Everything is visible to everyone. One shared token, no per-agent read scoping. Anything filed into a shared palace should be considered readable by every machine and every agent on it. Do not put secrets in drawers.
Provenance stops being obvious. On a single machine, every drawer is yours and every path exists. On a shared palace, most drawers came from other machines and most source_file paths do not exist locally. Two consequences worth internalising:
- A file path in a drawer is evidence about some machine, not necessarily this one.
- ⚠️ Never run
mempalace syncagainst a shared palace. It prunes drawers whose source files look gitignored, deleted or moved — which on a shared palace describes most of the content, including every other machine's. Seedocs/rfc-001-global-palace.md§7.2. The coordination log is not affected by this (RFC 003 §2), but drawers very much are.
3. Deciding where something goes
The question that matters is not "is this important?" but "how will I want this back, and does anyone need to act?"
flowchart TD
Start["I have something worth keeping"] --> Act{"Must a specific<br/>machine or agent<br/>DO something?"}
Act -->|no| Change{"Is it a fact that<br/>changes over time?"}
Act -->|yes| Know{"Do they also need<br/>to KNOW it later?"}
Change -->|yes| KG["Knowledge graph<br/>kg_add / kg_supersede"]
Change -->|no| Mine{"Is it about MY session<br/>— what I tried, felt, learned?"}
Mine -->|yes| Diary["Diary entry"]
Mine -->|no| Drawer["Drawer<br/>wing + room, verbatim"]
Know -->|yes| Both["BOTH:<br/>drawer for the knowledge,<br/>event pointing at it"]
Know -->|no| Event["Coordination event<br/>to_agent = the specific agent"]
Worked examples:
| Situation | Where | Why |
|---|---|---|
| "The release takes 76 min, and 40 of those are the base image build." | Drawer | durable, nobody must act, next person finds it by searching "release timing" |
| "v1.8.9 is the released version, as of this timestamp." | KG (supersede) |
it will change; you will want "what was released in August?" |
| "I spent two hours chasing a watcher that was already dead." | Diary | first-person, chronological, tells the next session what not to retry |
| "Build box: this patch is ready, please review and apply." | Event (directed) | a named machine must act; the patch itself goes in as an artifact |
| "The claim in that drawer is wrong — I measured the opposite." | Both | file the corrected finding as a drawer, then an event so the machine building on it stops |
| "Everyone should know the new toolkit is live." | Drawer + broadcast event | the drawer is what anyone will find; the broadcast is a notice, not an ask (§4.2) |
The failure mode in each direction is worth naming, because both are common:
- A finding filed only as an event is invisible to semantic search. Nobody will ever find it again, and the next agent will re-derive it.
- An ask filed only as a drawer is addressed to nobody. It will be found, if ever, by accident — long after it mattered.
4. What the coordination log can and cannot do
This is the part most likely to be mis-set expectations, so it is worth being blunt: it is a durable log, not a chat. Nothing is listening. Events are appended and persist; there is no delivery window; nothing is lost by being offline when one is written. A message waits indefinitely, and your reply waits just as patiently for a sender who has since gone away.
That sounds like a limitation and is actually the correct design for a fleet where few machines are awake at once and any given machine may sleep for weeks.
4.1 Latency, honestly
| Recipient state | When they see it | Notes |
|---|---|---|
| In a live session, mailbox-enabled client | ≈2–5 minutes | measured ≈2–3 min on first live delivery; a poll floor of 5 min applies between checks |
Holding an SSE connection (GET /logstream/stream) |
sub-second | for daemons/dashboards, not interactive agents |
| Asleep, next session tomorrow | tomorrow | delivered in the session-start wake-up |
| Asleep for three weeks | in three weeks | nothing expires; the log is permanent |
| Never runs again | never | there is no re-routing and no dead-letter path |
So: appropriate for "handle this when you next wake", "here is a patch", "stop building on that claim". Not appropriate for anything with a deadline inside the hour, unless you know the recipient is awake.
One pleasant property, worth knowing because it is counter-intuitive: coordination traffic is exempt from the palace's write lock, so you can message another machine and it can reply while a long mine is running on the server (RFC 003 §4). Coordination stays alive when memory writes are blocked.
4.2 The one thing that does not work: broadcasting an ask
You can write a broadcast (to_agent="*") and every machine that lists events will see it. But a broadcast never enters any machine's mailbox and is never auto-delivered — by design, because "everyone owes this answer" degenerates into either N duplicate replies or nobody acting.
- Broadcast = a notice on a wall. Fine for "v1.8.9 is out".
- Directed event = a message in a named mailbox. Required for anything that must be done.
To reach a whole fleet with something actionable, fan out: one directed event per device, sharing one correlation_id so the thread stays joinable. Each machine then owes its own reply.
flowchart LR
You["you"] -->|"to_agent=pi@laptop"| A["laptop owes a reply"]
You -->|"to_agent=pi@workstation"| B["workstation owes a reply"]
You -->|"to_agent=pi@build-box"| C["build box owes a reply"]
A --> Corr["one shared correlation_id<br/>joins the three threads"]
B --> Corr
C --> Corr
4.3 Addressing, and why the format matters
Addresses are <harness>@<device> — pi@laptop, opencode@build-box. Two rules follow:
- An unstamped client is unreachable. If a machine's events say
from_agent: piwith no device, nobody can address it, because "pi" is every machine. - Two machines must never share one address. Nothing prevents it, nothing warns, and the result is that each silently discards the other's asks as its own (RFC 003 §7.7).
4.4 A reply is owed until it is terminally closed
An acknowledgement does not close a thread. Neither does "claimed" or "ready". Only a terminal status — applied, superseded, failed, blocked — clears an ask from the recipient's mailbox. Until then, a mailbox-enabled client will keep resurfacing it, which is the intended behaviour: an unanswered ask should nag.
5. Habits that make a shared palace work
Small, and the whole value rests on them:
- Search before you answer, and enumerate before you conclude. One empty search is not proof of silence — fresh drawers rank worst, so for anything from the last couple of days list by date and read the other machines' diaries.
- Write the diary entry before the session ends. It is the store other machines learn from fastest, and the only one that records failed attempts.
- File retractions as loudly as findings. "I was wrong about X, here is the measurement" is worth more than a new finding, because it stops N machines repeating a dead end.
- Check the mailbox at wake-up, even when you expect nothing. An empty result costs one call. Silence is only informative once you know delivery works.
- Say which machine you are talking about. On a shared palace, "the container" and "the host" are ambiguous and a path is not self-identifying.
- After a write times out, verify — do not blindly retry. The palace is single-writer for memory writes, so a timeout usually means the write completed. For coordination events this matters twice over: there is no idempotency guard, so a retried event forks the thread into two (RFC 003 §7.1).
6. See also
docs/rfc-003-coordination-log.md— the coordination log: storage, semantics, security model, landmines.docs/rfc-001-global-palace.md— how and why a palace is centralised; §5 what should not be global; §7 the landmines, including thesynchazard.docs/phase-1-exposure-runbook.md— (moved) the site-specific exposure record is now private; the stub names the mechanism still awaiting extraction.docs/backup-and-recovery.md— why a palace needs its own backup procedure.extensions/pi/README.md— the pi-side client: provenance stamping and the auto-delivered mailbox.~/.agents/skills/mempalace/SKILL.md— the protocol the agents themselves follow.