# Fleet memory — what MemPalace stores, and what to put where **Audience:** the person running MemPalace on more than one machine, or thinking about it. **Companion documents:** `docs/rfc-003-coordination-log.md` for the coordination log's mechanism, `docs/rfc-001-global-palace.md` for the centralisation design, and `~/.agents/skills/mempalace/SKILL.md` for what the *agents* are told to do. MemPalace is usually described as "memory for agents", which is true and not very actionable. It is really **five stores with different retrieval models**, and most of the value — especially across a fleet — comes from putting each kind of thing in the store whose retrieval model matches how you will want it back. This document covers: what the stores are, what a central palace changes when several machines share one, and how to decide between filing a memory and sending a message. --- ## 1. The stores at a glance ```mermaid flowchart LR Q["ask by MEANING
'what do we know about X?'"] --> D["Drawers
chroma.sqlite3"] R["ask by ENTITY + TIME
'what was true in June?'"] --> G["Knowledge graph
knowledge_graph.sqlite3"] S["ask by ADDRESS
'what is waiting for me?'"] --> L["Coordination log
logstream.sqlite3"] T["ask by ASSOCIATION
'what else touches this?'"] --> P["Palace graph
tunnels.json + hallways.json"] D --> DI["Diaries live here too:
drawers with room=diary"] ``` All four stores are files inside **one palace directory**, so "the palace" is a directory you can back up in one go. Each holds a different *shape* of thing: drawers hold verbatim text, embedded for similarity; the knowledge graph holds typed facts that are time-bounded; the coordination log holds addressed events and their artifacts, read in append order; the palace graph holds links between rooms. Diaries are drawers, but they are read by recency rather than by similarity. | Store | You get things back by | Typical use | Wrong use | |---|---|---|---| | **Drawers** (wings → rooms) | semantic similarity | a verbatim finding, a decision and its reasoning, a runbook, a transcript excerpt | anything a specific machine must *act* on; anything whose value is its exact byte content | | **Diaries** (drawers with `room="diary"`) | agent + recency | "what did I do last session, and what did it feel like" — the continuity thread across sessions | facts other agents need to find by searching; a diary is read by *its author*, chronologically | | **Knowledge graph** | entity, relationship, point in time | facts that *change*: versions, employers, who owns what, an injury that heals | prose, reasoning, anything you'd want to read rather than query | | **Coordination log** | address, correlation, append order | "device B must review this patch"; "this claim is retracted, stop building on it" | durable knowledge — an event is invisible to semantic search | | **Palace graph** | traversal from a room | discovering that an API design in one project touches a schema in another | primary storage — it links drawers, it does not hold content | Two smaller files exist and are implementation detail, not user surface: `sqlite_exact.sqlite3` (an exact-match index over drawer metadata) and, if the daemon runs, `queue.sqlite3` (its job queue). ### 1.1 Drawers: wings and rooms A **wing** is a project or domain; a **room** is an aspect within it. `wing="pi-devbox", room="landmines"` is a good pair; `wing="misc", room="stuff"` is how a palace becomes a landfill. Content is stored **verbatim and chunked** — never summarised — and retrieved by embedding similarity, so a drawer is found by someone who *doesn't already know it exists*. That is the property to optimise for: write the drawer that the next person's search will match. The one counter-intuitive consequence: **fresh drawers rank worst.** A drawer filed an hour ago has no advantage in a similarity search, and a well-worn older drawer will outrank it. For anything recent, enumerate by date (`list_drawers(since=…)`) instead of searching. ### 1.2 Diaries A diary entry is a drawer with `room="diary"`, filed by default into `wing_`, tagged with the writing agent. It is the *first-person* record: what I did, what surprised me, what I would do differently. Agents are told to write one before a session ends, and to read the last few at session start. In a fleet this is the highest-signal store per byte, for a reason that is easy to miss: a diary entry is the only place that records **what did not work**. A drawer tends to record the conclusion; the diary records the three hours that produced it. ### 1.3 Knowledge graph Triples — subject, predicate, object — with `valid_from` / `valid_to`, so a fact can *stop* being true without being deleted. `supersede` replaces a single-valued fact at a shared boundary, so a point-in-time query at that instant returns exactly one value. Use it for anything you will later want to ask "what was true at time T?" about: which version was released when, who owned a service, what model an assistant was using. Do not use it for prose — a triple whose object is a paragraph is a drawer wearing a costume. ### 1.4 The coordination log Addressed, ordered, exact. `mempalace_event_*` carries messages between agents; `mempalace_artifact_*` carries byte-exact payloads (patches, logs, files) that events can reference. This is the only store where one machine can *reach* another. Its full mechanism, limits and landmines are in `docs/rfc-003-coordination-log.md`. §4 below covers what an operator needs to decide. --- ## 2. What changes when a fleet shares one palace A single machine's palace is a notebook. A shared palace is something different in kind: **the fleet stops being a set of independent agents that each learn the same lessons separately.** ```mermaid flowchart LR A["laptop
agent session"] -->|MCP over HTTPS| H(("central palace
one server")) B["workstation
agent session"] -->|MCP over HTTPS| H C["build box
agent session"] -->|MCP over HTTPS| H H --> D["drawers + diaries"] H --> G["knowledge graph"] H --> L["coordination log"] ``` Note the topology: in the common deployment the machines are **thin clients of one server**, not peer replicas. Everything one machine writes is immediately visible to the others — there is no sync delay to reason about, and equally no local copy to fall back on when the server is unreachable. (`docs/rfc-001-global-palace.md` §4 designs an edge proxy with a local palace and a durable outbox for deployments that need to keep working offline; the plain thin-client shape above does not.) ### 2.1 The three things this actually buys **Awareness.** "What has anyone been doing?" becomes answerable. Each machine's diary is readable by every other machine, so an agent starting work on a shared project can see that another machine spent yesterday on it, and how far it got. **Non-repetition of expensive work.** This is the biggest measurable win. Anything that cost real time to obtain — a scraped API surface, a spec read end to end, a bisect, a benchmark, a long investigation into why a build fails on one platform — is filed once and searchable everywhere. The second machine's cost drops from hours to one search. **Mistakes and retractions travel.** The subtle one, and the reason a shared palace is worth more than a shared wiki. When a machine discovers that a belief was *wrong*, it can file the retraction where every other machine will hit it. Without that, each machine independently rediscovers the same dead end — and worse, a machine can spend a day rebuilding something another machine already proved doesn't work. ```mermaid sequenceDiagram participant W as workstation participant P as central palace participant L as laptop, asleep 9 days W->>W: spends 3h finding why the build breaks W->>P: drawer — the finding, verbatim, with evidence W->>P: KG fact — "component X requires flag Y" W->>P: diary — what was tried and failed Note over L: ...9 days pass, laptop asleep... L->>P: wake-up: read diaries + search before starting P-->>L: the finding, the failed attempts, the fact Note over L: cost: one search instead of 3 hours ``` ### 2.2 The two costs, stated plainly **Everything is visible to everyone.** One shared token, no per-agent read scoping. Anything filed into a shared palace should be considered readable by every machine and every agent on it. Do not put secrets in drawers. **Provenance stops being obvious.** On a single machine, every drawer is yours and every path exists. On a shared palace, most drawers came from other machines and most `source_file` paths **do not exist locally**. Two consequences worth internalising: - A file path in a drawer is evidence about *some* machine, not necessarily this one. - ⚠️ **Never run `mempalace sync` against a shared palace.** It prunes drawers whose source files look gitignored, deleted or moved — which on a shared palace describes most of the content, including every other machine's. See `docs/rfc-001-global-palace.md` §7.2. The coordination log is *not* affected by this (RFC 003 §2), but drawers very much are. --- ## 3. Deciding where something goes The question that matters is not "is this important?" but **"how will I want this back, and does a specific machine or agent need to act?"** ```mermaid flowchart TD Start["I have something worth keeping"] --> Act{"Must someone specific
DO something?"} Act -->|no| Change{"Is it a fact that
changes over time?"} Act -->|yes| Know{"Do they also need
to KNOW it later?"} Change -->|yes| KG["Knowledge graph
kg_add / kg_supersede"] Change -->|no| Mine{"Is it about MY session
— tried, felt, learned?"} Mine -->|yes| Diary["Diary entry"] Mine -->|no| Drawer["Drawer
wing + room, verbatim"] Know -->|yes| Both["BOTH:
drawer + event pointing at it"] Know -->|no| Event["Coordination event
to_agent = that agent"] ``` Worked examples: | Situation | Where | Why | |---|---|---| | "The release takes 76 min, and 40 of those are the base image build." | Drawer | durable, nobody must act, next person finds it by searching "release timing" | | "v1.8.9 is the released version, as of this timestamp." | KG (`supersede`) | it will change; you will want "what was released in August?" | | "I spent two hours chasing a watcher that was already dead." | Diary | first-person, chronological, tells the next session what *not* to retry | | "Build box: this patch is ready, please review and apply." | Event (directed) | a named machine must act; the patch itself goes in as an artifact | | "The claim in that drawer is wrong — I measured the opposite." | Both | file the corrected finding as a drawer, then an event so the machine building on it stops | | "Everyone should know the new toolkit is live." | Drawer + broadcast event | the drawer is what anyone will *find*; the broadcast is a notice, not an ask (§4.2) | The failure mode in each direction is worth naming, because both are common: - **A finding filed only as an event** is invisible to semantic search. Nobody will ever find it again, and the next agent will re-derive it. - **An ask filed only as a drawer** is addressed to nobody. It will be found, if ever, by accident — long after it mattered. --- ## 4. What the coordination log can and cannot do This is the part most likely to be mis-set expectations, so it is worth being blunt: **it is a durable log, not a chat.** Nothing is listening. Events are appended and persist; there is no delivery window; nothing is lost by being offline when one is written. A message waits indefinitely, and your reply waits just as patiently for a sender who has since gone away. That sounds like a limitation and is actually the correct design for a fleet where few machines are awake at once and any given machine may sleep for weeks. ### 4.1 Latency, honestly | Recipient state | When they see it | Notes | |---|---|---| | In a live session, mailbox-enabled client | **≈2–5 minutes** | measured ≈2–3 min on first live delivery; a poll floor of 5 min applies between checks | | Holding an SSE connection (`GET /logstream/stream`) | sub-second | for daemons/dashboards, not interactive agents | | Asleep, next session tomorrow | tomorrow | delivered in the session-start wake-up | | Asleep for three weeks | in three weeks | nothing expires; the log is permanent | | Never runs again | never | there is no re-routing and no dead-letter path | So: appropriate for "handle this when you next wake", "here is a patch", "stop building on that claim". Not appropriate for anything with a deadline inside the hour, unless you know the recipient is awake. One pleasant property, worth knowing because it is counter-intuitive: coordination traffic is **exempt from the palace's write lock**, so you can message another machine and it can reply *while* a long mine is running on the server (RFC 003 §4). Coordination stays alive when memory writes are blocked. ### 4.2 The one thing that does not work: broadcasting an ask You can write a broadcast (`to_agent="*"`) and every machine that *lists* events will see it. But a broadcast **never enters any machine's mailbox** and is never auto-delivered — by design, because "everyone owes this answer" degenerates into either N duplicate replies or nobody acting. - **Broadcast** = a notice on a wall. Fine for "v1.8.9 is out". - **Directed event** = a message in a named mailbox. Required for anything that must be done. To reach a whole fleet with something actionable, **fan out**: one directed event per device, sharing one `correlation_id` so the thread stays joinable. Each machine then owes its own reply. ```mermaid flowchart LR You["you"] -->|"to_agent=pi@laptop"| A["laptop owes a reply"] You -->|"to_agent=pi@workstation"| B["workstation owes a reply"] You -->|"to_agent=pi@build-box"| C["build box owes a reply"] A --> Corr["one shared correlation_id
joins the three threads"] B --> Corr C --> Corr ``` ### 4.3 Addressing, and why the format matters Addresses are `@` — `pi@laptop`, `opencode@build-box`. Two rules follow: 1. **An unstamped client is unreachable.** If a machine's events say `from_agent: pi` with no device, nobody can address it, because "pi" is every machine. 2. **Two machines must never share one address.** Nothing prevents it, nothing warns, and the result is that each silently discards the other's asks as its own (RFC 003 §7.7). ### 4.4 A reply is owed until it is *terminally* closed An acknowledgement does not close a thread. Neither does "claimed" or "ready". Only a terminal status — `applied`, `superseded`, `failed`, `blocked` — clears an ask from the recipient's mailbox. Until then, a mailbox-enabled client will keep resurfacing it, which is the intended behaviour: an unanswered ask should nag. --- ## 5. Habits that make a shared palace work Small, and the whole value rests on them: 1. **Search before you answer, and enumerate before you conclude.** One empty search is not proof of silence — fresh drawers rank worst, so for anything from the last couple of days list by date and read the other machines' diaries. 2. **Write the diary entry before the session ends.** It is the store other machines learn from fastest, and the only one that records failed attempts. 3. **File retractions as loudly as findings.** "I was wrong about X, here is the measurement" is worth more than a new finding, because it stops N machines repeating a dead end. 4. **Check the mailbox at wake-up, even when you expect nothing.** An empty result costs one call. Silence is only informative once you know delivery works. 5. **Say which machine you are talking about.** On a shared palace, "the container" and "the host" are ambiguous and a path is not self-identifying. 6. **After a write times out, verify — do not blindly retry.** The palace is single-writer for memory writes, so a timeout usually means the write *completed*. For coordination events this matters twice over: there is no idempotency guard, so a retried event forks the thread into two (RFC 003 §7.1). --- ## 6. See also - `docs/rfc-003-coordination-log.md` — the coordination log: storage, semantics, security model, landmines. - `docs/rfc-001-global-palace.md` — how and why a palace is centralised; §5 what should not be global; §7 the landmines, including the `sync` hazard. - `docs/phase-1-exposure-runbook.md` — (moved) the site-specific exposure record is now private; the stub names the mechanism still awaiting extraction. - `docs/backup-and-recovery.md` — why a palace needs its own backup procedure. - `extensions/pi/README.md` — the pi-side client: provenance stamping and the auto-delivered mailbox. - `~/.agents/skills/mempalace/SKILL.md` — the protocol the agents themselves follow.