docs: write the RFC 003 that the code has been citing all along, plus an operator-facing fleet-memory guide

`logstream.py`'s module docstring is headed "Agent coordination event log for
MemPalace (RFC 003)" and enumerates five "Design constraints (RFC 003)".
Comments cite "RFC 003 phase 5", "RFC 003 suggested defaults" and "the first
RFC 003 dogfood". Every event/artifact tool description cites RFC 003. The
document has never existed — confirmed by searching this repo and the primary
host. So this is a retrospective spec: it transcribes what the implementation
already believes, and records what it does NOT do.

docs/rfc-003-coordination-log.md, verified line-by-line against mempalace
3.8.0 (every claim cites file.py:LINE, indexed in §10 for re-verification).
The parts that are not visible from the tool descriptions:

- No idempotency guard on event_append or put_artifact (§7.1). The replication
  path checks `id OR (origin_replica, origin_seq)` before applying; the client
  path checks nothing. So peer replay is safe and CLIENT RETRY IS NOT — a
  retried append forks a coordination thread into two ids. This makes the
  fleet's "a timeout is not a failure, verify before retrying" rule
  load-bearing rather than advisory.
- Coordination traffic is deliberately exempt from BOTH palace locks
  (§4) — _HTTP_LOCK_FREE_TOOLS and _PEER_WRITER_EXEMPT_TOOLS, each with its
  own rationale in-source. "One large mine blocks every client" is true of
  drawer writes and false of coordination writes.
- `mempalace sync` never touches the log, and no DELETE FROM events exists
  anywhere (§2, §9.1) — answering for the logstream a question RFC 001 §7.2
  left open, and making the log permanent and unbounded.
- from_agent is shape-validated and never authenticated; there is no read
  scoping at all (§6). The log authenticates the fleet, not the agent. That is
  simultaneously the security limitation and the only way to positive-control
  the mailbox.
- Three distinct orderings — seq (local arrival rowid), origin_seq (author's
  counter), hlc (fleet-wide, lexicographically sortable) — and origin_replica
  identifies the PALACE, not the writer, which is why from_agent/to_agent carry
  the whole distinction between machines (§3.2).
- GET /logstream/events does not exist and never did (§7.8). An earlier
  measurement saw it 404 and blamed the reverse proxy; that inference was right
  for /sync/* and /logstream/stream and wrong for this one. Corrected in
  extensions/pi/README.md §3 too, in place, dated.

docs/fleet-memory.md is the operator-facing companion the repo lacked
entirely: the front-door README had zero mentions of coordination, so the
channel was undiscoverable unless a message happened to arrive. It covers the
five stores and what each is for (drawers/wings/rooms, diaries, KG, palace
graph, coordination log), what a central palace buys a fleet — awareness,
non-repetition of expensive work, and retractions that travel — and a decision
flow for drawer vs KG vs event. Six mermaid diagrams, all rendered and
inspected as images, not merely parsed: the first pass produced a truncated
state label from an HTML entity and a self-loop that drew a meaningless dotted
lasso. Validation says "no syntax error"; only looking says "correct".

Also: a Documentation table in the top-level README so all of the above is
reachable from the front door.

Deliberately host-agnostic, per synlig-primary-runbook.md's precedent — no
device names, hostnames or operator names in either new document.
This commit is contained in:
2026-08-26 22:55:55 +02:00
parent 5b8d78f946
commit e1cc7592a1
4 changed files with 646 additions and 2 deletions
+218
View File
@@ -0,0 +1,218 @@
# Fleet memory — what MemPalace stores, and what to put where
**Audience:** the person running MemPalace on more than one machine, or thinking about it.
**Companion documents:** `docs/rfc-003-coordination-log.md` for the coordination log's mechanism, `docs/rfc-001-global-palace.md` for the centralisation design, and `~/.agents/skills/mempalace/SKILL.md` for what the *agents* are told to do.
MemPalace is usually described as "memory for agents", which is true and not very actionable. It is really **five stores with different retrieval models**, and most of the value — especially across a fleet — comes from putting each kind of thing in the store whose retrieval model matches how you will want it back.
This document covers: what the stores are, what a central palace changes when several machines share one, and how to decide between filing a memory and sending a message.
---
## 1. The stores at a glance
```mermaid
flowchart LR
Q["ask by MEANING<br/>'what do we know about X?'"] --> D["<b>Drawers</b><br/>chroma.sqlite3<br/>verbatim text, embedded"]
R["ask by ENTITY + TIME<br/>'what was true in June?'"] --> G["<b>Knowledge graph</b><br/>knowledge_graph.sqlite3<br/>typed facts, time-bounded"]
S["ask by ADDRESS + ORDER<br/>'what is waiting for me?'"] --> L["<b>Coordination log</b><br/>logstream.sqlite3<br/>addressed events + artifacts"]
T["ask by ASSOCIATION<br/>'what else touches this?'"] --> P["<b>Palace graph</b><br/>tunnels.json + hallways.json<br/>links between rooms"]
D --> DI["<b>Diaries</b> live here too:<br/>drawers with room=diary,<br/>read by recency, not similarity"]
```
All four stores are files inside **one palace directory**, so "the palace" is a directory you can back up in one go.
| Store | You get things back by | Typical use | Wrong use |
|---|---|---|---|
| **Drawers** (wings → rooms) | semantic similarity | a verbatim finding, a decision and its reasoning, a runbook, a transcript excerpt | anything a specific machine must *act* on; anything whose value is its exact byte content |
| **Diaries** (drawers with `room="diary"`) | agent + recency | "what did I do last session, and what did it feel like" — the continuity thread across sessions | facts other agents need to find by searching; a diary is read by *its author*, chronologically |
| **Knowledge graph** | entity, relationship, point in time | facts that *change*: versions, employers, who owns what, an injury that heals | prose, reasoning, anything you'd want to read rather than query |
| **Coordination log** | address, correlation, append order | "device B must review this patch"; "this claim is retracted, stop building on it" | durable knowledge — an event is invisible to semantic search |
| **Palace graph** | traversal from a room | discovering that an API design in one project touches a schema in another | primary storage — it links drawers, it does not hold content |
Two smaller files exist and are implementation detail, not user surface: `sqlite_exact.sqlite3` (an exact-match index over drawer metadata) and, if the daemon runs, `queue.sqlite3` (its job queue).
### 1.1 Drawers: wings and rooms
A **wing** is a project or domain; a **room** is an aspect within it. `wing="pi-devbox", room="landmines"` is a good pair; `wing="misc", room="stuff"` is how a palace becomes a landfill. Content is stored **verbatim and chunked** — never summarised — and retrieved by embedding similarity, so a drawer is found by someone who *doesn't already know it exists*. That is the property to optimise for: write the drawer that the next person's search will match.
The one counter-intuitive consequence: **fresh drawers rank worst.** A drawer filed an hour ago has no advantage in a similarity search, and a well-worn older drawer will outrank it. For anything recent, enumerate by date (`list_drawers(since=…)`) instead of searching.
### 1.2 Diaries
A diary entry is a drawer with `room="diary"`, filed by default into `wing_<agent>`, tagged with the writing agent. It is the *first-person* record: what I did, what surprised me, what I would do differently. Agents are told to write one before a session ends, and to read the last few at session start.
In a fleet this is the highest-signal store per byte, for a reason that is easy to miss: a diary entry is the only place that records **what did not work**. A drawer tends to record the conclusion; the diary records the three hours that produced it.
### 1.3 Knowledge graph
Triples — subject, predicate, object — with `valid_from` / `valid_to`, so a fact can *stop* being true without being deleted. `supersede` replaces a single-valued fact at a shared boundary, so a point-in-time query at that instant returns exactly one value.
Use it for anything you will later want to ask "what was true at time T?" about: which version was released when, who owned a service, what model an assistant was using. Do not use it for prose — a triple whose object is a paragraph is a drawer wearing a costume.
### 1.4 The coordination log
Addressed, ordered, exact. `mempalace_event_*` carries messages between agents; `mempalace_artifact_*` carries byte-exact payloads (patches, logs, files) that events can reference. This is the only store where one machine can *reach* another.
Its full mechanism, limits and landmines are in `docs/rfc-003-coordination-log.md`. §4 below covers what an operator needs to decide.
---
## 2. What changes when a fleet shares one palace
A single machine's palace is a notebook. A shared palace is something different in kind: **the fleet stops being a set of independent agents that each learn the same lessons separately.**
```mermaid
flowchart LR
A["laptop<br/>agent session"] -->|MCP over HTTPS| H(("central palace<br/>one server"))
B["workstation<br/>agent session"] -->|MCP over HTTPS| H
C["build box<br/>agent session"] -->|MCP over HTTPS| H
H --> D["drawers + diaries"]
H --> G["knowledge graph"]
H --> L["coordination log"]
```
Note the topology: in the common deployment the machines are **thin clients of one server**, not peer replicas. Everything one machine writes is immediately visible to the others — there is no sync delay to reason about, and equally no local copy to fall back on when the server is unreachable. (`docs/rfc-001-global-palace.md` §4 designs an edge proxy with a local palace and a durable outbox for deployments that need to keep working offline; the plain thin-client shape above does not.)
### 2.1 The three things this actually buys
**Awareness.** "What has anyone been doing?" becomes answerable. Each machine's diary is readable by every other machine, so an agent starting work on a shared project can see that another machine spent yesterday on it, and how far it got.
**Non-repetition of expensive work.** This is the biggest measurable win. Anything that cost real time to obtain — a scraped API surface, a spec read end to end, a bisect, a benchmark, a long investigation into why a build fails on one platform — is filed once and searchable everywhere. The second machine's cost drops from hours to one search.
**Mistakes and retractions travel.** The subtle one, and the reason a shared palace is worth more than a shared wiki. When a machine discovers that a belief was *wrong*, it can file the retraction where every other machine will hit it. Without that, each machine independently rediscovers the same dead end — and worse, a machine can spend a day rebuilding something another machine already proved doesn't work.
```mermaid
sequenceDiagram
participant W as workstation
participant P as central palace
participant L as laptop, asleep 9 days
W->>W: spends 3h finding why the build breaks
W->>P: drawer — the finding, verbatim, with evidence
W->>P: KG fact — "component X requires flag Y"
W->>P: diary — what was tried and failed
Note over L: ...9 days pass, laptop asleep...
L->>P: wake-up: read diaries + search before starting
P-->>L: the finding, the failed attempts, the fact
Note over L: cost: one search instead of 3 hours
```
### 2.2 The two costs, stated plainly
**Everything is visible to everyone.** One shared token, no per-agent read scoping. Anything filed into a shared palace should be considered readable by every machine and every agent on it. Do not put secrets in drawers.
**Provenance stops being obvious.** On a single machine, every drawer is yours and every path exists. On a shared palace, most drawers came from other machines and most `source_file` paths **do not exist locally**. Two consequences worth internalising:
- A file path in a drawer is evidence about *some* machine, not necessarily this one.
- ⚠️ **Never run `mempalace sync` against a shared palace.** It prunes drawers whose source files look gitignored, deleted or moved — which on a shared palace describes most of the content, including every other machine's. See `docs/rfc-001-global-palace.md` §7.2. The coordination log is *not* affected by this (RFC 003 §2), but drawers very much are.
---
## 3. Deciding where something goes
The question that matters is not "is this important?" but **"how will I want this back, and does anyone need to act?"**
```mermaid
flowchart TD
Start["I have something worth keeping"] --> Act{"Must a specific<br/>machine or agent<br/>DO something?"}
Act -->|no| Change{"Is it a fact that<br/>changes over time?"}
Act -->|yes| Know{"Do they also need<br/>to KNOW it later?"}
Change -->|yes| KG["Knowledge graph<br/>kg_add / kg_supersede"]
Change -->|no| Mine{"Is it about MY session<br/>— what I tried, felt, learned?"}
Mine -->|yes| Diary["Diary entry"]
Mine -->|no| Drawer["Drawer<br/>wing + room, verbatim"]
Know -->|yes| Both["BOTH:<br/>drawer for the knowledge,<br/>event pointing at it"]
Know -->|no| Event["Coordination event<br/>to_agent = the specific agent"]
```
Worked examples:
| Situation | Where | Why |
|---|---|---|
| "The release takes 76 min, and 40 of those are the base image build." | Drawer | durable, nobody must act, next person finds it by searching "release timing" |
| "v1.8.9 is the released version, as of this timestamp." | KG (`supersede`) | it will change; you will want "what was released in August?" |
| "I spent two hours chasing a watcher that was already dead." | Diary | first-person, chronological, tells the next session what *not* to retry |
| "Build box: this patch is ready, please review and apply." | Event (directed) | a named machine must act; the patch itself goes in as an artifact |
| "The claim in that drawer is wrong — I measured the opposite." | Both | file the corrected finding as a drawer, then an event so the machine building on it stops |
| "Everyone should know the new toolkit is live." | Drawer + broadcast event | the drawer is what anyone will *find*; the broadcast is a notice, not an ask (§4.2) |
The failure mode in each direction is worth naming, because both are common:
- **A finding filed only as an event** is invisible to semantic search. Nobody will ever find it again, and the next agent will re-derive it.
- **An ask filed only as a drawer** is addressed to nobody. It will be found, if ever, by accident — long after it mattered.
---
## 4. What the coordination log can and cannot do
This is the part most likely to be mis-set expectations, so it is worth being blunt: **it is a durable log, not a chat.** Nothing is listening. Events are appended and persist; there is no delivery window; nothing is lost by being offline when one is written. A message waits indefinitely, and your reply waits just as patiently for a sender who has since gone away.
That sounds like a limitation and is actually the correct design for a fleet where few machines are awake at once and any given machine may sleep for weeks.
### 4.1 Latency, honestly
| Recipient state | When they see it | Notes |
|---|---|---|
| In a live session, mailbox-enabled client | **≈2–5 minutes** | measured ≈2–3 min on first live delivery; a poll floor of 5 min applies between checks |
| Holding an SSE connection (`GET /logstream/stream`) | sub-second | for daemons/dashboards, not interactive agents |
| Asleep, next session tomorrow | tomorrow | delivered in the session-start wake-up |
| Asleep for three weeks | in three weeks | nothing expires; the log is permanent |
| Never runs again | never | there is no re-routing and no dead-letter path |
So: appropriate for "handle this when you next wake", "here is a patch", "stop building on that claim". Not appropriate for anything with a deadline inside the hour, unless you know the recipient is awake.
One pleasant property, worth knowing because it is counter-intuitive: coordination traffic is **exempt from the palace's write lock**, so you can message another machine and it can reply *while* a long mine is running on the server (RFC 003 §4). Coordination stays alive when memory writes are blocked.
### 4.2 The one thing that does not work: broadcasting an ask
You can write a broadcast (`to_agent="*"`) and every machine that *lists* events will see it. But a broadcast **never enters any machine's mailbox** and is never auto-delivered — by design, because "everyone owes this answer" degenerates into either N duplicate replies or nobody acting.
- **Broadcast** = a notice on a wall. Fine for "v1.8.9 is out".
- **Directed event** = a message in a named mailbox. Required for anything that must be done.
To reach a whole fleet with something actionable, **fan out**: one directed event per device, sharing one `correlation_id` so the thread stays joinable. Each machine then owes its own reply.
```mermaid
flowchart LR
You["you"] -->|"to_agent=pi@laptop"| A["laptop owes a reply"]
You -->|"to_agent=pi@workstation"| B["workstation owes a reply"]
You -->|"to_agent=pi@build-box"| C["build box owes a reply"]
A --> Corr["one shared correlation_id<br/>joins the three threads"]
B --> Corr
C --> Corr
```
### 4.3 Addressing, and why the format matters
Addresses are `<harness>@<device>` — `pi@laptop`, `opencode@build-box`. Two rules follow:
1. **An unstamped client is unreachable.** If a machine's events say `from_agent: pi` with no device, nobody can address it, because "pi" is every machine.
2. **Two machines must never share one address.** Nothing prevents it, nothing warns, and the result is that each silently discards the other's asks as its own (RFC 003 §7.7).
### 4.4 A reply is owed until it is *terminally* closed
An acknowledgement does not close a thread. Neither does "claimed" or "ready". Only a terminal status — `applied`, `superseded`, `failed`, `blocked` — clears an ask from the recipient's mailbox. Until then, a mailbox-enabled client will keep resurfacing it, which is the intended behaviour: an unanswered ask should nag.
---
## 5. Habits that make a shared palace work
Small, and the whole value rests on them:
1. **Search before you answer, and enumerate before you conclude.** One empty search is not proof of silence — fresh drawers rank worst, so for anything from the last couple of days list by date and read the other machines' diaries.
2. **Write the diary entry before the session ends.** It is the store other machines learn from fastest, and the only one that records failed attempts.
3. **File retractions as loudly as findings.** "I was wrong about X, here is the measurement" is worth more than a new finding, because it stops N machines repeating a dead end.
4. **Check the mailbox at wake-up, even when you expect nothing.** An empty result costs one call. Silence is only informative once you know delivery works.
5. **Say which machine you are talking about.** On a shared palace, "the container" and "the host" are ambiguous and a path is not self-identifying.
6. **After a write times out, verify — do not blindly retry.** The palace is single-writer for memory writes, so a timeout usually means the write *completed*. For coordination events this matters twice over: there is no idempotency guard, so a retried event forks the thread into two (RFC 003 §7.1).
---
## 6. See also
- `docs/rfc-003-coordination-log.md` — the coordination log: storage, semantics, security model, landmines.
- `docs/rfc-001-global-palace.md` — how and why a palace is centralised; §5 what should not be global; §7 the landmines, including the `sync` hazard.
- `docs/phase-1-exposure-runbook.md` — exposing a palace over HTTP.
- `docs/backup-and-recovery.md` — why a palace needs its own backup procedure.
- `extensions/pi/README.md` — the pi-side client: provenance stamping and the auto-delivered mailbox.
- `~/.agents/skills/mempalace/SKILL.md` — the protocol the agents themselves follow.