diff --git a/README.md b/README.md index 758c7cf..5cddad1 100644 --- a/README.md +++ b/README.md @@ -58,6 +58,27 @@ Who owns what: --- +## Documentation + +Start here if you are deciding **what to put where**, or running MemPalace on more than one machine: + +| Document | What it answers | +|---|---| +| [`docs/fleet-memory.md`](docs/fleet-memory.md) | **Start here.** The five stores MemPalace keeps and what each is for; what a *central* palace buys a fleet of machines; when to file a memory versus when to message another device. Diagrams, worked examples. | +| [`docs/rfc-003-coordination-log.md`](docs/rfc-003-coordination-log.md) | The inter-agent / inter-device coordination log (`logstream`): storage, append and query semantics, delivery and latency, the trust model, and its landmines. | +| [`docs/rfc-001-global-palace.md`](docs/rfc-001-global-palace.md) | Why and how a palace is centralised, what should *not* be global, and the Phase-0 landmines (including: never run `mempalace sync` against a shared palace). | +| [`docs/rfc-002-joiner.md`](docs/rfc-002-joiner.md) | Replaying a second palace into a shared primary, and what dedupes what. | +| [`docs/phase-1-exposure-runbook.md`](docs/phase-1-exposure-runbook.md) | Exposing a palace over HTTP: the Host/Origin pin, tokens, TLS. | +| [`docs/backup-and-recovery.md`](docs/backup-and-recovery.md) | Why a palace needs its own backup procedure, and how to restore one. | +| [`extensions/pi/README.md`](extensions/pi/README.md) | The pi-side client: provenance stamping at the edge, and the auto-delivered mailbox. | + +Behaviour for the *agents* is normative in the `mempalace` skill (`SKILL.md` here, +installed to `~/.agents/skills/mempalace/`), not in these documents — the split is +deliberate: the skill says what an agent should do, the docs say what the machinery +actually does. + +--- + ## Why this exists MemPalace is the agent memory layer. Its stock CLI has two gaps that bite on a machine running opencode with a docs-first palace policy: diff --git a/docs/fleet-memory.md b/docs/fleet-memory.md new file mode 100644 index 0000000..3561967 --- /dev/null +++ b/docs/fleet-memory.md @@ -0,0 +1,218 @@ +# Fleet memory — what MemPalace stores, and what to put where + +**Audience:** the person running MemPalace on more than one machine, or thinking about it. +**Companion documents:** `docs/rfc-003-coordination-log.md` for the coordination log's mechanism, `docs/rfc-001-global-palace.md` for the centralisation design, and `~/.agents/skills/mempalace/SKILL.md` for what the *agents* are told to do. + +MemPalace is usually described as "memory for agents", which is true and not very actionable. It is really **five stores with different retrieval models**, and most of the value — especially across a fleet — comes from putting each kind of thing in the store whose retrieval model matches how you will want it back. + +This document covers: what the stores are, what a central palace changes when several machines share one, and how to decide between filing a memory and sending a message. + +--- + +## 1. The stores at a glance + +```mermaid +flowchart LR + Q["ask by MEANING
'what do we know about X?'"] --> D["Drawers
chroma.sqlite3
verbatim text, embedded"] + R["ask by ENTITY + TIME
'what was true in June?'"] --> G["Knowledge graph
knowledge_graph.sqlite3
typed facts, time-bounded"] + S["ask by ADDRESS + ORDER
'what is waiting for me?'"] --> L["Coordination log
logstream.sqlite3
addressed events + artifacts"] + T["ask by ASSOCIATION
'what else touches this?'"] --> P["Palace graph
tunnels.json + hallways.json
links between rooms"] + D --> DI["Diaries live here too:
drawers with room=diary,
read by recency, not similarity"] +``` + +All four stores are files inside **one palace directory**, so "the palace" is a directory you can back up in one go. + +| Store | You get things back by | Typical use | Wrong use | +|---|---|---|---| +| **Drawers** (wings → rooms) | semantic similarity | a verbatim finding, a decision and its reasoning, a runbook, a transcript excerpt | anything a specific machine must *act* on; anything whose value is its exact byte content | +| **Diaries** (drawers with `room="diary"`) | agent + recency | "what did I do last session, and what did it feel like" — the continuity thread across sessions | facts other agents need to find by searching; a diary is read by *its author*, chronologically | +| **Knowledge graph** | entity, relationship, point in time | facts that *change*: versions, employers, who owns what, an injury that heals | prose, reasoning, anything you'd want to read rather than query | +| **Coordination log** | address, correlation, append order | "device B must review this patch"; "this claim is retracted, stop building on it" | durable knowledge — an event is invisible to semantic search | +| **Palace graph** | traversal from a room | discovering that an API design in one project touches a schema in another | primary storage — it links drawers, it does not hold content | + +Two smaller files exist and are implementation detail, not user surface: `sqlite_exact.sqlite3` (an exact-match index over drawer metadata) and, if the daemon runs, `queue.sqlite3` (its job queue). + +### 1.1 Drawers: wings and rooms + +A **wing** is a project or domain; a **room** is an aspect within it. `wing="pi-devbox", room="landmines"` is a good pair; `wing="misc", room="stuff"` is how a palace becomes a landfill. Content is stored **verbatim and chunked** — never summarised — and retrieved by embedding similarity, so a drawer is found by someone who *doesn't already know it exists*. That is the property to optimise for: write the drawer that the next person's search will match. + +The one counter-intuitive consequence: **fresh drawers rank worst.** A drawer filed an hour ago has no advantage in a similarity search, and a well-worn older drawer will outrank it. For anything recent, enumerate by date (`list_drawers(since=…)`) instead of searching. + +### 1.2 Diaries + +A diary entry is a drawer with `room="diary"`, filed by default into `wing_`, tagged with the writing agent. It is the *first-person* record: what I did, what surprised me, what I would do differently. Agents are told to write one before a session ends, and to read the last few at session start. + +In a fleet this is the highest-signal store per byte, for a reason that is easy to miss: a diary entry is the only place that records **what did not work**. A drawer tends to record the conclusion; the diary records the three hours that produced it. + +### 1.3 Knowledge graph + +Triples — subject, predicate, object — with `valid_from` / `valid_to`, so a fact can *stop* being true without being deleted. `supersede` replaces a single-valued fact at a shared boundary, so a point-in-time query at that instant returns exactly one value. + +Use it for anything you will later want to ask "what was true at time T?" about: which version was released when, who owned a service, what model an assistant was using. Do not use it for prose — a triple whose object is a paragraph is a drawer wearing a costume. + +### 1.4 The coordination log + +Addressed, ordered, exact. `mempalace_event_*` carries messages between agents; `mempalace_artifact_*` carries byte-exact payloads (patches, logs, files) that events can reference. This is the only store where one machine can *reach* another. + +Its full mechanism, limits and landmines are in `docs/rfc-003-coordination-log.md`. §4 below covers what an operator needs to decide. + +--- + +## 2. What changes when a fleet shares one palace + +A single machine's palace is a notebook. A shared palace is something different in kind: **the fleet stops being a set of independent agents that each learn the same lessons separately.** + +```mermaid +flowchart LR + A["laptop
agent session"] -->|MCP over HTTPS| H(("central palace
one server")) + B["workstation
agent session"] -->|MCP over HTTPS| H + C["build box
agent session"] -->|MCP over HTTPS| H + H --> D["drawers + diaries"] + H --> G["knowledge graph"] + H --> L["coordination log"] +``` + +Note the topology: in the common deployment the machines are **thin clients of one server**, not peer replicas. Everything one machine writes is immediately visible to the others — there is no sync delay to reason about, and equally no local copy to fall back on when the server is unreachable. (`docs/rfc-001-global-palace.md` §4 designs an edge proxy with a local palace and a durable outbox for deployments that need to keep working offline; the plain thin-client shape above does not.) + +### 2.1 The three things this actually buys + +**Awareness.** "What has anyone been doing?" becomes answerable. Each machine's diary is readable by every other machine, so an agent starting work on a shared project can see that another machine spent yesterday on it, and how far it got. + +**Non-repetition of expensive work.** This is the biggest measurable win. Anything that cost real time to obtain — a scraped API surface, a spec read end to end, a bisect, a benchmark, a long investigation into why a build fails on one platform — is filed once and searchable everywhere. The second machine's cost drops from hours to one search. + +**Mistakes and retractions travel.** The subtle one, and the reason a shared palace is worth more than a shared wiki. When a machine discovers that a belief was *wrong*, it can file the retraction where every other machine will hit it. Without that, each machine independently rediscovers the same dead end — and worse, a machine can spend a day rebuilding something another machine already proved doesn't work. + +```mermaid +sequenceDiagram + participant W as workstation + participant P as central palace + participant L as laptop, asleep 9 days + W->>W: spends 3h finding why the build breaks + W->>P: drawer — the finding, verbatim, with evidence + W->>P: KG fact — "component X requires flag Y" + W->>P: diary — what was tried and failed + Note over L: ...9 days pass, laptop asleep... + L->>P: wake-up: read diaries + search before starting + P-->>L: the finding, the failed attempts, the fact + Note over L: cost: one search instead of 3 hours +``` + +### 2.2 The two costs, stated plainly + +**Everything is visible to everyone.** One shared token, no per-agent read scoping. Anything filed into a shared palace should be considered readable by every machine and every agent on it. Do not put secrets in drawers. + +**Provenance stops being obvious.** On a single machine, every drawer is yours and every path exists. On a shared palace, most drawers came from other machines and most `source_file` paths **do not exist locally**. Two consequences worth internalising: + +- A file path in a drawer is evidence about *some* machine, not necessarily this one. +- ⚠️ **Never run `mempalace sync` against a shared palace.** It prunes drawers whose source files look gitignored, deleted or moved — which on a shared palace describes most of the content, including every other machine's. See `docs/rfc-001-global-palace.md` §7.2. The coordination log is *not* affected by this (RFC 003 §2), but drawers very much are. + +--- + +## 3. Deciding where something goes + +The question that matters is not "is this important?" but **"how will I want this back, and does anyone need to act?"** + +```mermaid +flowchart TD + Start["I have something worth keeping"] --> Act{"Must a specific
machine or agent
DO something?"} + Act -->|no| Change{"Is it a fact that
changes over time?"} + Act -->|yes| Know{"Do they also need
to KNOW it later?"} + Change -->|yes| KG["Knowledge graph
kg_add / kg_supersede"] + Change -->|no| Mine{"Is it about MY session
— what I tried, felt, learned?"} + Mine -->|yes| Diary["Diary entry"] + Mine -->|no| Drawer["Drawer
wing + room, verbatim"] + Know -->|yes| Both["BOTH:
drawer for the knowledge,
event pointing at it"] + Know -->|no| Event["Coordination event
to_agent = the specific agent"] +``` + +Worked examples: + +| Situation | Where | Why | +|---|---|---| +| "The release takes 76 min, and 40 of those are the base image build." | Drawer | durable, nobody must act, next person finds it by searching "release timing" | +| "v1.8.9 is the released version, as of this timestamp." | KG (`supersede`) | it will change; you will want "what was released in August?" | +| "I spent two hours chasing a watcher that was already dead." | Diary | first-person, chronological, tells the next session what *not* to retry | +| "Build box: this patch is ready, please review and apply." | Event (directed) | a named machine must act; the patch itself goes in as an artifact | +| "The claim in that drawer is wrong — I measured the opposite." | Both | file the corrected finding as a drawer, then an event so the machine building on it stops | +| "Everyone should know the new toolkit is live." | Drawer + broadcast event | the drawer is what anyone will *find*; the broadcast is a notice, not an ask (§4.2) | + +The failure mode in each direction is worth naming, because both are common: + +- **A finding filed only as an event** is invisible to semantic search. Nobody will ever find it again, and the next agent will re-derive it. +- **An ask filed only as a drawer** is addressed to nobody. It will be found, if ever, by accident — long after it mattered. + +--- + +## 4. What the coordination log can and cannot do + +This is the part most likely to be mis-set expectations, so it is worth being blunt: **it is a durable log, not a chat.** Nothing is listening. Events are appended and persist; there is no delivery window; nothing is lost by being offline when one is written. A message waits indefinitely, and your reply waits just as patiently for a sender who has since gone away. + +That sounds like a limitation and is actually the correct design for a fleet where few machines are awake at once and any given machine may sleep for weeks. + +### 4.1 Latency, honestly + +| Recipient state | When they see it | Notes | +|---|---|---| +| In a live session, mailbox-enabled client | **≈2–5 minutes** | measured ≈2–3 min on first live delivery; a poll floor of 5 min applies between checks | +| Holding an SSE connection (`GET /logstream/stream`) | sub-second | for daemons/dashboards, not interactive agents | +| Asleep, next session tomorrow | tomorrow | delivered in the session-start wake-up | +| Asleep for three weeks | in three weeks | nothing expires; the log is permanent | +| Never runs again | never | there is no re-routing and no dead-letter path | + +So: appropriate for "handle this when you next wake", "here is a patch", "stop building on that claim". Not appropriate for anything with a deadline inside the hour, unless you know the recipient is awake. + +One pleasant property, worth knowing because it is counter-intuitive: coordination traffic is **exempt from the palace's write lock**, so you can message another machine and it can reply *while* a long mine is running on the server (RFC 003 §4). Coordination stays alive when memory writes are blocked. + +### 4.2 The one thing that does not work: broadcasting an ask + +You can write a broadcast (`to_agent="*"`) and every machine that *lists* events will see it. But a broadcast **never enters any machine's mailbox** and is never auto-delivered — by design, because "everyone owes this answer" degenerates into either N duplicate replies or nobody acting. + +- **Broadcast** = a notice on a wall. Fine for "v1.8.9 is out". +- **Directed event** = a message in a named mailbox. Required for anything that must be done. + +To reach a whole fleet with something actionable, **fan out**: one directed event per device, sharing one `correlation_id` so the thread stays joinable. Each machine then owes its own reply. + +```mermaid +flowchart LR + You["you"] -->|"to_agent=pi@laptop"| A["laptop owes a reply"] + You -->|"to_agent=pi@workstation"| B["workstation owes a reply"] + You -->|"to_agent=pi@build-box"| C["build box owes a reply"] + A --> Corr["one shared correlation_id
joins the three threads"] + B --> Corr + C --> Corr +``` + +### 4.3 Addressing, and why the format matters + +Addresses are `@` — `pi@laptop`, `opencode@build-box`. Two rules follow: + +1. **An unstamped client is unreachable.** If a machine's events say `from_agent: pi` with no device, nobody can address it, because "pi" is every machine. +2. **Two machines must never share one address.** Nothing prevents it, nothing warns, and the result is that each silently discards the other's asks as its own (RFC 003 §7.7). + +### 4.4 A reply is owed until it is *terminally* closed + +An acknowledgement does not close a thread. Neither does "claimed" or "ready". Only a terminal status — `applied`, `superseded`, `failed`, `blocked` — clears an ask from the recipient's mailbox. Until then, a mailbox-enabled client will keep resurfacing it, which is the intended behaviour: an unanswered ask should nag. + +--- + +## 5. Habits that make a shared palace work + +Small, and the whole value rests on them: + +1. **Search before you answer, and enumerate before you conclude.** One empty search is not proof of silence — fresh drawers rank worst, so for anything from the last couple of days list by date and read the other machines' diaries. +2. **Write the diary entry before the session ends.** It is the store other machines learn from fastest, and the only one that records failed attempts. +3. **File retractions as loudly as findings.** "I was wrong about X, here is the measurement" is worth more than a new finding, because it stops N machines repeating a dead end. +4. **Check the mailbox at wake-up, even when you expect nothing.** An empty result costs one call. Silence is only informative once you know delivery works. +5. **Say which machine you are talking about.** On a shared palace, "the container" and "the host" are ambiguous and a path is not self-identifying. +6. **After a write times out, verify — do not blindly retry.** The palace is single-writer for memory writes, so a timeout usually means the write *completed*. For coordination events this matters twice over: there is no idempotency guard, so a retried event forks the thread into two (RFC 003 §7.1). + +--- + +## 6. See also + +- `docs/rfc-003-coordination-log.md` — the coordination log: storage, semantics, security model, landmines. +- `docs/rfc-001-global-palace.md` — how and why a palace is centralised; §5 what should not be global; §7 the landmines, including the `sync` hazard. +- `docs/phase-1-exposure-runbook.md` — exposing a palace over HTTP. +- `docs/backup-and-recovery.md` — why a palace needs its own backup procedure. +- `extensions/pi/README.md` — the pi-side client: provenance stamping and the auto-delivered mailbox. +- `~/.agents/skills/mempalace/SKILL.md` — the protocol the agents themselves follow. diff --git a/docs/rfc-003-coordination-log.md b/docs/rfc-003-coordination-log.md new file mode 100644 index 0000000..2c2d389 --- /dev/null +++ b/docs/rfc-003-coordination-log.md @@ -0,0 +1,396 @@ +# RFC 003 — The coordination log (`logstream`) + +**Status:** implemented and in production use since 3.7.x. This document is a *retrospective* specification, written after the fact. +**Author:** pi (agent), 2026-08-26. +**Applies to:** mempalace 3.8.0 (`logstream.py`, 1261 lines), mempalace-toolkit `5b8d78f` (the pi-side mailbox). +**Context:** every `mempalace_event_*` and `mempalace_artifact_*` tool description cites "RFC 003". `logstream.py`'s module docstring is headed *"Agent coordination event log for MemPalace (RFC 003)"* and enumerates five "Design constraints (RFC 003)". Inline comments cite "RFC 003 phase 5", "RFC 003 suggested defaults" and "the first RFC 003 dogfood". **No such document has ever existed** — verified 2026-08-26 by searching both the toolkit repository and the primary host. This RFC transcribes the spec the implementation already believes in, and — more usefully — records what it does *not* do. + +Read §7 before you build anything on this log; it is the part that is not obvious from the tool descriptions. Every mechanical claim below cites `file.py:LINE` in mempalace 3.8.0, and §10 indexes them so any claim can be re-verified without re-reading 1261 lines. Claims marked **measured** were executed against the live fleet on the date given; claims without that marker are reads of the source. Where the code and the shipped tool descriptions disagree, §10.1 says so explicitly rather than quietly siding with one. + +One scoping note up front. This RFC covers the **log**: its storage, its append and query semantics, delivery, and the trust model. It does not specify **replication**, which the source already attributes to a different document (`# ── Replication (RFC 004 step 0: logstream multi-master) ──`, `logstream.py:1096`). RFC 004 does not exist either; §8.2 states what it owes. + +--- + +## 1. The problem + +The palace stores what an agent *knows*: drawers are semantic, retrieved by meaning, and deliberately have no addressee. That shape is wrong for four things a fleet of agents actually needs: + +1. **Addressing.** "This is for the machine that owns the release" cannot be expressed as a drawer. A drawer is found by whoever happens to search for the right words. +2. **A reply that closes something.** Semantic memory has no notion of an outstanding question. Nothing in a drawer can be *owed*. +3. **Exact payloads.** A unified diff must survive byte-for-byte. Drawers are chunked and embedded; that is a feature for prose and a defect for patches. +4. **Order.** "What happened after this?" needs an append cursor, not a similarity score. + +The coordination log adds exactly those four properties and nothing else. It is a second store beside the palace, not a new kind of drawer. + +### 1.1 What it is not + +Stated first, because every misuse of this log so far has come from assuming one of these: + +- **Not a bus.** Nothing subscribes by default; nothing is delivered "live" unless a client is holding an SSE connection or polling. There is no delivery window and nothing is lost by being offline when an event is written. +- **Not a chat channel.** Latency is bounded by *when the recipient next runs*, which in a fleet of workstations is hours, days or weeks. §3.5 gives the measured numbers for the case where the recipient is awake. +- **Not authenticated per agent.** `from_agent` is a routing label with the trust properties of an e-mail `From:` header (§6). +- **Not a queue.** Nothing is consumed, acknowledged-and-removed, or retried. Events are permanent (§9.1) and "handled" is a *derived* property (§3.3). + +--- + +## 2. Verified starting point (2026-08-26, mempalace 3.8.0) + +| Piece | State | Evidence | +|---|---|---| +| Separate SQLite store, inside the palace dir | ✅ `logstream.sqlite3` beside `chroma.sqlite3`; WAL; dir `chmod 0700` best-effort | `logstream.py:51`, `mcp_server.py:1108-1117`, `logstream.py:430-433` | +| `events`, `artifacts`, `event_artifacts` tables | ✅ Created idempotently at open | `logstream.py:447-497` | +| Hybrid logical clock on every event | ✅ `hlc` populated on every local append | `logstream.py:668`, `hlc.py:1-21` | +| Seven MCP tools (append/list/wait/ack, artifact put/get, patch_submit) | ✅ | `mcp_server.py:4546-4771` | +| Server-sent events push | ✅ `GET /logstream/stream`, auth required, 15 s heartbeat, ≤8 clients | `mcp_server.py:7570-7671`, `7189-7194` | +| `GET /logstream/events` | ❌ **Never implemented.** Only `/logstream/stream` exists | see §7.8 | +| Auto-delivered mailbox (owed-set derivation + injection) | ✅ **Client-side, not server-side** — lives in the toolkit's pi extension | `extensions/pi/mempalace.ts` (toolkit `5b8d78f`) | +| Idempotency guard on append | ❌ **Absent.** See §7.1 | `logstream.py:625-716` | +| Retention / TTL / compaction | ❌ Absent by design-so-far. See §9.1 | no `DELETE FROM events` in the package | +| Interaction with `mempalace sync` | ✅ **None.** `sync` prunes drawers only | `cli.py:1057`, `sync.py` (no logstream references) | + +Two things that sound like this feature and are not: + +- **`mempalace_mesh_peers` is not a fleet roster.** It reports peer *replicas*. A hub-and-spoke deployment — many thin MCP clients of one server — correctly reports `peers: []` while every machine in the fleet is actively writing to the same log. Deciding "coordination does not apply to me" from an empty peer list is a measured failure mode, not a hypothetical one. +- **`origin_replica` is not the writing machine.** It is a property of the palace *directory* (§3.2). + +--- + +## 3. Design + +The five constraints in the module docstring (`logstream.py:1-18`) are the design, quoted verbatim because they are already normative in the implementation: + +> - No Chroma dependency, no vector index open — plain SQLite only. +> - Append-only: events are immutable; corrections are new events that reference prior events. +> - Exact payloads: event bodies and artifact content are stored verbatim. +> - Safe under concurrent HTTP requests (WAL + per-instance lock, same pattern as `knowledge_graph.py`). +> - Explicit size limits with clear errors, never silent truncation. + +The shape, in the same idiom as RFC 001 §4.1: + +``` +agent on device A ─┐ palace directory +agent on device B ─┼── MCP /mcp ──► server ─┬─ chroma.sqlite3 (drawers: what you know) +agent on device C ─┘ │ ├─ knowledge_graph.sqlite3 (facts, temporal) + │ └─ logstream.sqlite3 (events + artifacts) + GET /logstream/stream events ── append-only, permanent + (SSE, auth, ≤8 clients) artifacts ── verbatim, ≤4 MiB, sha256 + event_artifacts ── join, checked at append +``` + +### 3.1 Data model + +```sql +CREATE TABLE IF NOT EXISTS events ( + id TEXT PRIMARY KEY, + type TEXT NOT NULL, + stream TEXT NOT NULL, + room TEXT NOT NULL, + from_agent TEXT NOT NULL, + to_agent TEXT, + correlation_id TEXT, + branch TEXT, + base_commit TEXT, + status TEXT, + body TEXT NOT NULL DEFAULT '', + created_at TEXT NOT NULL, + metadata_json TEXT NOT NULL DEFAULT '{}', + origin_replica TEXT, + origin_seq INTEGER, + hlc TEXT +); +CREATE UNIQUE INDEX IF NOT EXISTS events_origin_seq_idx ON events(origin_replica, origin_seq); +``` +(`logstream.py:447-497`, unique index at `552-556`. `artifacts` and `event_artifacts` in the same script; `artifacts_sha256_idx` is **not** unique — see §3.4.) + +Validation, all server-side, all raising `ValueError` naming the allowed set: + +| Field | Required | Rule | Where | +|---|---|---|---| +| `type` | yes | `^[a-z0-9][a-z0-9_.-]{0,63}$` — **lowercase only, ≤64** | `logstream.py:76, 124-134` | +| `stream`, `room`, `from_agent` | yes | non-empty, ≤256, no control chars | `logstream.py:78, 103-122` | +| `status` | no | one of `open claimed ready applied blocked failed superseded` | `logstream.py:67-69, 136-143` | +| `body` | no (empty allowed) | ≤256 KiB, no NUL | `logstream.py:55, 145-158` | +| `metadata` | no | canonical JSON, **≤64 KiB** | `logstream.py:57, 160-177` | +| `artifact_ids` | no | each must already exist, else `ValueError` | `logstream.py:687-694` | + +Three consequences worth stating because they are invisible from the tool descriptions: an event body **may be empty** while artifact content **may not** (`logstream.py:766-767`); `to_agent` is **nullable**, and an event with `to_agent = NULL` is addressed to nobody and can never match a `to_agent=` query; and the 64 KiB `metadata` cap and the lowercase-`type` regex are **enforced but undocumented** in the shipped tool schemas. + +### 3.2 Three orderings, and one identity that is not what it looks like + +The log carries three distinct notions of order, and conflating them is the single most likely design error in any consumer: + +| Field | Meaning | Scope | Where | +|---|---|---|---| +| `seq` | **local arrival cursor** — the SQLite `rowid`, surfaced at read time | this replica only | `logstream.py:580-586` | +| `origin_seq` | the author replica's own gap-free counter, assigned by `UPDATE … SET origin_seq = rowid` in the insert transaction | per author | `logstream.py:700-706` | +| `hlc` | hybrid logical clock, `--`, fixed-width so TEXT comparison *is* causal comparison | fleet-wide | `logstream.py:668`, `hlc.py:1-21` | + +The rule, from `hlc.py:19-21`: *"Cursor semantics stay LOCAL (rowid arrival order) … a tail consumer must see late-arriving remote ops even though their HLC is older. HLC is the display/merge order; arrival is the delivery order."* + +**`origin_replica` identifies the palace, not the writer.** It comes from `get_replica_id(db_parent)` — the `replica.json` in the palace directory (`logstream.py:422-436`, `replica.py:31-49`). In a hub-and-spoke deployment every client writes through the server's single `Logstream`, so **every event from every machine carries the same `origin_replica`** and it cannot distinguish two devices. This is mechanically forced by the design, not a misconfiguration. + +Therefore: **`from_agent` / `to_agent` carry the entire distinction between machines.** The convention that makes them able to is `@` — e.g. `pi@laptop`, `opencode@build-box` — stamped at the client edge, which RFC 001 §7.3.5 specifies. An unstamped client addressed as bare `pi` is unreachable in a fleet, because nobody can name it. + +### 3.3 The ack contract: owed-ness is derived, never read + +`status` is written **once**, into an append-only row. Nothing updates it. So `status="open"` means *"the sender declared this an ask at the moment of writing"* — it does **not** mean unanswered, and a directed `open` event keeps matching a mailbox query forever, answered or not. + +`ack_event` (`logstream.py:812-845`) never touches the target row. It *appends*: + +```python +return self.append_event( + type=ACK_EVENT_TYPE, # "event.ack" + stream=target["stream"], room=target["room"], + from_agent=from_agent, + to_agent=target["from_agent"], # routes back to the sender + correlation_id=target["correlation_id"] or target["id"], # falls back to the target id + status=status, body=body, + metadata={"ack_of": event_id}, # the join key lives in metadata +) +``` + +So a consumer that wants "what do I still owe?" must derive it. The derivation the toolkit implements, and which this RFC adopts as the reference semantics: + +```mermaid +stateDiagram-v2 + [*] --> open : task.request written, addressed to you + open --> claimed : you announce you are working (does NOT clear) + open --> ready : work exists, thread still open (does NOT clear) + claimed --> applied : terminal, clears + ready --> applied : terminal, clears + open --> blocked : terminal, clears (with a reason) + open --> failed : terminal, clears (with a reason) + open --> superseded : terminal, clears (replaced by another thread) + applied --> [*] + blocked --> [*] + failed --> [*] + superseded --> [*] +``` + +An event is **still owed** when all of these hold: + +1. it is directed at you (`to_agent = `, and `to_agent = '*'` is excluded — §7.4); +2. it was not written by you (you cannot owe yourself); +3. **no** event of yours has *all* of: a strictly higher `seq`, a join to it (`metadata.ack_of = ` or a shared `correlation_id`), **and** a terminal status (`applied` / `superseded` / `failed` / `blocked`). + +The strictly-higher-`seq` test is load-bearing: without it, one terminal reply on a `correlation_id` suppresses every *later* ask on that same correlation, permanently. And because `seq` is local arrival order, this test is only sound on a single replica — §7.3. + +### 3.4 Artifacts + +Exact payloads for handoff: `patch`, `file`, `log`, `json`, `note`; UTF-8 text only; ≤4 MiB; `sha256` and `size_bytes` returned (`logstream.py:744-810`). + +⚠️ **The sha256 is for verification, not deduplication.** `artifacts_sha256_idx` is a plain index and `put_artifact` always INSERTs, so storing the same patch twice stores the content twice. Referential integrity in the other direction *is* enforced: `artifact_ids` naming an unknown artifact raises at append time, in both the local and the remote path (`logstream.py:687-694`, `1181-1189`), *"so readers never see a dangling reference"*. + +`kind="patch"` gets advisory warnings — never mutation — for a missing trailing newline or CRLF endings (`logstream.py:718-742`). The comment records why: *"the first RFC 003 dogfood: a patch stored without its final trailing newline truncates the last hunk line."* + +### 3.5 Delivery: three mechanisms, one client feature + +| Mechanism | Shape | Latency | Where | +|---|---|---|---| +| `event_list` | pull, filtered, `ORDER BY rowid ASC`, default 50 / max 500 (silently clamped) | whenever you ask | `logstream.py:914-990` | +| `event_wait` | in-request long poll, jittered backoff 0.25 s→1 s, default 60 s, hard cap 300 s, returns `{timed_out: true, events: []}` rather than raising | seconds, while you hold the call | `logstream.py:1004-1039` | +| `GET /logstream/stream` | SSE; resume by `?since_event_id=` or `Last-Event-ID`; `: ping` every 15 s; ≤8 concurrent clients then `503` + `Retry-After` | sub-second, while connected | `mcp_server.py:7570-7671` | + +None of these push anything into an agent that is not already running. That gap is closed **client-side** by the toolkit's *mailbox*: it derives the owed set per §3.3 and injects it at two points — once in the session-start wake-up, and mid-session on a settled-agent poll floored at 5 minutes, delivered as a queued steer rather than an interrupt. + +**Measured 2026-08-26** (first live delivery on a published image, positive control with a synthetic sender): planted → delivered in **≈2–3 minutes** to a session that was already running, as exactly one item; three already-answered asks were correctly excluded and a `*` broadcast was correctly ignored. The wake-up path and the resurface path were **not** exercised by that test and remain unmeasured. + +This is why the mailbox is a client concern and belongs in the toolkit rather than here: the server has no idea which agents exist, and no way to reach one that is not calling it. + +--- + +## 4. Availability: the logstream is deliberately exempt from both palace locks + +The rest of the palace is a single-writer service — `_HTTP_REQUEST_LOCK` wraps every dispatch (RFC 001 §7.5), and a Chroma peer-writer lease protects against a concurrent CLI mine. **Coordination traffic is exempt from both**, in two separate places, for two different reasons: + +- `_HTTP_LOCK_FREE_TOOLS` (`mcp_server.py:6887-6902`) — all seven event/artifact tools dispatch outside the request lock: *"Dispatching them outside `_HTTP_REQUEST_LOCK` keeps a five-minute `mempalace_event_wait` long-poll from stalling every other agent on a shared hub, and lets the SSE stream coexist with normal tool traffic."* +- `_PEER_WRITER_EXEMPT_TOOLS` (`mcp_server.py:434-449`) — the four mutating ones bypass the writer lease: *"Exempting them keeps agent coordination alive while a CLI mine or a peer stdio writer holds the palace lock."* + +They remain in `_MUTATING_TOOLS`, so `--read-only` still refuses them. + +✅ **Consequence, and it corrects a widely-held belief in this fleet:** "one large mine blocks every client for minutes" is true of drawer writes and **false of coordination writes**. You can message another device, and it can reply, while a mine is running. + +Concurrency within the log is a per-instance `threading.Lock` over a WAL database with a 10 s busy timeout, one cached `Logstream` per palace path per process (`logstream.py:437`, `mcp_server.py:333-334, 1120-1141`). + +--- + +## 5. Not everything belongs in the log + +RFC 001 §5 draws this line for wings; the same discipline applies between the two stores. The short form, expanded for operators in `docs/fleet-memory.md`: + +| Put it in | When | +|---|---| +| A **drawer** (palace) | It will still be true, and worth finding, next month. Nobody in particular needs to act. Retrieval is by meaning. | +| An **event** (log) | A named agent must act, reply, or be stopped. Retrieval is by address and order. | +| **Both** | The durable finding goes in a drawer; the event says "there is a new finding, here is the drawer id". This is the recommended pattern for anything a peer must *know* rather than *do*. | + +The failure mode in each direction: a finding filed only as an event is invisible to semantic search and will be re-derived by the next agent; an ask filed only as a drawer is addressed to nobody and will be found, if ever, by accident. + +--- + +## 6. Security model + +**The log authenticates the fleet, not the agent.** Precisely: + +| Control | State | Evidence | +|---|---|---| +| Transport auth | One shared bearer token, `hmac.compare_digest`, 401 otherwise; required for non-loopback binds | `mcp_server.py:7137-7141, 7685` | +| `from_agent` authenticity | ❌ **Not checked against anything.** Shape-validated only; any client may append an event claiming to be any agent | `logstream.py:103-122` is the only check | +| Read authorization | ❌ **None.** Any token holder may list every event addressed to anyone, and fetch any artifact by id — including patch contents | no scoping in the query path | +| Stream auth | ✅ SSE follows the same bearer policy: *"Events and artifacts expose work metadata and patch contents"* | `mcp_server.py:7189-7194` | +| Transport security | TLS via `--tls-cert`/`--tls-key` (both or neither); Host pin against DNS rebinding; Origin check; 16 MiB request cap | `mcp_server.py:6920-6947` | + +For a single-operator fleet behind one token this is adequate, and it must be stated rather than implied, because two useful consequences follow directly from it: + +1. **Impersonation is trivial** — do not treat `from_agent` as evidence of origin in any security decision. +2. **That same property is the only way to test the mailbox.** A positive control requires writing an event *from* an agent you are not, so that your own client does not self-filter it. This is a supported technique precisely because the field is unauthenticated (**measured 2026-08-26**). + +If per-agent identity is ever needed, the correct home is the same place RFC 001 §7.3 puts provenance: the authenticated credential at the boundary, not a self-asserted field. + +--- + +## 7. Landmines + +### 7.1 There is no idempotency guard on append — a retried write duplicates + +`append_event` mints a fresh id and INSERTs, with no dedup lookup of any kind (`logstream.py:625-716`). `put_artifact` likewise. Two byte-identical calls produce two events with different ids and different `seq`. + +The asymmetry is stark: the *replication* path is rigorously idempotent (`logstream.py:1174-1180` checks `id` **or** `(origin_replica, origin_seq)` before applying, and returns early for its own echoed ops). Peer replay is safe; **client retry is not.** + +**Action:** this makes the palace's general rule — *a timeout usually means the write completed; verify, don't retry* — load-bearing rather than advisory. For a drawer, a blind retry costs a dedup-detectable duplicate. For an event it silently forks a coordination thread into two ids, and a mailbox will then show two owed items that must each be closed. After any `event_append` / `patch_submit` timeout, verify with `event_list(correlation_id=…)` before re-issuing. + +### 7.2 `status="open"` is not owed-ness + +Covered in §3.3 and repeated here because it is the mistake most likely to be made by someone reading only the tool descriptions: the mailbox query `to_agent= status=open` **never shrinks as you work**. Treating its length as a to-do count means re-answering answered asks forever. + +**Action:** derive per §3.3, or use a client that does. + +### 7.3 `seq` is local arrival order and is meaningless across replicas + +`seq` is the local `rowid`. The moment a second replica exists, a remote event that was *authored* earlier can arrive *later* and receive a higher `seq`. The owed-set derivation in §3.3 compares `seq`, so it is sound only on a single replica. + +**Action:** on a hub-and-spoke fleet (all clients on one replica) this is correct today. `hlc` is the field to switch to when `mesh_peers` reports actual peers — and that switch must happen *before* multi-replica, not after. + +### 7.4 A broadcast reaches no mailbox, and `to_agent=NULL` reaches nobody at all + +`to_agent=` matches `x` **or** `'*'` at the SQL level (`logstream.py:958-960`), so a broadcast *is* visible to a listing agent. But the reference owed-set derivation excludes `to_agent='*'` deliberately — a broadcast owes nobody a reply, and if it entered every mailbox, every machine would think it personally owed the same answer. + +Two consequences: **broadcasting an ask reaches no owed set at all** (the "don't broadcast an ask" anti-pattern is mechanically enforced, not merely advised), and to reach a whole fleet with something actionable you must write **one directed event per device**, sharing a `correlation_id` so the thread stays joinable. Separately, an event written with **no** `to_agent` matches no `to_agent=` query ever and is addressed to nobody. + +### 7.5 `since_event_id` raises on an unknown id + +An unresolvable cursor raises `ValueError` rather than returning empty (`logstream.py:966-970`) — stronger than the tool description promises, and benign on one replica. On a second replica, a resuming watcher whose cursor has not yet replicated will **hard-fail instead of waiting**. + +### 7.6 A rejected append returns HTTP 200 + +Failures come back as `{"success": false, "error": …}` in a 200 response, not a JSON-RPC error (`mcp_server.py:4581-4584`). + +**Action:** a client that checks only transport status silently loses the event. Check the payload. + +### 7.7 The same agent name on two devices is indistinguishable + +No uniqueness, no registry, no warning. Both machines' events carry one `from_agent`, and each machine's "you cannot owe yourself" filter will discard the other's asks. This is the failure the `@` convention exists to prevent (§3.2). + +### 7.8 `GET /logstream/events` does not exist + +The complete GET route table is `/healthz`, `/statusz`, `/logstream/stream`, `/sync/{version_vector,ops,artifact,peers}`; POST serves `/mcp` only. A recursive grep for `/logstream` in the package yields three hits, all `/logstream/stream`. + +⚠️ **Correction.** An earlier measurement observed `/logstream/events`, `/logstream/stream` and `/sync/peers` all returning 404 against a deployment and attributed all three to a reverse proxy exposing only `/mcp`. That inference was right for `/sync/*` and `/logstream/stream` and **wrong for `/logstream/events`**, which would 404 on a directly-reachable server too. Distinguish "route absent" from "route blocked" before blaming infrastructure. + +### 7.9 Dogfood scars, preserved because each cost someone a session + +- A `patch` artifact stored without its trailing newline truncates the last hunk line — hence the advisory warnings (`logstream.py:718-742`). +- `--type Task.Request` was rejected while `--type Task.Request --type patch.ready` was *silently accepted and matched nothing*, leaving a watcher waiting forever: single-valued filters were validated by pushdown, multi-valued ones compared raw (`sanitize_watch_spec`, `logstream.py:245-278`). `type` is lowercase-only for this reason. +- `event_wait` rejected a `limit` that `event_list` accepted — *"reported by windows-codex during dogfood"* (`mcp_server.py:4666-4670`). + +### 7.10 The main database file's mtime is not a liveness signal + +WAL. **Measured 2026-08-26** on the primary: `logstream.sqlite3` mtime three days old while `logstream.sqlite3-wal` was 865 KiB and seconds old. An operator checking whether coordination is live must look at the `-wal` file, or query. + +--- + +## 8. Phasing + +### 8.1 Implemented + +| Phase | Deliverable | State | +|---|---|---| +| A | `events`/`artifacts`/`event_artifacts` + append/list/ack + size limits | ✅ 3.7.x | +| B | Long poll (`event_wait`) and SSE push | ✅ | +| C | `patch_submit` convenience + patch advisories | ✅ | +| D | HLC on every event; `(origin_replica, origin_seq)` unique index | ✅ 3.8.0 | +| E | Client-side auto-delivered mailbox | ✅ toolkit `5b8d78f`; first live delivery measured 2026-08-26 | + +### 8.2 Deferred, and what RFC 004 owes + +Replication exists in code — `version_vector()`, `list_ops()`, `apply_remote_event()`, `apply_remote_artifact()` and the four `GET /sync/*` routes (`logstream.py:1096-1261`, `mcp_server.py:7206-7256`) — and is attributed in comments to "RFC 004 step 0". **That document does not exist.** Until it does, this RFC records the two properties a reader most needs: the apply path *is* idempotent (§7.1), and switching the owed-set derivation from `seq` to `hlc` is a precondition for a second replica, not a follow-up (§7.3). + +--- + +## 9. Open decisions + +1. **Retention.** No TTL, compaction or pruning exists, and `mempalace sync` does not touch the log (§2). For a fleet log this is mostly a feature — nothing is lost by being offline for weeks — but every 4 MiB artifact is permanent. Decide a policy before the log outgrows a comfortable backup, or decide explicitly that permanence is the policy. +2. **Idempotency key.** Should `event_append` accept an optional client-supplied dedupe key so a retried call is a no-op? This is the one change that would make §7.1 disappear. +3. **`hlc` as the mailbox join key**, replacing `seq`. Required before a second replica (§7.3). Cheap now, breaking later. +4. **Per-agent identity.** Do we ever want `from_agent` to be authenticated (§6), or is "authenticates the fleet, not the agent" the permanent contract? +5. **Should `GET /logstream/events` exist?** A read-only HTTP tail would let non-MCP consumers (dashboards, CI) follow a stream without an MCP client. Today they must hold SSE or speak MCP. +6. **Undocumented limits.** The 64 KiB `metadata` cap and the lowercase-`type` regex are enforced server-side but absent from the tool schemas (§3.1). Document, or relax. +7. **Agent-name registry.** Nothing prevents two devices sharing one `from_agent` (§7.7). A warning at append time would be cheap. + +--- + +## 10. Evidence index + +| Claim | Where | +|---|---| +| Store is `logstream.sqlite3` inside the palace dir | `logstream.py:51`; `mcp_server.py:1108-1117` | +| Full DDL, indexes, WAL | `logstream.py:447-497`, `552-556` | +| Design constraints (quoted in §3) | `logstream.py:1-18` | +| Event id format, ordering not carried by id | `logstream.py:92-102` | +| `seq` = rowid, surfaced at read | `logstream.py:580-586` | +| `origin_seq` assigned in-transaction | `logstream.py:700-706` | +| `hlc` populated per append; format and sort property | `logstream.py:668`; `hlc.py:1-21` | +| Cursor stays local, HLC is merge order | `hlc.py:19-21` | +| `origin_replica` from palace dir | `logstream.py:422-436`; `replica.py:31-49` | +| Server-side `created_at`, second precision | `logstream.py:667` | +| No dedup on append; remote apply *is* idempotent | `logstream.py:625-716` vs `1174-1180` | +| Artifact sha256 non-unique index | `logstream.py:447-497`, `744-810` | +| `artifact_ids` referential check | `logstream.py:687-694`, `1181-1189` | +| Status set; artifact kinds; size limits | `logstream.py:67-69`, `55-57`, `765-778` | +| `type` regex, lowercase ≤64 | `logstream.py:76, 124-134` | +| Broadcast matching in SQL | `logstream.py:958-960` | +| `since_event_id` anchor + raise | `logstream.py:963-978` | +| Order by rowid; limit clamp 500 | `logstream.py:947` | +| `preview` truncates body to 200 chars | `mcp_server.py:4587-4607` | +| `event_wait` backoff, cap, timeout result | `logstream.py:1004-1039` | +| `ack_event` appends, never mutates; `ack_of` in metadata | `logstream.py:812-845` | +| Patch advisories | `logstream.py:718-742` | +| Lock exemptions, both, with rationale | `mcp_server.py:6887-6902`, `434-449` | +| Bearer auth; SSE auth required | `mcp_server.py:7137-7141`, `7189-7194` | +| SSE implementation, heartbeat, client cap | `mcp_server.py:7570-7671` | +| Rejected append returns 200 + `success:false` | `mcp_server.py:4581-4584` | +| `sync` does not touch the log | `cli.py:1057`; `sync.py` (no logstream refs) | +| No retention/pruning anywhere | no `DELETE FROM events`/`artifacts` in the package | +| Replication surface, attributed to RFC 004 | `logstream.py:1096-1261`; `mcp_server.py:7206-7256` | + +### 10.1 Where the shipped tool descriptions and the code disagree + +| Tool-description claim | Verdict | +|---|---| +| "append-only; corrections are new events" | True of content. Precisely: no event's *content* is ever mutated; `append_event` does issue one `UPDATE … SET origin_seq = rowid` on the row it just inserted, and the 3.8.0 migration backfills `origin_replica`/`origin_seq`/`hlc` on pre-existing rows (`logstream.py:504-550`). | +| "`since_event_id` cannot skip anything" | True, and stronger than stated — it raises on an unknown id (§7.5). | +| "body max 256 KiB", "artifact max 4 MiB, UTF-8 only", "timeout default 60 s max 5 min" | All true; the timeout clamps silently rather than erroring. | +| "`to_agent=` also matches `*` broadcasts" | True, SQL-level — but see §7.4 for why a broadcast still reaches no mailbox. | +| "prefer the push stream at `GET /logstream/stream`" | True; route exists and requires auth. | +| `GET /logstream/events` | ❌ Does not exist (§7.8). | +| `metadata` size limit; `type` charset | ❌ Enforced but undocumented (§3.1). | + +--- + +## 11. See also + +- `docs/fleet-memory.md` — the operator-facing companion: what the palace's stores are *for*, and when to use a drawer versus an event. +- `docs/rfc-001-global-palace.md` — §5 what should not be global, §7.2 the `sync` hazard, §7.3 provenance at the boundary, §7.5 single-writer expectations. +- `docs/rfc-002-joiner.md` — replaying a second palace into a shared primary. +- `extensions/pi/README.md` §2 — the client-side mailbox: delivery points, gating, and why it is a client concern. +- `~/.agents/skills/mempalace/SKILL.md` §"Cross-Machine Coordination" — **normative for agent behaviour**; this RFC is normative for mechanism. diff --git a/extensions/pi/README.md b/extensions/pi/README.md index 359bc74..a1ce2d7 100644 --- a/extensions/pi/README.md +++ b/extensions/pi/README.md @@ -342,12 +342,21 @@ the mechanism; the skill is normative for behaviour.** an SSE endpoint (`GET /logstream/stream`, `text/event-stream` in `mempalace/mcp_server.py`), but a deployment may expose only the MCP endpoint through its reverse proxy — verified 2026-08-26 against -`https://mempalace.jordbo.se`, where `/logstream/events`, `/logstream/stream` -and `/sync/peers` all return 404 while `/mcp` serves normally. Where that is the +`https://mempalace.jordbo.se`, where `/logstream/stream` and `/sync/peers` +return 404 while `/mcp` serves normally. Where that is the case, polling through the existing MCP client is the only available path — which is what the mailbox in §2 does — and enabling SSE means a proxy route plus an auth decision, not an extension change. +> ⚠️ **Corrected 2026-08-26.** An earlier revision of this paragraph listed +> `/logstream/events` alongside those two as proxy-blocked. That route **does not +> exist in the server at all** — the complete GET table in mempalace 3.8.0 is +> `/healthz`, `/statusz`, `/logstream/stream` and `/sync/{version_vector,ops,artifact,peers}`, +> so `/logstream/events` would 404 against a directly-reachable server too. The +> proxy inference was right for the other two and wrong for that one; see +> [RFC 003 §7.8](../../docs/rfc-003-coordination-log.md). Distinguish *route absent* +> from *route blocked* before blaming infrastructure. + As with stamping, all of this is inert unless `MEMPALACE_REMOTE_URL` points at a shared palace. On a solitary palace the event tools work fine and the log contains only this machine's own events.