docs: write the RFC 003 that the code has been citing all along, plus an operator-facing fleet-memory guide
`logstream.py`'s module docstring is headed "Agent coordination event log for MemPalace (RFC 003)" and enumerates five "Design constraints (RFC 003)". Comments cite "RFC 003 phase 5", "RFC 003 suggested defaults" and "the first RFC 003 dogfood". Every event/artifact tool description cites RFC 003. The document has never existed — confirmed by searching this repo and the primary host. So this is a retrospective spec: it transcribes what the implementation already believes, and records what it does NOT do. docs/rfc-003-coordination-log.md, verified line-by-line against mempalace 3.8.0 (every claim cites file.py:LINE, indexed in §10 for re-verification). The parts that are not visible from the tool descriptions: - No idempotency guard on event_append or put_artifact (§7.1). The replication path checks `id OR (origin_replica, origin_seq)` before applying; the client path checks nothing. So peer replay is safe and CLIENT RETRY IS NOT — a retried append forks a coordination thread into two ids. This makes the fleet's "a timeout is not a failure, verify before retrying" rule load-bearing rather than advisory. - Coordination traffic is deliberately exempt from BOTH palace locks (§4) — _HTTP_LOCK_FREE_TOOLS and _PEER_WRITER_EXEMPT_TOOLS, each with its own rationale in-source. "One large mine blocks every client" is true of drawer writes and false of coordination writes. - `mempalace sync` never touches the log, and no DELETE FROM events exists anywhere (§2, §9.1) — answering for the logstream a question RFC 001 §7.2 left open, and making the log permanent and unbounded. - from_agent is shape-validated and never authenticated; there is no read scoping at all (§6). The log authenticates the fleet, not the agent. That is simultaneously the security limitation and the only way to positive-control the mailbox. - Three distinct orderings — seq (local arrival rowid), origin_seq (author's counter), hlc (fleet-wide, lexicographically sortable) — and origin_replica identifies the PALACE, not the writer, which is why from_agent/to_agent carry the whole distinction between machines (§3.2). - GET /logstream/events does not exist and never did (§7.8). An earlier measurement saw it 404 and blamed the reverse proxy; that inference was right for /sync/* and /logstream/stream and wrong for this one. Corrected in extensions/pi/README.md §3 too, in place, dated. docs/fleet-memory.md is the operator-facing companion the repo lacked entirely: the front-door README had zero mentions of coordination, so the channel was undiscoverable unless a message happened to arrive. It covers the five stores and what each is for (drawers/wings/rooms, diaries, KG, palace graph, coordination log), what a central palace buys a fleet — awareness, non-repetition of expensive work, and retractions that travel — and a decision flow for drawer vs KG vs event. Six mermaid diagrams, all rendered and inspected as images, not merely parsed: the first pass produced a truncated state label from an HTML entity and a self-loop that drew a meaningless dotted lasso. Validation says "no syntax error"; only looking says "correct". Also: a Documentation table in the top-level README so all of the above is reachable from the front door. Deliberately host-agnostic, per synlig-primary-runbook.md's precedent — no device names, hostnames or operator names in either new document.
This commit is contained in:
@@ -58,6 +58,27 @@ Who owns what:
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## Documentation
|
||||||
|
|
||||||
|
Start here if you are deciding **what to put where**, or running MemPalace on more than one machine:
|
||||||
|
|
||||||
|
| Document | What it answers |
|
||||||
|
|---|---|
|
||||||
|
| [`docs/fleet-memory.md`](docs/fleet-memory.md) | **Start here.** The five stores MemPalace keeps and what each is for; what a *central* palace buys a fleet of machines; when to file a memory versus when to message another device. Diagrams, worked examples. |
|
||||||
|
| [`docs/rfc-003-coordination-log.md`](docs/rfc-003-coordination-log.md) | The inter-agent / inter-device coordination log (`logstream`): storage, append and query semantics, delivery and latency, the trust model, and its landmines. |
|
||||||
|
| [`docs/rfc-001-global-palace.md`](docs/rfc-001-global-palace.md) | Why and how a palace is centralised, what should *not* be global, and the Phase-0 landmines (including: never run `mempalace sync` against a shared palace). |
|
||||||
|
| [`docs/rfc-002-joiner.md`](docs/rfc-002-joiner.md) | Replaying a second palace into a shared primary, and what dedupes what. |
|
||||||
|
| [`docs/phase-1-exposure-runbook.md`](docs/phase-1-exposure-runbook.md) | Exposing a palace over HTTP: the Host/Origin pin, tokens, TLS. |
|
||||||
|
| [`docs/backup-and-recovery.md`](docs/backup-and-recovery.md) | Why a palace needs its own backup procedure, and how to restore one. |
|
||||||
|
| [`extensions/pi/README.md`](extensions/pi/README.md) | The pi-side client: provenance stamping at the edge, and the auto-delivered mailbox. |
|
||||||
|
|
||||||
|
Behaviour for the *agents* is normative in the `mempalace` skill (`SKILL.md` here,
|
||||||
|
installed to `~/.agents/skills/mempalace/`), not in these documents — the split is
|
||||||
|
deliberate: the skill says what an agent should do, the docs say what the machinery
|
||||||
|
actually does.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## Why this exists
|
## Why this exists
|
||||||
|
|
||||||
MemPalace is the agent memory layer. Its stock CLI has two gaps that bite on a machine running opencode with a docs-first palace policy:
|
MemPalace is the agent memory layer. Its stock CLI has two gaps that bite on a machine running opencode with a docs-first palace policy:
|
||||||
|
|||||||
@@ -0,0 +1,218 @@
|
|||||||
|
# Fleet memory — what MemPalace stores, and what to put where
|
||||||
|
|
||||||
|
**Audience:** the person running MemPalace on more than one machine, or thinking about it.
|
||||||
|
**Companion documents:** `docs/rfc-003-coordination-log.md` for the coordination log's mechanism, `docs/rfc-001-global-palace.md` for the centralisation design, and `~/.agents/skills/mempalace/SKILL.md` for what the *agents* are told to do.
|
||||||
|
|
||||||
|
MemPalace is usually described as "memory for agents", which is true and not very actionable. It is really **five stores with different retrieval models**, and most of the value — especially across a fleet — comes from putting each kind of thing in the store whose retrieval model matches how you will want it back.
|
||||||
|
|
||||||
|
This document covers: what the stores are, what a central palace changes when several machines share one, and how to decide between filing a memory and sending a message.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. The stores at a glance
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
Q["ask by MEANING<br/>'what do we know about X?'"] --> D["<b>Drawers</b><br/>chroma.sqlite3<br/>verbatim text, embedded"]
|
||||||
|
R["ask by ENTITY + TIME<br/>'what was true in June?'"] --> G["<b>Knowledge graph</b><br/>knowledge_graph.sqlite3<br/>typed facts, time-bounded"]
|
||||||
|
S["ask by ADDRESS + ORDER<br/>'what is waiting for me?'"] --> L["<b>Coordination log</b><br/>logstream.sqlite3<br/>addressed events + artifacts"]
|
||||||
|
T["ask by ASSOCIATION<br/>'what else touches this?'"] --> P["<b>Palace graph</b><br/>tunnels.json + hallways.json<br/>links between rooms"]
|
||||||
|
D --> DI["<b>Diaries</b> live here too:<br/>drawers with room=diary,<br/>read by recency, not similarity"]
|
||||||
|
```
|
||||||
|
|
||||||
|
All four stores are files inside **one palace directory**, so "the palace" is a directory you can back up in one go.
|
||||||
|
|
||||||
|
| Store | You get things back by | Typical use | Wrong use |
|
||||||
|
|---|---|---|---|
|
||||||
|
| **Drawers** (wings → rooms) | semantic similarity | a verbatim finding, a decision and its reasoning, a runbook, a transcript excerpt | anything a specific machine must *act* on; anything whose value is its exact byte content |
|
||||||
|
| **Diaries** (drawers with `room="diary"`) | agent + recency | "what did I do last session, and what did it feel like" — the continuity thread across sessions | facts other agents need to find by searching; a diary is read by *its author*, chronologically |
|
||||||
|
| **Knowledge graph** | entity, relationship, point in time | facts that *change*: versions, employers, who owns what, an injury that heals | prose, reasoning, anything you'd want to read rather than query |
|
||||||
|
| **Coordination log** | address, correlation, append order | "device B must review this patch"; "this claim is retracted, stop building on it" | durable knowledge — an event is invisible to semantic search |
|
||||||
|
| **Palace graph** | traversal from a room | discovering that an API design in one project touches a schema in another | primary storage — it links drawers, it does not hold content |
|
||||||
|
|
||||||
|
Two smaller files exist and are implementation detail, not user surface: `sqlite_exact.sqlite3` (an exact-match index over drawer metadata) and, if the daemon runs, `queue.sqlite3` (its job queue).
|
||||||
|
|
||||||
|
### 1.1 Drawers: wings and rooms
|
||||||
|
|
||||||
|
A **wing** is a project or domain; a **room** is an aspect within it. `wing="pi-devbox", room="landmines"` is a good pair; `wing="misc", room="stuff"` is how a palace becomes a landfill. Content is stored **verbatim and chunked** — never summarised — and retrieved by embedding similarity, so a drawer is found by someone who *doesn't already know it exists*. That is the property to optimise for: write the drawer that the next person's search will match.
|
||||||
|
|
||||||
|
The one counter-intuitive consequence: **fresh drawers rank worst.** A drawer filed an hour ago has no advantage in a similarity search, and a well-worn older drawer will outrank it. For anything recent, enumerate by date (`list_drawers(since=…)`) instead of searching.
|
||||||
|
|
||||||
|
### 1.2 Diaries
|
||||||
|
|
||||||
|
A diary entry is a drawer with `room="diary"`, filed by default into `wing_<agent>`, tagged with the writing agent. It is the *first-person* record: what I did, what surprised me, what I would do differently. Agents are told to write one before a session ends, and to read the last few at session start.
|
||||||
|
|
||||||
|
In a fleet this is the highest-signal store per byte, for a reason that is easy to miss: a diary entry is the only place that records **what did not work**. A drawer tends to record the conclusion; the diary records the three hours that produced it.
|
||||||
|
|
||||||
|
### 1.3 Knowledge graph
|
||||||
|
|
||||||
|
Triples — subject, predicate, object — with `valid_from` / `valid_to`, so a fact can *stop* being true without being deleted. `supersede` replaces a single-valued fact at a shared boundary, so a point-in-time query at that instant returns exactly one value.
|
||||||
|
|
||||||
|
Use it for anything you will later want to ask "what was true at time T?" about: which version was released when, who owned a service, what model an assistant was using. Do not use it for prose — a triple whose object is a paragraph is a drawer wearing a costume.
|
||||||
|
|
||||||
|
### 1.4 The coordination log
|
||||||
|
|
||||||
|
Addressed, ordered, exact. `mempalace_event_*` carries messages between agents; `mempalace_artifact_*` carries byte-exact payloads (patches, logs, files) that events can reference. This is the only store where one machine can *reach* another.
|
||||||
|
|
||||||
|
Its full mechanism, limits and landmines are in `docs/rfc-003-coordination-log.md`. §4 below covers what an operator needs to decide.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. What changes when a fleet shares one palace
|
||||||
|
|
||||||
|
A single machine's palace is a notebook. A shared palace is something different in kind: **the fleet stops being a set of independent agents that each learn the same lessons separately.**
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
A["laptop<br/>agent session"] -->|MCP over HTTPS| H(("central palace<br/>one server"))
|
||||||
|
B["workstation<br/>agent session"] -->|MCP over HTTPS| H
|
||||||
|
C["build box<br/>agent session"] -->|MCP over HTTPS| H
|
||||||
|
H --> D["drawers + diaries"]
|
||||||
|
H --> G["knowledge graph"]
|
||||||
|
H --> L["coordination log"]
|
||||||
|
```
|
||||||
|
|
||||||
|
Note the topology: in the common deployment the machines are **thin clients of one server**, not peer replicas. Everything one machine writes is immediately visible to the others — there is no sync delay to reason about, and equally no local copy to fall back on when the server is unreachable. (`docs/rfc-001-global-palace.md` §4 designs an edge proxy with a local palace and a durable outbox for deployments that need to keep working offline; the plain thin-client shape above does not.)
|
||||||
|
|
||||||
|
### 2.1 The three things this actually buys
|
||||||
|
|
||||||
|
**Awareness.** "What has anyone been doing?" becomes answerable. Each machine's diary is readable by every other machine, so an agent starting work on a shared project can see that another machine spent yesterday on it, and how far it got.
|
||||||
|
|
||||||
|
**Non-repetition of expensive work.** This is the biggest measurable win. Anything that cost real time to obtain — a scraped API surface, a spec read end to end, a bisect, a benchmark, a long investigation into why a build fails on one platform — is filed once and searchable everywhere. The second machine's cost drops from hours to one search.
|
||||||
|
|
||||||
|
**Mistakes and retractions travel.** The subtle one, and the reason a shared palace is worth more than a shared wiki. When a machine discovers that a belief was *wrong*, it can file the retraction where every other machine will hit it. Without that, each machine independently rediscovers the same dead end — and worse, a machine can spend a day rebuilding something another machine already proved doesn't work.
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
sequenceDiagram
|
||||||
|
participant W as workstation
|
||||||
|
participant P as central palace
|
||||||
|
participant L as laptop, asleep 9 days
|
||||||
|
W->>W: spends 3h finding why the build breaks
|
||||||
|
W->>P: drawer — the finding, verbatim, with evidence
|
||||||
|
W->>P: KG fact — "component X requires flag Y"
|
||||||
|
W->>P: diary — what was tried and failed
|
||||||
|
Note over L: ...9 days pass, laptop asleep...
|
||||||
|
L->>P: wake-up: read diaries + search before starting
|
||||||
|
P-->>L: the finding, the failed attempts, the fact
|
||||||
|
Note over L: cost: one search instead of 3 hours
|
||||||
|
```
|
||||||
|
|
||||||
|
### 2.2 The two costs, stated plainly
|
||||||
|
|
||||||
|
**Everything is visible to everyone.** One shared token, no per-agent read scoping. Anything filed into a shared palace should be considered readable by every machine and every agent on it. Do not put secrets in drawers.
|
||||||
|
|
||||||
|
**Provenance stops being obvious.** On a single machine, every drawer is yours and every path exists. On a shared palace, most drawers came from other machines and most `source_file` paths **do not exist locally**. Two consequences worth internalising:
|
||||||
|
|
||||||
|
- A file path in a drawer is evidence about *some* machine, not necessarily this one.
|
||||||
|
- ⚠️ **Never run `mempalace sync` against a shared palace.** It prunes drawers whose source files look gitignored, deleted or moved — which on a shared palace describes most of the content, including every other machine's. See `docs/rfc-001-global-palace.md` §7.2. The coordination log is *not* affected by this (RFC 003 §2), but drawers very much are.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Deciding where something goes
|
||||||
|
|
||||||
|
The question that matters is not "is this important?" but **"how will I want this back, and does anyone need to act?"**
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart TD
|
||||||
|
Start["I have something worth keeping"] --> Act{"Must a specific<br/>machine or agent<br/>DO something?"}
|
||||||
|
Act -->|no| Change{"Is it a fact that<br/>changes over time?"}
|
||||||
|
Act -->|yes| Know{"Do they also need<br/>to KNOW it later?"}
|
||||||
|
Change -->|yes| KG["Knowledge graph<br/>kg_add / kg_supersede"]
|
||||||
|
Change -->|no| Mine{"Is it about MY session<br/>— what I tried, felt, learned?"}
|
||||||
|
Mine -->|yes| Diary["Diary entry"]
|
||||||
|
Mine -->|no| Drawer["Drawer<br/>wing + room, verbatim"]
|
||||||
|
Know -->|yes| Both["BOTH:<br/>drawer for the knowledge,<br/>event pointing at it"]
|
||||||
|
Know -->|no| Event["Coordination event<br/>to_agent = the specific agent"]
|
||||||
|
```
|
||||||
|
|
||||||
|
Worked examples:
|
||||||
|
|
||||||
|
| Situation | Where | Why |
|
||||||
|
|---|---|---|
|
||||||
|
| "The release takes 76 min, and 40 of those are the base image build." | Drawer | durable, nobody must act, next person finds it by searching "release timing" |
|
||||||
|
| "v1.8.9 is the released version, as of this timestamp." | KG (`supersede`) | it will change; you will want "what was released in August?" |
|
||||||
|
| "I spent two hours chasing a watcher that was already dead." | Diary | first-person, chronological, tells the next session what *not* to retry |
|
||||||
|
| "Build box: this patch is ready, please review and apply." | Event (directed) | a named machine must act; the patch itself goes in as an artifact |
|
||||||
|
| "The claim in that drawer is wrong — I measured the opposite." | Both | file the corrected finding as a drawer, then an event so the machine building on it stops |
|
||||||
|
| "Everyone should know the new toolkit is live." | Drawer + broadcast event | the drawer is what anyone will *find*; the broadcast is a notice, not an ask (§4.2) |
|
||||||
|
|
||||||
|
The failure mode in each direction is worth naming, because both are common:
|
||||||
|
|
||||||
|
- **A finding filed only as an event** is invisible to semantic search. Nobody will ever find it again, and the next agent will re-derive it.
|
||||||
|
- **An ask filed only as a drawer** is addressed to nobody. It will be found, if ever, by accident — long after it mattered.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. What the coordination log can and cannot do
|
||||||
|
|
||||||
|
This is the part most likely to be mis-set expectations, so it is worth being blunt: **it is a durable log, not a chat.** Nothing is listening. Events are appended and persist; there is no delivery window; nothing is lost by being offline when one is written. A message waits indefinitely, and your reply waits just as patiently for a sender who has since gone away.
|
||||||
|
|
||||||
|
That sounds like a limitation and is actually the correct design for a fleet where few machines are awake at once and any given machine may sleep for weeks.
|
||||||
|
|
||||||
|
### 4.1 Latency, honestly
|
||||||
|
|
||||||
|
| Recipient state | When they see it | Notes |
|
||||||
|
|---|---|---|
|
||||||
|
| In a live session, mailbox-enabled client | **≈2–5 minutes** | measured ≈2–3 min on first live delivery; a poll floor of 5 min applies between checks |
|
||||||
|
| Holding an SSE connection (`GET /logstream/stream`) | sub-second | for daemons/dashboards, not interactive agents |
|
||||||
|
| Asleep, next session tomorrow | tomorrow | delivered in the session-start wake-up |
|
||||||
|
| Asleep for three weeks | in three weeks | nothing expires; the log is permanent |
|
||||||
|
| Never runs again | never | there is no re-routing and no dead-letter path |
|
||||||
|
|
||||||
|
So: appropriate for "handle this when you next wake", "here is a patch", "stop building on that claim". Not appropriate for anything with a deadline inside the hour, unless you know the recipient is awake.
|
||||||
|
|
||||||
|
One pleasant property, worth knowing because it is counter-intuitive: coordination traffic is **exempt from the palace's write lock**, so you can message another machine and it can reply *while* a long mine is running on the server (RFC 003 §4). Coordination stays alive when memory writes are blocked.
|
||||||
|
|
||||||
|
### 4.2 The one thing that does not work: broadcasting an ask
|
||||||
|
|
||||||
|
You can write a broadcast (`to_agent="*"`) and every machine that *lists* events will see it. But a broadcast **never enters any machine's mailbox** and is never auto-delivered — by design, because "everyone owes this answer" degenerates into either N duplicate replies or nobody acting.
|
||||||
|
|
||||||
|
- **Broadcast** = a notice on a wall. Fine for "v1.8.9 is out".
|
||||||
|
- **Directed event** = a message in a named mailbox. Required for anything that must be done.
|
||||||
|
|
||||||
|
To reach a whole fleet with something actionable, **fan out**: one directed event per device, sharing one `correlation_id` so the thread stays joinable. Each machine then owes its own reply.
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
flowchart LR
|
||||||
|
You["you"] -->|"to_agent=pi@laptop"| A["laptop owes a reply"]
|
||||||
|
You -->|"to_agent=pi@workstation"| B["workstation owes a reply"]
|
||||||
|
You -->|"to_agent=pi@build-box"| C["build box owes a reply"]
|
||||||
|
A --> Corr["one shared correlation_id<br/>joins the three threads"]
|
||||||
|
B --> Corr
|
||||||
|
C --> Corr
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4.3 Addressing, and why the format matters
|
||||||
|
|
||||||
|
Addresses are `<harness>@<device>` — `pi@laptop`, `opencode@build-box`. Two rules follow:
|
||||||
|
|
||||||
|
1. **An unstamped client is unreachable.** If a machine's events say `from_agent: pi` with no device, nobody can address it, because "pi" is every machine.
|
||||||
|
2. **Two machines must never share one address.** Nothing prevents it, nothing warns, and the result is that each silently discards the other's asks as its own (RFC 003 §7.7).
|
||||||
|
|
||||||
|
### 4.4 A reply is owed until it is *terminally* closed
|
||||||
|
|
||||||
|
An acknowledgement does not close a thread. Neither does "claimed" or "ready". Only a terminal status — `applied`, `superseded`, `failed`, `blocked` — clears an ask from the recipient's mailbox. Until then, a mailbox-enabled client will keep resurfacing it, which is the intended behaviour: an unanswered ask should nag.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Habits that make a shared palace work
|
||||||
|
|
||||||
|
Small, and the whole value rests on them:
|
||||||
|
|
||||||
|
1. **Search before you answer, and enumerate before you conclude.** One empty search is not proof of silence — fresh drawers rank worst, so for anything from the last couple of days list by date and read the other machines' diaries.
|
||||||
|
2. **Write the diary entry before the session ends.** It is the store other machines learn from fastest, and the only one that records failed attempts.
|
||||||
|
3. **File retractions as loudly as findings.** "I was wrong about X, here is the measurement" is worth more than a new finding, because it stops N machines repeating a dead end.
|
||||||
|
4. **Check the mailbox at wake-up, even when you expect nothing.** An empty result costs one call. Silence is only informative once you know delivery works.
|
||||||
|
5. **Say which machine you are talking about.** On a shared palace, "the container" and "the host" are ambiguous and a path is not self-identifying.
|
||||||
|
6. **After a write times out, verify — do not blindly retry.** The palace is single-writer for memory writes, so a timeout usually means the write *completed*. For coordination events this matters twice over: there is no idempotency guard, so a retried event forks the thread into two (RFC 003 §7.1).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. See also
|
||||||
|
|
||||||
|
- `docs/rfc-003-coordination-log.md` — the coordination log: storage, semantics, security model, landmines.
|
||||||
|
- `docs/rfc-001-global-palace.md` — how and why a palace is centralised; §5 what should not be global; §7 the landmines, including the `sync` hazard.
|
||||||
|
- `docs/phase-1-exposure-runbook.md` — exposing a palace over HTTP.
|
||||||
|
- `docs/backup-and-recovery.md` — why a palace needs its own backup procedure.
|
||||||
|
- `extensions/pi/README.md` — the pi-side client: provenance stamping and the auto-delivered mailbox.
|
||||||
|
- `~/.agents/skills/mempalace/SKILL.md` — the protocol the agents themselves follow.
|
||||||
@@ -0,0 +1,396 @@
|
|||||||
|
# RFC 003 — The coordination log (`logstream`)
|
||||||
|
|
||||||
|
**Status:** implemented and in production use since 3.7.x. This document is a *retrospective* specification, written after the fact.
|
||||||
|
**Author:** pi (agent), 2026-08-26.
|
||||||
|
**Applies to:** mempalace 3.8.0 (`logstream.py`, 1261 lines), mempalace-toolkit `5b8d78f` (the pi-side mailbox).
|
||||||
|
**Context:** every `mempalace_event_*` and `mempalace_artifact_*` tool description cites "RFC 003". `logstream.py`'s module docstring is headed *"Agent coordination event log for MemPalace (RFC 003)"* and enumerates five "Design constraints (RFC 003)". Inline comments cite "RFC 003 phase 5", "RFC 003 suggested defaults" and "the first RFC 003 dogfood". **No such document has ever existed** — verified 2026-08-26 by searching both the toolkit repository and the primary host. This RFC transcribes the spec the implementation already believes in, and — more usefully — records what it does *not* do.
|
||||||
|
|
||||||
|
Read §7 before you build anything on this log; it is the part that is not obvious from the tool descriptions. Every mechanical claim below cites `file.py:LINE` in mempalace 3.8.0, and §10 indexes them so any claim can be re-verified without re-reading 1261 lines. Claims marked **measured** were executed against the live fleet on the date given; claims without that marker are reads of the source. Where the code and the shipped tool descriptions disagree, §10.1 says so explicitly rather than quietly siding with one.
|
||||||
|
|
||||||
|
One scoping note up front. This RFC covers the **log**: its storage, its append and query semantics, delivery, and the trust model. It does not specify **replication**, which the source already attributes to a different document (`# ── Replication (RFC 004 step 0: logstream multi-master) ──`, `logstream.py:1096`). RFC 004 does not exist either; §8.2 states what it owes.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. The problem
|
||||||
|
|
||||||
|
The palace stores what an agent *knows*: drawers are semantic, retrieved by meaning, and deliberately have no addressee. That shape is wrong for four things a fleet of agents actually needs:
|
||||||
|
|
||||||
|
1. **Addressing.** "This is for the machine that owns the release" cannot be expressed as a drawer. A drawer is found by whoever happens to search for the right words.
|
||||||
|
2. **A reply that closes something.** Semantic memory has no notion of an outstanding question. Nothing in a drawer can be *owed*.
|
||||||
|
3. **Exact payloads.** A unified diff must survive byte-for-byte. Drawers are chunked and embedded; that is a feature for prose and a defect for patches.
|
||||||
|
4. **Order.** "What happened after this?" needs an append cursor, not a similarity score.
|
||||||
|
|
||||||
|
The coordination log adds exactly those four properties and nothing else. It is a second store beside the palace, not a new kind of drawer.
|
||||||
|
|
||||||
|
### 1.1 What it is not
|
||||||
|
|
||||||
|
Stated first, because every misuse of this log so far has come from assuming one of these:
|
||||||
|
|
||||||
|
- **Not a bus.** Nothing subscribes by default; nothing is delivered "live" unless a client is holding an SSE connection or polling. There is no delivery window and nothing is lost by being offline when an event is written.
|
||||||
|
- **Not a chat channel.** Latency is bounded by *when the recipient next runs*, which in a fleet of workstations is hours, days or weeks. §3.5 gives the measured numbers for the case where the recipient is awake.
|
||||||
|
- **Not authenticated per agent.** `from_agent` is a routing label with the trust properties of an e-mail `From:` header (§6).
|
||||||
|
- **Not a queue.** Nothing is consumed, acknowledged-and-removed, or retried. Events are permanent (§9.1) and "handled" is a *derived* property (§3.3).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Verified starting point (2026-08-26, mempalace 3.8.0)
|
||||||
|
|
||||||
|
| Piece | State | Evidence |
|
||||||
|
|---|---|---|
|
||||||
|
| Separate SQLite store, inside the palace dir | ✅ `logstream.sqlite3` beside `chroma.sqlite3`; WAL; dir `chmod 0700` best-effort | `logstream.py:51`, `mcp_server.py:1108-1117`, `logstream.py:430-433` |
|
||||||
|
| `events`, `artifacts`, `event_artifacts` tables | ✅ Created idempotently at open | `logstream.py:447-497` |
|
||||||
|
| Hybrid logical clock on every event | ✅ `hlc` populated on every local append | `logstream.py:668`, `hlc.py:1-21` |
|
||||||
|
| Seven MCP tools (append/list/wait/ack, artifact put/get, patch_submit) | ✅ | `mcp_server.py:4546-4771` |
|
||||||
|
| Server-sent events push | ✅ `GET /logstream/stream`, auth required, 15 s heartbeat, ≤8 clients | `mcp_server.py:7570-7671`, `7189-7194` |
|
||||||
|
| `GET /logstream/events` | ❌ **Never implemented.** Only `/logstream/stream` exists | see §7.8 |
|
||||||
|
| Auto-delivered mailbox (owed-set derivation + injection) | ✅ **Client-side, not server-side** — lives in the toolkit's pi extension | `extensions/pi/mempalace.ts` (toolkit `5b8d78f`) |
|
||||||
|
| Idempotency guard on append | ❌ **Absent.** See §7.1 | `logstream.py:625-716` |
|
||||||
|
| Retention / TTL / compaction | ❌ Absent by design-so-far. See §9.1 | no `DELETE FROM events` in the package |
|
||||||
|
| Interaction with `mempalace sync` | ✅ **None.** `sync` prunes drawers only | `cli.py:1057`, `sync.py` (no logstream references) |
|
||||||
|
|
||||||
|
Two things that sound like this feature and are not:
|
||||||
|
|
||||||
|
- **`mempalace_mesh_peers` is not a fleet roster.** It reports peer *replicas*. A hub-and-spoke deployment — many thin MCP clients of one server — correctly reports `peers: []` while every machine in the fleet is actively writing to the same log. Deciding "coordination does not apply to me" from an empty peer list is a measured failure mode, not a hypothetical one.
|
||||||
|
- **`origin_replica` is not the writing machine.** It is a property of the palace *directory* (§3.2).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Design
|
||||||
|
|
||||||
|
The five constraints in the module docstring (`logstream.py:1-18`) are the design, quoted verbatim because they are already normative in the implementation:
|
||||||
|
|
||||||
|
> - No Chroma dependency, no vector index open — plain SQLite only.
|
||||||
|
> - Append-only: events are immutable; corrections are new events that reference prior events.
|
||||||
|
> - Exact payloads: event bodies and artifact content are stored verbatim.
|
||||||
|
> - Safe under concurrent HTTP requests (WAL + per-instance lock, same pattern as `knowledge_graph.py`).
|
||||||
|
> - Explicit size limits with clear errors, never silent truncation.
|
||||||
|
|
||||||
|
The shape, in the same idiom as RFC 001 §4.1:
|
||||||
|
|
||||||
|
```
|
||||||
|
agent on device A ─┐ palace directory
|
||||||
|
agent on device B ─┼── MCP /mcp ──► server ─┬─ chroma.sqlite3 (drawers: what you know)
|
||||||
|
agent on device C ─┘ │ ├─ knowledge_graph.sqlite3 (facts, temporal)
|
||||||
|
│ └─ logstream.sqlite3 (events + artifacts)
|
||||||
|
GET /logstream/stream events ── append-only, permanent
|
||||||
|
(SSE, auth, ≤8 clients) artifacts ── verbatim, ≤4 MiB, sha256
|
||||||
|
event_artifacts ── join, checked at append
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3.1 Data model
|
||||||
|
|
||||||
|
```sql
|
||||||
|
CREATE TABLE IF NOT EXISTS events (
|
||||||
|
id TEXT PRIMARY KEY,
|
||||||
|
type TEXT NOT NULL,
|
||||||
|
stream TEXT NOT NULL,
|
||||||
|
room TEXT NOT NULL,
|
||||||
|
from_agent TEXT NOT NULL,
|
||||||
|
to_agent TEXT,
|
||||||
|
correlation_id TEXT,
|
||||||
|
branch TEXT,
|
||||||
|
base_commit TEXT,
|
||||||
|
status TEXT,
|
||||||
|
body TEXT NOT NULL DEFAULT '',
|
||||||
|
created_at TEXT NOT NULL,
|
||||||
|
metadata_json TEXT NOT NULL DEFAULT '{}',
|
||||||
|
origin_replica TEXT,
|
||||||
|
origin_seq INTEGER,
|
||||||
|
hlc TEXT
|
||||||
|
);
|
||||||
|
CREATE UNIQUE INDEX IF NOT EXISTS events_origin_seq_idx ON events(origin_replica, origin_seq);
|
||||||
|
```
|
||||||
|
(`logstream.py:447-497`, unique index at `552-556`. `artifacts` and `event_artifacts` in the same script; `artifacts_sha256_idx` is **not** unique — see §3.4.)
|
||||||
|
|
||||||
|
Validation, all server-side, all raising `ValueError` naming the allowed set:
|
||||||
|
|
||||||
|
| Field | Required | Rule | Where |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `type` | yes | `^[a-z0-9][a-z0-9_.-]{0,63}$` — **lowercase only, ≤64** | `logstream.py:76, 124-134` |
|
||||||
|
| `stream`, `room`, `from_agent` | yes | non-empty, ≤256, no control chars | `logstream.py:78, 103-122` |
|
||||||
|
| `status` | no | one of `open claimed ready applied blocked failed superseded` | `logstream.py:67-69, 136-143` |
|
||||||
|
| `body` | no (empty allowed) | ≤256 KiB, no NUL | `logstream.py:55, 145-158` |
|
||||||
|
| `metadata` | no | canonical JSON, **≤64 KiB** | `logstream.py:57, 160-177` |
|
||||||
|
| `artifact_ids` | no | each must already exist, else `ValueError` | `logstream.py:687-694` |
|
||||||
|
|
||||||
|
Three consequences worth stating because they are invisible from the tool descriptions: an event body **may be empty** while artifact content **may not** (`logstream.py:766-767`); `to_agent` is **nullable**, and an event with `to_agent = NULL` is addressed to nobody and can never match a `to_agent=` query; and the 64 KiB `metadata` cap and the lowercase-`type` regex are **enforced but undocumented** in the shipped tool schemas.
|
||||||
|
|
||||||
|
### 3.2 Three orderings, and one identity that is not what it looks like
|
||||||
|
|
||||||
|
The log carries three distinct notions of order, and conflating them is the single most likely design error in any consumer:
|
||||||
|
|
||||||
|
| Field | Meaning | Scope | Where |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `seq` | **local arrival cursor** — the SQLite `rowid`, surfaced at read time | this replica only | `logstream.py:580-586` |
|
||||||
|
| `origin_seq` | the author replica's own gap-free counter, assigned by `UPDATE … SET origin_seq = rowid` in the insert transaction | per author | `logstream.py:700-706` |
|
||||||
|
| `hlc` | hybrid logical clock, `<unix_ms>-<counter_hex>-<replica_id>`, fixed-width so TEXT comparison *is* causal comparison | fleet-wide | `logstream.py:668`, `hlc.py:1-21` |
|
||||||
|
|
||||||
|
The rule, from `hlc.py:19-21`: *"Cursor semantics stay LOCAL (rowid arrival order) … a tail consumer must see late-arriving remote ops even though their HLC is older. HLC is the display/merge order; arrival is the delivery order."*
|
||||||
|
|
||||||
|
**`origin_replica` identifies the palace, not the writer.** It comes from `get_replica_id(db_parent)` — the `replica.json` in the palace directory (`logstream.py:422-436`, `replica.py:31-49`). In a hub-and-spoke deployment every client writes through the server's single `Logstream`, so **every event from every machine carries the same `origin_replica`** and it cannot distinguish two devices. This is mechanically forced by the design, not a misconfiguration.
|
||||||
|
|
||||||
|
Therefore: **`from_agent` / `to_agent` carry the entire distinction between machines.** The convention that makes them able to is `<harness>@<device>` — e.g. `pi@laptop`, `opencode@build-box` — stamped at the client edge, which RFC 001 §7.3.5 specifies. An unstamped client addressed as bare `pi` is unreachable in a fleet, because nobody can name it.
|
||||||
|
|
||||||
|
### 3.3 The ack contract: owed-ness is derived, never read
|
||||||
|
|
||||||
|
`status` is written **once**, into an append-only row. Nothing updates it. So `status="open"` means *"the sender declared this an ask at the moment of writing"* — it does **not** mean unanswered, and a directed `open` event keeps matching a mailbox query forever, answered or not.
|
||||||
|
|
||||||
|
`ack_event` (`logstream.py:812-845`) never touches the target row. It *appends*:
|
||||||
|
|
||||||
|
```python
|
||||||
|
return self.append_event(
|
||||||
|
type=ACK_EVENT_TYPE, # "event.ack"
|
||||||
|
stream=target["stream"], room=target["room"],
|
||||||
|
from_agent=from_agent,
|
||||||
|
to_agent=target["from_agent"], # routes back to the sender
|
||||||
|
correlation_id=target["correlation_id"] or target["id"], # falls back to the target id
|
||||||
|
status=status, body=body,
|
||||||
|
metadata={"ack_of": event_id}, # the join key lives in metadata
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
So a consumer that wants "what do I still owe?" must derive it. The derivation the toolkit implements, and which this RFC adopts as the reference semantics:
|
||||||
|
|
||||||
|
```mermaid
|
||||||
|
stateDiagram-v2
|
||||||
|
[*] --> open : task.request written, addressed to you
|
||||||
|
open --> claimed : you announce you are working (does NOT clear)
|
||||||
|
open --> ready : work exists, thread still open (does NOT clear)
|
||||||
|
claimed --> applied : terminal, clears
|
||||||
|
ready --> applied : terminal, clears
|
||||||
|
open --> blocked : terminal, clears (with a reason)
|
||||||
|
open --> failed : terminal, clears (with a reason)
|
||||||
|
open --> superseded : terminal, clears (replaced by another thread)
|
||||||
|
applied --> [*]
|
||||||
|
blocked --> [*]
|
||||||
|
failed --> [*]
|
||||||
|
superseded --> [*]
|
||||||
|
```
|
||||||
|
|
||||||
|
An event is **still owed** when all of these hold:
|
||||||
|
|
||||||
|
1. it is directed at you (`to_agent = <you>`, and `to_agent = '*'` is excluded — §7.4);
|
||||||
|
2. it was not written by you (you cannot owe yourself);
|
||||||
|
3. **no** event of yours has *all* of: a strictly higher `seq`, a join to it (`metadata.ack_of = <id>` or a shared `correlation_id`), **and** a terminal status (`applied` / `superseded` / `failed` / `blocked`).
|
||||||
|
|
||||||
|
The strictly-higher-`seq` test is load-bearing: without it, one terminal reply on a `correlation_id` suppresses every *later* ask on that same correlation, permanently. And because `seq` is local arrival order, this test is only sound on a single replica — §7.3.
|
||||||
|
|
||||||
|
### 3.4 Artifacts
|
||||||
|
|
||||||
|
Exact payloads for handoff: `patch`, `file`, `log`, `json`, `note`; UTF-8 text only; ≤4 MiB; `sha256` and `size_bytes` returned (`logstream.py:744-810`).
|
||||||
|
|
||||||
|
⚠️ **The sha256 is for verification, not deduplication.** `artifacts_sha256_idx` is a plain index and `put_artifact` always INSERTs, so storing the same patch twice stores the content twice. Referential integrity in the other direction *is* enforced: `artifact_ids` naming an unknown artifact raises at append time, in both the local and the remote path (`logstream.py:687-694`, `1181-1189`), *"so readers never see a dangling reference"*.
|
||||||
|
|
||||||
|
`kind="patch"` gets advisory warnings — never mutation — for a missing trailing newline or CRLF endings (`logstream.py:718-742`). The comment records why: *"the first RFC 003 dogfood: a patch stored without its final trailing newline truncates the last hunk line."*
|
||||||
|
|
||||||
|
### 3.5 Delivery: three mechanisms, one client feature
|
||||||
|
|
||||||
|
| Mechanism | Shape | Latency | Where |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `event_list` | pull, filtered, `ORDER BY rowid ASC`, default 50 / max 500 (silently clamped) | whenever you ask | `logstream.py:914-990` |
|
||||||
|
| `event_wait` | in-request long poll, jittered backoff 0.25 s→1 s, default 60 s, hard cap 300 s, returns `{timed_out: true, events: []}` rather than raising | seconds, while you hold the call | `logstream.py:1004-1039` |
|
||||||
|
| `GET /logstream/stream` | SSE; resume by `?since_event_id=` or `Last-Event-ID`; `: ping` every 15 s; ≤8 concurrent clients then `503` + `Retry-After` | sub-second, while connected | `mcp_server.py:7570-7671` |
|
||||||
|
|
||||||
|
None of these push anything into an agent that is not already running. That gap is closed **client-side** by the toolkit's *mailbox*: it derives the owed set per §3.3 and injects it at two points — once in the session-start wake-up, and mid-session on a settled-agent poll floored at 5 minutes, delivered as a queued steer rather than an interrupt.
|
||||||
|
|
||||||
|
**Measured 2026-08-26** (first live delivery on a published image, positive control with a synthetic sender): planted → delivered in **≈2–3 minutes** to a session that was already running, as exactly one item; three already-answered asks were correctly excluded and a `*` broadcast was correctly ignored. The wake-up path and the resurface path were **not** exercised by that test and remain unmeasured.
|
||||||
|
|
||||||
|
This is why the mailbox is a client concern and belongs in the toolkit rather than here: the server has no idea which agents exist, and no way to reach one that is not calling it.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Availability: the logstream is deliberately exempt from both palace locks
|
||||||
|
|
||||||
|
The rest of the palace is a single-writer service — `_HTTP_REQUEST_LOCK` wraps every dispatch (RFC 001 §7.5), and a Chroma peer-writer lease protects against a concurrent CLI mine. **Coordination traffic is exempt from both**, in two separate places, for two different reasons:
|
||||||
|
|
||||||
|
- `_HTTP_LOCK_FREE_TOOLS` (`mcp_server.py:6887-6902`) — all seven event/artifact tools dispatch outside the request lock: *"Dispatching them outside `_HTTP_REQUEST_LOCK` keeps a five-minute `mempalace_event_wait` long-poll from stalling every other agent on a shared hub, and lets the SSE stream coexist with normal tool traffic."*
|
||||||
|
- `_PEER_WRITER_EXEMPT_TOOLS` (`mcp_server.py:434-449`) — the four mutating ones bypass the writer lease: *"Exempting them keeps agent coordination alive while a CLI mine or a peer stdio writer holds the palace lock."*
|
||||||
|
|
||||||
|
They remain in `_MUTATING_TOOLS`, so `--read-only` still refuses them.
|
||||||
|
|
||||||
|
✅ **Consequence, and it corrects a widely-held belief in this fleet:** "one large mine blocks every client for minutes" is true of drawer writes and **false of coordination writes**. You can message another device, and it can reply, while a mine is running.
|
||||||
|
|
||||||
|
Concurrency within the log is a per-instance `threading.Lock` over a WAL database with a 10 s busy timeout, one cached `Logstream` per palace path per process (`logstream.py:437`, `mcp_server.py:333-334, 1120-1141`).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Not everything belongs in the log
|
||||||
|
|
||||||
|
RFC 001 §5 draws this line for wings; the same discipline applies between the two stores. The short form, expanded for operators in `docs/fleet-memory.md`:
|
||||||
|
|
||||||
|
| Put it in | When |
|
||||||
|
|---|---|
|
||||||
|
| A **drawer** (palace) | It will still be true, and worth finding, next month. Nobody in particular needs to act. Retrieval is by meaning. |
|
||||||
|
| An **event** (log) | A named agent must act, reply, or be stopped. Retrieval is by address and order. |
|
||||||
|
| **Both** | The durable finding goes in a drawer; the event says "there is a new finding, here is the drawer id". This is the recommended pattern for anything a peer must *know* rather than *do*. |
|
||||||
|
|
||||||
|
The failure mode in each direction: a finding filed only as an event is invisible to semantic search and will be re-derived by the next agent; an ask filed only as a drawer is addressed to nobody and will be found, if ever, by accident.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Security model
|
||||||
|
|
||||||
|
**The log authenticates the fleet, not the agent.** Precisely:
|
||||||
|
|
||||||
|
| Control | State | Evidence |
|
||||||
|
|---|---|---|
|
||||||
|
| Transport auth | One shared bearer token, `hmac.compare_digest`, 401 otherwise; required for non-loopback binds | `mcp_server.py:7137-7141, 7685` |
|
||||||
|
| `from_agent` authenticity | ❌ **Not checked against anything.** Shape-validated only; any client may append an event claiming to be any agent | `logstream.py:103-122` is the only check |
|
||||||
|
| Read authorization | ❌ **None.** Any token holder may list every event addressed to anyone, and fetch any artifact by id — including patch contents | no scoping in the query path |
|
||||||
|
| Stream auth | ✅ SSE follows the same bearer policy: *"Events and artifacts expose work metadata and patch contents"* | `mcp_server.py:7189-7194` |
|
||||||
|
| Transport security | TLS via `--tls-cert`/`--tls-key` (both or neither); Host pin against DNS rebinding; Origin check; 16 MiB request cap | `mcp_server.py:6920-6947` |
|
||||||
|
|
||||||
|
For a single-operator fleet behind one token this is adequate, and it must be stated rather than implied, because two useful consequences follow directly from it:
|
||||||
|
|
||||||
|
1. **Impersonation is trivial** — do not treat `from_agent` as evidence of origin in any security decision.
|
||||||
|
2. **That same property is the only way to test the mailbox.** A positive control requires writing an event *from* an agent you are not, so that your own client does not self-filter it. This is a supported technique precisely because the field is unauthenticated (**measured 2026-08-26**).
|
||||||
|
|
||||||
|
If per-agent identity is ever needed, the correct home is the same place RFC 001 §7.3 puts provenance: the authenticated credential at the boundary, not a self-asserted field.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Landmines
|
||||||
|
|
||||||
|
### 7.1 There is no idempotency guard on append — a retried write duplicates
|
||||||
|
|
||||||
|
`append_event` mints a fresh id and INSERTs, with no dedup lookup of any kind (`logstream.py:625-716`). `put_artifact` likewise. Two byte-identical calls produce two events with different ids and different `seq`.
|
||||||
|
|
||||||
|
The asymmetry is stark: the *replication* path is rigorously idempotent (`logstream.py:1174-1180` checks `id` **or** `(origin_replica, origin_seq)` before applying, and returns early for its own echoed ops). Peer replay is safe; **client retry is not.**
|
||||||
|
|
||||||
|
**Action:** this makes the palace's general rule — *a timeout usually means the write completed; verify, don't retry* — load-bearing rather than advisory. For a drawer, a blind retry costs a dedup-detectable duplicate. For an event it silently forks a coordination thread into two ids, and a mailbox will then show two owed items that must each be closed. After any `event_append` / `patch_submit` timeout, verify with `event_list(correlation_id=…)` before re-issuing.
|
||||||
|
|
||||||
|
### 7.2 `status="open"` is not owed-ness
|
||||||
|
|
||||||
|
Covered in §3.3 and repeated here because it is the mistake most likely to be made by someone reading only the tool descriptions: the mailbox query `to_agent=<me> status=open` **never shrinks as you work**. Treating its length as a to-do count means re-answering answered asks forever.
|
||||||
|
|
||||||
|
**Action:** derive per §3.3, or use a client that does.
|
||||||
|
|
||||||
|
### 7.3 `seq` is local arrival order and is meaningless across replicas
|
||||||
|
|
||||||
|
`seq` is the local `rowid`. The moment a second replica exists, a remote event that was *authored* earlier can arrive *later* and receive a higher `seq`. The owed-set derivation in §3.3 compares `seq`, so it is sound only on a single replica.
|
||||||
|
|
||||||
|
**Action:** on a hub-and-spoke fleet (all clients on one replica) this is correct today. `hlc` is the field to switch to when `mesh_peers` reports actual peers — and that switch must happen *before* multi-replica, not after.
|
||||||
|
|
||||||
|
### 7.4 A broadcast reaches no mailbox, and `to_agent=NULL` reaches nobody at all
|
||||||
|
|
||||||
|
`to_agent=<x>` matches `x` **or** `'*'` at the SQL level (`logstream.py:958-960`), so a broadcast *is* visible to a listing agent. But the reference owed-set derivation excludes `to_agent='*'` deliberately — a broadcast owes nobody a reply, and if it entered every mailbox, every machine would think it personally owed the same answer.
|
||||||
|
|
||||||
|
Two consequences: **broadcasting an ask reaches no owed set at all** (the "don't broadcast an ask" anti-pattern is mechanically enforced, not merely advised), and to reach a whole fleet with something actionable you must write **one directed event per device**, sharing a `correlation_id` so the thread stays joinable. Separately, an event written with **no** `to_agent` matches no `to_agent=` query ever and is addressed to nobody.
|
||||||
|
|
||||||
|
### 7.5 `since_event_id` raises on an unknown id
|
||||||
|
|
||||||
|
An unresolvable cursor raises `ValueError` rather than returning empty (`logstream.py:966-970`) — stronger than the tool description promises, and benign on one replica. On a second replica, a resuming watcher whose cursor has not yet replicated will **hard-fail instead of waiting**.
|
||||||
|
|
||||||
|
### 7.6 A rejected append returns HTTP 200
|
||||||
|
|
||||||
|
Failures come back as `{"success": false, "error": …}` in a 200 response, not a JSON-RPC error (`mcp_server.py:4581-4584`).
|
||||||
|
|
||||||
|
**Action:** a client that checks only transport status silently loses the event. Check the payload.
|
||||||
|
|
||||||
|
### 7.7 The same agent name on two devices is indistinguishable
|
||||||
|
|
||||||
|
No uniqueness, no registry, no warning. Both machines' events carry one `from_agent`, and each machine's "you cannot owe yourself" filter will discard the other's asks. This is the failure the `<harness>@<device>` convention exists to prevent (§3.2).
|
||||||
|
|
||||||
|
### 7.8 `GET /logstream/events` does not exist
|
||||||
|
|
||||||
|
The complete GET route table is `/healthz`, `/statusz`, `/logstream/stream`, `/sync/{version_vector,ops,artifact,peers}`; POST serves `/mcp` only. A recursive grep for `/logstream` in the package yields three hits, all `/logstream/stream`.
|
||||||
|
|
||||||
|
⚠️ **Correction.** An earlier measurement observed `/logstream/events`, `/logstream/stream` and `/sync/peers` all returning 404 against a deployment and attributed all three to a reverse proxy exposing only `/mcp`. That inference was right for `/sync/*` and `/logstream/stream` and **wrong for `/logstream/events`**, which would 404 on a directly-reachable server too. Distinguish "route absent" from "route blocked" before blaming infrastructure.
|
||||||
|
|
||||||
|
### 7.9 Dogfood scars, preserved because each cost someone a session
|
||||||
|
|
||||||
|
- A `patch` artifact stored without its trailing newline truncates the last hunk line — hence the advisory warnings (`logstream.py:718-742`).
|
||||||
|
- `--type Task.Request` was rejected while `--type Task.Request --type patch.ready` was *silently accepted and matched nothing*, leaving a watcher waiting forever: single-valued filters were validated by pushdown, multi-valued ones compared raw (`sanitize_watch_spec`, `logstream.py:245-278`). `type` is lowercase-only for this reason.
|
||||||
|
- `event_wait` rejected a `limit` that `event_list` accepted — *"reported by windows-codex during dogfood"* (`mcp_server.py:4666-4670`).
|
||||||
|
|
||||||
|
### 7.10 The main database file's mtime is not a liveness signal
|
||||||
|
|
||||||
|
WAL. **Measured 2026-08-26** on the primary: `logstream.sqlite3` mtime three days old while `logstream.sqlite3-wal` was 865 KiB and seconds old. An operator checking whether coordination is live must look at the `-wal` file, or query.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Phasing
|
||||||
|
|
||||||
|
### 8.1 Implemented
|
||||||
|
|
||||||
|
| Phase | Deliverable | State |
|
||||||
|
|---|---|---|
|
||||||
|
| A | `events`/`artifacts`/`event_artifacts` + append/list/ack + size limits | ✅ 3.7.x |
|
||||||
|
| B | Long poll (`event_wait`) and SSE push | ✅ |
|
||||||
|
| C | `patch_submit` convenience + patch advisories | ✅ |
|
||||||
|
| D | HLC on every event; `(origin_replica, origin_seq)` unique index | ✅ 3.8.0 |
|
||||||
|
| E | Client-side auto-delivered mailbox | ✅ toolkit `5b8d78f`; first live delivery measured 2026-08-26 |
|
||||||
|
|
||||||
|
### 8.2 Deferred, and what RFC 004 owes
|
||||||
|
|
||||||
|
Replication exists in code — `version_vector()`, `list_ops()`, `apply_remote_event()`, `apply_remote_artifact()` and the four `GET /sync/*` routes (`logstream.py:1096-1261`, `mcp_server.py:7206-7256`) — and is attributed in comments to "RFC 004 step 0". **That document does not exist.** Until it does, this RFC records the two properties a reader most needs: the apply path *is* idempotent (§7.1), and switching the owed-set derivation from `seq` to `hlc` is a precondition for a second replica, not a follow-up (§7.3).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Open decisions
|
||||||
|
|
||||||
|
1. **Retention.** No TTL, compaction or pruning exists, and `mempalace sync` does not touch the log (§2). For a fleet log this is mostly a feature — nothing is lost by being offline for weeks — but every 4 MiB artifact is permanent. Decide a policy before the log outgrows a comfortable backup, or decide explicitly that permanence is the policy.
|
||||||
|
2. **Idempotency key.** Should `event_append` accept an optional client-supplied dedupe key so a retried call is a no-op? This is the one change that would make §7.1 disappear.
|
||||||
|
3. **`hlc` as the mailbox join key**, replacing `seq`. Required before a second replica (§7.3). Cheap now, breaking later.
|
||||||
|
4. **Per-agent identity.** Do we ever want `from_agent` to be authenticated (§6), or is "authenticates the fleet, not the agent" the permanent contract?
|
||||||
|
5. **Should `GET /logstream/events` exist?** A read-only HTTP tail would let non-MCP consumers (dashboards, CI) follow a stream without an MCP client. Today they must hold SSE or speak MCP.
|
||||||
|
6. **Undocumented limits.** The 64 KiB `metadata` cap and the lowercase-`type` regex are enforced server-side but absent from the tool schemas (§3.1). Document, or relax.
|
||||||
|
7. **Agent-name registry.** Nothing prevents two devices sharing one `from_agent` (§7.7). A warning at append time would be cheap.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. Evidence index
|
||||||
|
|
||||||
|
| Claim | Where |
|
||||||
|
|---|---|
|
||||||
|
| Store is `logstream.sqlite3` inside the palace dir | `logstream.py:51`; `mcp_server.py:1108-1117` |
|
||||||
|
| Full DDL, indexes, WAL | `logstream.py:447-497`, `552-556` |
|
||||||
|
| Design constraints (quoted in §3) | `logstream.py:1-18` |
|
||||||
|
| Event id format, ordering not carried by id | `logstream.py:92-102` |
|
||||||
|
| `seq` = rowid, surfaced at read | `logstream.py:580-586` |
|
||||||
|
| `origin_seq` assigned in-transaction | `logstream.py:700-706` |
|
||||||
|
| `hlc` populated per append; format and sort property | `logstream.py:668`; `hlc.py:1-21` |
|
||||||
|
| Cursor stays local, HLC is merge order | `hlc.py:19-21` |
|
||||||
|
| `origin_replica` from palace dir | `logstream.py:422-436`; `replica.py:31-49` |
|
||||||
|
| Server-side `created_at`, second precision | `logstream.py:667` |
|
||||||
|
| No dedup on append; remote apply *is* idempotent | `logstream.py:625-716` vs `1174-1180` |
|
||||||
|
| Artifact sha256 non-unique index | `logstream.py:447-497`, `744-810` |
|
||||||
|
| `artifact_ids` referential check | `logstream.py:687-694`, `1181-1189` |
|
||||||
|
| Status set; artifact kinds; size limits | `logstream.py:67-69`, `55-57`, `765-778` |
|
||||||
|
| `type` regex, lowercase ≤64 | `logstream.py:76, 124-134` |
|
||||||
|
| Broadcast matching in SQL | `logstream.py:958-960` |
|
||||||
|
| `since_event_id` anchor + raise | `logstream.py:963-978` |
|
||||||
|
| Order by rowid; limit clamp 500 | `logstream.py:947` |
|
||||||
|
| `preview` truncates body to 200 chars | `mcp_server.py:4587-4607` |
|
||||||
|
| `event_wait` backoff, cap, timeout result | `logstream.py:1004-1039` |
|
||||||
|
| `ack_event` appends, never mutates; `ack_of` in metadata | `logstream.py:812-845` |
|
||||||
|
| Patch advisories | `logstream.py:718-742` |
|
||||||
|
| Lock exemptions, both, with rationale | `mcp_server.py:6887-6902`, `434-449` |
|
||||||
|
| Bearer auth; SSE auth required | `mcp_server.py:7137-7141`, `7189-7194` |
|
||||||
|
| SSE implementation, heartbeat, client cap | `mcp_server.py:7570-7671` |
|
||||||
|
| Rejected append returns 200 + `success:false` | `mcp_server.py:4581-4584` |
|
||||||
|
| `sync` does not touch the log | `cli.py:1057`; `sync.py` (no logstream refs) |
|
||||||
|
| No retention/pruning anywhere | no `DELETE FROM events`/`artifacts` in the package |
|
||||||
|
| Replication surface, attributed to RFC 004 | `logstream.py:1096-1261`; `mcp_server.py:7206-7256` |
|
||||||
|
|
||||||
|
### 10.1 Where the shipped tool descriptions and the code disagree
|
||||||
|
|
||||||
|
| Tool-description claim | Verdict |
|
||||||
|
|---|---|
|
||||||
|
| "append-only; corrections are new events" | True of content. Precisely: no event's *content* is ever mutated; `append_event` does issue one `UPDATE … SET origin_seq = rowid` on the row it just inserted, and the 3.8.0 migration backfills `origin_replica`/`origin_seq`/`hlc` on pre-existing rows (`logstream.py:504-550`). |
|
||||||
|
| "`since_event_id` cannot skip anything" | True, and stronger than stated — it raises on an unknown id (§7.5). |
|
||||||
|
| "body max 256 KiB", "artifact max 4 MiB, UTF-8 only", "timeout default 60 s max 5 min" | All true; the timeout clamps silently rather than erroring. |
|
||||||
|
| "`to_agent=<you>` also matches `*` broadcasts" | True, SQL-level — but see §7.4 for why a broadcast still reaches no mailbox. |
|
||||||
|
| "prefer the push stream at `GET /logstream/stream`" | True; route exists and requires auth. |
|
||||||
|
| `GET /logstream/events` | ❌ Does not exist (§7.8). |
|
||||||
|
| `metadata` size limit; `type` charset | ❌ Enforced but undocumented (§3.1). |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. See also
|
||||||
|
|
||||||
|
- `docs/fleet-memory.md` — the operator-facing companion: what the palace's stores are *for*, and when to use a drawer versus an event.
|
||||||
|
- `docs/rfc-001-global-palace.md` — §5 what should not be global, §7.2 the `sync` hazard, §7.3 provenance at the boundary, §7.5 single-writer expectations.
|
||||||
|
- `docs/rfc-002-joiner.md` — replaying a second palace into a shared primary.
|
||||||
|
- `extensions/pi/README.md` §2 — the client-side mailbox: delivery points, gating, and why it is a client concern.
|
||||||
|
- `~/.agents/skills/mempalace/SKILL.md` §"Cross-Machine Coordination" — **normative for agent behaviour**; this RFC is normative for mechanism.
|
||||||
+11
-2
@@ -342,12 +342,21 @@ the mechanism; the skill is normative for behaviour.**
|
|||||||
an SSE endpoint (`GET /logstream/stream`, `text/event-stream` in
|
an SSE endpoint (`GET /logstream/stream`, `text/event-stream` in
|
||||||
`mempalace/mcp_server.py`), but a deployment may expose only the MCP endpoint
|
`mempalace/mcp_server.py`), but a deployment may expose only the MCP endpoint
|
||||||
through its reverse proxy — verified 2026-08-26 against
|
through its reverse proxy — verified 2026-08-26 against
|
||||||
`https://mempalace.jordbo.se`, where `/logstream/events`, `/logstream/stream`
|
`https://mempalace.jordbo.se`, where `/logstream/stream` and `/sync/peers`
|
||||||
and `/sync/peers` all return 404 while `/mcp` serves normally. Where that is the
|
return 404 while `/mcp` serves normally. Where that is the
|
||||||
case, polling through the existing MCP client is the only available path — which
|
case, polling through the existing MCP client is the only available path — which
|
||||||
is what the mailbox in §2 does — and enabling SSE means a proxy route plus an
|
is what the mailbox in §2 does — and enabling SSE means a proxy route plus an
|
||||||
auth decision, not an extension change.
|
auth decision, not an extension change.
|
||||||
|
|
||||||
|
> ⚠️ **Corrected 2026-08-26.** An earlier revision of this paragraph listed
|
||||||
|
> `/logstream/events` alongside those two as proxy-blocked. That route **does not
|
||||||
|
> exist in the server at all** — the complete GET table in mempalace 3.8.0 is
|
||||||
|
> `/healthz`, `/statusz`, `/logstream/stream` and `/sync/{version_vector,ops,artifact,peers}`,
|
||||||
|
> so `/logstream/events` would 404 against a directly-reachable server too. The
|
||||||
|
> proxy inference was right for the other two and wrong for that one; see
|
||||||
|
> [RFC 003 §7.8](../../docs/rfc-003-coordination-log.md). Distinguish *route absent*
|
||||||
|
> from *route blocked* before blaming infrastructure.
|
||||||
|
|
||||||
As with stamping, all of this is inert unless `MEMPALACE_REMOTE_URL` points at a
|
As with stamping, all of this is inert unless `MEMPALACE_REMOTE_URL` points at a
|
||||||
shared palace. On a solitary palace the event tools work fine and the log
|
shared palace. On a solitary palace the event tools work fine and the log
|
||||||
contains only this machine's own events.
|
contains only this machine's own events.
|
||||||
|
|||||||
Reference in New Issue
Block a user