d4d8bb6109
RFC 003 §9.1 — retention is settled in direction (rotate old traffic out of
the way, logrotate-style, moved aside rather than destroyed) and the sketch
records the parts that are not obvious:
- Tier artifacts before events. Events are a few KB; artifacts are capped at
4 MiB and stored in-row, so moving artifact CONTENT cold while keeping the
kind/sha256/size/created_by stub reclaims nearly all the space and keeps the
audit trail ("what was handed over, by whom, verified how") intact.
- If events rotate at all, the unit is the correlation_id THREAD whose latest
event is terminal — never the row. Archiving an ask while leaving its reply
(or the reverse) breaks the owed-set join, and both failure modes are bad:
an ask that can never be cleared resurfaces as owed forever, or a reply is
orphaned from what it answered.
- Never rotate an event that is still owed. Owed-ness is DERIVED at read time,
so an unanswered ask is indistinguishable from a stale one except by that
derivation — a purely time-based sweep would discard the live obligations of
a machine that has merely been offline for a month, which is precisely the
case this log exists to serve.
- Rotation invalidates held cursors: since_event_id RAISES on an unknown id
(§7.5), so archiving an event a watcher holds as its resume point turns its
next poll into an error. Either announce rotation ahead of live cursors, or
teach the anchor lookup to fall back to created_at/hlc.
- Archive → verify (row counts, artifact sha256) → only then DELETE + VACUUM,
with a --dry-run that reports in thread units.
RFC 002 §5: "Open decisions for ALC" → "Open decisions". RFC 001 never names
the operator anywhere; impersonal is the mature precedent and ALC was never a
real identifier in the first place (it is the AAAK spec's illustrative code for
"Alice", copy-forwarded into ~700 diary entries without verification).
Same edit also removes two device names and a hostname from §5.3, which is
host inventory and belongs in the private fleet repo, not a public one. The
mechanism it teaches — a palace in a Docker named volume dies on the next
container recreate, so census before flipping — is unchanged and is the part
that mattered.
414 lines
33 KiB
Markdown
414 lines
33 KiB
Markdown
# RFC 003 — The coordination log (`logstream`)
|
||
|
||
**Status:** implemented and in production use since 3.7.x. This document is a *retrospective* specification, written after the fact.
|
||
**Author:** pi (agent), 2026-08-26.
|
||
**Applies to:** mempalace 3.8.0 (`logstream.py`, 1261 lines), mempalace-toolkit `5b8d78f` (the pi-side mailbox).
|
||
**Context:** every `mempalace_event_*` and `mempalace_artifact_*` tool description cites "RFC 003". `logstream.py`'s module docstring is headed *"Agent coordination event log for MemPalace (RFC 003)"* and enumerates five "Design constraints (RFC 003)". Inline comments cite "RFC 003 phase 5", "RFC 003 suggested defaults" and "the first RFC 003 dogfood". **No such document has ever existed** — verified 2026-08-26 by searching both the toolkit repository and the primary host. This RFC transcribes the spec the implementation already believes in, and — more usefully — records what it does *not* do.
|
||
|
||
Read §7 before you build anything on this log; it is the part that is not obvious from the tool descriptions. Every mechanical claim below cites `file.py:LINE` in mempalace 3.8.0, and §10 indexes them so any claim can be re-verified without re-reading 1261 lines. Claims marked **measured** were executed against the live fleet on the date given; claims without that marker are reads of the source. Where the code and the shipped tool descriptions disagree, §10.1 says so explicitly rather than quietly siding with one.
|
||
|
||
One scoping note up front. This RFC covers the **log**: its storage, its append and query semantics, delivery, and the trust model. It does not specify **replication**, which the source already attributes to a different document (`# ── Replication (RFC 004 step 0: logstream multi-master) ──`, `logstream.py:1096`). RFC 004 does not exist either; §8.2 states what it owes.
|
||
|
||
---
|
||
|
||
## 1. The problem
|
||
|
||
The palace stores what an agent *knows*: drawers are semantic, retrieved by meaning, and deliberately have no addressee. That shape is wrong for four things a fleet of agents actually needs:
|
||
|
||
1. **Addressing.** "This is for the machine that owns the release" cannot be expressed as a drawer. A drawer is found by whoever happens to search for the right words.
|
||
2. **A reply that closes something.** Semantic memory has no notion of an outstanding question. Nothing in a drawer can be *owed*.
|
||
3. **Exact payloads.** A unified diff must survive byte-for-byte. Drawers are chunked and embedded; that is a feature for prose and a defect for patches.
|
||
4. **Order.** "What happened after this?" needs an append cursor, not a similarity score.
|
||
|
||
The coordination log adds exactly those four properties and nothing else. It is a second store beside the palace, not a new kind of drawer.
|
||
|
||
### 1.1 What it is not
|
||
|
||
Stated first, because every misuse of this log so far has come from assuming one of these:
|
||
|
||
- **Not a bus.** Nothing subscribes by default; nothing is delivered "live" unless a client is holding an SSE connection or polling. There is no delivery window and nothing is lost by being offline when an event is written.
|
||
- **Not a chat channel.** Latency is bounded by *when the recipient next runs*, which in a fleet of workstations is hours, days or weeks. §3.5 gives the measured numbers for the case where the recipient is awake.
|
||
- **Not authenticated per agent.** `from_agent` is a routing label with the trust properties of an e-mail `From:` header (§6).
|
||
- **Not a queue.** Nothing is consumed, acknowledged-and-removed, or retried. Events are permanent (§9.1) and "handled" is a *derived* property (§3.3).
|
||
|
||
---
|
||
|
||
## 2. Verified starting point (2026-08-26, mempalace 3.8.0)
|
||
|
||
| Piece | State | Evidence |
|
||
|---|---|---|
|
||
| Separate SQLite store, inside the palace dir | ✅ `logstream.sqlite3` beside `chroma.sqlite3`; WAL; dir `chmod 0700` best-effort | `logstream.py:51`, `mcp_server.py:1108-1117`, `logstream.py:430-433` |
|
||
| `events`, `artifacts`, `event_artifacts` tables | ✅ Created idempotently at open | `logstream.py:447-497` |
|
||
| Hybrid logical clock on every event | ✅ `hlc` populated on every local append | `logstream.py:668`, `hlc.py:1-21` |
|
||
| Seven MCP tools (append/list/wait/ack, artifact put/get, patch_submit) | ✅ | `mcp_server.py:4546-4771` |
|
||
| Server-sent events push | ✅ `GET /logstream/stream`, auth required, 15 s heartbeat, ≤8 clients | `mcp_server.py:7570-7671`, `7189-7194` |
|
||
| `GET /logstream/events` | ❌ **Never implemented.** Only `/logstream/stream` exists | see §7.8 |
|
||
| Auto-delivered mailbox (owed-set derivation + injection) | ✅ **Client-side, not server-side** — lives in the toolkit's pi extension | `extensions/pi/mempalace.ts` (toolkit `5b8d78f`) |
|
||
| Idempotency guard on append | ❌ **Absent.** See §7.1 | `logstream.py:625-716` |
|
||
| Retention / TTL / compaction | ❌ Absent by design-so-far. See §9.1 | no `DELETE FROM events` in the package |
|
||
| Interaction with `mempalace sync` | ✅ **None.** `sync` prunes drawers only | `cli.py:1057`, `sync.py` (no logstream references) |
|
||
|
||
Two things that sound like this feature and are not:
|
||
|
||
- **`mempalace_mesh_peers` is not a fleet roster.** It reports peer *replicas*. A hub-and-spoke deployment — many thin MCP clients of one server — correctly reports `peers: []` while every machine in the fleet is actively writing to the same log. Deciding "coordination does not apply to me" from an empty peer list is a measured failure mode, not a hypothetical one.
|
||
- **`origin_replica` is not the writing machine.** It is a property of the palace *directory* (§3.2).
|
||
|
||
---
|
||
|
||
## 3. Design
|
||
|
||
The five constraints in the module docstring (`logstream.py:1-18`) are the design, quoted verbatim because they are already normative in the implementation:
|
||
|
||
> - No Chroma dependency, no vector index open — plain SQLite only.
|
||
> - Append-only: events are immutable; corrections are new events that reference prior events.
|
||
> - Exact payloads: event bodies and artifact content are stored verbatim.
|
||
> - Safe under concurrent HTTP requests (WAL + per-instance lock, same pattern as `knowledge_graph.py`).
|
||
> - Explicit size limits with clear errors, never silent truncation.
|
||
|
||
The shape, in the same idiom as RFC 001 §4.1:
|
||
|
||
```
|
||
agent on device A ─┐ palace directory
|
||
agent on device B ─┼── MCP /mcp ──► server ─┬─ chroma.sqlite3 (drawers: what you know)
|
||
agent on device C ─┘ │ ├─ knowledge_graph.sqlite3 (facts, temporal)
|
||
│ └─ logstream.sqlite3 (events + artifacts)
|
||
GET /logstream/stream events ── append-only, permanent
|
||
(SSE, auth, ≤8 clients) artifacts ── verbatim, ≤4 MiB, sha256
|
||
event_artifacts ── join, checked at append
|
||
```
|
||
|
||
### 3.1 Data model
|
||
|
||
```sql
|
||
CREATE TABLE IF NOT EXISTS events (
|
||
id TEXT PRIMARY KEY,
|
||
type TEXT NOT NULL,
|
||
stream TEXT NOT NULL,
|
||
room TEXT NOT NULL,
|
||
from_agent TEXT NOT NULL,
|
||
to_agent TEXT,
|
||
correlation_id TEXT,
|
||
branch TEXT,
|
||
base_commit TEXT,
|
||
status TEXT,
|
||
body TEXT NOT NULL DEFAULT '',
|
||
created_at TEXT NOT NULL,
|
||
metadata_json TEXT NOT NULL DEFAULT '{}',
|
||
origin_replica TEXT,
|
||
origin_seq INTEGER,
|
||
hlc TEXT
|
||
);
|
||
CREATE UNIQUE INDEX IF NOT EXISTS events_origin_seq_idx ON events(origin_replica, origin_seq);
|
||
```
|
||
(`logstream.py:447-497`, unique index at `552-556`. `artifacts` and `event_artifacts` in the same script; `artifacts_sha256_idx` is **not** unique — see §3.4.)
|
||
|
||
Validation, all server-side, all raising `ValueError` naming the allowed set:
|
||
|
||
| Field | Required | Rule | Where |
|
||
|---|---|---|---|
|
||
| `type` | yes | `^[a-z0-9][a-z0-9_.-]{0,63}$` — **lowercase only, ≤64** | `logstream.py:76, 124-134` |
|
||
| `stream`, `room`, `from_agent` | yes | non-empty, ≤256, no control chars | `logstream.py:78, 103-122` |
|
||
| `status` | no | one of `open claimed ready applied blocked failed superseded` | `logstream.py:67-69, 136-143` |
|
||
| `body` | no (empty allowed) | ≤256 KiB, no NUL | `logstream.py:55, 145-158` |
|
||
| `metadata` | no | canonical JSON, **≤64 KiB** | `logstream.py:57, 160-177` |
|
||
| `artifact_ids` | no | each must already exist, else `ValueError` | `logstream.py:687-694` |
|
||
|
||
Three consequences worth stating because they are invisible from the tool descriptions: an event body **may be empty** while artifact content **may not** (`logstream.py:766-767`); `to_agent` is **nullable**, and an event with `to_agent = NULL` is addressed to nobody and can never match a `to_agent=` query; and the 64 KiB `metadata` cap and the lowercase-`type` regex are **enforced but undocumented** in the shipped tool schemas.
|
||
|
||
### 3.2 Three orderings, and one identity that is not what it looks like
|
||
|
||
The log carries three distinct notions of order, and conflating them is the single most likely design error in any consumer:
|
||
|
||
| Field | Meaning | Scope | Where |
|
||
|---|---|---|---|
|
||
| `seq` | **local arrival cursor** — the SQLite `rowid`, surfaced at read time | this replica only | `logstream.py:580-586` |
|
||
| `origin_seq` | the author replica's own gap-free counter, assigned by `UPDATE … SET origin_seq = rowid` in the insert transaction | per author | `logstream.py:700-706` |
|
||
| `hlc` | hybrid logical clock, `<unix_ms>-<counter_hex>-<replica_id>`, fixed-width so TEXT comparison *is* causal comparison | fleet-wide | `logstream.py:668`, `hlc.py:1-21` |
|
||
|
||
The rule, from `hlc.py:19-21`: *"Cursor semantics stay LOCAL (rowid arrival order) … a tail consumer must see late-arriving remote ops even though their HLC is older. HLC is the display/merge order; arrival is the delivery order."*
|
||
|
||
**`origin_replica` identifies the palace, not the writer.** It comes from `get_replica_id(db_parent)` — the `replica.json` in the palace directory (`logstream.py:422-436`, `replica.py:31-49`). In a hub-and-spoke deployment every client writes through the server's single `Logstream`, so **every event from every machine carries the same `origin_replica`** and it cannot distinguish two devices. This is mechanically forced by the design, not a misconfiguration.
|
||
|
||
Therefore: **`from_agent` / `to_agent` carry the entire distinction between machines.** The convention that makes them able to is `<harness>@<device>` — e.g. `pi@laptop`, `opencode@build-box` — stamped at the client edge, which RFC 001 §7.3.5 specifies. An unstamped client addressed as bare `pi` is unreachable in a fleet, because nobody can name it.
|
||
|
||
### 3.3 The ack contract: owed-ness is derived, never read
|
||
|
||
`status` is written **once**, into an append-only row. Nothing updates it. So `status="open"` means *"the sender declared this an ask at the moment of writing"* — it does **not** mean unanswered, and a directed `open` event keeps matching a mailbox query forever, answered or not.
|
||
|
||
`ack_event` (`logstream.py:812-845`) never touches the target row. It *appends*:
|
||
|
||
```python
|
||
return self.append_event(
|
||
type=ACK_EVENT_TYPE, # "event.ack"
|
||
stream=target["stream"], room=target["room"],
|
||
from_agent=from_agent,
|
||
to_agent=target["from_agent"], # routes back to the sender
|
||
correlation_id=target["correlation_id"] or target["id"], # falls back to the target id
|
||
status=status, body=body,
|
||
metadata={"ack_of": event_id}, # the join key lives in metadata
|
||
)
|
||
```
|
||
|
||
So a consumer that wants "what do I still owe?" must derive it. The derivation the toolkit implements, and which this RFC adopts as the reference semantics:
|
||
|
||
```mermaid
|
||
stateDiagram-v2
|
||
[*] --> open : task.request written, addressed to you
|
||
open --> claimed : you announce you are working (does NOT clear)
|
||
open --> ready : work exists, thread still open (does NOT clear)
|
||
claimed --> applied : terminal, clears
|
||
ready --> applied : terminal, clears
|
||
open --> blocked : terminal, clears (with a reason)
|
||
open --> failed : terminal, clears (with a reason)
|
||
open --> superseded : terminal, clears (replaced by another thread)
|
||
applied --> [*]
|
||
blocked --> [*]
|
||
failed --> [*]
|
||
superseded --> [*]
|
||
```
|
||
|
||
An event is **still owed** when all of these hold:
|
||
|
||
1. it is directed at you (`to_agent = <you>`, and `to_agent = '*'` is excluded — §7.4);
|
||
2. it was not written by you (you cannot owe yourself);
|
||
3. **no** event of yours has *all* of: a strictly higher `seq`, a join to it (`metadata.ack_of = <id>` or a shared `correlation_id`), **and** a terminal status (`applied` / `superseded` / `failed` / `blocked`).
|
||
|
||
The strictly-higher-`seq` test is load-bearing: without it, one terminal reply on a `correlation_id` suppresses every *later* ask on that same correlation, permanently. And because `seq` is local arrival order, this test is only sound on a single replica — §7.3.
|
||
|
||
### 3.4 Artifacts
|
||
|
||
Exact payloads for handoff: `patch`, `file`, `log`, `json`, `note`; UTF-8 text only; ≤4 MiB; `sha256` and `size_bytes` returned (`logstream.py:744-810`).
|
||
|
||
⚠️ **The sha256 is for verification, not deduplication.** `artifacts_sha256_idx` is a plain index and `put_artifact` always INSERTs, so storing the same patch twice stores the content twice. Referential integrity in the other direction *is* enforced: `artifact_ids` naming an unknown artifact raises at append time, in both the local and the remote path (`logstream.py:687-694`, `1181-1189`), *"so readers never see a dangling reference"*.
|
||
|
||
`kind="patch"` gets advisory warnings — never mutation — for a missing trailing newline or CRLF endings (`logstream.py:718-742`). The comment records why: *"the first RFC 003 dogfood: a patch stored without its final trailing newline truncates the last hunk line."*
|
||
|
||
### 3.5 Delivery: three mechanisms, one client feature
|
||
|
||
| Mechanism | Shape | Latency | Where |
|
||
|---|---|---|---|
|
||
| `event_list` | pull, filtered, `ORDER BY rowid ASC`, default 50 / max 500 (silently clamped) | whenever you ask | `logstream.py:914-990` |
|
||
| `event_wait` | in-request long poll, jittered backoff 0.25 s→1 s, default 60 s, hard cap 300 s, returns `{timed_out: true, events: []}` rather than raising | seconds, while you hold the call | `logstream.py:1004-1039` |
|
||
| `GET /logstream/stream` | SSE; resume by `?since_event_id=` or `Last-Event-ID`; `: ping` every 15 s; ≤8 concurrent clients then `503` + `Retry-After` | sub-second, while connected | `mcp_server.py:7570-7671` |
|
||
|
||
None of these push anything into an agent that is not already running. That gap is closed **client-side** by the toolkit's *mailbox*: it derives the owed set per §3.3 and injects it at two points — once in the session-start wake-up, and mid-session on a settled-agent poll floored at 5 minutes, delivered as a queued steer rather than an interrupt.
|
||
|
||
**Measured 2026-08-26** (first live delivery on a published image, positive control with a synthetic sender): planted → delivered in **≈2–3 minutes** to a session that was already running, as exactly one item; three already-answered asks were correctly excluded and a `*` broadcast was correctly ignored. The wake-up path and the resurface path were **not** exercised by that test and remain unmeasured.
|
||
|
||
This is why the mailbox is a client concern and belongs in the toolkit rather than here: the server has no idea which agents exist, and no way to reach one that is not calling it.
|
||
|
||
---
|
||
|
||
## 4. Availability: the logstream is deliberately exempt from both palace locks
|
||
|
||
The rest of the palace is a single-writer service — `_HTTP_REQUEST_LOCK` wraps every dispatch (RFC 001 §7.5), and a Chroma peer-writer lease protects against a concurrent CLI mine. **Coordination traffic is exempt from both**, in two separate places, for two different reasons:
|
||
|
||
- `_HTTP_LOCK_FREE_TOOLS` (`mcp_server.py:6887-6902`) — all seven event/artifact tools dispatch outside the request lock: *"Dispatching them outside `_HTTP_REQUEST_LOCK` keeps a five-minute `mempalace_event_wait` long-poll from stalling every other agent on a shared hub, and lets the SSE stream coexist with normal tool traffic."*
|
||
- `_PEER_WRITER_EXEMPT_TOOLS` (`mcp_server.py:434-449`) — the four mutating ones bypass the writer lease: *"Exempting them keeps agent coordination alive while a CLI mine or a peer stdio writer holds the palace lock."*
|
||
|
||
They remain in `_MUTATING_TOOLS`, so `--read-only` still refuses them.
|
||
|
||
✅ **Consequence, and it corrects a widely-held belief in this fleet:** "one large mine blocks every client for minutes" is true of drawer writes and **false of coordination writes**. You can message another device, and it can reply, while a mine is running.
|
||
|
||
Concurrency within the log is a per-instance `threading.Lock` over a WAL database with a 10 s busy timeout, one cached `Logstream` per palace path per process (`logstream.py:437`, `mcp_server.py:333-334, 1120-1141`).
|
||
|
||
---
|
||
|
||
## 5. Not everything belongs in the log
|
||
|
||
RFC 001 §5 draws this line for wings; the same discipline applies between the two stores. The short form, expanded for operators in `docs/fleet-memory.md`:
|
||
|
||
| Put it in | When |
|
||
|---|---|
|
||
| A **drawer** (palace) | It will still be true, and worth finding, next month. Nobody in particular needs to act. Retrieval is by meaning. |
|
||
| An **event** (log) | A named agent must act, reply, or be stopped. Retrieval is by address and order. |
|
||
| **Both** | The durable finding goes in a drawer; the event says "there is a new finding, here is the drawer id". This is the recommended pattern for anything a peer must *know* rather than *do*. |
|
||
|
||
The failure mode in each direction: a finding filed only as an event is invisible to semantic search and will be re-derived by the next agent; an ask filed only as a drawer is addressed to nobody and will be found, if ever, by accident.
|
||
|
||
---
|
||
|
||
## 6. Security model
|
||
|
||
**The log authenticates the fleet, not the agent.** Precisely:
|
||
|
||
| Control | State | Evidence |
|
||
|---|---|---|
|
||
| Transport auth | One shared bearer token, `hmac.compare_digest`, 401 otherwise; required for non-loopback binds | `mcp_server.py:7137-7141, 7685` |
|
||
| `from_agent` authenticity | ❌ **Not checked against anything.** Shape-validated only; any client may append an event claiming to be any agent | `logstream.py:103-122` is the only check |
|
||
| Read authorization | ❌ **None.** Any token holder may list every event addressed to anyone, and fetch any artifact by id — including patch contents | no scoping in the query path |
|
||
| Stream auth | ✅ SSE follows the same bearer policy: *"Events and artifacts expose work metadata and patch contents"* | `mcp_server.py:7189-7194` |
|
||
| Transport security | TLS via `--tls-cert`/`--tls-key` (both or neither); Host pin against DNS rebinding; Origin check; 16 MiB request cap | `mcp_server.py:6920-6947` |
|
||
|
||
For a single-operator fleet behind one token this is adequate, and it must be stated rather than implied, because two useful consequences follow directly from it:
|
||
|
||
1. **Impersonation is trivial** — do not treat `from_agent` as evidence of origin in any security decision.
|
||
2. **That same property is the only way to test the mailbox.** A positive control requires writing an event *from* an agent you are not, so that your own client does not self-filter it. This is a supported technique precisely because the field is unauthenticated (**measured 2026-08-26**).
|
||
|
||
If per-agent identity is ever needed, the correct home is the same place RFC 001 §7.3 puts provenance: the authenticated credential at the boundary, not a self-asserted field.
|
||
|
||
---
|
||
|
||
## 7. Landmines
|
||
|
||
### 7.1 There is no idempotency guard on append — a retried write duplicates
|
||
|
||
`append_event` mints a fresh id and INSERTs, with no dedup lookup of any kind (`logstream.py:625-716`). `put_artifact` likewise. Two byte-identical calls produce two events with different ids and different `seq`.
|
||
|
||
The asymmetry is stark: the *replication* path is rigorously idempotent (`logstream.py:1174-1180` checks `id` **or** `(origin_replica, origin_seq)` before applying, and returns early for its own echoed ops). Peer replay is safe; **client retry is not.**
|
||
|
||
**Action:** this makes the palace's general rule — *a timeout usually means the write completed; verify, don't retry* — load-bearing rather than advisory. For a drawer, a blind retry costs a dedup-detectable duplicate. For an event it silently forks a coordination thread into two ids, and a mailbox will then show two owed items that must each be closed. After any `event_append` / `patch_submit` timeout, verify with `event_list(correlation_id=…)` before re-issuing.
|
||
|
||
### 7.2 `status="open"` is not owed-ness
|
||
|
||
Covered in §3.3 and repeated here because it is the mistake most likely to be made by someone reading only the tool descriptions: the mailbox query `to_agent=<me> status=open` **never shrinks as you work**. Treating its length as a to-do count means re-answering answered asks forever.
|
||
|
||
**Action:** derive per §3.3, or use a client that does.
|
||
|
||
### 7.3 `seq` is local arrival order and is meaningless across replicas
|
||
|
||
`seq` is the local `rowid`. The moment a second replica exists, a remote event that was *authored* earlier can arrive *later* and receive a higher `seq`. The owed-set derivation in §3.3 compares `seq`, so it is sound only on a single replica.
|
||
|
||
**Action:** on a hub-and-spoke fleet (all clients on one replica) this is correct today. `hlc` is the field to switch to when `mesh_peers` reports actual peers — and that switch must happen *before* multi-replica, not after.
|
||
|
||
### 7.4 A broadcast reaches no mailbox, and `to_agent=NULL` reaches nobody at all
|
||
|
||
`to_agent=<x>` matches `x` **or** `'*'` at the SQL level (`logstream.py:958-960`), so a broadcast *is* visible to a listing agent. But the reference owed-set derivation excludes `to_agent='*'` deliberately — a broadcast owes nobody a reply, and if it entered every mailbox, every machine would think it personally owed the same answer.
|
||
|
||
Two consequences: **broadcasting an ask reaches no owed set at all** (the "don't broadcast an ask" anti-pattern is mechanically enforced, not merely advised), and to reach a whole fleet with something actionable you must write **one directed event per device**, sharing a `correlation_id` so the thread stays joinable. Separately, an event written with **no** `to_agent` matches no `to_agent=` query ever and is addressed to nobody.
|
||
|
||
### 7.5 `since_event_id` raises on an unknown id
|
||
|
||
An unresolvable cursor raises `ValueError` rather than returning empty (`logstream.py:966-970`) — stronger than the tool description promises, and benign on one replica. On a second replica, a resuming watcher whose cursor has not yet replicated will **hard-fail instead of waiting**.
|
||
|
||
### 7.6 A rejected append returns HTTP 200
|
||
|
||
Failures come back as `{"success": false, "error": …}` in a 200 response, not a JSON-RPC error (`mcp_server.py:4581-4584`).
|
||
|
||
**Action:** a client that checks only transport status silently loses the event. Check the payload.
|
||
|
||
### 7.7 The same agent name on two devices is indistinguishable
|
||
|
||
No uniqueness, no registry, no warning. Both machines' events carry one `from_agent`, and each machine's "you cannot owe yourself" filter will discard the other's asks. This is the failure the `<harness>@<device>` convention exists to prevent (§3.2).
|
||
|
||
### 7.8 `GET /logstream/events` does not exist
|
||
|
||
The complete GET route table is `/healthz`, `/statusz`, `/logstream/stream`, `/sync/{version_vector,ops,artifact,peers}`; POST serves `/mcp` only. A recursive grep for `/logstream` in the package yields three hits, all `/logstream/stream`.
|
||
|
||
⚠️ **Correction.** An earlier measurement observed `/logstream/events`, `/logstream/stream` and `/sync/peers` all returning 404 against a deployment and attributed all three to a reverse proxy exposing only `/mcp`. That inference was right for `/sync/*` and `/logstream/stream` and **wrong for `/logstream/events`**, which would 404 on a directly-reachable server too. Distinguish "route absent" from "route blocked" before blaming infrastructure.
|
||
|
||
### 7.9 Dogfood scars, preserved because each cost someone a session
|
||
|
||
- A `patch` artifact stored without its trailing newline truncates the last hunk line — hence the advisory warnings (`logstream.py:718-742`).
|
||
- `--type Task.Request` was rejected while `--type Task.Request --type patch.ready` was *silently accepted and matched nothing*, leaving a watcher waiting forever: single-valued filters were validated by pushdown, multi-valued ones compared raw (`sanitize_watch_spec`, `logstream.py:245-278`). `type` is lowercase-only for this reason.
|
||
- `event_wait` rejected a `limit` that `event_list` accepted — *"reported by windows-codex during dogfood"* (`mcp_server.py:4666-4670`).
|
||
|
||
### 7.10 The main database file's mtime is not a liveness signal
|
||
|
||
WAL. **Measured 2026-08-26** on the primary: `logstream.sqlite3` mtime three days old while `logstream.sqlite3-wal` was 865 KiB and seconds old. An operator checking whether coordination is live must look at the `-wal` file, or query.
|
||
|
||
---
|
||
|
||
## 8. Phasing
|
||
|
||
### 8.1 Implemented
|
||
|
||
| Phase | Deliverable | State |
|
||
|---|---|---|
|
||
| A | `events`/`artifacts`/`event_artifacts` + append/list/ack + size limits | ✅ 3.7.x |
|
||
| B | Long poll (`event_wait`) and SSE push | ✅ |
|
||
| C | `patch_submit` convenience + patch advisories | ✅ |
|
||
| D | HLC on every event; `(origin_replica, origin_seq)` unique index | ✅ 3.8.0 |
|
||
| E | Client-side auto-delivered mailbox | ✅ toolkit `5b8d78f`; first live delivery measured 2026-08-26 |
|
||
|
||
### 8.2 Deferred, and what RFC 004 owes
|
||
|
||
Replication exists in code — `version_vector()`, `list_ops()`, `apply_remote_event()`, `apply_remote_artifact()` and the four `GET /sync/*` routes (`logstream.py:1096-1261`, `mcp_server.py:7206-7256`) — and is attributed in comments to "RFC 004 step 0". **That document does not exist.** Until it does, this RFC records the two properties a reader most needs: the apply path *is* idempotent (§7.1), and switching the owed-set derivation from `seq` to `hlc` is a precondition for a second replica, not a follow-up (§7.3).
|
||
|
||
---
|
||
|
||
## 9. Open decisions
|
||
|
||
1. **Retention.** No TTL, compaction or pruning exists, and `mempalace sync` does not touch the log (§2). For a fleet log this is mostly a feature — nothing is lost by being offline for weeks — but every 4 MiB artifact is permanent. Decide a policy before the log outgrows a comfortable backup, or decide explicitly that permanence is the policy.
|
||
2. **Idempotency key.** Should `event_append` accept an optional client-supplied dedupe key so a retried call is a no-op? This is the one change that would make §7.1 disappear.
|
||
3. **`hlc` as the mailbox join key**, replacing `seq`. Required before a second replica (§7.3). Cheap now, breaking later.
|
||
4. **Per-agent identity.** Do we ever want `from_agent` to be authenticated (§6), or is "authenticates the fleet, not the agent" the permanent contract?
|
||
5. **Should `GET /logstream/events` exist?** A read-only HTTP tail would let non-MCP consumers (dashboards, CI) follow a stream without an MCP client. Today they must hold SSE or speak MCP.
|
||
6. **Undocumented limits.** The 64 KiB `metadata` cap and the lowercase-`type` regex are enforced server-side but absent from the tool schemas (§3.1). Document, or relax.
|
||
7. **Agent-name registry.** Nothing prevents two devices sharing one `from_agent` (§7.7). A warning at append time would be cheap.
|
||
|
||
### 9.1 Retention — decided direction (2026-08-26)
|
||
|
||
Decision 1 above is **settled in direction**: old traffic should eventually move out of the way, `logrotate`-style. Moved aside, not necessarily destroyed — the audit trail of who asked whom for what is worth more than the disk it occupies.
|
||
|
||
The economics point at a two-tier design rather than one sweep. Events are tiny (a body cap of 256 KiB, and in practice a few KB); artifacts are capped at **4 MiB each** and are stored as content in-row. So:
|
||
|
||
- **Tier the artifacts first.** Keep every `events` row indefinitely — they are the index and the audit trail — and move artifact *content* to cold storage past the window, keeping the `kind`, `sha256`, `size_bytes` and `created_by` stub in place. A stub still answers "what was handed over, by whom, verified how"; only the bytes go cold. This reclaims nearly all of the space with none of the join risk below.
|
||
- **Rotate events by thread, never by row.** If rotation is wanted for events too, the unit must be the **`correlation_id` thread**, and only a thread whose latest event is terminal (§3.3) and older than the window. Archiving an *ask* while leaving its *reply* behind — or the reverse — breaks the owed-set join, and the two failure modes are both bad: an ask that can never be cleared resurfaces as owed forever, or a reply is orphaned from what it answered.
|
||
|
||
Three constraints any implementation has to respect:
|
||
|
||
1. **Never rotate an event that is still owed.** The owed set is derived at read time (§3.3), so an unanswered ask has no marker distinguishing it from a stale one except that derivation. A time-based sweep alone would silently discard live obligations from a machine that has simply been offline for a month — exactly the case this log exists to serve.
|
||
2. **Rotation invalidates held cursors.** `since_event_id` raises on an unknown id rather than returning empty (§7.5), so archiving an event that a watcher still holds as its resume point turns that watcher's next poll into an error. Either rotation is announced far enough ahead of any live cursor, or the anchor lookup learns to fall back to `created_at`/`hlc` when the id is gone.
|
||
3. **Archive before delete, and verify.** Same discipline as any palace destructive op: write the cold copy, verify row counts and `sha256` for artifacts, and only then `DELETE` + `VACUUM` the live database. A `--dry-run` that reports what would move, in thread units, is the minimum interface.
|
||
|
||
Open sub-questions: the window length (a quarter is the obvious first guess); whether cold storage is a sibling `logstream-archive-<period>.sqlite3` or plain files on disk; and whether archives are queryable through the same tools behind an explicit opt-in flag, or simply left as files for a human to open when a question reaches back that far.
|
||
|
||
---
|
||
|
||
## 10. Evidence index
|
||
|
||
| Claim | Where |
|
||
|---|---|
|
||
| Store is `logstream.sqlite3` inside the palace dir | `logstream.py:51`; `mcp_server.py:1108-1117` |
|
||
| Full DDL, indexes, WAL | `logstream.py:447-497`, `552-556` |
|
||
| Design constraints (quoted in §3) | `logstream.py:1-18` |
|
||
| Event id format, ordering not carried by id | `logstream.py:92-102` |
|
||
| `seq` = rowid, surfaced at read | `logstream.py:580-586` |
|
||
| `origin_seq` assigned in-transaction | `logstream.py:700-706` |
|
||
| `hlc` populated per append; format and sort property | `logstream.py:668`; `hlc.py:1-21` |
|
||
| Cursor stays local, HLC is merge order | `hlc.py:19-21` |
|
||
| `origin_replica` from palace dir | `logstream.py:422-436`; `replica.py:31-49` |
|
||
| Server-side `created_at`, second precision | `logstream.py:667` |
|
||
| No dedup on append; remote apply *is* idempotent | `logstream.py:625-716` vs `1174-1180` |
|
||
| Artifact sha256 non-unique index | `logstream.py:447-497`, `744-810` |
|
||
| `artifact_ids` referential check | `logstream.py:687-694`, `1181-1189` |
|
||
| Status set; artifact kinds; size limits | `logstream.py:67-69`, `55-57`, `765-778` |
|
||
| `type` regex, lowercase ≤64 | `logstream.py:76, 124-134` |
|
||
| Broadcast matching in SQL | `logstream.py:958-960` |
|
||
| `since_event_id` anchor + raise | `logstream.py:963-978` |
|
||
| Order by rowid; limit clamp 500 | `logstream.py:947` |
|
||
| `preview` truncates body to 200 chars | `mcp_server.py:4587-4607` |
|
||
| `event_wait` backoff, cap, timeout result | `logstream.py:1004-1039` |
|
||
| `ack_event` appends, never mutates; `ack_of` in metadata | `logstream.py:812-845` |
|
||
| Patch advisories | `logstream.py:718-742` |
|
||
| Lock exemptions, both, with rationale | `mcp_server.py:6887-6902`, `434-449` |
|
||
| Bearer auth; SSE auth required | `mcp_server.py:7137-7141`, `7189-7194` |
|
||
| SSE implementation, heartbeat, client cap | `mcp_server.py:7570-7671` |
|
||
| Rejected append returns 200 + `success:false` | `mcp_server.py:4581-4584` |
|
||
| `sync` does not touch the log | `cli.py:1057`; `sync.py` (no logstream refs) |
|
||
| No retention/pruning anywhere | no `DELETE FROM events`/`artifacts` in the package |
|
||
| Replication surface, attributed to RFC 004 | `logstream.py:1096-1261`; `mcp_server.py:7206-7256` |
|
||
|
||
### 10.1 Where the shipped tool descriptions and the code disagree
|
||
|
||
| Tool-description claim | Verdict |
|
||
|---|---|
|
||
| "append-only; corrections are new events" | True of content. Precisely: no event's *content* is ever mutated; `append_event` does issue one `UPDATE … SET origin_seq = rowid` on the row it just inserted, and the 3.8.0 migration backfills `origin_replica`/`origin_seq`/`hlc` on pre-existing rows (`logstream.py:504-550`). |
|
||
| "`since_event_id` cannot skip anything" | True, and stronger than stated — it raises on an unknown id (§7.5). |
|
||
| "body max 256 KiB", "artifact max 4 MiB, UTF-8 only", "timeout default 60 s max 5 min" | All true; the timeout clamps silently rather than erroring. |
|
||
| "`to_agent=<you>` also matches `*` broadcasts" | True, SQL-level — but see §7.4 for why a broadcast still reaches no mailbox. |
|
||
| "prefer the push stream at `GET /logstream/stream`" | True; route exists and requires auth. |
|
||
| `GET /logstream/events` | ❌ Does not exist (§7.8). |
|
||
| `metadata` size limit; `type` charset | ❌ Enforced but undocumented (§3.1). |
|
||
|
||
---
|
||
|
||
## 11. See also
|
||
|
||
- `docs/fleet-memory.md` — the operator-facing companion: what the palace's stores are *for*, and when to use a drawer versus an event.
|
||
- `docs/rfc-001-global-palace.md` — §5 what should not be global, §7.2 the `sync` hazard, §7.3 provenance at the boundary, §7.5 single-writer expectations.
|
||
- `docs/rfc-002-joiner.md` — replaying a second palace into a shared primary.
|
||
- `extensions/pi/README.md` §2 — the client-side mailbox: delivery points, gating, and why it is a client concern.
|
||
- `~/.agents/skills/mempalace/SKILL.md` §"Cross-Machine Coordination" — **normative for agent behaviour**; this RFC is normative for mechanism.
|