From d4d8bb6109972d8a02caeee236cfabba664f48b3 Mon Sep 17 00:00:00 2001 From: Joakim Persson Date: Wed, 26 Aug 2026 23:07:40 +0200 Subject: [PATCH] =?UTF-8?q?docs:=20retention=20direction=20for=20the=20coo?= =?UTF-8?q?rdination=20log,=20and=20drop=20the=20operator's=20name=20from?= =?UTF-8?q?=20RFC=20002=20=C2=A75?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit RFC 003 §9.1 — retention is settled in direction (rotate old traffic out of the way, logrotate-style, moved aside rather than destroyed) and the sketch records the parts that are not obvious: - Tier artifacts before events. Events are a few KB; artifacts are capped at 4 MiB and stored in-row, so moving artifact CONTENT cold while keeping the kind/sha256/size/created_by stub reclaims nearly all the space and keeps the audit trail ("what was handed over, by whom, verified how") intact. - If events rotate at all, the unit is the correlation_id THREAD whose latest event is terminal — never the row. Archiving an ask while leaving its reply (or the reverse) breaks the owed-set join, and both failure modes are bad: an ask that can never be cleared resurfaces as owed forever, or a reply is orphaned from what it answered. - Never rotate an event that is still owed. Owed-ness is DERIVED at read time, so an unanswered ask is indistinguishable from a stale one except by that derivation — a purely time-based sweep would discard the live obligations of a machine that has merely been offline for a month, which is precisely the case this log exists to serve. - Rotation invalidates held cursors: since_event_id RAISES on an unknown id (§7.5), so archiving an event a watcher holds as its resume point turns its next poll into an error. Either announce rotation ahead of live cursors, or teach the anchor lookup to fall back to created_at/hlc. - Archive → verify (row counts, artifact sha256) → only then DELETE + VACUUM, with a --dry-run that reports in thread units. RFC 002 §5: "Open decisions for ALC" → "Open decisions". RFC 001 never names the operator anywhere; impersonal is the mature precedent and ALC was never a real identifier in the first place (it is the AAAK spec's illustrative code for "Alice", copy-forwarded into ~700 diary entries without verification). Same edit also removes two device names and a hostname from §5.3, which is host inventory and belongs in the private fleet repo, not a public one. The mechanism it teaches — a palace in a Docker named volume dies on the next container recreate, so census before flipping — is unchanged and is the part that mattered. --- docs/rfc-002-joiner.md | 6 +++--- docs/rfc-003-coordination-log.md | 17 +++++++++++++++++ 2 files changed, 20 insertions(+), 3 deletions(-) diff --git a/docs/rfc-002-joiner.md b/docs/rfc-002-joiner.md index f9be2db..5f276be 100644 --- a/docs/rfc-002-joiner.md +++ b/docs/rfc-002-joiner.md @@ -183,12 +183,12 @@ correctness, not for this join. `filed_at` spread over the same parent drawers — 12 in May, 52 in June, 13,619 in July, 882 in August — is the concrete case for regime A: an MCP replay would restamp all of it to the join date. -## 5. Open decisions for ALC +## 5. Open decisions 1. **Chronology: keep it or flatten it?** Regime A keeps it and costs a service stop plus an rsync of the - source palace onto synlig. Regime B is simpler and loses it. This is the fork in the road. + source palace onto the primary host. Regime B is simpler and loses it. This is the fork in the road. 2. **Hallways for joined content:** re-mine to rebuild, or accept degraded `traverse`? -3. **Do MBP-M1-2020 and tor-ms22 keep their palaces on persistent storage?** If either is a Docker named +3. **Do the secondary machines keep their palaces on persistent storage?** If a palace lives in a Docker named volume rather than a bind mount, its un-migrated content dies on the next container recreate — so the census (Phase A) is time-sensitive there, and flipping before censusing is risky. 4. **Does §7.6 get fixed client-side (in the joiner) or upstream (probe-and-skip in `diary_write`)?** The diff --git a/docs/rfc-003-coordination-log.md b/docs/rfc-003-coordination-log.md index 2c2d389..be192cb 100644 --- a/docs/rfc-003-coordination-log.md +++ b/docs/rfc-003-coordination-log.md @@ -337,6 +337,23 @@ Replication exists in code — `version_vector()`, `list_ops()`, `apply_remote_e 6. **Undocumented limits.** The 64 KiB `metadata` cap and the lowercase-`type` regex are enforced server-side but absent from the tool schemas (§3.1). Document, or relax. 7. **Agent-name registry.** Nothing prevents two devices sharing one `from_agent` (§7.7). A warning at append time would be cheap. +### 9.1 Retention — decided direction (2026-08-26) + +Decision 1 above is **settled in direction**: old traffic should eventually move out of the way, `logrotate`-style. Moved aside, not necessarily destroyed — the audit trail of who asked whom for what is worth more than the disk it occupies. + +The economics point at a two-tier design rather than one sweep. Events are tiny (a body cap of 256 KiB, and in practice a few KB); artifacts are capped at **4 MiB each** and are stored as content in-row. So: + +- **Tier the artifacts first.** Keep every `events` row indefinitely — they are the index and the audit trail — and move artifact *content* to cold storage past the window, keeping the `kind`, `sha256`, `size_bytes` and `created_by` stub in place. A stub still answers "what was handed over, by whom, verified how"; only the bytes go cold. This reclaims nearly all of the space with none of the join risk below. +- **Rotate events by thread, never by row.** If rotation is wanted for events too, the unit must be the **`correlation_id` thread**, and only a thread whose latest event is terminal (§3.3) and older than the window. Archiving an *ask* while leaving its *reply* behind — or the reverse — breaks the owed-set join, and the two failure modes are both bad: an ask that can never be cleared resurfaces as owed forever, or a reply is orphaned from what it answered. + +Three constraints any implementation has to respect: + +1. **Never rotate an event that is still owed.** The owed set is derived at read time (§3.3), so an unanswered ask has no marker distinguishing it from a stale one except that derivation. A time-based sweep alone would silently discard live obligations from a machine that has simply been offline for a month — exactly the case this log exists to serve. +2. **Rotation invalidates held cursors.** `since_event_id` raises on an unknown id rather than returning empty (§7.5), so archiving an event that a watcher still holds as its resume point turns that watcher's next poll into an error. Either rotation is announced far enough ahead of any live cursor, or the anchor lookup learns to fall back to `created_at`/`hlc` when the id is gone. +3. **Archive before delete, and verify.** Same discipline as any palace destructive op: write the cold copy, verify row counts and `sha256` for artifacts, and only then `DELETE` + `VACUUM` the live database. A `--dry-run` that reports what would move, in thread units, is the minimum interface. + +Open sub-questions: the window length (a quarter is the obvious first guess); whether cold storage is a sibling `logstream-archive-.sqlite3` or plain files on disk; and whether archives are queryable through the same tools behind an explicit opt-in flag, or simply left as files for a human to open when a question reaches back that far. + --- ## 10. Evidence index