d2764bf78e202aedf41b18ca9ac0bca3fe5a59d6
2 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
d4d8bb6109 |
docs: retention direction for the coordination log, and drop the operator's name from RFC 002 §5
RFC 003 §9.1 — retention is settled in direction (rotate old traffic out of
the way, logrotate-style, moved aside rather than destroyed) and the sketch
records the parts that are not obvious:
- Tier artifacts before events. Events are a few KB; artifacts are capped at
4 MiB and stored in-row, so moving artifact CONTENT cold while keeping the
kind/sha256/size/created_by stub reclaims nearly all the space and keeps the
audit trail ("what was handed over, by whom, verified how") intact.
- If events rotate at all, the unit is the correlation_id THREAD whose latest
event is terminal — never the row. Archiving an ask while leaving its reply
(or the reverse) breaks the owed-set join, and both failure modes are bad:
an ask that can never be cleared resurfaces as owed forever, or a reply is
orphaned from what it answered.
- Never rotate an event that is still owed. Owed-ness is DERIVED at read time,
so an unanswered ask is indistinguishable from a stale one except by that
derivation — a purely time-based sweep would discard the live obligations of
a machine that has merely been offline for a month, which is precisely the
case this log exists to serve.
- Rotation invalidates held cursors: since_event_id RAISES on an unknown id
(§7.5), so archiving an event a watcher holds as its resume point turns its
next poll into an error. Either announce rotation ahead of live cursors, or
teach the anchor lookup to fall back to created_at/hlc.
- Archive → verify (row counts, artifact sha256) → only then DELETE + VACUUM,
with a --dry-run that reports in thread units.
RFC 002 §5: "Open decisions for ALC" → "Open decisions". RFC 001 never names
the operator anywhere; impersonal is the mature precedent and ALC was never a
real identifier in the first place (it is the AAAK spec's illustrative code for
"Alice", copy-forwarded into ~700 diary entries without verification).
Same edit also removes two device names and a hostname from §5.3, which is
host inventory and belongs in the private fleet repo, not a public one. The
mechanism it teaches — a palace in a Docker named volume dies on the next
container recreate, so census before flipping — is unchanged and is the part
that mattered.
|
||
|
|
e1cc7592a1 |
docs: write the RFC 003 that the code has been citing all along, plus an operator-facing fleet-memory guide
`logstream.py`'s module docstring is headed "Agent coordination event log for MemPalace (RFC 003)" and enumerates five "Design constraints (RFC 003)". Comments cite "RFC 003 phase 5", "RFC 003 suggested defaults" and "the first RFC 003 dogfood". Every event/artifact tool description cites RFC 003. The document has never existed — confirmed by searching this repo and the primary host. So this is a retrospective spec: it transcribes what the implementation already believes, and records what it does NOT do. docs/rfc-003-coordination-log.md, verified line-by-line against mempalace 3.8.0 (every claim cites file.py:LINE, indexed in §10 for re-verification). The parts that are not visible from the tool descriptions: - No idempotency guard on event_append or put_artifact (§7.1). The replication path checks `id OR (origin_replica, origin_seq)` before applying; the client path checks nothing. So peer replay is safe and CLIENT RETRY IS NOT — a retried append forks a coordination thread into two ids. This makes the fleet's "a timeout is not a failure, verify before retrying" rule load-bearing rather than advisory. - Coordination traffic is deliberately exempt from BOTH palace locks (§4) — _HTTP_LOCK_FREE_TOOLS and _PEER_WRITER_EXEMPT_TOOLS, each with its own rationale in-source. "One large mine blocks every client" is true of drawer writes and false of coordination writes. - `mempalace sync` never touches the log, and no DELETE FROM events exists anywhere (§2, §9.1) — answering for the logstream a question RFC 001 §7.2 left open, and making the log permanent and unbounded. - from_agent is shape-validated and never authenticated; there is no read scoping at all (§6). The log authenticates the fleet, not the agent. That is simultaneously the security limitation and the only way to positive-control the mailbox. - Three distinct orderings — seq (local arrival rowid), origin_seq (author's counter), hlc (fleet-wide, lexicographically sortable) — and origin_replica identifies the PALACE, not the writer, which is why from_agent/to_agent carry the whole distinction between machines (§3.2). - GET /logstream/events does not exist and never did (§7.8). An earlier measurement saw it 404 and blamed the reverse proxy; that inference was right for /sync/* and /logstream/stream and wrong for this one. Corrected in extensions/pi/README.md §3 too, in place, dated. docs/fleet-memory.md is the operator-facing companion the repo lacked entirely: the front-door README had zero mentions of coordination, so the channel was undiscoverable unless a message happened to arrive. It covers the five stores and what each is for (drawers/wings/rooms, diaries, KG, palace graph, coordination log), what a central palace buys a fleet — awareness, non-repetition of expensive work, and retractions that travel — and a decision flow for drawer vs KG vs event. Six mermaid diagrams, all rendered and inspected as images, not merely parsed: the first pass produced a truncated state label from an HTML entity and a self-loop that drew a meaningless dotted lasso. Validation says "no syntax error"; only looking says "correct". Also: a Documentation table in the top-level README so all of the above is reachable from the front door. Deliberately host-agnostic, per synlig-primary-runbook.md's precedent — no device names, hostnames or operator names in either new document. |