mailbox: join the owed set on hlc, not seq, before a second replica exists
deriveOwed decides "is this ask still owed?" by asking whether one of my own terminal replies is LATER than the ask. It compared `seq` — this database's arrival rowid. On a single hub that is global order, so it was correct; the comment above it already said hlc was the durable key "once mesh_peers reports actual peers". Making the switch now, while one replica means the two orderings agree, costs nothing; making it later means changing the rule while two machines already disagree about order. The bug being pre-empted is specific: with a second replica the same event gets a different `seq` in each database, because arrival order is not authorship order. A reply authored after its ask can arrive first and take the lower seq; the join then concludes "no later reply exists" and an already-answered ask reappears as owed — permanently, on that machine. isStrictlyAfter() prefers `hlc` when both events carry one and falls back to `seq` otherwise (a server predating the field, or an un-backfilled row). hlc is rendered fixed-width, <unix_ms:13 digits>-<counter:6 hex>-<replica_id>, so a plain string comparison IS the causal comparison, with the replica id as final tiebreak. Verified before writing the code that the field is actually on the wire — event_list returns it per event (logstream.py:593) — because a fallback that never fires would have made this a no-op dressed as a fix. `created_at` stays rejected, and the reason is now written down where the decision is: it is server-generated at second precision, so ties are routine, and a tie can suppress an UNANSWERED ask. Both remaining failure modes are on the noisy-but-visible side — an answered item resurfacing is annoying, an unanswered ask going silent defeats the mailbox. Tested: 18 cases against the extracted comparator — hlc later/earlier/equal, same-ms counter ties in hex (0x10 vs 0x9, which is where a non-padded format would break), cross-replica tiebreak, both mesh reorderings, every seq-fallback path, and degenerate input (null seq, numeric hlc, empty strings, no keys) which must never claim "answered". All pass. Real-data check on the positive-control pair: seq 26/27 carry hlc ...3071857/...3574085, so the orderings agree today and the switch is a no-op now and correct later. The cursor keeps the opposite ordering ON PURPOSE — since_event_id is local-arrival ordered so a tail consumer still sees late-arriving remote ops whose hlc is older (hlc.py:19-21). RFC 003 §9.3 now says so explicitly, because that asymmetry looks like a bug worth "fixing" and is not.
This commit is contained in:
@@ -267,9 +267,11 @@ Covered in §3.3 and repeated here because it is the mistake most likely to be m
|
||||
|
||||
### 7.3 `seq` is local arrival order and is meaningless across replicas
|
||||
|
||||
`seq` is the local `rowid`. The moment a second replica exists, a remote event that was *authored* earlier can arrive *later* and receive a higher `seq`. The owed-set derivation in §3.3 compares `seq`, so it is sound only on a single replica.
|
||||
`seq` is the local `rowid`. The moment a second replica exists, a remote event that was *authored* earlier can arrive *later* and receive a higher `seq`. An owed-set derivation that compares `seq` is therefore sound only on a single replica: on a mesh, a reply can land before the ask it answers, the join concludes "no later reply exists", and an already-answered ask reappears as owed — permanently, on that machine.
|
||||
|
||||
**Action:** on a hub-and-spoke fleet (all clients on one replica) this is correct today. `hlc` is the field to switch to when `mesh_peers` reports actual peers — and that switch must happen *before* multi-replica, not after.
|
||||
**✅ Resolved 2026-08-26 (toolkit `extensions/pi/mempalace.ts`).** The derivation now compares `hlc` when both events carry one, falling back to `seq` only when either lacks it (a server predating the field, or an un-backfilled row). `hlc` is rendered fixed-width — `<unix_ms:13 digits>-<counter:6 hex>-<replica_id>` (`hlc.py:1-21`) — so a plain string comparison *is* the causal comparison, with the replica id as final tiebreak. The field is present on the wire: `event_list` returns it per event (`logstream.py:593`), verified against the live hub the same day.
|
||||
|
||||
Measured on real data: the positive-control pair carries `seq` 26/27 and `hlc` `1787773071857-…`/`1787773574085-…`, so on today's single replica the two orderings agree and the switch is a **no-op now and correct later** — which is the whole reason to make it before a second replica exists rather than after. `created_at` remains the wrong key in both worlds: server-generated at second precision, so ties are routine and a tie can suppress an *unanswered* ask outright.
|
||||
|
||||
### 7.4 A broadcast reaches no mailbox, and `to_agent=NULL` reaches nobody at all
|
||||
|
||||
@@ -331,7 +333,7 @@ Replication exists in code — `version_vector()`, `list_ops()`, `apply_remote_e
|
||||
|
||||
1. **Retention.** No TTL, compaction or pruning exists, and `mempalace sync` does not touch the log (§2). For a fleet log this is mostly a feature — nothing is lost by being offline for weeks — but every 4 MiB artifact is permanent. Decide a policy before the log outgrows a comfortable backup, or decide explicitly that permanence is the policy.
|
||||
2. **Idempotency key.** Should `event_append` accept an optional client-supplied dedupe key so a retried call is a no-op? This is the one change that would make §7.1 disappear.
|
||||
3. **`hlc` as the mailbox join key**, replacing `seq`. Required before a second replica (§7.3). Cheap now, breaking later.
|
||||
3. ~~**`hlc` as the mailbox join key**, replacing `seq`.~~ **Done 2026-08-26** — see §7.3. What remains open is the *reverse* direction: `mempalace_event_list`'s `since_event_id` cursor is deliberately local-arrival-ordered (`hlc.py:19-21` — "a tail consumer must see late-arriving remote ops even though their HLC is older"), so cursor semantics and join semantics use different orderings on purpose. That is correct, and worth stating loudly before someone "fixes" the cursor to match the join.
|
||||
4. **Per-agent identity.** Do we ever want `from_agent` to be authenticated (§6), or is "authenticates the fleet, not the agent" the permanent contract?
|
||||
5. **Should `GET /logstream/events` exist?** A read-only HTTP tail would let non-MCP consumers (dashboards, CI) follow a stream without an MCP client. Today they must hold SSE or speak MCP.
|
||||
6. **Undocumented limits.** The 64 KiB `metadata` cap and the lowercase-`type` regex are enforced server-side but absent from the tool schemas (§3.1). Document, or relax.
|
||||
|
||||
Reference in New Issue
Block a user