mailbox: join the owed set on hlc, not seq, before a second replica exists
deriveOwed decides "is this ask still owed?" by asking whether one of my own terminal replies is LATER than the ask. It compared `seq` — this database's arrival rowid. On a single hub that is global order, so it was correct; the comment above it already said hlc was the durable key "once mesh_peers reports actual peers". Making the switch now, while one replica means the two orderings agree, costs nothing; making it later means changing the rule while two machines already disagree about order. The bug being pre-empted is specific: with a second replica the same event gets a different `seq` in each database, because arrival order is not authorship order. A reply authored after its ask can arrive first and take the lower seq; the join then concludes "no later reply exists" and an already-answered ask reappears as owed — permanently, on that machine. isStrictlyAfter() prefers `hlc` when both events carry one and falls back to `seq` otherwise (a server predating the field, or an un-backfilled row). hlc is rendered fixed-width, <unix_ms:13 digits>-<counter:6 hex>-<replica_id>, so a plain string comparison IS the causal comparison, with the replica id as final tiebreak. Verified before writing the code that the field is actually on the wire — event_list returns it per event (logstream.py:593) — because a fallback that never fires would have made this a no-op dressed as a fix. `created_at` stays rejected, and the reason is now written down where the decision is: it is server-generated at second precision, so ties are routine, and a tie can suppress an UNANSWERED ask. Both remaining failure modes are on the noisy-but-visible side — an answered item resurfacing is annoying, an unanswered ask going silent defeats the mailbox. Tested: 18 cases against the extracted comparator — hlc later/earlier/equal, same-ms counter ties in hex (0x10 vs 0x9, which is where a non-padded format would break), cross-replica tiebreak, both mesh reorderings, every seq-fallback path, and degenerate input (null seq, numeric hlc, empty strings, no keys) which must never claim "answered". All pass. Real-data check on the positive-control pair: seq 26/27 carry hlc ...3071857/...3574085, so the orderings agree today and the switch is a no-op now and correct later. The cursor keeps the opposite ordering ON PURPOSE — since_event_id is local-arrival ordered so a tail consumer still sees late-arriving remote ops whose hlc is older (hlc.py:19-21). RFC 003 §9.3 now says so explicitly, because that asymmetry looks like a bug worth "fixing" and is not.
This commit is contained in:
@@ -132,6 +132,7 @@ const num = (envVal: string | undefined, fallback: number): number => {
|
||||
type LogEvent = {
|
||||
id?: string;
|
||||
seq?: number;
|
||||
hlc?: string;
|
||||
type?: string;
|
||||
status?: string;
|
||||
from_agent?: string;
|
||||
@@ -1007,6 +1008,40 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
}
|
||||
};
|
||||
|
||||
/**
|
||||
* Is `later` strictly after `earlier` in the fleet's ordering?
|
||||
*
|
||||
* Prefer `hlc` — a hybrid logical clock rendered fixed-width
|
||||
* (`<unix_ms:13 digits>-<counter:6 hex>-<replica_id>`), so a plain string
|
||||
* comparison IS the causal comparison, and it stays correct once a second
|
||||
* replica exists. `seq` is the arrival rowid of THIS database, so on a mesh
|
||||
* the same event carries a different `seq` per replica and a reply can land
|
||||
* before the ask it answers.
|
||||
*
|
||||
* Fall back to `seq` only when either side lacks an `hlc` (a server predating
|
||||
* the field, or a row whose backfill did not run). On a single replica the two
|
||||
* agree, so the fallback is not a downgrade today — it is the single-replica
|
||||
* case still being handled after the mesh case became primary.
|
||||
*
|
||||
* NEVER compare `created_at`: it is server-generated at second precision, so
|
||||
* ties are routine, and a tie or a skew there can suppress an UNANSWERED ask
|
||||
* outright. Every failure mode here is deliberately kept on the
|
||||
* noisy-but-visible side — an answered item resurfacing is annoying, an
|
||||
* unanswered ask going silent defeats the mailbox.
|
||||
*/
|
||||
const isStrictlyAfter = (later: LogEvent, earlier: LogEvent): boolean => {
|
||||
if (
|
||||
typeof later.hlc === "string" &&
|
||||
typeof earlier.hlc === "string" &&
|
||||
later.hlc !== "" &&
|
||||
earlier.hlc !== ""
|
||||
) {
|
||||
return later.hlc > earlier.hlc;
|
||||
}
|
||||
if (typeof later.seq !== "number" || typeof earlier.seq !== "number") return false;
|
||||
return later.seq > earlier.seq;
|
||||
};
|
||||
|
||||
/**
|
||||
* Has one of MY events closed this candidate?
|
||||
*
|
||||
@@ -1023,15 +1058,8 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
// it one terminal reply suppresses every LATER ask on the same
|
||||
// correlation_id forever — silently, permanently, and worst on exactly
|
||||
// the long-running threads the correlation join is for.
|
||||
//
|
||||
// Compare `seq`, NEVER `created_at`. `seq` is replica-local (it equals
|
||||
// origin_seq only while a single replica authors for every machine); the
|
||||
// durable key once `mempalace_mesh_peers` reports real peers is `hlc`,
|
||||
// which is total and causally consistent. Local-seq skew can only make an
|
||||
// ANSWERED item resurface (noise, visible), whereas a timestamp
|
||||
// comparison can suppress an UNANSWERED ask outright.
|
||||
if (typeof m.seq !== "number" || typeof candidate.seq !== "number") return false;
|
||||
if (m.seq <= candidate.seq) return false;
|
||||
// See isStrictlyAfter for why the key is `hlc` and not `seq`/`created_at`.
|
||||
if (!isStrictlyAfter(m, candidate)) return false;
|
||||
// (b) join on the exact ack (written for us by event_ack), else on a
|
||||
// shared correlation_id — which is why correlation_id is mandatory on a
|
||||
// directed open: without it there is no key to join a reply back to.
|
||||
|
||||
Reference in New Issue
Block a user