diff --git a/CHANGELOG.md b/CHANGELOG.md index 0a5f4e7..37d9249 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -11,7 +11,7 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`). --- -## Unreleased +## v1.8.8 — 2026-08-26 The vendored `mempalace` skill snapshot stops being anonymous, and the container starts saying which copy of each skill it is actually reading. @@ -45,6 +45,23 @@ being fixed.** only, `2` cannot determine (ref absent from this clone). `AGENTS.md` step 2 rewritten to state all three, since its promise that "the message distinguishes the two" was exactly what the branch was breaking. +- **The staleness `NOTICE` then asserted a direction it had never tested** — the + same defect one layer down, found by pi@emb-7kj4vr4g against the real state of + its own host. The branch fired the notice on "recorded ≠ HEAD" and announced + that HEAD was the newer side, so a clone that was merely *behind* was told + *"has moved to 82a8d3c; the snapshot describes the older 5fd0d5c"* when + `82a8d3c` is `5fd0d5c`'s **ancestor**. Harmless to the verdict (`rc` stayed 0, + nothing was mis-verified) but it points the operator at a refresh — a + ~67-minute base rebuild — when the real remedy is `git pull`. It now tests + ancestry with the `merge-base --is-ancestor` primitive the refresh path two + sections above already used, and reports three distinct verdicts: **stale** + (recorded is an ancestor — refresh), **your clone is behind** (HEAD is an + ancestor — pull, do not refresh), **diverged** (neither). All three verified by + execution; only the first was right before. +- **`--help` died with `unknown option: --help`.** The strict argument loop that + closed the silent-ignore hole never added a `--help` case, so the one script + whose argument *order* was itself a landmine had an erroring discoverability + path. It now prints its own header block. - **`VENDORED.md` contradicted itself, in the release whose stated invariant is non-contradiction.** Its hand-maintained "Snapshot provenance at last refresh" line named skillset `670f7f1` — seven commits behind the ARG, and *the very @@ -88,8 +105,22 @@ earlier. Fixed in skillset `5fd0d5c`, which derives owed-ness by joining on terminal reply suppresses every later ask on the same thread forever. Because the skillset is mounted live on every enrolled host, that correction was already deployed fleet-wide before this image was built; the vendored snapshot is -resynced to it (`c04cd15` → `5fd0d5c`) so the no-clone fallback does not ship -the withdrawn rule. Canary re-verified bidirectionally against the new bytes. +resynced to it (`c04cd15` → `5fd0d5c` → `6eb20af`) so the no-clone fallback does +not ship the withdrawn rule. Canary re-verified bidirectionally against the new +bytes. + +`6eb20af` adds the limit of that ordering test, found when pi@emb-7kj4vr4g +verified it rather than adopting it: **`seq` is replica-local.** It equals +`origin_seq` today only because one replica authors events for all four machines, +so a second replica could order the same pair differently and derive a different +owed-set from the same log — use `hlc` (already on every event, total and +causally consistent) once `mesh_peers` reports any peer. Documented as reasoning, +not measurement, since a second replica cannot be stood up to test it. The part +worth keeping is the **asymmetry**: local-`seq` skew makes an answered item +*resurface* (noise, self-correcting, visible), while a timestamp comparison +*suppresses an unanswered ask forever* (silent, permanent) — so anyone tempted to +"fix" a resurfacing item with `created_at` would be trading the safe failure for +the dangerous one. **Also carried, previously undocumented:** `dbb7879` resynced the vendored `mempalace` snapshot to skillset `c04cd15` ("the withdrawal only holds where the @@ -275,14 +306,17 @@ the skillset actually is — a maintainer's clone, or any running container. because when every host is a thin client of one palace the stamped agent name is the only thing distinguishing them. Set both or neither: a container without the device var can read the log but is addressable by nobody. - - the skillset's `mempalace` skill (`82a8d3c`, live on every host that mounts + - the skillset's `mempalace` skill (`6eb20af`, live on every host that mounts the skillset, no rebuild needed) — the **norms**: a mailbox query at wake-up, and the sender-declared ack contract, where a *directed* event with - `status="open"` is owed a reply and a `*` broadcast owes nothing. Measured - while designing it: an unfiltered mailbox returned 5 events, 4 of them - finished broadcasts from eight days earlier, where `status="open"` returned - exactly the 1 that needed an answer — an unfiltered mailbox trains you to - ignore it, so the filter is the feature. + `status="open"` is owed a reply and a `*` broadcast owes nothing. The + `status` filter earns its place by dropping broadcast noise — measured, an + unfiltered mailbox returned 5 events, 4 of them finished broadcasts from + eight days earlier — but that is **all** it does; it does not compute + owed-ness, and the version of this entry that claimed otherwise is withdrawn + above. Measured today, both machines: the raw filter returns 2 asks here and + 1 there, **every one already answered**, while the derivation returns 0 for + both. Dropping noise and deciding what is owed are two different jobs. - mempalace-toolkit `extensions/pi/README.md` (`e70bef2`) — the **mechanism**, including that the bridge is *write-only* today (it stamps events going out and never reads the log, so nothing in this image polls on the agent's @@ -290,16 +324,19 @@ the skillset actually is — a maintainer's clone, or any running container. implements `GET /logstream/stream`, but a reverse proxy exposing only `/mcp` makes it unreachable — verified by 404s against the real endpoint. - ⚠️ **This makes the vendored snapshot stale on purpose.** The skill edit is in - the skillset (`82a8d3c`), so `SKILLSET_SNAPSHOT_REF` still records `c04cd15` - and `scripts/vendor-mempalace-skill.sh --check` now exits 1 with - *"has moved to 82a8d3c; the snapshot describes the older c04cd15"*. Refreshing - it is a deliberate release-day decision, not an oversight — hence the new - step 2 in `AGENTS.md` § *Release-day checklist*, which states both that the - refresh costs a base rebuild and that skipping it is legitimate because every - enrolled host reads its live clone. What is not legitimate is skipping it - *silently*, which is precisely what the new manifest fields and - `pi-devbox-version` output make impossible. + ⚠️ **The snapshot was refreshed rather than left stale.** The skill edits landed + in the skillset (`5fd0d5c`, then `6eb20af`), so `SKILLSET_SNAPSHOT_REF` was + resynced to match and `scripts/vendor-mempalace-skill.sh --check` is a clean + `OK` with no notice: the no-clone fallback carries the **corrected** protocol, + not the withdrawn one. That mattered more than currency usually does, because + the superseded copy contained an instruction — the `mesh_peers` gate — that + actively told a reader to skip the feature. Refreshing remains a deliberate + release-day decision rather than an automatic one: it costs a base rebuild, and + skipping it is legitimate because every enrolled host reads its live clone. + What is not legitimate is skipping it *silently*, which is what the new manifest + fields and `pi-devbox-version` output make impossible — hence step 2 in + `AGENTS.md` § *Release-day checklist*. In this release the refresh was free: + `rootfs/` was already changing, so the base rebuild was forced anyway. --- diff --git a/Dockerfile.variant b/Dockerfile.variant index 2e1232a..76716ef 100644 --- a/Dockerfile.variant +++ b/Dockerfile.variant @@ -305,7 +305,7 @@ ARG MEMPALACE_TOOLKIT_REF=main # no ~67-minute base rebuild. (scripts/check-base-hash.sh scans only # Dockerfile.base, so no folding into the base hash is required — nor would # it be correct, since this ARG changes nothing about the base's contents.) -ARG SKILLSET_SNAPSHOT_REF=5fd0d5c406df506fd0e70977b6bd87f5fbc3b803 +ARG SKILLSET_SNAPSHOT_REF=6eb20af181f0147cb8c1377f6e36a6a47a68e8e5 # Dockerfile.base sets description="pi-devbox — base image (variant-independent)" # and every variant INHERITS it, so both published images used to advertise diff --git a/rootfs/usr/local/share/pi-devbox/skills/mempalace/SKILL.md b/rootfs/usr/local/share/pi-devbox/skills/mempalace/SKILL.md index eca8708..44b0a0e 100644 --- a/rootfs/usr/local/share/pi-devbox/skills/mempalace/SKILL.md +++ b/rootfs/usr/local/share/pi-devbox/skills/mempalace/SKILL.md @@ -462,6 +462,24 @@ every later ask on the same `correlation_id` for good; verified on a live thread where a `ready` reply at `seq` 16 sits *before* the request at `seq` 17 that it obviously cannot have answered. +**On a real mesh, compare `hlc` instead.** `seq` is *replica-local*: it equals +`origin_seq` today only because a single replica authors events for every +machine. Enrol a second replica and a late-syncing peer event gets a late local +`seq`, so two replicas can order the same pair differently and derive different +owed-sets from the same log. Every event already carries `hlc` +(`--`), which is total and causally consistent. +So: compare `seq` while `mempalace_mesh_peers` reports no peers, `hlc` once it +reports any, and `created_at` never. (This is a legitimate use of `mesh_peers` — +choosing an ordering key — not the discredited gate on *whether* to read your +mailbox at all.) + +**The failure directions are not symmetric, which is why this is safe to get +slightly wrong.** Local-`seq` skew can make an already-answered item *resurface* +as owed: noise, self-correcting, and visible. A timestamp comparison can +*suppress an unanswered ask forever*: silent and permanent. So if you ever see an +item you know you answered come back, do **not** "fix" it by reaching for +`created_at` — you would be trading the safe failure for the dangerous one. + This also supplies the "taken, not finished" state that looked missing: `claimed` and `ready` are deliberately **not** terminal, so work you have picked up keeps resurfacing until you close it out. No extra convention, no new field. diff --git a/scripts/vendor-mempalace-skill.sh b/scripts/vendor-mempalace-skill.sh index 385295e..086a1b5 100755 --- a/scripts/vendor-mempalace-skill.sh +++ b/scripts/vendor-mempalace-skill.sh @@ -86,7 +86,13 @@ for arg in "$@"; do case "$arg" in --check) MODE="check" ;; --force) FORCE=1 ;; - --*) die "unknown option: $arg" ;; + # This is the one script whose argument ORDER was itself a landmine, so the + # path that documents the trap must not be the path that errors. + -h|--help) + awk 'NR>1 && /^#/ { sub(/^# ?/, ""); print; next } NR>1 { exit }' "$0" + exit 0 + ;; + --*) die "unknown option: $arg (try --help)" ;; *) [ -z "$ROOT" ] || die "unexpected extra argument: $arg (root already set to $ROOT)" ROOT="$arg" @@ -110,6 +116,12 @@ fi [ -d "$ROOT/.git" ] || die "not a git clone: $ROOT" [ -f "$ROOT/$REL_PATH" ] || die "no $REL_PATH in $ROOT" [ -f "$VENDORED" ] || die "vendored snapshot missing: $VENDORED" +# LOAD-BEARING, DO NOT DELETE AS "REDUNDANT WITH THE EXISTENCE PROBES": -f +# accepts an empty file, and sha256 of an empty file equals sha256 of a failed +# pipeline's empty stdin. Guarding it HERE, before mode dispatch, makes that +# collision unreachable by construction rather than by a probe further down -- +# which also means no test below exercises the collision any more. Remove this +# line and the false "OK" for a nonexistent ref returns with nothing failing. [ -s "$VENDORED" ] || die "vendored snapshot is empty: $VENDORED" head_sha=$(git -C "$ROOT" rev-parse HEAD 2>/dev/null) || die "cannot read HEAD of $ROOT" @@ -200,8 +212,20 @@ if [ "$MODE" = "check" ]; then # clean and the ref simply moved on — a message that names the wrong cause # is the same defect class as a canary pinned to a deleted phrase. if [ "$recorded" != "$head_sha" ] && [ "$blob_sha" = "$upstream_sha" ]; then - printf 'NOTICE: %s has moved to %s; the snapshot describes the older %s (stale, not untruthful)\n' \ - "$ROOT" "${head_sha:0:7}" "${recorded:0:7}" >&2 + # Do not ASSERT which side is newer — test it. Asserting that HEAD is the + # newer side points the operator at a refresh (which costs a ~67-minute + # base rebuild) when the real remedy may be `git pull` in this clone. The + # refresh path below already uses this primitive; reuse it here. + if git -C "$ROOT" merge-base --is-ancestor "$recorded" "$head_sha" 2>/dev/null; then + printf 'NOTICE: %s has moved on to %s; the snapshot describes the older %s (stale, not untruthful — refresh to catch up)\n' \ + "$ROOT" "${head_sha:0:7}" "${recorded:0:7}" >&2 + elif git -C "$ROOT" merge-base --is-ancestor "$head_sha" "$recorded" 2>/dev/null; then + printf 'NOTICE: %s is BEHIND at %s; the snapshot describes the newer %s — pull this clone, do NOT refresh the snapshot\n' \ + "$ROOT" "${head_sha:0:7}" "${recorded:0:7}" >&2 + else + printf 'NOTICE: %s (HEAD %s) and the recorded %s have DIVERGED — neither is an ancestor of the other; reconcile the clone before refreshing\n' \ + "$ROOT" "${head_sha:0:7}" "${recorded:0:7}" >&2 + fi else printf 'NOTICE: the working tree of %s differs from the snapshot (HEAD %s)\n' \ "$ROOT" "${head_sha:0:7}" >&2