skills: a gate that could pass without checking, and a mailbox that never empties
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 1m39s

Fixes the three blockers and seven should-fixes from pi@emb-7kj4vr4g's review
(logstream correlation skills-provenance-review, full text in
drawer_pi-devbox_reviews_e43e766641c9ec85217bc6ce). Every finding was
reproduced by execution here before being fixed; two were refined by that
reproduction rather than taken as given.

BLOCKER 1 — the provenance gate could print OK and exit 0 without verifying.
`git show <ref>:<path> | sha256sum` hashes EMPTY STDIN when the ref does not
resolve, so at_ref was never empty and the UNKNOWN branch was dead code.
Measured: a bogus ref reported MISMATCH — accusing the snapshot of lying when
the real cause was an incomplete clone, and the operator's natural remedy for
MISMATCH is to re-run the refresh, which rewrites provenance to silence the
complaint; and with a 0-byte snapshot against a 0-byte upstream file it printed
"OK: exactly skillset@aaaaaaa" with exit 0 for a ref that does not exist. The
script already had the sha_empty idiom and had applied it to blob_sha but not
to at_ref. Existence is now PROVEN with git cat-file -e before anything is
hashed, at two levels (ref resolves / path exists at it) because those deserve
different messages. Same defect class as the canary it replaces: a check that
can succeed without checking. A second, unflagged instance of the same pipeline
shape in blob_sha was found and fixed too.

Exit codes split, because the old contract failed the sanctioned case: 0
truthful (including stale, with a NOTICE), 1 a lying record only, 2 cannot
determine. AGENTS.md step 2 promised "the message distinguishes the two" and
was the thing this branch was breaking; rewritten to state all three.

BLOCKER 3 — VENDORED.md contradicted itself in the release whose stated
invariant is non-contradiction: its hand-maintained provenance line named
skillset 670f7f1, seven commits behind the ARG and itself the commit that told
agents to hand-stamp added_by — the withdrawn instruction this work exists to
stop shipping — while its cp recipe contradicted the "not cp" rule 20 lines
above. Line removed (nothing forced it to move when the ARGs did); 670f7f1 kept
only as a labelled cautionary example. The pi-extensions half was verified
redundant (CI require_sha resolves PI_EXTENSIONS_REF) before removal.

SHOULD-FIXES: `<root> --check`, the spelling VENDORED.md documented, silently
ran a REFRESH because only $1 was parsed (both tools now parse all args and
reject unknown ones); refresh at a detached/older HEAD silently rewound ref and
bytes (now refused unless the recorded ref is an ancestor, --force to override);
upstream_dirty was computed and never used in check mode; --no-skills --json
printed human text and broke jq; --help was a hardcoded sed range this branch
had already made stale; the fingerprint hashed SKILL.md alone so a live skill
differing only in a sibling file reported "identical", and pi-extensions already
ships two files, so it is now a per-skill TREE hash with the manifest field
renamed skillset_snapshot_tree_sha256; the --no-skills smoke assertion was
negative-only and passed on a crashed binary. mktemp+mv left files 0600 — CI was
unaffected since the index records 100644, so the blast radius was local builds
only, narrower than the review inferred.

Snapshot resynced c04cd15 -> 5fd0d5c so the no-clone fallback carries the
CORRECTED coordination protocol rather than the withdrawn one; --check is now OK
with no staleness notice, and the bidirectional canary re-verified against the
new bytes. Local validation is bash -n only (shellcheck, hadolint and actionlint
are all absent in this container) — CI remains the shellcheck gate.
This commit is contained in:
2026-08-26 14:38:07 +02:00
parent 49a6534093
commit e8ddeaf89f
8 changed files with 526 additions and 91 deletions
@@ -114,15 +114,31 @@ mounted, plus a fabricated-skillset run of the reconciler).
cp <pi-extensions-pkg>/skill/SKILL.md pi-extensions/SKILL.md
cp <pi-extensions-pkg>/skill/evaluate-extension-usage.py pi-extensions/
cp <skillset>/skills/mempalace/SKILL.md mempalace/SKILL.md
Copy each snapshot **from its owner in the table above** — `pi-extensions` from
the package repo's `skill/` (since `a7f3044` co-located it there; `skillset`
also carries a copy, but it is a downstream duplicate and can lag), and
`mempalace` from `skillset`. Copying `pi-extensions` from `skillset` would
regress the snapshot to whatever that repo last mirrored.
Copy `pi-extensions` **from its owner in the table above** — the package
repo's `skill/` (since `a7f3044` co-located it there; `skillset` also carries a
copy, but it is a downstream duplicate and can lag). Copying `pi-extensions`
from `skillset` would regress the snapshot to whatever that repo last mirrored.
Snapshot provenance at last refresh: skillset `670f7f1`, pi-extensions pkg `e73cb9f`.
`mempalace` is **not** refreshed by `cp` — see the *Freshness model* section
above: `scripts/vendor-mempalace-skill.sh <skillset-root>` is the only thing
that should ever touch that snapshot, because a bare copy can update the bytes
without updating the ref that claims to describe them, which produces a
manifest that confidently lies.
Neither vendored skill has a hand-maintained "last refreshed at" line here on
purpose — one previously existed (skillset `670f7f1`, pi-extensions pkg
`e73cb9f`) and went stale within hours, because nothing forced it to move
when the ARGs did. `670f7f1` is now a cautionary example rather than a fact
worth recording: it is the commit that told agents to hand-stamp `added_by`,
which a later skillset commit (and the pi-devbox edge stamper) withdrew — so a
reader trusting that line would have been pointed at superseded guidance.
Both facts it tried to capture now live somewhere that cannot drift by hand:
| Fact | Where |
|---|---|
| which skillset commit `mempalace`'s bytes came from | `ARG SKILLSET_SNAPSHOT_REF` (Dockerfile.variant) + `skillset_snapshot_ref` in `build-manifest.json`, written *only* by `vendor-mempalace-skill.sh` |
| which pi-extensions package commit was vendored | `ARG PI_EXTENSIONS_REF` (Dockerfile.variant, CI-resolved to a 40-hex commit) → OCI label `se.jordbo.pi-devbox.pi-extensions-ref` and `build-manifest.json`'s `components.pi-extensions`, both read from the actual `/opt/pi-extensions` checkout, not from intent |
When you refresh the `mempalace` snapshot, also update the phrase asserted by
the "mempalace skill snapshot is current" smoke test — it deliberately pins the
@@ -41,6 +41,21 @@ Run these immediately when a session begins, before responding to the user:
mempalace_kg_query(entity="<project_or_person>")
```
4. **Check your mailbox.** Just run it — an empty result is a fine answer and
costs one call. Do not try to decide first whether coordination "applies to
you"; that test is what used to be wrong here (see *Cross-Machine
Coordination* below):
```
mempalace_event_list(to_agent="<harness>@<device>", status="open")
```
This is a candidate list, not a to-do list — `status` never changes after an
event is written, so finished asks keep matching. Subtract the ones you have
already answered using the rule in *What you actually owe*, below.
Another machine may have asked you something, or corrected something you are
about to rely on. This costs one call and is the only way you will find out:
nothing pushes an event into your session unless your bridge delivers it for
you, and if it does you will already have seen it before reading this.
Do NOT announce this to the user. Just do it silently to orient yourself.
### Temporal grounding — compute time deltas, don't guess
@@ -340,6 +355,158 @@ mempalace_kg_invalidate(subject="...", predicate="...", object="...", ended="<to
mempalace_kg_add(subject="...", predicate="...", object="...", valid_from="<today>")
```
## Cross-Machine Coordination — the logstream
The palace stores what you *know*. The logstream (`mempalace_event_*`,
`mempalace_artifact_*`) carries what you want to *say to another agent* —
delegation, review, patch handoff, retraction. It is the only channel on which
another machine can reach you.
**Does this apply to you at all? Do not use `mempalace_mesh_peers` to decide.**
It answers a different question than it appears to. A shared palace can be
*hub-and-spoke* — many machines as thin clients of one central replica — and
then `mesh_peers` reports `peers: []` because there are no peer *replicas*,
even while four machines are actively writing to the same log. Measured on this
fleet: `peers: []`, one replica authoring every event from every machine. An
earlier version of this section told you to read `mesh_peers` and skip the
mailbox when it came back empty, which disabled the mailbox on precisely the
fleet it was written for.
The honest discriminators, cheapest first: **just run the mailbox query** (empty
is a fine answer); check whether `MEMPALACE_REMOTE_URL` is set, which is what
actually selects a shared palace; or look for any event whose `from_agent` is
not you. On a solitary palace the event tools still work — you are writing to
yourself and your mailbox stays empty. That is not a fault to debug.
**It is a durable log, not a bus — nobody is "listening".** Events are appended
and persist; there is no subscription, no delivery window, and nothing is lost
by being offline when one is written. A message waits indefinitely for you, and
your reply waits just as patiently for a sender who has since gone away. Machines
in a fleet are rarely awake at the same time, which is exactly why this is a log
and not a chat.
**Agent name is the only identity the log has.** Depending on deployment, every
client may share one `origin_replica` — on the fleet this skill was written for,
all machines are thin MCP clients of a single central replica, so `origin_replica`
is identical for every event and cannot tell two machines apart. `from_agent` /
`to_agent` carry the whole distinction, which is why the `<harness>@<device>`
stamping in *Provenance is stamped for you* is load-bearing here and not mere
tidiness.
### Reading your mailbox
```
mempalace_event_list(to_agent="<harness>@<device>", status="open")
```
- `to_agent=<you>` **also matches `*` broadcasts**, so one call covers both. No
second query needed.
- `status="open"` narrows the mailbox to what a sender *said was an ask at the
time of writing* — that is all it can do. It is a good first filter (on a real
stream it cut 5 events to 2), but it is **not** a list of what you owe, and it
never shrinks as you work. Treating it as owed-ness is the mistake this
section previously made: an earlier draft cited "5 unfiltered, exactly 1
filtered — the one that needed a reply" as proof the filter tracked
obligation. It did not. That single result was an event which had *already
been acked* half an hour earlier; the filter looked decisive only because the
stream happened to contain one directed `open` event. **Unfiltered mailboxes
train you to ignore them — and so does a filter that keeps showing you
finished work.**
- To resume where you left off, use `since_event_id`, **never**
`since_created_at`. A timestamp cursor permanently skips an event that synced
in late — it is a time window ("what happened today"), not a cursor.
- Read `metadata` before acting: senders put the load-bearing specifics there
(which host verified what, which run failed, what a change retracts).
### The ack contract — the sender declares whether a reply is owed
An obligation you never agreed to is noise, so the sender states it:
| Sender writes | Means | Recipient owes |
|---|---|---|
| `to_agent="<specific agent>"` + `status="open"` | an ask | an ack or a reply (the event itself keeps matching forever — see below) |
| `to_agent="*"` (any status) | broadcast FYI | nothing |
| any other status (`ready`, `applied`, `blocked`, …) | a statement of fact | nothing |
Ack with `mempalace_event_ack(event_id=…, from_agent="<you>", status=…)`. It
**appends a new event** and never mutates the original; the correlation id is
copied for you, and `metadata.ack_of` is set to the event you answered.
#### What you actually owe — derive it, do not read it off `status`
The log is append-only and `status` is written **once**, so it is an honest
statement about an item *at the moment it was written* and nothing more. It is
not mutable state, and asking it to carry mutable state is what breaks:
acking appends a new event and changes nothing about the old one, so **a
directed `open` event matches your mailbox query forever, answered or not.**
Nothing is ever "dismissed" — which also means a deferred ask cannot be
accidentally lost, only that you must compute what is outstanding:
```
candidates = mempalace_event_list(to_agent="<you>", status="open")
mine = mempalace_event_list(from_agent="<you>")
```
A candidate is **answered** when one of your own events
1. has a **higher `seq`** than the candidate, and
2. joins to it — `metadata.ack_of == candidate.id` (exact, written for you by
`event_ack`) or the same `correlation_id` (the fallback), and
3. carries a **terminal** status: `applied`, `superseded`, `failed`, `blocked`.
Everything else is still owed. Two calls, constant cost.
**Compare `seq`, never `created_at`** — the same reason you resume with
`since_event_id`. Without the ordering test, one terminal reply would suppress
every later ask on the same `correlation_id` for good; verified on a live thread
where a `ready` reply at `seq` 16 sits *before* the request at `seq` 17 that it
obviously cannot have answered.
This also supplies the "taken, not finished" state that looked missing:
`claimed` and `ready` are deliberately **not** terminal, so work you have picked
up keeps resurfacing until you close it out. No extra convention, no new field.
Two consequences worth internalising:
- **"Seen, not doing it" is a legitimate ack** — `status="blocked"` or
`"superseded"` plus the reason. Silence is not, and it is not merely rude:
with no terminal event of yours to join to, the ask stays in the owed set
indefinitely and there is nothing anyone can do about it from the other end.
- **Nothing expires, and it should not.** An `open` with no terminal reply is
still live by definition, and the finished threads are valuable history. If
content is genuinely perishable ("do not push to main for the next hour"), say
so in `metadata.expires_at` — metadata is stored verbatim — and honour it as a
hint when reading. An old `open` that the derivation still counts as owed is a
signal, not garbage: it means somebody asked and nobody answered.
### Writing to another machine
- **Address the stamped name you actually saw** in a `from_agent` field, e.g.
`pi@tor-ms22`. A bare `pi` reaches nobody's mailbox once stamping is live, and
older events in the log still carry bare names — do not copy them.
- **Use `status="open"` only when you truly need an answer.** It places an
obligation on another machine.
- **Never broadcast an ask.** `to_agent="*"` + `status="open"` obliges everyone
and therefore no one.
- **Always set a `correlation_id` on a directed `open`,** and reply with the
same one. It is not just for reconstructing a conversation later: it is the
join the owed-set derivation depends on. An uncorrelated ask can only ever be
closed by an `event_ack` (which sets `ack_of` for you) — a plain reply cannot
be matched to it at all.
- **Corrections are new events, never edits.** Say explicitly what you retract
and name the id — drawer or event — that carried the withdrawn claim.
- **Put a retraction where the reader will look.** An event reaches a live agent;
a *drawer* is what a future semantic search finds. If you filed advice as a
drawer and later withdraw it, file the withdrawal as a drawer too — otherwise
the next agent finds your original confident advice and no trace of the
correction. (This is a real incident, not a hypothetical.)
- **Hand over exact content as an artifact**, not prose: `mempalace_artifact_put`
or `mempalace_patch_submit` store bytes with a sha256, and the event references
the id. Never paste a diff into a body and hope it survives.
- **Waiting on a specific reply?** `mempalace_event_wait` blocks with backoff —
do not poll `event_list` in a loop. A timeout there is a normal result, not an
error.
## Palace Structure
### Wings
@@ -413,4 +580,6 @@ Entity-relationship triples with temporal validity. Query with `mempalace_kg_que
- **Don't mine .git directories or node_modules.** The CLI miner respects .gitignore by default.
- **Don't create duplicate drawers.** Use `mempalace_check_duplicate` before adding manually.
- **Don't treat the palace as a task list.** It's for knowledge and context, not todos.
- **Don't broadcast an ask, and don't leave one unanswered.** On a shared palace, `to_agent="*"` + `status="open"` obliges every machine and therefore none of them. And don't expect acking to tidy your mailbox: `status` is immutable, so the event keeps matching either way — what a terminal reply buys you is that the *derived* owed set (see *What you actually owe*) stops counting it. Leave asks unanswered and that set only grows, until everyone learns to stop looking. "Seen, not doing it" is a complete answer — silence is not.
- **Don't assume you would have heard.** Nothing pushes another machine's message into your session. If you did not run the mailbox query at wake-up, a correction addressed to you by name can sit unread while you confidently rebuild the thing it warned you about.
- **Don't invent provenance metadata, and don't hand-stamp it either.** An earlier version of this list told you to set `added_by="<harness>@<device>"` by hand; that instruction has been withdrawn, because RFC 001 §7.3.2 places provenance at the client/server boundary and the pi bridge now does it uniformly (see *Provenance is stamped for you* above) — but the withdrawal only holds where the bridge is live, so run the one-line check in that bullet first; on an older image hand-stamping is still the only signal a hand-filed drawer gets. DO NOT invent values for the palace's own metadata fields (`device`, `agent_kind`, `origin_device`): those are stamped by infrastructure that also records *how* each was determined, and a fabricated value is worse than none because it silently corrupts a future merge. DO pass `source_drawer_id` on `kg_add`. And never put a machine name in a diary's `agent_name` — it becomes the wing name and hides your entries from `diary_read`.