Compare commits

...

15 Commits

Author SHA1 Message Date
joakimp 36e65fe657 skills: add credential-incident-response, and assert it stays baked
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 18s
Carries the facts a two-day credential incident produced, not the discipline:
probe the issuer FIRST (11 of 13 "exposed" credentials were already dead at the
provider, which cost five HTTP requests to learn and was never checked), the
403-vs-401 trap that scoped tokens introduce into liveness probes, revocation
beats deletion for anything already replicated, the three places a secret hides
in a Chroma palace (FTS content, metadata, raw bytes) in coverage order, scope
derivation from measured consumers, and this fleet's age store with its
single-recipient weakness.

Facts transfer between sessions; exhortations do not — hence a separate skill
for the domain knowledge and a one-line pointer in the always-loaded block.

Authored here, so baked is canonical and it is NOT added to skillset-owned.txt.
Skill dirs are picked up by a glob in entrypoint-user.sh, so no registration is
needed — verified rather than assumed, since an enumerated list would have left
the skill inert, a fitting failure given its subject. Three smoke assertions
extended so a future rebuild cannot silently drop it.
2026-08-30 00:50:11 +02:00
joakimp f0ebea2d98 skills: fix the half of the negative-result rule that was wrong
pi-devbox-environment already warned that "a negative result is usually your
own filter" — baked, symlinked in at every container start, authored by an
earlier session. It survived every recreate, was available all of a later
session, and was violated five times anyway. So the gap was never persistence.

That section also closed with "a positive result needs no such scepticism — it
carries its own evidence." That is false, and it aimed the scepticism budget
one way only. Three of those five errors were positives:

  - an SSH handshake SUCCEEDED and greeted me as joakimp while I believed I was
    probing gitea.egl.lan — `Host gitea*` had rewritten HostName
  - a 401 that was a genuine answer from an issuer which never minted the token
  - a "regression" produced by diffing against a value my own -p 2222 flag set

Adds the three missing false-negative rows, replaces the false claim with the
"a positive result only proves what you actually asked" subsection, and records
the two habits that actually caught these: `ssh -G` to learn which rule
captured a hostname, and declaring the expected result before running a check.

The cross-cutting form goes in pi-global-AGENTS.append.md rather than in a
skill, because it has to fire without a task description matching it — being
loadable on demand is exactly what failed. Across all five errors, none was
caught by re-reading my reasoning; every one was caught by a second measurement
that disagreed.
2026-08-30 00:50:11 +02:00
joakimp b615571913 changelog: release v1.8.11 — the mail nobody could deliver, and the ship that skipped a file
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 19s
Publish Docker Image / base-decide (push) Successful in 10s
Publish Docker Image / build-base (push) Successful in 53m20s
Publish Docker Image / smoke (push) Successful in 4m48s
Publish Docker Image / smoke-studio (push) Successful in 4m58s
Publish Docker Image / build-variant-studio (push) Successful in 17m0s
Publish Docker Image / build-variant (push) Successful in 26m43s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 24s
Names three things riding in via floating refs, since the rule this
fleet adopted (v1.8.9) is to name a behaviour change before tagging,
not after:

- mempalace-toolkit b2b50af -> 21023e7: the rsync --checksum ship fix
  (a real behaviour change to every remote-mode client) and RFC 003
  §7.13 (docs only).
- skillset 6eb20af -> a12fe5e: mermaid-diagrams cutU normalisation and
  the mempalace skill's from_agent identity rule, plus the vendored
  fallback snapshot re-pinned to match (this pi-devbox commit).

Dependency audit run by direct command against every component, not
assumed: pi/mempalace/pi-atelier/pi-studio/pi-toolkit/pi-extensions/
pi-fork/pi-observational-memory all confirmed unchanged since v1.8.10.
pi-studio's apparent SHA drift on ls-remote turned out to be an
annotated-tag-object hash vs its underlying commit hash, not real
drift -- checked with ^{commit} before writing it down as 'none'.
2026-08-27 23:39:39 +02:00
joakimp 495b7e3859 vendor: resync mempalace skill snapshot to skillset a12fe5e
scripts/vendor-mempalace-skill.sh, real refresh not --check: skillset
moved 6eb20af -> a12fe5e (mermaid-diagrams cutU normalisation, and the
from_agent identity rule this same release ships in RFC 003). --check
reported stale-but-truthful (exit 0, the sanctioned skip) but this
release's point is getting today's fixes live fleet-wide, and the base
rebuild is already forced by the entrypoint change and the floating
mempalace-toolkit ref moving -- so the incremental cost of also
bumping this pin is zero. Phrase canary in scripts/smoke-test.sh
unaffected: neither pinned phrase ('Provenance is stamped for you'
present, 'Attribute what you file yourself' absent) is in the section
that changed; both verified still correct in the new snapshot.
2026-08-27 23:36:03 +02:00
joakimp 45850bc973 entrypoint: put back the shell state a recreate eats
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 25s
cli_utils' install.sh reaches PATH by symlinking bin/ into ~/.local/bin.
That is persistent on a host and ephemeral in a container, so the same
installer produced opposite durability and every --force-recreate sent the
human back to typing /workspace/cli_utils/bin/git-status-all. Re-link at
start instead, and add a per-device boot hook so the next question of this
shape needs no image change at all.

Symlinks rather than a PATH edit in an rc file, deliberately: ~/.local/bin
is already ahead of /usr/local/bin in ENV PATH, so links resolve in
NON-interactive shells too (docker exec, agent tool shells, scripts). An
rc-file PATH edit cannot reach those because ~/.bashrc returns early when
not interactive — measured, that asymmetry is exactly why `command -v
git-status-all` failed in one shell and worked in another on the same box.
Shell FUNCTIONS remain the sourced file's job; a symlink cannot carry them.

Guards, because ~/.local/bin is shared: a real file is never clobbered, a
symlink pointing elsewhere is never stolen, ours are refreshed, and links
into a cli_utils/bin whose target vanished are pruned — a dangling link on
PATH reads as a broken container rather than a removed script.

The hook (~/.config/devbox-shell/init.sh) adds no trust boundary: that dir
is already sourced into every interactive shell by the baked bash_aliases,
so it is already arbitrary code from the same owner. Only WHEN it runs is
new. bash <file>, never sourced, exit status ignored, output to a log.

Caught before commit, and the reason the loops use `if` bodies instead of
`&&` chains: under `set -euo pipefail` a for-loop whose last command is a
false test exits non-zero, and with no match the /workspace/*/cli_utils
glob stays literal — so the first draft would have failed to START a
container on every machine that does not have this repo, rather than merely
skipping the links. Re-tested with set -e in place: no-cli_utils/empty-HOME
no-op, guards, idempotence, CLI_UTILS_LINK=0, a hook that exits 7, and a
hook that tries to mutate CLI_UTILS_BIN — all exit 0 with intact state.

Not covered by CI: docker-publish.yml runs only on tags and lint.yml lints
workflow run: steps, so neither executes this file. Validated by extracting
both sections and running them against fixtures, then for real in a live
v1.8.10 container. No smoke assertion added on purpose — the positive path
needs a /workspace mount smoke does not have, and asserting it there would
repeat the v1.8.0 mistake of a smoke check written against a stage that
does not exist at run time.

Moves the base hash (base-decide folds `cat entrypoint.sh
entrypoint-user.sh`), so this rides along with the next tag's ~40-minute
base rebuild rather than justifying a tag of its own.
2026-08-27 21:52:51 +02:00
joakimp 6891dc32b8 changelog: release v1.8.10 — deploy the scrubber, and name the check that must not be skipped
Publish Docker Image / build-variant-studio (push) Successful in 17m4s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 14s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-variant (push) Successful in 18m45s
Publish Docker Image / build-base (push) Successful in 41m59s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / smoke (push) Successful in 4m55s
Lint / hadolint (push) Successful in 8s
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Audited every component against upstream rather than assuming. Only two moved
since v1.8.9:

- mempalace-toolkit 5b8d78f -> b2b50af (ours): feeder scrubber, the symlink fix
  that stops it refusing to stage, hlc owed-set join, queued-delivery note,
  explicit MEMPALACE_MAILBOX_NOTIFY protocol modes, RFC 003 + fleet-memory docs.
- pi-studio v0.9.48 -> v0.9.52 (upstream, studio variant only): 22 commits, all
  additive — PDF previews, hideable header, contextual side questions. No removals
  or renames. Its pi floor is >=0.84.3 against our exact 0.84.3 pin: satisfied,
  zero headroom, named as a watch item because the next floor bump breaks the
  studio job only, after core has already published.

Verified unchanged, with dates proving they predate the v1.8.9 bake: pi 0.84.3 (=
npm latest), mempalace 3.8.0 (= PyPI latest), pi-atelier v0.8.2, playwright 1.62.1
(no drift this cycle despite floating on latest), pi-fork bf702b4, pi-obsmem
ce9fc98, pi-toolkit 0e1369e, pi-extensions 2022887. No breaking changes anywhere,
so everything can ship in one tag.

The reason for tagging now is not features. The scrubber has been live on exactly
ONE device since this morning; every other device has kept staging unscrubbed
transcripts into a shared palace, and cleanup after the fact is manual redaction
(done twice today, with freed-page residue left behind by choice).

The entry leads with the acceptance check instead of burying it, because this
release's worst failure is silent: the feeder is fail-closed, so a packaging or
path mistake stops the fleet's memory feed and nothing complains — refusing to
stage looks exactly like a quiet session. f0bffd1 exists because that very bug was
real (BASH_SOURCE reports the symlink path, not the target). First client on the
new image must confirm a [scrub] summary line appears, the drawer count moves, and
exit 3 did not fire. Silence is the failure signal, not success.
2026-08-27 17:51:46 +02:00
Joakim Persson 8a673ec143 docs: unclip the diagrams, and answer what compaction leaves behind
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 15s
Two problems reported against docs/observational-memory.md, one cosmetic and one
substantive. Both turned out to be worth more than the fix.

Clipping. Several boxes lost their bottom line of text in the viewer, and neither
the source nor my own render showed it. Cause: Mermaid measures a node label with
its own font metrics, commits to a box size, then renders that label as real HTML
in a <foreignObject> — so any host stylesheet touching line-height or font-size
inflates the text past a box that is already fixed, and the overflow is clipped.
The error accumulates per line, which is why it always eats the last line of the
tallest labels.

Rejected the obvious fix after testing it rather than assuming it:
%%{init: {'flowchart': {'htmlLabels': false}}}%% *is* honoured (labels switch from
16 foreignObject to 7 tspan) and clips identically, because inflated font-size
inherits into SVG text too. The fix that works is a hard limit of two short lines
per node, with the detail moved into prose under each diagram — one- and two-line
boxes have the vertical slack to absorb inflation, three- and four-line boxes do
not. The nodes were carrying paragraph-sized text; the diagrams are better for
losing it.

Regression harness: render every block with line-height: 1.7 !important forced
onto the label HTML, screenshot, read it. That caught two survivors of the rewrite
that looked fine in the clean render — a long unbreakable /opt path wrapping to a
third line, and a cylinder shape whose curved bottom leaves less room than a
rectangle for the same two lines.

New §4, because the document explained that compaction folds the ledger and never
said what that leaves in the context. Asked directly: is the session back to
knowing nothing? No. A verbatim tail survives, sized by keepRecentTokens (20k) and
cut only at turn boundaries; the system prompt and AGENTS.md were never in the
compacted region because they are rebuilt from disk each request; nothing is
deleted from disk, since compaction appends a compaction entry rather than
rewriting lines; and recall keeps resolving ids whose sources left the context
because it reads the full branch via sessionManager.getBranch() and never consults
the context window. Repeated compaction renders from live records, not from the
previous summary's prose, so there is no generation-loss spiral.

Also corrects this repo's own "compaction calls no model" to the steady-state
claim it actually is: with an empty ledger the hook returns nothing and explicitly
declines ownership, and pi's native model-based summariser runs. Snippet quoted in
the doc.

And a bug shipped in the first version: §9 said the ledger entries are
custom_message and specifically not custom. Exactly backwards, so the one grep
that section existed to get right was the one it got wrong. Verified empirically
against the live session file — 11 om.observations.recorded and 6
om.reflections.recorded, all "type":"custom", next to "type":"custom_message"
entries whose customType is mempalace-mailbox and mempalace-wakeup. That is where
the confusion came from, and the distinction is load-bearing rather than
cosmetic: the mailbox uses the context-visible append API, om's ledger uses the
invisible one. Which makes it a feature the doc now advertises — the ledger costs
zero context until it is folded.
2026-08-27 14:24:19 +02:00
joakimp cdb6fc0950 changelog: name what the floating toolkit ref will pull into the next tag
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 16s
MEMPALACE_TOOLKIT_REF=main floats and docker-publish.yml resolves it to a SHA at
build time, so toolkit main moving 5b8d78f -> f0bffd1 (10 commits) ships in the
next tagged image whether or not this repo has a commit. v1.8.9 adopted the rule
after that shape bit twice (553d865, 5b8d78f): name the behaviour change BEFORE
tagging. This is that rule obeyed rather than re-learned — the work was pushed to
toolkit main earlier today and this entry was missing, which is exactly the gap
that caused a cross-host misattribution in v1.8.7.

Contents: the feeder-side secret scrubber (three tiers, T3 report-only after a
measured 403 -> 29 false-positive calibration on 52 MB of real transcripts, fail
closed); the symlink near-miss that fail-closed would have turned into a
fleet-wide silent memory outage at bake time, caught before tagging; the mailbox
work (hlc owed-set join, queued-until-next-turn note, explicit notify protocol
modes, tmux path documented unverified); and the documentation set (RFC 003,
fleet-memory.md, secret-hygiene.md). Also states what remains unscrubbed: the
opencode bridge write path and the unbuilt server-side layer.

No tag pushed — per the release protocol, no tag means no build.
2026-08-27 14:20:35 +02:00
Joakim Persson 14371e2da6 docs: explain the memory that runs itself, and fix a claim the mailbox falsified
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 16s
`pi-observational-memory` is baked, registered in the seeded settings.json, and
handed a cheaper model than the session it serves — and the only description in
this repo was five words in a feature list. Someone meeting `/om:status` or a
"compacted memory" block had nothing to read that said whether to leave any of it
on. New docs/observational-memory.md, 272 lines and five diagrams.

Scoped by who is authoritative, so there is one copy of each claim:

- Upstream (/opt/pi-observational-memory/docs/) already documents the mechanism
  well — concepts.md, how-it-works.md, configuration.md, including a v3 lifecycle
  diagram that matches the deployed code. Linked, not re-derived.
- This document takes the four facts pi-devbox owns and can change: the pinned
  commit it bakes (v3.0.4 ce9fc98, matching build-manifest.json), the packages[]
  entry that decides which copy loads, the Haiku-workers-vs-Opus-session split,
  and the devbox-pi-config volume that makes the ledger outlive the container.
- Plus the confusion this image creates by shipping two things called memory: a
  section contrasting it with MemPalace, on the line "observational memory keeps
  a session coherent, the palace keeps the fleet coherent".

Placement follows the split fleet-ops states for itself — reusable mechanism is
not deployment data — so a "why is this in my container" document belongs in the
repo that pins and wires the component. Linked twice from the README, because
until this commit the README referenced docs/ zero times and the file already
sitting there was reachable only by listing the directory.

Numbers were read out of the live container and the baked tree rather than out of
release notes, which caught one thing the pi-extensions skill still has wrong:
the dropper is gated on a successful same-turn reflection, not on a token
threshold of its own.

Also fixes a README sentence that v1.8.9 made false. § Cross-machine agent
coordination ended with "Nothing in this image polls the log on the agent's
behalf"; the mailbox has shipped since aac4a1c. Replaced with the three knobs and
their defaults, derived-not-read owed-ness, and the queued-into-the-next-turn
delivery measured on two devices — and dated to "as baked in v1.8.9
(mempalace-toolkit 5b8d78f)", pointing at RFC 003 §7.11–§7.12 for the mechanism,
because toolkit main is already ahead (a92c75d pings the human who is not
looking) and describing that here would trade a stale-behind claim for a
stale-ahead one.

The cause is worth more than the fix: the behaviour arrived through the floating
MEMPALACE_TOOLKIT_REF, so no diff in this repo ever touched the paragraph making
the claim. v1.8.9's own rule fired for the CHANGELOG and nobody swept the README.
The CHANGELOG records what changed, the README asserts what is true, and only the
first is reviewed at release time — so the rule now extends to grepping the README
for absolute claims (nothing, never, does not, only) about a component whose SHA
moved.

Diagrams were verified by rendering, not by parsing. Both comparison diagrams
parsed clean and rendered with their meaning reversed: Mermaid laid the second
declared subgraph out first, putting "with observational memory" before "without"
and MemPalace before observational memory in the diagram whose entire job was
that contrast. Rebuilt as declaration-ordered chains and re-rendered at mermaid@11
— the version pi-studio pins — in the baked headless browser, then read back as an
image. mermaid.parse() proves syntax and says nothing about layout.
2026-08-27 13:55:55 +02:00
joakimp aac4a1c323 release: v1.8.9 — the version flag that blamed the wrong component
Lint / hadolint (push) Successful in 15s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 41m49s
Publish Docker Image / smoke (push) Successful in 4m51s
Publish Docker Image / smoke-studio (push) Successful in 4m59s
Publish Docker Image / build-variant-studio (push) Successful in 16m58s
Publish Docker Image / build-variant (push) Successful in 28m35s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 17s
Two versions, two flags. `--expected-version` has only ever asserted
`pi --version`, but AGENTS.md step 4 spelled it `X.Y.Z` inside a checklist where
every other X.Y.Z is the pi-devbox tag. Run as documented for v1.8.8 the final
runtime gate of the release printed

    ✗ pi version mismatch: expected 1.8.8, got 0.84.3

and exited 1 — a red accusing the image of being the wrong version. Not one
reader's slip: the v1.8.8 release-readiness handoff from pi@emb-7kj4vr4g
propagated the same wrong spelling twice while correctly calling step 4 "not
ceremonial", so two independent readers converged on it. README.md had it right
all along, which means the two documents disagreed.

- new --expected-image-version asserts the pi-devbox release tag, read from
  release_tag in /etc/pi-devbox/build-manifest.json (no checkout, no network);
  leading `v` optional on either side
- both flags detect being handed the other one's value, and the test is exact
  rather than heuristic: the value is compared against the other quantity the
  image itself reports, so it can only fire on a real mix-up
- neither flag is required now. With none, live `pi --version` is asserted
  against the manifest's pi_version — not a tautology, since a stale pi in the
  ~/.pi/npm-global volume can shadow the baked one, exactly as a stale
  npm:pi-atelier can in packages[]
- the header note replaced was stale and load-bearing: it claimed pi is resolved
  from 'latest' and cannot be self-derived, while Dockerfile.variant pins
  ARG PI_VERSION=0.84.3 and docker-publish.yml reads that ARG as its source of
  truth. The same withdrawn claim also sat in cli_utils' pi-devbox-sanity --help
- argument parsing: a missing value, or a value that is another flag, is a usage
  error instead of silently consuming the next argument; --help works

All fourteen flag combinations exercised by execution, including the two
manifest-absent branches and the shadowing branch a healthy container cannot
reach — mutation-tested with a doctored manifest so each failure branch was
observed firing rather than assumed present.

CHANGELOG also names what no commit here causes: mempalace-toolkit main moved
e70bef2 -> 5b8d78f, so this tag ships the auto-delivered logstream mailbox
because base_tag folds the resolved toolkit SHA. It would have landed either
way; going unnamed is the 553d865 shape that already caused one cross-host
misattribution. Component audit found nothing else to bump — pi, mempalace,
pi-atelier all equal their upstream latest, and every other floating ref
resolves to the commit already baked.
2026-08-26 18:47:03 +02:00
joakimp 34cf1e3810 release: v1.8.8, and a notice that named the wrong remedy
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 19s
Publish Docker Image / build-base (push) Successful in 1h3m31s
Publish Docker Image / smoke (push) Successful in 4m47s
Publish Docker Image / smoke-studio (push) Successful in 8m8s
Publish Docker Image / build-variant-studio (push) Successful in 20m4s
Publish Docker Image / build-variant (push) Successful in 26m9s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / update-description (push) Successful in 14s
Freezes the v1.8.8 section and clears the two non-code checklist items
pi@emb-7kj4vr4g handed over (evt_20260826T134919_a614ecfc2d4f), plus the two
carried nits from its round-2 verification (evt_20260826T133356_d56792791a49).
Every claim below was re-measured here rather than taken from the handoff.

THE STALENESS NOTICE ASSERTED A DIRECTION IT NEVER TESTED — Blocker 1's shape,
one layer down, in the message I added to replace the message that named the
wrong cause. The notice fired on "recorded != HEAD" and then announced HEAD as
the newer side without testing ancestry, so a clone that was merely BEHIND got
"has moved to 82a8d3c; the snapshot describes the older 5fd0d5c" when 82a8d3c is
5fd0d5c's ANCESTOR. Found by EMB against the real state of its own host, not a
fabrication. The verdict was never wrong (rc 0, nothing mis-verified) but the
remedy it implies is a ~67-minute base rebuild when the actual fix is `git pull`
— the only one of the two carried nits with a price tag, which is why it went
first. Now tests ancestry with the merge-base --is-ancestor primitive the refresh
path 60 lines below already used, and reports three verdicts: stale (refresh),
clone behind (pull, do NOT refresh), diverged (reconcile). All three verified by
execution; only the first was correct before. --help no longer errors on the one
script whose argument order was itself a landmine. The `-s "$VENDORED"` guard is
now commented as load-bearing: it makes the empty-stdin collision unreachable by
construction, which also means no test below exercises it any more, so deleting
it as "redundant with the probes" would silently restore the false OK.

CHANGELOG: retitled, and three stale spots fixed in what becomes the permanent
record. It cited the skill at 82a8d3c (twice superseded); it RE-ASSERTED the
retracted mailbox measurement as live evidence 200 lines after withdrawing it,
which is the exact non-contradiction failure this release exists to fix; and its
warning block still described the pre-e8ddeaf world ("still records c04cd15",
"now exits 1") while quoting as exemplary the very notice whose direction was
unverified. Per EMB's steer the conclusion was kept and only the evidence
replaced: the status filter does drop broadcast noise, it just never computed
owed-ness. Honest replacement, measured on both machines: raw filter returns 2
here and 1 there, EVERY ONE already answered, derivation returns 0 for both.

SNAPSHOT RESYNCED AGAIN, 5fd0d5c -> 6eb20af, because skillset 6eb20af adds the
limit of my own seq ordering test: seq is REPLICA-LOCAL, equal to origin_seq only
because one replica authors for all four machines, so use hlc once mesh_peers
reports a peer. Recorded as reasoning not measurement — a second replica cannot
be stood up here. The durable half is the asymmetry: seq skew makes an ANSWERED
item resurface (noise, visible, self-correcting) while created_at SUPPRESSES AN
UNANSWERED ask forever (silent, permanent), so the skill now says outright that
"fixing" a resurfacing item with a timestamp trades the safe failure for the
dangerous one. The resync was free: rootfs/ was already changing, so the base
rebuild was forced regardless — the ordering warning about accidental staleness
does not apply to a deliberate refresh before the tag.

--check is a clean OK at 6eb20af with no notice, canary re-verified bidirectionally
(present 3, withdrawn 0), baked snapshot 0644, tree hash recomputed at build time
and re-verified in-container. bash -n clean; shellcheck/hadolint/actionlint remain
absent locally, so CI is still the only evidence for those.
2026-08-26 15:56:03 +02:00
joakimp e8ddeaf89f skills: a gate that could pass without checking, and a mailbox that never empties
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 1m39s
Fixes the three blockers and seven should-fixes from pi@emb-7kj4vr4g's review
(logstream correlation skills-provenance-review, full text in
drawer_pi-devbox_reviews_e43e766641c9ec85217bc6ce). Every finding was
reproduced by execution here before being fixed; two were refined by that
reproduction rather than taken as given.

BLOCKER 1 — the provenance gate could print OK and exit 0 without verifying.
`git show <ref>:<path> | sha256sum` hashes EMPTY STDIN when the ref does not
resolve, so at_ref was never empty and the UNKNOWN branch was dead code.
Measured: a bogus ref reported MISMATCH — accusing the snapshot of lying when
the real cause was an incomplete clone, and the operator's natural remedy for
MISMATCH is to re-run the refresh, which rewrites provenance to silence the
complaint; and with a 0-byte snapshot against a 0-byte upstream file it printed
"OK: exactly skillset@aaaaaaa" with exit 0 for a ref that does not exist. The
script already had the sha_empty idiom and had applied it to blob_sha but not
to at_ref. Existence is now PROVEN with git cat-file -e before anything is
hashed, at two levels (ref resolves / path exists at it) because those deserve
different messages. Same defect class as the canary it replaces: a check that
can succeed without checking. A second, unflagged instance of the same pipeline
shape in blob_sha was found and fixed too.

Exit codes split, because the old contract failed the sanctioned case: 0
truthful (including stale, with a NOTICE), 1 a lying record only, 2 cannot
determine. AGENTS.md step 2 promised "the message distinguishes the two" and
was the thing this branch was breaking; rewritten to state all three.

BLOCKER 3 — VENDORED.md contradicted itself in the release whose stated
invariant is non-contradiction: its hand-maintained provenance line named
skillset 670f7f1, seven commits behind the ARG and itself the commit that told
agents to hand-stamp added_by — the withdrawn instruction this work exists to
stop shipping — while its cp recipe contradicted the "not cp" rule 20 lines
above. Line removed (nothing forced it to move when the ARGs did); 670f7f1 kept
only as a labelled cautionary example. The pi-extensions half was verified
redundant (CI require_sha resolves PI_EXTENSIONS_REF) before removal.

SHOULD-FIXES: `<root> --check`, the spelling VENDORED.md documented, silently
ran a REFRESH because only $1 was parsed (both tools now parse all args and
reject unknown ones); refresh at a detached/older HEAD silently rewound ref and
bytes (now refused unless the recorded ref is an ancestor, --force to override);
upstream_dirty was computed and never used in check mode; --no-skills --json
printed human text and broke jq; --help was a hardcoded sed range this branch
had already made stale; the fingerprint hashed SKILL.md alone so a live skill
differing only in a sibling file reported "identical", and pi-extensions already
ships two files, so it is now a per-skill TREE hash with the manifest field
renamed skillset_snapshot_tree_sha256; the --no-skills smoke assertion was
negative-only and passed on a crashed binary. mktemp+mv left files 0600 — CI was
unaffected since the index records 100644, so the blast radius was local builds
only, narrower than the review inferred.

Snapshot resynced c04cd15 -> 5fd0d5c so the no-clone fallback carries the
CORRECTED coordination protocol rather than the withdrawn one; --check is now OK
with no staleness notice, and the bidirectional canary re-verified against the
new bytes. Local validation is bash -n only (shellcheck, hadolint and actionlint
are all absent in this container) — CI remains the shellcheck gate.
2026-08-26 14:38:07 +02:00
joakimp 49a6534093 docs: write down the coordination channel the fleet already runs on
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 15s
The logstream has carried cross-machine work since 2026-08-18 — patch handoff,
review, a v1->v2 supersede — and nothing in this repo said it existed. That gap
had a measurable cost this morning: another host addressed a retraction to
pi@tor-ms22 by name and it was read only because the human said "read the
logstream", while the agent was actively rebuilding the thing it warned about.

Split by what each document is authoritative for, so there is one copy of each
claim rather than three that drift:

- README § Cross-machine agent coordination — what the CONTAINER needs.
  MEMPALACE_REMOTE_URL selects the shared palace; MEMPALACE_PI_DEVICE is what
  makes this machine reachable, because where every host is a thin client of one
  palace the stamped agent name is the only thing that distinguishes them. Stated
  as a rule with teeth: set both or neither, since a container missing the device
  var can read the log but is addressable by nobody.
- AGENTS.md release checklist step 2 — the vendored-snapshot refresh, as a
  MECHANISM in the document a releasing agent actually reads, not a comment
  hoping to be noticed. It says the refresh costs a base rebuild, that skipping
  it is legitimate (every enrolled host reads its live clone), and that skipping
  it silently is not.
- CHANGELOG — the three-way split itself, plus the measurement that shaped the
  ack contract: unfiltered, the mailbox returned 5 events, 4 of them finished
  broadcasts from eight days earlier; with status="open", exactly the 1 that
  needed an answer.

Norms live in the skillset skill (82a8d3c, already live on every host that mounts
the skillset — no rebuild) and mechanism in mempalace-toolkit's
extensions/pi/README.md (e70bef2, which also documents the edge stamper that
553d8657 shipped undocumented). Deliberately NOT duplicated here.

Consequence recorded rather than hidden: the skill edit lands in the skillset, so
this repo's SKILLSET_SNAPSHOT_REF now honestly reports itself behind, and
--check exits 1 with "has moved to 82a8d3c; the snapshot describes the older
c04cd15". That message is also fixed in this commit — it previously blamed "the
working tree" even when the tree was clean and only the ref had moved, which is
the same defect class as a canary pinned to a phrase the release deleted: a
message that names the wrong cause. Now distinguishes moved-HEAD from dirty-tree,
verified against both plus the in-sync case.
2026-08-26 12:51:26 +02:00
joakimp e070e0bcbf skills: record the vendored snapshot's provenance, and report which copy wins
Lint / actionlint (push) Successful in 16s
Lint / hadolint (push) Successful in 16s
Found while verifying v1.8.7 from inside a fresh container: the baked mempalace
snapshot is read by no host on this fleet. devbox-skill-reconcile repoints
~/.agents/skills/mempalace at the mounted live clone (the v1.8.5 fix working as
designed), and all four compose stacks mount a workspace containing the
skillset. So the phrase canary that blocked v1.8.7's first tag polices a file
nobody opens, while the drift that could actually mislead an agent — a git pull
nobody ran in /workspace/skillset — was invisible from inside the container and
is invisible to CI by construction.

Record provenance instead of policing it, and move the check to where the
skillset actually is:

- Dockerfile.variant: ARG SKILLSET_SNAPSHOT_REF (the claim) + a sha256 of the
  shipped bytes measured in the manifest layer (the fact), as manifest siblings
  rather than components{} members, plus an OCI label. An ARG default, not a
  CI-resolved output: no credential for the private skillset, no change at any
  of the four variant build call sites, and a local docker build records what CI
  does. Variant-only, so no base rebuild — check-base-hash.sh scans
  Dockerfile.base alone, verified by running it.
- pi-devbox-version: a skills: section naming baked vs live <repo> @ <sha> per
  vendored skill, and for mempalace whether the live copy is identical to the
  baked fingerprint, at the same commit with uncommitted edits, or divergent.
  entrypoint-user.sh passes the new --no-skills, because the banner prints
  before the links exist and long before the reconcile runs.
- scripts/vendor-mempalace-skill.sh: refresh the file and rewrite the ref
  together (a cp without an ARG bump makes the manifest lie, which is worse than
  anonymity); --check verifies the claim against a real clone.
- 5 new smoke assertions (78 -> 83), mutation-tested through the real sh -c
  path: 6 fabricated manifests, where a well-formed hash of the wrong file
  proves the two manifest assertions are not redundant; the all-baked reporting
  test verified to FAIL against a live-skillset environment.

Reviewed mid-flight by pi@emb-7kj4vr4g over the logstream (correlation
skillset-vendor-drift), which retracted its own earlier recommendation of a
build-time byte-compare against skillset HEAD and supplied the better framing:
the invariant is NON-CONTRADICTION, not currency. Byte parity on a fallback
would have cost a resync commit plus a ~67-min base rebuild for each of the four
skillset commits pushed in one evening. Its warning also found a real bug here:
the script now CONSTRUCTS the snapshot from `git show HEAD:<path>` instead of
copying the working tree, because a clean `git diff` says nothing about an
untracked file — the one input the first draft would have recorded a false ref
for. Tested: untracked, unstaged and staged-but-uncommitted all refuse, atomically.

Also fixes three stale in-repo markers of the same class the canary belongs to
(true when written, silently false at release): two dangling "Unreleased"
pointers and a typst line still marked Unreleased five releases after v1.4.0.
2026-08-26 10:30:27 +02:00
pi dbb78798fb vendor: resync mempalace skill snapshot to skillset c04cd15
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 40s
c04cd15 ('the withdrawal only holds where the bridge is live') landed after the
v1.8.7 snapshot was taken, so the baked fallback was already 4 lines behind the
skillset within hours of publishing. It adds the caveat this fleet is currently
living in: the bridge is baked at image build, so a container on an image older
than the stamping commit satisfies both env gates while stamping nothing, and
hand-stamping is still the only signal a hand-filed drawer gets there. It also
gives the one-line test —
  grep -c MEMPALACE_PI_DEVICE "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)"
which returns 0 on this v1.8.6 container, confirming the gap empirically.

Note what this instance proves about the canary fixed one commit ago: it still
PASSES on the refreshed copy, because both pinned phrases survived the edit. A
phrase canary cannot detect 'older than skillset main' — only a diff can. This is
the second drift in 24h and is the argument for the Still-open item (a CI job
diffing this file against the skillset repo, blocked on a clone credential for a
private repo). No re-pin was needed here.
2026-08-26 09:41:02 +02:00
16 changed files with 2891 additions and 54 deletions
+8
View File
@@ -146,6 +146,14 @@ GIT_USER_EMAIL=
# Detection is automatic if the skillset lives at WORKSPACE_PATH/skillset.
# SKILLSET_CONTAINER_PATH=
# ── cli_utils (standalone commands from a mounted checkout) ──────────
# If a cli_utils repo is mounted, the entrypoint symlinks its bin/ commands
# into ~/.local/bin on every start, so they survive container recreate and
# resolve in non-interactive shells too (docker exec, agent tool shells).
# Detection is automatic at WORKSPACE_PATH/cli_utils (or one level below).
# CLI_UTILS_CONTAINER_PATH=
# CLI_UTILS_LINK=0 # disable the linking entirely
# ── Locale ───────────────────────────────────────────────────────────
# LANG=sv_SE.UTF-8
# LANGUAGE=sv_SE:sv
+51 -9
View File
@@ -64,17 +64,59 @@ re-brand of opencode-devbox's `pi-only` variant.
(`curl -sf 'https://registry.npmjs.org/@earendil-works%2Fpi-coding-agent/latest' | jq -r .version`).
Check release notes at https://github.com/earendil-works/pi/releases for
the upstream changelog to include in `CHANGELOG.md`.
2. Update `CHANGELOG.md` Unreleased → vX.Y.Z section.
3. Verify `docker compose up` works locally with the current `latest` image
2. **Refresh the vendored mempalace skill snapshot if the skillset moved:**
`scripts/vendor-mempalace-skill.sh --check` (reads a real skillset clone,
writes nothing). Three exit codes, not two — a stale-but-truthful record is
**not** a release blocker, so don't treat any non-zero exit as "must
refresh" without reading which one it was:
- **0** — the record is truthful. This includes stale-but-truthful
(upstream has moved past the recorded ref, or the local clone has
uncommitted changes) — a `NOTICE` is printed, but nothing is lying.
**Skipping the refresh in this case is the legitimate, sanctioned
outcome** — every enrolled host reads its own live skillset clone, so
the baked copy is only a no-mount fallback. What is not legitimate is
skipping it *silently*: the drift is visible here, in
`pi-devbox-version`, and in the manifest, so decide rather than forget.
- **1** — a confirmed problem: the vendored bytes provably do NOT match
the file at the recorded ref (a lying record), or the recorded ref
doesn't even resolve to that path in this clone. Refresh.
- **2** — cannot determine (the recorded ref itself isn't resolvable in
this clone — commonly a shallow checkout missing history). Fetch full
history and re-check before deciding; don't refresh blind.
Refresh with `scripts/vendor-mempalace-skill.sh`, which rewrites the file
**and** the ARG together so they cannot drift apart, and refuses (exit 1)
rather than silently rewinding provenance if the skillset clone's HEAD is
behind the already-recorded ref (detached HEAD, older checkout) — pass
`--force` only if that rewind is genuinely intended.
Two consequences to accept deliberately on an actual refresh: the snapshot
is hashed into `base_tag`, so it costs a base rebuild (~67 min); and if the
section the phrase canary names has changed, re-pin it in
`scripts/smoke-test.sh`.
3. Update `CHANGELOG.md` Unreleased → vX.Y.Z section.
4. Verify `docker compose up` works locally with the current `latest` image
if you're upgrading users from a previous version. Then run the
**post-recreate sanity check** inside the running container to confirm
persisted volumes survived and the pi runtime wiring re-deployed (not just
that the container booted):
`docker compose exec devbox bash scripts/recreate-sanity-check.sh --expected-version X.Y.Z`
(or just `pi-devbox-sanity --expected-version X.Y.Z` if `cli_utils/bin` is
on PATH). This is the runtime peer of the build-time `smoke-test.sh` gate.
4. Push tag: `git tag vX.Y.Z && git push origin vX.Y.Z`.
5. Watch CI: smoke job builds amd64 only and asserts size + extensions +
`docker compose exec devbox bash scripts/recreate-sanity-check.sh --expected-image-version X.Y.Z`
(or just `pi-devbox-sanity --expected-image-version X.Y.Z` if
`cli_utils/bin` is on PATH). This is the runtime peer of the build-time
`smoke-test.sh` gate.
**`X.Y.Z` here is the pi-devbox release tag** you are shipping (e.g.
`1.8.9`), which is what the rest of this checklist means by `vX.Y.Z`.
`--expected-image-version` is the flag that asserts it. There is also an
`--expected-version`, and it means something else — the **pi coding agent**
version (e.g. `0.84.3`, the `ARG PI_VERSION` pin). Handing the release tag
to that one used to report *"pi version mismatch: expected 1.8.8, got
0.84.3"*, i.e. a red on the final gate of the release accusing the wrong
component; it now tells you to use `--expected-image-version` instead, and
the reverse mix-up is caught too. Both flags are optional — with neither,
the live pi version is asserted against the version recorded in the image's
own build manifest (which catches a stale `pi` in the `~/.pi/npm-global`
volume shadowing the baked one) and the image tag is reported
informationally.
5. Push tag: `git tag vX.Y.Z && git push origin vX.Y.Z`.
6. Watch CI: smoke job builds amd64 only and asserts size + extensions +
pi version + new-base-tooling presence. Variant build is multi-arch
(amd64 + arm64) only after smoke passes. A tag push fires **only**
`docker-publish.yml` — `lint.yml` is scoped to `branches: ['**']`, which
@@ -85,9 +127,9 @@ re-brand of opencode-devbox's `pi-only` variant.
discovery on `head_sha` **and** the workflow `path` — see *Gitea API access*
below — because that guard costs nothing and a future workflow added on `v*`
would silently reintroduce the ambiguity.
6. Verify the Hub tags appear (latest + vX.Y.Z, the `-studio` pair, plus
7. Verify the Hub tags appear (latest + vX.Y.Z, the `-studio` pair, plus
base-latest if the base was rebuilt this run).
7. **Revoke any short-lived Gitea PAT** used during the release at
8. **Revoke any short-lived Gitea PAT** used during the release at
`gitea.jordbo.se/user/settings/applications`. N/A if you used the
`GITEA_ACCESS_TOKEN` env var instead (see *Gitea API access* below) —
its lifecycle is managed host-side, nothing to revoke.
+976 -1
View File
@@ -11,6 +11,981 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
---
## v1.8.11 — 2026-08-27
**Shell state that the writable layer eats on every recreate now gets rebuilt at
start.** Two additions to `entrypoint-user.sh`, both idempotent, both silent
no-ops when the thing they wire up is absent.
**`cli_utils` commands are linked onto `PATH`.** If a `cli_utils` checkout is
mounted, every executable in its `bin/` is symlinked into `~/.local/bin` at
container start — `git-status-all`, `git-pull-all`, `devbox-sanity`,
`pi-devbox-sanity`, `pi-session-repair`, `docker-clean`, `vpn-status`. Detection:
`CLI_UTILS_CONTAINER_PATH` → `/workspace/cli_utils` → `$HOME/cli_utils` →
`/workspace/*/cli_utils`; `CLI_UTILS_LINK=0` disables it.
The reason this is an *image* concern and not the user's problem to re-solve: on a
host, `cli_utils/install.sh` puts those commands on `PATH` by symlinking them into
`~/.local/bin`, which is persistent there — and **ephemeral here**. Same installer,
same repo, opposite durability, so the fix died on every `--force-recreate` and the
next session was back to typing `/workspace/cli_utils/bin/git-status-all`. Running
`install.sh` *inside* a container is the trap rather than the fix: it re-creates
the same disposable state.
**Symlinks rather than a `PATH` edit in an rc file, deliberately.** `~/.local/bin`
is already ahead of `/usr/local/bin` in `ENV PATH`, so links resolve in
**non-interactive** shells too — `docker exec <c> git-status-all`, agent tool
shells, scripts. An rc-file `PATH` edit cannot reach those: `~/.bashrc` returns
early when the shell is not interactive. Measured on tor-ms22 2026-08-27,
`command -v git-status-all` failed in a non-interactive shell while succeeding in
an interactive one, from exactly that asymmetry. Guards, because `~/.local/bin` is
shared with other tooling: a real file is never clobbered, a symlink pointing
somewhere else is never stolen, our own links are refreshed, and links into a
`cli_utils/bin` whose target vanished are pruned — a dangling link on `PATH`
reports "No such file or directory" and reads as a broken container rather than a
removed script.
**A per-device boot hook: `~/.config/devbox-shell/init.sh`.** If the host provides
one, it runs once at start with output to `~/.pi/agent/devbox-init.log`. That
directory is the host-owned bind-mount already sourced into every interactive
shell by `/etc/skel-devbox/.bash_aliases`, so this is its boot-time twin — the
same ownership and the same persistence, but running *before any shell*, which is
what non-interactive fixups (symlinks, directories, one-off migrations) need. **It
introduces no new trust boundary**: that path is already arbitrary code from the
same owner; only *when* it runs is new. Invoked as `bash <file>`, never sourced,
and its exit status is ignored — a hook must not be able to mutate the
entrypoint's own shell state or stop a container from starting.
With the hook in place, the next "can this run on every recreate?" question needs
no image change at all — which is the point, given what the next paragraph costs.
**This moves the base hash.** `base-decide` folds `cat entrypoint.sh
entrypoint-user.sh` into it, so this change forces the ~40-minute base rebuild at
the next tag whether or not anything else in the base moved. It is a rider, not a
reason to tag.
**How it was validated, since CI cannot.** `docker-publish.yml` runs only on
`push: tags: v*`, and `lint.yml` runs `actionlint` over workflow `run:` steps —
neither one executes `entrypoint-user.sh`. So both sections were extracted and run
against fixtures in a throwaway `$HOME` before commit: real file not clobbered,
foreign symlink respected, stale link pruned, new command picked up, second run
byte-identical, `CLI_UTILS_LINK=0` honoured, and "no `cli_utils` anywhere" a silent
`exit 0`. Then run for real in a live v1.8.10 container, after which
`command -v git-status-all` resolved in a *non-interactive* shell. No
`smoke-test.sh` assertion was added on purpose: the positive path needs a
`/workspace` mount that smoke does not have, and asserting it there would repeat
the v1.8.0 mistake of a smoke assertion written against a stage that does not
exist at run time. `workflow_dispatch` with `smoke_only` remains the way to
exercise this against `HEAD` before a tag.
**Also carried by the floating `mempalace-toolkit` main ref** (resolved at build
time, not by a pi-devbox commit — `MEMPALACE_TOOLKIT_REF=main`):
**A scrubbed re-export of a dormant session could silently never reach the
palace host.** `bin/mempalace-pi-session` ships to the palace with
`rsync -a --update`, and the stage file's mtime is deliberately the SOURCE
transcript's mtime (`os.utime()`, "preserve session mtime for dedup
stability"). Re-exporting a session that has not been appended to since its
last ship therefore produces a mtime that is *not newer* than the receiver's —
exactly the case a redactor upgrade needs to ship, since content differs while
mtime does not. `--update` reported success and sent nothing. Found and
patched by `pi@mbp-m1-2020` (mempalace-toolkit `a361b71`): `--update` →
`--checksum`, which compares content and ignores size/mtime entirely.
Dropping `--update` outright was considered and rejected — rsync's default
quick check already transfers on a size difference alone, which would have
masked the *next* instance of this (a redaction whose placeholder happens to
match the secret's length) as fixed. `os.utime()` is untouched; its backdating
is a separate, load-bearing design call for dedup stability. New regression
test, `scripts/test-rsync-ship-idempotency.sh`, runs fully offline (a local
rsync destination exercises the same size/mtime/checksum comparison as the ssh
transfer) and is built to *discriminate*: it must fail against `--update` and
pass against `--checksum,` not merely exercise the code path — the first draft
of the test used fixture strings of different lengths and passed for the wrong
reason (rsync's quick check transfers on size difference alone regardless of
`--update`), which is the same trap the patch itself was written to avoid.
**Acceptance line for this class of change going forward:** "receiver sha256
matches sender for every staged file", not "local stage is clean" — a clean
local stage says nothing about what a dormant session already sent.
**An event addressed to an identity no session runs as is delivered to
nobody, and this fleet has now hit it three separate ways.** RFC 003 gains
§7.13 and open-decision 10 (mempalace-toolkit `21023e7`, docs only, no image
behaviour change): the owed-set derivation — the log's only push channel — is
keyed on `to_agent`, and a reply is always addressed back to whatever string
the *original writer* put in `from_agent`. Nothing validates that string
against a live session identity, so authoring under a synthetic or foreign
name makes every reply to that event write-only. Measured cost this cycle: a
directed ask planted under a synthetic sender drew a correct reply containing
an urgent security finding, and it sat unread for ~2h20m, found only because a
human asked whether mail had arrived. Permitted exception, unchanged: a
synthetic sender is fine for a deliberate control experiment, provided the
body names the real identity to reply to.
**Also carried by the live `skillset` mount** (each device's own clone, not
baked — except the `mempalace` skill's fallback snapshot, re-vendored below):
**The mermaid-diagrams checker's cut gate moved from client pixels to a
per-SVG user-space unit.** `CUT_PX` was calibrated against one live page at
one render scale; sweeping `--viewport` 500→1600 on an *unchanged* document
moved the worst overflow −1.0px → −3.0px, near-proportional to the viewport,
i.e. a constant geometric overflow viewed through a changing scale. `cutU =
cutPx / scale` (scale taken per-SVG, never a page average — one page mixes
scales 0.643–0.988) recovers that invariant: the sweep now collapses to
exactly −3.0u at every width. Re-deriving the threshold against the live host
surfaced a real false negative the old pixel gate had: a label at
`cutPx=0.4, scale=0.678` read as healthy under `CUT_PX=0.5` but is `0.59u` —
a genuine cut hiding behind a compressed render scale. `CUT_U` stays `0.5`;
`cutPx` and `scale` are still printed on every issue so a devtools ruler still
confirms the number on the actual page. A new, explicitly-deferred finding
from the same review: `cut` only measures vertically, so an unbreakable token
wider than its box (a long URL, a `snake_case` identifier) is invisible to
soft-wrap, tall, *and* cut simultaneously — filed as a backlog item, not
implemented, pending a fifth acceptance control.
**The `from_agent`-identity finding above is also now in the `mempalace`
skill itself** ("Writing to another machine", and Anti-Patterns), and the
baked fallback snapshot of that skill was refreshed to match
(`vendor-mempalace-skill.sh`, `6eb20af` → `a12fe5e`) — sanctioned to skip on
its own (`--check` reported stale-but-truthful), done anyway because this
release's point is getting today's fixes live, and the base rebuild below was
already forced regardless.
### Dependency audit (2026-08-27)
Every component checked against upstream by direct command, not assumed:
| Component | Baked in v1.8.10 | Upstream now | Action |
|---|---|---|---|
| **mempalace-toolkit** | `b2b50af` | **`21023e7`** | ships the rsync ship-fix + RFC 003 §7.13 (both above) |
| **skillset** (mempalace fallback snapshot) | `6eb20af` | **`a12fe5e`** | re-vendored (above); live-mounted devices already had it |
| pi | `0.84.3` (pinned) | `0.84.3` is npm latest | none |
| mempalace | `3.8.0` (pinned) | `3.8.0` is PyPI latest | none |
| pi-atelier | `v0.8.2` (pinned) | `v0.8.2` highest tag | none |
| pi-studio (studio variant) | `v0.9.52` | `v0.9.52` — `main`'s commit and the tag's commit are identical (0 either direction) | none |
| pi-toolkit | `0e1369e` | `0e1369e` (local clone HEAD == `origin/main`) | none |
| pi-extensions | `2022887` | `2022887` (local clone HEAD == `origin/main`) | none |
| pi-fork | `bf702b4` | `bf702b4` | none |
| pi-observational-memory | `ce9fc98` | `ce9fc98` | none |
pi-toolkit / pi-extensions checked against their actual Gitea origin (the
Dockerfile's `PI_TOOLKIT_REPO` / `PI_EXTENSIONS_REPO`), not a GitHub mirror —
querying `api.github.com` for those two returned nothing (rate-limited or
blocked; not investigated, the local clones are the source of truth anyway).
No SHA above is fork-supplied; a fork inventing plausible-looking upstream SHAs
is a recorded failure mode (v1.8.9), so every value here came from
`git ls-remote`, a local clone's own `origin/HEAD`, `npm view`/registry JSON,
or the PyPI JSON API, run directly.
---
## v1.8.10 — 2026-08-27
**This tag exists to deploy a fix and a safety net that are currently running on
exactly one machine.** The feeder scrubber has been hand-copied to `/opt` on one
device since this morning; every other device has kept staging unscrubbed
transcripts into the shared palace. Nothing here is a new capability for its own
sake.
`MEMPALACE_TOOLKIT_REF=main` floats: `docker-publish.yml` resolves it to a
concrete SHA at build time, so whatever is on toolkit `main` when the tag is
pushed ships in that image whether or not this repo has a commit. That is the
rule v1.8.9 adopted after `553d865`/`5b8d78f` shipped undocumented twice — *name
the behaviour change before tagging, not after* — and this entry is that rule
being obeyed rather than re-learned.
**`mempalace-toolkit` main moves `5b8d78f` → `b2b50af`** (13 commits, ~2100
insertions / ~520 deletions). No pi-devbox commit implements any of it.
### ⚠️ THE ONE CHECK THIS RELEASE MUST NOT SKIP
**On the first client that runs this image, verify the memory feed still stages.**
The feeder is *fail-closed* by design: no redactor module, no staging (`exit 3`).
That is correct behaviour and it is also the failure mode with no alarm — a
packaging or path mistake stops the fleet's entire transcript feed and nothing
complains loudly, because refusing to stage looks exactly like a quiet session.
This is not hypothetical. `f0bffd1` exists because the feeder is installed as a
symlink (`/usr/local/bin/mempalace-pi-session` → `/opt/mempalace-toolkit/bin/…`)
and `${BASH_SOURCE[0]}` reports the *symlink* path, so the module lookup landed in
a directory where it does not exist. Had that shipped, every device would have
refused to stage on first boot. It was caught by execution, not by review.
Acceptance, in order, on the first recreated client:
1. Run a session, then confirm the feeder logged a scrub summary — a
`[scrub]` line with tier-tagged counts (`T1:env-value=…`, `T2:github-pat=…`),
or an explicit "zero redactions". **Silence is the failure signal**, not success.
2. Confirm the palace drawer count *moved* for that session (the feed reached the
server, not just the stager).
3. Confirm `exit 3` did **not** fire: `mempalace-pi-session` invoked through the
`/usr/local/bin` symlink must find `mempalace_redact.py`.
4. Only then trust the rest of this release.
If step 1 or 3 fails, the memory feed is down fleet-wide until it is fixed, and
sessions that ran in the meantime are not recoverable from the palace — they were
never staged. `MEMPALACE_FEED_ALLOW_UNSCRUBBED=1` is the loud escape hatch, and
using it means accepting unscrubbed transcripts until the packaging is repaired.
### Dependency audit (2026-08-27)
Every component checked against upstream, not assumed:
| Component | Baked in v1.8.9 | Upstream now | Action |
|---|---|---|---|
| **mempalace-toolkit** | `5b8d78f` | **`b2b50af`** | ships the scrubber + symlink fix + `hlc` join |
| **pi-studio** (studio variant) | `v0.9.48` | **`v0.9.52`** | 22 commits, additive only — see below |
| pi | `0.84.3` (pinned) | `0.84.3` is npm latest | none |
| mempalace | `3.8.0` (pinned) | `3.8.0` is PyPI latest | none |
| pi-atelier | `v0.8.2` (pinned) | `v0.8.2` highest tag | none |
| playwright | `1.62.1` (floats `latest`) | `1.62.1` | none — no drift this cycle |
| pi-fork | `bf702b4` | `bf702b4` (2026-08-24) | none |
| pi-observational-memory | `ce9fc98` (v3.0.4) | `ce9fc98` | none |
| pi-toolkit | `0e1369e` | `0e1369e` (2026-08-07) | none |
| pi-extensions | `2022887` | `2022887` (2026-08-17) | none |
**`pi-studio` `v0.9.48` → `v0.9.52`** — four releases, 22 commits, all additive:
PDFs open directly in Studio with watched previews, the header can hide, and
contextual *side questions* arrive (selected-tool use, frozen git context, export,
keyboard shortcuts). No removals or renames in the diff; the changes are
concentrated in `client/studio-client.js`, `index.ts` and three new `shared/`
helpers.
**Its `pi` floor is `>=0.84.3` and we pin exactly `0.84.3` — satisfied with zero
headroom.** Worth naming as a watch item rather than a problem: the next studio
release that raises the floor breaks the studio variant until `PI_VERSION` moves,
and that failure surfaces at build time in the studio job only, after the core
variant has already published.
### Also pulled in by the floating toolkit ref (documentation only)
RFC 003 gains **§9.2**, a proposed direction for the one open decision this
fleet keeps tripping over — that a report addressed to a device is never
delivered, because mailbox candidacy requires exactly `status="open"`. It records
a negative result worth keeping: widening the owed set to include terminal events
cannot work, since the asserting shape and the clearing shape must be disjoint or
every closure mints a fresh obligation. No code implements §9.2 in this release.
The skillset snapshot also moves, so this image bakes the mermaid-diagrams skill's
Playwright driver and the honest note that a `claimed` ack notifies nobody.
### Transcripts get scrubbed before they are staged (`3d47937`, `836e35b`, `f0bffd1`)
`bin/mempalace_redact.py`, called from `mempalace-pi-session` at the moment the
staged transcript is written — one hook covering both transports, because local
mode mines that file and remote mode rsyncs the same bytes.
- **Why it exists, measured rather than argued.** One leaked bearer token had
reached 3 drawers, 13 feeder inbox files across all three devices, and 10 local
files spanning 10 days, from an agent printing an env var while debugging. A
second sweep then found `GITEA_ACCESS_TOKEN` in 2 more drawers and
`GITEA_EGL_ACCESS_TOKEN` in 3. This is routine agent behaviour, so the fix
belongs in the pipeline, not in discipline.
- **Detection is name-anchored, never entropy-anchored.** A palace's own primary
keys — drawer ids, chunk ids, event ids, replica ids, HLCs, commit SHAs — *are*
its high-entropy strings, so an entropy detector eats the memory it protects,
silently and unrecoverably. Three tiers instead: T1 literal values from this
process's env whose name says secret (zero false positives by construction);
T2 vendor shapes (`ghp_`, `glpat-`, `xox*-`, `sk-`, `AKIA`, JWT, PEM, URL
credentials, `Authorization:`); T3 key-name-says-secret.
- **T3 is report-only, because the false-positive rate was measured.** On 52 MB
of real fleet transcripts T3 fired 403 times, mostly `${VAR}` interpolation in
compose files, TypeScript identifiers, a *type annotation*
(`credentials: Credentials`), an IPA attribute holding a date
(`krbPasswordExpiration`), AAAK diary shorthand, and terminal output following
an ssh `Password:` prompt. With interpolation/code-context/key-suffix guards the
enforced count fell **403 → 29** on the same corpus. `MEMPALACE_REDACT_STRICT=1`
makes T3 enforce.
- **Operational shape.** Fail closed — no redactor, no staging (`exit 3`),
overridable with `MEMPALACE_FEED_ALLOW_UNSCRUBBED=1`. Every run prints a count
*including* `0 redaction(s)`, because silence is indistinguishable from a
scrubber that never ran. Findings carry rule, label, length and `sha256[:8]` —
never the value.
**Near-miss this image would have shipped, caught before tagging (`f0bffd1`).**
The image installs `/usr/local/bin/mempalace-pi-session` as a **symlink** into
`/opt/mempalace-toolkit/bin`, and `${BASH_SOURCE[0]}` reports the invoked path,
not the target — so the sibling-module lookup resolved to `/usr/local/bin`, the
redactor was absent, and fail-closed did as instructed: `[FATAL] ... refusing to
stage`. Measured side by side, the symlinked invocation FATALed while the direct
one scrubbed 40 findings. **At the next bake that would have stopped every
feeder tick on every device — a silent fleet-wide memory outage, worse than the
leak the scrubber prevents.** Fixed by chasing the symlink chain in portable
shell (`readlink -f` avoided: GNU/newer-BSD only, and this script also runs
directly on macOS hosts) with colon-separated fallback candidates. Verified via
the symlink, the direct path, and a second-hop symlink. General lesson: fail-closed
converts "module not found" into an outage, which makes the module lookup
load-bearing infrastructure that must be tested through the invocation path the
fleet actually uses — not the convenient one from a checkout.
### The mailbox becomes explainable and mesh-safe (`bfe9c5c`, `a92c75d`, `e917662`, `ecc2a9c`)
- **Owed-set derivation joins on `hlc`, not `seq`** (`bfe9c5c`). `seq` is a
replica-local arrival counter — the same event is `#7` in one database and `#12`
in another — so a second replica would let already-answered asks resurrect.
`hlc` is immutable and replicated, fixed-width, so string comparison *is* causal
comparison. A safe no-op on today's single replica (verified: the positive-control
pair orders identically under both keys), correct once a mesh exists.
- **Delivered text now says it is queued** (`a92c75d`). Delivery uses `steer`
with no `triggerTurn`, and the poll fires on `agent_settled`, so nothing wakes
the model — a delivered ask sits until a human starts the next turn. Measured
case: a directed report sat unread for 2.5 hours. The note explains the agent is
not ignoring the ask, it is not running.
- **`MEMPALACE_MAILBOX_NOTIFY` gains explicit `=kitty` / `=osc777` modes**
(`e917662`). Terminal autodetection inside a container is not unreliable, it is
*blind*: `docker exec` forwards neither `KITTY_WINDOW_ID` nor `TERM_PROGRAM`, and
`TMUX` is unset because tmux runs on the host. Verified on a live process:
`TERM=xterm-256color` and nothing else.
- **The terminal path through tmux is documented as UNVERIFIED** (`ecc2a9c`).
Test sequences written to the pty produced no notification on a remote client;
tmux likely drops unknown OSC types without `allow-passthrough`, and multi-client
routing (one ask pinging every attached client) is an open question.
### Documentation (`e1cc759`, `982b001`, `d4d8bb6`, `d2764bf`)
- **RFC 003, the coordination-log spec the code had been citing all along** — it
did not exist anywhere. 11 sections, retrospective against mempalace 3.8.0
`logstream.py`, incl. owed-set derivation, ten dogfooded landmines and seven open
decisions. Non-obvious findings: `event_append` has **no** idempotency guard on
the write path (verify-before-retry; the replication path *is* guarded),
coordination tools are exempt from both palace locks by design,
`GET /logstream/events` **never existed** in 3.8.0 (not proxy-blocked), and
`mempalace sync` never touches the logstream — the log is permanent and unbounded.
- **`docs/fleet-memory.md`**, operator-facing: five storage types, a decision tree,
latency expectations (~2–5 min live session; next session while offline),
broadcast exclusion by design, fan-out, and the search-before-answer /
diary-at-session-end / verify-don't-retry habits.
- **`docs/secret-hygiene.md`**, incl. the tier definitions, the measured FP data,
stated false negatives, and the three server-side call sites (specified, not
built — tier 2 only there, since the hub cannot see a client's env).
- Phase 1 exposure record moved to the private fleet repo with a moved-note stub;
retention direction for the unbounded log (logrotate-style: never rotate
still-owed events, rotation invalidates held cursors, archive-verify-delete).
### The other memory system finally gets explained — `docs/observational-memory.md`
`pi-observational-memory` has been baked for several releases and described in
one line of the feature list (*"the `recall` tool for session compaction"*),
which is enough to name it and not nearly enough to use it. New 272-line
explainer with five diagrams, aimed at someone who has seen `/om:status` or a
"compacted memory" block and wondered whether to leave any of it switched on.
**Scoped to what this repo is authoritative for, because upstream already
documents the mechanism well.** `/opt/pi-observational-memory/docs/` ships
`concepts.md`, `how-it-works.md` and `configuration.md`, including a correct v3
lifecycle diagram — so the new document links those for depth and spends its own
words on the four facts pi-devbox owns and can change: the pinned commit it bakes
(v3.0.4 `ce9fc98`, the value in `build-manifest.json`), the `packages[]` entry
that decides which copy loads, the Haiku-workers-vs-Opus-session split seeded
into `~/.pi/agent/settings.json`, and the `devbox-pi-config` volume that makes
the ledger survive `--force-recreate`. Plus the confusion this image creates by
shipping two things called memory: a section contrasting it with MemPalace, on
the line *observational memory keeps a session coherent, the palace keeps the
fleet coherent*.
Every stated number was read out of the live container or the baked tree rather
than copied from release notes — including the correction that the dropper is
gated on a **successful same-turn reflection** and not on a token threshold of
its own, which is the one detail `pi-extensions/SKILL.md` still gets wrong.
Placement follows the audience split fleet-ops states for itself: reusable
mechanism is not deployment data, so a "why is this in my container" document
belongs in the repo that **pins and wires** the component, pointing upstream for
depth. Linked twice from the README, because before this commit the README
referenced `docs/` zero times and the one file already there
(`mempalace-broker-design.md`) was reachable only by listing the directory.
### A README claim that v1.8.9 made false, and how it got there
**§ Cross-machine agent coordination ended with "Nothing in this image polls the
log on the agent's behalf." The mailbox shipped in v1.8.9, so that sentence has
been wrong since `aac4a1c`.** Replaced with the three knobs and their defaults
(`MEMPALACE_MAILBOX`, `MEMPALACE_MAILBOX_POLL_MS` 300000,
`MEMPALACE_MAILBOX_RESURFACE_MS` 3600000), the fact that owed-ness is *derived*
rather than read off `status`, and the queued-into-the-next-turn delivery
semantics measured on two devices.
The cause is the one v1.8.9 wrote a rule about: the behaviour arrived through the
floating `MEMPALACE_TOOLKIT_REF`, so **no diff in this repo ever touched the
paragraph that made the claim**. v1.8.9's rule ("name a floating-ref behaviour
change in the CHANGELOG before tagging") fired for the CHANGELOG and nobody swept
the README. The CHANGELOG records what *changed*; the README asserts what is
*true*, and only the first is reviewed at release time. Extending the rule
accordingly: grep the README for absolute claims — *nothing*, *never*, *does
not*, *only* — about any component whose SHA moved.
**The replacement is dated on purpose.** It says it describes the bridge *as baked
in v1.8.9* (`mempalace-toolkit` `5b8d78f`) and points at that repo's
`docs/rfc-003-coordination-log.md` §7.11–§7.12 for the mechanism, because toolkit
main is already ahead of the baked copy (`a92c75d` makes delivery say it is queued
and ping the human who is not looking; `e917662` and `ecc2a9c` refine that notify
path) and none of it reaches a container until a base rebuild. Documenting those
here would have swapped a stale-behind claim for a stale-ahead one — the same
defect with the sign flipped.
### Diagrams verified by rendering, not by parsing
Both comparison diagrams **parsed clean and rendered with their meaning
reversed**: Mermaid laid the second declared `subgraph` out first, so "with
observational memory" appeared before "without", and MemPalace before
observational memory in the diagram whose entire job was that contrast. A third
was legible only at 1280px. Rebuilt as declaration-ordered node chains, then
re-rendered at mermaid@11 — the version `pi-studio` pins — in the baked headless
browser and read back as an image. Recorded because it generalises:
`mermaid.parse()` proves syntax and says nothing about layout, so a diagram is
unverified until someone has looked at it.
### … and rendering it in *my* browser was still not enough
Reported from a real viewer: several boxes had their bottom line of text sliced
off. Reproduced and root-caused rather than nudged — **Mermaid measures a node
label with its own font metrics, computes the box, then renders the label as real
HTML inside a `<foreignObject>`.** Any host stylesheet that touches the
`line-height` or `font-size` of that HTML makes the text taller than the box
already committed to, and the overflow is clipped at the box edge. Error
accumulates per line, so the loss always lands on the last line of the tallest
labels — which is exactly what was reported.
Two fixes were tried and only the second works:
- `%%{init: {'flowchart': {'htmlLabels': false}}}%%` — **rejected, and verified
ineffective rather than assumed so.** The directive *is* honoured (label
elements switch from 16 `foreignObject` to 7 `tspan`), and the clipping is
identical, because the inflated font-size still inherits into SVG text.
- **A hard limit of two short lines per node, with the detail moved into the prose
under each diagram.** One- and two-line boxes have enough vertical slack to
absorb the inflation; three- and four-line boxes do not. This is also better
documentation — the old nodes were carrying paragraph-sized text.
The regression harness is now the interesting artefact: render every block with a
deliberately inflated `line-height: 1.7 !important` on the label HTML, screenshot,
and read it. Two survivors of the rewrite were caught only by that harness — a
long unbreakable `/opt/pi-observational-memory` path silently wrapping to a third
line, and a cylinder (`[( )]`) shape, whose curved bottom leaves less room than a
rectangle for the same two lines.
### §4 answers the question the document left hanging: what compaction does to your context
Asked directly and worth writing down: *if the old conversation is folded away, is
the session back to knowing nothing?* No — and the specifics are all checkable
against pi 0.84.3's own `docs/compaction.md` and the extension's source:
- **A verbatim tail survives, sized by a token budget rather than a message
count.** Pi walks back from the newest entry until `keepRecentTokens` [20000],
and everything from that `firstKeptEntryId` onward is kept **unchanged**. Cut
points land on turn boundaries, never mid-tool-call.
- **The system prompt and `AGENTS.md` are not in the compacted region at all** —
they are rebuilt from disk on every request, so compaction cannot lose them.
- **Nothing is deleted from disk.** Compaction *appends* a `compaction` entry
carrying the summary and the cut pointer; no session line is rewritten in place.
- **`recall` therefore still resolves ids whose sources left the context**, because
it reads the full branch via `sessionManager.getBranch()` and never consults the
context window.
- **Repeated compaction does not summarise the summary.** The text is always
rendered from live observation/reflection records, so there is no
generation-loss spiral; the projection is incremental against the last full-fold
boundary and escalates to a true re-fold from the branch root at
`observationsPoolMaxTokens` [20000].
And one correction to this repo's own earlier claim: **"compaction calls no model"
is a steady-state property, not an absolute.** If the ledger is empty — compaction
firing before the observer has ever run — the hook returns nothing and explicitly
declines ownership (`// Decline ownership so Pi's native summarizer preserves the
pre-cut context.`), and pi's own model-based summariser runs. The doc now says so,
with the snippet.
### A shipped doc bug: the ledger entry type was stated exactly backwards
§9 told readers the entries are `custom_message` and specifically *not* `custom`.
It is the other way round, so the one grep the section existed to get right was
the one it got wrong. Corrected against the live session file — 11
`om.observations.recorded` and 6 `om.reflections.recorded` entries, all
`"type":"custom"`, alongside `"type":"custom_message"` entries whose `customType`
is `mempalace-mailbox` and `mempalace-wakeup`, which is precisely where the
confusion came from: **the mailbox uses the context-visible API, om's ledger uses
the invisible one.**
That is not a typo but a load-bearing distinction, and the fix turns it into a
feature the doc now advertises: `custom` entries *"do not participate in LLM
context"* (pi `docs/session-format.md`), so **the ledger costs zero context until
it is folded** — now a row in the cost table.
### Not covered by any of this
The opencode bridge is a separate write path the feeder hook never sees, and the
server-side layer is unbuilt — so a secret typed straight into `add_drawer`, or
staged by a non-pi client, still lands unscrubbed.
---
## v1.8.9 — 2026-08-26
The coordination log gets a reader, and the release checklist's last gate stops
accusing the wrong component.
### The mailbox arrives — named here *because nothing in this repo caused it*
**`mempalace-toolkit` main moves `e70bef2` → `5b8d78f` (exactly one commit, 281
insertions / 9 deletions across `extensions/pi/mempalace.ts` and
`extensions/pi/README.md`), and that is what actually ships the auto-delivered
logstream mailbox.** No pi-devbox commit implements it. `docker-publish.yml`
resolves `MEMPALACE_TOOLKIT_REF=main` to a concrete SHA at build time and folds
that SHA into `base_tag`, so the mailbox would have landed in the next tagged
image **whether or not this section existed** — which is precisely why it exists.
That is the same shipped-undocumented shape as `553d865` in v1.8.7, and that one
caused a cross-host misattribution: an agent on another machine reasoned about
which image contained which behaviour from a CHANGELOG that never mentioned it.
The rule this release adopts: **if a floating ref will pull a behaviour change
into the image, name it in the CHANGELOG before tagging, not after.**
What the mailbox does, from the shipped code rather than from the design
discussion:
- **The bridge was write-only.** It stamped provenance on the way *out* and never
read the log back, so a directed ask reached an agent only if that agent
happened to run `mempalace_event_list` itself. The channel carried real
cross-machine traffic from 2026-08-18 onward with **zero readers** — every
delivery in that window happened because a human said "check your mailbox".
- **Doubly gated, exactly like the provenance stamper:** inert unless *both*
`MEMPALACE_PI_DEVICE` and `MEMPALACE_REMOTE_URL` are set. An unstamped client
has no address to be reached at, so there is nothing for it to read.
- **On by default, opt out with `MEMPALACE_MAILBOX=0`.** Deliberate: an opt-in
fix for a nobody-remembers-to-do-it problem only relocates the forgetting.
Tunables: `MEMPALACE_MAILBOX_POLL_MS` (min gap between mid-session polls,
default 300000) and `MEMPALACE_MAILBOX_RESURFACE_MS` (re-announce a still-owed
ask after, default 3600000).
- **Owed-ness is derived, never read off `status`.** `event_ack` appends and never
mutates, and `status` is written once, so a directed `open` keeps matching the
mailbox query forever — answered or not. A candidate counts as answered only
when one of this device's own events has a **higher `seq`**, joins via
`metadata.ack_of` or a shared `correlation_id`, and carries a terminal status
(`applied`, `superseded`, `failed`, `blocked`). `claimed` and `ready` are
deliberately **not** terminal — that is how "taken, but not finished" keeps
resurfacing.
- **`*` broadcasts are excluded from the owed set.** `to_agent: <me>` also matches
broadcasts per the tool contract, so without this a broadcast written with
`status="open"` would make every machine believe it personally owed the same
answer — and the code would contradict the skill that documents it.
- **The dedup map is in memory on purpose.** A restart forgets, so an already-seen
ask can resurface: visible noise a human corrects in one turn. The opposite
failure — suppressing an unanswered ask — is silent and permanent. Do not
"fix" the noise by persisting it.
- **Delivery queues, it never interrupts.** A sections push at
`before_agent_start` plus a second `agent_settled` handler behind the 300 s
floor, using `steer` and *not* `triggerTurn`: `agent_settled` means idle, so
nothing wakes a model on inbound fleet traffic.
Measured on v1.8.8 (which bakes `e70bef2`, i.e. no mailbox) immediately before
this release: the wake-up mailbox query had to be run by hand, returned **3**
directed asks with `status="open"`, and the derivation above resolved **all
three** as already answered — the third independent confirmation that the raw
`status` filter never shrinks, and the first taken on a fresh container with no
memory of having answered them.
### `--expected-image-version`: two versions, two flags
**`scripts/recreate-sanity-check.sh --expected-version 1.8.8` reported
`✗ pi version mismatch: expected 1.8.8, got 0.84.3` and exit 1** — a red on the
final runtime gate of a release, accusing the image of being the wrong version,
when the flag had only ever asserted `pi --version`. `AGENTS.md` step 4 spelled
it `--expected-version X.Y.Z` inside a checklist where every *other* `X.Y.Z` is
the pi-devbox tag; `README.md` got it right, so the two documents disagreed.
Not hypothetical, and not one reader's slip: the v1.8.8 release-readiness handoff
from `pi@emb-7kj4vr4g` (`evt_20260826T134919_a614ecfc2d4f`) propagated
`--expected-version 1.8.8` twice, in its body and in
`metadata.cannot_check_here`, while correctly calling step 4 "the runtime peer of
the smoke gate, so it is not ceremonial". Two independent readers, one on another
machine, converged on the wrong meaning. Left alone it puts a spurious red on
every release, and the intuitive remedy — re-pull, re-recreate — is pure waste.
- **New `--expected-image-version X.Y.Z`** asserts the pi-devbox release tag,
read from `release_tag` in `/etc/pi-devbox/build-manifest.json` (the image's
own build-time ground truth — no checkout, no network, no Docker socket). A
leading `v` is optional on either side, so `1.8.9` and `v1.8.9` both work.
- **Both flags now detect being handed the other one's value**, and the test is
exact rather than heuristic: the value is compared against the *other*
quantity this image actually reports, so it can only fire when the mix-up is
real. `--expected-version 1.8.9` now says *"is the pi-devbox IMAGE version,
not the pi version — use `--expected-image-version`"*, and the reverse mix-up
is caught the same way.
- **Neither flag is required any more.** With none, the live `pi --version` is
asserted against `pi_version` in the build manifest. That is not a tautology:
`pi` resolves through `PATH`, and a stale install in the `~/.pi/npm-global`
volume can shadow the baked one — the same shadowing this script already
guards against for `npm:pi-atelier` in `packages[]`. Verified by mutating the
manifest to a different version, which made the new check fail as intended.
- **The header note it replaced was stale and load-bearing.** It claimed pi "is
resolved from `latest` at CI build time and is NOT pinned … cannot self-derive
an expected version". `Dockerfile.variant` pins `ARG PI_VERSION=0.84.3`, and
`docker-publish.yml` *reads that ARG* as its source of truth (refusing to
build on a floating value, checking it is published on npm, warning when npm
is ahead). The same withdrawn claim also sat in `cli_utils`'s
`pi-devbox-sanity --help`, the third place this confusion lived; fixed there
too, in that repo.
- Argument parsing hardened while in there: a flag whose value is missing — or
is another flag — is now a usage error (exit 2) instead of silently consuming
the next argument, and `--help` works.
All fourteen flag combinations were exercised by execution, including the two
manifest-absent branches and the shadowing branch, which a healthy container
cannot reach naturally — mutation-tested with a doctored manifest path so that
each failure branch was observed *firing* rather than assumed present.
### Component audit: no bumps, and that is the finding
Checked before tagging, since a base rebuild was already forced:
| Component | In v1.8.8 | Upstream now | Action |
|---|---|---|---|
| pi (npm) | `0.84.3` (pinned) | `0.84.3` is `latest` | none |
| mempalace (PyPI) | `3.8.0` (pinned) | `3.8.0` | none |
| pi-atelier | `v0.8.2` (pinned) | `v0.8.2` highest tag | none |
| pi-toolkit, pi-extensions, pi-fork, pi-observational-memory, pi-studio | floating | **identical to baked** | none |
| skillset snapshot | `6eb20af` | `6eb20af` | none |
| **mempalace-toolkit** | `e70bef2` | **`5b8d78f`** | ships the mailbox |
So the whole ~67-minute base rebuild this tag pays for is attributable to the
toolkit SHA alone — `base_tag` folds it, and it moved. Every other floating ref
resolved to the commit already baked (verified with `git ls-remote` per repo, not
by reading a cached clone).
One claim in this audit came from a fork that had fabricated its findings — six
plausible-looking toolkit commits with five nonexistent SHAs, a pi `0.84.4` that
npm has never published, a pi-studio commit `ls-remote` says does not exist, and
a compatibility floor of `0.8.2` where the code says `0.7.1`. Every row above was
therefore re-measured directly. Recorded because the failure mode is specific:
none of it looked wrong, and `git cat-file -e` is what caught it.
---
## v1.8.8 — 2026-08-26
The vendored `mempalace` skill snapshot stops being anonymous, and the
container starts saying which copy of each skill it is actually reading.
**Peer review (pi@emb-7kj4vr4g, logstream correlation
`skills-provenance-review`, full text in
`drawer_pi-devbox_reviews_e43e766641c9ec85217bc6ce`) found three blockers before
this was tagged. All three were the same species: a record asserting something
it had not verified. Every finding below was reproduced by execution here before
being fixed.**
- **The verification gate could print `OK` and exit 0 without verifying
anything.** `git show <ref>:<path> | sha256sum` hashes *empty stdin* when the
ref does not resolve, yielding a real-looking `sha256("")` rather than an
empty string — so the `UNKNOWN` branch in `--check` was dead code. Reproduced:
a bogus ref reported `MISMATCH` (accusing the snapshot of lying when the true
cause was an incomplete clone — and the operator's natural remedy for
MISMATCH is to re-run the refresh, which *rewrites provenance to silence the
complaint*); with a 0-byte snapshot against a 0-byte upstream file it printed
`OK: … exactly skillset@aaaaaaa` and exited 0 for a ref that does not exist.
The script already had the right idiom (`sha_empty`) and had applied it to
`blob_sha` but not to `at_ref`. Now existence is *proven* with `git cat-file
-e` before anything is hashed, at two levels (does the ref resolve; does the
path exist at it) because those are different failures. This was the same
defect class as the canary it replaces: a check that can succeed without
checking. A second, unflagged instance of the identical pipeline shape was
found in `blob_sha` and fixed too.
- **`--check`'s exit codes conflated "stale" with "lying",** so the release step
failed in the case `AGENTS.md` step 2 explicitly calls legitimate. Now: `0`
truthful (including stale-but-truthful, with a `NOTICE`), `1` a lying record
only, `2` cannot determine (ref absent from this clone). `AGENTS.md` step 2
rewritten to state all three, since its promise that "the message
distinguishes the two" was exactly what the branch was breaking.
- **The staleness `NOTICE` then asserted a direction it had never tested** — the
same defect one layer down, found by pi@emb-7kj4vr4g against the real state of
its own host. The branch fired the notice on "recorded ≠ HEAD" and announced
that HEAD was the newer side, so a clone that was merely *behind* was told
*"has moved to 82a8d3c; the snapshot describes the older 5fd0d5c"* when
`82a8d3c` is `5fd0d5c`'s **ancestor**. Harmless to the verdict (`rc` stayed 0,
nothing was mis-verified) but it points the operator at a refresh — a
~67-minute base rebuild — when the real remedy is `git pull`. It now tests
ancestry with the `merge-base --is-ancestor` primitive the refresh path two
sections above already used, and reports three distinct verdicts: **stale**
(recorded is an ancestor — refresh), **your clone is behind** (HEAD is an
ancestor — pull, do not refresh), **diverged** (neither). All three verified by
execution; only the first was right before.
- **`--help` died with `unknown option: --help`.** The strict argument loop that
closed the silent-ignore hole never added a `--help` case, so the one script
whose argument *order* was itself a landmine had an erroring discoverability
path. It now prints its own header block.
- **`VENDORED.md` contradicted itself, in the release whose stated invariant is
non-contradiction.** Its hand-maintained "Snapshot provenance at last refresh"
line named skillset `670f7f1` — seven commits behind the ARG, and *the very
commit that told agents to hand-stamp `added_by`*, i.e. the withdrawn
instruction this line of work exists to stop shipping — while its `cp` recipe
still contradicted the "not `cp`" rule 20 lines above. The hand-maintained
line is gone (nothing forced it to move when the ARGs did); `670f7f1` is kept
only as a labelled cautionary example. The `pi-extensions` half was verified
redundant (CI resolves `PI_EXTENSIONS_REF` via `require_sha`) before removal,
rather than silently dropped.
**Should-fixes from the same review, all reproduced:** `--check` given the
documented positional spelling (`<root> --check`) silently ran a *refresh*,
because only `$1` was parsed — both tools now parse all arguments and reject
unknown ones; a refresh at a detached or older `HEAD` silently rewound ref and
bytes, now refused unless the recorded ref is an ancestor (`--force` to
override); `upstream_dirty` was computed and never used in check mode, now
reported; `pi-devbox-version --no-skills --json` printed human text and broke
`jq`; `--help` was a hardcoded `sed -n '2,22p'` range that this branch had
already made stale; the skill fingerprint hashed `SKILL.md` alone, so a live
skill dir differing only in a sibling file still reported "identical" — and
`pi-extensions` already ships two files — so it is now a per-skill **tree** hash
and the manifest field is renamed `skillset_snapshot_tree_sha256` to say what it
measures; and the `--no-skills` smoke assertion was negative-only, passing on a
crashed binary, now anchored positively. `mktemp`+`mv` left written files at
`0600` (a `mv` takes the temp file's mode) — CI was unaffected because the git
index records `100644`, but a local build from a dirty tree would have baked it;
now `chmod 0644` before the `mv`.
**The skill fix ships outside this release, because it had to.** The review also
found that skillset `82a8d3c` — the coordination protocol itself — told every
machine on this fleet to *skip* the mailbox it introduced: it gated the mailbox
on `mempalace_mesh_peers`, and a hub-and-spoke palace reports `peers: []`
precisely because every machine is a thin client of one replica. It also
asserted that a directed `open` event "stays in their mailbox until" acked —
false, because `event_ack` appends and `status` is written once, so an answered
ask matches forever. The headline measurement behind that claim ("exactly 1 —
the one that needed a reply") was of an event already acked half an hour
earlier. Fixed in skillset `5fd0d5c`, which derives owed-ness by joining on
`ack_of`/`correlation_id` with a **`seq` ordering test** — without which one
terminal reply suppresses every later ask on the same thread forever. Because
the skillset is mounted live on every enrolled host, that correction was already
deployed fleet-wide before this image was built; the vendored snapshot is
resynced to it (`c04cd15` → `5fd0d5c` → `6eb20af`) so the no-clone fallback does
not ship the withdrawn rule. Canary re-verified bidirectionally against the new
bytes.
`6eb20af` adds the limit of that ordering test, found when pi@emb-7kj4vr4g
verified it rather than adopting it: **`seq` is replica-local.** It equals
`origin_seq` today only because one replica authors events for all four machines,
so a second replica could order the same pair differently and derive a different
owed-set from the same log — use `hlc` (already on every event, total and
causally consistent) once `mesh_peers` reports any peer. Documented as reasoning,
not measurement, since a second replica cannot be stood up to test it. The part
worth keeping is the **asymmetry**: local-`seq` skew makes an answered item
*resurface* (noise, self-correcting, visible), while a timestamp comparison
*suppresses an unanswered ask forever* (silent, permanent) — so anyone tempted to
"fix" a resurfacing item with `created_at` would be trading the safe failure for
the dangerous one.
**Also carried, previously undocumented:** `dbb7879` resynced the vendored
`mempalace` snapshot to skillset `c04cd15` ("the withdrawal only holds where the
bridge is live"), landed after the v1.8.7 tag and so absent from that image.
⚠️ **A base rebuild is forced** (~67 min): both that resync and the
`pi-devbox-version` / `entrypoint-user.sh` changes below touch inputs to
`base_tag` (`rootfs/` and `entrypoint*.sh`). The provenance recording itself
adds nothing to that cost — it lives entirely in `Dockerfile.variant`.
Both come from one finding, made while verifying v1.8.7 from inside a freshly
recreated container: **the baked `mempalace` snapshot is read by no host on this
fleet.** `~/.agents/skills/mempalace` is a symlink to `/workspace/skillset/skills/mempalace`
— `entrypoint-user.sh` links the baked skill only `if [ ! -e ]`, and
`devbox-skill-reconcile` then repoints the skillset-owned ones at the live clone
(that is the v1.8.5 fix working as designed). All four compose stacks in
`docker-compose-repo` mount a workspace containing the skillset, so the vendored
copy is a CI/no-mount **fallback** and nothing else. Which means the
`mempalace skill snapshot is current` canary — the assertion that blocked
v1.8.7's first tag — polices a file that no agent on this fleet ever opens,
while the drift that *could* actually mislead an agent (a `git pull` nobody ran
in `/workspace/skillset`) was invisible from inside the container and is
invisible to CI by construction.
**The rejected fix is worth recording, because it was the obvious one.** The
old comment in `scripts/smoke-test.sh` said the real answer was "a CI job
diffing this file against the skillset repo". It isn't:
| Objection | Detail |
|---|---|
| needs a credential CI does not have | the skillset is **private** (`ssh://git@gitea.jordbo.se:2222/joakimp/skillset.git`); every build-time clone in this image uses anonymous HTTPS, and `resolve-versions`' `gitea_sha()` is explicitly documented as public-repo-only — its 401/403 path exists to survive a *stale token against a public repo*, so a private 403 would return empty and `require_sha` would hard-abort the release |
| makes another repo's branch able to fail this build | the same pi-devbox commit would go green today and red tomorrow, and a release could be blocked by an edit in an unrelated repo — precisely the shape of the run 589 failure, but automated and permanent |
| pure churn, and it is measurable | pi@emb-7kj4vr4g pushed **four** skillset commits in one evening (`d9dbbbd`, `b740d51`, `3324bd0`, `c04cd15`); a byte-parity gate would have demanded a pi-devbox resync commit **and a ~67-minute base rebuild for each one**, to keep current a copy almost nobody resolves |
| guards the wrong artefact | see above: on this fleet, nobody reads it |
**The invariant is not currency, it is non-contradiction** — the framing comes
from pi@emb-7kj4vr4g's review (logstream `project/pi-devbox`, correlation
`skillset-vendor-drift`, which also **retracted** its own earlier build-time
byte-compare recommendation). A stale-but-self-consistent fallback is harmless;
a stale fallback carrying a **withdrawn instruction** is a live footgun, and
this project has already paid for that one — through v1.8.4 the baked snapshot
*shadowed* the live clone, which is how superseded attribution guidance kept
reaching agents. That is precisely what the bidirectional canary asserts, and
why it stays.
So provenance is **recorded** rather than policed, and the check moves to where
the skillset actually is — a maintainer's clone, or any running container.
### Added
- **`build-manifest.json` now records the vendored snapshot's provenance:
`skillset_snapshot_ref` (which skillset commit the bytes are claimed to come
from) and `skillset_snapshot_sha256` (the bytes that actually shipped).** The
ref is a plain `ARG` **default in `Dockerfile.variant`**, deliberately not a
CI-resolved output, which buys three things at once: it needs no credential
for a private repo; it keeps a local `docker build` and CI identical by
construction (the same reasoning that put `MEMPALACE_VERSION` in
`Dockerfile.base` rather than duplicating it in the workflow); and it requires
**no change at any of the four `Dockerfile.variant` call sites** (`smoke`,
`smoke-studio`, `build-variant`, `build-variant-studio`), whose `--build-arg`
lists are hand-duplicated and therefore easy to under-apply to only two.
Also emitted as OCI label `se.jordbo.pi-devbox.skillset-snapshot-ref`, so it
is readable off the registry without pulling the image.
Two design points, each arrived at from the file's own rules:
- **The ref is a claim; the hash is measured.** `Dockerfile.variant` writes
the manifest from ground truth (`rev()` on each `/opt` clone, the live
`pi --version`), so the snapshot hash is computed with `sha256sum` in that
same layer rather than passed in. A build where the two disagree is exactly
what the new smoke assertions catch.
- **They are siblings, not members of `components{}`.** That map means "HEAD
of a clone present in this image" and the skillset is not cloned here —
calling it a component would be a lie a future reader would act on. It is
also load-bearing mechanically: `pi-devbox-version` renders every
`components{}` value with `.value[0:12]`, which would truncate a 64-hex
digest into something that looks like a short commit. Same reasoning as
`mempalace_version`'s existing comment.
⚠️ **Costs no base rebuild.** `base_tag` hashes `Dockerfile.base` + `rootfs/`
+ `entrypoint*.sh` + the mempalace-toolkit SHA; `Dockerfile.variant` is in
none of it. `scripts/check-base-hash.sh` scans `Dockerfile.base` **only**
(`DF="Dockerfile.base"`, single hardcoded path), so a new `*_REF` ARG in the
variant is invisible to that guard — correctly, since it changes nothing
about the base's contents.
- **`pi-devbox-version` gained a `skills:` section** reporting, per vendored
skill, whether the live copy is `baked` or a `live <repo> @ <sha>` clone —
and for `mempalace`, whether that live copy matches the baked fingerprint:
`(identical to baked snapshot)`, `(baked snapshot <ref> + uncommitted edits)`
when the clone is at the recorded commit but the bytes differ, or
`(baked snapshot <ref> — live copy differs)`. Same live-vs-baked shape as the
existing `pi:`/`palace:` drift annotations. **This is the check CI cannot do
and a container can, for free**, since every host that matters already has the
skillset mounted. The list iterates the baked tree rather than a hardcoded
name list, so vendoring a fourth skill needs no edit here.
`entrypoint-user.sh` calls it with the new **`--no-skills`** flag: the banner
is printed FIRST, before the baked links exist and long before the skillset
deploy and reconcile run last, so anything it said about skill sources would
describe a state that is about to change. Wrong-but-plausible is worse than
absent. (This is the one part of the change that touches `rootfs/` and
`entrypoint-user.sh`, so it does cost a base rebuild — already sunk, since
`dbb7879` refreshed the vendored snapshot.)
- **`scripts/vendor-mempalace-skill.sh`** — refreshes the snapshot and rewrites
the recorded ref *together*, because a `cp` without a matching ARG bump
produces a manifest that confidently lies, which is worse than the anonymous
snapshot it replaced. Refuses to record a ref when the upstream file has
uncommitted modifications (no commit describes those bytes, so recording one
would be a fabrication) — checked on that one file, not the whole tree, so
unrelated work in progress in the skillset does not block a vendoring.
`--check` answers "is the committed snapshot really `skillset@<recorded
ref>`?" and separately reports staleness against the clone's HEAD.
Counterfactual-tested rather than reasoned about, against throwaway clones:
a tampered snapshot reports `MISMATCH` **and** `STALE` (rc 1); a ref rolled
back to the previous skillset commit reports `MISMATCH` with content
unchanged (rc 1) and a subsequent refresh fixes only the ref, leaving the
bytes alone; unstaged and staged-but-uncommitted upstream edits are refused
with distinct messages and the snapshot left byte-identical, i.e. the refusal
is atomic.
**Hardened after review** by pi@emb-7kj4vr4g, whose warning was that a resync
script must "write the ref it ACTUALLY copied from, or the provenance field
inherits the same class of bug the canary just had". The first draft copied
the working tree and guarded it with `git diff` — which says nothing about an
**untracked** file, and can be clean on a detached or behind checkout while
`HEAD` names something else. The snapshot is now *constructed* from
`git show HEAD:<path>`, so the recorded pair cannot be a lie by construction,
and the untracked case is refused explicitly (tested: it was the one input the
first draft would have silently recorded a false ref for). Both new scripts are
`bash -n` clean and `shellcheck -S error` clean — the gate v1.8.7 added.
### Fixed
- **Three stale in-repo markers, all the same failure class.** Two "Unreleased"
pointers — `scripts/smoke-test.sh` pointed the reader at "the Unreleased
changelog note", and the v1.8.6 correction at "the Unreleased entry above";
that section became the `## v1.8.7` heading at release time and neither
back-reference was updated. The third: `scripts/smoke-test.sh`'s own coverage
list still advertised "typst PDF engine for pandoc **(Unreleased)**", five
releases after typst shipped in v1.4.0. Same class as the canary they sit
next to: true when written, silently false at release, with nothing checking
them. The smoke comment now describes the mechanism that actually shipped
(and why the CI-diff idea it advertised was rejected); the changelog one names
v1.8.7; the typst line names v1.4.0.
### Not fixed, deliberately
- **CI still cannot tell you the vendored snapshot is behind `skillset` main.**
That needs a read-only deploy key for a private repo threaded into
`resolve-versions`, to warn about a file no host on this fleet reads. Revisit
when a no-skillset container becomes a real deployment (shipping the image
outside the fleet, or a CI-only agent) — at which point the honest gate is a
**warning**, matching the existing `PI_VERSION`/`MEMPALACE_VERSION` policy
(concreteness → error, newer-release-exists → warning), never a build
failure.
- **The phrase canary stays.** It is orthogonal and free: it pins *content*
where the new fields pin *provenance*, so it still catches a re-vendored
snapshot whose ref was bumped correctly but whose bytes came from the wrong
place — and, per the review above, asserting the **absence of withdrawn
guidance** is the half of it that earns its keep. Its comment now states the
limit instead of promising a fix.
- **v1.8.7's published image has no recorded ref**, and that is expected: the
field arrives here. Worth knowing when reading one, since the tag move
`ebd0de0` → `f645e66` means the published v1.8.7 carries a pre-`dbb7879`
snapshot, i.e. its baked mempalace skill lacks c04cd15's "confirm the bridge
actually stamps" caveat. Harmless — on v1.8.7 the bridge *is* live, so that
caveat self-retires, and every enrolled host reads the live clone anyway.
`pi-devbox-version` degrades quietly on such an image: no fingerprint, no
annotation, verified against the real v1.8.7 manifest.
### Documented
- **The fleet's cross-machine coordination, which was working and unwritten.**
The RFC 003 logstream has carried real work between hosts since 2026-08-18 —
patch handoff, design review, a v1→v2 supersede — and no document in this repo
or the toolkit said so. Written up in three places, split by what each is
authoritative for:
- `README.md` § *Cross-machine agent coordination* — what the **container**
needs: `MEMPALACE_REMOTE_URL` selects the shared palace, and
`MEMPALACE_PI_DEVICE` is what makes this machine *reachable* on the log,
because when every host is a thin client of one palace the stamped agent name
is the only thing distinguishing them. Set both or neither: a container
without the device var can read the log but is addressable by nobody.
- the skillset's `mempalace` skill (`6eb20af`, live on every host that mounts
the skillset, no rebuild needed) — the **norms**: a mailbox query at wake-up,
and the sender-declared ack contract, where a *directed* event with
`status="open"` is owed a reply and a `*` broadcast owes nothing. The
`status` filter earns its place by dropping broadcast noise — measured, an
unfiltered mailbox returned 5 events, 4 of them finished broadcasts from
eight days earlier — but that is **all** it does; it does not compute
owed-ness, and the version of this entry that claimed otherwise is withdrawn
above. Measured today, both machines: the raw filter returns 2 asks here and
1 there, **every one already answered**, while the derivation returns 0 for
both. Dropping noise and deciding what is owed are two different jobs.
- mempalace-toolkit `extensions/pi/README.md` (`e70bef2`) — the **mechanism**,
including that the bridge is *write-only* today (it stamps events going out
and never reads the log, so nothing in this image polls on the agent's
behalf), and that live SSE push is a palace-deployment question: the server
implements `GET /logstream/stream`, but a reverse proxy exposing only `/mcp`
makes it unreachable — verified by 404s against the real endpoint.
⚠️ **The snapshot was refreshed rather than left stale.** The skill edits landed
in the skillset (`5fd0d5c`, then `6eb20af`), so `SKILLSET_SNAPSHOT_REF` was
resynced to match and `scripts/vendor-mempalace-skill.sh --check` is a clean
`OK` with no notice: the no-clone fallback carries the **corrected** protocol,
not the withdrawn one. That mattered more than currency usually does, because
the superseded copy contained an instruction — the `mesh_peers` gate — that
actively told a reader to skip the feature. Refreshing remains a deliberate
release-day decision rather than an automatic one: it costs a base rebuild, and
skipping it is legitimate because every enrolled host reads its live clone.
What is not legitimate is skipping it *silently*, which is what the new manifest
fields and `pi-devbox-version` output make impossible — hence step 2 in
`AGENTS.md` § *Release-day checklist*. In this release the refresh was free:
`rootfs/` was already changing, so the base rebuild was forced anyway.
---
## v1.8.7 — 2026-08-25
Patch release, and the fastest turnaround in the series (~9 h after v1.8.6) for
@@ -457,7 +1432,7 @@ and did not require a toolkit-side change.
**CORRECTION (2026-08-25, post-tag):** this bullet is wrong and was never
true of the tagged tree. `pi-devbox-version` *does* print a `palace:` line
in human mode, with live-vs-baked drift detection, degrading quietly on
pre-v1.8.6 manifests. Nothing is open here. See the Unreleased entry above.
pre-v1.8.6 manifests. Nothing is open here. See the v1.8.7 entry above.
**Resolved during this release, not left open:** the feeder `--agent`
default behavioural hook initially looked like it might need a
+70 -1
View File
@@ -277,6 +277,36 @@ ARG SOURCE_REVISION=
# MEMPALACE_TOOLKIT_REF is consumed in Dockerfile.base; re-declared here
# only so its intended ref lands in the label set alongside the others.
ARG MEMPALACE_TOOLKIT_REF=main
# ── Vendored skill provenance ─────────────────────────────────────────
# The vendored mempalace SKILL.md is the ONLY baked artefact with no /opt
# clone behind it: its upstream (the skillset repo) is PRIVATE, so the
# image cannot clone it and CI cannot resolve its HEAD (see VENDORED.md).
# Consequence through v1.8.7: the snapshot was ANONYMOUS — nothing in the
# image or the repo recorded which skillset commit it was taken from, so
# the only staleness check available was a hand-maintained phrase canary in
# scripts/smoke-test.sh, which by construction can only detect "older than
# what I remembered to pin", never "older than skillset main".
#
# Recording the ref costs nothing and makes the question answerable. It is
# deliberately a plain ARG DEFAULT rather than a CI-resolved output:
# * the value is a fact about the committed snapshot, so it belongs in
# the tree next to it — not in a workflow that a local `docker build`
# never runs (same reasoning as MEMPALACE_VERSION living in
# Dockerfile.base rather than being duplicated in docker-publish.yml);
# * CI therefore needs NO new build-arg at any of its four
# Dockerfile.variant call sites (smoke, smoke-studio, build-variant,
# build-variant-studio) — a plumbing change that is easy to
# under-apply to only two of them;
# * and it needs no credential for a private repo.
# Bump it with scripts/vendor-mempalace-skill.sh, which refreshes the file
# and rewrites this line together, so the pair cannot drift apart by hand.
# This ARG lives in Dockerfile.variant ON PURPOSE: Dockerfile.base and
# rootfs/ are both hashed into base_tag, so recording provenance here costs
# no ~67-minute base rebuild. (scripts/check-base-hash.sh scans only
# Dockerfile.base, so no folding into the base hash is required — nor would
# it be correct, since this ARG changes nothing about the base's contents.)
ARG SKILLSET_SNAPSHOT_REF=a12fe5ecc71e60feb24791e3e33571105f1afba7
# Dockerfile.base sets description="pi-devbox — base image (variant-independent)"
# and every variant INHERITS it, so both published images used to advertise
# themselves on Docker Hub as the base image. A LABEL cannot branch on
@@ -301,7 +331,8 @@ LABEL org.opencontainers.image.version="${RELEASE_TAG}" \
se.jordbo.pi-devbox.pi-atelier-version="${PI_ATELIER_VERSION}" \
se.jordbo.pi-devbox.mempalace-toolkit-ref="${MEMPALACE_TOOLKIT_REF}" \
se.jordbo.pi-devbox.pi-studio-ref="${PI_STUDIO_REF}" \
se.jordbo.pi-devbox.pi-studio-version="${PI_STUDIO_VERSION}"
se.jordbo.pi-devbox.pi-studio-version="${PI_STUDIO_VERSION}" \
se.jordbo.pi-devbox.skillset-snapshot-ref="${SKILLSET_SNAPSHOT_REF}"
# The manifest is written from GROUND TRUTH — the actual checked-out HEAD
# of each /opt clone and the live `pi --version` — not merely the intended
@@ -327,6 +358,34 @@ RUN set -e; \
case "$MP_V" in [0-9]*) MP_CORE="\"${MP_V}\"" ;; *) MP_CORE='null' ;; esac; \
STUDIO_REV='null'; \
if [ -d /opt/pi-studio/.git ]; then STUDIO_REV="\"$(rev /opt/pi-studio)\""; fi; \
# The vendored skill snapshot's fingerprint is MEASURED here, not passed
# in as a build-arg, per the ground-truth rule above: SKILLSET_SNAPSHOT_REF
# is a CLAIM about which skillset commit the file came from, while this
# hash is what the image actually ships. Recorded together they let any
# reader with the skillset checked out — which on this fleet is every
# host, since all four compose stacks mount it — verify the claim at
# RUNTIME, without CI ever needing access to the private repo. Degrades
# to JSON null rather than failing the build if the directory is absent;
# the smoke assertion is what turns that into a loud failure.
#
# Hashes the whole DIRECTORY, not just SKILL.md: a single-file hash
# answers "did this one file change", not "is the live copy the same
# skill" — a live checkout that added or edited a SIBLING file (a
# reference/ doc, a helper script) would still report "identical to
# baked snapshot" against a file-only hash. pi-extensions already ships
# two files for exactly this reason (SKILL.md + evaluate-extension-usage.py),
# so this is not a hypothetical. Deterministic over `find | sort`, never
# readdir order: relative paths + per-file sha256, folded into one hash.
# pi-devbox-version mirrors this exact pipeline over the live directory so
# the two sides are comparable — if you change this, change that too.
tree_sha256() { \
( cd "$1" && find . -type f -print | LC_ALL=C sort | xargs -r sha256sum ) 2>/dev/null | sha256sum | cut -d' ' -f1; \
}; \
SKILL_SNAP='null'; \
_snap_dir=/usr/local/share/pi-devbox/skills/mempalace; \
if [ -d "$_snap_dir" ] && [ -n "$(find "$_snap_dir" -type f -print -quit)" ]; then \
SKILL_SNAP="\"$(tree_sha256 "$_snap_dir")\""; \
fi; \
{ \
echo '{'; \
echo " \"release_tag\": \"${RELEASE_TAG}\","; \
@@ -337,6 +396,16 @@ RUN set -e; \
# SHAs and `pi-devbox-version` renders it with .value[0:12], which would
# silently truncate a longer version string.
echo " \"mempalace_version\": ${MP_CORE},"; \
# Siblings, NOT members of components{}, for two independent reasons:
# that map means "HEAD of a clone present in this image" and the
# skillset is not cloned here (calling it a component would be a
# lie a future reader would act on), and `pi-devbox-version` renders
# every components{} value with .value[0:12] — which would truncate
# a 64-hex sha256 into something that looks like a short commit.
# Named `_tree_sha256`, not `_sha256`: it measures every file under the
# vendored skill directory, not one file — see tree_sha256() above.
echo " \"skillset_snapshot_ref\": \"${SKILLSET_SNAPSHOT_REF}\","; \
echo " \"skillset_snapshot_tree_sha256\": ${SKILL_SNAP},"; \
echo " \"components\": {"; \
echo " \"pi-toolkit\": \"$(rev /opt/pi-toolkit)\","; \
echo " \"pi-extensions\": \"$(rev /opt/pi-extensions)\","; \
+114 -3
View File
@@ -20,7 +20,9 @@ on the host.
- `pi-extensions` — TypeScript extensions for pi (preview, MCP bridges,
mempalace integration, etc.)
- `pi-fork` — the `fork` tool for spawning sub-agents
- `pi-observational-memory` — the `recall` tool for session compaction
- `pi-observational-memory` — durable session memory: the ledger that makes
compaction cheap, plus the `recall` tool. See
[`docs/observational-memory.md`](docs/observational-memory.md)
- `pi-atelier` — TUI sidebar: ordered panels, split-pane, themes. Pinned to an
audited tag; see [Version pins](#version-pins-pi-pi-atelier-mempalace)
@@ -536,6 +538,35 @@ to refresh.
Anything not on a volume is on the writable layer and is lost on
container recreate.
### Rebuilding ephemeral shell state at start
Two entrypoint steps put back the kind of state that the writable layer eats, so a
recreate does not cost you a manual re-install:
- **`cli_utils` commands.** If a `cli_utils` checkout is mounted, every
executable in its `bin/` is symlinked into `~/.local/bin` on start, so
`git-status-all` and friends are on `PATH` without a path prefix. Detection:
`CLI_UTILS_CONTAINER_PATH` → `/workspace/cli_utils` → `$HOME/cli_utils` →
`/workspace/*/cli_utils`. Set `CLI_UTILS_LINK=0` to disable. Existing real files
in `~/.local/bin` and symlinks pointing elsewhere are left alone, so a
deliberate override still wins; links whose target disappeared are pruned.
Do **not** run a host installer's `install.sh` inside the container to achieve
this — it writes to the ephemeral home and dies on the next recreate.
- **A per-device boot hook.** If `~/.config/devbox-shell/init.sh` exists it is run
once at start (`bash`, never sourced, exit status ignored), with output in
`~/.pi/agent/devbox-init.log`. `~/.config/devbox-shell/` is the host-owned
bind-mount whose `bash_aliases` is already sourced into every interactive shell,
so a hook there persists across recreates with no image change. Use it for
fixups that must exist *before any shell* — symlinks, directories, one-off
migrations.
The distinction that decides which mechanism you want: `~/.local/bin` is on `ENV
PATH`, so symlinks there work in **non-interactive** shells too (`docker exec <c>
<cmd>`, agent tool shells, scripts). A `PATH` edit in `bash_aliases` reaches only
*interactive* shells, because `~/.bashrc` returns early when non-interactive —
which is also why shell **functions** (fzf helpers and the like) can only come
from the sourced file, never from a symlink.
## MemPalace integration
MemPalace is installed in the base image and pre-warmed with the
@@ -555,6 +586,75 @@ session/docs mining; the 29 MCP tools (search, kg-query, drawer-add,
diary-write, etc.) are wired into pi automatically by the pi-extensions
mempalace bridge.
### Cross-machine agent coordination
When `MEMPALACE_REMOTE_URL` points at a *shared* palace, the container gets more
than shared search: it joins an append-only coordination log (RFC 003) that other
machines' agents can address it on — used here for design review, patch handoff
and retraction between hosts.
Two container-side settings make it work:
| Variable | Why it matters |
|---|---|
| `MEMPALACE_REMOTE_URL` | selects the shared palace; unset means a purely local palace, and the log then contains only this machine's own events |
| `MEMPALACE_PI_DEVICE` | the bridge stamps `pi@<device>` as the writer, which is the **only** way the log can tell two machines apart when both are thin clients of one palace |
So a container with no `MEMPALACE_PI_DEVICE` can read the log but is not
reachable *on* it: messages addressed to a bare `pi` match nobody. Set both, or
neither.
What the agent is expected to *do* with this lives in the mempalace skill
(`~/.agents/skills/mempalace/SKILL.md`) — the mailbox query at wake-up, and the
convention that a directed event with `status="open"` is a request owed a reply
while a `*` broadcast owes nothing. The mechanism side (what the bridge stamps,
and why live SSE push depends on the palace deployment's reverse proxy rather
than on this image) is documented in the toolkit's `extensions/pi/README.md`.
**Since v1.8.9 the bridge reads the log for you.** Earlier images were write-only
— they stamped provenance on the way out and never read back, so a directed ask
reached an agent only if that agent happened to run `mempalace_event_list`
itself. The mailbox is gated on the same two variables as the stamper, is on by
default, and derives what is *owed* rather than trusting `status` (an acked event
keeps matching a `status="open"` query forever, because the log is append-only):
| Variable | Default | Effect |
|---|---|---|
| `MEMPALACE_MAILBOX` | unset (on) | `0` disables mailbox reads entirely |
| `MEMPALACE_MAILBOX_POLL_MS` | `300000` | minimum gap between mid-session polls |
| `MEMPALACE_MAILBOX_RESURFACE_MS` | `3600000` | re-announce a still-owed ask after this long |
Delivery **queues, it never interrupts**: the poll runs when pi goes idle and the
message is steered into the *next* turn, so nothing wakes the model on inbound
fleet traffic. The practical consequence, measured on two devices: the message
appears in your session window and the agent acts on it when the next turn
starts — you are the trigger. (That describes the bridge **as baked in v1.8.9**,
`mempalace-toolkit` `5b8d78f`; the mailbox's own mechanism and landmines live in
the toolkit's `docs/rfc-003-coordination-log.md` §7.11–§7.12, which moves ahead of
whatever this image has baked.)
## Observational memory (in-session memory)
The image also bakes [pi-observational-memory](https://github.com/elpapi42/pi-observational-memory),
which is memory of a *different kind* from the palace and is easy to confuse with
it. It keeps a small branch-local ledger of observations and reflections while a
session runs, so when pi compacts the conversation the summary is a
**deterministic fold of that ledger rather than a model call**, and every item
keeps a 12-character id that `recall(<id>)` resolves back to the exact source.
In one line: **observational memory keeps a session coherent; the palace keeps
the fleet coherent.**
It is on by default, needs no habit from you, and sends its background work to a
cheaper model than your session (Haiku while the session runs Opus, in the seeded
`~/.pi/agent/settings.json`). Inspect it from inside pi with `/om:status` and
`/om:view`; turn all proactive work off for one run with
`PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi`.
What it is for, how the lifecycle works, what it costs, every setting and its
default, and how it differs from MemPalace:
[`docs/observational-memory.md`](docs/observational-memory.md).
## Agent skills
pi discovers skills under `~/.agents/skills/`. Two delivery paths feed that
@@ -991,10 +1091,21 @@ After `docker compose up -d --force-recreate`, run the **runtime** peer of
persisted volumes survived, and pi runtime wiring is intact:
```bash
./scripts/recreate-sanity-check.sh # auto-detects variant
./scripts/recreate-sanity-check.sh --expected-version 0.79.4 # assert pi version
./scripts/recreate-sanity-check.sh # auto-detects variant
./scripts/recreate-sanity-check.sh --expected-image-version 1.8.9 # assert the pi-devbox release tag
./scripts/recreate-sanity-check.sh --expected-version 0.84.3 # assert the pi coding agent version
```
Those are **two different versions**, and the flags are not interchangeable:
`--expected-image-version` takes the pi-devbox release tag (`v` optional),
`--expected-version` takes `pi --version`. Hand one the other's value and it
says so by name instead of reporting a mismatch against the wrong component.
With neither flag, both values are read from the image's own build manifest
(`/etc/pi-devbox/build-manifest.json`): the live pi version is asserted against
the one recorded at build time — which catches a stale `pi` in the
`~/.pi/npm-global` volume shadowing the baked one — and the release tag is
reported informationally.
If `cli_utils` is on your PATH, the `pi-devbox-sanity` wrapper runs the same
check by short name and locates the repo automatically (override with
`PI_DEVBOX_REPO=/path/to/pi-devbox`). Like `smoke-test.sh`, this script is
+371
View File
@@ -0,0 +1,371 @@
# Observational memory — why this image has it, and what it does for you
**Audience:** anyone using this container for long pi sessions who has wondered
what `recall`, `/om:status` and "compacted memory" are, or whether they should
leave any of it switched on.
**Companion documents:** the extension ships its own reference docs at
`/opt/pi-observational-memory/docs/` —
[`concepts.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/concepts.md)
(the model),
[`how-it-works.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/how-it-works.md)
(hooks and internals) and
[`configuration.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/configuration.md)
(every setting). Pi's own compaction mechanics are in
`/usr/lib/node_modules/@earendil-works/pi-coding-agent/docs/compaction.md`.
Those are normative; this document is the **deployment** view — what is pinned
here, how it is wired, what it costs, and how it differs from MemPalace. For the
palace, see
[`mempalace-toolkit/docs/fleet-memory.md`](https://gitea.jordbo.se/joakimp/mempalace-toolkit/src/branch/main/docs/fleet-memory.md).
> Verified on pi-devbox **v1.8.9** (`release_tag v1.8.9`, source `aac4a1c`),
> which bakes pi-observational-memory **v3.0.4** at commit `ce9fc98` — the value
> in `/etc/pi-devbox/build-manifest.json` → `components.pi-observational-memory`.
> Every number below was read from that tree, from pi 0.84.3's own docs, or from
> the live container.
---
## 1. The problem it solves
A long pi session outgrows the model's context window. Pi's answer is
**compaction**: fold the older part of the conversation into a summary and keep
recent messages verbatim. That is unavoidable, and it is where sessions go
wrong — the summary is produced *at the moment of pressure*, by a model, about a
transcript that is about to leave the context.
Observational memory changes *when* the remembering happens. Instead of
summarising in a panic at the end, it keeps a small **ledger** up to date while
the session runs, and compaction then just folds that ledger.
```mermaid
flowchart LR
A0["plain compaction"] --> A1["context fills"]
A1 --> A2["a model summarises<br/>under pressure"]
A2 --> A3["prose summary,<br/>no way back"]
B0["with observational<br/>memory"] --> B1["context fills"]
B1 --> B2["ledger written<br/>as you work"]
B2 --> B3["compaction folds<br/>the ledger"]
B3 --> B4["ids you can<br/>recall"]
```
Top row is pi on its own: one model call at the worst possible moment, detail
chosen in a hurry, and the original wording gone from view. Bottom row is this
image's default: the thinking happened earlier on a cheap model, the fold is
deterministic, and every line in the result carries an id that resolves back to
the exact source.
## 2. The mental model: three layers and a ledger
| Layer | What it is | Example |
|---|---|---|
| **Observation** | a timestamped, source-backed event from the conversation | "user rejected option B because it needs a base rebuild" |
| **Reflection** | a durable conclusion *backed by* observations | "the user optimises for avoiding 67-minute rebuilds" |
| **Drop** | a tombstone retiring an observation from active memory | the superseded detail of a bug that is now fixed |
These are appended to the session as silent ledger entries
(`om.observations.recorded`, `om.reflections.recorded`,
`om.observations.dropped`) and **folded** — replayed in order — to produce the
memory state. The ledger is the source of truth; what you see in a compacted
session is a rendering of it.
Two properties follow, and both matter later:
- **The ledger itself costs no context.** Those entries are pi `custom` entries,
which *"do not participate in LLM context"* (pi `docs/session-format.md`). They
sit in the session file and reach the model only via the fold at compaction.
- **Memory is branch-local.** A pi session is a tree (resume, fork), and the fold
follows the current branch only, so a forked branch does not inherit another
branch's view.
## 3. The lifecycle
Three background workers and one compaction hook, driven by *raw token
progress* rather than wall-clock time. Defaults in brackets.
```mermaid
flowchart TD
T(["turn_end"]) --> O{"10k raw tokens<br/>since observing?"}
O -- yes --> OBS["<b>observer</b> runs"]
O -- "no" --> R{"20k tokens<br/>since reflecting?"}
R -- yes --> REF["<b>reflector</b> runs"]
REF -- "if pool over 10k" --> DR["<b>dropper</b> prunes"]
S(["agent_settled"]) --> C{"81k tokens<br/>since compacting?"}
C -- yes --> CP["ctx.compact()"]
CP --> H(["session_before_compact"])
H --> F["fold the ledger<br/>no model call"]
F --> VIS["compacted memory"]
```
- **observer** — `observeAfterTokens` [10000]: writes observations for the
conversation it has not covered yet.
- **reflector** — `reflectAfterTokens` [20000]: promotes patterns across
observations into durable reflections.
- **dropper** — no clock of its own. It is post-reflection maintenance, gated on
a *successful same-turn* reflection **and** an active pool above
`observationsPoolTargetTokens` [10000]. Not a third worker on a third
threshold.
- **compaction** — `compactAfterTokens` [81000], checked when pi goes idle, so it
never interrupts a turn. Pi will also compact on its own when the context is
nearly full (`contextTokens > contextWindow - reserveTokens`, `reserveTokens`
[16384]).
## 4. What compaction actually does to your context
This is the question the rest of the document used to leave hanging: if the old
conversation is folded away, is the session back to knowing nothing?
**No.** Compaction replaces *part* of the context, not all of it, and it deletes
nothing at all from disk.
```mermaid
flowchart LR
SYS["system prompt<br/>+ AGENTS.md"] --> CTX["what the model sees<br/>on the next turn"]
SUM["folded memory:<br/>reflections + observations"] --> CTX
TAIL["recent turns,<br/>verbatim"] --> CTX
DISK[("session .jsonl: all of it")] -. "recall(id)" .-> CTX
```
Where each piece comes from:
- **System prompt and `AGENTS.md` — never compacted, because they were never
conversation.** Pi rebuilds them from disk on every request
(`loadContextFileFromDir`), so they cannot be lost by compaction.
- **The verbatim tail — sized by a token budget, not a message count.** Pi walks
backwards from the newest entry accumulating token estimates until
`keepRecentTokens` [20000] is reached; that entry becomes `firstKeptEntryId`,
and *everything from there on is kept unchanged*. Cut points land on turn
boundaries, never mid-tool-call. So the most recent ~20k tokens of real work —
your last instructions, the diffs, the test output — survive word for word.
- **The folded memory — replaces only what came before that cut.** Rendered from
the ledger's records: reflections and observations, each with its 12-hex id.
- **The session file — untouched.** Compaction *appends* a `compaction` entry
(`{"type":"compaction", summary, firstKeptEntryId, tokensBefore, …}`) and
rebuilds context from it on later turns. Nothing is rewritten in place; the
only documented way to remove session content is deleting the whole `.jsonl`.
That last point is what makes the answer to "is the detail gone?" *no* rather
than *mostly*: `recall` does not read the context window at all. It calls
`sessionManager.getBranch()` — the full branch from the root — and resolves an
observation id back to the original entries. Detail that left the model's view
an hour ago is still one `recall` away.
**Repeated compaction does not summarise the summary.** The rendered text is
always built from live observation/reflection *records*, never from the previous
compaction's prose, so there is no generation-loss spiral. (Mechanically the
projection is incremental — it re-derives back to the last full-fold boundary and
carries the rest forward, escalating to a genuine re-fold from the branch root
when the observation pool reaches `observationsPoolMaxTokens` [20000].)
So the honest summary of the state after compaction: **the model keeps its
instructions, keeps recent work verbatim, trades older turns for a dense
id-carrying digest of them, and can pull any of it back on demand.** Not a fresh
start — a smaller, cheaper, still-navigable one.
### One caveat about "no model call"
If the ledger is empty — compaction fires before the observer has ever run — the
hook returns nothing and *declines ownership*, and pi's own model-based
summariser runs instead:
```ts
const summary = renderSummary(projection.reflections, projection.observations);
if (summary.length === 0) {
// Decline ownership so Pi's native summarizer preserves the pre-cut context.
return;
}
```
In steady state (any session old enough to have produced one observation) om's
hook wins and compaction is model-free. "Never calls a model" is true in practice
and false in principle; the fallback is deliberate, so an empty ledger degrades
to normal pi rather than to no summary at all.
## 5. What you actually get
- **Compaction stops being a stall.** In steady state the latency path is
deterministic work over ledger entries, not a summarisation call.
- **Nothing important vanishes silently.** Compaction is lossy by design, but
every item keeps a 12-character id, and `recall(<id>)` returns the exact
evidence — original wording, reasoning, file path, error text.
- **The bookkeeping runs on a cheaper model than your session.** In this image
that is deliberate and visible (§7): background workers on Haiku, session on
Opus.
- **It is automatic.** No habit to maintain, unlike the palace protocol — which is
exactly why the two complement each other (§11).
- **Forks stay clean.** Branch-local memory means a `fork` sub-agent's noise does
not leak into the parent's folded memory.
## 6. `recall` is not a search tool
`recall` takes **one specific 12-hex id** that already appears in compacted
memory or in `/om:view`. It cannot be given a topic. It can return an observation
(marked `active` or `dropped`), or a reflection together with the observations
supporting it.
```mermaid
sequenceDiagram
participant M as compacted memory
participant A as agent
participant L as ledger
M->>A: "[high] user rejected option B (a1b2c3d4e5f6)"
A->>L: recall("a1b2c3d4e5f6")
L-->>A: exact observation + source ids
Note over A: acts on the original wording
```
The rule of thumb the agent skill uses: recall **before a load-bearing action**
that rests on a compressed memory — shipping a change, asserting a fact,
answering "why do you believe that". One recall is cheap; redoing finished work
is not.
## 7. How it is wired in this image
```mermaid
flowchart TB
IMG["baked in the image:<br/>v3.0.4 @ ce9fc98"] --> REG["settings.json<br/>packages[]"]
REG --> SESS["your pi session"]
SESS -- "your turns" --> SM["session model:<br/>Opus"]
SESS -- "observer, reflector,<br/>dropper" --> WM["memory model:<br/>Haiku"]
SESS -- "ledger entries" --> JL["session .jsonl"]
JL --> VOL[("devbox-pi-config<br/>volume")]
```
Four consequences of that wiring:
1. **There is no separate database.** Memory *is* entries inside the ordinary pi
session file (`~/.pi/agent/sessions/<project>/<timestamp>_<uuid>.jsonl`).
Nothing extra to back up, nothing to migrate.
2. **It survives container recreate**, because `~/.pi` is the `devbox-pi-config`
named volume (`docker-compose.yml`) — the same one holding your pi config and
session history.
3. **`packages[]` is the only source of truth for which copy is loaded.** A clone
at `/workspace/pi-observational-memory` may exist (and today matches `/opt`
byte-for-byte at `ce9fc98`) — its presence proves nothing. To run a patched
build you point `packages[]` at it explicitly and start a new session.
4. **The worker model is a deliberate choice, and it is yours to change.** The
seeded config sends background work to Haiku while your session runs Opus:
```json
"observational-memory": {
"model": { "provider": "amazon-bedrock", "id": "eu.anthropic.claude-haiku-4-5-20251001-v1:0" },
"debugLog": false
}
```
## 8. What it costs
| Resource | Cost |
|---|---|
| Model calls | up to **three** background calls per consolidation pass (observer, reflector, dropper), each capped at `agentMaxTurns` [16], on the configured memory model — not your session model |
| Latency in your turns | none by construction: workers run from `turn_end`, compaction runs when pi is idle, and the fold itself does no model work |
| Disk | negligible — JSON lines inside a session file that would exist anyway (measured here: `~/.pi/agent/sessions` = 30 MB total, tens of `om.*` entries per session) |
| Context window | **zero until compaction.** `custom` entries do not enter LLM context; only the folded summary does |
| Attention | none once configured; there is no protocol for you or the agent to remember |
If that is still more than you want on a given run, §9's `passive` switch turns
off all proactive work while keeping `recall` and `/om:*` usable.
## 9. Configuration
Global: `~/.pi/agent/settings.json` (persisted in the volume). Per project:
`<project>/.pi/settings.json`, which overrides global. Precedence is
project → global → environment, and the environment can only override `passive`.
| Key | Default | What it changes |
|---|---|---|
| `observeAfterTokens` | `10000` | observer cadence — lower means smaller chunks and more calls |
| `reflectAfterTokens` | `20000` | reflector cadence (and thereby dropper opportunities) |
| `observerChunkMaxTokens` | 20% of the memory model's context window, else `60000` | cap on one observer run's input |
| `compactAfterTokens` | `81000` | when proactive auto-compaction fires |
| `observationsPoolMaxTokens` | `20000` | pool size at which compaction does a full re-fold from the branch root |
| `observationsPoolTargetTokens` | half of max (`10000`) | what the dropper aims back down to |
| `agentMaxTurns` | `16` | shared turn cap for the three workers |
| `model` | unset → session model | send background work to a cheaper/faster model |
| `showWorkerNotifications` | `true` | routine "observer ran" notices |
| `passive` | `false` | **kill switch** for all proactive background work; `recall` and `/om:*` still work |
| `debugLog` | `false` | per-session NDJSON trace at `~/.pi/agent/observational-memory/debug/<session-id>.ndjson` |
Pi's own compaction knobs live under a separate `compaction` key —
`keepRecentTokens` [20000] sets the verbatim tail from §4, `reserveTokens`
[16384] the headroom that triggers pi's own compaction.
One-off passive run, no config edit:
```bash
PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi
```
Invalid values are ignored rather than fatal, so a typo degrades to the default
instead of breaking your session — which also means a typo is silent. Check with
`/om:status`.
## 10. Confirming it is actually working
Do not infer health from the absence of a warning; look:
```bash
# 1. inside pi — the authoritative view
/om:status # visible-vs-full drift, thresholds, worker state
/om:view # what the agent currently sees
/om:view full # full ledger truth at the branch tip
# 2. from a shell — are ledger entries being written, and has it compacted?
grep -o '"customType":"om\.[a-z.]*"' \
"$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)" | sort | uniq -c
grep -c '"type":"compaction"' "$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)"
# 3. which copy is loaded, and at what commit
python3 -c "import json;print(json.load(open('$HOME/.pi/agent/settings.json'))['packages'])"
git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory rev-parse HEAD
```
Ledger entries are `"type":"custom"` with `"customType":"om.…"`. Do not grep for
`custom_message` — that is a *different* pi API for entries that **do** enter LLM
context, used here by the MemPalace mailbox (`customType: "mempalace-mailbox"`),
not by om.
## 11. It is not the same thing as MemPalace
Both are called "memory" and they solve different problems. Nothing is wrong with
running both — this image does, and they cover each other's failure modes.
```mermaid
flowchart LR
O0["observational<br/>memory"] --> O1["horizon:<br/>this session"]
O1 --> O2["scope: one branch,<br/>one machine"]
O2 --> O3["automatic"]
O3 --> O4["retrieval:<br/>recall(id)"]
P0["MemPalace"] --> P1["horizon: months,<br/>machines"]
P1 --> P2["scope:<br/>the fleet"]
P2 --> P3["protocol-driven"]
P3 --> P4["retrieval:<br/>search, KG, mailbox"]
```
| Question | Answer |
|---|---|
| "What did we decide 200 turns ago in *this* session?" | observational memory (and `recall` for the exact wording) |
| "What did we decide last month, or on another machine?" | MemPalace (`mempalace_search`, diaries) |
| "What is true *right now* about version X?" | MemPalace knowledge graph |
| "Does another machine need something from me?" | MemPalace coordination log — see [Cross-machine agent coordination](../README.md#cross-machine-agent-coordination) |
| "Why is compaction not losing my session?" | observational memory |
The crisp version: **observational memory keeps a session coherent; the palace
keeps the fleet coherent.** A container recreate wipes neither — but only because
`~/.pi` and the palace both live outside the container filesystem.
## 12. Gotchas
- **Branch-local means branch-local.** Resuming or forking changes which ledger
is folded. Memory that "disappeared" is usually on another branch.
- **`recall` needs an id, not a topic.** If you only have a topic, that is a
palace search, not a recall.
- **A `/workspace` clone is not evidence of what is loaded** — see §7.3.
- **`showWorkerNotifications: true` is not proof of work**; it reports runs, and
an observer that deliberately emits nothing writes no ledger entry and simply
retries after another `observeAfterTokens`.
- **A turn bigger than `keepRecentTokens` splits.** The cut then lands mid-turn at
an assistant message and pi merges two summaries — rare, but it is why a very
large single turn can lose more verbatim detail than you would expect.
- **`git log` in the baked tree needs `safe.directory`** (`/opt` is root-owned):
`git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory log`.
+110 -1
View File
@@ -7,7 +7,11 @@ set -euo pipefail
# so this reaches the same stream as the interactive shell the user lands
# in). Reads the ground-truth manifest baked in Dockerfile.variant; a no-op
# with a short stderr notice on images built before it existed.
command -v pi-devbox-version >/dev/null 2>&1 && pi-devbox-version || true
# `--no-skills`: this runs FIRST, before the baked skill links are created
# below and long before the skillset deploy + devbox-skill-reconcile run at the
# end of this script, so the skill-source section would report a pre-reconcile
# state that is about to change. Wrong-but-plausible is worse than absent.
command -v pi-devbox-version >/dev/null 2>&1 && pi-devbox-version --no-skills || true
# ── SSH ControlMaster socket dir ────────────────────────────────
# Companion to /etc/ssh/ssh_config.d/00-devbox-controlmaster.conf in the
@@ -184,6 +188,111 @@ if [ "${MEMPALACE_FEED:-1}" != "0" ] && [ -n "$MEMPALACE_FEEDER" ]; then
fi
fi
# ── cli_utils: link workspace bin/ commands onto PATH ────────────────
# Standalone commands from a mounted cli_utils checkout (git-status-all,
# git-pull-all, devbox-sanity, pi-session-repair, ...) live in <repo>/bin. On a
# host they reach PATH via cli_utils' own install.sh, whose install_bin step
# symlinks them into ~/.local/bin — but that home is on the container's WRITABLE
# LAYER, so every recreate loses them and the human is back to typing
# /workspace/cli_utils/bin/git-status-all. This is the container equivalent of
# that install step, re-run at every start.
#
# WHY SYMLINKS RATHER THAN A PATH EDIT IN AN rc FILE: ~/.local/bin is already
# ahead of /usr/local/bin in ENV PATH (Dockerfile.base), so links here resolve in
# NON-interactive shells too — `docker exec <c> git-status-all`, agent tool
# shells, scripts. An rc-file PATH edit cannot reach those, because ~/.bashrc
# returns early when the shell is not interactive. Measured 2026-08-27 on
# tor-ms22: `command -v git-status-all` failed in a non-interactive shell while
# working in an interactive one, from exactly that asymmetry.
#
# Detection order (first hit wins):
# 1. CLI_UTILS_CONTAINER_PATH explicit, for non-standard layouts
# 2. /workspace/cli_utils repo directly in the workspace root
# 3. $HOME/cli_utils dedicated mount
# 4. /workspace/*/cli_utils workspace root holds several repo groups
# CLI_UTILS_LINK=0 disables. Absent repo = silent no-op, which is the common
# case for anyone who does not use cli_utils.
if [ "${CLI_UTILS_LINK:-1}" != "0" ]; then
CLI_UTILS_BIN=""
if [ -n "${CLI_UTILS_CONTAINER_PATH:-}" ] && [ -d "${CLI_UTILS_CONTAINER_PATH}/bin" ]; then
CLI_UTILS_BIN="${CLI_UTILS_CONTAINER_PATH}/bin"
elif [ -d /workspace/cli_utils/bin ]; then
CLI_UTILS_BIN=/workspace/cli_utils/bin
elif [ -d "$HOME/cli_utils/bin" ]; then
CLI_UTILS_BIN="$HOME/cli_utils/bin"
else
# `if` bodies, not `&&` chains: under `set -e` a loop whose LAST command is a
# false test exits non-zero and would abort the entrypoint. With no match the
# glob stays literal, so that is the normal case on any machine without this
# repo — i.e. the bug would have been "container will not start", not "links
# missing".
for _cu in /workspace/*/cli_utils/bin; do
if [ -d "$_cu" ]; then
CLI_UTILS_BIN="$_cu"
break
fi
done
unset _cu
fi
if [ -n "$CLI_UTILS_BIN" ]; then
mkdir -p "$HOME/.local/bin" 2>/dev/null || true
# Never clobber a real file, and never steal a link that points elsewhere: a
# deliberate user override in ~/.local/bin must win, and silently shadowing
# an image-provided command is worse than the missing command.
for _f in "$CLI_UTILS_BIN"/*; do
if [ ! -f "$_f" ] || [ ! -x "$_f" ]; then
continue
fi
_link="$HOME/.local/bin/$(basename "$_f")"
if [ -e "$_link" ] && [ ! -L "$_link" ]; then
continue
fi
if [ -L "$_link" ]; then
case "$(readlink "$_link")" in
"$CLI_UTILS_BIN"/*) ;;
*) continue ;;
esac
fi
ln -sf "$_f" "$_link" 2>/dev/null || true
done
# Prune links we own whose target vanished (command renamed, repo moved),
# mirroring the skillset deploy's --prune-stale. A dangling link on PATH
# reports "No such file or directory" for a command that simply no longer
# exists, which reads as a broken container rather than a removed script.
for _link in "$HOME/.local/bin"/*; do
[ -L "$_link" ] || continue
case "$(readlink "$_link")" in
*/cli_utils/bin/*) [ -e "$_link" ] || rm -f "$_link" ;;
esac
done
unset _f _link
fi
unset CLI_UTILS_BIN
fi
# ── Per-device boot hook ─────────────────────────────────────────────
# Runs ~/.config/devbox-shell/init.sh if the host provides one. That directory is
# the host-owned, bind-mounted shell-sharing dir (see "Volumes and persistence"),
# so a hook placed there survives every recreate WITHOUT an image change — the
# boot-time twin of the interactive bridge in /etc/skel-devbox/.bash_aliases,
# which sources ~/.config/devbox-shell/bash_aliases for every interactive shell.
#
# NO NEW TRUST BOUNDARY: that same directory is already sourced into every
# interactive shell, i.e. it is already arbitrary code from the same owner. What
# is new is only WHEN it runs — once at start, before any shell — which is what
# non-interactive fixups (symlinks, dirs, one-off migrations) need.
#
# Deliberately `bash <file>`, not `.` — a hook must not be able to mutate this
# entrypoint's own shell state, and its exit status must not matter. Output goes
# to a log rather than the container's start output, so a chatty hook cannot
# masquerade as a startup error.
if [ -r "$HOME/.config/devbox-shell/init.sh" ]; then
mkdir -p "$HOME/.pi/agent" 2>/dev/null || true
bash "$HOME/.config/devbox-shell/init.sh" \
>"$HOME/.pi/agent/devbox-init.log" 2>&1 || true
fi
# ── Git config defaults ──────────────────────────────────────────────
if [ -n "${GIT_USER_NAME:-}" ] && ! git config --global user.name &>/dev/null; then
git config --global user.name "$GIT_USER_NAME"
+135 -8
View File
@@ -14,6 +14,8 @@
# pi-devbox-version human-readable summary (default)
# pi-devbox-version --json raw manifest JSON (for scripting)
# pi-devbox-version --quiet one-line "release_tag (source_revision)" form
# pi-devbox-version --no-skills skip the skill-source section (used at
# container start, where it would be premature)
#
# EXIT STATUS
# 0 on success. 1 if the manifest is missing (e.g. an image built before
@@ -24,15 +26,35 @@ set -euo pipefail
MANIFEST=/etc/pi-devbox/build-manifest.json
MODE="human"
SHOW_SKILLS="yes"
case "${1:-}" in
--json) MODE="json" ;;
--quiet|-q) MODE="quiet" ;;
--help|-h)
sed -n '2,20p' "$0" | sed 's/^# \?//'
exit 0
;;
esac
# A `case "${1:-}"` here only ever looked at the FIRST argument, so
# `--no-skills --json` matched --no-skills, silently dropped --json, and
# printed human text to a caller expecting JSON (a real failure: a jq
# consumer piping that output gets a parse error, not a wrong-but-parseable
# answer). Loop over every argument instead, and reject anything unknown
# rather than silently ignoring it the same way.
for _arg in "$@"; do
case "$_arg" in
--json) MODE="json" ;;
--quiet|-q) MODE="quiet" ;;
--no-skills) SHOW_SKILLS="no" ;;
--help|-h)
# Print the leading `#`-comment block verbatim, stopping at the first
# non-comment line, rather than a hardcoded line range: `sed -n
# '2,22p'` was silently truncating --help because this file has grown
# usage lines since that range was written, and a fixed range will
# drift again the next time a comment is added above it.
awk 'NR==1{next} /^#/{sub(/^# ?/,""); print; next} {exit}' "$0"
exit 0
;;
*)
echo "pi-devbox-version: unknown option: $_arg" >&2
echo " try --help" >&2
exit 2
;;
esac
done
if [ ! -f "$MANIFEST" ]; then
echo "pi-devbox-version: no build manifest at $MANIFEST" >&2
@@ -105,3 +127,108 @@ fi
printf ' components:\n'
jq -r '.components | to_entries[] | select(.value != null) | " \(.key): \(.value[0:12])"' "$MANIFEST"
# ── Which copy of each vendored skill is actually being read? ─────────
# The image bakes fallback skills under /usr/local/share/pi-devbox/skills/,
# but for skills the skillset repo OWNS (skillset-owned.txt) a mounted live
# clone takes over at container start via devbox-skill-reconcile. Nothing
# reported which copy won, so a stale baked snapshot and a current live clone
# looked identical from inside — and on this fleet the baked mempalace copy is
# read by NOBODY (all four compose stacks mount a workspace containing the
# skillset), which is exactly the sort of fact that should be visible rather
# than reasoned about. Same "drift detected" shape as the pi/palace lines
# above: what is live, annotated with what was baked, when they disagree.
#
# Skipped with --no-skills at container start (entrypoint-user.sh calls this
# FIRST, before the baked links exist and long before the skillset deploy and
# reconcile run last), because a section that is accurate only after boot
# finishes is worse than no section at all.
BAKED_SKILLS=/usr/local/share/pi-devbox/skills
SKILLS_DIR="${HOME:-/home/developer}/.agents/skills"
if [ "$SHOW_SKILLS" = "yes" ] && [ -d "$BAKED_SKILLS" ] && [ -d "$SKILLS_DIR" ]; then
# Recorded provenance of the vendored mempalace snapshot (absent on images
# built before this existed — `// empty` so a JSON null never prints as the
# 4-char string "null", the same trap noted for mempalace_version above).
# `_tree_sha256`, not `_sha256`: it is a hash over every file in the
# vendored skill DIRECTORY (see tree_sha256() below), not one file, because
# a single-file hash reports "identical" against a live checkout that added
# or edited a sibling file — pi-extensions already ships two files, so this
# is not hypothetical.
snap_ref=$(jq -r '.skillset_snapshot_ref // empty' "$MANIFEST")
snap_sha=$(jq -r '.skillset_snapshot_tree_sha256 // empty' "$MANIFEST")
# Same pipeline Dockerfile.variant uses to measure the baked directory at
# build time: relative paths in `find | sort` order, each hashed, the whole
# listing folded into one sha256. Keep the two definitions identical — they
# run in different processes (image build vs. this container) and are
# meaningless to compare unless they agree byte-for-byte on the algorithm.
tree_sha256() {
( cd "$1" && find . -type f -print | LC_ALL=C sort | xargs -r sha256sum ) 2>/dev/null | sha256sum | cut -d' ' -f1
}
# Iterate the baked tree rather than a hardcoded name list, so vendoring a
# fourth skill needs no edit here. The header prints only if the tree is
# non-empty, so this can never emit a dangling "skills:" label.
_printed_header="no"
for _dir in "$BAKED_SKILLS"/*/; do
[ -d "$_dir" ] || continue
if [ "$_printed_header" = "no" ]; then
printf ' skills:\n'
_printed_header="yes"
fi
_name=$(basename "$_dir")
_link="$SKILLS_DIR/$_name"
if [ ! -e "$_link" ]; then
printf ' %-22s not linked\n' "$_name"
continue
fi
_target=$(readlink -f "$_link" 2>/dev/null || echo "$_link")
case "$_target" in
"$BAKED_SKILLS"/*|"$BAKED_SKILLS")
printf ' %-22s baked\n' "$_name"
continue
;;
esac
# Outside the baked tree: a mounted skillset clone, or a user override.
# The link target is <repo>/skills/<name>, so the repo root is two up.
# Everything here is guarded: this script runs on the container-start path
# and must never fail, and `set -e` is in force.
_root=$(cd "$_target/../.." 2>/dev/null && pwd) || _root=""
_head=""
if [ -n "$_root" ]; then
_head=$(git -C "$_root" rev-parse HEAD 2>/dev/null || echo "")
fi
_where="live ${_root:-$_target}"
[ -n "$_head" ] && _where="$_where @ ${_head:0:7}"
# For the one skill whose baked fingerprint we recorded, say plainly
# whether the live copy differs from what shipped. This is the check CI
# cannot perform (the skillset is private) and the container can, free.
# Hash the whole live DIRECTORY with the same tree_sha256() used to
# measure the baked one in Dockerfile.variant — a SKILL.md-only compare
# would silently ignore a changed or added sibling file.
_live_sha=""
if [ -n "$snap_sha" ] && [ "$_name" = "mempalace" ] && [ -d "$_target" ]; then
_live_sha=$(tree_sha256 "$_target")
fi
if [ -z "$_live_sha" ]; then
printf ' %-22s %s\n' "$_name" "$_where"
elif [ "$_live_sha" = "$snap_sha" ]; then
printf ' %-22s %s (identical to baked snapshot)\n' "$_name" "$_where"
elif [ -n "$_head" ] && [ "$_head" = "$snap_ref" ]; then
# Same commit, different bytes — i.e. uncommitted edits in the live
# checkout. Distinguished from plain drift because otherwise the line
# reads as a self-contradiction ("@ c04cd15 ... baked snapshot c04cd15
# — live copy differs") and a reader would suspect the tool, not the
# working tree.
printf ' %-22s %s \033[33m(baked snapshot %s + uncommitted edits)\033[0m\n' \
"$_name" "$_where" "${snap_ref:0:7}"
else
printf ' %-22s %s \033[33m(baked snapshot %s — live copy differs)\033[0m\n' \
"$_name" "$_where" "${snap_ref:0:7}"
fi
done
fi
@@ -70,3 +70,41 @@ rather than merely confusing you:
local disk, so `mempalace search` can return older and different results than
the MCP tools while both look correct. Use the MCP tools for the central
palace; the CLI only for a local one.
## Before you file a finding: second measurement, different route
This is here rather than in a skill because it has to fire *without* a matching
task description, and because the version of it that lived only in a skill was
violated five times in one session by an agent that had the skill available.
**Any claim you are about to record as fact — in a drawer, a diary entry, a
coordination event, or a report to the user — needs a second measurement taken
by a different route.** Not a re-read of your reasoning: re-reading has caught
zero of these. A disagreeing measurement has caught all of them.
The two shapes that get filed as fact and are not:
- **A negative result** (`401`, connection refused, zero rows, "not found") is
first a claim about *your filter*, not about the world. Wrong host, wrong port,
wrong table, capped output.
- **A positive result** proves only what your command *actually asked*. An SSH
handshake can succeed against the wrong host (`ssh -G` tells you which rule
captured the name); a `401` can be a real answer from an issuer that never
minted the credential.
Cheapest habit that works: **write the expected result next to each check before
running it**, then diff. Expectations declared up front turn a silent wrong
assumption into a visible mismatch. And if you cannot think of a second route to
the same fact, you do not have a finding — you have a hypothesis, so label it as
one.
## Handling an exposed credential
If a task touches a leaked secret, a token rotation, "is this credential still
live?", whether to delete stored content, or which scopes a new token needs:
**read `~/.agents/skills/credential-incident-response/SKILL.md` first.** One rule
is load-bearing enough to state here: **probe the issuing provider before doing
anything else** — most "exposed" credentials in a long-lived fleet are already
dead, and the ones that are live are often far more privileged than assumed.
Severity first, cleanup second, and prefer **revocation over deletion** for
anything already replicated.
@@ -9,6 +9,7 @@ one", which was a bug).
| skill | owner | how it gets here |
|-------|-------|------------------|
| `pi-devbox-environment` | pi-devbox (this repo) | authored here; the canonical copy |
| `credential-incident-response` | pi-devbox (this repo) | authored here; the canonical copy |
| `pi-extensions` | the `pi-extensions` package repo (`skill/`) | **vendored fallback** + refreshed at build |
| `mempalace` | the `skillset` repo | **vendored fallback** (snapshot only) |
@@ -39,6 +40,33 @@ its skill file needed baking.
*different* skill, `opencode-mempalace-bridge`), so there is no public
package source to copy from. This snapshot is refreshed manually per release.
**Refresh it with `scripts/vendor-mempalace-skill.sh <skillset-root>`, not
`cp`.** Because the image cannot clone the private upstream, the snapshot used
to be *anonymous* — nothing recorded which skillset commit the bytes came
from, so the only staleness check possible was a hand-maintained phrase canary
in `scripts/smoke-test.sh`, which by construction detects "older than the
phrase I remembered to pin", never "older than skillset main". Two facts now
travel with the file:
| Fact | Where | Kind |
|---|---|---|
| `ARG SKILLSET_SNAPSHOT_REF` in `Dockerfile.variant` | manifest `skillset_snapshot_ref` + OCI label `se.jordbo.pi-devbox.skillset-snapshot-ref` | a **claim** about which commit these bytes are |
| `sha256sum` of this file, measured in the manifest layer | manifest `skillset_snapshot_sha256` | the bytes that **actually shipped** |
The script writes both together, refuses when the upstream file has
uncommitted modifications (no commit describes those bytes), and
`--check` verifies the claim against a real clone. Deliberately an `ARG`
default rather than a CI-resolved value: no credential for a private repo, no
change at any of the four `Dockerfile.variant` build call sites, and a local
`docker build` records the same thing CI does.
Verifying "is this snapshot current?" is **not** a CI job and was deliberately
not made one — see the Unreleased CHANGELOG entry for why (private repo;
another repo's branch must not be able to fail this build; and the artefact it
would guard is read by no host on this fleet). The check belongs where the
skillset actually is: `vendor-mempalace-skill.sh --check` for a maintainer,
and `pi-devbox-version`'s `skills:` section for an agent inside a container.
## Runtime precedence (v1.8.5+)
The baked links are created **early** in `entrypoint-user.sh` (before pi-deploy,
@@ -59,6 +87,21 @@ repoints the links for skills the **skillset owns**, listed one per line in
3. **baked snapshot** — everything else, and every skill when no skillset is
mounted
**Which one won is now reportable from inside the container:**
`pi-devbox-version` prints a `skills:` section naming, per vendored skill,
`baked` or `live <repo> @ <sha>` — and for `mempalace` whether that live copy is
identical to the baked fingerprint, at the same commit but with uncommitted
edits, or genuinely divergent. Before that, a stale baked snapshot and a current
live clone were indistinguishable from inside, which is how the freshness of
this file went unexamined for three releases. The section is suppressed with
`--no-skills` on the container-start banner, because `entrypoint-user.sh` prints
the version *before* the links exist and long before the reconcile below runs.
On this fleet, precedence 2 wins for `mempalace` on **every** host — all four
compose stacks mount a workspace containing the skillset — so the baked copy is
exercised only by CI and by a hypothetical no-mount container. Worth
remembering before spending effort on its freshness.
Ownership is per-skill on purpose: `pi-extensions`' authoritative source is the
package repo (copied over the snapshot at build), and `skillset` carries a
downstream copy that can lag, so handing it to the clone would *regress* the
@@ -72,15 +115,31 @@ mounted, plus a fabricated-skillset run of the reconciler).
cp <pi-extensions-pkg>/skill/SKILL.md pi-extensions/SKILL.md
cp <pi-extensions-pkg>/skill/evaluate-extension-usage.py pi-extensions/
cp <skillset>/skills/mempalace/SKILL.md mempalace/SKILL.md
Copy each snapshot **from its owner in the table above** — `pi-extensions` from
the package repo's `skill/` (since `a7f3044` co-located it there; `skillset`
also carries a copy, but it is a downstream duplicate and can lag), and
`mempalace` from `skillset`. Copying `pi-extensions` from `skillset` would
regress the snapshot to whatever that repo last mirrored.
Copy `pi-extensions` **from its owner in the table above** — the package
repo's `skill/` (since `a7f3044` co-located it there; `skillset` also carries a
copy, but it is a downstream duplicate and can lag). Copying `pi-extensions`
from `skillset` would regress the snapshot to whatever that repo last mirrored.
Snapshot provenance at last refresh: skillset `670f7f1`, pi-extensions pkg `e73cb9f`.
`mempalace` is **not** refreshed by `cp` — see the *Freshness model* section
above: `scripts/vendor-mempalace-skill.sh <skillset-root>` is the only thing
that should ever touch that snapshot, because a bare copy can update the bytes
without updating the ref that claims to describe them, which produces a
manifest that confidently lies.
Neither vendored skill has a hand-maintained "last refreshed at" line here on
purpose — one previously existed (skillset `670f7f1`, pi-extensions pkg
`e73cb9f`) and went stale within hours, because nothing forced it to move
when the ARGs did. `670f7f1` is now a cautionary example rather than a fact
worth recording: it is the commit that told agents to hand-stamp `added_by`,
which a later skillset commit (and the pi-devbox edge stamper) withdrew — so a
reader trusting that line would have been pointed at superseded guidance.
Both facts it tried to capture now live somewhere that cannot drift by hand:
| Fact | Where |
|---|---|
| which skillset commit `mempalace`'s bytes came from | `ARG SKILLSET_SNAPSHOT_REF` (Dockerfile.variant) + `skillset_snapshot_ref` in `build-manifest.json`, written *only* by `vendor-mempalace-skill.sh` |
| which pi-extensions package commit was vendored | `ARG PI_EXTENSIONS_REF` (Dockerfile.variant, CI-resolved to a 40-hex commit) → OCI label `se.jordbo.pi-devbox.pi-extensions-ref` and `build-manifest.json`'s `components.pi-extensions`, both read from the actual `/opt/pi-extensions` checkout, not from intent |
When you refresh the `mempalace` snapshot, also update the phrase asserted by
the "mempalace skill snapshot is current" smoke test — it deliberately pins the
@@ -0,0 +1,157 @@
---
name: credential-incident-response
description: >-
Respond correctly when a live credential is found somewhere it should not be —
in a chat transcript, a MemPalace drawer, a log, a git-tracked config, or an
agent-authored note. Load this whenever a task involves a leaked/exposed
secret, a token rotation, a "is this credential still live?" question, deciding
whether to delete or scrub stored content, or choosing scopes for a new API
token. Covers the mandatory order of operations (probe the issuer FIRST —
severity before cleanliness), leak-free credential identity via sha256[:8]
fingerprints, why revocation beats deletion for anything already replicated,
deriving least-privilege scopes from measured consumers instead of guessing,
where this fleet's secrets actually live (age-encrypted .env.age in
docker-compose-repo, plus gitignored plaintext .env drift), the three places a
secret hides in a Chroma palace, and the exposures that rotation does NOT fix.
---
# Credential incident response
A leaked credential is a **severity** question before it is a cleanliness
question. Two days of scrubbing, redaction plumbing and deletion planning were
once spent on a set of 13 credentials of which **11 were already dead at the
provider** — a fact that cost five HTTP requests to establish and was never
checked. Meanwhile the two live ones turned out to be instance-owner **admin**
tokens, which nobody had looked at either.
## 1. Order of operations — do not reorder this
1. **Is it still accepted?** Probe the issuing provider. Dead credential →
hygiene item, stop panicking. Live → incident, continue.
2. **What can it do?** Read the identity back. `is_admin`, `id=1`, scopes,
which account. A read-only repo token and an instance-owner admin token are
not the same finding.
3. **What consumes it?** Grep for real consumers before assuming breakage.
4. **Where does it live?** Enumerate copies (store, palace, transcripts, git).
5. **Then** rotate/revoke, and only then consider cleanup.
Doing 4→3→1 in reverse produces confident, wrong severity calls and wasted
cleanup. If you only have time for one step, do step 1.
## 2. Leak-free identity: fingerprint, never the value
Publishing an 8-hex fingerprint lets you compare a credential across machines,
files, drawers and peers without ever materialising the secret. Same formula as
`mempalace_redact.py`:
```sh
printf '%s' "$SECRET" | sha256sum | cut -c1-8 # printf, NOT echo (no newline)
printf '%s' 'test' | sha256sum | cut -c1-8 # self-test -> 9f86d081
```
Report as `(variable, fp, length)`. Equal fingerprints across hosts prove a
shared credential; that is usually the important part. **Never** paste a live
value into a search query, a palace drawer, an event body, or a chat message —
in an agent context your own tool output is itself captured and re-filed.
## 3. Liveness probes, and the trap that scoping creates
```sh
# Gitea
curl -sS -m 10 -o /dev/null -w '%{http_code}\n' -H "Authorization: token $T" \
"$GITEA_HOST/api/v1/repos/<owner>/<repo>/actions/runs?limit=1"
# GitHub
curl -sS -m 10 -o /dev/null -w '%{http_code}\n' -H "Authorization: token $T" \
https://api.github.com/user
```
- `200` live · `401` revoked/invalid · **`403` = wrong question, not a dead token**
- **Probe the issuer that minted it.** A 401 from an unrelated instance says
nothing. Resolve the host from config (`GITEA_EGL_HOST` etc.), do not assume.
- **Under scoped tokens, `/api/v1/user` returns 403 for a perfectly live token**
unless `user` scope was granted. So it cannot distinguish *revoked* from
*merely scoped*. Use a **repository route the token is authorised for**.
- Verify **both directions** after a rotation: old → 401, new → 200. The second
check is what catches "deleted the wrong token".
- Port/scheme come from config, not habit: one instance here is
`http://gitea.egl.lan:3000` — plain HTTP, with 443 refused.
## 4. Revocation beats deletion — the load-bearing rule
Once revoked, stored copies are **inert**; you may leave them. Deleting them is
best-effort over an *unbounded* copy set: FTS shadow rows, feed inbox `.jsonl`
files on every host, sqlite free pages after the delete, mesh replicas that
already synced, and backups. **Revocation invalidates every copy everywhere at
once, including copies nobody enumerated.**
So: **rotate + revoke first.** Treat drawer deletion as optional hygiene, never
as the remedy. Then record the retired fingerprints as *known-dead* so the next
census recognises them instead of reopening the investigation.
Corollary: never reach for `mempalace_sync` or a bulk `delete_by_source` on a
shared palace as incident response. High blast radius, low actual benefit.
## 5. Finding a secret in a Chroma palace — three targets, in this order
1. `embedding_fulltext_search_content.c0` — **where document text actually is**
2. `embedding_metadata.string_value` — metadata fields only
3. raw byte scan of every `*.sqlite3` — backstop, covers FTS pages and free space
Scanning only (2) is the classic false clean: hundreds of thousands of rows,
zero hits, and the secret sitting in (1) the whole time. Semantic search proves
nothing about absence — it returns top-k. For completeness, enumerate by filing
window (`list_drawers(since=T, before=T+1m)`), since one mine shares a minute.
Value-agnostic sweeps (uuid / 40-hex / `NAME=VALUE`) drown in false positives at
fleet scale — 608 candidates, mostly session UUIDs and git SHAs. Name-anchoring
plus entropy plus provenance, applied to **document text**, is what works.
## 6. Choosing scopes: derive them from measured consumers
Before creating a replacement token, find out what actually uses it:
```sh
git -C <repo> remote get-url origin # ssh:// ? then git needs NO token
git config --global --list | grep -iE 'credential|insteadof' # and no helper?
grep -rhoE 'api/v1/[A-Za-z0-9/{}$_.-]+' <consumers> | sort -u # exact routes
grep -rhoE '\-X [A-Z]+' <consumers> # any writes?
```
Real outcome here: git used SSH keys throughout, and the token's only consumer
read three CI-run routes with `GET`. So `repository: Read` and nothing else
replaced two admin tokens. **Scoping shrinks the blast radius of the next leak
far more than any redaction pipeline does** — a read-only token in a transcript
is a hygiene event, not an instance compromise.
Then prove the scope with an acceptance suite that declares expectations first:
must-work routes → `200`; `/admin/*`, `/user`, `/user/repos` → `403`.
## 7. What rotation does *not* fix
- **A cleartext channel.** If the endpoint is `http://`, the *new* token is
exposed identically from first use. Raise TLS separately.
- **Git history.** A secret committed and pushed cannot be fixed by any store or
palace operation — it needs rotation *and* history surgery.
- **Agent-authored content.** Stage-write redactors see transcripts only, never
`add_drawer` / `checkpoint` / `diary_write` output. Never type a secret into
the palace yourself; nothing downstream will catch it.
- **Plaintext/encrypted drift.** Gitignored plaintext `.env` files go stale while
`.env.age` moves on, so old values linger on disk (and in backups) long after
rotation. They are a common source of "mystery" fingerprints in a census.
## 8. This fleet's secret store (verify, do not assume)
- All `*.env.age` live in **one** repo: `joakimp/docker-compose-repo`. `myconfigs`
has none.
- Every `.age` file has **one X25519 recipient** — a single key tracked in
`myconfigs` under git-crypt. Unlocking git-crypt therefore decrypts the entire
fleet's secrets, including hosts you have no access to. The age layer adds no
isolation beyond git-crypt.
- Flow: `./fetch-secrets.sh <host>` (decrypt → `.env`) → edit → `./encrypt-secrets.sh <host>`
→ commit → push → `docker compose up -d --force-recreate`.
- **Always pass the host argument** to `encrypt-secrets.sh`. Bare, it walks the
whole tree and re-encrypts every `.env` it finds, re-nonced, including stale
ones — silently rolling back other hosts' secrets.
- After any re-encrypt, check the header still shows exactly **one X25519
recipient**; a hand-rolled `age -r` locks the rest of the fleet out, and the
failure only appears on another machine, later.
@@ -41,6 +41,21 @@ Run these immediately when a session begins, before responding to the user:
mempalace_kg_query(entity="<project_or_person>")
```
4. **Check your mailbox.** Just run it — an empty result is a fine answer and
costs one call. Do not try to decide first whether coordination "applies to
you"; that test is what used to be wrong here (see *Cross-Machine
Coordination* below):
```
mempalace_event_list(to_agent="<harness>@<device>", status="open")
```
This is a candidate list, not a to-do list — `status` never changes after an
event is written, so finished asks keep matching. Subtract the ones you have
already answered using the rule in *What you actually owe*, below.
Another machine may have asked you something, or corrected something you are
about to rely on. This costs one call and is the only way you will find out:
nothing pushes an event into your session unless your bridge delivers it for
you, and if it does you will already have seen it before reading this.
Do NOT announce this to the user. Just do it silently to orient yourself.
### Temporal grounding — compute time deltas, don't guess
@@ -340,6 +355,234 @@ mempalace_kg_invalidate(subject="...", predicate="...", object="...", ended="<to
mempalace_kg_add(subject="...", predicate="...", object="...", valid_from="<today>")
```
## Cross-Machine Coordination — the logstream
The palace stores what you *know*. The logstream (`mempalace_event_*`,
`mempalace_artifact_*`) carries what you want to *say to another agent* —
delegation, review, patch handoff, retraction. It is the only channel on which
another machine can reach you.
**Does this apply to you at all? Do not use `mempalace_mesh_peers` to decide.**
It answers a different question than it appears to. A shared palace can be
*hub-and-spoke* — many machines as thin clients of one central replica — and
then `mesh_peers` reports `peers: []` because there are no peer *replicas*,
even while four machines are actively writing to the same log. Measured on this
fleet: `peers: []`, one replica authoring every event from every machine. An
earlier version of this section told you to read `mesh_peers` and skip the
mailbox when it came back empty, which disabled the mailbox on precisely the
fleet it was written for.
The honest discriminators, cheapest first: **just run the mailbox query** (empty
is a fine answer); check whether `MEMPALACE_REMOTE_URL` is set, which is what
actually selects a shared palace; or look for any event whose `from_agent` is
not you. On a solitary palace the event tools still work — you are writing to
yourself and your mailbox stays empty. That is not a fault to debug.
**It is a durable log, not a bus — nobody is "listening".** Events are appended
and persist; there is no subscription, no delivery window, and nothing is lost
by being offline when one is written. A message waits indefinitely for you, and
your reply waits just as patiently for a sender who has since gone away. Machines
in a fleet are rarely awake at the same time, which is exactly why this is a log
and not a chat.
**Agent name is the only identity the log has.** Depending on deployment, every
client may share one `origin_replica` — on the fleet this skill was written for,
all machines are thin MCP clients of a single central replica, so `origin_replica`
is identical for every event and cannot tell two machines apart. `from_agent` /
`to_agent` carry the whole distinction, which is why the `<harness>@<device>`
stamping in *Provenance is stamped for you* is load-bearing here and not mere
tidiness.
### Reading your mailbox
```
mempalace_event_list(to_agent="<harness>@<device>", status="open")
```
- `to_agent=<you>` **also matches `*` broadcasts**, so one call covers both. No
second query needed.
- `status="open"` narrows the mailbox to what a sender *said was an ask at the
time of writing* — that is all it can do. It is a good first filter (on a real
stream it cut 5 events to 2), but it is **not** a list of what you owe, and it
never shrinks as you work. Treating it as owed-ness is the mistake this
section previously made: an earlier draft cited "5 unfiltered, exactly 1
filtered — the one that needed a reply" as proof the filter tracked
obligation. It did not. That single result was an event which had *already
been acked* half an hour earlier; the filter looked decisive only because the
stream happened to contain one directed `open` event. **Unfiltered mailboxes
train you to ignore them — and so does a filter that keeps showing you
finished work.**
- To resume where you left off, use `since_event_id`, **never**
`since_created_at`. A timestamp cursor permanently skips an event that synced
in late — it is a time window ("what happened today"), not a cursor.
- Read `metadata` before acting: senders put the load-bearing specifics there
(which host verified what, which run failed, what a change retracts).
### The ack contract — the sender declares whether a reply is owed
An obligation you never agreed to is noise, so the sender states it:
| Sender writes | Means | Recipient owes |
|---|---|---|
| `to_agent="<specific agent>"` + `status="open"` | an ask | an ack or a reply (the event itself keeps matching forever — see below) |
| `to_agent="*"` (any status) | broadcast FYI | nothing |
| any other status (`ready`, `applied`, `blocked`, …) | a statement of fact | nothing |
**That table says what you *owe*. Delivery is stricter, and the difference bites:
the mailbox is an obligation channel, not a news channel.** Mailbox candidates are
drawn with `status="open"`, so an event carrying any **terminal** status
(`applied`, `superseded`, `failed`, `blocked`) is never a candidate — *whoever it
is addressed to*. A `task.reply` written to a named machine to share a finding is
delivered to nobody, ever, and neither is any `event_ack`. It sits in the log
until somebody reads the log.
So the most natural inter-machine message — *"here is something you should
know"* — is exactly the shape that gets no delivery. Pick deliberately:
| You want the peer to… | Write |
|---|---|
| **do something**, and you need it tracked until done | directed `status="open"` ask, with a `correlation_id` |
| **know something**, no response needed | terminal-status event **plus a drawer** — the drawer is what actually reaches them, via search |
What does **not** work is a terminal report plus an expectation of attention.
Measured 2026-08-26: a detailed report addressed to `pi@<peer>` with
`status="applied"` went unread for two and a half hours until the operator quoted
the event id by hand, with the mailbox working correctly the whole time. Full
mechanism in the toolkit's `docs/rfc-003-coordination-log.md` §7.12.
One more timing fact, because it looks like negligence and is not: a delivered
ask is queued into the agent's **next turn** (`deliverAs: "steer"`, deliberately
no `triggerTurn`), and the poll fires when the agent is *idle*. Between delivery
and the next turn no inference runs, so **a human starting a turn is the
trigger** (§7.11). An agent that "has not reacted" has usually not been running.
Ack with `mempalace_event_ack(event_id=…, from_agent="<you>", status=…)`. It
**appends a new event** and never mutates the original; the correlation id is
copied for you, and `metadata.ack_of` is set to the event you answered.
**Claiming, and what it does not do.** `status="claimed"` announces that you have
picked work up. Nothing requires it — a directed open ask owes "an ack *or* a
reply", and finishing the work is a complete answer. Do it anyway when the work is
long or the machine is unreliable, because it is the only thing that later
distinguishes *nobody started this* from *someone started and their container
died mid-task*. Be clear about its limits, both of which follow from candidacy
requiring exactly `status="open"`:
- **It does not notify the requester.** `claimed` is not `open`, so a claim is no
more deliverable than a finished report is (see the delivery table above). Its
reader is whoever pulls the log.
- **It does not quiet your own mailbox.** The ask stays owed until a *terminal*
event of yours joins it, so a claimed-then-silent thread keeps resurfacing —
correctly.
Prefer a prompt terminal reply over a claim plus a long silence; claim *in
addition*, when the gap between pickup and finish is where a machine might die.
#### What you actually owe — derive it, do not read it off `status`
The log is append-only and `status` is written **once**, so it is an honest
statement about an item *at the moment it was written* and nothing more. It is
not mutable state, and asking it to carry mutable state is what breaks:
acking appends a new event and changes nothing about the old one, so **a
directed `open` event matches your mailbox query forever, answered or not.**
Nothing is ever "dismissed" — which also means a deferred ask cannot be
accidentally lost, only that you must compute what is outstanding:
```
candidates = mempalace_event_list(to_agent="<you>", status="open")
mine = mempalace_event_list(from_agent="<you>")
```
A candidate is **answered** when one of your own events
1. has a **higher `seq`** than the candidate, and
2. joins to it — `metadata.ack_of == candidate.id` (exact, written for you by
`event_ack`) or the same `correlation_id` (the fallback), and
3. carries a **terminal** status: `applied`, `superseded`, `failed`, `blocked`.
Everything else is still owed. Two calls, constant cost.
**Compare `seq`, never `created_at`** — the same reason you resume with
`since_event_id`. Without the ordering test, one terminal reply would suppress
every later ask on the same `correlation_id` for good; verified on a live thread
where a `ready` reply at `seq` 16 sits *before* the request at `seq` 17 that it
obviously cannot have answered.
**On a real mesh, compare `hlc` instead.** `seq` is *replica-local*: it equals
`origin_seq` today only because a single replica authors events for every
machine. Enrol a second replica and a late-syncing peer event gets a late local
`seq`, so two replicas can order the same pair differently and derive different
owed-sets from the same log. Every event already carries `hlc`
(`<millis>-<counter>-<replica_id>`), which is total and causally consistent.
So: compare `seq` while `mempalace_mesh_peers` reports no peers, `hlc` once it
reports any, and `created_at` never. (This is a legitimate use of `mesh_peers` —
choosing an ordering key — not the discredited gate on *whether* to read your
mailbox at all.)
**The failure directions are not symmetric, which is why this is safe to get
slightly wrong.** Local-`seq` skew can make an already-answered item *resurface*
as owed: noise, self-correcting, and visible. A timestamp comparison can
*suppress an unanswered ask forever*: silent and permanent. So if you ever see an
item you know you answered come back, do **not** "fix" it by reaching for
`created_at` — you would be trading the safe failure for the dangerous one.
This also supplies the "taken, not finished" state that looked missing:
`claimed` and `ready` are deliberately **not** terminal, so work you have picked
up keeps resurfacing until you close it out. No extra convention, no new field.
Two consequences worth internalising:
- **"Seen, not doing it" is a legitimate ack** — `status="blocked"` or
`"superseded"` plus the reason. Silence is not, and it is not merely rude:
with no terminal event of yours to join to, the ask stays in the owed set
indefinitely and there is nothing anyone can do about it from the other end.
- **Nothing expires, and it should not.** An `open` with no terminal reply is
still live by definition, and the finished threads are valuable history. If
content is genuinely perishable ("do not push to main for the next hour"), say
so in `metadata.expires_at` — metadata is stored verbatim — and honour it as a
hint when reading. An old `open` that the derivation still counts as owed is a
signal, not garbage: it means somebody asked and nobody answered.
### Writing to another machine
- **Address the stamped name you actually saw** in a `from_agent` field, e.g.
`pi@tor-ms22`. A bare `pi` reaches nobody's mailbox once stamping is live, and
older events in the log still carry bare names — do not copy them.
- **The rule runs in reverse too: what you put in YOUR OWN `from_agent` decides
where every reply to your event goes.** Nothing stops you writing a synthetic
or borrowed identity there, and a reply is always addressed back to exactly
that string — so if no live session ever runs as it, the reply is stored,
searchable, and delivered to no one. Measured cost: a directed ask sent under
a synthetic sender got two correct replies, one of them an urgent security
finding, and both sat unread for ~2h20m because nobody's mailbox was that
identity (RFC 003 §7.13). Authoring under a synthetic name is fine for a
deliberate control experiment — this fleet does it on purpose — but then
**name the real identity to reply to inside the body**, because the address
line is not a safe place to also carry provenance.
- **Use `status="open"` only when you truly need an answer.** It places an
obligation on another machine.
- **Never broadcast an ask.** `to_agent="*"` + `status="open"` obliges everyone
and therefore no one.
- **Always set a `correlation_id` on a directed `open`,** and reply with the
same one. It is not just for reconstructing a conversation later: it is the
join the owed-set derivation depends on. An uncorrelated ask can only ever be
closed by an `event_ack` (which sets `ack_of` for you) — a plain reply cannot
be matched to it at all.
- **Corrections are new events, never edits.** Say explicitly what you retract
and name the id — drawer or event — that carried the withdrawn claim.
- **Put a retraction where the reader will look.** A *directed open ask* reaches a
live agent's mailbox; a **terminal-status event reaches no mailbox at all**, and
a *drawer* is what a future semantic search finds. If you filed advice as a
drawer and later withdraw it, file the withdrawal as a drawer too — otherwise
the next agent finds your original confident advice and no trace of the
correction. (This is a real incident, not a hypothetical.)
- **Hand over exact content as an artifact**, not prose: `mempalace_artifact_put`
or `mempalace_patch_submit` store bytes with a sha256, and the event references
the id. Never paste a diff into a body and hope it survives.
- **Waiting on a specific reply?** `mempalace_event_wait` blocks with backoff —
do not poll `event_list` in a loop. A timeout there is a normal result, not an
error.
## Palace Structure
### Wings
@@ -366,7 +609,7 @@ Zechner's pi-coding-agent). Implications:
When the palace is **central** (shared across machines), these further things apply:
- **Check which machine a conversation came from.** Transcripts are fed per device, so `source_path` reads `…/mempalace-feed/<device>/pi_<uuid>.jsonl` while the displayed `source_file` is only the basename. One search can legitimately return hits from several machines at once — look at the device segment before attributing a decision to *this* project.
- **Provenance is stamped for you — leave it alone.** Drawers carry `device` and `agent_kind` metadata (plus `device_source`/`agent_kind_source` recording *how* each was determined, so an inference is never mistaken for a fact). You do **not** set these, and you no longer set `added_by` either: the pi bridge defaults the writer field to `<harness>@<device>` on `add_drawer`/`checkpoint`/`mine`/`event_append`/`artifact_put`, and prefixes diary entries with `HOST:<device>|`, from host-supplied `$MEMPALACE_PI_DEVICE`. RFC 001 §7.3.2 ranks "agent stamps it via a skill instruction" as the *worst possible* place for exactly the reason you would expect — it is per-call boilerplate that gets forgotten, and it did: the agent who wrote the previous version of this bullet then filed its own provenance drawer as `added_by=checkpoint`. Two things remain yours: pass `source_drawer_id` on `kg_add` (triples have no provenance field, so that pointer is the only path back to a device), and pass an explicit `added_by` **only** when deliberately filing on behalf of another device. Never invent values for `device`/`agent_kind`/`origin_device` — a fabricated value is worse than a blank, because it silently corrupts a future merge.
- **Provenance is stamped for you — leave it alone.** Drawers carry `device` and `agent_kind` metadata (plus `device_source`/`agent_kind_source` recording *how* each was determined, so an inference is never mistaken for a fact). You do **not** set these, and you no longer set `added_by` either: the pi bridge defaults the writer field to `<harness>@<device>` on `add_drawer`/`checkpoint`/`mine`/`event_append`/`artifact_put`, and prefixes diary entries with `HOST:<device>|`, from host-supplied `$MEMPALACE_PI_DEVICE`. RFC 001 §7.3.2 ranks "agent stamps it via a skill instruction" as the *worst possible* place for exactly the reason you would expect — it is per-call boilerplate that gets forgotten, and it did: the agent who wrote the previous version of this bullet then filed its own provenance drawer as `added_by=checkpoint`. **Confirm the bridge in your image actually stamps before trusting it:** the extension is baked at image build time, so a container on an image older than the stamping commit (pi-devbox < v1.8.7) stamps nothing while still satisfying both gates — the env vars are set and the code is simply absent. Check with `grep -c MEMPALACE_PI_DEVICE "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)"`; zero means keep passing `added_by="<harness>@<device>"` and a manual `HOST:<device>|` diary prefix until the container is recreated on a newer image. Two things remain yours: pass `source_drawer_id` on `kg_add` (triples have no provenance field, so that pointer is the only path back to a device), and pass an explicit `added_by` **only** when deliberately filing on behalf of another device. Never invent values for `device`/`agent_kind`/`origin_device` — a fabricated value is worse than a blank, because it silently corrupts a future merge.
- **Metadata is invisible to search — so check the text, not the fields.** `search` results are built from a fixed key list and `diary_read` returns content, so neither ever shows `device`/`added_by`. Only `mempalace_get_drawer` reveals them. This is why diary entries carry an in-text `HOST:<device>` marker: it is the only attribution a reader actually sees. **A diary entry with no `HOST:` marker predates the convention and may be from any machine — do not assume it is this one's history.**
- **Mined drawers carry the MINE date, not the session date.** When history is imported, or re-mined on the palace host, `filed_at`/`created_at` is the *import* time — so sorting by them does not give chronological order. Real session time is recoverable from the UUIDv7 in `pi_<uuid>.jsonl`: the first 12 hex digits are milliseconds since the epoch (and UUIDv7 sorts lexicographically in time order, so a plain filename sort is already chronological). Agent-authored drawers and diaries have no such backdoor — for those `filed_at` is the only chronology, which is why it must never be restamped.
- **Beware the timezone mismatch when you combine those.** Palace `filed_at`/`created_at` are naive timestamps in the palace host's local time, while a UUIDv7 decodes to UTC. Comparing them directly introduces a silent offset (2 h for a CEST host). Normalise before drawing conclusions about ordering.
@@ -413,4 +656,7 @@ Entity-relationship triples with temporal validity. Query with `mempalace_kg_que
- **Don't mine .git directories or node_modules.** The CLI miner respects .gitignore by default.
- **Don't create duplicate drawers.** Use `mempalace_check_duplicate` before adding manually.
- **Don't treat the palace as a task list.** It's for knowledge and context, not todos.
- **Don't invent provenance metadata, and don't hand-stamp it either.** An earlier version of this list told you to set `added_by="<harness>@<device>"` by hand; that instruction has been withdrawn, because RFC 001 §7.3.2 places provenance at the client/server boundary and the pi bridge now does it uniformly (see *Provenance is stamped for you* above). DO NOT invent values for the palace's own metadata fields (`device`, `agent_kind`, `origin_device`): those are stamped by infrastructure that also records *how* each was determined, and a fabricated value is worse than none because it silently corrupts a future merge. DO pass `source_drawer_id` on `kg_add`. And never put a machine name in a diary's `agent_name` — it becomes the wing name and hides your entries from `diary_read`.
- **Don't broadcast an ask, and don't leave one unanswered.** On a shared palace, `to_agent="*"` + `status="open"` obliges every machine and therefore none of them. And don't expect acking to tidy your mailbox: `status` is immutable, so the event keeps matching either way — what a terminal reply buys you is that the *derived* owed set (see *What you actually owe*) stops counting it. Leave asks unanswered and that set only grows, until everyone learns to stop looking. "Seen, not doing it" is a complete answer — silence is not.
- **Don't assume you would have heard.** Nothing pushes another machine's message into your session. If you did not run the mailbox query at wake-up, a correction addressed to you by name can sit unread while you confidently rebuild the thing it warned you about.
- **Don't author an ask under an identity nobody runs as, including your own throwaway labels.** The failure is symmetric to the one above: it is not that you missed a message, it is that nothing could ever have delivered the reply to you, because you addressed it at a name instead of an agent. If you must use a synthetic sender for a control or an experiment, say inside the body who should actually receive the reply.
- **Don't invent provenance metadata, and don't hand-stamp it either.** An earlier version of this list told you to set `added_by="<harness>@<device>"` by hand; that instruction has been withdrawn, because RFC 001 §7.3.2 places provenance at the client/server boundary and the pi bridge now does it uniformly (see *Provenance is stamped for you* above) — but the withdrawal only holds where the bridge is live, so run the one-line check in that bullet first; on an older image hand-stamping is still the only signal a hand-filed drawer gets. DO NOT invent values for the palace's own metadata fields (`device`, `agent_kind`, `origin_device`): those are stamped by infrastructure that also records *how* each was determined, and a fabricated value is worse than none because it silently corrupts a future merge. DO pass `source_drawer_id` on `kg_add`. And never put a machine name in a diary's `agent_name` — it becomes the wing name and hides your entries from `diary_read`.
@@ -143,6 +143,9 @@ mine:
| "`tor-ms22` is not in the SSH config" | `grep … \| head -20` — the entry was at **line 454**. `~/.ssh/config` here is ~500 lines. |
| "the Docker host has no `docker`" | non-interactive SSH `PATH` lacks `/usr/local/bin` (§2, §3). It was at `/usr/local/bin/docker`. |
| "no ControlMaster is running" | pattern `ssh ` (trailing space) cannot match a master: those processes **rename themselves** to `ssh: <controlpath> [mux]`. |
| "the credential is not in the palace" | scanned `embedding_metadata.string_value` only. Drawer **text** lives in `embedding_fulltext_search_content.c0`; 554k metadata rows proved nothing. |
| "this token is dead — 401" | probed it against the **wrong issuer**. A 401 from an instance that never issued the credential is not evidence about the credential. |
| "that host is unreachable, can't test" | tried ports 443 and 80. It was on **3000**, and the env var I already held (`GITEA_EGL_HOST`) stated the scheme and port. |
Habits that would have caught all three:
@@ -157,8 +160,48 @@ ssh -F "$HOME/.ssh-local/config" mac 'command -v docker || ls /usr/local/bin/doc
ps -eo pid,etime,args | grep -Ei 'mux|mosh|ssh'
```
A positive result needs no such scepticism — it carries its own evidence. Only
absence has to be *earned*, so spend the extra command there.
Absence has to be *earned*, so spend the extra command there.
### …and a positive result only proves what you *actually asked*
An earlier version of this section claimed "a positive result needs no such
scepticism — it carries its own evidence." **That is false, and believing it
cost a later session three more wrong findings.** A positive result is evidence
about the question your command really posed, which may not be the question you
meant. The failure is invisible precisely *because* the command succeeded.
| Claim | The command succeeded — at answering something else |
|---|---|
| "EGL git over SSH works" | `ssh git@gitea.egl.lan` greeted me as `joakimp`. `~/.ssh/config` had `Host gitea*` → `HostName gitea.jordbo.se`, so I authenticated **to the wrong instance**. The real EGL account is `ecsjper`. |
| "the port config regressed" | compared `ssh -G` output against `2222` — a value produced by **my own earlier `-p 2222` flag**, not by the config. I reported the user's edit as a regression it never caused. |
| "the CI runners authenticate with this token" | pure fabrication, contradicted by my own scan output already on screen. The runners use per-runner `REGISTRATION_TOKEN`. |
Two habits that actually catch this class, both cheap:
```sh
# 1. ask which RULE captured your hostname before trusting any ssh result.
# ssh_config is first-obtained-value-wins PER KEYWORD, not per block: a
# specific block only wins the keywords it declares, so a later `Host gitea*`
# still supplies HostName unless the specific block restates it.
ssh -G git@thehost | grep -E '^(hostname|port|user|identityfile)'
# 2. state the expected result BEFORE running the check, and diff against it.
# This is the single technique that separated the one verification that went
# right (10/10, expectations declared per probe) from five that went wrong
# (results interpreted after the fact, each time in the direction I expected).
probe "/repos/.../actions/runs" 200 # must work
probe "/admin/users" 403 # must be denied
```
And the meta-observation, which is the reason this subsection exists: across all
five errors, **not one was caught by re-reading my own reasoning.** Every one was
caught by a second measurement that disagreed — the SSH lie surfaced only because
the greeting said `joakimp` while a token probe minutes earlier had said
`ecsjper`; the fabrication surfaced only because the user read my own output back
to me. So the operational rule is not "be careful". It is: **for a load-bearing
claim, produce a second measurement by a different route, and expect it to
disagree.** If you cannot think of a second route, you do not yet have a finding
— you have a hypothesis.
**`dscp`/`scp` with accented filenames on a macOS host.** macOS stores filenames
in Unicode **NFD** (decomposed — e.g. `ä` is `a` + combining U+0308), while the
+125 -15
View File
@@ -2,8 +2,8 @@
# Runtime post-recreate verification for pi-devbox.
#
# Verifies that after `docker compose up -d --force-recreate`:
# - The new image is actually live (pi version matches, when an expected
# version is supplied — see the version note below)
# - The new image is actually live (both the pi version and — when asked —
# the pi-devbox image release tag; see the two version notes below)
# - Persisted named volumes survived (~/.pi config, shell history, zoxide,
# nvim data, uv cache, ssh-local)
# - pi runtime wiring is intact: keybindings symlink, AGENTS.md symlink,
@@ -25,13 +25,33 @@
# the pi-devbox repo (which a maintainer already has for CI builds). A plain
# `docker pull` consumer is not the audience and will not have this file.
#
# Version note: pi's version is resolved from `latest` at CI build time and is
# NOT pinned to a concrete value in Dockerfile.variant (ARG PI_VERSION=latest).
# So unlike opencode-devbox, this script cannot self-derive an expected version
# from the Dockerfile. Pass --expected-version to assert a match; without it the
# live pi version is reported as an informational WARN, not a failure.
# TWO DIFFERENT VERSIONS, TWO DIFFERENT FLAGS. This distinction has already
# cost a release day, so it is spelled out here and in AGENTS.md step 4:
#
# Usage: ./scripts/recreate-sanity-check.sh [--expected-version X.Y.Z] [--variant studio|plain]
# --expected-version the PI CODING AGENT version, e.g. 0.84.3
# (`pi --version`; pinned as ARG PI_VERSION in
# Dockerfile.variant, which CI reads as the source
# of truth)
# --expected-image-version the PI-DEVBOX IMAGE release tag, e.g. 1.8.9 or
# v1.8.9 (the `release_tag` baked into
# /etc/pi-devbox/build-manifest.json)
#
# Passing a release tag to --expected-version used to report
# "pi version mismatch: expected 1.8.8, got 0.84.3" — an accusation aimed at
# the wrong component, on the last gate of a release. Both flags now detect
# being handed the other one's value and say so instead.
#
# Neither flag is required. Both values are derivable from the image's own
# build manifest, so by default the script asserts the LIVE pi version against
# the version recorded at build time — which is not a tautology: a stale
# `pi` in the ~/.pi/npm-global volume can shadow the baked one, exactly the
# way a stale npm:pi-atelier can (see the packages[] check below). Pass the
# flags when you want an assertion against a value you name yourself, which
# is what a release checklist wants.
#
# Usage: ./scripts/recreate-sanity-check.sh [--expected-version X.Y.Z]
# [--expected-image-version X.Y.Z]
# [--variant studio|plain]
#
# Exit codes:
# 0 all checks passed
@@ -41,22 +61,61 @@
set -euo pipefail
EXPECTED_VERSION=""
EXPECTED_IMAGE_VERSION=""
VARIANT=""
REPO_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
MANIFEST=/etc/pi-devbox/build-manifest.json
# Parse arguments
usage() {
cat >&2 <<'EOF'
usage: recreate-sanity-check.sh [--expected-version X.Y.Z]
[--expected-image-version X.Y.Z]
[--variant studio|plain]
--expected-version pi coding agent version, e.g. 0.84.3 (`pi --version`)
--expected-image-version pi-devbox image release tag, e.g. 1.8.9 or v1.8.9
--variant studio|plain (auto-detected when omitted)
These are two different versions. Both are read from the image's own build
manifest when the corresponding flag is omitted.
EOF
}
# Parse arguments. Every flag takes a value, so reject a missing one rather
# than swallowing the next flag as if it were the value.
need_value() {
case "${2:-}" in
""|-*)
echo "$1 requires a value" >&2
usage
exit 2
;;
esac
}
while [[ $# -gt 0 ]]; do
case "$1" in
--expected-version)
need_value "$@"
EXPECTED_VERSION="$2"
shift 2
;;
--expected-image-version)
need_value "$@"
EXPECTED_IMAGE_VERSION="$2"
shift 2
;;
--variant)
need_value "$@"
VARIANT="$2"
shift 2
;;
--help|-h)
usage
exit 0
;;
*)
echo "usage: $0 [--expected-version X.Y.Z] [--variant studio|plain]" >&2
echo "unknown option: $1" >&2
usage
exit 2
;;
esac
@@ -67,6 +126,19 @@ pass() { echo " ✓ $1"; }
fail() { echo " ✗ $1" >&2; FAILED=$((FAILED + 1)); }
warn() { echo " ⚠ $1" >&2; }
# Read one top-level field from the build manifest, or print nothing. The
# manifest is the image's own ground truth (written at `docker build` time by
# Dockerfile.variant), so it needs no checkout and no network. Absent on an
# image built before it existed, hence every caller treats "" as unknown.
manifest_field() {
[ -f "$MANIFEST" ] || return 0
command -v jq >/dev/null 2>&1 || return 0
jq -r --arg k "$1" '.[$k] // empty' "$MANIFEST" 2>/dev/null || true
}
# Release tags are written with a leading v in the manifest and quoted without
# one in checklists; compare on the bare number so both spellings work.
strip_v() { printf '%s' "${1#v}"; }
# Auto-detect variant if not provided. The studio variant vendors pi-studio to
# /opt/pi-studio; the plain variant does not.
if [ -z "$VARIANT" ]; then
@@ -86,21 +158,59 @@ else
fi
echo
echo "-- pi version --"
MANIFEST_PI_VERSION=$(manifest_field pi_version)
MANIFEST_RELEASE_TAG=$(manifest_field release_tag)
echo "-- pi (coding agent) version --"
if ACTUAL_VERSION=$(pi --version 2>&1 | head -1); then
if [ -n "$EXPECTED_VERSION" ]; then
if [ "$ACTUAL_VERSION" = "$EXPECTED_VERSION" ]; then
pass "pi version $ACTUAL_VERSION"
if [ "$(strip_v "$EXPECTED_VERSION")" = "$(strip_v "$ACTUAL_VERSION")" ]; then
pass "pi version $ACTUAL_VERSION (matches --expected-version)"
elif [ -n "$MANIFEST_RELEASE_TAG" ] &&
[ "$(strip_v "$EXPECTED_VERSION")" = "$(strip_v "$MANIFEST_RELEASE_TAG")" ]; then
# Exact, not heuristic: the value handed over IS this image's release
# tag, so it cannot be a pi version anyone meant.
fail "--expected-version $EXPECTED_VERSION is the pi-devbox IMAGE version, not the pi version — use --expected-image-version $EXPECTED_VERSION (live pi is $ACTUAL_VERSION)"
else
fail "pi version mismatch: expected $EXPECTED_VERSION, got $ACTUAL_VERSION"
fail "pi version mismatch: expected $EXPECTED_VERSION, got $ACTUAL_VERSION (this flag asserts the pi coding agent version; for the image release tag use --expected-image-version)"
fi
elif [ -n "$MANIFEST_PI_VERSION" ]; then
# Not a tautology: the manifest records what pi reported at BUILD time,
# while `pi --version` resolves through PATH, which a stale npm-global
# volume install can shadow.
if [ "$MANIFEST_PI_VERSION" = "$ACTUAL_VERSION" ]; then
pass "pi version $ACTUAL_VERSION (matches this image's build manifest)"
else
fail "live pi $ACTUAL_VERSION != $MANIFEST_PI_VERSION recorded in $MANIFEST — a stale pi in the ~/.pi/npm-global volume is shadowing the baked one"
fi
else
warn "pi version $ACTUAL_VERSION (no --expected-version given; pi is built from 'latest', cannot self-derive — informational only)"
warn "pi version $ACTUAL_VERSION (no --expected-version and no build manifest to compare against — informational only)"
fi
else
fail "pi --version failed"
fi
echo
echo "-- pi-devbox image version --"
if [ -z "$MANIFEST_RELEASE_TAG" ]; then
if [ -n "$EXPECTED_IMAGE_VERSION" ]; then
fail "cannot verify --expected-image-version $EXPECTED_IMAGE_VERSION: no readable release_tag in $MANIFEST (image built before the manifest existed, or jq missing)"
else
warn "image release tag unknown (no readable $MANIFEST) — pi-devbox-version would say the same"
fi
elif [ -n "$EXPECTED_IMAGE_VERSION" ]; then
if [ "$(strip_v "$EXPECTED_IMAGE_VERSION")" = "$(strip_v "$MANIFEST_RELEASE_TAG")" ]; then
pass "image version $MANIFEST_RELEASE_TAG (matches --expected-image-version)"
elif [ -n "$MANIFEST_PI_VERSION" ] &&
[ "$(strip_v "$EXPECTED_IMAGE_VERSION")" = "$MANIFEST_PI_VERSION" ]; then
fail "--expected-image-version $EXPECTED_IMAGE_VERSION is the pi version, not the image release tag — use --expected-version $EXPECTED_IMAGE_VERSION (this image is $MANIFEST_RELEASE_TAG)"
else
fail "image version mismatch: expected $EXPECTED_IMAGE_VERSION, got $MANIFEST_RELEASE_TAG — the recreate did not pick up the intended image"
fi
else
warn "image version $MANIFEST_RELEASE_TAG (no --expected-image-version given — informational only)"
fi
echo
echo "-- Persisted named volumes (must survive --force-recreate) --"
+84 -5
View File
@@ -7,7 +7,7 @@
# - pi binary present and (if EXPECTED_PI_VERSION set) matches CI's resolved version
# - mempalace core matches the audited pin (if EXPECTED_MEMPALACE_VERSION set)
# - new v1.0.0 base additions (pandoc, graphviz, imagemagick, yq, tealdeer)
# - typst PDF engine for pandoc (Unreleased) — `pandoc --pdf-engine=typst`
# - typst PDF engine for pandoc (v1.4.0) — `pandoc --pdf-engine=typst`
# - non-modal editors nano + micro (alongside nvim)
# - terminfo for modern emulators: xterm-kitty, xterm-ghostty, wezterm,
# alacritty, foot (kitty-terminfo + ncurses-term + compiled ghostty alias)
@@ -245,6 +245,8 @@ run "socat" "socat -V"
run "studio-expose helper" "test -x /usr/local/bin/studio-expose"
run "image-baked pi-devbox-environment skill" \
"test -f /usr/local/share/pi-devbox/skills/pi-devbox-environment/SKILL.md"
run "image-baked credential-incident-response skill" \
"test -f /usr/local/share/pi-devbox/skills/credential-incident-response/SKILL.md"
run "global-AGENTS append snippet present" \
"test -f /usr/local/share/pi-devbox/pi-global-AGENTS.append.md"
run "pi-devbox block merged into pi-global-AGENTS.md" \
@@ -463,6 +465,43 @@ run "pi-devbox-version --json round-trips the manifest byte-for-byte" '
'
run_expect "pi-devbox-version --quiet is a compact one-liner" \
"pi-devbox-version --quiet | wc -l" "1"
# ── Vendored skill snapshot provenance ─────────────────────────────────
# The vendored mempalace skill is the one baked artefact with no /opt clone
# behind it (private upstream — see VENDORED.md), so until now the manifest
# could not say which skillset commit it came from. Two fields now travel with
# it: the CLAIMED ref (ARG default in Dockerfile.variant) and the MEASURED
# sha256 of the shipped bytes. Assert both are well-formed, and — separately —
# that the measurement still describes the file in the image.
#
# Kept as two assertions for the same reason the component checks are: one
# proves the fields are not empty/garbage, the other proves they are not merely
# self-consistent. A single combined check could pass on a manifest whose hash
# was computed from a file that was later overwritten (the pi-extensions skill
# copy at Dockerfile.variant:165 does exactly that kind of overwrite, one stage
# earlier), which is the failure this second one exists to catch.
run "manifest records the vendored skill snapshot provenance" '
j=/etc/pi-devbox/build-manifest.json
r=$(jq -r ".skillset_snapshot_ref // empty" $j)
s=$(jq -r ".skillset_snapshot_tree_sha256 // empty" $j)
echo "ref=[$r] tree_sha256=[$s]" >&2
printf "%s" "$r" | grep -qxE "[0-9a-f]{40}" \
|| { echo "skillset_snapshot_ref is not a 40-hex commit" >&2; exit 1; }
printf "%s" "$s" | grep -qxE "[0-9a-f]{64}" \
|| { echo "skillset_snapshot_tree_sha256 is not a 64-hex digest" >&2; exit 1; }
'
# Recomputes over the whole DIRECTORY with the same tree_sha256() pipeline
# Dockerfile.variant used to measure it, not a plain `sha256sum SKILL.md` —
# a file-only compare here would pass even if the manifest recorded a
# fingerprint over a directory that has since grown a second file (this is
# not hypothetical: pi-extensions already ships two files for its skill).
run "manifest skill fingerprint matches the baked snapshot" '
j=/etc/pi-devbox/build-manifest.json
d=/usr/local/share/pi-devbox/skills/mempalace
m=$(jq -r ".skillset_snapshot_tree_sha256 // empty" $j)
a=$( (cd "$d" && find . -type f -print | LC_ALL=C sort | xargs -r sha256sum) | sha256sum | cut -d" " -f1)
echo "manifest=[$m] actual=[$a]" >&2
[ -n "$m" ] && [ "$m" = "$a" ]
'
# OCI labels live in the image config, not the container fs — inspect them
# from the host docker rather than via `docker run`.
LBL=$(docker inspect --format '{{ index .Config.Labels "se.jordbo.pi-devbox.pi-extensions-ref" }}' "$IMAGE" 2>/dev/null || true)
@@ -539,18 +578,58 @@ exec_test "mempalace skill linked (fallback)" 'test -L $HOME/.agents/skills
# one absent, so a re-vendored stale snapshot fails just as loudly as a
# forgotten bump. A one-way canary only catches half the drift.
# * a phrase canary can only ever detect "older than what I remembered to pin",
# never "older than skillset main". The real fix is a CI job diffing this
# file against the skillset repo — see the Unreleased changelog note.
# never "older than skillset main".
#
# That structural limit is now addressed, but NOT by the "CI job diffing this
# file against the skillset repo" this comment used to point at (that pointer
# also dangled: it referenced an Unreleased changelog note that had become the
# v1.8.7 heading). A CI diff cannot be done without granting CI a credential
# for the PRIVATE skillset repo, and it would guard a file that on this fleet
# NO host reads — all four compose stacks mount a workspace containing the
# skillset, so devbox-skill-reconcile repoints this link at the live clone and
# the baked copy is a CI/no-mount fallback only. Instead the snapshot now
# carries its provenance (skillset_snapshot_ref + a measured
# skillset_snapshot_sha256 in build-manifest.json, written by
# scripts/vendor-mempalace-skill.sh), which moves the check to where the
# skillset actually IS: `scripts/vendor-mempalace-skill.sh --check` for a
# maintainer, and `pi-devbox-version` for an agent inside any container.
# This assertion is kept because it is orthogonal and free: it pins content,
# not provenance, so it still catches a re-vendored snapshot whose ref was
# bumped correctly but whose bytes came from the wrong place.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Provenance is stamped for you" "$f" && ! grep -q "Attribute what you file yourself" "$f" && echo ok'
# Link TARGETS, not just link existence: with no skillset mounted (as here) the
# baked tree must be what resolves, for all three vendored skills.
# baked tree must be what resolves, for all four vendored skills.
exec_test "vendored skills resolve to the baked tree (no skillset mounted)" \
'for s in mempalace pi-extensions pi-devbox-environment; do
'for s in mempalace pi-extensions pi-devbox-environment credential-incident-response; do
case "$(readlink -f $HOME/.agents/skills/$s)" in
/usr/local/share/pi-devbox/skills/$s) ;;
*) echo "$s resolves to $(readlink -f $HOME/.agents/skills/$s)" >&2; exit 1 ;;
esac
done; echo ok'
# ... and that the tool REPORTS that resolution, which is the half that was
# missing: a stale baked snapshot and a current live clone were
# indistinguishable from inside the container. CI mounts no skillset, so every
# vendored skill must report "baked" here — which also makes this a real test of
# the fallback path rather than of the environment it happens to run in.
exec_test "pi-devbox-version reports skill sources (all baked, no skillset here)" \
'out=$(pi-devbox-version)
echo "$out" | grep -q "skills:" || { echo "no skills section" >&2; exit 1; }
for s in mempalace pi-extensions pi-devbox-environment credential-incident-response; do
echo "$out" | grep -qE "^ $s +baked$" \
|| { echo "$s not reported as baked" >&2; exit 1; }
done; echo ok'
# The boot banner must NOT carry the section: entrypoint-user.sh prints the
# version FIRST, before the baked links exist and long before the skillset
# deploy + reconcile run last, so anything it said about skill sources would be
# a pre-reconcile state that is about to change.
# A bare negative (`! grep -q "skills:"`) passes if the tool crashes or
# prints nothing at all — it cannot tell "correctly omitted the section"
# apart from "the binary is broken". Anchor it positively: the command must
# still succeed and still print its normal release-tag line.
exec_test "pi-devbox-version --no-skills omits the skills section" \
'out=$(pi-devbox-version --no-skills) && echo "$out" | grep -q "^pi-devbox " && ! echo "$out" | grep -q "skills:"'
exec_test "entrypoint prints the version banner with --no-skills" \
'grep -q "pi-devbox-version --no-skills" /usr/local/bin/entrypoint-user.sh'
# The handover path itself. CI never mounts a skillset, so without this the
# v1.8.5 fix would ship untested: fabricate a skillset + a skills dir holding
# baked-style links, run the reconciler, and assert all three outcomes —
+293
View File
@@ -0,0 +1,293 @@
#!/usr/bin/env bash
# vendor-mempalace-skill.sh — refresh the vendored mempalace skill snapshot
# AND its recorded provenance, together, so the two cannot drift apart.
#
# WHY THIS EXISTS
# ---------------
# rootfs/usr/local/share/pi-devbox/skills/mempalace/SKILL.md is a snapshot of a
# file owned by the PRIVATE skillset repo (see VENDORED.md). Because the image
# cannot clone that repo, refreshing the snapshot was a manual `cp` — and the
# result was anonymous: nothing recorded WHICH skillset commit the bytes came
# from. The only staleness check available was a hand-maintained phrase canary
# in scripts/smoke-test.sh, which by construction detects "older than the phrase
# I remembered to pin", never "older than skillset main".
#
# Two facts now travel with the snapshot: the skillset commit it was taken from
# (ARG SKILLSET_SNAPSHOT_REF in Dockerfile.variant) and the sha256 of the bytes
# themselves (measured at build time into build-manifest.json). This script is
# the only thing that should ever write the first one, because a `cp` without a
# matching ARG bump produces a manifest that CONFIDENTLY LIES — worse than the
# anonymous snapshot it replaced.
#
# HARDENED after peer review (pi@emb-7kj4vr4g, logstream correlation
# skills-provenance-review, 2026-08-26) proved the original --check could print
# OK and exit 0 without actually verifying anything: `git show <ref>:<path>`
# emits NOTHING when the ref/path doesn't resolve, and `sha256sum` still hashes
# that empty stdin, so "ref not found" silently collided with "the file really
# is 0 bytes". Depending on which side of the comparison hit the collision this
# fell through as either a false MISMATCH (blaming provenance for what was
# really an incomplete clone) or, worse, a false OK. See EXIT STATUS below —
# "cannot determine" is now its own outcome, distinct from "confirmed wrong",
# which is the same distinction the phrase canary this script replaced lacked.
#
# USAGE
# scripts/vendor-mempalace-skill.sh [skillset-root] [--force]
# refresh: rewrite the snapshot and the ARG together.
# scripts/vendor-mempalace-skill.sh --check [skillset-root]
# verify only, writes nothing. The root path and any flag may appear in
# either order — a positional-only parser previously made `<root>
# --check` silently run a refresh instead of the verification asked for.
#
# skillset-root defaults to /workspace/skillset, then $HOME/skillset.
#
# --force (refresh mode only) proceed even when the recorded ref cannot be
# proven to be an ancestor of the skillset's current HEAD — i.e.
# skip the guard against silently REWINDING provenance, which a
# detached HEAD, an older checkout, or a shallow clone lacking the
# recorded commit can all trigger. Meant to be used deliberately,
# not habitually: each use is a human deciding a rewind is fine.
#
# --check answers "is the committed snapshot really skillset@<recorded ref>?"
# — the question CI cannot answer without a credential for the private repo,
# and which anyone with the skillset checked out can answer for free.
#
# EXIT STATUS (same three codes in both modes)
# 0 the operation succeeded, or (--check) the record is verified truthful.
# This INCLUDES a truthful record that is merely stale — upstream has
# moved on since the recorded ref, or the local working tree has since
# diverged. A NOTICE is printed to stderr, but the snapshot is not being
# accused of lying, so this is not a release-blocking failure. Skipping a
# refresh is a legitimate release-day choice (see AGENTS.md); this exit
# code is what makes that choice checkable rather than merely asserted.
# 1 refused: a CONFIRMED problem. Dirty upstream file; a refresh that would
# rewind past the recorded ref; or (--check) the vendored bytes provably
# do NOT match the file at the recorded ref — a lying record.
# 2 cannot determine: the recorded ref, or the path at that ref, is not
# resolvable in this clone. Commonly a shallow clone missing history, or
# a ref that was rewritten or never pushed. Deliberately NOT the same as
# 1 — "I can't tell" must never be reported as "it's wrong".
set -euo pipefail
cd "$(dirname "$0")/.."
DOCKERFILE="Dockerfile.variant"
VENDORED="rootfs/usr/local/share/pi-devbox/skills/mempalace/SKILL.md"
ARG_NAME="SKILLSET_SNAPSHOT_REF"
REL_PATH="skills/mempalace/SKILL.md"
die() { printf '%s: %s\n' "$(basename "$0")" "$1" >&2; exit 1; }
# Parse flags and the optional root path in either order, and reject anything
# unrecognised rather than silently absorbing it.
MODE="refresh"
FORCE=0
ROOT=""
for arg in "$@"; do
case "$arg" in
--check) MODE="check" ;;
--force) FORCE=1 ;;
# This is the one script whose argument ORDER was itself a landmine, so the
# path that documents the trap must not be the path that errors.
-h|--help)
awk 'NR>1 && /^#/ { sub(/^# ?/, ""); print; next } NR>1 { exit }' "$0"
exit 0
;;
--*) die "unknown option: $arg (try --help)" ;;
*)
[ -z "$ROOT" ] || die "unexpected extra argument: $arg (root already set to $ROOT)"
ROOT="$arg"
;;
esac
done
if [ "$MODE" = "check" ] && [ "$FORCE" = 1 ]; then
die "--force has no effect with --check (nothing is written); remove it"
fi
if [ -z "$ROOT" ]; then
for candidate in /workspace/skillset "$HOME/skillset"; do
if [ -d "$candidate/.git" ]; then
ROOT="$candidate"
break
fi
done
fi
[ -n "$ROOT" ] || die "no skillset clone found (pass one: $(basename "$0") /path/to/skillset)"
[ -d "$ROOT/.git" ] || die "not a git clone: $ROOT"
[ -f "$ROOT/$REL_PATH" ] || die "no $REL_PATH in $ROOT"
[ -f "$VENDORED" ] || die "vendored snapshot missing: $VENDORED"
# LOAD-BEARING, DO NOT DELETE AS "REDUNDANT WITH THE EXISTENCE PROBES": -f
# accepts an empty file, and sha256 of an empty file equals sha256 of a failed
# pipeline's empty stdin. Guarding it HERE, before mode dispatch, makes that
# collision unreachable by construction rather than by a probe further down --
# which also means no test below exercises the collision any more. Remove this
# line and the false "OK" for a nonexistent ref returns with nothing failing.
[ -s "$VENDORED" ] || die "vendored snapshot is empty: $VENDORED"
head_sha=$(git -C "$ROOT" rev-parse HEAD 2>/dev/null) || die "cannot read HEAD of $ROOT"
recorded=$(grep -oE "^ARG ${ARG_NAME}=[0-9a-f]{40}$" "$DOCKERFILE" | cut -d= -f2 || true)
[ -n "$recorded" ] || die "no 'ARG ${ARG_NAME}=<40-hex>' line in $DOCKERFILE"
sha_of() { sha256sum "$1" | cut -d' ' -f1; }
vendored_sha=$(sha_of "$VENDORED")
upstream_sha=$(sha_of "$ROOT/$REL_PATH")
# Does $REL_PATH exist at HEAD at all? Proven with `cat-file -e` BEFORE
# hashing anything. Piping a failed `git show` straight into sha256sum, as
# this script used to, hashes an EMPTY stream and produces sha256(""): a real,
# collidable value — not a representation of absence. That collapsed "doesn't
# exist" and "exists and happens to be empty" into the same signal, which is
# exactly the defect class the peer review found in --check's at_ref, below.
blob_sha=""
if git -C "$ROOT" cat-file -e "HEAD:$REL_PATH" 2>/dev/null; then
blob_sha=$(git -C "$ROOT" show "HEAD:$REL_PATH" | sha256sum | cut -d' ' -f1)
fi
upstream_dirty=""
if [ -z "$blob_sha" ]; then
upstream_dirty="not present at HEAD (untracked, or absent at this commit)"
elif [ "$blob_sha" != "$upstream_sha" ]; then
if ! git -C "$ROOT" diff --quiet -- "$REL_PATH" 2>/dev/null; then
upstream_dirty="modified but not committed"
elif ! git -C "$ROOT" diff --cached --quiet -- "$REL_PATH" 2>/dev/null; then
upstream_dirty="staged but not committed"
else
upstream_dirty="different at HEAD than in the working tree"
fi
fi
if [ "$MODE" = "check" ]; then
# Resolve the recorded ref the same careful way: existence is proven with
# `cat-file -e` before anything is hashed, and "the ref itself is missing"
# is reported distinctly from "the ref resolves but the path isn't there
# at it" — both used to be silently swallowed into a plausible sha256("").
ref_exists=0
path_at_ref_exists=0
at_ref=""
if git -C "$ROOT" cat-file -e "${recorded}^{commit}" 2>/dev/null; then
ref_exists=1
if git -C "$ROOT" cat-file -e "${recorded}:${REL_PATH}" 2>/dev/null; then
path_at_ref_exists=1
at_ref=$(git -C "$ROOT" show "${recorded}:${REL_PATH}" | sha256sum | cut -d' ' -f1)
fi
fi
printf 'recorded ref: %s\n' "$recorded"
printf 'vendored sha256: %s\n' "$vendored_sha"
if [ "$path_at_ref_exists" = 1 ]; then
printf 'sha256 at ref: %s\n' "$at_ref"
elif [ "$ref_exists" = 1 ]; then
printf 'sha256 at ref: <%s not present at %s>\n' "$REL_PATH" "${recorded:0:7}"
else
printf 'sha256 at ref: <%s not present in this clone>\n' "${recorded:0:7}"
fi
printf 'skillset HEAD: %s (%s)\n' "$head_sha" "$upstream_sha"
if [ -n "$upstream_dirty" ]; then
printf 'live working tree: %s\n' "$upstream_dirty"
fi
rc=0
if [ "$ref_exists" != 1 ]; then
printf 'CANNOT-DETERMINE: %s is not present in %s — fetch, or check against a complete clone\n' "$recorded" "$ROOT" >&2
rc=2
elif [ "$path_at_ref_exists" != 1 ]; then
printf 'MISMATCH: %s does not exist at %s in this clone — the recorded ref cannot be describing these bytes\n' "$REL_PATH" "$recorded" >&2
rc=1
elif [ "$at_ref" != "$vendored_sha" ]; then
printf 'MISMATCH: the vendored snapshot is NOT the file at the recorded ref\n' >&2
rc=1
else
printf 'OK: the vendored snapshot is exactly skillset@%s:%s\n' "${recorded:0:7}" "$REL_PATH"
fi
# Staleness is orthogonal to truthfulness: a record can correctly describe
# an old commit even after upstream has moved on, and a dirty local working
# tree in $ROOT doesn't rewrite git history either — it says nothing about
# whether the RECORDED, committed ref describes the RECORDED, committed
# bytes. Only worth reporting once we already know rc=0 (truthful) — a
# MISMATCH or CANNOT-DETERMINE is the dominant fact and a staleness note
# would only muddy it.
if [ "$rc" = 0 ] && [ "$vendored_sha" != "$upstream_sha" ]; then
# Name the ACTUAL cause. "working tree differs" is wrong when the tree is
# clean and the ref simply moved on — a message that names the wrong cause
# is the same defect class as a canary pinned to a deleted phrase.
if [ "$recorded" != "$head_sha" ] && [ "$blob_sha" = "$upstream_sha" ]; then
# Do not ASSERT which side is newer — test it. Asserting that HEAD is the
# newer side points the operator at a refresh (which costs a ~67-minute
# base rebuild) when the real remedy may be `git pull` in this clone. The
# refresh path below already uses this primitive; reuse it here.
if git -C "$ROOT" merge-base --is-ancestor "$recorded" "$head_sha" 2>/dev/null; then
printf 'NOTICE: %s has moved on to %s; the snapshot describes the older %s (stale, not untruthful — refresh to catch up)\n' \
"$ROOT" "${head_sha:0:7}" "${recorded:0:7}" >&2
elif git -C "$ROOT" merge-base --is-ancestor "$head_sha" "$recorded" 2>/dev/null; then
printf 'NOTICE: %s is BEHIND at %s; the snapshot describes the newer %s — pull this clone, do NOT refresh the snapshot\n' \
"$ROOT" "${head_sha:0:7}" "${recorded:0:7}" >&2
else
printf 'NOTICE: %s (HEAD %s) and the recorded %s have DIVERGED — neither is an ancestor of the other; reconcile the clone before refreshing\n' \
"$ROOT" "${head_sha:0:7}" "${recorded:0:7}" >&2
fi
else
printf 'NOTICE: the working tree of %s differs from the snapshot (HEAD %s)\n' \
"$ROOT" "${head_sha:0:7}" >&2
fi
fi
exit "$rc"
fi
[ -z "$upstream_dirty" ] || die "$ROOT/$REL_PATH is $upstream_dirty — commit it first, or the recorded ref would not describe these bytes"
if [ "$vendored_sha" = "$upstream_sha" ] && [ "$recorded" = "$head_sha" ]; then
printf 'already current: snapshot == skillset@%s\n' "${head_sha:0:7}"
exit 0
fi
# Refuse to silently REWIND provenance. `git checkout <tag>`, a detached HEAD,
# or an older checkout can all leave $ROOT's HEAD behind the already-recorded
# ref; without this guard a refresh there would happily rewrite both the ARG
# and the bytes backwards and report it as an ordinary update.
if [ "$recorded" != "$head_sha" ]; then
if git -C "$ROOT" cat-file -e "${recorded}^{commit}" 2>/dev/null; then
if ! git -C "$ROOT" merge-base --is-ancestor "$recorded" "$head_sha" 2>/dev/null; then
if [ "$FORCE" != 1 ]; then
die "refusing: $ROOT's HEAD ($head_sha) is not a descendant of the recorded ref ($recorded) — this looks like a rewind. Pass --force if this is intentional."
fi
printf 'WARNING: --force set; %s is not an ancestor of HEAD %s — proceeding anyway\n' "${recorded:0:7}" "${head_sha:0:7}" >&2
fi
else
if [ "$FORCE" != 1 ]; then
printf 'CANNOT-DETERMINE: %s is not present in %s (shallow clone?) — fetch full history to verify this refresh moves forward, or pass --force to proceed without that guarantee\n' "$recorded" "$ROOT" >&2
exit 2
fi
printf 'WARNING: --force set; %s could not be resolved in %s — proceeding without verifying forward motion\n' "${recorded:0:7}" "$ROOT" >&2
fi
fi
# Written FROM THE REF, not copied from the working tree, so the pair cannot
# be a lie by construction. Via a temp file so a failed write cannot leave a
# half-vendored snapshot behind.
snap_tmp=$(mktemp)
if ! git -C "$ROOT" show "HEAD:$REL_PATH" > "$snap_tmp" 2>/dev/null; then
rm -f -- "$snap_tmp"
die "cannot read HEAD:$REL_PATH from $ROOT"
fi
chmod 0644 -- "$snap_tmp"
mv -- "$snap_tmp" "$VENDORED"
[ "$(sha_of "$VENDORED")" = "$blob_sha" ] \
|| die "internal: written snapshot does not match HEAD:$REL_PATH"
# In-place, and only the exact pinned line: a broad sed on this Dockerfile
# could rewrite one of the other *_REF ARGs.
tmp=$(mktemp)
sed "s|^ARG ${ARG_NAME}=.*\$|ARG ${ARG_NAME}=${head_sha}|" "$DOCKERFILE" > "$tmp"
chmod 0644 -- "$tmp"
mv -- "$tmp" "$DOCKERFILE"
new_recorded=$(grep -oE "^ARG ${ARG_NAME}=[0-9a-f]{40}$" "$DOCKERFILE" | cut -d= -f2 || true)
[ "$new_recorded" = "$head_sha" ] || die "failed to rewrite ${ARG_NAME} in $DOCKERFILE"
printf 'snapshot: %s -> %s\n' "${vendored_sha:0:12}" "$(sha_of "$VENDORED" | cut -c1-12)"
printf 'ref: %s -> %s\n' "${recorded:0:7}" "${head_sha:0:7}"
printf '\nNOTE: %s is hashed into base_tag, so this costs a base rebuild\n' "$VENDORED"
printf 'on the next tag (~67 min). Also re-pin the phrase canary in\n'
printf 'scripts/smoke-test.sh if the section it names changed.\n'