Compare commits

..

54 Commits

Author SHA1 Message Date
Joakim Persson 6c13f43ac8 release: date the v1.9.3 section for the tag
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 16s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 41m49s
Publish Docker Image / smoke (push) Successful in 8m45s
Publish Docker Image / smoke-studio (push) Successful in 9m8s
Publish Docker Image / build-variant (push) Successful in 17m42s
Publish Docker Image / update-description (push) Successful in 11s
Publish Docker Image / promote-base-latest (push) Successful in 26s
Publish Docker Image / build-variant-studio (push) Successful in 21m28s
CI checks out the tag, so the CHANGELOG heading has to name the version
before the tag exists (same convention as v1.9.1/v1.9.2). check-doc-drift
check 9 was sabotage-tested for exactly this rename and still passes:
everything above the published v1.9.2 heading is the pending text.

Contents of v1.9.3 (all measured, nothing inherited):
  - f25efa0  recreate-sanity-check resolves the ControlPath ssh actually
             uses (ssh -G); the ~/.pi/ssh/config hyphen typo deleted
  - b8d818e  Hub size claims corrected; check 8 gates them against full_size
  - c7d369f  promote-base-latest's conditional re-tag measured (run 669)
  - 50153e6  skill floor -> pi-extensions 25c1265 (task tool + fork-gate)
  - ea89605  changelog for the two floating-ref changes: pi-extensions
             25c1265 and mempalace-toolkit 817b3a8 (mine deadline)
  - cb6d9e5  check 9: what the next build bakes differently must be named
             in the CHANGELOG; first run found pi-obsmem 7b397f4->cba0334
  - b9057fd  LABEL se.jordbo.pi-devbox.mempalace-version (inherited)
  - 960aada  dependency audit 2026-09-19; mempalace 3.10.0 deferred

Floating refs this tag will bake, named per check 9: pi-toolkit 9c87ee8,
pi-extensions 25c1265, mempalace-toolkit 817b3a8, pi-observational-memory
cba0334 (3.1.3). Base REBUILDS (Dockerfile.base + rootfs/ changed), so the
BASE_REBUILD_DATE marker is refreshed now, when it is free -- it had been
stale since 2026-07-13 through two rebuilds.

Deliberately NOT in this tag: mempalace 3.10.0 (see the audit section for
the three measured preconditions).
2026-09-19 17:45:50 +02:00
Joakim Persson 960aada769 changelog: dependency audit 2026-09-19 — mempalace 3.10.0 measured and deliberately deferred
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 17s
Lint / skill-floor (push) Successful in 8s
Lint / doc-drift (push) Successful in 13s
Every row by direct command: npm/PyPI/GitHub release APIs, git ls-remote,
the v1.9.2 config labels via the registry API, and binaries in a running
v1.9.2 container for the floating tools. Only pin with a newer upstream is
mempalace 3.9.0 -> 3.10.0, and it is NOT a small bump: new installs move to
~/.config/mempalace (entrypoint-user.sh tests ~/.mempalace/palace and
entrypoint.sh persists ~/.mempalace, so a fresh container would initialise
outside the volume), and mempalace_event_list flips to newest-first by
default (the toolkit's deriveClosed passes no order, deriveOwed on 2 of 3
calls — server-default-dependent). Deferred to its own release with the
three measured preconditions listed.
2026-09-19 17:38:02 +02:00
Joakim Persson b9057fdc8c Dockerfile.base: LABEL se.jordbo.pi-devbox.mempalace-version, inherited by both variants
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 10s
Lint / doc-drift (push) Successful in 16s
Lint / actionlint (push) Successful in 21s
Closes the blind spot check 9 (cb6d9e5) named: no label recorded the palace
pin, so a MEMPALACE_VERSION bump — the one component whose skew against the
shared central palace is fleet-wide — could ship without a CHANGELOG line.

In Dockerfile.base, not Dockerfile.variant, deliberately: the value sits next
to the ARG that defines it (a copy in the variant is one more pin able to
drift); labels are inherited by every image built FROM the base, so no
build-arg to plumb through the variant's four call sites; and inheritance
means the label states the pin of the base the image ACTUALLY built on,
which is the question when base-decide cache-hits an older base. Both
mechanisms measured on the published v1.9.2 config blob rather than assumed:
maintainer + image.source (set only in Dockerfile.base) are present on the
variant image, and pi-version=0.85.1 is an ARG expanded inside a LABEL.

Intent, like every se.jordbo.pi-devbox.* label; the manifest's
mempalace_version (read from the installed binary) stays the ground truth,
and smoke-test.sh now asserts label == installed core — the one way they
diverge is a base built with INSTALL_MEMPALACE=false or an off-pin install,
both invisible to a label-only check.

check-doc-drift check 9 gains the component (literal, against
ARG MEMPALACE_VERSION in Dockerfile.base); the label-key rule generalises to
"names ending in -version are the label itself". Until a release carries the
label it reports a counted SKIP, not OK — measured: "v1.9.2 carries no
se.jordbo.pi-devbox.mempalace-version label", summary says 1 SKIPPED.

Costs nothing extra: this Unreleased already forces a base rebuild
(50153e6 rootfs/ skill floor). check-base-hash unchanged (no new *_REF).
2026-09-19 17:27:34 +02:00
Joakim Persson cb6d9e5dd0 check-doc-drift: check 9 — what the next build bakes differently must be named in the CHANGELOG
Lint / hadolint (push) Successful in 8s
Lint / skill-floor (push) Successful in 10s
Lint / doc-drift (push) Successful in 14s
Lint / actionlint (push) Successful in 19s
The two most recent Unreleased entries (pi-extensions 25c1265: task tool +
fork-gate; mempalace-toolkit 817b3a8: the mine deadline that never reached
the transport) reached this image through floating *_REF=main ARGs. Neither
produced a diff here, so nothing asked for a CHANGELOG line, and neither had
one until a reader asked. Same class as check 8 — a claim with no in-repo
anchor rots — and the hand practice that covered it (the per-release
"Dependency audit" table, Baked-in vs Upstream-now) is a someone-remembers
mechanism that had lapsed.

Method, no docker/crane/token: last published vX.Y.Z from Hub's tag list
(one request, now shared with check 8); its amd64 config blob via the
anonymous registry API carries one se.jordbo.pi-devbox.<name>-ref label per
component with the SHA the build-args baked. "Would bake" is resolved the way
resolve-versions does it: 40-hex ARG is itself, branch/tag via git ls-remote
with the peeled ^{} form preferred, pi-studio = highest semver tag (label on
<tag>-studio), PI_VERSION as a literal against the pi-version label. Nine
components, ~7.5 s. Unchanged: OK. Moved: the new value's 7-char prefix (tag
for studio, version for pi) must appear above the last published version's
"## " heading — Unreleased plus any not-yet-published section, which is what
the release commit turns Unreleased into, so the tag build passes on the same
text. A published tag with no heading is a failure, not a skip.

First run found a move nobody had recorded: pi-observational-memory 7b397f4
-> cba0334 (6 commits, 3.1.1 -> 3.1.3; memory workers now routed by the
model's exact provider instead of the first registered provider with a
matching api — upstream #70). Named in the CHANGELOG with the honest scope:
no expected effect on the shipped bedrock-only configuration, unmeasured.

Sabotage-tested, expectation written before each run:
  obsmem SHA removed from CHANGELOG ............ rc=1  DRIFT
  Unreleased renamed to "## v1.9.3 — …" ........ rc=0  (release-commit shape)
  "## v1.9.2" heading mangled to v1.9.2-typo .... rc=1  — this one FAILED first:
      \b accepted "v1.9.2-typo" as v1.9.2's heading; now (\s|$)
  bare "## v1.9.2" (no date) ................... rc=0
  ARG PI_OBSMEM_REPO renamed ................... rc=2  (blind gate must not pass;
      read_arg calls are bare top-level assignments on purpose — nested in a
      heredoc's $(...) the exit 2 is swallowed by cat)
  offline (proxy to a closed port) ............. rc=0, 2 SKIPs counted
  SKIP_REF_CHECK=1 ............................. rc=0, 1 SKIP counted
  restored ..................................... rc=0, files byte-identical
GIT_TERMINAL_PROMPT=0 on ls-remote: a repo flipped private fails in 0.3 s as
a SKIP instead of waiting for a username until the job times out.

Header, lint.yml job comment and AGENTS.md all still said this gate "needs
no network" — false since check 8 (2026-09-14). Now: two classes, hermetic
1-7 and published-state 8-9 that SKIP loudly and counted when offline.
Considered and not added: compose <-> .env.example variable cross-check (the
four mismatches are commented-out lines, the mempalace-server compose file's
own variables, and entrypoint-consumed ones — it would fire on nothing
wrong); "documented tag exists on Hub" (check 8 already SKIPs by name; a
hard fail would misreport the tag-to-publish window).
2026-09-19 17:15:12 +02:00
Joakim Persson ea896054df changelog: task tool + fork-gate (pi-extensions 25c1265), and the mine deadline that never reached the transport (mempalace-toolkit 817b3a8)
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 18s
Lint / doc-drift (push) Successful in 6s
Both land in the image through floating refs (PI_EXTENSIONS_REF=main,
MEMPALACE_TOOLKIT_REF=main), so neither shows up as a diff in this repo — the
changelog is the only place a reader of a pi-devbox tag learns that the
container's delegation and feed behaviour changed. Recorded under Unreleased
with the measurements: 5/5 real fork briefs redirected, live pi -p proof for
gate and tool, 2-of-6 → 6-of-6 on the toolkit's new timeout test.

Dockerfile.base: the "Stall protection" comment listed MEMPALACE_MCP_TIMEOUT_MS
as THE tool-call deadline; it now says the feed's mine carries its own.
Comment-only, but Dockerfile.base is hashed into base_tag; the rebuild was
already forced by 50153e6 (rootfs/ skill floor).
2026-09-19 17:00:59 +02:00
Joakim Persson 50153e65b7 skill floor: refresh vendored pi-extensions skill to pi-extensions@25c1265 (task tool + fork-gate)
Lint / hadolint (push) Successful in 11s
Lint / doc-drift (push) Successful in 10s
Lint / skill-floor (push) Successful in 10s
Lint / actionlint (push) Successful in 21s
check-skill-floor.sh: OK, tree_sha256 9b85a633… matches the package at main (25c1265).
Forces one base rebuild (base_tag hashes rootfs/), which is also what bakes
extensions/task.ts and fork-gate.ts into /opt/pi-extensions via the floating
PI_EXTENSIONS_REF=main; install.sh symlinks every extensions/*.ts on start.
2026-09-19 16:42:06 +02:00
Joakim Persson f25efa074d fix(sanity): verify the ControlPath ssh RESOLVES, not the dir the image creates
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 13s
Lint / doc-drift (push) Successful in 12s
Lint / actionlint (push) Successful in 24s
recreate-sanity-check.sh asserted `/tmp/sshcm exists with mode 700` and printed
a green tick while every ssh in the container died rc=255. It was right about
what it checked: the breakage was in the directory a CONFIG named, not the one
the image creates, and the old check could not see the disagreement between them.

Found on the v1.9.2 first boot on emb-7kj4vr4g. A durable ~/.pi/ssh/config,
hand-written into the ~/.pi named volume by the previous session so it would
survive the recreate, declared `ControlPath /tmp/ssh-cm/%C` — with a hyphen.
Nothing here creates that path; the canonical spelling is /tmp/sshcm, identical
in Dockerfile.base, entrypoint-user.sh, recreate-sanity-check.sh and
smoke-test.sh. ControlMaster auto with an unusable ControlPath does not degrade
to an unmultiplexed connection — it fails hard:

    unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory

rc=255, remote command never runs.

Now: resolve the ControlPath ssh itself would use via `ssh -G` and require its
parent to exist and be writable. -G applies real config precedence (first-value-
wins, the system drop-in, Include, an -F override), so it answers "which rule
captured this host" instead of re-implementing the guess, and it never opens a
connection — 0.116s for 48 hosts.

This also puts a check under a caveat documented in prose in Dockerfile.base
("SSH client defaults") and verified nowhere: a per-host ControlPath under a
read-only bind-mounted ~/.ssh gives the identical failure, "Read-only file
system". On the machine this was built on that is 16 of 50 hosts — freeipa-1..6,
gitea.egl.lan, runner-1..3, tor-ms22 — never reported by anything before.

Severity split is deliberate. Default ssh precedence legitimately lands in the
read-only ~/.ssh for any host whose own config pins it there, and the supported
workaround (ssh -F ~/.ssh-local/config, generated every start by
setup-lan-access.sh) already exists, so that is a warn. Failing it would paint
the script red on every run of every device, and a check that fires benignly
every time is one you learn to ignore — the same reasoning that keeps
lint-shell.sh at -S error. The sidecar route is prescribed, so there it is a
hard fail. Host lists cap at six names plus a count: unreadable output is
ignored output.

Teeth proven both directions, each sabotage confirmed by diff BEFORE the result
was believed: ControlPath -> nonexistent dir => rc=1; -> existing-but-read-only
dir => rc=1; restored => rc=0 with the sidecar byte-identical.

No config was added to fix the original problem — the fix was to DELETE
~/.pi/ssh/config, a third hand-maintained copy of what setup-lan-access.sh
already generates from version control on every container start, with fewer
features and one typo. The host-owned ~/.ssh/config is left alone on purpose:
those ~/.ssh/cm paths are correct on the host, where ~/.ssh is writable.

Also lint-shell.sh: the SC2088 count in the severity-choice rationale said 19
and is now 20, with the reproduce command recorded. That command needs a `$ `
prefix — a comment whose first word is "shellcheck" is parsed as a directive,
and the malformed one tripped SC1072/SC1073 at severity error. The gate caught
it on the very commit that introduced it.
2026-09-15 22:59:43 +02:00
joakimp c7d369f28d docs(changelog): measure promote-base-latest's conditional re-tag
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 20s
Closes the open question from v1.9.2's release verification. crane copy DOES
bump Hub's last_updated when it runs (run 669: digests differed, copy ran
17:10:04->17:10:08, base-latest.last_updated moved to 17:10). The
identical-digest path stays unproven by construction -- the job prints
'base-latest already current; nothing to promote.' and never copies -- so a
freshness assertion on the alias is conditional on base-decide's need_build,
not a blanket upgrade. Operational rule lives in the ci-release-watcher skill
(skillset c8034af); the pipeline property is recorded here.
2026-09-14 20:49:23 +02:00
joakimp b8d818ed99 fix(docs): correct the Hub size claims, and gate them so they cannot rot again
Lint / hadolint (push) Successful in 9s
Lint / doc-drift (push) Successful in 13s
Lint / skill-floor (push) Successful in 15s
Lint / actionlint (push) Successful in 18s
DOCKER_HUB.md claimed ~1.1 GB for :latest while Docker Hub served 1.37 GB --
20% wrong, drifting quietly across eight releases, and POSTed to Docker Hub by
update-description every time. It is the first number a stranger reads about
this image.

Root cause is structural, not carelessness: every OTHER claim check-doc-drift.sh
guards is anchored to a build file (a pin, an ARG, a placeholder), so it cannot
rot without someone editing the thing it describes. Nothing in this repo states
the image size, so nothing measured it.

Corrected against measured full_size after v1.9.2 published (amd64/arm64):
  :latest         ~1.1  -> ~1.23 GB   (1.228 / 1.211)
  :latest-studio  ~1.15 -> ~1.25 GB   (1.255 / 1.238)
  :base-latest    ~1.0  -> ~1.17 GB   (1.167 / 1.151)
full_size tracks the FIRST manifest entry (amd64), not the sum across arches --
measured: full_size=1.228, amd64=1.228, arm64=1.211, sum=2.439 -- which is what
the table's per-arch "Size (compressed)" column claims.

New check 8 gates those claims against Hub. Fails on drift; SKIPS LOUDLY via a
new skip() helper (counted, named in the summary) when curl/python3 are absent,
the API is unreachable, or SKIP_SIZE_CHECK=1. A skip is deliberately neither OK
nor a failure: printing an unverified claim as OK is the habit this file exists
to break, and failing on Docker Hub's uptime would hold releases hostage to a
third party. Not wired into hooks/pre-push, so pushes stay offline.

Two bugs caught by writing the expected exit code down before running the check:

  1. The tolerance would have missed its own motivating case. The percentage is
     computed against the MEASURED size, but the 20% first chosen came from the
     claim-relative figure; the real drift was |1.1-1.37|/1.37 = 19.7% and would
     have passed. Now 15%, inside a window with both bounds measured: above the
     largest legitimate skew (11.4%, a claim describing the published release
     while the next tag changes the size) and below the rot it must catch.

  2. A `|| true` on the python invocation made the gate FAIL OPEN -- it printed
     DRIFT and exited 0. Removed; the outer `|| SIZE_RC=$?` satisfies set -e
     without swallowing the code. Verified two-sided: 19% drift exits 1 at the
     default tolerance and 0 at SIZE_TOLERANCE_PCT=25, so the threshold does the
     work rather than the ordering.

NOT covered, and the script says so where a reader will see it: README.md's
~3.2 GB figures are UNCOMPRESSED and the registry exposes compressed sizes only,
so measuring them needs a real pull. A green check 8 says nothing about them.

Verified: check-doc-drift.sh rc=0 on the corrected tree, rc=1 on injected drift,
rc=0 with --warn-only, SKIP+rc=0 on the offline path (proxy to a dead port);
shellcheck clean at default severity (the sed single-quote SC2016 carries a
disable with its reason); lint-shell.sh 16 files clean.
2026-09-14 20:33:40 +02:00
joakimp f5c53b8693 release: date the v1.9.2 section for the tag
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 27s
Publish Docker Image / base-decide (push) Successful in 13s
Publish Docker Image / build-base (push) Successful in 43m59s
Publish Docker Image / smoke (push) Successful in 5m29s
Publish Docker Image / smoke-studio (push) Successful in 5m50s
Publish Docker Image / build-variant (push) Successful in 17m29s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / update-description (push) Successful in 10s
Publish Docker Image / build-variant-studio (push) Successful in 21m50s
CI checks out the tag, so the CHANGELOG heading has to name the version
before the tag exists, not after. v1.9.1's tagged tree carried no
"## Unreleased" either -- same convention, made explicit here.

Contents of v1.9.2 (all measured, nothing inherited):
  - 1baba79  the +131 MB v1.9.1 residual: /root/.npm (110 MB) plus
             @mariozechner/clipboard-* foreign natives at BOTH install
             sites (global + /opt/pi-fork's nested pi-coding-agent copy)
  - 852f900  smoke sentinels that name the residue instead of letting
             the size gate's ~225 MB margin swallow it
  - dab989b  mempalace-toolkit: the owed-withdrawal suite gate
             (node --check never type-strips)
  - 9aaff26  vendored mempalace skill snapshot -> e9e45f7, which is what
             turns the currently-red canary green

Deliberately NOT in this tag: DOCKER_HUB.md's "~1.1 GB" size claim is
low (Hub says v1.9.1 = 1.37 GB, v1.8.14 = 1.23 GB). This release
*changes* the size, so no number is right both before and after it --
writing the predicted ~1.24 GB would be an unmeasured claim in a
user-visible page. Measure post-build, fix on the next tag, and extend
check-doc-drift.sh to gate size claims against Hub's full_size so the
number cannot rot silently again.
2026-09-14 17:55:16 +02:00
joakimp 735565b9be docs: retire the deployed-and-unproven label on isWithdrawn, and name dab989b
Lint / skill-floor (push) Successful in 8s
Lint / hadolint (push) Successful in 17s
Lint / doc-drift (push) Successful in 7s
Lint / actionlint (push) Successful in 20s
Two things, and the second was found by doing the first.

v1.9.1's notes closed with an explicit open item: isWithdrawn had 17 assertions
and 4 mutation kills but had never been exercised on a released image against the
live logstream, which is a different state from untested and a worse one to leave
unlabelled. It has now been exercised, from a container recreated onto v1.9.1, and
the label is retired with the measurement rather than with an assurance.

What the entry records is the instrument as much as the result, because the result
is only as good as the thing that produced it: the shipped extractor was lifted
out of the toolkit's own test script and used to cut the four predicates out of
the BAKED mempalace.ts, and deriveOwed's three queries were replayed with their
exact shipped parameters. No second copy of the logic was written, which is the
same discipline the suite itself is built on. Baseline owed was established by
three independent routes BEFORE any probe was planted, because without a recorded
baseline a withdrawal that appears to work proves nothing. Predictions were
written down before each measurement; the table reports both columns.

The per-predicate attribution is in the entry deliberately. "The count went back
to baseline" is a claim about arithmetic and would have been satisfied by a
pre-filter or by an unrelated reply; isAnswered=false with isWithdrawn=true on a
candidate that was still in the raw set is a claim about mechanism. The negative
controls carry the same weight as the positive one -- a rule that lets a third
party or a prose sentence empty someone's mailbox would have passed a
count-only check.

Also recorded: the rule fires on the real incident it was written for (seq 112),
which is worth more than any synthetic probe, and the one path still unwitnessed
(the extension's own in-process poll, which only fires when the agent is idle) is
named as unwitnessed rather than folded into the pass.

Second: reaching for the suite as corroboration showed it cannot run on this image
at all -- its own precondition gate rejects the source it exists to check, because
`node --check` does not type-strip on any node and v1.9.1 moved to node 24. All 17
assertions had been skipped since the bump. Fixed in toolkit dab989b, named here
per the rule this repo adopted two commits ago after the withdrawal fix shipped
unnamed in v1.9.1. The gate now also distinguishes "cannot run" from "does not
parse", since collapsing those is how the defect pointed readers at the wrong
file.

The rebuild cost is stated and, having been measured, is smaller than the earlier
draft of this entry claimed: 9aaff26 already moved base_tag by refreshing the
vendored skill snapshot under rootfs/, so the base rebuild was pending before this
fix existed and the toolkit pickup rides along with it.
2026-09-14 16:29:52 +02:00
joakimp 9aaff26e3a docs: the withdrawal fix shipped in v1.9.1 unnamed — say so, and re-pin the canary
Lint / hadolint (push) Successful in 11s
Lint / doc-drift (push) Successful in 12s
Lint / skill-floor (push) Successful in 12s
Lint / actionlint (push) Successful in 16s
v1.9.1 bakes mempalace-toolkit e68ee20, which contains e2b060a. Requester-side
ask withdrawal has therefore been LIVE on every v1.9.1 device since 2026-09-10,
while the v1.9.0 section of CHANGELOG.md still read "not yet pinned ... this
image still pins e45f6b4", and the mempalace skill still told every agent at
session start that a withdrawal is impossible.

Measured, two independent routes, expectation recorded before looking:

  - published image label, :v1.9.1-studio and :latest-studio (same digest):
      se.jordbo.pi-devbox.mempalace-toolkit-ref = e68ee2071ca3ad39...
  - ancestry: e2b060a is an ancestor of e68ee20
  - baked mempalace.ts sha 7c16fe14 != v1.8.14's dfca71e9 (different bytes)
  - grep -c isWithdrawn on this container's baked copy = 0 (v1.8.14, so
    tor-ms22 cannot exercise the behaviour it is documenting)

The first label read came back EMPTY, and that empty was a claim about the
request rather than the image: Hub redirects blob fetches to a CDN and curl
without -L returns 0 bytes at exit 0. Recorded in the notes, because a registry
audit reporting "no labels" is missing -L until proven otherwise.

Changes:

  - CHANGELOG Unreleased: the floating-ref mechanism, which is the reusable
    part. ARG MEMPALACE_TOOLKIT_REF=main + CI resolving it to a SHA at build
    time means a release absorbs whatever toolkit main holds, and "what
    behaviour did this image gain" is a question nobody is forced to answer.
    The fleet rule (name the pickup before tagging) was honoured for the
    feed-tick commits v1.9.1 names, and missed for one in the same range.
  - CHANGELOG v1.9.0: the stale caveat is ANNOTATED, not rewritten. The
    sentence is the evidence for how the drift happened; deleting it would
    destroy the only trace. Same discipline as fleet-ops d42412c.
  - CHANGELOG Unreleased: corrected the size entry's "needs no base rebuild"
    claim. True of that entry alone, false of the release now that the
    vendored skill snapshot -- a base_tag input -- moves with it (~67 min).
  - rootfs mempalace skill snapshot 44472 -> 46045 B, refreshed via
    scripts/vendor-mempalace-skill.sh so the bytes and SKILLSET_SNAPSHOT_REF
    (4d7c0ea -> e9e45f7) move together; --check confirms exact match.
  - scripts/smoke-test.sh canary RE-PINNED. The retired pair was still green
    against the new snapshot, i.e. blind to this refresh for the same reason
    the pre-v1.8.13 pair was blind to that one. The replacement is stronger
    than any predecessor here because BOTH witnesses come from the same
    upstream commit: e9e45f7 added "Withdrawing an ask you sent" and deleted
    "nothing anyone can do about it from the other end", the sentence the new
    bullet contradicts. Directions measured against both files, then the
    canary body EXECUTED against each: new -> rc=0 "ok", old -> rc=1 empty.
    A canary whose negative witness was removed by the commit it pins fails
    loudly on stale bytes instead of merely failing to notice them.

Gates: check-doc-drift OK, check-skill-floor OK, vendor --check OK,
hooks/pre-push OK (16 shell files clean at severity error).

Not fixable here: isWithdrawn is deployed and UNPROVEN. 17 assertions and 4
mutation kills, never once exercised on a released image against the live
logstream. Ask routed to a v1.9.1 device.
2026-09-14 15:07:26 +02:00
Joakim Persson 852f900b53 test(smoke): make CI prove the natives still work — the runbook check didn't
Lint / doc-drift (push) Successful in 6s
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 9s
Lint / actionlint (push) Successful in 18m11s
The post-boot check v1.9.1 left for the next machine was

  node -e 'require("esbuild").transformSync("const x:number=1",{loader:"ts"})'

with "if this fails, the prune removed something needed on arm64 -> revert to
v1.8.14 and fix" attached. Run from /workspace on the freshly recreated v1.9.1
image it fails with MODULE_NOT_FOUND — on a perfectly healthy image. require()
resolves by walking up from the CURRENT DIRECTORY, esbuild is nested inside the
two pi trees, and a global install is not on node's require path (NODE_PATH is
unset). The version that "verified the prune" last night only passed because the
shell happened to sit inside the tree. So the check was CWD-dependent, its red
meant nothing, and the remedy it prescribed was reverting a good release.

Path-qualified, both sites compile TS today. Rather than leave a fragile command
in a note for a human to paste, CI now owns it:

  - esbuild must transformSync at EVERY install site found in the image
  - @mariozechner/clipboard must load with its native binding attached at every
    site — for the clipboard prune that IS the proof, since napi-rs resolves the
    platform package at require() time

Sites are discovered with find, so the studio variant's third site is covered
without naming it.

Verified as a four-way matrix, not just a green run: both assertions pass
against the real image FROM /workspace (the cwd that broke the old command), and
both fail (exit 1, "The package \"@esbuild/linux-arm64\" could not be found, and
is needed by esbuild") against copies of the same packages with the host
platform binary removed. An assertion never observed failing is not evidence.

Also documents both failure shapes in AGENTS.md next to the size gate: this one,
and the mode-700 /root one where `test ! -d` passes on a permission error.
2026-09-11 11:06:18 +02:00
Joakim Persson 1baba79c96 fix(size): the v1.9.1 residual was npm's own cache, not the platform binaries
Lint / doc-drift (push) Successful in 6s
Lint / hadolint (push) Successful in 10s
Lint / skill-floor (push) Successful in 9s
Lint / actionlint (push) Successful in 4m26s
v1.9.1's @esbuild prune fixed the 431 MB size-gate failure but still shipped
+131 MB compressed over v1.8.14, nearly all of it in the pi/extensions install
layer (87 -> 206 MB). That leftover was filed as an open item with an explicit
hypothesis — the @mariozechner/clipboard-* family, same npm 11 behaviour, a
different package — and an explicit warning that the hypothesis was not a
measured cause. Measured now, after recreating onto v1.9.1, and the hypothesis
accounted for one sixth of it:

  +110 MB  /root/.npm/_cacache  (35.2 -> 145.3 MB)   the build's npm cache
  + 21 MB  clipboard foreign platform packages, both install sites
  = 131 MB  i.e. the whole delta, no unexplained remainder

Method, since there is no docker CLI inside the container: pulled both variant
layer blobs straight from the registry with a token + manifest + blob fetch and
listed the tarballs (29 789 vs 29 889 entries, 270.8 vs 401.9 MB uncompressed),
then aggregated per package. The file COUNT barely moved, which is what said
"few large files", not "npm installed more packages".

  - purge_build_caches: npm cache clean --force + rm -rf /root/.npm, in the SAME
    layer as the installs, in both the main RUN and the studio RUN. npm 11
    caches every platform tarball it downloads, including the ones the prune
    then deletes, so the cache grew faster than the tree. Nothing at runtime
    reads it: build is root, container is developer with its own cache in $HOME.
  - prune_foreign_esbuild -> prune_foreign_natives: now covers both MEASURED
    families. Clipboard keeps linux-$arch-gnu AND -musl because its napi-rs
    loader picks between them at runtime via its own isMusl() probe; the musl
    package is a 420-byte stub. The bare @mariozechner/clipboard wrapper has no
    hyphen suffix and cannot match the pattern.

Verified on arm64 before writing the glob — a widened rm -rf against a tree you
cannot inspect is the one change shape not to write blind, which is why the
order was update-then-patch. Exercised against a copy of the real trees with
foreign dirs fabricated back in (aix-ppc64, android-arm64, darwin-arm64,
win32-x64, linux-x64): all removed, host linux-arm64 kept at both sites, 21 MB
freed, require('@mariozechner/clipboard') still loads and exports all 18
functions, esbuild.transformSync still compiles TS at both sites. v1.9.1's own
arm64 validation also passed here; CI could only smoke amd64.

Two sentinel assertions, because the size gate did not catch this: it has
~225 MB of margin, so 131 MB of residue stayed green. Both were verified RED
against the running v1.9.1 image and GREEN against a pruned tree. The cache one
refuses to run as non-root: test ! -d /root/.npm on mode-700 /root would
otherwise pass for the wrong reason. Size-failure diagnostics now list cache
paths too — they previously enumerated only node_modules and /opt, where these
bytes were not.

Also fixed: -printf '%f\\n' reaches the shell with both backslashes (confirmed
from the published image's recorded created_by), so v1.9.1's progress line
printed a mangled "li ux-arm64" — find emitted a literal backslash and tr ate
the n out of the name. Single backslash now.

Deliberately not purged: /tmp/node-compile-cache (1.3 MB). The manifest RUN
calls pi --version again, so deleting it earlier only relocates those bytes into
that layer — today's manifest layer is 128 kB precisely because it finds the
cache warm.

Dockerfile.base untouched, so no base rebuild: this rides the next release.
2026-09-11 09:39:54 +02:00
Joakim Persson 42bd29d654 fix(ci): unblock the release — npm 11 esbuild bloat + two self-inflicted assertions
Lint / hadolint (push) Successful in 8s
Lint / skill-floor (push) Successful in 9s
Lint / doc-drift (push) Successful in 10s
Lint / actionlint (push) Successful in 19s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 14s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Successful in 5m4s
Publish Docker Image / smoke-studio (push) Successful in 13m17s
Publish Docker Image / build-variant (push) Successful in 33m49s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 14s
Publish Docker Image / build-variant-studio (push) Successful in 30m58s
v1.9.0 was tagged but never published: smoke failed 90-passed/3-failed and
build-variant needs smoke, so nothing reached the registry. All three are fixed.

1. Size, 431 MB over threshold. Node 24 brings npm 11, which installs EVERY
   @esbuild/<platform> optional binary instead of the matching one: 26 dirs,
   284 MB per pi-coding-agent copy. Measured on pi-fork: npm 10.9.8 -> 165 MB
   (exactly what v1.8.14 shipped), npm 11.19.0 -> 449 MB. npm 11 ignores the
   os/cpu constraints AND --os/--cpu AND an npmrc carrying them, all measured,
   so Dockerfile.variant prunes explicitly, keeping linux-$(node -p
   process.arch) so one line is right on both arches. Pruned in the SAME layer
   as each install, or the bytes survive in the earlier layer. Three sites:
   global pi, pi-fork, pi-studio. Verified esbuild still transforms TS after.

2. om's node_modules assertion tested an npm artefact. om has zero runtime deps;
   its node_modules held ONE file (.package-lock.json) and 20 empty scope dirs.
   npm 11 stopped creating it. Now asserts the entry point pi actually loads,
   read from package.json -> pi.extensions.

3. The skill-source annotation added in v1.9.0 ("baked (package copy)") broke
   the assertion matching "baked$". Pattern now allows an optional suffix; which
   copy shipped stays authoritatively asserted against the manifest + tree hash.

Also: a failed size check now prints the largest layers, largest directories and
an @esbuild sentinel, so this class attributes itself next time instead of
costing a CI dig plus a local npm bisect.

Threshold stays 3800 MB: it caught a real regression and raising it would have
thrown the signal away. No Dockerfile.base/rootfs change, so base-0fb1256c7f99
is reused and build-base is skipped.
2026-09-10 23:39:40 +02:00
Joakim Persson 3a44e81cad feat(ci): gate documentation drift, and make docs a pre-tag release step
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Five doc claims had rotted by v1.9.0, all the same shape: a value written once
by hand, in a file nothing verifies, about a number that lives elsewhere and
moved. README pin table wrong on all three rows; a "Planned" section describing
something already shipped; DOCKER_HUB.md claiming Node v22 against Node 24.

DOCKER_HUB.md is why this is a gate and not a resolution to be careful: it is
PUBLISHED (update-description POSTs it as Docker Hub full_description on every
tag), it had gone eight releases untouched, nothing generates it, and it is read
from the TAG -- so the stale page shipped with v1.9.0 regardless.

scripts/check-doc-drift.sh: seven checks, all repo-local (no network, token,
image, or sibling clone). Exit 0/1/2 matching lint-shell.sh; a renamed ARG is a
red 2, not a green tick. Wired as a fourth lint.yml job so "the docs lie" is its
own red name.

Verified with 15 controls, including two false-positive controls: the first
placeholder check flagged README.md:900, a Go template in a legitimate
`docker inspect --format` example. The gate was wrong, not the doc, so the
pattern is now anchored to the UPPER_SNAKE convention CI substitutes.

Not gated, deliberately: counts/sizes needing a running image (they belong in
smoke-test.sh -- a guessing gate is worse than none), and Dockerfile.base
BASE_REBUILD_DATE, because base_tag hashes that file content-wise and demanding
it be current would force a ~60 min rebuild on releases that touch no base
files. Free during a rebuild, expensive otherwise.

AGENTS.md step 3 rewritten around the mechanism: checkout@v4 with no ref: means
every job reads github.ref, the tag. Docs must be right BEFORE tagging.
2026-09-10 22:01:49 +02:00
Joakim Persson 35964abd01 docs(hub): the Docker Hub page claimed Node v22; v1.9.0 ships Node 24
DOCKER_HUB.md is hand-maintained (no generator: CI only substitutes
{{PI_VERSION}} and POSTs the file as Docker Hub full_description), and it was
last touched at v1.8.6 -- eight releases ago. The Node claim is now false as of
this release, and unlike README.md this file IS published, so a stale claim here
is user-visible rather than internal.

Verified countable claims rather than assuming: "7 user-facing extensions" is
correct (7 files in pi-extensions/extensions/). The "29 mempalace_* tools" claim
is suspect -- this session sees 45 -- but left alone because I cannot attribute
that count to the baked 3.9.0 server without measuring it.
2026-09-10 21:43:17 +02:00
Joakim Persson 6353d59e63 docs(readme): correct the three stale version pins and the shipped-as-planned claim
The version-pin table existed precisely to be the reviewable record of what is
deliberately frozen, and it was wrong on every row: pi 0.84.4 -> 0.85.1,
pi-atelier v0.10.0 -> v0.10.1, mempalace 3.8.0 -> 3.9.0. Verified by parsing the
table and comparing against the ARGs it names rather than by eye.

Also: the "Planned for an upcoming minor release" section listed typst PDF
export, which shipped long ago and even carried a self-contradicting
"(shipped in Unreleased/base)" marker -- the fourth instance of the stale
in-repo Unreleased-pointer class this CHANGELOG already documents. typst 0.15.1
confirmed live in the running image, so the item is now stated as current fact.

The pi-devbox-version sample was v1.5.0-era and structurally outdated: it
predates the palace line the surrounding prose advertises, the pi-atelier
component, and the whole skills: block. Replaced with real observed output
rather than hand-written text.

README has no CI coupling (no workflow or gate reads it; the Docker Hub page
comes from DOCKER_HUB.md), so this cannot affect the in-flight v1.9.0 build.
2026-09-10 21:33:07 +02:00
Joakim Persson 8f0960e134 release: v1.9.0
Lint / skill-floor (push) Successful in 8s
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / lint-gate (push) Successful in 22s
Publish Docker Image / resolve-versions (push) Successful in 15s
Publish Docker Image / base-decide (push) Successful in 8s
Publish Docker Image / build-base (push) Successful in 42m19s
Publish Docker Image / smoke-studio (push) Failing after 6m17s
Publish Docker Image / build-variant-studio (push) Has been skipped
Publish Docker Image / smoke (push) Failing after 9m7s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Rename the Unreleased changelog section to its release heading, matching the
established `## vX.Y.Z — YYYY-MM-DD` format (em dash), per AGENTS.md release
step 3.

Minor rather than patch: CHANGELOG.md scopes minor to "significant base
additions", and this release bumps the Node runtime under every baked JS tool
(pi, agent-browser, playwright, mempalace) from 22 to 24 — the first
NODE_VERSION change since it was introduced at v1.0.0 — alongside five new
base packages (shellcheck, bind9-dnsutils, ldap-utils, xxd, python3-yaml).
Package additions alone have been patch here (v1.8.12, v1.8.13); the runtime
major is what lifts this one.
2026-09-10 21:21:09 +02:00
Joakim Persson 1c905480e3 docs(changelog): note the mempalace feed-tick fix this image carries
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 10s
Lint / actionlint (push) Successful in 19s
The recurring "[mempalace ext] feed (tick) failed: mine timed out after 30000ms"
message is fixed in mempalace-toolkit (309980b + e68ee20) and this image is what
delivers it, since CI resolves MEMPALACE_TOOLKIT_REF to a commit SHA at build
time. Worth a changelog entry rather than leaving it implicit in a ref bump: it
is the most visible symptom operators on this fleet have been living with, and
the entry records that it was a genuine defect (overlapping mines on a
single-writer palace) rather than the cosmetic annoyance it was parked as.
2026-09-10 20:58:42 +02:00
Joakim Persson ff6fd1492a feat(manifest): record WHICH pi-extensions skill copy shipped
Closes the half deliberately left open by cac5e00's skill-floor gate, and the
more important half: "the floor is currently fresh" is a fact with a shelf
life, whereas "the image says which copy it got" keeps working.

The refresh in Dockerfile.variant is guarded by
`[ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates the
co-located skill keeps the vendored floor and still succeeds GREEN, with nothing
in the manifest, labels or logs separating that from a normal build. Afterwards
the two are indistinguishable by inspection -- same path, same filenames, same
permissions -- which is exactly how the floor went unnoticed from 2026-07-30 to
2026-09-10.

build-manifest.json gains pi_extensions_skill_source and
pi_extensions_skill_tree_sha256, MEASURED rather than passed as build-args, per
the ground-truth rule the surrounding block already follows -- and necessarily
so, since the outcome depends on the clone's contents and no ARG could express
it. Three values, because two would force a lie: package (served bytes equal
the clone's skill/), vendored-floor (clone had no skill/ at this ref), and
divergent (both exist but differ -- e.g. the clone ships SKILL.md but not
evaluate-extension-usage.py, so the served directory is a genuine MIX). No OCI
label mirrors these deliberately: LABEL cannot take a RUN-computed value, and a
label fed from an ARG would be the claim-not-measurement being removed here.

Two smoke assertions make the record a gate: the source must be named and be
`package` -- vendored-floor FAILS rather than warns, since these images track
main where the package has shipped skill/ since fa04d20, so a fallback means
the clone did not resolve as intended -- and the tree hash is recomputed over
the served directory, because a recorded hash never recompared is a claim.

pi-devbox-version annotates the line too: "baked (package copy)" normally, or a
yellow "(FALLBACK: vendored floor)". Its existing section reports which copy is
READ at runtime; this is the one fact decided at BUILD time and unrecoverable
later. Old images degrade cleanly -- field absent, jq // empty yields nothing,
line prints plain "baked" as before (verified against this v1.8.14 manifest).

Tested by running the exact logic against this container's real layout, with
the expected value written down before each: package (served == clone),
vendored-floor (clone path absent), divergent (clone lacking the .py while the
served dir has it), and null (empty served dir) -- all four as predicted. The
five pi-devbox-version render branches likewise, including the absent-field
case. Emitted JSON validated with jq for both the populated and null forms.

Gates green: lint-shell.sh (15 files), hadolint 2.15.1, actionlint 1.7.12,
check-base-hash.sh, check-skill-floor.sh, vendor-mempalace-skill.sh --check.
2026-09-10 20:30:16 +02:00
Joakim Persson edc7659add chore(deps): node 22->24, actionlint 1.7.12, hadolint 2.15.1, skillset ref
Audited every component the image obtains OUTSIDE debian/apt. Of ~23, the 19
that resolve `latest` at build time were already current or refresh themselves
on the next rebuild, and the hard pins for pi (0.85.1), mempalace (3.9.0) and
pi-atelier (v0.10.1) were already newest. Four needed a human.

NODE_VERSION 22 -> 24 (LTS "Krypton"). This was a latent defect rather than
housekeeping: agent-browser publishes engines.node ">=24.0.0", so the image sat
BELOW a declared requirement -- v1.8.14 shipped node 22.23.2 with agent-browser
0.37.1, so every build installed it with an npm EBADENGINE warning and ran the
baked browser automation outside its supported range. pi (">=22.19.0") and
playwright (">=20") are satisfied either way. Verified before bumping, since a
missing NodeSource suite breaks every arch at once: setup_24.x returns HTTP 200
and node_24.x advertises `Architectures: amd64 arm64 armhf x86_64`, covering
the arm64 fleet and the amd64 CI runners. Nothing else pinned the node major.

actionlint 1.7.7 -> 1.7.12 and hadolint 2.14.0 -> 2.15.1, each RUN AGAINST THIS
TREE at the new version before being pinned -- both clean, no new findings. A
linter bump is the one dependency update that can turn CI red on unchanged
code, so it is verified locally rather than discovered on a round trip.

SKILLSET_SNAPSHOT_REF e9e09d9 -> 4d7c0ea via scripts/vendor-mempalace-skill.sh,
never by hand: that script is the only thing permitted to write the ARG,
because a cp without a matching bump yields a manifest that confidently lies.
This proved PROVENANCE-ONLY -- the ref was 6 commits behind, but
skills/mempalace/SKILL.md is byte-identical at both (3675bfab), so the snapshot
was already correct and only its recorded origin was stale. No rootfs/ bytes
changed, the smoke-test phrase canary stays valid, and this ARG alone would not
force a base rebuild (the node bump does).

Two measurement traps worth recording, since both would have produced a wrong
answer: GitHub's releases/latest reports pi-atelier v0.10.0 as newest because
v0.10.1 is a TAG WITH NO RELEASE OBJECT -- the pin was already current, and
`git ls-remote --tags` is the instrument that shows it. And gitea-mcp is hosted
on gitea.com, not GitHub, so querying api.github.com returned nothing at all
rather than an error.

Verified with every gate this repo owns, all green, using the NEW linter pins:
lint-shell.sh (15 files), check-workflow-shell.sh, check-base-hash.sh,
actionlint 1.7.12, hadolint 2.15.1, check-skill-floor.sh, and
vendor-mempalace-skill.sh --check.
2026-09-10 20:17:32 +02:00
Joakim Persson cac5e00a31 feat(ci): gate the vendored pi-extensions skill floor, and bake python3-yaml
Follows ecfd2fc, which refreshed the stale floor by hand. A one-off refresh
fixes the symptom; this makes the drift impossible to reintroduce silently.

scripts/check-skill-floor.sh compares the repo floor
(rootfs/usr/local/share/pi-devbox/skills/pi-extensions/) against the package
repo it is a snapshot of, wired in as a new `skill-floor` job in lint.yml.

DIRECTORY hash, not `sha256sum SKILL.md`, using the same tree_sha256 pipeline
Dockerfile.variant uses for skillset_snapshot_tree_sha256 and for the reason
already documented there: a file-only compare answers "did this one file
change", not "is this the same skill". Verified by NEGATIVE CONTROL rather
than asserted -- with SKILL.md left byte-identical and only
evaluate-extension-usage.py edited, the directory check fails (rc=1) where a
file-only compare would have passed. Seven behaviour tests, each with its
expected rc written down before running: in-sync via local dir (0), in-sync
via anonymous remote clone (0), missing --package-dir (2), bad argument (2),
content drift (1), the sibling-file case (1), and --warn-only over drift (0).

Exit codes 0 in sync / 1 drift / 2 cannot-run, matching scripts/lint-shell.sh:
a gate that cannot run must not pass, so an unreachable package repo is a red
2 and never a green tick. A ref with no skill/ is NOT drift -- that is the
documented fallback -- but it emits ::warning:: because it is precisely the
condition under which the floor ships.

Gating on another repo is normally a smell. It is proportionate here because
the check can only fire when skill/ itself changed, which is exactly when the
floor has gone stale; pi-extensions commits that leave skill/ alone cannot
turn this red. It also needs no secret: pi-extensions is anonymously clonable
(verified with `git ls-remote` and no credentials), so it cannot start failing
when a token expires.

Also bakes python3-yaml (552 KB, zero extra deps) into Dockerfile.base. This
is the shellcheck story repeating exactly: scripts/check-workflow-shell.sh --
the guard against the Gitea sh/dash footgun that broke resolve-versions
(ed49b8d) and promote-base-latest (b7197e8) -- hard-exits with "python3 yaml
module missing", so a gate this repo already owns could not be run locally by
anyone. lint.yml installing it explicitly in CI was the evidence. Found while
wiring the job above: the guard could not be run before pushing.

CHANGELOG Unreleased updated for both this and ecfd2fc, including an explicit
note on what is NOT fixed -- the silent-fallback half still has no manifest
flag recording which copy was served.

Verified locally with every gate this repo owns, all green: lint-shell.sh (15
files clean), check-workflow-shell.sh, check-base-hash.sh, actionlint 1.7.7
(pinned, same version as CI), hadolint 2.14.0, and the new check itself.
2026-09-10 19:14:01 +02:00
Joakim Persson ecfd2fc2e5 feat: bake dig/ldapsearch/xxd and refresh the stale pi-extensions rootfs floor
Two changes that share one forced base rebuild, hence one commit.

1. THREE PACKAGES, each closing a capability gap measured during the
   gitea.egl.lan/FreeIPA work on 2026-09-09..10 rather than a preference:

   bind9-dnsutils (~6.1 MB measured) -- dig/host/nslookup were ALL absent,
   so the container could resolve names but had no way to interrogate a
   SPECIFIC nameserver. `getent hosts` only follows the resolver's default
   path, so diagnosing "gateway 172.16.88.1 NXDOMAINs the egl.lan zone
   while 10.20.253.1 is authoritative for it" had to be hand-rolled in
   python3. Split-horizon DNS is a recurring class of bug on this fleet.
   Note the package name: plain `dnsutils` is transitional in trixie.

   ldap-utils (1244 KB, pulls nothing extra) -- the fleet authenticates
   against FreeIPA, yet every LDAP probe had to be run by SSHing to an
   already-enrolled host. Simple binds only; GSSAPI would additionally
   need krb5-user + libsasl2-modules-gssapi-mit, deliberately not added
   as that is a Kerberos-client decision, not a tool.

   xxd (198 KB) -- convenience for verifying git-crypt blob magic in
   myconfigs; `od -c` from coreutils already does the same job.

   netcat-openbsd was in the original proposal and is deliberately NOT
   here: measured redundant, because socat is already baked and bash's
   /dev/tcp does reachability checks with zero packages (verified against
   gitea.egl.lan:3000). Recorded in the Dockerfile so the omission reads
   as a decision rather than an oversight.

2. ROOTFS FLOOR REFRESH: rootfs/.../pi-extensions/SKILL.md was 34284 B,
   unchanged since fa04d20 (2026-07-30), while the canonical package copy
   is 38973 B. Dockerfile.variant copies the fresh package copy over the
   SERVED path at build time but never writes back to this floor, so the
   floor is a silent fallback: if that build-time copy is ever absent it
   ships the July skill with no log line or manifest flag to say which
   version deployed. Refreshed from pi-extensions@c64c122, verified
   byte-identical to both the canonical and the runtime-served copies.

Why one commit: the base_tag hash folds in `cat Dockerfile.base` AND
`find rootfs -type f | xargs cat` (.gitea/workflows/docker-publish.yml),
so either change alone forces the same full base rebuild -- and that
rebuild is precisely what re-bakes rootfs/ as it then stands. Emulating
the workflow hash with a fixed toolkit ref: f3d6462c7416 -> fc4edda03c54.

Verified: scripts/check-base-hash.sh passes (no new ARG *_REF added), and
no shell scripts are touched so the lint-shell gate is unaffected. Sizes
and dependency fan-out measured via apt-get --no-install-recommends
--dry-run on Debian 13 trixie.
2026-09-10 18:56:11 +02:00
joakimp 15a3728ae9 feat: bake shellcheck and add a client-side pre-push lint gate
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 19s
v1.8.14 made shell lint a RELEASE gate (scripts/lint-shell.sh, shared by lint.yml
and the new lint-gate job that resolve-versions depends on), and that script
correctly exits 2 when shellcheck is absent -- "a gate that cannot run must not
pass". Measured on v1.8.14 on 2026-09-09 by three routes (command -v, dpkg -l, a
filesystem search): shellcheck was NOT IN THE IMAGE AT ALL. So the gate could not
be run by a developer in any container, only in CI, and the loop stayed
write-shell -> push -> wait for CI -> discover. That is the loop the gate was
added to shorten, after v1.8.14's first attempt burned ~46 min on a tree whose
lint had already been red for 24 hours.

shellcheck 0.10.0-1 added to the Dockerfile.base apt block: ~39 MB installed
(Installed-Size 40112 KB), measured to pull ZERO additional packages under
--no-install-recommends because libc6/libffi8/libgmp10 are already present.
NOTE this forces one full base rebuild -- base-decide hashes Dockerfile.base +
rootfs/, so unlike a scripts/ change it cannot reuse the existing base- layer.

hooks/pre-push is opt-in per clone (git config core.hooksPath hooks), bypassable
with --no-verify, and execs scripts/lint-shell.sh rather than reimplementing it
-- one copy, because a duplicated check that drifts is the failure this repo
keeps paying for. Matches the idiom skillset/ and myconfigs/ already use.

WHY THIS REPO HAD NO HOOKS, since it was reported as drift and is not: a peer
asked tor-ms22 for core.hooksPath per clone on the premise that unset meant the
gates were unverified there. Measured: pi-devbox unset, skillset hooks, myconfigs
common/hooks, pi-toolkit unset -- but `git ls-files | grep -i hook` is EMPTY in
both pi-devbox and pi-toolkit, so there was nothing to point at on any machine
and unset was the only correct value. This closes the real half for pi-devbox;
pi-toolkit still ships none.

Verified, expected result written down before each check:
  * refusal paths -- shellcheck absent => rc 2 with the remedy named; linter
    missing => rc 2. Never waved through on the assumption CI will catch it.
  * the hook is IN the scan set -- "Checking 14 shell file(s)" with it present,
    13 with it moved aside, so the extensionless file is found by the shebang
    half of the linter's two-signal union. This check exists because the first
    attempt was ambiguous: a planted `[ $UNSET_VAR = "x" ]` was not reported,
    which could equally have meant "not scanned" or "below -S error". It was the
    latter. A count that moves is unambiguous; a clean run is not.
  * it catches the REAL v1.8.14 defect -- planting `echo 'the fleet\'s thing'`
    in hooks/pre-push yields SC1073/SC1072 at severity error, rc=1.

And the gate earned its keep inside this commit: the first version of the
smoke-test assertion carried a comment beginning "# shellcheck is a GATE
DEPENDENCY", and a comment whose first word is the tool's name is parsed as a
DIRECTIVE, not a comment. The new gate failed it with SC1073/SC1072 before the
push -- same family as the v1.8.14 apostrophe, a line that reads as prose to a
human and as syntax to the parser.
2026-09-09 08:57:34 +02:00
joakimp 361babd4fd ci: gate the release on shell lint, from one shared script
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Publish Docker Image / smoke (push) Successful in 7m39s
Publish Docker Image / build-variant-studio (push) Successful in 17m15s
Publish Docker Image / build-variant (push) Successful in 17m58s
Publish Docker Image / update-description (push) Successful in 8s
Publish Docker Image / promote-base-latest (push) Successful in 12s
v1.8.14's first attempt spent ~46 minutes building a base image for a tree whose
own lint had been failing for 24 hours. shellcheck had already flagged the
defect (SC2289, severity error) on the push that introduced it; the lint
workflow went red at run 186 and nobody read it.

lint.yml deliberately skips tag pushes and its reasoning is sound -- the tagged
tree was already linted on main, and a tag-ref lint run sorts above the publish
run, making a release look finished before anything ships. The missing invariant
was never "lint the tag". It was "do not RELEASE a tree whose lint failed", and
only a job inside the publish workflow can enforce that.

So: extract the shell-lint logic from lint.yml into scripts/lint-shell.sh and
call it from both places, then add a lint-gate job that resolve-versions depends
on. resolve-versions is the graph root, so gating it gates everything. Cost is
~40 s at the front of a release; the alternative already cost fifty minutes.

Extracted rather than copied on purpose. A second copy of a check is the drift
this repo keeps paying for -- the same evening produced a skillset mirror that
had sat 9579 B behind its upstream through two consecutive edits.

The script adds one behaviour the inline version lacked: if shellcheck is not
installed it exits 2 rather than silently finding nothing, inheriting the
existing "a gate that cannot run must not pass" rule from hooks/pre-commit in
the skillset repo. Without that, reordering the install step away would turn the
gate into a green tick over zero checks.

Verified locally with a stubbed shellcheck (the real binary is not in the
devbox), five cases, each with its expectation stated first: absent shellcheck
-> rc=2; stub pass -> rc=0 and a non-zero file count; stub fail -> rc=1; a
deliberately unterminated `if` planted in scripts/ -> rc=1 via the bash -n half,
naming the file; removal -> rc=0 again. Discovery cross-checks against CI's own
number: the inline version reported 12 files, the extracted one reports 13, the
difference being lint-shell.sh itself. YAML re-parsed (10 jobs, was 9) with an
assertion that resolve-versions needs lint-gate, and the repo's
check-workflow-shell.sh guard still passes.
2026-09-08 23:41:44 +02:00
joakimp 70e675afee fix(smoke): keep prose out of the single-quoted exec_test body
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 22s
The agent-browser execution guard added on 2026-09-07 carried its explanation
INSIDE the single-quoted script body, and the explanation contained an
apostrophe ("the fleet\'s only recurring amd64 runtime proof"). Inside '...'
bash treats a backslash as literal, so \' does not escape the quote -- it CLOSES
the string. The body truncated at that point and the remaining lines were parsed
by the calling shell.

Consequences, both measured rather than inferred:
  - exec_test received 12 arguments instead of 2 (verified two-sided: the fixed
    tree yields argc=2, HEAD yields argc=12).
  - the leaked `v=$(agent-browser --version)` ran on the CI RUNNER instead of
    inside the image. The runner has no agent-browser, so smoke and
    smoke-studio both failed with "line 770: command not found" after
    build-base had already spent ~46 minutes. Every downstream job was skipped.
  - the truncated body still passed inside the container and printed its green
    tick first, so the log shows a PASS immediately followed by the failure --
    the tick was real, it just no longer covered the assertion.

The prose now sits above the exec_test call, where an apostrophe cannot
terminate anything, and a comment at that spot records why it must stay there.

Not a new failure class: shellcheck flagged it as SC2289 at severity error the
same day, so the lint job has been red since run 186 (2026-09-07 21:21) and was
not read. The gate did its job; nobody looked.
2026-09-08 23:31:48 +02:00
joakimp 601fc98a49 docs(changelog): release v1.8.14
Lint / hadolint (push) Successful in 9s
Publish Docker Image / resolve-versions (push) Successful in 14s
Lint / actionlint (push) Failing after 24s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 50m32s
Publish Docker Image / smoke-studio (push) Failing after 5m13s
Publish Docker Image / build-variant-studio (push) Has been skipped
Publish Docker Image / smoke (push) Failing after 7m40s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Converts the Unreleased section and records what this build carries beyond it:
the mempalace-toolkit bump that makes closing replies reach the mailbox
(deriveClosed, 21023e7 -> e45f6b4), and the L0-L4 subtask documentation landing
via pi-toolkit adfb553 + pi-extensions c64c122.

Notes the mechanism that makes the toolkit fix land at all — the resolved
toolkit SHA is folded into the content-addressed base tag, so the toolkit
moving forces a base rebuild rather than waiting for one — and the consequence
for the amd64 item already in this section: v1.8.13's base was cached, so
Dockerfile.base:607's agent-browser assertion never ran. This base is not
cached, so the native-amd64 proof is finally collected instead of discarded.
2026-09-08 22:09:31 +02:00
joakimp 7e0e66997d docs(env): name MEMPALACE_MAILBOX_NOTIFY — auto-detect cannot work in a container
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Failing after 16s
Unset means the mailbox is silent outside the pi TUI, and the reason is
structural: docker exec does not forward KITTY_WINDOW_ID/TERM_PROGRAM, so
'desktop' detection always falls through to OSC 777, which Kitty does not
implement — the notification then silently does nothing, the worst failure for a
feature whose only job is to break a silence. Documents the four modes, and that
MEMPALACE_MAILBOX_POLL_MS is a FLOOR BETWEEN activity-coupled polls rather than a
wall-clock interval (an idle session polls zero times) — the exact expectation
mismatch reported today.
2026-09-07 21:43:01 +02:00
joakimp 6bd8b79d3a test(smoke): assert agent-browser EXECUTES — it was the discarded amd64 proof
Lint / hadolint (push) Successful in 11s
Lint / actionlint (push) Failing after 23s
Second instance of the same bug class as the node line, in the same file, found
the same way. The agent-browser guard captured the version inside an echo with
2>/dev/null:

  echo "resolved=[$r] version=[$(agent-browser --version 2>/dev/null|head -n1)]" >&2

so the exit code was discarded and a binary that could not execute at all still
PASSED, printing version=[]. Verified two-sided: a stub exiting 127 passes the old
form and is caught by the new one.

Why this exit code matters more than most: smoke runs platforms: linux/amd64 on an
x86 runner, i.e. NATIVE amd64, so this line is the fleet's only recurring amd64
runtime proof for agent-browser's linux-x64 ELF.

NO DEVBOX CAN EVER SUPPLY THAT PROOF. Every machine in the pi fleet is an Apple
Silicon Mac: mbp-m1-2020; tor-ms22 = Mac Studio Mac13,1 M1 Max (fleet-ops
hosts/tor-ms22.md, verified 2026-08-17 with system_profiler); emb-7kj4vr4g =
Apple Silicon, verified 4 routes 2026-09-07. The open "amd64 runtime proof still
needed" ask sent to two devices was asking for the impossible, and emb's reply
naming tor-ms22 as "the only remaining candidate" is wrong for the same reason.
CI had the answer all along and was throwing it away.

Dockerfile.base:607 DOES assert it (`agent-browser --version && \`), but only when
the base rebuilds, and v1.8.13's base was cached — so smoke is where the recurring
gate belongs.
2026-09-07 21:21:16 +02:00
joakimp fabf1274aa docs(changelog): Unreleased section for the smoke node assertion + agent-browser correction
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 19s
Summarises what changed since v1.8.13: the node-major assertion (a bump would
have passed the suite silently), the two-sided verification of the derivation,
and the v1.8.13 agent-browser 0.35.2 -> 0.36.0 correction. No image content
changes; NODE_VERSION still 22.
2026-09-07 21:17:06 +02:00
joakimp 5972a2c535 test+docs: assert the node major in smoke, and correct v1.8.13's agent-browser version
Two findings from a delegated read-only audit of this repo, both verified from the
filesystem before patching.

1. No test asserted the node major, so a node-24 bump would have passed the smoke
   suite SILENTLY. scripts/smoke-test.sh:94 was a bare `run "node" "node --version"`
   — exit-0 and non-empty output only, the printed version compared to nothing —
   while the line above it uses run_expect against $EXPECTED_PI_VERSION for pi. A
   reader skimming the suite would reasonably assume node regressions were covered.
   Worse, this is where the "node v22.23.2 verified" line in the v1.8.13 recreate
   notes came from: printed output, not an assertion.

   Now gated on EXPECTED_NODE_MAJOR, which CI derives from Dockerfile.base's ARG
   NODE_VERSION — the single source of truth (Dockerfile.base:557 is the ONLY hard
   pin in the repo; Dockerfile.variant has no node install at all). That also
   catches a stale cached layer whose node disagrees with the declared ARG.
   Unset => previous behaviour, so this is backward compatible.

   Verified two-sided rather than assumed: the sed derivation yields 22 (empty
   would have silently disabled the assertion, reintroducing the bug); grep -Fq
   "v22." matches v22.23.2; "v24." does NOT match, so a wrong major is caught; and
   "v2." does not prefix-collide. Workflow YAML re-parsed after editing (9 jobs).

2. The v1.8.13 entry claimed "the image's own 0.35.2" for agent-browser. The image
   ships 0.36.0: /usr/lib/node_modules/agent-browser/package.json says version
   0.36.0, engines.node >=24.0.0, and no 0.35.2 exists anywhere in the image. The
   claim was also internally incoherent, contrasting 0.36.0 against a version that
   is not present. Corrected in place with a visible note, since the entry is
   already released. The reasoning survives untouched: the engines floor really is
   vestigial, because /usr/bin/agent-browser is a prebuilt aarch64 ELF invoked
   directly and never through node — which is why 0.36.0 runs fine on 22.23.2.
2026-09-07 21:05:24 +02:00
joakimp aa0fbc5ec0 fix: correct the pi-studio claim — CI publishes v0.9.59, not the v0.9.60-rc.0 label
Lint / actionlint (push) Successful in 17s
Lint / hadolint (push) Successful in 14s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Successful in 4m44s
Publish Docker Image / smoke-studio (push) Successful in 5m8s
Publish Docker Image / build-variant (push) Successful in 15m52s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / build-variant-studio (push) Successful in 21m22s
Measured at the wrong layer during the v1.8.13 audit. I read `ARG
PI_STUDIO_REF=main` in Dockerfile.variant, concluded the release would adopt
main (= v0.9.60-rc.0), set PI_STUDIO_VERSION to that, and wrote a comment plus a
CHANGELOG entry describing deliberate RC adoption. A Dockerfile default cannot
answer "what will CI publish?" when CI overrides it, and it does: build-variant
passes PI_STUDIO_REF=studio_ref and PI_STUDIO_VERSION=studio_tag (lines 598-599
and 787-788), and resolve-versions picks the newest STABLE semver tag via
`^v?[0-9]+\.[0-9]+\.[0-9]+$`, which excludes pre-releases.

Caught by reading run 639's own resolve-versions output rather than the
Dockerfile: studio_tag=v0.9.59, studio_ref=9eed84f = refs/tags/v0.9.59^{}, while
main/v0.9.60-rc.0 is 658536f and never gets built. So published v1.8.13 studio
images carry pi-studio v0.9.59.

ARG restored to `none` rather than pinned to v0.9.59: the local-build default
should not hardcode a tag that goes stale as soon as main moves, which is how the
previous value came to lie. The comment now leads with the override so the next
reader starts at the layer that decides. Upstream's tag-over-main policy is
deliberate (Releases stopped at v0.5.55, main receives half-finished commits), so
adopting an RC from CI would mean changing that filter, not this ARG.

Consequence kept on purpose: the RC's opt-in Studio network binding is in NO
published v1.8.13 image, so it needs no audit this release.

Doc/label-only: base_tag hashes Dockerfile.base + rootfs/** + both entrypoints +
mempalace_toolkit_ref, none of which this touches, so the in-flight base build
(base-ad9faf00f2b2) stays valid and the tag run will reuse it. Verified with CI's
pinned linters: hadolint 2.14.0 exit 0 on both Dockerfiles, actionlint 1.7.7 exit
0, shellcheck 0.10.0 -S error exit 0.
2026-09-06 23:51:00 +02:00
joakimp 702dd71f4c ci: declare workflow_dispatch input types so Gitea renders the dispatch form
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 16s
Gitea (1.26.2) builds the "Run workflow" dialog from each input's `type:`.
With no type declared, the form renders a branch selector and NO input fields,
so a manual run silently takes every default -- and for release_tag: '' that
means env.RELEASE_TAG resolves EMPTY, the variant tag list becomes `<image>:`,
and the run dies on an invalid docker reference only AFTER paying the full base
+ smoke cost (~70 min). Net effect: the `smoke_only` escape hatch documented in
this file's own header has been unreachable from the UI for its entire
existence. Found 2026-09-06 while trying to use it to validate three new smoke
assertions before cutting v1.8.13.

Typed as `string`, deliberately, even though promote_latest/smoke_only read as
booleans: all six consumption sites compare strings against 'true'
(inputs.smoke_only != 'true' at both build-variant gates,
inputs.promote_latest == 'true' at both promote gates) or interpolate into
env.PROMOTE_LATEST. A boolean-typed input yields a real boolean, so `!= 'true'`
would compare across types and could invert a publish gate silently rather than
fail loudly. This keeps the change a pure rendering fix with zero semantic
delta; switching to boolean would require re-auditing all six call sites.

Validated locally with CI's own pinned tools before pushing, because lint is
the only gate on this file: actionlint 1.7.7 exit 0 (clean baseline before the
edit, clean after), shellcheck 0.10.0 -S error exit 0 across all 17 shell
files, and a pyyaml structural check confirming the three inputs still carry
string defaults, the `v*` tag trigger is intact, and all 9 jobs still parse.
2026-09-06 23:40:26 +02:00
joakimp f561acc89a skills: refresh vendored mempalace snapshot a12fe5e -> e9e09d9, re-pin the canary
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Failing after 31m52s
Folded into v1.8.13 at zero marginal cost: the snapshot is hashed into
base_tag, but Dockerfile.base already changed this release, so the ~67 min base
rebuild was already being paid. vendor-mempalace-skill.sh --check reported exit
0 (stale-but-truthful) beforehand, so skipping was sanctioned -- this is the
deliberate call the release checklist asks for. Upstream content: the bare
project-name wing convention and the <harness>@<device> added_by rule, both
downstream of the attribution defect measured on this device 2026-09-06.

The canary re-pin matters more than the refresh. Its old pair ("Provenance is
stamped for you" present / "Attribute what you file yourself" absent) still
PASSED against the new snapshot, so leaving it would have yielded a canary
green on both old and new bytes -- blind to exactly the refresh it exists to
witness, the same false-green family as the pre-v1.8.5 canary. New pair chosen
by measuring direction against both files rather than reading the diff
("Diaries self-heal; plain drawers do not" new=1/old=0; "Agent diaries live in"
new=0/old=1), then tested two-sided: PASS on refreshed bytes, FAIL on the old
bytes recovered from git.

Gates after the change: smoke-test.sh parses, vendor --check exit 0,
check-base-hash exit 0.
2026-09-06 22:31:47 +02:00
joakimp 0d984b1414 changelog: cut v1.8.13 section
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 18s
2026-09-06 22:14:20 +02:00
joakimp adcf56f829 release: audited bumps (pi 0.85.1, mempalace 3.9.0, atelier v0.10.1) + two guards
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 17s
Version audit for the next release. pi 0.84.4 -> 0.85.1, deliberately skipping
0.85.0 (it published internal experimental code and broke SDK imports,
upstream #9132). mempalace 3.8.0 -> 3.9.0. pi-atelier v0.10.0 -> v0.10.1.
PI_STUDIO_VERSION relabelled none -> v0.9.60-rc.0 so the floating main ref's
RC status is visible at docker-inspect time instead of discovered later.
PI_FORK_REF stays floating and adopts e69725c.

The pi bump was verified by running it under a pty in five combinations rather
than by reading the changelog, because this repo has already shipped a version
pair no changelog flagged (atelier < 0.7.1 hangs pi >= 0.84). CPU delta
0.00-0.01s over 5s against a ~5s sustained-CPU hang signature, two-sided via
the atelier sidebar painting identically to the 0.84.4 control.

NODE_VERSION stays 22 on purpose: node 24 is technically safe (pi's five
prebuilt addons are all NAPI, nothing declares a ceiling, agent-browser's
engines.node >=24 is vestigial for the shipped aarch64 ELF), but this release
already moves two minors and bakes an RC, and a node major would leave four
suspects if the image misbehaves. Own release, smoke suite as the gate.

Also corrects a stale claim at the mempalace ARG: synlig serves 3.8.0
server-side, not 3.7.1 (measured over ssh 2026-09-06).

agent-browser volume shadowing: the image has shipped 0.35.2, but every
session on mbp-m1-2020 ran 0.27.0 from a 2026-07-17 hand-install in
~/.pi/npm-global (a VOLUME, at PATH position 2 vs /usr/bin at 8). Third
package hit by this hazard after pi and pi-atelier, so the guard is now
generalised: entrypoint-user.sh retires the copy by moving it aside
(reversible, only when the image ships its own), recreate-sanity-check.sh
asserts resolution under /usr where the volume is real, smoke-test.sh carries
the build-time half and says in the source why it is weak. The real damage was
the stale BUNDLED SKILL (3 skillsets/17.6 KB vs 8/31.5 KB, ten subcommands
undocumented to the agent) - a stale tool errors, a stale skill quietly
teaches wrong commands.

pi-fork capability floor (extensions: []): forks were measured across four
dispatches ignoring their brief, answering in the user's voice, fabricating
self-referential measurements, and once filing a diary entry as agent_name=pi.
Cause is upstream by design - the child gets getHeader()+getBranch(), the
whole active session branch, with the brief as the final user message. Not a
model-capability problem: the same model as the fast profile obeyed the
identical brief perfectly with a fresh session and no inherited context.
extensions: [] runs children with --no-extensions, so the mempalace bridge is
absent and palace writes are impossible by construction (verified by asking a
child to enumerate its tools: read, bash, edit, write). Removes palace writes,
not filesystem writes.
2026-09-06 20:40:02 +02:00
joakimp c8622ece9d skills: correct the credential-incident-response §5 premise about chroma metadata
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 17s
§5 said embedding_metadata.string_value holds "metadata fields only". False,
measured directly: chroma also stores a copy of the document text there, under
key chroma:document. Confirmed with a disposable sentinel drawer (pi@tor-ms22,
2026-08-30): one row in fts_content AND one row in embedding_metadata for the
same drawer.

This was a real mistake in shipped guidance, not a nitpick: this section's own
scanning advice was written to guard against explaining a zero with a
mechanism nobody verified from source, and the section itself did exactly
that -- I downgraded a census to "a floor" on the strength of a metadata-blind
claim I never checked against chroma's actual storage layout. The practical
scan order is unchanged (fts_content is still the direct target, raw bytes are
still the backstop); only the stated REASON for a metadata zero changes: it
needs a different explanation now (key filter, query shape, escaping), not
"structurally absent".

§6's row-gone/bytes-gone claim is upgraded from asserted to measured, same
sentinel: delete_by_source took both fts_content and embedding_metadata 1->0,
raw bytes stayed 4->4 (freed pages persist until VACUUM). Also records the
method that unblocked the measurement: not a better instrument, a disposable
sentinel drawer instead of risking real fleet data.

No image behaviour changes.
2026-09-01 22:37:46 +02:00
joakimp 05843ecfae changelog: reopen an Unreleased section after v1.8.12
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 16s
v1.8.12's release retitled the previous Unreleased heading, leaving the file
with no place to put the next change — so the next contributor either invents a
heading or appends to a released section. The note under it points at the
release checklist step that renames it, so the convention is discoverable from
the file rather than only from AGENTS.md.

Also the first push after moving CI off synlig: lint.yml should now run on
runner-a1 (8 vCPU / 16 GB, Debian 13, upstream Docker CE) instead of the box
that hosts the palace.
2026-09-01 00:29:29 +02:00
joakimp a2846a5f7e release: adopt pi 0.84.4 + pi-atelier v0.10.0, and fix the doc claim the pi bump invalidates
Lint / hadolint (push) Successful in 11s
Lint / actionlint (push) Successful in 21s
Publish Docker Image / resolve-versions (push) Successful in 19s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Successful in 59m39s
Publish Docker Image / smoke-studio (push) Successful in 5m22s
Publish Docker Image / smoke (push) Successful in 17m53s
Publish Docker Image / build-variant-studio (push) Successful in 17m14s
Publish Docker Image / build-variant (push) Successful in 18m22s
Publish Docker Image / promote-base-latest (push) Successful in 12s
Publish Docker Image / update-description (push) Successful in 20s
pi 0.84.3 -> 0.84.4 (no Breaking Changes / Removed heading in that section,
grepped). Adopted for three fixes that land on machinery this fleet runs:
#6879 (large tool results crossing the auto-compaction threshold were sent to
the provider before compacting), #8345 (a resumed session corrupted its next
appended entry when the JSONL lacked a trailing newline -- that file is the
memory feeder's input; measured 49/49 clean here beforehand), and #8537
(triggerTurn:false messages sent mid-run were inserted between a tool call and
its result). The mempalace mailbox is outside #8537's precondition: it delivers
at agent_settled with deliverAs:"steer" and no triggerTurn, and 0.84.4 leaves
the documented steer semantics unchanged.

pi-atelier v0.8.2 -> v0.10.0: two minor releases, both UI-only, no BREAKING
notice. v0.9.0 raises its minimum pi to 0.84.0 and, unlike the
0.7.1-under-pi-0.84 startup-hang precedent, encodes it in peerDependencies
(>=0.84.0). Satisfied by PI_VERSION=0.84.4. Both executable floors compare with
sort -V, so 0.10.0 >= 0.7.1 evaluates correctly.

docs/observational-memory.md: pi's own compaction.md gained one paragraph in
0.84.4 -- autoCompact is now also checked mid-run, after a tool batch's results
are appended. Our text said compaction is checked only when pi goes idle and so
"never interrupts a turn"; that was only ever true of the OM trigger. The
section now states both entry points into session_before_compact and the
diagram carries the second edge (mermaid checker re-run: 6 blocks, 44 labels,
0 soft-wrapped, no cut glyphs at 1280px and 800px).

README: the version-pin table had been wrong since v1.8.6 -- 93f986e moved
ARG PI_VERSION to 0.84.3 and MEMPALACE_VERSION to 3.8.0 and neither table row,
so it advertised pi 0.84.2 / mempalace 3.7.1. Corrected, plus the
--expected-version example that would now fail against a 0.84.4 image.

CHANGELOG: Unreleased retitled v1.8.12 (2026-08-31) with the audits above and a
dependency-audit table -- every other component measured SAME (skillset
snapshot --check OK at a12fe5e, 0 commits since baked).
2026-08-31 07:08:41 +02:00
joakimp 58c22afb04 skills: the fingerprint advice was missing its precondition, and the skill had no section on proving absence
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 23s
Docs only; no image behaviour changes.

WHY THIS AND NOT A PRIVATE NOTE. pi@emb-7kj4vr4g reported itself for printing
sha256[:8] fingerprints of GIT_USER_EMAIL, GIT_USER_NAME and HOST_SSH_USER, and
wrote a private rule forbidding it. It had not broken a rule. It followed §2 of
this skill as written, and §2 is incomplete: it says a fingerprint lets you
compare a credential "without ever materialising the secret" with no condition
attached. When two agents independently make the same mistake, the artifact that
taught them both is the bug.

§2 NOW CARRIES THE PRECONDITION. A fingerprint is 32 bits over its INPUT SPACE,
so publishing fp8(x) hands anyone a MEMBERSHIP ORACLE: they can test x == v for
every candidate v they can generate. Safe for a 40-char random token; a wordlist
for a hostname, username, e-mail, port, path, commit SHA or weak password. "High
entropy" is the usual sufficient condition, NOT the test — a commit SHA is
160-bit and still fully enumerable from the repo. Operationally: if you can
imagine writing the wordlist, you cannot publish the fingerprint. Also added:
candidate fingerprints are working memory and never output (an extractor hashes
hostnames and paths too, so the tempting "print what the scanner saw" debug step
leaks low-entropy fingerprints wholesale), and a plain statement that a
fingerprint register is a CONFIRMATION ORACLE for anyone already holding a
candidate corpus — which is exactly how a retired token is identified in old
transcripts, and works identically for someone else holding those same files.

NEW §6, "Proving absence: instrument strength, and four ways a scan lies clean",
placed next to §5 on purpose: §5 optimises against false POSITIVES, and every
failure in §6 is a false NEGATIVE. Triage optimises precision, a gate optimises
recall, and conflating them is what produced three clean reports over secrets
that were really there. Contents: instrument ranking (exact-byte value search >
class/structure pass > fingerprint census) with the instruction to state which
one produced your zero; census vs class passes as different questions, both
failure modes measured on this fleet; the tokenisation trap where quoting alone
decided detectability; scan the index or pushed tree, never the working tree;
git filters never run on symlinks while check-attr claims they do; two-sided
self-tests that abort, incl. the fixture-interaction artifact; row-gone is not
bytes-gone.

Attribution kept per finding: the census/class split and the instrument
ranking's provenance are pi@emb-7kj4vr4g's; exact-byte search over index blobs
is pi@tor-ms22's. The credential sense of "census" originated in this skill, not
with either agent.

TRAP FOR THE NEXT EDITOR, also in the CHANGELOG: the frontmatter description is
now 1022 of 1024 characters. Trim before adding, or the skill silently fails to
load. Verified by parsing the frontmatter (1022 chars, name intact, every prior
trigger phrase retained).

Deployment: baked skill -> needs an image rebuild AND a container recreate to
reach a running container.
2026-08-30 23:32:53 +02:00
joakimp 30094782df shell: source cli_utils' functions, and install the iproute2 that one of them needs
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 30s
v1.8.11 linked cli_utils' bin/ COMMANDS onto PATH and stopped there. Nothing ever
sourced cli_utils.sh, so its 14 FUNCTIONS were missing from every interactive shell
whose $HOME has no zsh rc -- which is the normal case, not an edge case: the
container's interactive shell is bash and zsh is not installed in the image. A
symlink cannot carry a shell function and a function cannot be reached from a
non-interactive shell, so the two mechanisms are disjoint and both are required.
The image was already paying this layer's dependency cost (fzf, bat, fd, rg, jq are
baked partly FOR these functions) while delivering none of its benefit.

The two changes ship together because they are coupled: portcheck is one of the 14,
and it was a hard stub in every image up to v1.8.11 -- neither ss nor ip nor lsof
nor netstat was present, so it printed "portcheck requires at least one of: ss,
lsof, netstat" and exited. Wiring the functions in without iproute2 would have
shipped a visibly broken one.

MEASURED, not assumed:
  - the loader is bash-safe despite the *.zsh filenames: `bash --noprofile --norc`
    exits 0, defines all 14, and they run (pathls, mkcd, up, extract, agents-sync,
    fhist verified). The tree's one zsh-only construct (print -z in fzf/fhist.zsh)
    is already guarded by [[ -n $ZSH_VERSION ]] with a bash fallback.
  - fresh-$HOME seeding resolves 14/14; CLI_UTILS_SOURCE=0 is honoured; an absent
    checkout is a genuinely silent no-op (no output, no leaked _cu).
  - interactive shell startup 12 ms -> 17 ms.
  - iproute2 is ~5.5 MB (4.2 MB itself + 6 libs under --no-install-recommends;
    libpam-cap is a Recommends and correctly dropped). ss lands at /usr/bin/ss,
    ip at /usr/sbin/ip, both already on the developer PATH, and `portcheck --all`
    then correctly identifies the socat listener on 8765.
  - hadolint clean on both Dockerfiles; repo-wide shellcheck -S error and bash -n
    clean. .bash_aliases is outside CI's discovery (no shebang, not *.sh), so it
    was checked by hand with -s bash at error AND warning level.

Named explicitly per this repo's floating-ref rule: /workspace/cli_utils is a HOST
BIND MOUNT, not a pinned ref, so the image now executes unpinned content in every
interactive shell. Errors are left visible rather than sent to /dev/null so that a
future zsh-only file in that repo is diagnosable rather than mysterious, and
CLI_UTILS_SOURCE=0 is the documented escape hatch. It is deliberately independent
of CLI_UTILS_LINK=0: the two disable independent mechanisms.

Deployment: needs a rebuild AND a recreate. $HOME is the container's writable layer
rather than a named volume (verified -- ~/.bash_aliases carries the container start
mtime while ~/.bashrc carries the image's), so the skel file is re-seeded on every
recreate; a host-bind-mounted ~/.bash_aliases is still never overwritten.
2026-08-30 11:48:03 +02:00
joakimp 9b5783f9dd skills: a sixth instance, found by the repo owner within the hour
Lint / actionlint (push) Successful in 16s
Lint / hadolint (push) Successful in 56s
Claimed "## Unreleased is a new convention in this repo" after reading
CHANGELOG.md once — minutes after b615571 had renamed that very section to
`## v1.8.11`. 33 commits touch the heading; the convention is that new entries
land under Unreleased and the heading is renamed at tag time, exactly as it
was renamed out from under my snapshot.

Same root cause as the five instances the section above already records: an
absence observed in one frame, promoted to a fact about the world, with no
second measurement. `git log -S'## Unreleased' -- CHANGELOG.md` was the oracle
and costs one command.

The generalisation is worth more than the instance, so it goes in the habits
block: to learn a repeating PROCESS, read history, not the file. A file's
current content is one frame of a cycle, and the frame you catch may be the one
where the thing you are looking for has just been consumed.
2026-08-30 10:50:20 +02:00
joakimp d9a7fe101b changelog: an Unreleased section for a rule that was already there
Lint / actionlint (push) Successful in 16s
Lint / hadolint (push) Successful in 1m22s
Records the two skill commits ahead of tomorrow's build, and states the finding
that shaped them: the "a negative result is usually your own filter" rule was
already baked, already symlinked in at every container start, and already
survived every recreate — then was violated five times by a session that had it
available. The gap was activation, not persistence, which is why the
cross-cutting form went into the always-appended AGENTS block instead of into a
skill that only loads when a task description matches.

Also notes what the entry's own subject implies for the reader: neither change
reaches a running container until the image is rebuilt AND the container
recreated, since ~/.agents/skills and the global AGENTS.md both live in the
image rather than in a volume or a mount.
2026-08-30 00:52:36 +02:00
joakimp 36e65fe657 skills: add credential-incident-response, and assert it stays baked
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 18s
Carries the facts a two-day credential incident produced, not the discipline:
probe the issuer FIRST (11 of 13 "exposed" credentials were already dead at the
provider, which cost five HTTP requests to learn and was never checked), the
403-vs-401 trap that scoped tokens introduce into liveness probes, revocation
beats deletion for anything already replicated, the three places a secret hides
in a Chroma palace (FTS content, metadata, raw bytes) in coverage order, scope
derivation from measured consumers, and this fleet's age store with its
single-recipient weakness.

Facts transfer between sessions; exhortations do not — hence a separate skill
for the domain knowledge and a one-line pointer in the always-loaded block.

Authored here, so baked is canonical and it is NOT added to skillset-owned.txt.
Skill dirs are picked up by a glob in entrypoint-user.sh, so no registration is
needed — verified rather than assumed, since an enumerated list would have left
the skill inert, a fitting failure given its subject. Three smoke assertions
extended so a future rebuild cannot silently drop it.
2026-08-30 00:50:11 +02:00
joakimp f0ebea2d98 skills: fix the half of the negative-result rule that was wrong
pi-devbox-environment already warned that "a negative result is usually your
own filter" — baked, symlinked in at every container start, authored by an
earlier session. It survived every recreate, was available all of a later
session, and was violated five times anyway. So the gap was never persistence.

That section also closed with "a positive result needs no such scepticism — it
carries its own evidence." That is false, and it aimed the scepticism budget
one way only. Three of those five errors were positives:

  - an SSH handshake SUCCEEDED and greeted me as joakimp while I believed I was
    probing gitea.egl.lan — `Host gitea*` had rewritten HostName
  - a 401 that was a genuine answer from an issuer which never minted the token
  - a "regression" produced by diffing against a value my own -p 2222 flag set

Adds the three missing false-negative rows, replaces the false claim with the
"a positive result only proves what you actually asked" subsection, and records
the two habits that actually caught these: `ssh -G` to learn which rule
captured a hostname, and declaring the expected result before running a check.

The cross-cutting form goes in pi-global-AGENTS.append.md rather than in a
skill, because it has to fire without a task description matching it — being
loadable on demand is exactly what failed. Across all five errors, none was
caught by re-reading my reasoning; every one was caught by a second measurement
that disagreed.
2026-08-30 00:50:11 +02:00
joakimp b615571913 changelog: release v1.8.11 — the mail nobody could deliver, and the ship that skipped a file
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 19s
Publish Docker Image / base-decide (push) Successful in 10s
Publish Docker Image / build-base (push) Successful in 53m20s
Publish Docker Image / smoke (push) Successful in 4m48s
Publish Docker Image / smoke-studio (push) Successful in 4m58s
Publish Docker Image / build-variant-studio (push) Successful in 17m0s
Publish Docker Image / build-variant (push) Successful in 26m43s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 24s
Names three things riding in via floating refs, since the rule this
fleet adopted (v1.8.9) is to name a behaviour change before tagging,
not after:

- mempalace-toolkit b2b50af -> 21023e7: the rsync --checksum ship fix
  (a real behaviour change to every remote-mode client) and RFC 003
  §7.13 (docs only).
- skillset 6eb20af -> a12fe5e: mermaid-diagrams cutU normalisation and
  the mempalace skill's from_agent identity rule, plus the vendored
  fallback snapshot re-pinned to match (this pi-devbox commit).

Dependency audit run by direct command against every component, not
assumed: pi/mempalace/pi-atelier/pi-studio/pi-toolkit/pi-extensions/
pi-fork/pi-observational-memory all confirmed unchanged since v1.8.10.
pi-studio's apparent SHA drift on ls-remote turned out to be an
annotated-tag-object hash vs its underlying commit hash, not real
drift -- checked with ^{commit} before writing it down as 'none'.
2026-08-27 23:39:39 +02:00
joakimp 495b7e3859 vendor: resync mempalace skill snapshot to skillset a12fe5e
scripts/vendor-mempalace-skill.sh, real refresh not --check: skillset
moved 6eb20af -> a12fe5e (mermaid-diagrams cutU normalisation, and the
from_agent identity rule this same release ships in RFC 003). --check
reported stale-but-truthful (exit 0, the sanctioned skip) but this
release's point is getting today's fixes live fleet-wide, and the base
rebuild is already forced by the entrypoint change and the floating
mempalace-toolkit ref moving -- so the incremental cost of also
bumping this pin is zero. Phrase canary in scripts/smoke-test.sh
unaffected: neither pinned phrase ('Provenance is stamped for you'
present, 'Attribute what you file yourself' absent) is in the section
that changed; both verified still correct in the new snapshot.
2026-08-27 23:36:03 +02:00
joakimp 45850bc973 entrypoint: put back the shell state a recreate eats
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 25s
cli_utils' install.sh reaches PATH by symlinking bin/ into ~/.local/bin.
That is persistent on a host and ephemeral in a container, so the same
installer produced opposite durability and every --force-recreate sent the
human back to typing /workspace/cli_utils/bin/git-status-all. Re-link at
start instead, and add a per-device boot hook so the next question of this
shape needs no image change at all.

Symlinks rather than a PATH edit in an rc file, deliberately: ~/.local/bin
is already ahead of /usr/local/bin in ENV PATH, so links resolve in
NON-interactive shells too (docker exec, agent tool shells, scripts). An
rc-file PATH edit cannot reach those because ~/.bashrc returns early when
not interactive — measured, that asymmetry is exactly why `command -v
git-status-all` failed in one shell and worked in another on the same box.
Shell FUNCTIONS remain the sourced file's job; a symlink cannot carry them.

Guards, because ~/.local/bin is shared: a real file is never clobbered, a
symlink pointing elsewhere is never stolen, ours are refreshed, and links
into a cli_utils/bin whose target vanished are pruned — a dangling link on
PATH reads as a broken container rather than a removed script.

The hook (~/.config/devbox-shell/init.sh) adds no trust boundary: that dir
is already sourced into every interactive shell by the baked bash_aliases,
so it is already arbitrary code from the same owner. Only WHEN it runs is
new. bash <file>, never sourced, exit status ignored, output to a log.

Caught before commit, and the reason the loops use `if` bodies instead of
`&&` chains: under `set -euo pipefail` a for-loop whose last command is a
false test exits non-zero, and with no match the /workspace/*/cli_utils
glob stays literal — so the first draft would have failed to START a
container on every machine that does not have this repo, rather than merely
skipping the links. Re-tested with set -e in place: no-cli_utils/empty-HOME
no-op, guards, idempotence, CLI_UTILS_LINK=0, a hook that exits 7, and a
hook that tries to mutate CLI_UTILS_BIN — all exit 0 with intact state.

Not covered by CI: docker-publish.yml runs only on tags and lint.yml lints
workflow run: steps, so neither executes this file. Validated by extracting
both sections and running them against fixtures, then for real in a live
v1.8.10 container. No smoke assertion added on purpose — the positive path
needs a /workspace mount smoke does not have, and asserting it there would
repeat the v1.8.0 mistake of a smoke check written against a stage that
does not exist at run time.

Moves the base hash (base-decide folds `cat entrypoint.sh
entrypoint-user.sh`), so this rides along with the next tag's ~40-minute
base rebuild rather than justifying a tag of its own.
2026-08-27 21:52:51 +02:00
joakimp 6891dc32b8 changelog: release v1.8.10 — deploy the scrubber, and name the check that must not be skipped
Publish Docker Image / build-variant-studio (push) Successful in 17m4s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 14s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-variant (push) Successful in 18m45s
Publish Docker Image / build-base (push) Successful in 41m59s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / smoke (push) Successful in 4m55s
Lint / hadolint (push) Successful in 8s
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Audited every component against upstream rather than assuming. Only two moved
since v1.8.9:

- mempalace-toolkit 5b8d78f -> b2b50af (ours): feeder scrubber, the symlink fix
  that stops it refusing to stage, hlc owed-set join, queued-delivery note,
  explicit MEMPALACE_MAILBOX_NOTIFY protocol modes, RFC 003 + fleet-memory docs.
- pi-studio v0.9.48 -> v0.9.52 (upstream, studio variant only): 22 commits, all
  additive — PDF previews, hideable header, contextual side questions. No removals
  or renames. Its pi floor is >=0.84.3 against our exact 0.84.3 pin: satisfied,
  zero headroom, named as a watch item because the next floor bump breaks the
  studio job only, after core has already published.

Verified unchanged, with dates proving they predate the v1.8.9 bake: pi 0.84.3 (=
npm latest), mempalace 3.8.0 (= PyPI latest), pi-atelier v0.8.2, playwright 1.62.1
(no drift this cycle despite floating on latest), pi-fork bf702b4, pi-obsmem
ce9fc98, pi-toolkit 0e1369e, pi-extensions 2022887. No breaking changes anywhere,
so everything can ship in one tag.

The reason for tagging now is not features. The scrubber has been live on exactly
ONE device since this morning; every other device has kept staging unscrubbed
transcripts into a shared palace, and cleanup after the fact is manual redaction
(done twice today, with freed-page residue left behind by choice).

The entry leads with the acceptance check instead of burying it, because this
release's worst failure is silent: the feeder is fail-closed, so a packaging or
path mistake stops the fleet's memory feed and nothing complains — refusing to
stage looks exactly like a quiet session. f0bffd1 exists because that very bug was
real (BASH_SOURCE reports the symlink path, not the target). First client on the
new image must confirm a [scrub] summary line appears, the drawer count moves, and
exit 3 did not fire. Silence is the failure signal, not success.
2026-08-27 17:51:46 +02:00
Joakim Persson 8a673ec143 docs: unclip the diagrams, and answer what compaction leaves behind
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 15s
Two problems reported against docs/observational-memory.md, one cosmetic and one
substantive. Both turned out to be worth more than the fix.

Clipping. Several boxes lost their bottom line of text in the viewer, and neither
the source nor my own render showed it. Cause: Mermaid measures a node label with
its own font metrics, commits to a box size, then renders that label as real HTML
in a <foreignObject> — so any host stylesheet touching line-height or font-size
inflates the text past a box that is already fixed, and the overflow is clipped.
The error accumulates per line, which is why it always eats the last line of the
tallest labels.

Rejected the obvious fix after testing it rather than assuming it:
%%{init: {'flowchart': {'htmlLabels': false}}}%% *is* honoured (labels switch from
16 foreignObject to 7 tspan) and clips identically, because inflated font-size
inherits into SVG text too. The fix that works is a hard limit of two short lines
per node, with the detail moved into prose under each diagram — one- and two-line
boxes have the vertical slack to absorb inflation, three- and four-line boxes do
not. The nodes were carrying paragraph-sized text; the diagrams are better for
losing it.

Regression harness: render every block with line-height: 1.7 !important forced
onto the label HTML, screenshot, read it. That caught two survivors of the rewrite
that looked fine in the clean render — a long unbreakable /opt path wrapping to a
third line, and a cylinder shape whose curved bottom leaves less room than a
rectangle for the same two lines.

New §4, because the document explained that compaction folds the ledger and never
said what that leaves in the context. Asked directly: is the session back to
knowing nothing? No. A verbatim tail survives, sized by keepRecentTokens (20k) and
cut only at turn boundaries; the system prompt and AGENTS.md were never in the
compacted region because they are rebuilt from disk each request; nothing is
deleted from disk, since compaction appends a compaction entry rather than
rewriting lines; and recall keeps resolving ids whose sources left the context
because it reads the full branch via sessionManager.getBranch() and never consults
the context window. Repeated compaction renders from live records, not from the
previous summary's prose, so there is no generation-loss spiral.

Also corrects this repo's own "compaction calls no model" to the steady-state
claim it actually is: with an empty ledger the hook returns nothing and explicitly
declines ownership, and pi's native model-based summariser runs. Snippet quoted in
the doc.

And a bug shipped in the first version: §9 said the ledger entries are
custom_message and specifically not custom. Exactly backwards, so the one grep
that section existed to get right was the one it got wrong. Verified empirically
against the live session file — 11 om.observations.recorded and 6
om.reflections.recorded, all "type":"custom", next to "type":"custom_message"
entries whose customType is mempalace-mailbox and mempalace-wakeup. That is where
the confusion came from, and the distinction is load-bearing rather than
cosmetic: the mailbox uses the context-visible append API, om's ledger uses the
invisible one. Which makes it a feature the doc now advertises — the ledger costs
zero context until it is folded.
2026-08-27 14:24:19 +02:00
joakimp cdb6fc0950 changelog: name what the floating toolkit ref will pull into the next tag
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 16s
MEMPALACE_TOOLKIT_REF=main floats and docker-publish.yml resolves it to a SHA at
build time, so toolkit main moving 5b8d78f -> f0bffd1 (10 commits) ships in the
next tagged image whether or not this repo has a commit. v1.8.9 adopted the rule
after that shape bit twice (553d865, 5b8d78f): name the behaviour change BEFORE
tagging. This is that rule obeyed rather than re-learned — the work was pushed to
toolkit main earlier today and this entry was missing, which is exactly the gap
that caused a cross-host misattribution in v1.8.7.

Contents: the feeder-side secret scrubber (three tiers, T3 report-only after a
measured 403 -> 29 false-positive calibration on 52 MB of real transcripts, fail
closed); the symlink near-miss that fail-closed would have turned into a
fleet-wide silent memory outage at bake time, caught before tagging; the mailbox
work (hlc owed-set join, queued-until-next-turn note, explicit notify protocol
modes, tmux path documented unverified); and the documentation set (RFC 003,
fleet-memory.md, secret-hygiene.md). Also states what remains unscrubbed: the
opencode bridge write path and the unbuilt server-side layer.

No tag pushed — per the release protocol, no tag means no build.
2026-08-27 14:20:35 +02:00
Joakim Persson 14371e2da6 docs: explain the memory that runs itself, and fix a claim the mailbox falsified
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 16s
`pi-observational-memory` is baked, registered in the seeded settings.json, and
handed a cheaper model than the session it serves — and the only description in
this repo was five words in a feature list. Someone meeting `/om:status` or a
"compacted memory" block had nothing to read that said whether to leave any of it
on. New docs/observational-memory.md, 272 lines and five diagrams.

Scoped by who is authoritative, so there is one copy of each claim:

- Upstream (/opt/pi-observational-memory/docs/) already documents the mechanism
  well — concepts.md, how-it-works.md, configuration.md, including a v3 lifecycle
  diagram that matches the deployed code. Linked, not re-derived.
- This document takes the four facts pi-devbox owns and can change: the pinned
  commit it bakes (v3.0.4 ce9fc98, matching build-manifest.json), the packages[]
  entry that decides which copy loads, the Haiku-workers-vs-Opus-session split,
  and the devbox-pi-config volume that makes the ledger outlive the container.
- Plus the confusion this image creates by shipping two things called memory: a
  section contrasting it with MemPalace, on the line "observational memory keeps
  a session coherent, the palace keeps the fleet coherent".

Placement follows the split fleet-ops states for itself — reusable mechanism is
not deployment data — so a "why is this in my container" document belongs in the
repo that pins and wires the component. Linked twice from the README, because
until this commit the README referenced docs/ zero times and the file already
sitting there was reachable only by listing the directory.

Numbers were read out of the live container and the baked tree rather than out of
release notes, which caught one thing the pi-extensions skill still has wrong:
the dropper is gated on a successful same-turn reflection, not on a token
threshold of its own.

Also fixes a README sentence that v1.8.9 made false. § Cross-machine agent
coordination ended with "Nothing in this image polls the log on the agent's
behalf"; the mailbox has shipped since aac4a1c. Replaced with the three knobs and
their defaults, derived-not-read owed-ness, and the queued-into-the-next-turn
delivery measured on two devices — and dated to "as baked in v1.8.9
(mempalace-toolkit 5b8d78f)", pointing at RFC 003 §7.11–§7.12 for the mechanism,
because toolkit main is already ahead (a92c75d pings the human who is not
looking) and describing that here would trade a stale-behind claim for a
stale-ahead one.

The cause is worth more than the fix: the behaviour arrived through the floating
MEMPALACE_TOOLKIT_REF, so no diff in this repo ever touched the paragraph making
the claim. v1.8.9's own rule fired for the CHANGELOG and nobody swept the README.
The CHANGELOG records what changed, the README asserts what is true, and only the
first is reviewed at release time — so the rule now extends to grepping the README
for absolute claims (nothing, never, does not, only) about a component whose SHA
moved.

Diagrams were verified by rendering, not by parsing. Both comparison diagrams
parsed clean and rendered with their meaning reversed: Mermaid laid the second
declared subgraph out first, putting "with observational memory" before "without"
and MemPalace before observational memory in the diagram whose entire job was
that contrast. Rebuilt as declaration-ordered chains and re-rendered at mermaid@11
— the version pi-studio pins — in the baked headless browser, then read back as an
image. mermaid.parse() proves syntax and says nothing about layout.
2026-08-27 13:55:55 +02:00
joakimp aac4a1c323 release: v1.8.9 — the version flag that blamed the wrong component
Lint / hadolint (push) Successful in 15s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 41m49s
Publish Docker Image / smoke (push) Successful in 4m51s
Publish Docker Image / smoke-studio (push) Successful in 4m59s
Publish Docker Image / build-variant-studio (push) Successful in 16m58s
Publish Docker Image / build-variant (push) Successful in 28m35s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 17s
Two versions, two flags. `--expected-version` has only ever asserted
`pi --version`, but AGENTS.md step 4 spelled it `X.Y.Z` inside a checklist where
every other X.Y.Z is the pi-devbox tag. Run as documented for v1.8.8 the final
runtime gate of the release printed

    ✗ pi version mismatch: expected 1.8.8, got 0.84.3

and exited 1 — a red accusing the image of being the wrong version. Not one
reader's slip: the v1.8.8 release-readiness handoff from pi@emb-7kj4vr4g
propagated the same wrong spelling twice while correctly calling step 4 "not
ceremonial", so two independent readers converged on it. README.md had it right
all along, which means the two documents disagreed.

- new --expected-image-version asserts the pi-devbox release tag, read from
  release_tag in /etc/pi-devbox/build-manifest.json (no checkout, no network);
  leading `v` optional on either side
- both flags detect being handed the other one's value, and the test is exact
  rather than heuristic: the value is compared against the other quantity the
  image itself reports, so it can only fire on a real mix-up
- neither flag is required now. With none, live `pi --version` is asserted
  against the manifest's pi_version — not a tautology, since a stale pi in the
  ~/.pi/npm-global volume can shadow the baked one, exactly as a stale
  npm:pi-atelier can in packages[]
- the header note replaced was stale and load-bearing: it claimed pi is resolved
  from 'latest' and cannot be self-derived, while Dockerfile.variant pins
  ARG PI_VERSION=0.84.3 and docker-publish.yml reads that ARG as its source of
  truth. The same withdrawn claim also sat in cli_utils' pi-devbox-sanity --help
- argument parsing: a missing value, or a value that is another flag, is a usage
  error instead of silently consuming the next argument; --help works

All fourteen flag combinations exercised by execution, including the two
manifest-absent branches and the shadowing branch a healthy container cannot
reach — mutation-tested with a doctored manifest so each failure branch was
observed firing rather than assumed present.

CHANGELOG also names what no commit here causes: mempalace-toolkit main moved
e70bef2 -> 5b8d78f, so this tag ships the auto-delivered logstream mailbox
because base_tag folds the resolved toolkit SHA. It would have landed either
way; going unnamed is the 553d865 shape that already caused one cross-host
misattribution. Component audit found nothing else to bump — pi, mempalace,
pi-atelier all equal their upstream latest, and every other floating ref
resolves to the commit already baked.
2026-08-26 18:47:03 +02:00
25 changed files with 5861 additions and 124 deletions
+37
View File
@@ -87,6 +87,35 @@ SSH_KEY_PATH=~/.ssh
# MEMPALACE_PI_REMOTE_PATH=/data/feed
# MEMPALACE_PI_DEVICE=
# ── Mailbox notification: MUST BE NAMED, auto-detect CANNOT work here ──
# The mempalace extension polls the logstream for fleet asks addressed to this
# device and queues them into the next turn. That part needs no config. The
# NOTIFICATION that tells the human it happened does, and unset means SILENT
# outside the pi TUI.
#
# Why there is no working default: terminal identity lives in env vars set by
# the emulator (KITTY_WINDOW_ID, TERM_PROGRAM) and `docker exec` does NOT
# forward them — inside the container pi sees only TERM=xterm-256color no matter
# what is rendering it. So "desktop" auto-detection always falls through to
# OSC 777, which Kitty does not implement, and the notification silently does
# nothing: the worst outcome for a feature whose only job is to break a silence.
# Naming the protocol is what makes it fire.
#
# kitty OSC 99 desktop notification (correct for Kitty, incl. over SSH)
# osc777 OSC 777 (tmux/iTerm2/foot and others)
# desktop OSC 99 if KITTY_WINDOW_ID is visible, else OSC 777 — inside a
# container that means effectively always OSC 777, so prefer naming
# 0 / off suppress entirely (in-TUI notify still shows)
# MEMPALACE_MAILBOX_NOTIFY=kitty
#
# Cadence, if the delivery ever feels late: the poll is coupled to session
# activity (it runs when the agent settles), NOT to a wall clock.
# MEMPALACE_MAILBOX_POLL_MS is therefore a FLOOR BETWEEN POLLS (default 300000),
# not a promise of one every 5 minutes — an idle session polls zero times, and
# session start does the first look.
# MEMPALACE_MAILBOX_POLL_MS=300000
# MEMPALACE_MAILBOX_RESURFACE_MS=3600000
# ── LAN access from the container (host-OS-agnostic) ─────────────────
# On VM-backed hosts (macOS OrbStack / Docker Desktop) the container can't
# reach the host's directly-attached LAN peers by default. The entrypoint
@@ -146,6 +175,14 @@ GIT_USER_EMAIL=
# Detection is automatic if the skillset lives at WORKSPACE_PATH/skillset.
# SKILLSET_CONTAINER_PATH=
# ── cli_utils (standalone commands from a mounted checkout) ──────────
# If a cli_utils repo is mounted, the entrypoint symlinks its bin/ commands
# into ~/.local/bin on every start, so they survive container recreate and
# resolve in non-interactive shells too (docker exec, agent tool shells).
# Detection is automatic at WORKSPACE_PATH/cli_utils (or one level below).
# CLI_UTILS_CONTAINER_PATH=
# CLI_UTILS_LINK=0 # disable the linking entirely
# ── Locale ───────────────────────────────────────────────────────────
# LANG=sv_SE.UTF-8
# LANGUAGE=sv_SE:sv
+68 -3
View File
@@ -33,18 +33,39 @@ on:
- 'v*'
workflow_dispatch:
inputs:
# `type:` is REQUIRED for Gitea to render these fields in the "Run
# workflow" dialog. Without it (Gitea 1.26.2) the dispatch form shows a
# branch selector and NO inputs at all, so a manual run silently uses
# every default — which for `release_tag: ''` means RELEASE_TAG resolves
# empty, the variant tag list becomes `<image>:`, and the run dies on an
# invalid reference AFTER paying the full base + smoke cost (~70 min).
# That made the documented `smoke_only` escape hatch below unreachable
# from the UI for its whole existence; found 2026-09-06 trying to use it.
#
# Deliberately `string` and not `boolean`, even though these two read as
# flags: every consumption is a STRING comparison against 'true'
# (`inputs.smoke_only != 'true'` at the build-variant gates,
# `inputs.promote_latest == 'true'` at the promote gates) plus string
# interpolation into env.PROMOTE_LATEST. A boolean-typed input yields a
# real boolean, so `!= 'true'` would compare across types and could
# invert a publish gate rather than fail loudly. Changing the type here
# would mean re-auditing all six call sites; keeping it string is a
# rendering fix with provably zero semantic change.
release_tag:
description: 'Release tag to publish (e.g. v1.0.0). Used only for workflow_dispatch runs.'
required: false
default: ''
type: string
promote_latest:
description: 'Update latest aliases (default true for tag-push, false for manual test runs)'
required: false
default: 'false'
type: string
smoke_only:
description: 'Build base + run both smoke jobs against HEAD, then stop. Publishes nothing. Use to validate smoke assertions without cutting a tag.'
description: 'Build base + run both smoke jobs against HEAD, then stop. Publishes nothing. Use to validate smoke assertions without cutting a tag. Set to the literal string true.'
required: false
default: 'false'
type: string
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
@@ -136,7 +157,41 @@ jobs:
# buildcache silently reuses the layer from whatever pi version was
# current when the cache was first populated. Same class of bug as
# pi-devbox v0.74.0..v0.75.5 (fixed in v0.75.5b 2026-05-23).
# ── release gate ──────────────────────────────────────────────
# Refuse to spend a base build on a tree whose own shell scripts do not lint.
#
# v1.8.14's first attempt is why this exists. smoke and smoke-studio both failed
# at scripts/smoke-test.sh:770 AFTER build-base had already spent ~46 minutes,
# on a defect shellcheck had flagged as SC2289 (severity error) a day earlier:
# the lint workflow went red on the very push that introduced it (run 186) and
# stayed red for runs 187 and 188, unread.
#
# lint.yml deliberately does not run on tag pushes, and its reasoning is sound
# (the tagged tree was already linted on main; a tag-ref lint run sorts above
# the publish run and makes a release look finished before anything ships). The
# missing invariant was never "lint the tag" -- it was "do not RELEASE a tree
# whose lint failed", and only a job inside THIS workflow can enforce that.
#
# ~40 s, ahead of everything expensive, and it runs scripts/lint-shell.sh --
# the same file lint.yml calls, not a second copy that drifts.
lint-gate:
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
steps:
- uses: actions/checkout@v4
- name: Install shellcheck
run: |
apt-get update
apt-get install -y --no-install-recommends shellcheck
- name: "Shellcheck + syntax-check repository scripts (severity: error)"
run: bash scripts/lint-shell.sh
resolve-versions:
# Gated: a defective tree must not reach a 46-minute base build.
needs: [lint-gate]
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
@@ -522,7 +577,12 @@ jobs:
env:
EXPECTED_PI_VERSION: ${{ needs.resolve-versions.outputs.pi_version }}
EXPECTED_MEMPALACE_VERSION: ${{ needs.resolve-versions.outputs.mempalace_version }}
run: bash scripts/smoke-test.sh pi-devbox:smoke
run: |
# Single source of truth for the node major is Dockerfile.base's ARG.
# Asserting the BUILT image matches it also catches a stale cached layer.
EXPECTED_NODE_MAJOR=$(sed -n 's/^ARG NODE_VERSION=\([0-9][0-9]*\).*/\1/p' Dockerfile.base)
export EXPECTED_NODE_MAJOR
bash scripts/smoke-test.sh pi-devbox:smoke
# ── Phase 3b: amd64 smoke for the studio variant ────────────────────
# Additive + independent of the core `smoke` job: gates ONLY
@@ -585,7 +645,12 @@ jobs:
env:
EXPECTED_PI_VERSION: ${{ needs.resolve-versions.outputs.pi_version }}
EXPECTED_MEMPALACE_VERSION: ${{ needs.resolve-versions.outputs.mempalace_version }}
run: bash scripts/smoke-test.sh pi-devbox:smoke-studio
run: |
# Single source of truth for the node major is Dockerfile.base's ARG.
# Asserting the BUILT image matches it also catches a stale cached layer.
EXPECTED_NODE_MAJOR=$(sed -n 's/^ARG NODE_VERSION=\([0-9][0-9]*\).*/\1/p' Dockerfile.base)
export EXPECTED_NODE_MAJOR
bash scripts/smoke-test.sh pi-devbox:smoke-studio
# ── Phase 4: multi-arch publish ─────────────────────────────────────
build-variant:
+92 -27
View File
@@ -75,31 +75,11 @@ jobs:
# are shell scripts with no extension. -print0/mapfile -d '' so a path
# with a space cannot silently split, and the file count is asserted
# non-zero — a green tick over an empty file set is not a check.
run: |
# Union of two signals, because either alone misses a real case:
# a shebang scan misses a sourced fragment with no shebang, and a
# *.sh glob misses the extensionless tools in rootfs/usr/local/bin/.
# Silent skipping is precisely the failure mode this gate exists to
# prevent, so err toward over-collecting.
mapfile -d '' -t all_files < <(find . -not -path './.git/*' -type f -print0)
sh_files=()
for f in "${all_files[@]}"; do
case "$f" in *.sh) sh_files+=("$f"); continue;; esac
if head -n1 "$f" 2>/dev/null | grep -qE '^#!.*\b(bash|sh)\b'; then
sh_files+=("$f")
fi
done
echo "Checking ${#sh_files[@]} shell file(s)"
if [ "${#sh_files[@]}" -eq 0 ]; then
echo "::error::no shell files found — the shebang scan or the checkout is wrong"
exit 1
fi
shellcheck -S error -f gcc "${sh_files[@]}"
rc=0
for f in "${sh_files[@]}"; do
bash -n "$f" || { echo "::error file=$f::bash -n failed"; rc=1; }
done
exit "$rc"
#
# The implementation moved to scripts/lint-shell.sh on 2026-09-08 so the
# release gate in docker-publish.yml runs the SAME code rather than a
# second copy that drifts. Edit the script, not a copy of it.
run: bash scripts/lint-shell.sh
- name: Gitea shell guard (catches the actionlint blind spot)
# actionlint models GitHub Actions, where the default run shell is
@@ -112,7 +92,7 @@ jobs:
- name: Install actionlint (pinned)
env:
ACTIONLINT_VERSION: 1.7.7
ACTIONLINT_VERSION: 1.7.12
run: |
curl -fsSL \
"https://github.com/rhysd/actionlint/releases/download/v${ACTIONLINT_VERSION}/actionlint_${ACTIONLINT_VERSION}_linux_amd64.tar.gz" \
@@ -147,7 +127,7 @@ jobs:
- name: Install hadolint (pinned)
env:
HADOLINT_VERSION: 2.14.0
HADOLINT_VERSION: 2.15.1
run: |
curl -fsSL \
"https://github.com/hadolint/hadolint/releases/download/v${HADOLINT_VERSION}/hadolint-Linux-x86_64" \
@@ -157,3 +137,88 @@ jobs:
- name: Run hadolint
run: hadolint Dockerfile.base Dockerfile.variant
skill-floor:
# Gate the VENDORED pi-extensions skill snapshot in rootfs/ against the
# package repo it is a snapshot of. Its own job rather than a step in
# `actionlint`, so "the floor is stale" is a distinct red name in the runs
# list instead of being buried in a lint job that is about something else.
#
# The gap it closes, measured 2026-09-10: the floor sat at 34284 B, untouched
# since fa04d20 (2026-07-30), while the package copy was 38973 B.
# Dockerfile.variant copies the fresh package copy over the SERVED path but
# never writes back to the floor, so nothing in the repo ever noticed. That
# matters because the floor is a FALLBACK: the copy is guarded by
# `if [ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone
# yields no skill/ ships the vendored snapshot and still goes green, with no
# manifest flag or label saying which copy was served.
#
# Gating on another repo is normally a smell; it is proportionate here
# because the check compares the skill DIRECTORY hash, so it can only fire
# when that directory actually changed — which is exactly when the floor has
# gone stale. pi-extensions commits that leave skill/ alone cannot turn this
# red. No secret is needed either: the repo is anonymously clonable (verified
# 2026-09-10 with `git ls-remote` and no credentials), so this cannot start
# failing when a token expires.
#
# Exit codes are 0 in sync / 1 drift / 2 cannot-run, matching
# scripts/lint-shell.sh: a gate that cannot run must not pass, so an
# unreachable package repo is a red 2 rather than a green tick.
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
steps:
- uses: actions/checkout@v4
- name: Vendored pi-extensions skill floor matches the package
run: bash scripts/check-skill-floor.sh
doc-drift:
# Gate hand-maintained doc claims against the build files they describe.
# Its own job for the same reason as skill-floor: "the docs lie" should be a
# distinct red name, not a line buried in a job about workflow syntax.
#
# The gap it closes, measured 2026-09-10 while preparing v1.9.0 — five
# claims had rotted, every one of them a fact written by hand in a file
# nothing verified:
# * README.md's "Version pins" table was wrong on ALL THREE rows (pi
# 0.84.4 vs 0.85.1, pi-atelier v0.10.0 vs v0.10.1, mempalace 3.8.0 vs
# 3.9.0) — and that table exists specifically to be the reviewable
# record of what the repo freezes on purpose, so a wrong row destroys
# the only thing it is for.
# * README.md listed already-shipped typst PDF export under "Planned for
# an upcoming minor release", marked "(shipped in Unreleased/base)".
# * DOCKER_HUB.md claimed "Node.js v22" while v1.9.0 ships Node 24.
#
# DOCKER_HUB.md is why this is a gate and not a habit. It is PUBLISHED —
# update-description POSTs it to Docker Hub as full_description on every tag
# — and it had gone eight releases (v1.8.6 -> v1.9.0) untouched. Nothing
# generates it and nothing checked it, so the only thing keeping it true was
# someone remembering. It is also read from the TAG, so a fix pushed to main
# after tagging never reaches the published page.
#
# Two classes of check. 1-7 are hermetic: each compares a doc string against
# a value that exists in this repo — no network, no token, no built image.
# 8-9 compare against what is PUBLISHED, because those claims have no
# in-repo anchor and rotted for exactly that reason: 8 reads Docker Hub's
# measured sizes; 9 reads the ref labels baked into the last released image
# (anonymous registry API, no docker/crane) and `git ls-remote`s each
# floating upstream, then requires every component the next build would
# bake differently to be NAMED in the CHANGELOG above that release's
# heading. Both SKIP loudly and counted when offline — a skip is neither OK
# nor a failure. Claims that genuinely need a running container (the "N
# mempalace_* tools" count, uncompressed sizes) are still left out — a gate
# that cannot evaluate a claim honestly would have to guess, and a guessing
# gate is worse than none. Assert those in scripts/smoke-test.sh instead.
#
# Exit codes 0 in sync / 1 drift / 2 cannot-run, matching lint-shell.sh and
# check-skill-floor.sh. A renamed ARG makes the gate blind, so that is a red
# 2, not a green tick.
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
steps:
- uses: actions/checkout@v4
- name: Doc claims match the build files
run: bash scripts/check-doc-drift.sh
+92 -6
View File
@@ -92,15 +92,78 @@ re-brand of opencode-devbox's `pi-only` variant.
is hashed into `base_tag`, so it costs a base rebuild (~67 min); and if the
section the phrase canary names has changed, re-pin it in
`scripts/smoke-test.sh`.
3. Update `CHANGELOG.md` Unreleased → vX.Y.Z section.
3. **Update the docs this release makes stale — BEFORE you tag.** Rename
`CHANGELOG.md`'s `## Unreleased` to `## vX.Y.Z — YYYY-MM-DD` (em dash, as
every prior release heading uses), then run the gate:
```bash
bash scripts/check-doc-drift.sh # 0 in sync / 1 drift / 2 cannot run
```
It compares README.md's version-pin table against the ARGs it names, and
DOCKER_HUB.md's Node claim against `ARG NODE_VERSION`, plus Hub's
25 000-char limit, unsubstituted `{{PLACEHOLDERS}}`, and stale `Unreleased`
pointers in user-facing docs. With network it also checks DOCKER_HUB.md's
size claims against Hub's measured sizes (check 8) and — check 9 — that
**every component the next build would bake differently from the last
published release is named in the CHANGELOG** above that release's heading:
it reads the `se.jordbo.pi-devbox.*-ref` labels off the published image and
`git ls-remote`s each floating `*_REF`. A red check 9 means an upstream
(pi-toolkit, pi-extensions, mempalace-toolkit, pi-fork,
pi-observational-memory, pi-studio) moved and no entry names the new SHA;
the failure prints the compare URL; a `PI_VERSION` or `MEMPALACE_VERSION`
bump is caught the same way via the `pi-version` / `mempalace-version`
labels. Name the 7-char SHA (or version) where you describe
the change — that is what the old "Dependency audit" tables recorded by
hand, now required.
**Why before and not after:** `docker-publish.yml` runs `actions/checkout@v4`
with no `ref:`, so every job reads `github.ref` — the **tag**. A doc fix
pushed to `main` after tagging does not reach the release, and for
`DOCKER_HUB.md` it does not reach the published Hub page either, because
`update-description` POSTs that file as Docker Hub's `full_description` from
the tag's tree. Getting it in afterwards means re-pointing the tag, which is
its own hazard (v1.8.14 went `601fc98` → `361babd` and broke deploy
verification until `git fetch --tags --force`).
The gate is deliberately narrow — it only checks claims verifiable from files
in this repo. Still eyeball, because these are NOT gated:
- counts and sizes (`~1.1 GB`, "N `mempalace_*` tools", "7 extensions") —
they need a running image; assert them in `scripts/smoke-test.sh` instead
- feature prose that quietly became false, e.g. a "Planned for an upcoming
release" section describing something that already shipped
- `Dockerfile.base`'s `# BASE_REBUILD_DATE:` marker. Ungated on purpose:
`base_tag` hashes Dockerfile.base's content, comments included, so
demanding it be current would force a ~60 min base rebuild on a release
that touched no base files. **Fix it when the base is already rebuilding —
then it is free.**
Measured cost of skipping this, 2026-09-10 (v1.9.0): five stale claims, one
of them published. README's pin table was wrong on all three rows, and
DOCKER_HUB.md — untouched for eight releases — still said Node v22 while the
image shipped Node 24.
4. Verify `docker compose up` works locally with the current `latest` image
if you're upgrading users from a previous version. Then run the
**post-recreate sanity check** inside the running container to confirm
persisted volumes survived and the pi runtime wiring re-deployed (not just
that the container booted):
`docker compose exec devbox bash scripts/recreate-sanity-check.sh --expected-version X.Y.Z`
(or just `pi-devbox-sanity --expected-version X.Y.Z` if `cli_utils/bin` is
on PATH). This is the runtime peer of the build-time `smoke-test.sh` gate.
`docker compose exec devbox bash scripts/recreate-sanity-check.sh --expected-image-version X.Y.Z`
(or just `pi-devbox-sanity --expected-image-version X.Y.Z` if
`cli_utils/bin` is on PATH). This is the runtime peer of the build-time
`smoke-test.sh` gate.
**`X.Y.Z` here is the pi-devbox release tag** you are shipping (e.g.
`1.8.9`), which is what the rest of this checklist means by `vX.Y.Z`.
`--expected-image-version` is the flag that asserts it. There is also an
`--expected-version`, and it means something else — the **pi coding agent**
version (e.g. `0.84.3`, the `ARG PI_VERSION` pin). Handing the release tag
to that one used to report *"pi version mismatch: expected 1.8.8, got
0.84.3"*, i.e. a red on the final gate of the release accusing the wrong
component; it now tells you to use `--expected-image-version` instead, and
the reverse mix-up is caught too. Both flags are optional — with neither,
the live pi version is asserted against the version recorded in the image's
own build manifest (which catches a stale `pi` in the `~/.pi/npm-global`
volume shadowing the baked one) and the image tag is reported
informationally.
5. Push tag: `git tag vX.Y.Z && git push origin vX.Y.Z`.
6. Watch CI: smoke job builds amd64 only and asserts size + extensions +
pi version + new-base-tooling presence. Variant build is multi-arch
@@ -229,8 +292,10 @@ shipped the same image bytes); preventatively fixed for `PI_VERSION` +
image. Verifies binaries, repo clones, runtime deployment (waits for
keybindings + mempalace bridge + ≥4 extensions before sampling — fixes
the parallel-build-load race documented in opencode-devbox c6f9d11
2026-06-08), and image size threshold (3500 MB; revisit after a few
releases as actuals settle).
2026-06-08), build-time leftovers (see below), and image size threshold
(3800 MB in `SIZE_THRESHOLD_MB`; revisit after a few releases as actuals
settle — this doc said 3500 until 2026-09-11, after the bar had already
moved twice).
If smoke fails on size threshold but build is otherwise fine: bump
`SIZE_THRESHOLD_MB` in scripts/smoke-test.sh in a follow-up commit and
@@ -238,6 +303,27 @@ re-run. The threshold exists to catch *runaway* growth (an accidental
texlive bake-in, a forgotten chrome dependency), not to block ordinary
upstream bumps.
**The size gate is not a substitute for naming the residue.** It carries
~225 MB of deliberate margin, so v1.9.1 shipped +131 MB of pure build
residue — 110 MB of it npm's own download cache under `/root/.npm`, the
rest foreign platform packages — and stayed green. Four named assertions
now cover that ground: no foreign npm-11 platform packages beyond the
host arch (`@esbuild/*`, `@mariozechner/clipboard-*`), no `/root/.npm` in
the image, and — because the prune's real risk is *removing something
needed*, not size — esbuild must compile TS and clipboard must load its
native binding at **every** install site.
Two failure shapes to copy from those, both of which bit here:
- `test ! -d /root/.npm` on mode-700 `/root` passes for a **permission**
error, so the cache assertion refuses to run as non-root. Watch for
this in any assertion about a path you may not be allowed to read.
- `node -e 'require("esbuild")'` resolves by walking up from the CURRENT
DIRECTORY, so it fails with `MODULE_NOT_FOUND` from `/workspace` on a
perfectly healthy image (esbuild is nested inside the pi trees;
`NODE_PATH` is unset). Always path-qualify: `require("<abs>/esbuild")`.
A runbook shipped the bare form with "if this fails, revert the
release" attached, and it duly went red for the wrong reason.
## Build pipeline notes
- **Two-phase**: base + variant. Base is rebuilt only when
+2315
View File
File diff suppressed because it is too large Load Diff
+5 -5
View File
@@ -8,12 +8,12 @@ A self-contained Docker container for the [pi coding-agent](https://github.com/e
| Tag | Architectures | Size (compressed) | What you get |
|---|---|---|---|
| `joakimp/pi-devbox:latest` | amd64, arm64 | ~1.1 GB | Self-contained: base + pi `{{PI_VERSION}}` + companions |
| `joakimp/pi-devbox:latest` | amd64, arm64 | ~1.23 GB | Self-contained: base + pi `{{PI_VERSION}}` + companions |
| `joakimp/pi-devbox:vX.Y.Z` | amd64, arm64 | same | Pinned semver release |
| `joakimp/pi-devbox:latest-studio` | amd64, arm64 | ~1.15 GB | `latest` + [pi-studio](https://github.com/omaclaren/pi-studio): browser prompt editor, KaTeX/Mermaid preview, tmux-backed literate REPLs |
| `joakimp/pi-devbox:latest-studio` | amd64, arm64 | ~1.25 GB | `latest` + [pi-studio](https://github.com/omaclaren/pi-studio): browser prompt editor, KaTeX/Mermaid preview, tmux-backed literate REPLs |
| `joakimp/pi-devbox:vX.Y.Z-studio` | amd64, arm64 | same | Pinned semver studio release |
| `joakimp/pi-devbox:base-latest` | amd64, arm64 | ~1.0 GB | Base layer alias (internal building block; pull `:latest` instead) |
| `joakimp/pi-devbox:base-<hash>` | amd64, arm64 | ~1.0 GB | Content-addressed base; immutable. Stable parent for variant rebuilds. |
| `joakimp/pi-devbox:base-latest` | amd64, arm64 | ~1.17 GB | Base layer alias (internal building block; pull `:latest` instead) |
| `joakimp/pi-devbox:base-<hash>` | amd64, arm64 | ~1.17 GB | Content-addressed base; immutable. Stable parent for variant rebuilds. |
> **pi-studio (`-studio` tags):** launch with `/studio --no-browser --port 8765` inside a pi session. The server binds `127.0.0.1` **inside the container**, so reach it via host networking or a loopback bridge (and `ssh -L` for a remote host; mosh needs a parallel `ssh -L`). Full recipe: [README → Using pi-studio](https://gitea.jordbo.se/joakimp/pi-devbox#using-pi-studio--studio-variant).
@@ -94,7 +94,7 @@ The entrypoint deploys/registers all of these on first container start. Re-runni
uv run --with jupyterlab jupyter lab --no-browser --port 8888
uv run --with marimo marimo edit
```
- **Node.js** v22 + npm (used by pi itself)
- **Node.js** v24 LTS + npm (used by pi itself)
- **Rust** — `rustup-init` is on PATH; install toolchains on demand
- **Go** — opt-in via `--build-arg INSTALL_GO=true` if rebuilding from source
+148 -10
View File
@@ -14,7 +14,7 @@
# content-addressed over this file, so any byte change invalidates the
# cache. Recommended cadence: once per release for security updates.
#
# BASE_REBUILD_DATE: 2026-07-13 (Unreleased — agent-browser CLI + Playwright Chromium for headless browser automation; prior: typst PDF engine + xz-utils + pandoc typst-template default-font patch)
# BASE_REBUILD_DATE: 2026-09-19 (v1.9.3 — mempalace-version label, mempalace-toolkit 817b3a8 mine deadline, pi-extensions skill floor 25c1265; the marker had been stale since 2026-07-13 through the v1.9.0 Node 24 and v1.9.2 npm-residue rebuilds, updated now because the base is rebuilding anyway)
#
# ── Lineage note ─────────────────────────────────────────────────────
# Adapted from opencode-devbox/Dockerfile.base (commit before v1.16.2).
@@ -83,6 +83,98 @@ ENV DEBIAN_FRONTEND=noninteractive
# above); TERM=xterm-ghostty is compiled from an alias further
# down (ncurses ships `ghostty`, not `xterm-ghostty`). iTerm2
# defaults to xterm-256color (ncurses-base), so needs nothing.
# iproute2 — `ss` (socket statistics) and `ip`. Measured 2026-08-30 on
# v1.8.11: NEITHER was present, so the container could not
# answer "what is listening in here" by any means, and
# cli_utils' `portcheck` was a hard stub — it prints
# "portcheck requires at least one of: ss, lsof, netstat" and
# all three were absent. `ss` satisfies its preferred branch
# (`ss -tlnp`), which is also the branch that reports the
# owning PID, so nothing further is needed: net-tools is
# deliberately NOT added (`netstat` is deprecated and only a
# fallback branch) and neither is lsof (~500 KB for a third
# path to the same answer). ~5.5 MB total: iproute2 itself is
# 4.2 MB and pulls 6 libs under --no-install-recommends
# (libbpf1, libmnl0, libtirpc-common, libtirpc3t64,
# libxtables12, libcap2-bin — libpam-cap is a Recommends and
# is correctly dropped). Verified end-to-end in a live
# container: `ss` lands at /usr/bin/ss, `ip` at /usr/sbin/ip
# (both already on the developer PATH), and `portcheck --all`
# then correctly identifies the socat listener on 8765.
# shellcheck — shell linter. Added 2026-09-09 to close a CAPABILITY gap, not
# a style preference. `scripts/lint-shell.sh` is the release
# GATE (the `lint-gate` job that `resolve-versions` depends
# on), and it correctly refuses to pass when shellcheck is
# missing — "a gate that cannot run must not pass". Measured on
# v1.8.14: shellcheck was absent from this image by all three
# routes (PATH, dpkg, filesystem), so `bash
# scripts/lint-shell.sh` exited 2 in EVERY devbox container and
# no developer could run the release gate locally at all. The
# loop was therefore write-shell → push → wait for CI → discover,
# which is the loop the gate was added to shorten: v1.8.14's
# first attempt burned ~46 min on a tree whose lint had already
# been red for 24 h. This is also what makes a client-side
# pre-push hook possible (see hooks/pre-push); without the
# binary that hook would refuse every push. ~39 MB installed
# (Installed-Size 40112 KB, shellcheck 0.10.0-1) and measured
# to pull ZERO additional packages under
# --no-install-recommends: its deps (libc6, libffi8, libgmp10)
# are already present. NOTE this file feeds the base-decide
# hash (Dockerfile.base + rootfs/), so adding it forces one
# full base rebuild.
# bind9-dnsutils — `dig` and `nslookup`. Added 2026-09-10 to close a
# DIAGNOSTIC gap measured during the gitea.egl.lan/FreeIPA
# work: the container could resolve names but had NO way to
# ask a SPECIFIC nameserver anything. `getent hosts` only
# follows the resolver's default path, so the whole "gateway
# 172.16.88.1 returns NXDOMAIN for the egl.lan zone while
# 10.20.253.1 is authoritative for it" diagnosis had to be
# hand-rolled in python3 — dig, host AND nslookup were all
# absent. `dig @10.20.253.1 freeipa-4.egl.lan` is the
# one-liner that replaces it, and split-horizon DNS is a
# recurring class of bug on this fleet, not a one-off. NOTE
# the package to name is bind9-dnsutils: plain `dnsutils` is
# a transitional package in trixie. ~6.1 MB total (6210 KB
# measured): bind9-dnsutils 721 KB + bind9-host 161 KB +
# bind9-libs 3804 KB plus 7 small libs (libfstrm0,
# libjson-c5, liblmdb0, libmaxminddb0, libprotobuf-c1,
# liburcu8t64, libuv1t64) under --no-install-recommends.
# ldap-utils — `ldapsearch`/`ldapmodify`. Added 2026-09-10. This fleet
# authenticates against FreeIPA (EGL.LAN), and every LDAP
# probe during the Gitea auth work had to be run by SSHing to
# an already-enrolled host because the container had no LDAP
# client at all. 1244 KB and pulls NOTHING extra under
# --no-install-recommends — its deps (libldap, libsasl2) are
# already present. CAVEAT: this gives SIMPLE binds only,
# which is what Gitea itself uses and what most probes need.
# GSSAPI binds (`ldapsearch -Y GSSAPI`) additionally require
# krb5-user + libsasl2-modules-gssapi-mit, deliberately NOT
# added here — that is a Kerberos-client decision with
# /etc/krb5.conf implications, not just a tool.
# xxd — hex dump. 198 KB, no extra deps. Convenience, and honestly
# marginal: `od -c` from coreutils is always present and does
# the same job. Earned its place because verifying that
# git-crypt actually encrypted a staged blob (the \0GITCRYPT\0
# magic) is a recurring check in myconfigs and xxd is the
# muscle-memory command for it.
# NOT added — netcat-openbsd (133 KB): measured redundant on
# 2026-09-10, because socat is already baked above AND bash's
# /dev/tcp does reachability checks with zero packages
# (verified against gitea.egl.lan:3000). Recorded here so the
# omission reads as a decision rather than an oversight.
# python3-yaml — PyYAML. Added 2026-09-10 for precisely the same reason as
# shellcheck above: a gate this repo ALREADY OWNS could not be
# run locally by anyone. scripts/check-workflow-shell.sh — the
# guard that catches the "bash-only syntax under Gitea's default
# sh/dash shell" footgun that broke resolve-versions (ed49b8d)
# and promote-base-latest (b7197e8) — hard-exits with "ERROR:
# python3 yaml module missing" without it. lint.yml installs it
# explicitly in CI (`shellcheck python3-yaml`), which is itself
# the evidence that the image lacked it. Measured 2026-09-10
# while wiring the skill-floor job: the guard could not be run
# before pushing — the same write → push → wait-for-CI loop that
# shellcheck was baked to shorten. 552 KB, and pulls ZERO extra
# packages under --no-install-recommends.
RUN apt-get update && \
apt-get upgrade -y --no-install-recommends && \
apt-get install -y --no-install-recommends \
@@ -102,6 +194,7 @@ RUN apt-get update && \
make \
patch \
diffutils \
shellcheck \
git-crypt \
age \
file \
@@ -122,6 +215,11 @@ RUN apt-get update && \
nano \
kitty-terminfo \
ncurses-term \
iproute2 \
bind9-dnsutils \
ldap-utils \
xxd \
python3-yaml \
&& ln -s /usr/bin/fdfind /usr/local/bin/fd \
&& apt-get clean \
&& rm -rf /var/lib/apt/lists/*
@@ -354,7 +452,10 @@ RUN ARCH=$(case "${TARGETARCH}" in amd64) echo "x86_64" ;; arm64) echo "aarch64"
# uninterruptibly. A stall-kill is no longer a permanent latch either: the
# next tool call respawns the server with capped exponential backoff (the
# budget resets on any successful response). Tunables:
# MEMPALACE_MCP_TIMEOUT_MS (default 60000), MEMPALACE_MCP_INIT_TIMEOUT_MS
# MEMPALACE_MCP_TIMEOUT_MS (default 60000; the feed's `mempalace_mine` carries
# its own longer MEMPALACE_FEED_MINE_TIMEOUT_MS, default 300000, since toolkit
# 817b3a8 — before that the 60 s deadline cut every honest mine off),
# MEMPALACE_MCP_INIT_TIMEOUT_MS
# (default 300000 — generous so a genuine first cold-open isn't killed),
# MEMPALACE_MCP_MAX_RESPAWNS (default 2; 0 disables self-heal),
# MEMPALACE_MCP_RESPAWN_BACKOFF_MS (default 1000); timeouts of 0 disable.
@@ -431,13 +532,40 @@ ARG INSTALL_MEMPALACE=true
# the part that should stay manual.
#
# Deployment sequencing note for whoever ships this bump: synlig (the shared
# central palace host) currently serves mempalace 3.7.1 SERVER-SIDE via
# docker-compose.mempalace.yml, which reuses this same devbox image. Bumping
# this ARG changes only the CLIENT version baked into pi-devbox images: it
# introduces client/server skew until synlig's compose stack is separately
# rebuilt/redeployed with the new pin. Not something to code around here —
# just sequence the redeploy.
ARG MEMPALACE_VERSION=3.8.0
# central palace host) serves mempalace 3.8.0 SERVER-SIDE via
# docker-compose.mempalace.yml, which reuses this same devbox image. (Measured
# 2026-09-06 over ssh: synlig's UV_TOOL_DIR mempalace entry last changed
# 2026-08-25 15:33 — this comment previously said 3.7.1, which was stale.)
# Bumping this ARG changes only the CLIENT version baked into pi-devbox
# images: it introduces client/server skew until synlig's compose stack is
# separately rebuilt/redeployed with the new pin. Not something to code around
# here — just sequence the redeploy.
#
# v1.8.13: 3.8.0 -> 3.9.0. Audited: no Breaking/Removed changelog headings.
# Adopted mainly for #2281 (`mempalace_mine` accepts a single conversation
# file again) — though note that does NOT unblock this image's own feeder,
# which was measured to mine DIRECTORIES, not files, so it was never hitting
# that bug. Four behaviour changes ride along and are skew-relevant while
# synlig stays on 3.8.0: hub-forward escaping, an HTTP lock split, similarity
# score semantics, and parsed-output compatibility. 3.9.0-only features
# (release awareness, `task create`/`task launch` MCP tools) are SERVER-side,
# so they stay dark until synlig is redeployed — a client bump alone cannot
# light them up.
ARG MEMPALACE_VERSION=3.9.0
# Recorded as a label HERE, not in Dockerfile.variant, for three reasons: the
# value lives next to the ARG that defines it (a second copy in the variant
# would be one more pin able to drift, which is the class check-doc-drift.sh
# exists to catch); labels are inherited by every image built FROM this one, so
# both variants carry it with no build-arg to plumb through four call sites;
# and inheritance means the label states the pin of the base the variant
# ACTUALLY built on — which is the question when base-decide cache-hits an
# older base. Like every se.jordbo.pi-devbox.* label this records INTENT; the
# ground truth is /etc/pi-devbox/build-manifest.json's mempalace_version, read
# from the installed binary, and scripts/smoke-test.sh asserts the two agree.
# check-doc-drift.sh check 9 reads this off the last published image so that a
# pin bump must be named in the CHANGELOG — until this label ships, that
# component reports SKIP (label absent on the published release), not OK.
LABEL se.jordbo.pi-devbox.mempalace-version="${MEMPALACE_VERSION}"
ENV UV_TOOL_DIR=/opt/uv-tools
ENV UV_TOOL_BIN_DIR=/usr/local/bin
RUN if [ "${INSTALL_MEMPALACE}" = "true" ]; then \
@@ -522,7 +650,17 @@ ENV COLORTERM=truecolor
ENV PATH="/home/developer/.local/bin:/home/developer/.cargo/bin:${PATH}"
# ── Node.js (required for pi + MCP servers + tldr) ──
ARG NODE_VERSION=22
# 24 (LTS "Krypton"), raised from 22 on 2026-09-10 because the image was BELOW a
# DECLARED requirement, not merely behind the newest release: `agent-browser`
# publishes engines.node ">=24.0.0", so every build on 22 installed it with an npm
# EBADENGINE warning and then ran it outside its supported range — measured on
# v1.8.14, which shipped node 22.23.2 with agent-browser 0.37.1. The other two npm
# consumers are satisfied either way: pi declares ">=22.19.0" and playwright
# ">=20". Verified before bumping, because a missing NodeSource suite would break
# the build for every arch at once: deb.nodesource.com/setup_24.x returns HTTP 200
# and the node_24.x suite advertises `Architectures: amd64 arm64 armhf x86_64`, so
# both the arm64 fleet and the amd64 CI runners resolve.
ARG NODE_VERSION=24
RUN curl -fsSL --retry 5 --retry-delay 5 --retry-all-errors https://deb.nodesource.com/setup_${NODE_VERSION}.x | bash - && \
apt-get install -y --no-install-recommends nodejs && \
rm -rf /var/lib/apt/lists/*
+222 -6
View File
@@ -57,6 +57,32 @@ ARG USER_NAME=developer
# v0.74.0..v0.75.5; discovered + fixed in v0.75.5b, 2026-05-23). The `latest`
# branch below is kept only for a deliberate local `docker build` override.
#
# AUDITED AT 0.84.4 (2026-08-31, was 0.84.3): NO "Breaking Changes" and no
# "Removed" heading in the 0.84.4 section (grepped, 0 matches) — unlike 0.84.3,
# whose heading is described in the paragraph below and stays audited. Adopted
# for three fixes that land on machinery this fleet actually runs:
# - #6879 large tool results crossing the auto-compaction threshold were sent
# to the provider BEFORE compacting; pi now compacts between tool execution
# and the next assistant response in the same run. This is the shape of
# nearly every session here (multi-hundred-KB logstream/palace tool output).
# - #8345 a resumed session corrupted its next appended entry when the JSONL
# lacked a trailing newline. That file is the memory feeder's own input.
# Measured on tor-ms22 before the bump: 49/49 transcripts end in a newline,
# 0 lines fail json.loads — the bug had not bitten this corpus.
# - #8537 extension messages sent with `triggerTurn: false` WHILE THE AGENT IS
# RUNNING were inserted between a tool call and its result, so
# order-validating providers rejected the replayed history. The mempalace
# mailbox is outside that precondition — it delivers at `agent_settled`
# (idle) with `{deliverAs:"steer"}` and deliberately no `triggerTurn` — and
# 0.84.4 leaves the documented steer semantics unchanged, so RFC 003 §7.11
# still holds. Recorded because the fix is what would make a future mid-run
# delivery safe, which is the only reason we would ever change that call.
# One doc consequence, fixed in this same release: pi's own docs/compaction.md
# gained exactly one paragraph — the autoCompact threshold is now ALSO checked
# mid-run, after a tool batch's results are appended. See
# docs/observational-memory.md §3, which had said compaction is only checked
# when pi goes idle.
#
# AUDITED AT 0.84.3 (2026-08-25, was 0.84.2): upstream's notes carry a
# "Breaking Changes" heading — `GoogleThinkingLevel` renamed to
# `GoogleApiThinkingLevel`. INERT FOR THIS IMAGE: all four vendored companions
@@ -69,9 +95,25 @@ ARG USER_NAME=developer
# `.agents/skills/<group>/` directories were not discovered, and root Markdown
# files such as README.md / AGENTS.md inside a skill dir were reported as
# broken skills unless they declared valid skill frontmatter.
# pi-atelier needs no companion bump: v0.8.2 clears the >=0.7.1 floor that
# pi >= 0.84 requires (see PI_ATELIER_REF below).
ARG PI_VERSION=0.84.3
#
# v1.8.13: 0.84.4 -> 0.85.1. SKIP 0.85.0 deliberately — it accidentally
# published internal experimental code and extra subpaths, breaking SDK
# imports (upstream #9132); 0.85.1 exists specifically to undo that, with the
# supported SDK and stdio RPC API unchanged. Audited: no Breaking/Removed
# changelog headings in either release, engine floor unchanged (>=22.19.0,
# container runs 22.23.2), runtime deps 20 -> 19. User-visible changes are the
# streaming indicator moving into the editor border and faster fullscreen
# transcript search; no deprecation language anywhere.
#
# Verified EMPIRICALLY rather than from the changelog, because a pi bump has
# hung the TUI before (pi-atelier < 0.7.1 + pi >= 0.84): 0.85.1 was
# side-installed and driven under a pty against all four companion extensions,
# with atelier v0.10.0 AND v0.10.1 — five combinations, each rendering alive
# with a CPU delta of 0.00-0.01s over a 5s window, where the known hang
# signature is ~5s of sustained CPU. Two-sided check: the atelier sidebar
# painted ACTIVITY+WORKSPACE identically to the 0.84.4 control, so the test
# could distinguish "loaded" from "silently absent".
ARG PI_VERSION=0.85.1
ARG PI_TOOLKIT_REF=main
ARG PI_EXTENSIONS_REF=main
# Repo URLs default to the canonical gitea origin but are overridable so a
@@ -101,15 +143,36 @@ ARG PI_OBSMEM_REF=master
# pin and PI_VERSION together, checking atelier's CHANGELOG for the pi
# version it claims to track.
#
# AUDITED AT v0.10.0 (2026-08-31, was v0.8.2 — two minor releases): no
# BREAKING notice in either release, and both are UI-only (Sidebar calm during
# an active Turn, composer frame + Status Rail, fullscreen-copy-safe Sidebar,
# Windows path normalisation, Workspace Pulse deferred until pi trusts the
# project). The one coupling that matters runs the OPPOSITE way to the floor
# above: v0.9.0 renders the Sidebar as a separate split-layout child and
# therefore "raises the minimum supported Pi version to 0.84.0", which its
# peerDependencies do encode this time (`>=0.84.0`, up from `>=0.80.7`).
# Satisfied with room to spare by PI_VERSION 0.84.4 above — and note that both
# executable floors (scripts/smoke-test.sh, scripts/recreate-sanity-check.sh)
# compare with `sort -V`, so 0.10.0 >= 0.7.1 is evaluated correctly rather than
# as the string comparison that would read 0.10.0 as older than 0.7.1.
# Pairs deliberately with pi 0.84.4's own fullscreen selection-copy controls:
# atelier keeps Sidebar content out of the transcript selection, pi adds
# `fullscreenCopyOnSelect` + Ctrl+X for the selection itself.
#
# No `npm install` step, unlike pi-fork/pi-observational-memory/pi-studio:
# pi-atelier declares ZERO runtime dependencies (only peerDeps, satisfied by
# the baked pi) and has no build step — pi loads its TypeScript directly from
# the /opt checkout. Adding an install here would be a no-op that only costs
# build time.
ARG PI_ATELIER_REPO=https://github.com/michaelmjhhhh/pi-atelier.git
ARG PI_ATELIER_REF=v0.8.2
# v1.8.13: v0.10.0 -> v0.10.1. Refactor-only upstream (formatters, tests,
# panel identity); peerDependencies declare pi >=0.84.0, so it spans both the
# old and new pin. Included because it was already exercised: the pty matrix
# for PI_VERSION above ran atelier v0.10.1 against pi 0.85.1 and painted the
# sidebar identically to v0.10.0.
ARG PI_ATELIER_REF=v0.10.1
# Human-readable tag PI_ATELIER_REF was resolved from; recorded as a label.
ARG PI_ATELIER_VERSION=v0.8.2
ARG PI_ATELIER_VERSION=v0.10.1
RUN set -e && \
# git_fetch_ref: clone-equivalent helper that accepts EITHER a branch name
@@ -133,6 +196,71 @@ RUN set -e && \
done; \
return 1; \
} && \
# prune_foreign_natives: npm 11 (shipped with Node 24) installs EVERY optional
# platform package of a native dependency, not just the one matching the host.
# TWO families are affected in this image, and BOTH have been measured — add a
# family here only after measuring it, never by widening the pattern on a hunch:
#
# @esbuild/<platform> 26 dirs, 284 MB (found first, v1.9.1)
# @mariozechner/clipboard-<triple> 11 dirs, 12 MB per site, 10 MB foreign
#
# esbuild declares those with os/cpu constraints, but npm 11 ignores them and
# ALSO ignores --os/--cpu and an npmrc carrying os=/cpu= (all three measured).
# So prune explicitly, keeping only the host platform, computed from
# `node -p process.arch` so one line stays correct on amd64 and arm64.
# Measured on pi-fork's tree: npm 10.9.8 -> 165 MB, npm 11.19.0 -> 449 MB,
# and the 165 MB figure reproduces what v1.9.0's predecessor actually shipped.
#
# WHY THE CLIPBOARD FAMILY WAS ADDED (2026-09-11): v1.9.1 pruned @esbuild only
# and still shipped +131 MB compressed over v1.8.14. That residual was
# attributed by listing the PUBLISHED arm64 layer tarballs straight from the
# registry (there is no docker CLI inside the container, so `docker history`
# was not available): +110 MB /root/.npm/_cacache (purged below) and +21 MB of
# clipboard platform packages across the two install sites — 131 MB total, so
# the delta is now fully accounted for with no unexplained remainder.
#
# Keeping linux-$arch-{gnu,musl} is deliberate: clipboard's napi-rs loader
# tries ./<name>.node then the platform package, per platform in try/catch, and
# chooses gnu vs musl at runtime from its own isMusl() probe — so both host-arch
# branches must survive. The musl package is a 420-byte stub, i.e. free. The
# bare wrapper `@mariozechner/clipboard` has no hyphen suffix and therefore
# cannot match the regex below. Verified on arm64 against a copy of the real
# tree before this was written: after pruning to those two,
# require('@mariozechner/clipboard') still loads and exports all 18 functions.
# esbuild likewise still compiles TS via transformSync at both install sites.
# This removes dead weight, not function.
#
# MUST run in the SAME layer as the npm installs above: deleting in a later RUN
# leaves the bytes in this layer and shrinks the image by nothing.
# NOTE the single backslash in -printf '%f\n': Docker passes '\\n' through
# verbatim, so v1.9.1's doubled version printed a mangled "li ux-arm64"
# (find emitted a literal backslash, then `tr` translated the n out of the
# name). Confirmed from the published image's own recorded created_by.
prune_foreign_natives() { \
arch="$(node -p process.arch)"; \
find /usr/lib/node_modules /opt -type d -regex '.*/@esbuild/[^/]+' \
! -name "linux-$arch" -prune -exec rm -rf {} + ; \
find /usr/lib/node_modules /opt -type d -regex '.*/@mariozechner/clipboard-[^/]+' \
! -name "clipboard-linux-$arch-gnu" ! -name "clipboard-linux-$arch-musl" \
-prune -exec rm -rf {} + ; \
echo "native platform dirs kept: $(find /usr/lib/node_modules /opt -type d \( -regex '.*/@esbuild/[^/]+' -o -regex '.*/@mariozechner/clipboard-[^/]+' \) -printf '%f\n' 2>/dev/null | sort | uniq -c | tr '\n' ' ')"; \
} && \
# purge_build_caches: the build's own download caches are NOT free — they land
# in whichever layer created them. Measured on the published v1.9.1 arm64
# variant layer: root/.npm/_cacache was 145.2 MB of a 401.9 MB layer (35.2 MB
# in v1.8.14), the single biggest item in the +131 MB residual, because npm 11
# caches every platform tarball it fetched — including the ones just pruned.
# Nothing at runtime reads it: the build runs as root, the container runs as
# `developer` with its own cache under $HOME (and $HOME/.pi is a volume).
# DELIBERATELY NOT purged here: /tmp/node-compile-cache (1.3 MB, written by
# `pi --version` below). The manifest RUN at the end of this file calls
# `pi --version` again, so deleting it here only relocates those bytes into
# that layer instead of removing them from the image — measured, not assumed:
# today the manifest layer is 128 kB precisely because it finds the cache warm.
purge_build_caches() { \
npm cache clean --force >/dev/null 2>&1 || true; \
rm -rf /root/.npm; \
} && \
if [ "${PI_VERSION}" = "latest" ]; then \
NPM_CONFIG_PREFIX=/usr npm install -g @earendil-works/pi-coding-agent ; \
else \
@@ -146,6 +274,8 @@ RUN set -e && \
git_fetch_ref "${PI_ATELIER_REPO}" "${PI_ATELIER_REF}" /opt/pi-atelier && \
(cd /opt/pi-fork && npm install --omit=dev --no-audit --no-fund) && \
(cd /opt/pi-observational-memory && npm install --omit=dev --no-audit --no-fund) && \
prune_foreign_natives && \
purge_build_caches && \
echo "pi-toolkit at $(cd /opt/pi-toolkit && git rev-parse --short HEAD)" && \
echo "pi-extensions at $(cd /opt/pi-extensions && git rev-parse --short HEAD)" && \
echo "pi-fork at $(cd /opt/pi-fork && git rev-parse --short HEAD)" && \
@@ -219,9 +349,53 @@ ARG PI_STUDIO_REF=main
# PI_STUDIO_VERSION is the human-readable tag (e.g. v0.9.36) that PI_STUDIO_REF
# was resolved from; recorded as a label below for at-a-glance identification.
# Only meaningful for the studio variant (default `none` otherwise).
#
# v1.8.13 — READ THIS BEFORE REASONING ABOUT WHICH pi-studio SHIPS. Neither
# default below survives a CI build. `resolve-versions` in
# .gitea/workflows/docker-publish.yml passes BOTH as build-args (studio_ref and
# studio_tag), and it deliberately selects the newest STABLE semver tag: its
# filter is `^v?[0-9]+\.[0-9]+\.[0-9]+$`, which excludes pre-releases. So a
# PUBLISHED v1.8.13 studio image contains pi-studio v0.9.59 (commit 9eed84f,
# = refs/tags/v0.9.59^{}), NOT the v0.9.60-rc.0 that `main` currently points at
# (658536f). The `main` default here only applies to a local `docker build`
# that passes no studio args.
#
# That upstream-tag-over-main choice is intentional and documented at the
# resolve step: pi-studio keeps tagging every version but stopped publishing
# GitHub Releases at v0.5.55 and pushes freely to main, so pinning main risked
# baking half-finished commits that land after a tag.
#
# Corrected here on 2026-09-06 after reading the run-639 resolve-versions
# output: the v1.8.13 audit had recorded "RC adopted deliberately" and set this
# ARG to v0.9.60-rc.0, which was measured at the wrong layer — a Dockerfile
# default cannot answer "what will CI publish?" when CI overrides it. Left at
# `none` rather than pinned to a tag, because a hardcoded pre-release here goes
# stale the moment main moves and would re-tell the same lie to the next reader.
# Consequence worth keeping: the RC's opt-in Studio network binding is NOT in
# any published v1.8.13 image, so it needs no audit for this release.
ARG PI_STUDIO_VERSION=none
RUN if [ "${INSTALL_STUDIO}" = "true" ]; then \
set -e; \
# Same prune + cache purge as the main install RUN — see the comments there.
# They have to be redefined because shell functions do not survive across
# layers, and they have to run in THIS layer because pi-studio's npm install
# happens here: deleting in a later RUN would leave the bytes in this layer
# and shrink nothing. pi-studio pulls its own pi-coding-agent copy, so it is
# a third ~274 MB site on top of the two in the non-studio variant — and its
# npm install refills /root/.npm, which the main RUN emptied in ITS layer.
prune_foreign_natives() { \
arch="$(node -p process.arch)"; \
find /usr/lib/node_modules /opt -type d -regex '.*/@esbuild/[^/]+' \
! -name "linux-$arch" -prune -exec rm -rf {} + ; \
find /usr/lib/node_modules /opt -type d -regex '.*/@mariozechner/clipboard-[^/]+' \
! -name "clipboard-linux-$arch-gnu" ! -name "clipboard-linux-$arch-musl" \
-prune -exec rm -rf {} + ; \
echo "native platform dirs kept: $(find /usr/lib/node_modules /opt -type d \( -regex '.*/@esbuild/[^/]+' -o -regex '.*/@mariozechner/clipboard-[^/]+' \) -printf '%f\n' 2>/dev/null | sort | uniq -c | tr '\n' ' ')"; \
}; \
purge_build_caches() { \
npm cache clean --force >/dev/null 2>&1 || true; \
rm -rf /root/.npm; \
}; \
rm -rf /opt/pi-studio && mkdir -p /opt/pi-studio && \
git -C /opt/pi-studio init -q && \
git -C /opt/pi-studio remote add origin "${PI_STUDIO_REPO}" && \
@@ -233,6 +407,8 @@ RUN if [ "${INSTALL_STUDIO}" = "true" ]; then \
done; \
[ "$ok" = "1" ] && \
(cd /opt/pi-studio && npm install --omit=dev --no-audit --no-fund) && \
prune_foreign_natives && \
purge_build_caches && \
echo "pi-studio at $(cd /opt/pi-studio && git rev-parse --short HEAD)"; \
fi
@@ -305,7 +481,7 @@ ARG MEMPALACE_TOOLKIT_REF=main
# no ~67-minute base rebuild. (scripts/check-base-hash.sh scans only
# Dockerfile.base, so no folding into the base hash is required — nor would
# it be correct, since this ARG changes nothing about the base's contents.)
ARG SKILLSET_SNAPSHOT_REF=6eb20af181f0147cb8c1377f6e36a6a47a68e8e5
ARG SKILLSET_SNAPSHOT_REF=e9e45f7acdde490c3b5d24ce5f508bff8785c2c7
# Dockerfile.base sets description="pi-devbox — base image (variant-independent)"
# and every variant INHERITS it, so both published images used to advertise
@@ -386,6 +562,44 @@ RUN set -e; \
if [ -d "$_snap_dir" ] && [ -n "$(find "$_snap_dir" -type f -print -quit)" ]; then \
SKILL_SNAP="\"$(tree_sha256 "$_snap_dir")\""; \
fi; \
# ── WHICH pi-extensions skill copy actually shipped ──
# Closes the silent-fallback hole. The refresh step above is guarded by
# `[ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates
# the co-located skill (or a fork pointing at a mirror without it) keeps the
# vendored floor and still succeeds — GREEN, with nothing anywhere recording
# that a snapshot shipped instead of the package copy. Measured 2026-09-10:
# the floor had been stale since 2026-07-30, so that fallback would have
# shipped a six-week-old skill silently. The floor is fresh now and gated by
# the skill-floor CI job, but "the fallback is currently harmless" is not the
# same as "you can tell which copy you got", and only the second survives.
#
# MEASURED, never claimed, per the ground-truth rule above: the branch
# condition is re-derived from the same test the refresh step used, and the
# served bytes are then compared against the clone. A build-arg could not
# express this at all, since the outcome depends on the clone's contents.
# package served bytes == the clone's skill/ (the normal path)
# vendored-floor the clone has no skill/ at this ref (fallback shipped)
# divergent both exist but differ — e.g. the clone ships SKILL.md but
# not evaluate-extension-usage.py, so the served directory is
# a MIX of package and floor. Worth its own value: it is the
# one state neither of the other two names honestly.
# No OCI label mirrors this, deliberately: LABEL cannot take a value computed
# in a RUN, and a label fed from an ARG would be exactly the claim-not-
# measurement this block exists to avoid.
_px_dir=/usr/local/share/pi-devbox/skills/pi-extensions; \
PIEXT_SRC='null'; PIEXT_HASH='null'; \
if [ -d "$_px_dir" ] && [ -n "$(find "$_px_dir" -type f -print -quit)" ]; then \
PIEXT_HASH="\"$(tree_sha256 "$_px_dir")\""; \
if [ -f /opt/pi-extensions/skill/SKILL.md ]; then \
if [ "$(tree_sha256 "$_px_dir")" = "$(tree_sha256 /opt/pi-extensions/skill)" ]; then \
PIEXT_SRC='"package"'; \
else \
PIEXT_SRC='"divergent"'; \
fi; \
else \
PIEXT_SRC='"vendored-floor"'; \
fi; \
fi; \
{ \
echo '{'; \
echo " \"release_tag\": \"${RELEASE_TAG}\","; \
@@ -406,6 +620,8 @@ RUN set -e; \
# vendored skill directory, not one file — see tree_sha256() above.
echo " \"skillset_snapshot_ref\": \"${SKILLSET_SNAPSHOT_REF}\","; \
echo " \"skillset_snapshot_tree_sha256\": ${SKILL_SNAP},"; \
echo " \"pi_extensions_skill_source\": ${PIEXT_SRC},"; \
echo " \"pi_extensions_skill_tree_sha256\": ${PIEXT_HASH},"; \
echo " \"components\": {"; \
echo " \"pi-toolkit\": \"$(rev /opt/pi-toolkit)\","; \
echo " \"pi-extensions\": \"$(rev /opt/pi-extensions)\","; \
+116 -24
View File
@@ -20,7 +20,9 @@ on the host.
- `pi-extensions` — TypeScript extensions for pi (preview, MCP bridges,
mempalace integration, etc.)
- `pi-fork` — the `fork` tool for spawning sub-agents
- `pi-observational-memory` — the `recall` tool for session compaction
- `pi-observational-memory` — durable session memory: the ledger that makes
compaction cheap, plus the `recall` tool. See
[`docs/observational-memory.md`](docs/observational-memory.md)
- `pi-atelier` — TUI sidebar: ordered panels, split-pane, themes. Pinned to an
audited tag; see [Version pins](#version-pins-pi-pi-atelier-mempalace)
@@ -173,12 +175,10 @@ Currently published:
| `joakimp/pi-devbox:latest-studio` | `latest` + [pi-studio](https://github.com/omaclaren/pi-studio) (browser prompt editor, KaTeX/Mermaid preview, tmux-backed literate REPLs) | ~3.25 GB |
| `joakimp/pi-devbox:vX.Y.Z-studio` | pinned-version studio equivalent | ~3.25 GB |
Planned for an upcoming minor release:
- *(shipped in Unreleased/base)* **PDF export from Studio/pandoc** now works:
the base image ships **`typst`** as the PDF engine (`pandoc --pdf-engine=typst`),
a single ~30 MB static binary — no separate `-tex` variant needed.
`texlive-xetex` stays the higher-fidelity fallback (install on demand).
Both variants ship **`typst`** as the pandoc PDF engine
(`pandoc --pdf-engine=typst`), a single ~30 MB static binary, so PDF export from
Studio/pandoc works out of the box — no separate `-tex` variant needed.
`texlive-xetex` stays the higher-fidelity fallback (install on demand).
## Using pi-studio (`-studio` variant)
@@ -536,6 +536,35 @@ to refresh.
Anything not on a volume is on the writable layer and is lost on
container recreate.
### Rebuilding ephemeral shell state at start
Two entrypoint steps put back the kind of state that the writable layer eats, so a
recreate does not cost you a manual re-install:
- **`cli_utils` commands.** If a `cli_utils` checkout is mounted, every
executable in its `bin/` is symlinked into `~/.local/bin` on start, so
`git-status-all` and friends are on `PATH` without a path prefix. Detection:
`CLI_UTILS_CONTAINER_PATH` → `/workspace/cli_utils` → `$HOME/cli_utils` →
`/workspace/*/cli_utils`. Set `CLI_UTILS_LINK=0` to disable. Existing real files
in `~/.local/bin` and symlinks pointing elsewhere are left alone, so a
deliberate override still wins; links whose target disappeared are pruned.
Do **not** run a host installer's `install.sh` inside the container to achieve
this — it writes to the ephemeral home and dies on the next recreate.
- **A per-device boot hook.** If `~/.config/devbox-shell/init.sh` exists it is run
once at start (`bash`, never sourced, exit status ignored), with output in
`~/.pi/agent/devbox-init.log`. `~/.config/devbox-shell/` is the host-owned
bind-mount whose `bash_aliases` is already sourced into every interactive shell,
so a hook there persists across recreates with no image change. Use it for
fixups that must exist *before any shell* — symlinks, directories, one-off
migrations.
The distinction that decides which mechanism you want: `~/.local/bin` is on `ENV
PATH`, so symlinks there work in **non-interactive** shells too (`docker exec <c>
<cmd>`, agent tool shells, scripts). A `PATH` edit in `bash_aliases` reaches only
*interactive* shells, because `~/.bashrc` returns early when non-interactive —
which is also why shell **functions** (fzf helpers and the like) can only come
from the sourced file, never from a symlink.
## MemPalace integration
MemPalace is installed in the base image and pre-warmed with the
@@ -579,7 +608,50 @@ convention that a directed event with `status="open"` is a request owed a reply
while a `*` broadcast owes nothing. The mechanism side (what the bridge stamps,
and why live SSE push depends on the palace deployment's reverse proxy rather
than on this image) is documented in the toolkit's `extensions/pi/README.md`.
Nothing in this image polls the log on the agent's behalf.
**Since v1.8.9 the bridge reads the log for you.** Earlier images were write-only
— they stamped provenance on the way out and never read back, so a directed ask
reached an agent only if that agent happened to run `mempalace_event_list`
itself. The mailbox is gated on the same two variables as the stamper, is on by
default, and derives what is *owed* rather than trusting `status` (an acked event
keeps matching a `status="open"` query forever, because the log is append-only):
| Variable | Default | Effect |
|---|---|---|
| `MEMPALACE_MAILBOX` | unset (on) | `0` disables mailbox reads entirely |
| `MEMPALACE_MAILBOX_POLL_MS` | `300000` | minimum gap between mid-session polls |
| `MEMPALACE_MAILBOX_RESURFACE_MS` | `3600000` | re-announce a still-owed ask after this long |
Delivery **queues, it never interrupts**: the poll runs when pi goes idle and the
message is steered into the *next* turn, so nothing wakes the model on inbound
fleet traffic. The practical consequence, measured on two devices: the message
appears in your session window and the agent acts on it when the next turn
starts — you are the trigger. (That describes the bridge **as baked in v1.8.9**,
`mempalace-toolkit` `5b8d78f`; the mailbox's own mechanism and landmines live in
the toolkit's `docs/rfc-003-coordination-log.md` §7.11–§7.12, which moves ahead of
whatever this image has baked.)
## Observational memory (in-session memory)
The image also bakes [pi-observational-memory](https://github.com/elpapi42/pi-observational-memory),
which is memory of a *different kind* from the palace and is easy to confuse with
it. It keeps a small branch-local ledger of observations and reflections while a
session runs, so when pi compacts the conversation the summary is a
**deterministic fold of that ledger rather than a model call**, and every item
keeps a 12-character id that `recall(<id>)` resolves back to the exact source.
In one line: **observational memory keeps a session coherent; the palace keeps
the fleet coherent.**
It is on by default, needs no habit from you, and sends its background work to a
cheaper model than your session (Haiku while the session runs Opus, in the seeded
`~/.pi/agent/settings.json`). Inspect it from inside pi with `/om:status` and
`/om:view`; turn all proactive work off for one run with
`PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi`.
What it is for, how the lifecycle works, what it costs, every setting and its
default, and how it differs from MemPalace:
[`docs/observational-memory.md`](docs/observational-memory.md).
## Agent skills
@@ -829,8 +901,10 @@ docker inspect --format '{{json .Config.Labels}}' joakimp/pi-devbox:latest | jq
```
`org.opencontainers.image.{version,revision,created}` plus
`se.jordbo.pi-devbox.*-ref` record the intended pi version and companion
refs. The on-disk `/etc/pi-devbox/build-manifest.json` records **ground
`se.jordbo.pi-devbox.*-ref` and `se.jordbo.pi-devbox.*-version` record the
intended pi and mempalace versions and companion refs (`mempalace-version` is
set in `Dockerfile.base` and inherited, so it names the pin of the base the
image actually built on). The on-disk `/etc/pi-devbox/build-manifest.json` records **ground
truth** — the actual checked-out commit of each `/opt` clone, the live
`pi --version`, and (from v1.8.6) the live `mempalace --version` of the
installed palace core — so a tag is reconstructable after CI logs rotate:
@@ -845,16 +919,23 @@ through `jq` yourself:
```console
$ pi-devbox-version
pi-devbox v1.5.0
built: 2026-07-13T17:53:16Z (source d68674d11e06)
pi: 0.80.6
pi-devbox v1.8.14
built: 2026-09-08T21:54:07Z (source 361babd4fd61)
pi: 0.85.1
palace: 3.9.0
components:
pi-toolkit: 9a8f6faeaa08
pi-extensions: 61c98e004e3d
pi-fork: 4a09af4ef527
pi-observational-memory: 27a5195eaf90
mempalace-toolkit: 96699f2a1781
pi-studio: 2ef38ef31cea
pi-toolkit: adfb553f5c8a
pi-extensions: 2610545c83bb
pi-fork: e69725c39603
pi-observational-memory: ce9fc982b3a2
pi-atelier: 734258bbcb62
mempalace-toolkit: e45f6b430181
pi-studio: e04fc7aa3275
skills:
credential-incident-response baked
mempalace live /workspace/skillset @ 4d7c0ea (identical to baked snapshot)
pi-devbox-environment baked
pi-extensions baked
```
It also flags **live drift** — if `pi --version` no longer matches what was
@@ -1017,10 +1098,21 @@ After `docker compose up -d --force-recreate`, run the **runtime** peer of
persisted volumes survived, and pi runtime wiring is intact:
```bash
./scripts/recreate-sanity-check.sh # auto-detects variant
./scripts/recreate-sanity-check.sh --expected-version 0.79.4 # assert pi version
./scripts/recreate-sanity-check.sh # auto-detects variant
./scripts/recreate-sanity-check.sh --expected-image-version 1.8.9 # assert the pi-devbox release tag
./scripts/recreate-sanity-check.sh --expected-version 0.85.1 # assert the pi coding agent version
```
Those are **two different versions**, and the flags are not interchangeable:
`--expected-image-version` takes the pi-devbox release tag (`v` optional),
`--expected-version` takes `pi --version`. Hand one the other's value and it
says so by name instead of reporting a mismatch against the wrong component.
With neither flag, both values are read from the image's own build manifest
(`/etc/pi-devbox/build-manifest.json`): the live pi version is asserted against
the one recorded at build time — which catches a stale `pi` in the
`~/.pi/npm-global` volume shadowing the baked one — and the release tag is
reported informationally.
If `cli_utils` is on your PATH, the `pi-devbox-sanity` wrapper runs the same
check by short name and locates the repo automatically (override with
`PI_DEVBOX_REPO=/path/to/pi-devbox`). Like `smoke-test.sh`, this script is
@@ -1047,9 +1139,9 @@ resolved to `latest` at build time:
| Component | Pin | Where |
|---|---|---|
| pi | `0.84.2` | `ARG PI_VERSION` — `Dockerfile.variant` |
| pi-atelier | `v0.8.2` | `ARG PI_ATELIER_REF` — `Dockerfile.variant` |
| mempalace | `3.7.1` | `ARG MEMPALACE_VERSION` — `Dockerfile.base` |
| pi | `0.85.1` | `ARG PI_VERSION` — `Dockerfile.variant` |
| pi-atelier | `v0.10.1` | `ARG PI_ATELIER_REF` — `Dockerfile.variant` |
| mempalace | `3.9.0` | `ARG MEMPALACE_VERSION` — `Dockerfile.base` |
The objective is **not** to freeze versions. Bumping is routine — usually one
line plus a changelog note. The objective is that adopting a new upstream
+379
View File
@@ -0,0 +1,379 @@
# Observational memory — why this image has it, and what it does for you
**Audience:** anyone using this container for long pi sessions who has wondered
what `recall`, `/om:status` and "compacted memory" are, or whether they should
leave any of it switched on.
**Companion documents:** the extension ships its own reference docs at
`/opt/pi-observational-memory/docs/` —
[`concepts.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/concepts.md)
(the model),
[`how-it-works.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/how-it-works.md)
(hooks and internals) and
[`configuration.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/configuration.md)
(every setting). Pi's own compaction mechanics are in
`/usr/lib/node_modules/@earendil-works/pi-coding-agent/docs/compaction.md`.
Those are normative; this document is the **deployment** view — what is pinned
here, how it is wired, what it costs, and how it differs from MemPalace. For the
palace, see
[`mempalace-toolkit/docs/fleet-memory.md`](https://gitea.jordbo.se/joakimp/mempalace-toolkit/src/branch/main/docs/fleet-memory.md).
> Verified on pi-devbox **v1.8.9** (`release_tag v1.8.9`, source `aac4a1c`),
> which bakes pi-observational-memory **v3.0.4** at commit `ce9fc98` — the value
> in `/etc/pi-devbox/build-manifest.json` → `components.pi-observational-memory`.
> Every number below was read from that tree, from pi's own docs, or from the
> live container. The pi-side mechanics were first read at pi **0.84.3** and
> re-checked at **0.84.4** (v1.8.12), which moved one of them — see §3.
---
## 1. The problem it solves
A long pi session outgrows the model's context window. Pi's answer is
**compaction**: fold the older part of the conversation into a summary and keep
recent messages verbatim. That is unavoidable, and it is where sessions go
wrong — the summary is produced *at the moment of pressure*, by a model, about a
transcript that is about to leave the context.
Observational memory changes *when* the remembering happens. Instead of
summarising in a panic at the end, it keeps a small **ledger** up to date while
the session runs, and compaction then just folds that ledger.
```mermaid
flowchart LR
A0["plain compaction"] --> A1["context fills"]
A1 --> A2["a model summarises<br/>under pressure"]
A2 --> A3["prose summary,<br/>no way back"]
B0["with observational<br/>memory"] --> B1["context fills"]
B1 --> B2["ledger written<br/>as you work"]
B2 --> B3["compaction folds<br/>the ledger"]
B3 --> B4["ids you can<br/>recall"]
```
Top row is pi on its own: one model call at the worst possible moment, detail
chosen in a hurry, and the original wording gone from view. Bottom row is this
image's default: the thinking happened earlier on a cheap model, the fold is
deterministic, and every line in the result carries an id that resolves back to
the exact source.
## 2. The mental model: three layers and a ledger
| Layer | What it is | Example |
|---|---|---|
| **Observation** | a timestamped, source-backed event from the conversation | "user rejected option B because it needs a base rebuild" |
| **Reflection** | a durable conclusion *backed by* observations | "the user optimises for avoiding 67-minute rebuilds" |
| **Drop** | a tombstone retiring an observation from active memory | the superseded detail of a bug that is now fixed |
These are appended to the session as silent ledger entries
(`om.observations.recorded`, `om.reflections.recorded`,
`om.observations.dropped`) and **folded** — replayed in order — to produce the
memory state. The ledger is the source of truth; what you see in a compacted
session is a rendering of it.
Two properties follow, and both matter later:
- **The ledger itself costs no context.** Those entries are pi `custom` entries,
which *"do not participate in LLM context"* (pi `docs/session-format.md`). They
sit in the session file and reach the model only via the fold at compaction.
- **Memory is branch-local.** A pi session is a tree (resume, fork), and the fold
follows the current branch only, so a forked branch does not inherit another
branch's view.
## 3. The lifecycle
Three background workers and one compaction hook, driven by *raw token
progress* rather than wall-clock time. Defaults in brackets.
```mermaid
flowchart TD
T(["turn_end"]) --> O{"10k raw tokens<br/>since observing?"}
O -- yes --> OBS["<b>observer</b> runs"]
O -- "no" --> R{"20k tokens<br/>since reflecting?"}
R -- yes --> REF["<b>reflector</b> runs"]
REF -- "if pool over 10k" --> DR["<b>dropper</b> prunes"]
S(["agent_settled"]) --> C{"81k tokens<br/>since compacting?"}
C -- yes --> CP["ctx.compact()"]
CP --> H(["session_before_compact"])
A(["pi autoCompact<br/>idle, or mid-run<br/>after a tool batch"]) --> H
H --> F["fold the ledger<br/>no model call"]
F --> VIS["compacted memory"]
```
- **observer** — `observeAfterTokens` [10000]: writes observations for the
conversation it has not covered yet.
- **reflector** — `reflectAfterTokens` [20000]: promotes patterns across
observations into durable reflections.
- **dropper** — no clock of its own. It is post-reflection maintenance, gated on
a *successful same-turn* reflection **and** an active pool above
`observationsPoolTargetTokens` [10000]. Not a third worker on a third
threshold.
- **compaction** — `compactAfterTokens` [81000], checked at `agent_settled`, so
*this* trigger never interrupts a turn. Pi will also compact on its own when
the context is nearly full (`contextTokens > contextWindow - reserveTokens`,
`reserveTokens` [16384]), and **from pi 0.84.4 that check also runs mid-run** —
after a tool batch's results are appended, before the next assistant response,
skipped only when the batch ends the run and no queued message needs another
response. So `session_before_compact` has **two** entry points and the second
one can fire *inside* a turn. Harmless for the fold itself, which makes no
model call, but worth stating plainly: "never interrupts a turn" was only ever
true of the observational-memory trigger, and reads as a promise about pi's.
## 4. What compaction actually does to your context
This is the question the rest of the document used to leave hanging: if the old
conversation is folded away, is the session back to knowing nothing?
**No.** Compaction replaces *part* of the context, not all of it, and it deletes
nothing at all from disk.
```mermaid
flowchart LR
SYS["system prompt<br/>+ AGENTS.md"] --> CTX["what the model sees<br/>on the next turn"]
SUM["folded memory:<br/>reflections + observations"] --> CTX
TAIL["recent turns,<br/>verbatim"] --> CTX
DISK[("session .jsonl: all of it")] -. "recall(id)" .-> CTX
```
Where each piece comes from:
- **System prompt and `AGENTS.md` — never compacted, because they were never
conversation.** Pi rebuilds them from disk on every request
(`loadContextFileFromDir`), so they cannot be lost by compaction.
- **The verbatim tail — sized by a token budget, not a message count.** Pi walks
backwards from the newest entry accumulating token estimates until
`keepRecentTokens` [20000] is reached; that entry becomes `firstKeptEntryId`,
and *everything from there on is kept unchanged*. Cut points land on turn
boundaries, never mid-tool-call. So the most recent ~20k tokens of real work —
your last instructions, the diffs, the test output — survive word for word.
- **The folded memory — replaces only what came before that cut.** Rendered from
the ledger's records: reflections and observations, each with its 12-hex id.
- **The session file — untouched.** Compaction *appends* a `compaction` entry
(`{"type":"compaction", summary, firstKeptEntryId, tokensBefore, …}`) and
rebuilds context from it on later turns. Nothing is rewritten in place; the
only documented way to remove session content is deleting the whole `.jsonl`.
That last point is what makes the answer to "is the detail gone?" *no* rather
than *mostly*: `recall` does not read the context window at all. It calls
`sessionManager.getBranch()` — the full branch from the root — and resolves an
observation id back to the original entries. Detail that left the model's view
an hour ago is still one `recall` away.
**Repeated compaction does not summarise the summary.** The rendered text is
always built from live observation/reflection *records*, never from the previous
compaction's prose, so there is no generation-loss spiral. (Mechanically the
projection is incremental — it re-derives back to the last full-fold boundary and
carries the rest forward, escalating to a genuine re-fold from the branch root
when the observation pool reaches `observationsPoolMaxTokens` [20000].)
So the honest summary of the state after compaction: **the model keeps its
instructions, keeps recent work verbatim, trades older turns for a dense
id-carrying digest of them, and can pull any of it back on demand.** Not a fresh
start — a smaller, cheaper, still-navigable one.
### One caveat about "no model call"
If the ledger is empty — compaction fires before the observer has ever run — the
hook returns nothing and *declines ownership*, and pi's own model-based
summariser runs instead:
```ts
const summary = renderSummary(projection.reflections, projection.observations);
if (summary.length === 0) {
// Decline ownership so Pi's native summarizer preserves the pre-cut context.
return;
}
```
In steady state (any session old enough to have produced one observation) om's
hook wins and compaction is model-free. "Never calls a model" is true in practice
and false in principle; the fallback is deliberate, so an empty ledger degrades
to normal pi rather than to no summary at all.
## 5. What you actually get
- **Compaction stops being a stall.** In steady state the latency path is
deterministic work over ledger entries, not a summarisation call.
- **Nothing important vanishes silently.** Compaction is lossy by design, but
every item keeps a 12-character id, and `recall(<id>)` returns the exact
evidence — original wording, reasoning, file path, error text.
- **The bookkeeping runs on a cheaper model than your session.** In this image
that is deliberate and visible (§7): background workers on Haiku, session on
Opus.
- **It is automatic.** No habit to maintain, unlike the palace protocol — which is
exactly why the two complement each other (§11).
- **Forks stay clean.** Branch-local memory means a `fork` sub-agent's noise does
not leak into the parent's folded memory.
## 6. `recall` is not a search tool
`recall` takes **one specific 12-hex id** that already appears in compacted
memory or in `/om:view`. It cannot be given a topic. It can return an observation
(marked `active` or `dropped`), or a reflection together with the observations
supporting it.
```mermaid
sequenceDiagram
participant M as compacted memory
participant A as agent
participant L as ledger
M->>A: "[high] user rejected option B (a1b2c3d4e5f6)"
A->>L: recall("a1b2c3d4e5f6")
L-->>A: exact observation + source ids
Note over A: acts on the original wording
```
The rule of thumb the agent skill uses: recall **before a load-bearing action**
that rests on a compressed memory — shipping a change, asserting a fact,
answering "why do you believe that". One recall is cheap; redoing finished work
is not.
## 7. How it is wired in this image
```mermaid
flowchart TB
IMG["baked in the image:<br/>v3.0.4 @ ce9fc98"] --> REG["settings.json<br/>packages[]"]
REG --> SESS["your pi session"]
SESS -- "your turns" --> SM["session model:<br/>Opus"]
SESS -- "observer, reflector,<br/>dropper" --> WM["memory model:<br/>Haiku"]
SESS -- "ledger entries" --> JL["session .jsonl"]
JL --> VOL[("devbox-pi-config<br/>volume")]
```
Four consequences of that wiring:
1. **There is no separate database.** Memory *is* entries inside the ordinary pi
session file (`~/.pi/agent/sessions/<project>/<timestamp>_<uuid>.jsonl`).
Nothing extra to back up, nothing to migrate.
2. **It survives container recreate**, because `~/.pi` is the `devbox-pi-config`
named volume (`docker-compose.yml`) — the same one holding your pi config and
session history.
3. **`packages[]` is the only source of truth for which copy is loaded.** A clone
at `/workspace/pi-observational-memory` may exist (and today matches `/opt`
byte-for-byte at `ce9fc98`) — its presence proves nothing. To run a patched
build you point `packages[]` at it explicitly and start a new session.
4. **The worker model is a deliberate choice, and it is yours to change.** The
seeded config sends background work to Haiku while your session runs Opus:
```json
"observational-memory": {
"model": { "provider": "amazon-bedrock", "id": "eu.anthropic.claude-haiku-4-5-20251001-v1:0" },
"debugLog": false
}
```
## 8. What it costs
| Resource | Cost |
|---|---|
| Model calls | up to **three** background calls per consolidation pass (observer, reflector, dropper), each capped at `agentMaxTurns` [16], on the configured memory model — not your session model |
| Latency in your turns | none by construction: workers run from `turn_end`, compaction runs when pi is idle, and the fold itself does no model work |
| Disk | negligible — JSON lines inside a session file that would exist anyway (measured here: `~/.pi/agent/sessions` = 30 MB total, tens of `om.*` entries per session) |
| Context window | **zero until compaction.** `custom` entries do not enter LLM context; only the folded summary does |
| Attention | none once configured; there is no protocol for you or the agent to remember |
If that is still more than you want on a given run, §9's `passive` switch turns
off all proactive work while keeping `recall` and `/om:*` usable.
## 9. Configuration
Global: `~/.pi/agent/settings.json` (persisted in the volume). Per project:
`<project>/.pi/settings.json`, which overrides global. Precedence is
project → global → environment, and the environment can only override `passive`.
| Key | Default | What it changes |
|---|---|---|
| `observeAfterTokens` | `10000` | observer cadence — lower means smaller chunks and more calls |
| `reflectAfterTokens` | `20000` | reflector cadence (and thereby dropper opportunities) |
| `observerChunkMaxTokens` | 20% of the memory model's context window, else `60000` | cap on one observer run's input |
| `compactAfterTokens` | `81000` | when proactive auto-compaction fires |
| `observationsPoolMaxTokens` | `20000` | pool size at which compaction does a full re-fold from the branch root |
| `observationsPoolTargetTokens` | half of max (`10000`) | what the dropper aims back down to |
| `agentMaxTurns` | `16` | shared turn cap for the three workers |
| `model` | unset → session model | send background work to a cheaper/faster model |
| `showWorkerNotifications` | `true` | routine "observer ran" notices |
| `passive` | `false` | **kill switch** for all proactive background work; `recall` and `/om:*` still work |
| `debugLog` | `false` | per-session NDJSON trace at `~/.pi/agent/observational-memory/debug/<session-id>.ndjson` |
Pi's own compaction knobs live under a separate `compaction` key —
`keepRecentTokens` [20000] sets the verbatim tail from §4, `reserveTokens`
[16384] the headroom that triggers pi's own compaction.
One-off passive run, no config edit:
```bash
PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi
```
Invalid values are ignored rather than fatal, so a typo degrades to the default
instead of breaking your session — which also means a typo is silent. Check with
`/om:status`.
## 10. Confirming it is actually working
Do not infer health from the absence of a warning; look:
```bash
# 1. inside pi — the authoritative view
/om:status # visible-vs-full drift, thresholds, worker state
/om:view # what the agent currently sees
/om:view full # full ledger truth at the branch tip
# 2. from a shell — are ledger entries being written, and has it compacted?
grep -o '"customType":"om\.[a-z.]*"' \
"$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)" | sort | uniq -c
grep -c '"type":"compaction"' "$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)"
# 3. which copy is loaded, and at what commit
python3 -c "import json;print(json.load(open('$HOME/.pi/agent/settings.json'))['packages'])"
git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory rev-parse HEAD
```
Ledger entries are `"type":"custom"` with `"customType":"om.…"`. Do not grep for
`custom_message` — that is a *different* pi API for entries that **do** enter LLM
context, used here by the MemPalace mailbox (`customType: "mempalace-mailbox"`),
not by om.
## 11. It is not the same thing as MemPalace
Both are called "memory" and they solve different problems. Nothing is wrong with
running both — this image does, and they cover each other's failure modes.
```mermaid
flowchart LR
O0["observational<br/>memory"] --> O1["horizon:<br/>this session"]
O1 --> O2["scope: one branch,<br/>one machine"]
O2 --> O3["automatic"]
O3 --> O4["retrieval:<br/>recall(id)"]
P0["MemPalace"] --> P1["horizon: months,<br/>machines"]
P1 --> P2["scope:<br/>the fleet"]
P2 --> P3["protocol-driven"]
P3 --> P4["retrieval:<br/>search, KG, mailbox"]
```
| Question | Answer |
|---|---|
| "What did we decide 200 turns ago in *this* session?" | observational memory (and `recall` for the exact wording) |
| "What did we decide last month, or on another machine?" | MemPalace (`mempalace_search`, diaries) |
| "What is true *right now* about version X?" | MemPalace knowledge graph |
| "Does another machine need something from me?" | MemPalace coordination log — see [Cross-machine agent coordination](../README.md#cross-machine-agent-coordination) |
| "Why is compaction not losing my session?" | observational memory |
The crisp version: **observational memory keeps a session coherent; the palace
keeps the fleet coherent.** A container recreate wipes neither — but only because
`~/.pi` and the palace both live outside the container filesystem.
## 12. Gotchas
- **Branch-local means branch-local.** Resuming or forking changes which ledger
is folded. Memory that "disappeared" is usually on another branch.
- **`recall` needs an id, not a topic.** If you only have a topic, that is a
palace search, not a recall.
- **A `/workspace` clone is not evidence of what is loaded** — see §7.3.
- **`showWorkerNotifications: true` is not proof of work**; it reports runs, and
an observer that deliberately emits nothing writes no ledger entry and simply
retries after another `observeAfterTokens`.
- **A turn bigger than `keepRecentTokens` splits.** The cut then lands mid-turn at
an assistant message and pi merges two summaries — rare, but it is why a very
large single turn can lose more verbatim detail than you would expect.
- **`git log` in the baked tree needs `safe.directory`** (`/opt` is root-owned):
`git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory log`.
+137
View File
@@ -188,6 +188,111 @@ if [ "${MEMPALACE_FEED:-1}" != "0" ] && [ -n "$MEMPALACE_FEEDER" ]; then
fi
fi
# ── cli_utils: link workspace bin/ commands onto PATH ────────────────
# Standalone commands from a mounted cli_utils checkout (git-status-all,
# git-pull-all, devbox-sanity, pi-session-repair, ...) live in <repo>/bin. On a
# host they reach PATH via cli_utils' own install.sh, whose install_bin step
# symlinks them into ~/.local/bin — but that home is on the container's WRITABLE
# LAYER, so every recreate loses them and the human is back to typing
# /workspace/cli_utils/bin/git-status-all. This is the container equivalent of
# that install step, re-run at every start.
#
# WHY SYMLINKS RATHER THAN A PATH EDIT IN AN rc FILE: ~/.local/bin is already
# ahead of /usr/local/bin in ENV PATH (Dockerfile.base), so links here resolve in
# NON-interactive shells too — `docker exec <c> git-status-all`, agent tool
# shells, scripts. An rc-file PATH edit cannot reach those, because ~/.bashrc
# returns early when the shell is not interactive. Measured 2026-08-27 on
# tor-ms22: `command -v git-status-all` failed in a non-interactive shell while
# working in an interactive one, from exactly that asymmetry.
#
# Detection order (first hit wins):
# 1. CLI_UTILS_CONTAINER_PATH explicit, for non-standard layouts
# 2. /workspace/cli_utils repo directly in the workspace root
# 3. $HOME/cli_utils dedicated mount
# 4. /workspace/*/cli_utils workspace root holds several repo groups
# CLI_UTILS_LINK=0 disables. Absent repo = silent no-op, which is the common
# case for anyone who does not use cli_utils.
if [ "${CLI_UTILS_LINK:-1}" != "0" ]; then
CLI_UTILS_BIN=""
if [ -n "${CLI_UTILS_CONTAINER_PATH:-}" ] && [ -d "${CLI_UTILS_CONTAINER_PATH}/bin" ]; then
CLI_UTILS_BIN="${CLI_UTILS_CONTAINER_PATH}/bin"
elif [ -d /workspace/cli_utils/bin ]; then
CLI_UTILS_BIN=/workspace/cli_utils/bin
elif [ -d "$HOME/cli_utils/bin" ]; then
CLI_UTILS_BIN="$HOME/cli_utils/bin"
else
# `if` bodies, not `&&` chains: under `set -e` a loop whose LAST command is a
# false test exits non-zero and would abort the entrypoint. With no match the
# glob stays literal, so that is the normal case on any machine without this
# repo — i.e. the bug would have been "container will not start", not "links
# missing".
for _cu in /workspace/*/cli_utils/bin; do
if [ -d "$_cu" ]; then
CLI_UTILS_BIN="$_cu"
break
fi
done
unset _cu
fi
if [ -n "$CLI_UTILS_BIN" ]; then
mkdir -p "$HOME/.local/bin" 2>/dev/null || true
# Never clobber a real file, and never steal a link that points elsewhere: a
# deliberate user override in ~/.local/bin must win, and silently shadowing
# an image-provided command is worse than the missing command.
for _f in "$CLI_UTILS_BIN"/*; do
if [ ! -f "$_f" ] || [ ! -x "$_f" ]; then
continue
fi
_link="$HOME/.local/bin/$(basename "$_f")"
if [ -e "$_link" ] && [ ! -L "$_link" ]; then
continue
fi
if [ -L "$_link" ]; then
case "$(readlink "$_link")" in
"$CLI_UTILS_BIN"/*) ;;
*) continue ;;
esac
fi
ln -sf "$_f" "$_link" 2>/dev/null || true
done
# Prune links we own whose target vanished (command renamed, repo moved),
# mirroring the skillset deploy's --prune-stale. A dangling link on PATH
# reports "No such file or directory" for a command that simply no longer
# exists, which reads as a broken container rather than a removed script.
for _link in "$HOME/.local/bin"/*; do
[ -L "$_link" ] || continue
case "$(readlink "$_link")" in
*/cli_utils/bin/*) [ -e "$_link" ] || rm -f "$_link" ;;
esac
done
unset _f _link
fi
unset CLI_UTILS_BIN
fi
# ── Per-device boot hook ─────────────────────────────────────────────
# Runs ~/.config/devbox-shell/init.sh if the host provides one. That directory is
# the host-owned, bind-mounted shell-sharing dir (see "Volumes and persistence"),
# so a hook placed there survives every recreate WITHOUT an image change — the
# boot-time twin of the interactive bridge in /etc/skel-devbox/.bash_aliases,
# which sources ~/.config/devbox-shell/bash_aliases for every interactive shell.
#
# NO NEW TRUST BOUNDARY: that same directory is already sourced into every
# interactive shell, i.e. it is already arbitrary code from the same owner. What
# is new is only WHEN it runs — once at start, before any shell — which is what
# non-interactive fixups (symlinks, dirs, one-off migrations) need.
#
# Deliberately `bash <file>`, not `.` — a hook must not be able to mutate this
# entrypoint's own shell state, and its exit status must not matter. Output goes
# to a log rather than the container's start output, so a chatty hook cannot
# masquerade as a startup error.
if [ -r "$HOME/.config/devbox-shell/init.sh" ]; then
mkdir -p "$HOME/.pi/agent" 2>/dev/null || true
bash "$HOME/.config/devbox-shell/init.sh" \
>"$HOME/.pi/agent/devbox-init.log" 2>&1 || true
fi
# ── Git config defaults ──────────────────────────────────────────────
if [ -n "${GIT_USER_NAME:-}" ] && ! git config --global user.name &>/dev/null; then
git config --global user.name "$GIT_USER_NAME"
@@ -366,6 +471,38 @@ if command -v pi &>/dev/null; then
done
fi
# ── agent-browser: retire a stale volume copy that shadows the image ───
# Same hazard class as the pi-atelier retirement above, different delivery
# path — and this block exists because that guard did not generalise.
# ~/.pi/npm-global lives on the devbox-pi-config VOLUME, so anything ever
# installed there with `npm i -g` survives every image upgrade, and PATH puts
# it AHEAD of /usr/bin (position 2 vs 8).
#
# Measured on mbp-m1-2020, 2026-09-06: a 2026-07-17 hand-install pinned
# agent-browser 0.27.0 in the volume while the image shipped 0.35.2, so every
# session for ~7 weeks ran a stale CLI. The damaging part was not the binary
# but its BUNDLED SKILL, which is what the agent actually reads: 3 skillsets /
# 17.6 KB core in 0.27.0 vs 8 skillsets / 31.5 KB core in 0.35.2, with ten
# subcommands present in the image and undocumented to the agent (a11y,
# browser, data, mcp, page, plugin, read, selectors, to, webmcp). A stale tool
# announces itself; a stale skill quietly teaches the wrong commands.
#
# MOVE rather than delete (reversible, same instinct as the settings backups
# above), and only when the image ships its own copy — a machine that
# deliberately hand-installs agent-browser on an image WITHOUT one keeps it.
_ab_vol="$HOME/.pi/npm-global/lib/node_modules/agent-browser"
if [ -d "$_ab_vol" ] && [ -d /usr/lib/node_modules/agent-browser ]; then
_ab_park="$HOME/.pi/npm-global/.retired-agent-browser-$(date +%Y%m%d-%H%M%S)"
if mkdir -p "$_ab_park" 2>/dev/null && mv "$_ab_vol" "$_ab_park/" 2>/dev/null; then
# The bin shim is what PATH actually hits; leaving it behind would give a
# dangling symlink, which is a worse failure than a stale version.
rm -f "$HOME/.pi/npm-global/bin/agent-browser" 2>/dev/null || true
echo "agent-browser: retired stale volume copy -> ${_ab_park} (image copy now wins; delete the parked dir when satisfied)"
else
echo "WARN: agent-browser: stale volume copy at $_ab_vol shadows the image copy and could not be moved; retire it by hand"
fi
fi
# ── pi-studio: optional loopback bridge (opt-in) ──────────────────────
# pi-studio binds its server to 127.0.0.1 inside the container, which a
# published Docker port cannot reach. When STUDIO_EXPOSE is truthy (set in
Executable
+64
View File
@@ -0,0 +1,64 @@
#!/usr/bin/env bash
# Pre-push gate for pi-devbox: shellcheck every shell script before it leaves
# this clone. Thin wrapper — all logic lives in scripts/lint-shell.sh, which is
# the SAME script the CI release gate runs. One copy, not two: a duplicated
# check that drifts is the failure this repo keeps paying for.
#
# Install per clone: git config core.hooksPath hooks
# Bypass this gate: git push --no-verify (a guard, not a wall)
#
# WHY THIS HOOK EXISTS
# v1.8.14's first release attempt died at scripts/smoke-test.sh:770 after
# build-base had already spent ~46 minutes. shellcheck had ALREADY caught the
# defect — SC2289 at severity error, on the very push that introduced it — and
# the lint job stayed red for 24 hours, unread, across three runs. The fix at
# the time was to gate the release on the same script (the `lint-gate` job).
# This hook is the cheaper end of that: the same finding, before the push,
# in seconds rather than after a 40 s CI gate or a 46 min build.
#
# WHY IT COULD NOT EXIST UNTIL NOW
# Measured on v1.8.14 (2026-09-09): shellcheck was absent from the devbox
# image by all three routes — PATH, dpkg and a filesystem search. So
# lint-shell.sh exited 2 in every container, and a hook calling it would have
# refused EVERY push rather than gating anything. `shellcheck` was added to
# Dockerfile.base in the same change that added this file; on an image built
# before that, enable this hook and you will simply be told the gate cannot
# run. That is the correct behaviour, but it is not a working hook — so do not
# set core.hooksPath on a container older than the release that bakes it.
#
# NOTE ON SCOPE: this lints the WORKING TREE, not the exact commit range being
# pushed. That is deliberate and matches what the CI gate does to the tagged
# tree. It means a defect you have staged-but-not-committed is also reported,
# which is noisy in the safe direction.
set -euo pipefail
HOOK_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
REPO_ROOT="$(cd "$HOOK_DIR/.." && pwd)"
LINTER="$REPO_ROOT/scripts/lint-shell.sh"
tag="[lint-shell]"
# Same rule the gate itself applies, applied one level up: a missing check is
# not a pass. If the script is gone, the push is refused rather than waved
# through on the assumption that CI will catch it.
if [ ! -r "$LINTER" ]; then
echo "$tag refusing the push: $LINTER is missing, so the gate cannot" >&2
echo "$tag run. A gate that cannot run must not pass." >&2
exit 2
fi
# Point the message at the actual remedy when the binary is absent, because the
# linter's own message ("install it or run this in CI") is written for a CI
# runner and is misleading inside a container the developer cannot apt-install
# into persistently.
if ! command -v shellcheck >/dev/null 2>&1; then
echo "$tag refusing the push: shellcheck is not installed, so the gate" >&2
echo "$tag cannot run. A gate that cannot run must not pass." >&2
echo "$tag" >&2
echo "$tag This container predates the image that bakes shellcheck." >&2
echo "$tag Either recreate onto an image that has it, or unset the hook:" >&2
echo "$tag git config --unset core.hooksPath" >&2
echo "$tag To push this once without the gate: git push --no-verify" >&2
exit 2
fi
exec bash "$LINTER" "$REPO_ROOT"
+44
View File
@@ -116,6 +116,50 @@ if command -v fzf >/dev/null 2>&1; then
eval "$(fzf --bash)" 2>/dev/null || true
fi
# cli_utils — shell FUNCTIONS (fgit, fhist, fssh, portcheck, up, mkcd, extract,
# agents-sync, …). This is the OTHER HALF of the cli_utils wiring, and until
# v1.8.11 the image shipped only one half. entrypoint-user.sh symlinks the repo's
# bin/ COMMANDS into ~/.local/bin, which is what makes them resolve in
# NON-interactive shells (docker exec, agent tool shells, scripts). A symlink
# cannot carry a shell function, and a function cannot be reached from a
# non-interactive shell, so the two mechanisms are disjoint and both are
# required. Nothing sourced the loader: measured 2026-08-30 on v1.8.11, all 14
# functions were simply missing on a device whose $HOME has no zsh rc — which is
# the normal case, since the container's interactive shell is bash and zsh is not
# installed in the image. The image was already paying this layer's dependency
# cost (fzf, bat, fd, rg, jq are all baked partly FOR these functions) while
# delivering none of its benefit.
#
# Detection order deliberately mirrors the symlink block in entrypoint-user.sh so
# that commands and functions can never come from two different clones.
# CLI_UTILS_SOURCE=0 opts out. That is independent of CLI_UTILS_LINK=0 on purpose:
# they disable independent mechanisms, and someone who wants PATH commands
# without 14 extra functions in every prompt (or vice versa) should be able to
# say so.
#
# THE LOADER IS BASH-SAFE, MEASURED, NOT ASSUMED: despite every function file
# being named *.zsh, sourcing cli_utils.sh under `bash --noprofile --norc` exits
# 0 with no errors and defines all 14, and they run (pathls, mkcd, up, extract,
# agents-sync, fhist all verified). The single zsh-only construct in the tree
# (`print -z` in fzf/fhist.zsh) is already guarded by [[ -n $ZSH_VERSION ]] with
# a bash fallback, and the loader's own header states "bash & zsh compatible".
# ACCEPTED RISK, stated plainly: /workspace/cli_utils is a HOST BIND MOUNT, so
# unlike a pinned git ref this content floats outside the image's control. A
# future cli_utils commit that adds a genuinely zsh-only file would surface as
# parse errors at every prompt on every device. Errors are left VISIBLE rather
# than sent to /dev/null so that failure is diagnosable instead of mysterious,
# and CLI_UTILS_SOURCE=0 is the documented one-line escape hatch.
if [ "${CLI_UTILS_SOURCE:-1}" != "0" ]; then
for _cu in "${CLI_UTILS_CONTAINER_PATH:-}" /workspace/cli_utils "$HOME/cli_utils" /workspace/*/cli_utils; do
[ -n "$_cu" ] || continue
if [ -r "$_cu/cli_utils.sh" ]; then
. "$_cu/cli_utils.sh" || true
break
fi
done
unset _cu
fi
# ── PROMPT_COMMAND: flush history every prompt ───────────────────────
# Installed AFTER zoxide init so zoxide's hook is already in place;
# we append with a newline separator to avoid the ';;' parse error
+27 -1
View File
@@ -157,6 +157,11 @@ if [ "$SHOW_SKILLS" = "yes" ] && [ -d "$BAKED_SKILLS" ] && [ -d "$SKILLS_DIR" ];
# is not hypothetical.
snap_ref=$(jq -r '.skillset_snapshot_ref // empty' "$MANIFEST")
snap_sha=$(jq -r '.skillset_snapshot_tree_sha256 // empty' "$MANIFEST")
# Which pi-extensions copy the BUILD baked. Distinct from everything else in
# this section, which reports which copy is being READ at runtime: for
# pi-extensions the baked tree is itself one of two possible copies, and that
# choice was made at build time and is not recoverable by inspection.
px_src=$(jq -r '.pi_extensions_skill_source // empty' "$MANIFEST")
# Same pipeline Dockerfile.variant uses to measure the baked directory at
# build time: relative paths in `find | sort` order, each hashed, the whole
# listing folded into one sha256. Keep the two definitions identical — they
@@ -187,7 +192,28 @@ if [ "$SHOW_SKILLS" = "yes" ] && [ -d "$BAKED_SKILLS" ] && [ -d "$SKILLS_DIR" ];
_target=$(readlink -f "$_link" 2>/dev/null || echo "$_link")
case "$_target" in
"$BAKED_SKILLS"/*|"$BAKED_SKILLS")
printf ' %-22s baked\n' "$_name"
# "baked" alone used to be the whole story. For pi-extensions it is not:
# the baked tree holds EITHER the package copy that Dockerfile.variant
# lays over the snapshot, OR the vendored floor, when the clone had no
# skill/ at that ref. The two are indistinguishable by inspection — same
# path, same filenames, same permissions — so the build records which one
# it used and this reports it. Without this line a six-week-stale
# fallback skill looks exactly like a current one, which is precisely how
# the floor went unnoticed from 2026-07-30 to 2026-09-10.
if [ "$_name" = "pi-extensions" ] && [ -n "$px_src" ]; then
case "$px_src" in
package)
printf ' %-22s baked (package copy)\n' "$_name" ;;
vendored-floor)
printf ' %-22s baked \033[33m(FALLBACK: vendored floor — clone had no skill/)\033[0m\n' "$_name" ;;
divergent)
printf ' %-22s baked \033[33m(MIXED: part package, part floor)\033[0m\n' "$_name" ;;
*)
printf ' %-22s baked\n' "$_name" ;;
esac
else
printf ' %-22s baked\n' "$_name"
fi
continue
;;
esac
@@ -70,3 +70,41 @@ rather than merely confusing you:
local disk, so `mempalace search` can return older and different results than
the MCP tools while both look correct. Use the MCP tools for the central
palace; the CLI only for a local one.
## Before you file a finding: second measurement, different route
This is here rather than in a skill because it has to fire *without* a matching
task description, and because the version of it that lived only in a skill was
violated five times in one session by an agent that had the skill available.
**Any claim you are about to record as fact — in a drawer, a diary entry, a
coordination event, or a report to the user — needs a second measurement taken
by a different route.** Not a re-read of your reasoning: re-reading has caught
zero of these. A disagreeing measurement has caught all of them.
The two shapes that get filed as fact and are not:
- **A negative result** (`401`, connection refused, zero rows, "not found") is
first a claim about *your filter*, not about the world. Wrong host, wrong port,
wrong table, capped output.
- **A positive result** proves only what your command *actually asked*. An SSH
handshake can succeed against the wrong host (`ssh -G` tells you which rule
captured the name); a `401` can be a real answer from an issuer that never
minted the credential.
Cheapest habit that works: **write the expected result next to each check before
running it**, then diff. Expectations declared up front turn a silent wrong
assumption into a visible mismatch. And if you cannot think of a second route to
the same fact, you do not have a finding — you have a hypothesis, so label it as
one.
## Handling an exposed credential
If a task touches a leaked secret, a token rotation, "is this credential still
live?", whether to delete stored content, or which scopes a new token needs:
**read `~/.agents/skills/credential-incident-response/SKILL.md` first.** One rule
is load-bearing enough to state here: **probe the issuing provider before doing
anything else** — most "exposed" credentials in a long-lived fleet are already
dead, and the ones that are live are often far more privileged than assumed.
Severity first, cleanup second, and prefer **revocation over deletion** for
anything already replicated.
@@ -9,6 +9,7 @@ one", which was a bug).
| skill | owner | how it gets here |
|-------|-------|------------------|
| `pi-devbox-environment` | pi-devbox (this repo) | authored here; the canonical copy |
| `credential-incident-response` | pi-devbox (this repo) | authored here; the canonical copy |
| `pi-extensions` | the `pi-extensions` package repo (`skill/`) | **vendored fallback** + refreshed at build |
| `mempalace` | the `skillset` repo | **vendored fallback** (snapshot only) |
@@ -0,0 +1,271 @@
---
name: credential-incident-response
description: >-
Respond correctly when a live credential is found where it should not be —
in a chat transcript, a MemPalace drawer, a log, a git-tracked config, or an
agent-authored note. Load this whenever a task involves a leaked/exposed
secret, a token rotation, a "is this credential still live?" question, deciding
whether to delete or scrub stored content, proving a corpus is clean, or
choosing scopes for a new API token. Covers the mandatory order of operations
(probe the issuer FIRST — severity before cleanliness), leak-free identity via
sha256[:8] fingerprints and when publishing one is safe,
why revocation beats deletion for anything already replicated, scopes derived
from measured consumers, the three places a secret hides in a Chroma palace, how to prove ABSENCE rather than assume it (instrument strength,
census vs class passes, the tokenisation trap where quoting decides detectability, why git filters never run on symlinks, self-tests that abort),
where this fleet's secrets live, and what rotation does NOT fix.
---
# Credential incident response
A leaked credential is a **severity** question before it is a cleanliness
question. Two days of scrubbing, redaction plumbing and deletion planning were
once spent on a set of 13 credentials of which **11 were already dead at the
provider** — a fact that cost five HTTP requests to establish and was never
checked. Meanwhile the two live ones turned out to be instance-owner **admin**
tokens, which nobody had looked at either.
## 1. Order of operations — do not reorder this
1. **Is it still accepted?** Probe the issuing provider. Dead credential →
hygiene item, stop panicking. Live → incident, continue.
2. **What can it do?** Read the identity back. `is_admin`, `id=1`, scopes,
which account. A read-only repo token and an instance-owner admin token are
not the same finding.
3. **What consumes it?** Grep for real consumers before assuming breakage.
4. **Where does it live?** Enumerate copies (store, palace, transcripts, git).
5. **Then** rotate/revoke, and only then consider cleanup.
Doing 4→3→1 in reverse produces confident, wrong severity calls and wasted
cleanup. If you only have time for one step, do step 1.
## 2. Leak-free identity: fingerprint, never the value
Publishing an 8-hex fingerprint lets you compare a credential across machines,
files, drawers and peers without ever materialising the secret. Same formula as
`mempalace_redact.py`:
```sh
printf '%s' "$SECRET" | sha256sum | cut -c1-8 # printf, NOT echo (no newline)
printf '%s' 'test' | sha256sum | cut -c1-8 # self-test -> 9f86d081
```
Report as `(variable, fp, length)`. Equal fingerprints across hosts prove a
shared credential; that is usually the important part. **Never** paste a live
value into a search query, a palace drawer, an event body, or a chat message —
in an agent context your own tool output is itself captured and re-filed.
**Precondition — only fingerprint what an adversary cannot enumerate.** An 8-hex
fingerprint is 32 bits over its *input space*, so publishing `fp8(x)` hands
anyone a **membership oracle**: they can test `x == v` for every candidate `v`
they can generate. For a 40-char random token that space is unreachable. For a
hostname, username, e-mail, port, path, commit SHA or weak password it is a
wordlist. **If you can imagine writing the wordlist, you cannot publish the
fingerprint** — reference those by name and location instead. "High entropy" is
the usual *sufficient condition*, not the test: a commit SHA is 160-bit and still
fully enumerable from the repo. `sha256("")` = `e3b0c442` is the degenerate case,
recognisable on sight precisely because its input space has one member.
**Candidate fingerprints are working memory, never output.** A scanner that hashes
every token in a file also hashes hostnames, paths and e-mails. Print only
fingerprints that *matched* a known entry — the tempting debug step when a scan
returns zero ("print what it saw") publishes low-entropy fingerprints wholesale.
And say plainly what a fingerprint register *is*, so nobody rediscovers it later
as an alarm: even for an unguessable secret, a published fingerprint is a
**confirmation oracle** for anyone who already holds a candidate corpus. That is
exactly how a long-retired token gets identified in old transcripts — and it works
identically for someone else holding those same files. Net positive, since they
would already hold the value; state it rather than leaving it implicit.
## 3. Liveness probes, and the trap that scoping creates
```sh
# Gitea
curl -sS -m 10 -o /dev/null -w '%{http_code}\n' -H "Authorization: token $T" \
"$GITEA_HOST/api/v1/repos/<owner>/<repo>/actions/runs?limit=1"
# GitHub
curl -sS -m 10 -o /dev/null -w '%{http_code}\n' -H "Authorization: token $T" \
https://api.github.com/user
```
- `200` live · `401` revoked/invalid · **`403` = wrong question, not a dead token**
- **Probe the issuer that minted it.** A 401 from an unrelated instance says
nothing. Resolve the host from config (`GITEA_EGL_HOST` etc.), do not assume.
- **Under scoped tokens, `/api/v1/user` returns 403 for a perfectly live token**
unless `user` scope was granted. So it cannot distinguish *revoked* from
*merely scoped*. Use a **repository route the token is authorised for**.
- Verify **both directions** after a rotation: old → 401, new → 200. The second
check is what catches "deleted the wrong token".
- Port/scheme come from config, not habit: one instance here is
`http://gitea.egl.lan:3000` — plain HTTP, with 443 refused.
## 4. Revocation beats deletion — the load-bearing rule
Once revoked, stored copies are **inert**; you may leave them. Deleting them is
best-effort over an *unbounded* copy set: FTS shadow rows, feed inbox `.jsonl`
files on every host, sqlite free pages after the delete, mesh replicas that
already synced, and backups. **Revocation invalidates every copy everywhere at
once, including copies nobody enumerated.**
So: **rotate + revoke first.** Treat drawer deletion as optional hygiene, never
as the remedy. Then record the retired fingerprints as *known-dead* so the next
census recognises them instead of reopening the investigation.
Corollary: never reach for `mempalace_sync` or a bulk `delete_by_source` on a
shared palace as incident response. High blast radius, low actual benefit.
## 5. Finding a secret in a Chroma palace — three targets, in this order
1. `embedding_fulltext_search_content.c0` — document text
2. `embedding_metadata.string_value` — metadata fields, **and a second copy of
the document text** under key `chroma:document`
3. raw byte scan of every `*.sqlite3` — backstop, covers FTS pages and free space
**Correction, measured on chroma 1.5.9 with a sentinel drawer:** one row in (1)
AND one row in (2) for the same drawer, so **(2) is not structurally
content-blind** — an earlier version of this section said it held "metadata
fields only", and that was wrong. Scan (1) and (3) regardless: (1) is the direct
target. But if a `string_value` query returns zero for a value you know is in a
drawer, the cause is a key filter, a query shape or escaping — *not* structural
absence, and the difference matters because the false explanation is what makes
the zero feel safe. See §6: do not explain a zero with a mechanism you have not
read from source.
Semantic search proves nothing about absence — it returns top-k. For
completeness, enumerate by filing window (`list_drawers(since=T, before=T+1m)`),
since one mine shares a minute.
Value-agnostic sweeps (uuid / 40-hex / `NAME=VALUE`) drown in false positives at
fleet scale — 608 candidates, mostly session UUIDs and git SHAs. Name-anchoring
plus entropy plus provenance, applied to **document text**, is what works.
## 6. Proving absence: instrument strength, and four ways a scan lies clean
Section 5's warning is about false *positives* — name-anchoring and provenance are
what stop a triage sweep drowning in session UUIDs. **A gate is the opposite job.**
Triage optimises precision; proving absence optimises recall. Every failure below
reported a reassuring zero over a secret that was really there.
**Rank the instrument, and state which one produced your zero.**
| Instrument | Needs | Blind to |
|---|---|---|
| exact-byte value search | you hold the value | nothing — no tokeniser to fool |
| class/structure pass | a header pattern | anything without a recognisable shape |
| fingerprint census | a fingerprint list | any secret not listed; tokenisation |
A census is deliberately value-free, so it must *extract candidates and hash them*
— which makes its sensitivity a property of the tokeniser, not of the corpus. If
you hold the value, search the bytes instead, and search the value's JSON-escaped
rendering too when the corpus is `.jsonl`.
**1. Census and class answer different questions; neither substitutes.** A census
answers *"has a KNOWN secret leaked?"*, a class pass *"is there secret-SHAPED
material here?"* Both failure modes were measured on this fleet: a class-only
pre-commit hook passed plaintext UUID API credentials to a shared repo twice,
because a UUID carries no key header — while a census-only gate reported 0 hits
with freshly-synced SSH private keys and an age identity in the tree, because no
key is in the census. Run both passes.
**2. Tokenisation — quoting alone can decide detectability.** Maximal-run
extraction swallows the value of an *unquoted* assignment:
```
PROXMOX_SECRET=<uuid> # ONE run; the uuid is never hashed alone -> MISS
export SECRET="<uuid>" # the quote ends the run; bare uuid hashed -> HIT
```
Take the **union** of three strategies, because each fails in a different
direction — (2) is the one that recovers the unquoted case:
~~~python
runs = re.findall(r'[^\s"\'`]{12,}', text) # 1. maximal runs
split = [p for r in runs for p in re.split(r'[=!,;:@|()\[\]{}<>]', r) if len(p) >= 12]
shape = re.findall(UUID_RE, text) + re.findall(r'[0-9a-f]{32,64}', text)
candidates = set(runs) | set(split) | set(shape)
~~~
**3. Scan the index or the pushed tree, never the working tree.** The working tree
is not what gets published. And for an rsync-published mirror a repo-only fix is
not weaker, it is *temporary*: the next sync re-publishes the live disk. Fix the
live file first, verify it clean **by fingerprint**, then sync. Read blobs with
`git ls-tree -r <sha>` plus one `git cat-file --batch` (thousands of `git show`
calls is the slow way).
**4. Git filters never run on symlinks — and `check-attr` will not tell you.** A
symlink's blob is the *target path*, so `filter=git-crypt` can never encrypt it,
yet `git check-attr filter` cheerfully answers `git-crypt` for that path. **A
symlinked secret stays plaintext no matter what `.gitattributes` says.** Join the
attribute against the **file mode** (`git ls-files -s`, mode `120000`) and verify
the index blob really begins `\0GITCRYPT\0`. Report encrypted / symlinked /
scanned as three separate numbers and assert they sum — encrypted and symlinked
blobs are *skipped*, not certified clean.
**Self-test two-sided, and abort if it cannot discriminate.** Require a synthetic
positive to fire AND a negative to stay silent before trusting any zero. Keep the
fixtures in *structurally separate buffers*: put a quoted and an unquoted probe in
one buffer and the quote terminates the run, handing the bare token to the weak
extractor and making it look as strong as the union — a self-test artifact that
has already fooled an agent here. And never gate on `$?` when the tool has a
lock-skip or no-op path that also exits 0; judge the reported line.
**Row-gone is not bytes-gone.** Measured, same sentinel drawer: after
`delete_by_source` the row count went 1 -> 0 in *both* the FTS content table and
`embedding_metadata`, while the raw byte count stayed 4 -> 4 — sqlite does not
zero freed pages, so the payload sits in free space until `VACUUM`. Deletion
effectiveness is therefore *two* numbers, and each direction has a trap: one
aggregate figure reported as "erased" has only measured "unretrievable", while a
raw byte scan used as the acceptance gate reads a CORRECT, complete deletion as a
failure. (Note how this was measured: the blocker was never a better instrument,
it was the subject — file your own disposable sentinel and delete that, instead
of testing deletion on real data.)
## 7. Choosing scopes: derive them from measured consumers
Before creating a replacement token, find out what actually uses it:
```sh
git -C <repo> remote get-url origin # ssh:// ? then git needs NO token
git config --global --list | grep -iE 'credential|insteadof' # and no helper?
grep -rhoE 'api/v1/[A-Za-z0-9/{}$_.-]+' <consumers> | sort -u # exact routes
grep -rhoE '\-X [A-Z]+' <consumers> # any writes?
```
Real outcome here: git used SSH keys throughout, and the token's only consumer
read three CI-run routes with `GET`. So `repository: Read` and nothing else
replaced two admin tokens. **Scoping shrinks the blast radius of the next leak
far more than any redaction pipeline does** — a read-only token in a transcript
is a hygiene event, not an instance compromise.
Then prove the scope with an acceptance suite that declares expectations first:
must-work routes → `200`; `/admin/*`, `/user`, `/user/repos` → `403`.
## 8. What rotation does *not* fix
- **A cleartext channel.** If the endpoint is `http://`, the *new* token is
exposed identically from first use. Raise TLS separately.
- **Git history.** A secret committed and pushed cannot be fixed by any store or
palace operation — it needs rotation *and* history surgery.
- **Agent-authored content.** Stage-write redactors see transcripts only, never
`add_drawer` / `checkpoint` / `diary_write` output. Never type a secret into
the palace yourself; nothing downstream will catch it.
- **Plaintext/encrypted drift.** Gitignored plaintext `.env` files go stale while
`.env.age` moves on, so old values linger on disk (and in backups) long after
rotation. They are a common source of "mystery" fingerprints in a census.
## 9. This fleet's secret store (verify, do not assume)
- All `*.env.age` live in **one** repo: `joakimp/docker-compose-repo`. `myconfigs`
has none.
- Every `.age` file has **one X25519 recipient** — a single key tracked in
`myconfigs` under git-crypt. Unlocking git-crypt therefore decrypts the entire
fleet's secrets, including hosts you have no access to. The age layer adds no
isolation beyond git-crypt.
- Flow: `./fetch-secrets.sh <host>` (decrypt → `.env`) → edit → `./encrypt-secrets.sh <host>`
→ commit → push → `docker compose up -d --force-recreate`.
- **Always pass the host argument** to `encrypt-secrets.sh`. Bare, it walks the
whole tree and re-encrypts every `.env` it finds, re-nonced, including stale
ones — silently rolling back other hosts' secrets.
- After any re-encrypt, check the header still shows exactly **one X25519
recipient**; a hand-rolled `age -r` locks the rest of the fleet out, and the
failure only appears on another machine, later.
@@ -428,10 +428,56 @@ An obligation you never agreed to is noise, so the sender states it:
| `to_agent="*"` (any status) | broadcast FYI | nothing |
| any other status (`ready`, `applied`, `blocked`, …) | a statement of fact | nothing |
**That table says what you *owe*. Delivery is stricter, and the difference bites:
the mailbox is an obligation channel, not a news channel.** Mailbox candidates are
drawn with `status="open"`, so an event carrying any **terminal** status
(`applied`, `superseded`, `failed`, `blocked`) is never a candidate — *whoever it
is addressed to*. A `task.reply` written to a named machine to share a finding is
delivered to nobody, ever, and neither is any `event_ack`. It sits in the log
until somebody reads the log.
So the most natural inter-machine message — *"here is something you should
know"* — is exactly the shape that gets no delivery. Pick deliberately:
| You want the peer to… | Write |
|---|---|
| **do something**, and you need it tracked until done | directed `status="open"` ask, with a `correlation_id` |
| **know something**, no response needed | terminal-status event **plus a drawer** — the drawer is what actually reaches them, via search |
What does **not** work is a terminal report plus an expectation of attention.
Measured 2026-08-26: a detailed report addressed to `pi@<peer>` with
`status="applied"` went unread for two and a half hours until the operator quoted
the event id by hand, with the mailbox working correctly the whole time. Full
mechanism in the toolkit's `docs/rfc-003-coordination-log.md` §7.12.
One more timing fact, because it looks like negligence and is not: a delivered
ask is queued into the agent's **next turn** (`deliverAs: "steer"`, deliberately
no `triggerTurn`), and the poll fires when the agent is *idle*. Between delivery
and the next turn no inference runs, so **a human starting a turn is the
trigger** (§7.11). An agent that "has not reacted" has usually not been running.
Ack with `mempalace_event_ack(event_id=…, from_agent="<you>", status=…)`. It
**appends a new event** and never mutates the original; the correlation id is
copied for you, and `metadata.ack_of` is set to the event you answered.
**Claiming, and what it does not do.** `status="claimed"` announces that you have
picked work up. Nothing requires it — a directed open ask owes "an ack *or* a
reply", and finishing the work is a complete answer. Do it anyway when the work is
long or the machine is unreliable, because it is the only thing that later
distinguishes *nobody started this* from *someone started and their container
died mid-task*. Be clear about its limits, both of which follow from candidacy
requiring exactly `status="open"`:
- **It does not notify the requester.** `claimed` is not `open`, so a claim is no
more deliverable than a finished report is (see the delivery table above). Its
reader is whoever pulls the log.
- **It does not quiet your own mailbox.** The ask stays owed until a *terminal*
event of yours joins it, so a claimed-then-silent thread keeps resurfacing —
correctly.
Prefer a prompt terminal reply over a claim plus a long silence; claim *in
addition*, when the gap between pickup and finish is where a machine might die.
#### What you actually owe — derive it, do not read it off `status`
The log is append-only and `status` is written **once**, so it is an honest
@@ -489,7 +535,10 @@ Two consequences worth internalising:
- **"Seen, not doing it" is a legitimate ack** — `status="blocked"` or
`"superseded"` plus the reason. Silence is not, and it is not merely rude:
with no terminal event of yours to join to, the ask stays in the owed set
indefinitely and there is nothing anyone can do about it from the other end.
indefinitely. The *original requester* — and nobody else — can release it from
the other end, but only by saying so explicitly: see **Withdrawing an ask you
sent** below. That is a release by the asker, not an escape for the answerer.
While the ask still stands, only *your* terminal event clears it.
- **Nothing expires, and it should not.** An `open` with no terminal reply is
still live by definition, and the finished threads are valuable history. If
content is genuinely perishable ("do not push to main for the next hour"), say
@@ -502,6 +551,17 @@ Two consequences worth internalising:
- **Address the stamped name you actually saw** in a `from_agent` field, e.g.
`pi@tor-ms22`. A bare `pi` reaches nobody's mailbox once stamping is live, and
older events in the log still carry bare names — do not copy them.
- **The rule runs in reverse too: what you put in YOUR OWN `from_agent` decides
where every reply to your event goes.** Nothing stops you writing a synthetic
or borrowed identity there, and a reply is always addressed back to exactly
that string — so if no live session ever runs as it, the reply is stored,
searchable, and delivered to no one. Measured cost: a directed ask sent under
a synthetic sender got two correct replies, one of them an urgent security
finding, and both sat unread for ~2h20m because nobody's mailbox was that
identity (RFC 003 §7.13). Authoring under a synthetic name is fine for a
deliberate control experiment — this fleet does it on purpose — but then
**name the real identity to reply to inside the body**, because the address
line is not a safe place to also carry provenance.
- **Use `status="open"` only when you truly need an answer.** It places an
obligation on another machine.
- **Never broadcast an ask.** `to_agent="*"` + `status="open"` obliges everyone
@@ -513,7 +573,26 @@ Two consequences worth internalising:
be matched to it at all.
- **Corrections are new events, never edits.** Say explicitly what you retract
and name the id — drawer or event — that carried the withdrawn claim.
- **Put a retraction where the reader will look.** An event reaches a live agent;
- **Withdrawing an ask you sent: state it, never imply it.** Your release only
counts when the terminal event (a) comes from the same `from_agent` that sent
the ask, (b) is directed at that recipient exactly — never `*`, so a broadcast
can neither oblige nor release, (c) carries a terminal status (`claimed` and
`ready` are not terminal and do not release anything), (d) is strictly after
the ask, (e) joins it via `ack_of` or the same `correlation_id`, **and (f)
names that ask in `metadata.withdraws` or `metadata.closes`.** Prose in the
body does not count, and neither does a bare terminal event on the
correlation: inferring release from *any* terminal would let your own
bookkeeping silently delete a real obligation, so the release must be stated.
Needs toolkit ≥ `e2b060a` (image ≥ `v1.9.1`) — check with
`grep -c isWithdrawn /opt/mempalace-toolkit/extensions/pi/mempalace.ts` and
read `0` as "my withdrawal will have no effect on their mailbox". Measured
cost of getting it wrong: a `v1.8.13` rollout ask was withdrawn by its sender,
who recorded it as done; the recipient's derivation never saw the release and
still reported the ask owed **41 hours later**, for a release that device
never installed — and the asymmetry was invisible from the sender's side
(RFC 003 §3.3 clause 4).
- **Put a retraction where the reader will look.** A *directed open ask* reaches a
live agent's mailbox; a **terminal-status event reaches no mailbox at all**, and
a *drawer* is what a future semantic search finds. If you filed advice as a
drawer and later withdraw it, file the withdrawal as a drawer too — otherwise
the next agent finds your original confident advice and no trace of the
@@ -529,9 +608,31 @@ Two consequences worth internalising:
### Wings
Wings are top-level categories, typically one per project or domain:
- Named after the project directory (e.g., `cli_utils`, `opencode_devbox`)
- Agent diaries live in `wing_<agent_name>` (e.g., `wing_orchestrator`, `wing_pi`)
Wings are top-level categories, typically one per project or domain.
**NAMING CONVENTION — decided 2026-09-06 by Joakim: bare project names, no `wing_`
prefix.** `home-network`, `pi-devbox`, `mempalace-toolkit` — *not* `wing_pi-devbox`. The
mass is already there (`pi-devbox` 2061 drawers vs `wing_pi-devbox` 25), and a prefix
present on some wings and absent on others turns every read into a guess about which
spelling holds the content.
- Named after the project directory or domain (e.g., `cli_utils`, `home-network`)
- **Always pass `wing` explicitly to `diary_write`.** Omitting it defaults to
`wing_{agent_name}`, which mints or feeds a *parallel* wing — this tool default, not
anyone's sloppiness, is the mechanism that produced the drift. Measured harm
(2026-09-06, `pi@mbp-m1-2020`): a diary entry written with `agent_name=pi` and no
`wing` landed in `wing_pi` while that agent's history lives in `pi-devbox`, so a
`diary_read` scoped to `pi-devbox` showed **no trace of it**. A wing-scoped read that
silently returns an incomplete history is the worst failure mode a memory store has.
- **Legacy `wing_*` wings are frozen and documented, not renamed.** `wing_conversations`
(written by the session feeders), `wing_pi`, `wing_pi-devbox`, `wing_pi-tor-ms22`,
`wing_pi-devbox-emb7kj`, `wing_mempalace`, `wing_orchestrator`, `wing_code` all still
hold real content. **When searching for history, check both spellings** — this is the
practical cost of the drift and it does not go away by decree.
- If a migration is ever done, the acceptance criterion must be at the **relationship**
level: chunk ids still resolve to their parent, and `diary_read` returns the same entry
set before and after. Per-wing drawer counts can look correct while the relationships
underneath are broken, because a count query never touches them.
#### Shared palace: multiple harnesses, and possibly multiple machines
@@ -551,7 +652,7 @@ Zechner's pi-coding-agent). Implications:
When the palace is **central** (shared across machines), these further things apply:
- **Check which machine a conversation came from.** Transcripts are fed per device, so `source_path` reads `…/mempalace-feed/<device>/pi_<uuid>.jsonl` while the displayed `source_file` is only the basename. One search can legitimately return hits from several machines at once — look at the device segment before attributing a decision to *this* project.
- **Provenance is stamped for you — leave it alone.** Drawers carry `device` and `agent_kind` metadata (plus `device_source`/`agent_kind_source` recording *how* each was determined, so an inference is never mistaken for a fact). You do **not** set these, and you no longer set `added_by` either: the pi bridge defaults the writer field to `<harness>@<device>` on `add_drawer`/`checkpoint`/`mine`/`event_append`/`artifact_put`, and prefixes diary entries with `HOST:<device>|`, from host-supplied `$MEMPALACE_PI_DEVICE`. RFC 001 §7.3.2 ranks "agent stamps it via a skill instruction" as the *worst possible* place for exactly the reason you would expect — it is per-call boilerplate that gets forgotten, and it did: the agent who wrote the previous version of this bullet then filed its own provenance drawer as `added_by=checkpoint`. **Confirm the bridge in your image actually stamps before trusting it:** the extension is baked at image build time, so a container on an image older than the stamping commit (pi-devbox < v1.8.7) stamps nothing while still satisfying both gates — the env vars are set and the code is simply absent. Check with `grep -c MEMPALACE_PI_DEVICE "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)"`; zero means keep passing `added_by="<harness>@<device>"` and a manual `HOST:<device>|` diary prefix until the container is recreated on a newer image. Two things remain yours: pass `source_drawer_id` on `kg_add` (triples have no provenance field, so that pointer is the only path back to a device), and pass an explicit `added_by` **only** when deliberately filing on behalf of another device. Never invent values for `device`/`agent_kind`/`origin_device` — a fabricated value is worse than a blank, because it silently corrupts a future merge.
- **Provenance is stamped for you — leave it alone.** Drawers carry `device` and `agent_kind` metadata (plus `device_source`/`agent_kind_source` recording *how* each was determined, so an inference is never mistaken for a fact). You do **not** set these, and you no longer set `added_by` either: the pi bridge defaults the writer field to `<harness>@<device>` on `add_drawer`/`checkpoint`/`mine`/`event_append`/`artifact_put`, and prefixes diary entries with `HOST:<device>|`, from host-supplied `$MEMPALACE_PI_DEVICE`. RFC 001 §7.3.2 ranks "agent stamps it via a skill instruction" as the *worst possible* place for exactly the reason you would expect — it is per-call boilerplate that gets forgotten, and it did: the agent who wrote the previous version of this bullet then filed its own provenance drawer as `added_by=checkpoint`. **Confirm the bridge in your image actually stamps before trusting it:** the extension is baked at image build time, so a container on an image older than the stamping commit (pi-devbox < v1.8.7) stamps nothing while still satisfying both gates — the env vars are set and the code is simply absent. Check with `grep -c MEMPALACE_PI_DEVICE "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)"`; zero means keep passing `added_by="<harness>@<device>"` and a manual `HOST:<device>|` diary prefix until the container is recreated on a newer image. Two things remain yours: pass `source_drawer_id` on `kg_add` (triples have no provenance field, so that pointer is the only path back to a device), and pass an explicit `added_by` **only** when deliberately filing on behalf of another device — and when you do, it **must** be `<harness>@<device>`. A bare nickname (`pi-devbox-claude`) has no `@device` to parse, so `agent_at_device` cannot attribute it and the drawer is unattributable *by rule*, not by lag: it survives every future stamp run with no `device`, and on a shared palace a device-less drawer is one nobody can later scope, audit or clean up per machine. Measured 2026-09-06: 11 drawers on `tor-ms22` were filed this way — including the credential rows, i.e. exactly where "which machine measured this?" matters most — by an agent that had passed its own chosen nickname on every call. Its *diary* entries escaped, because `HOST:<device>|` in the AAAK text recovers the device. **Diaries self-heal; plain drawers do not.** The safest habit is the one above: pass nothing and let the bridge stamp. Never invent values for `device`/`agent_kind`/`origin_device` — a fabricated value is worse than a blank, because it silently corrupts a future merge.
- **Metadata is invisible to search — so check the text, not the fields.** `search` results are built from a fixed key list and `diary_read` returns content, so neither ever shows `device`/`added_by`. Only `mempalace_get_drawer` reveals them. This is why diary entries carry an in-text `HOST:<device>` marker: it is the only attribution a reader actually sees. **A diary entry with no `HOST:` marker predates the convention and may be from any machine — do not assume it is this one's history.**
- **Mined drawers carry the MINE date, not the session date.** When history is imported, or re-mined on the palace host, `filed_at`/`created_at` is the *import* time — so sorting by them does not give chronological order. Real session time is recoverable from the UUIDv7 in `pi_<uuid>.jsonl`: the first 12 hex digits are milliseconds since the epoch (and UUIDv7 sorts lexicographically in time order, so a plain filename sort is already chronological). Agent-authored drawers and diaries have no such backdoor — for those `filed_at` is the only chronology, which is why it must never be restamped.
- **Beware the timezone mismatch when you combine those.** Palace `filed_at`/`created_at` are naive timestamps in the palace host's local time, while a UUIDv7 decodes to UTC. Comparing them directly introduces a silent offset (2 h for a CEST host). Normalise before drawing conclusions about ordering.
@@ -600,4 +701,5 @@ Entity-relationship triples with temporal validity. Query with `mempalace_kg_que
- **Don't treat the palace as a task list.** It's for knowledge and context, not todos.
- **Don't broadcast an ask, and don't leave one unanswered.** On a shared palace, `to_agent="*"` + `status="open"` obliges every machine and therefore none of them. And don't expect acking to tidy your mailbox: `status` is immutable, so the event keeps matching either way — what a terminal reply buys you is that the *derived* owed set (see *What you actually owe*) stops counting it. Leave asks unanswered and that set only grows, until everyone learns to stop looking. "Seen, not doing it" is a complete answer — silence is not.
- **Don't assume you would have heard.** Nothing pushes another machine's message into your session. If you did not run the mailbox query at wake-up, a correction addressed to you by name can sit unread while you confidently rebuild the thing it warned you about.
- **Don't author an ask under an identity nobody runs as, including your own throwaway labels.** The failure is symmetric to the one above: it is not that you missed a message, it is that nothing could ever have delivered the reply to you, because you addressed it at a name instead of an agent. If you must use a synthetic sender for a control or an experiment, say inside the body who should actually receive the reply.
- **Don't invent provenance metadata, and don't hand-stamp it either.** An earlier version of this list told you to set `added_by="<harness>@<device>"` by hand; that instruction has been withdrawn, because RFC 001 §7.3.2 places provenance at the client/server boundary and the pi bridge now does it uniformly (see *Provenance is stamped for you* above) — but the withdrawal only holds where the bridge is live, so run the one-line check in that bullet first; on an older image hand-stamping is still the only signal a hand-filed drawer gets. DO NOT invent values for the palace's own metadata fields (`device`, `agent_kind`, `origin_device`): those are stamped by infrastructure that also records *how* each was determined, and a fabricated value is worse than none because it silently corrupts a future merge. DO pass `source_drawer_id` on `kg_add`. And never put a machine name in a diary's `agent_name` — it becomes the wing name and hides your entries from `diary_read`.
@@ -143,6 +143,10 @@ mine:
| "`tor-ms22` is not in the SSH config" | `grep … \| head -20` — the entry was at **line 454**. `~/.ssh/config` here is ~500 lines. |
| "the Docker host has no `docker`" | non-interactive SSH `PATH` lacks `/usr/local/bin` (§2, §3). It was at `/usr/local/bin/docker`. |
| "no ControlMaster is running" | pattern `ssh ` (trailing space) cannot match a master: those processes **rename themselves** to `ssh: <controlpath> [mux]`. |
| "the credential is not in the palace" | scanned `embedding_metadata.string_value` only. Drawer **text** lives in `embedding_fulltext_search_content.c0`; 554k metadata rows proved nothing. |
| "this token is dead — 401" | probed it against the **wrong issuer**. A 401 from an instance that never issued the credential is not evidence about the credential. |
| "that host is unreachable, can't test" | tried ports 443 and 80. It was on **3000**, and the env var I already held (`GITEA_EGL_HOST`) stated the scheme and port. |
| "this repo has no `## Unreleased` convention" | read `CHANGELOG.md` **once**, minutes after a release commit had renamed that section to a version heading. 33 commits touch `## Unreleased`. A snapshot cannot show you a cycle. |
Habits that would have caught all three:
@@ -155,10 +159,55 @@ ssh -F "$HOME/.ssh-local/config" mac 'command -v docker || ls /usr/local/bin/doc
# match a process's ACTUAL argv, not the name you imagine
ps -eo pid,etime,args | grep -Ei 'mux|mosh|ssh'
# to learn a repeating PROCESS or convention, read history, not the file. A
# file's current content is one frame of a cycle, and the frame you happen to
# catch may be the one where the thing you are looking for was just consumed.
git log -S'## Unreleased' -- CHANGELOG.md # not `head -60 CHANGELOG.md`
```
A positive result needs no such scepticism — it carries its own evidence. Only
absence has to be *earned*, so spend the extra command there.
Absence has to be *earned*, so spend the extra command there.
### …and a positive result only proves what you *actually asked*
An earlier version of this section claimed "a positive result needs no such
scepticism — it carries its own evidence." **That is false, and believing it
cost a later session three more wrong findings.** A positive result is evidence
about the question your command really posed, which may not be the question you
meant. The failure is invisible precisely *because* the command succeeded.
| Claim | The command succeeded — at answering something else |
|---|---|
| "EGL git over SSH works" | `ssh git@gitea.egl.lan` greeted me as `joakimp`. `~/.ssh/config` had `Host gitea*` → `HostName gitea.jordbo.se`, so I authenticated **to the wrong instance**. The real EGL account is `ecsjper`. |
| "the port config regressed" | compared `ssh -G` output against `2222` — a value produced by **my own earlier `-p 2222` flag**, not by the config. I reported the user's edit as a regression it never caused. |
| "the CI runners authenticate with this token" | pure fabrication, contradicted by my own scan output already on screen. The runners use per-runner `REGISTRATION_TOKEN`. |
Two habits that actually catch this class, both cheap:
```sh
# 1. ask which RULE captured your hostname before trusting any ssh result.
# ssh_config is first-obtained-value-wins PER KEYWORD, not per block: a
# specific block only wins the keywords it declares, so a later `Host gitea*`
# still supplies HostName unless the specific block restates it.
ssh -G git@thehost | grep -E '^(hostname|port|user|identityfile)'
# 2. state the expected result BEFORE running the check, and diff against it.
# This is the single technique that separated the one verification that went
# right (10/10, expectations declared per probe) from five that went wrong
# (results interpreted after the fact, each time in the direction I expected).
probe "/repos/.../actions/runs" 200 # must work
probe "/admin/users" 403 # must be denied
```
And the meta-observation, which is the reason this subsection exists: across all
five errors, **not one was caught by re-reading my own reasoning.** Every one was
caught by a second measurement that disagreed — the SSH lie surfaced only because
the greeting said `joakimp` while a token probe minutes earlier had said
`ecsjper`; the fabrication surfaced only because the user read my own output back
to me. So the operational rule is not "be careful". It is: **for a load-bearing
claim, produce a second measurement by a different route, and expect it to
disagree.** If you cannot think of a second route, you do not yet have a finding
— you have a hypothesis.
**`dscp`/`scp` with accented filenames on a macOS host.** macOS stores filenames
in Unicode **NFD** (decomposed — e.g. `ä` is `a` + combining U+0308), while the
@@ -1,17 +1,17 @@
---
name: pi-extensions
description: >-
Use the pi extensions (pi-fork, pi-observational-memory, ssh-controlmaster) effectively in the pi coding agent harness. Load this skill only when running inside pi (detection - `fork` and `recall` are present in your tool list, or `pi --ssh` was used to start the session). pi-fork dispatches focused subtasks to forked agents at fast/balanced/deep effort tiers; pi-observational-memory compacts long sessions into recallable observations + reflections; ssh-controlmaster rewires pi's read/write/edit/bash tools to execute on a remote host over a multiplexed SSH connection. This skill covers tier selection, task design, boundary discipline, when to use recall, and remote-pi mechanics.
Use the pi extensions (pi-fork, pi-observational-memory, ssh-controlmaster) effectively in the pi coding agent harness. Load this skill only when running inside pi (detection - `fork` and `recall` are present in your tool list, or `pi --ssh` was used to start the session). pi-fork dispatches focused subtasks to forked agents at fast/balanced/deep effort tiers; pi-observational-memory compacts long sessions into recallable observations + reflections; ssh-controlmaster rewires pi's read/write/edit/bash tools to execute on a remote host over a multiplexed SSH connection. Also covers the context ladder L0-L4 and the `task` tool (pi-task: isolated child, immutable spec, machine-checked envelope, write-boundary diff) that is the DEFAULT for delegated work, with `fork` reserved for read-only exploration and parallel opinions, plus the `fork-gate` hook that enforces the split. This skill covers rung and tier selection, task design, boundary discipline, when to use recall, and remote-pi mechanics.
---
# Pi Extensions: pi-fork, pi-observational-memory, ssh-controlmaster
# Pi Extensions: pi-fork + task/fork-gate, pi-observational-memory, ssh-controlmaster
## When to Load This Skill
Load only when **both** of these are true:
1. You are running inside the **pi coding agent harness** (not Claude Code, not opencode, not any other harness).
2. The `fork` and/or `recall` tools appear in your available tool list, **or** the session was started with `pi --ssh ...`.
2. The `fork`, `task` and/or `recall` tools appear in your available tool list, **or** the session was started with `pi --ssh ...`.
If you do not see those tools, this skill does not apply — skip it. Other harnesses do not have these extensions and the patterns below will not work there.
@@ -71,7 +71,78 @@ ssh-controlmaster is orthogonal but composes cleanly: when pi is operating remot
---
## Part 1: pi-fork
## Part 1: delegating work — `task` and `fork`
### Decide the rung BEFORE the brief (read this first)
Two tools run a child agent. They differ in one thing, and it decides the
quality of what comes back: **what the child sees.**
| tool | child sees | rung | gives you | use for |
|---|---|---|---|---|
| **`task`** (pi-extensions `task.ts`, wraps `pi-task`) | **only your spec** — goal, named files, curated facts | L0–L2 | immutable spec, PASS/FAIL envelope, per-root boundary diff, audit dir, budgets | **any delegated work that changes files or must obey a rule** — the default |
| **`fork`** (pi-fork) | **your entire branch**, brief appended last | L4 | prose report, effort tiers, parallel dispatch from one message | read-only exploration that needs this conversation; N independent opinions |
**Pre-flight before any `fork(...)` — one *yes* makes it a `task`:**
1. Does the brief say *do not / only / never / must not*?
2. Will the child write, edit, commit or push anything?
3. Do I want a PASS/FAIL I can check, rather than prose?
Why the text alone did not work (and why this is now enforced): the rule above
lived in this skill and in the global AGENTS.md for months and was still
violated by agents that had just read it — five fork briefs in one session on
2026-09-17, all carrying "do not", one of which returned confident verbatim
quotes that did not exist. `fork` is a *tool*: its self-recommending
description ("implementation, testing, review…") is in the tool list every turn
and survives compaction; this skill is gone after the first compaction, and
`pi-task` was a CLI to be remembered. Two structural fixes shipped 2026-09-19:
- **`task` is a tool** (`extensions/task.ts`), so both rungs sit in the tool
list with the decision rule in their descriptions, and `promptGuidelines`
puts the rule in the system prompt where compaction cannot remove it. It
rejects, before spending a model run, the two spec errors that make a
violation certain (see "roots" below) and serialises overlapping tasks.
- **`fork-gate`** (`extensions/fork-gate.ts`) is a `tool_call` hook that BLOCKS
a fork whose brief contains a prohibition, a write boundary or a clause-initial
file-changing imperative, and returns the `task(...)` call to make instead.
It matches wording, not intent: a genuinely read-only exploration brief that
trips it is rephrased, and one that cannot be rephrased without its
prohibition needed `task` all along. `PI_FORK_GATE=off` logs instead of
blocking; `/ext` disables it entirely.
Minimal call:
```
task(id="slug", goal="…verbatim; the child has NO other context…",
deliverable="…exact shape wanted…", effort="fast|balanced|deep",
read_only=false,
roots=["/abs/repo/docs", "/abs/repo/src"], # WATCHED, each diffed alone
write_allowed=["/abs/repo/docs"], # exact subset of roots
facts=["verified fact"], files=["/abs/path/to/read"])
```
**Roots — the two errors the tool refuses up front.** Every root is diffed on
its own and a delta is allowed only if that *exact root string* is in
`write_allowed`. So (a) `write_allowed` must be a subset of `roots`, not a
subdirectory of one, and (b) a writable root must not lie inside a watched-only
root — the parent's porcelain would change and register a violation every time.
List the writable part as its own root and leave the enclosing repo out. This is
the shape the 2026-09-17 migration tasks used (sibling roots, `write_allowed`
naming four of them) and it passed cleanly.
**Overlap.** Sibling `task` calls whose roots overlap would see each other's
writes as violations; the tool runs them one after another automatically. Do
not rely on that for ordering *semantics* — if B needs A's output, call B after
A returns.
**What isolation does not fix.** L0 removes the *narrative* failures (parent
voice, invented continuity, ignored prohibitions). It does not remove
confabulation: an under-specified spec still gets a confident deliverable. The
report prints the evidence pointers under a "SPOT-CHECK THESE" heading for a
reason.
Everything below about tiers, brief design and boundary discipline applies to
**both** tools — a `task` spec is a brief too.
### Effort tier mapping
@@ -85,15 +156,15 @@ Configured in `~/.pi/agent/settings.json` under `pi-fork.effortProfiles`. The co
**Rule of thumb:** start at `balanced` unless you have a specific reason to go up or down. Going too cheap on a deep task wastes a fork; going too expensive on a mechanical task is just slow.
### When to fork vs. do it yourself
### When to delegate vs. do it yourself
Fork when **any** of:
Having chosen the rung above, delegate (either tool) when **any** of:
- The task requires reading many files whose contents you don't need to keep in your main context afterwards (the fork returns a dense summary; raw file contents stay in the fork's context and are discarded).
- You want to run multiple analyses in **parallel** (especially: comparing N options, where independent reasoning is itself a signal — see "parallel forks" below).
- The task is well-scoped enough to specify completely up front and well-bounded enough that returning a dense report is more useful than continuing the dialogue.
- You are about to do something that would burn a lot of tokens on tool calls (long file reads, many bash invocations) whose output you will mostly discard.
Don't fork when:
Don't delegate when:
- The work fits in your current context budget without crowding out what comes next.
- The task is exploratory and you'll need to iterate based on what you find (forking turns iteration into round-trips with full task-spec rewrites).
- You need to make decisions during the work that depend on context only the main thread has.
@@ -161,6 +232,69 @@ The "three" things it completed were exactly the main thread's pending todos, vi
- Distrust **quantities** and **provenance claims** in fork prose specifically ("all N sessions", "shipped with the image", "as expected") — those are the slots confabulation fills.
- The fact that the fork was "right anyway" is not the same as the fork having followed instructions.
### The context ladder — and the second dispatch mechanism (`pi-task`)
Everything above describes a child that inherits everything. That is not a fixed
cost of delegation — **how much context a child gets is a choice**, and `fork`
sits at one extreme of it. Five rungs:
| rung | what the child sees | mechanism | built? |
|---|---|---|---|
| **L0** | nothing but the goal | `pi-task` default: fresh `--session-id pitask-<id>-<stamp>` in a private `--session-dir` | yes |
| **L1** | goal + **names** of files/commands to read itself | `pi-task` spec `context.files` / `context.commands` (`bin/pi-task:154,157`) | yes |
| **L2** | goal + an **excerpt the parent curated** | `pi-task` spec `context.facts`, pasted verbatim (`bin/pi-task:151`) | yes |
| **L3** | a **truncated tail** of the parent branch | *nothing implements this* — would need a new spec key plus `--session <trimmed snapshot>` | **no** |
| **L4** | the **entire** parent branch | `fork(task=…)` — `getHeader()+getBranch()`, no offset or limit anywhere in the call chain | yes |
**`pi-task` is a CLI (`/opt/pi-toolkit/bin/pi-task`, `schema` prints the spec
fields); the `task` tool from pi-extensions wraps it** so it appears in your tool
list next to `fork`. If the tool is absent, invoke the CLI with `bash`:
`/opt/pi-toolkit/bin/pi-task run <spec.json>`. Either way it reads an immutable
JSON spec, and "inherit the session" is not expressible in that schema — the
isolation is structural, not a request.
**Choose the lowest rung that can do the job:**
- **`fork` (L4)** when the subtask only makes sense against this conversation,
when you want several independent opinions in parallel from one message, or for
read-only exploration whose detail you will discard. Everything in "Boundary
discipline" above applies in full.
- **`pi-task` (L0–L2)** when the brief contains a **prohibition** (the inherited
transcript is exactly what overrides those), when you want a **pass/fail**
result instead of prose, when you need an **audit trail**, or when writes
outside an authorised set must be caught.
- **Neither** for trivial work, iterative work (both are one-shot), or judgement
that needs context only you have.
**What `pi-task` gets you that no brief can.** The envelope must parse or the run
FAILED, however fluent the prose. `roots[]` is the WATCHED set and
`write_allowed` the CHANGEABLE subset, diffed before and after with git
`--porcelain --ignored`. That `--ignored` flag is load-bearing: in the T4 test the
child obeyed its brief perfectly and still tripped the detector, because
`py_compile` wrote `__pycache__` into a watched-but-not-writable root — a
gitignored path that plain `--porcelain` reports as clean. Note the structural
point that test exposed: under `read_only: true` a write is *defiance*, so a
well-behaved child never produces a delta and the detector is never exercised.
Splitting WATCHED from WRITABLE is what lets an **obedient** child reveal a
violation, which is the realistic hazard.
**What it does not fix.** `--no-extensions` removes extensions, not the core
`read`/`write`/`edit`/`bash` tools — exactly as described above — so the boundary
diff is post-hoc **detection, not prevention**. And a fresh L0 context removes the
*narrative* failures (parent voice, invented continuity) without removing
confabulation: given an under-specified spec built on a false premise, the child
still filled the `deliverable` slot with a confident shape. The envelope's own
structure creates that pressure. Verify decisive claims from the filesystem
regardless of which rung you used.
**Trap — the capability floor is inverted from intuition.** `runner.ts:188` reads
`if (extensions !== null) args.push("--no-extensions")`. So `pi-fork.extensions:
[]` passes the flag and the floor is **on**; setting it to `null` — documented in
`settings.json` as the way to "restore normal extension loading" — passes nothing
and the floor is **off**, restoring palace writes inside every fork child.
Changing `[]` to `null` as a tidy-up re-arms what was deliberately disarmed.
`pi-task` hardcodes the flag and cannot drift this way.
### Anti-patterns
- **Forking trivial work.** A fork has overhead. If the task takes < 30 seconds in your main thread, just do it.
@@ -230,7 +364,17 @@ When entries conflict, **the most recent observation reflects the latest known s
## Quick Reference
```
fork(task=..., effort=fast|balanced|deep)
task(id, goal, deliverable, effort, read_only, roots, write_allowed, facts, files, commands, wall_s, usd)
- L0-L2: isolated child sees ONLY the spec — DEFAULT for work that writes or has rules
- roots[] = WATCHED (each diffed alone); write_allowed[] = exact subset of roots,
never nested inside a watched-only root (the tool rejects both errors up front)
- envelope must parse or the run FAILED; spot-check evidence pointers
- overlapping-root tasks are serialised; audit: ~/.pi/agent/pi-task/<stamp>-<id>/result.json
- CLI fallback: bash /opt/pi-toolkit/bin/pi-task run <spec.json> (schema | selftest | run --dry-run)
fork(task=..., effort=fast|balanced|deep) # L4: child inherits your WHOLE branch
- ONLY for read-only exploration needing this conversation, or N parallel opinions
- fork-gate BLOCKS briefs with do-not/only/never, write boundaries, or "edit/commit/fix …"
- state decision authority explicitly
- pass verified context up front
- specify deliverable shape
@@ -248,8 +392,9 @@ recall(id=<12-char-hex>)
```
~/.pi/agent/settings.json
pi-fork.effortProfiles — model + thinking-depth per tier
pi-fork.effortProfiles — model + thinking-depth per tier (used by BOTH fork and task)
pi-fork.defaultEffort — usually "balanced"
env PI_FORK_GATE=off — fork-gate logs instead of blocking (default: block)
observational-memory.* — token thresholds, model, agentMaxTurns
observational-memory.debugLog: true — opt-in NDJSON telemetry at
~/.pi/agent/observational-memory/debug/<session>.ndjson (off by default)
+698
View File
@@ -0,0 +1,698 @@
#!/usr/bin/env bash
# check-doc-drift.sh — fail when a hand-maintained doc claim contradicts the
# build files it describes.
#
# THE DEFECT CLASS THIS EXISTS TO CATCH, measured 2026-09-10 while preparing
# v1.9.0. Five separate claims had rotted, all of them the same shape: a fact
# written once by hand, in a file nothing verifies, about a value that lives
# somewhere else and moved.
#
# 1..3. README.md's "Version pins" table was wrong on EVERY row — pi `0.84.4`
# vs ARG PI_VERSION=0.85.1, pi-atelier `v0.10.0` vs v0.10.1, mempalace
# `3.8.0` vs 3.9.0. That table is the worst possible place for this: it
# exists precisely to be the reviewable record of what is deliberately
# frozen, so when it lies, the review it enables is worthless.
# 4. README.md carried a "Planned for an upcoming minor release" section
# listing typst PDF export, which had ALREADY SHIPPED, tagged with a
# self-contradicting "(shipped in Unreleased/base)" marker. The
# CHANGELOG had already documented three earlier instances of exactly
# this stale-"Unreleased"-pointer class (see its v1.8.7 notes).
# 5. DOCKER_HUB.md claimed "Node.js v22" while this release ships Node 24.
# This one is the reason the gate exists at all: DOCKER_HUB.md is
# PUBLISHED. `update-description` in docker-publish.yml POSTs it to Hub
# as full_description on every tag, so unlike README.md — which no
# workflow or gate reads — a stale claim here is what users see.
#
# WHY A GATE AND NOT "REMEMBER TO CHECK". DOCKER_HUB.md had gone eight releases
# (v1.8.6 → v1.9.0) without a touch. Nothing generates it and nothing verifies
# it; the only mechanism keeping it true was whoever remembered. That is the
# same failure mode check-skill-floor.sh was written for, and the same fix:
# convert "someone remembers" into "CI refuses".
#
# TWO CLASSES OF CHECK, DELIBERATELY. Checks 1-7 compare a doc string to a
# value that EXISTS IN THIS REPO, so they can never be wrong about the world and
# need no network, no token, and no built image. Checks 8-9 compare against what
# is PUBLISHED (Docker Hub's measured sizes; the ref labels baked into the last
# released image), because those claims have no in-repo anchor at all and had
# rotted for exactly that reason. They need the network and therefore SKIP,
# loudly and counted, when it is absent -- a skip is neither OK nor a failure,
# because printing an unverified claim as OK is the habit this file exists to
# break, while failing on a third party's uptime would make every release
# hostage to it. Claims that need a RUNNING CONTAINER (the "N mempalace_* tools"
# count, uncompressed on-disk sizes) are still not gated here; assert them in
# scripts/smoke-test.sh where a real image is available.
#
# DELIBERATELY NOT GATED: Dockerfile.base's `# BASE_REBUILD_DATE:` comment, which
# is also stale (2026-07-13, three base rebuilds ago). base_tag is a hash of
# Dockerfile.base's CONTENT plus rootfs/, comments included, so a gate that
# demanded that comment be current would force a ~60 min base rebuild on any
# release that touched no base files at all. Fix it when you are already
# rebuilding the base — then it is free. This is a real cost asymmetry, not
# laziness.
#
# EXIT CODES (same contract as lint-shell.sh and check-skill-floor.sh):
# 0 every checked claim matches
# 1 at least one claim has drifted
# 2 cannot run (a file or ARG this gate reads is missing/unparseable)
# A gate that cannot run must not pass, so a missing input is 2, never 0.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
cd "$REPO_ROOT"
README="README.md"
HUB="DOCKER_HUB.md"
DF_VARIANT="Dockerfile.variant"
DF_BASE="Dockerfile.base"
# Docker Hub rejects a full_description longer than this. docker-publish.yml has
# no size check of its own; it only notices via a non-200 from the API, i.e.
# after paying the whole build. Catching it here makes it a 2-second failure.
HUB_MAX_CHARS=25000
WARN_ONLY=0
FAILURES=0
SKIPS=0
# Tolerance for the published size claims (check 8), as a percentage OF THE
# MEASURED SIZE. The denominator matters: against the claim instead, the same
# drift reads as a different number, and an early draft of this gate took 20%
# from the claim-relative figure and would therefore have MISSED its own
# motivating case. Both bounds are measured, not guessed:
# - the rot that motivated this check: claimed 1.1 GB vs measured 1.37 GB
# = 19.7% off, so the threshold must sit BELOW that or the gate is theatre.
# - the largest legitimate skew, i.e. a claim describing the currently-published
# release while the next tag changes the size: v1.9.1's 1.37 GB against
# v1.9.2's measured 1.23 GB = 11.4% off, so the threshold must sit ABOVE that
# or every size-changing release trips it.
# 15% sits in that 11.4%-19.7% window. Widen it only with a measured reason, and
# re-derive both bounds if you do.
SIZE_TOLERANCE_PCT="${SIZE_TOLERANCE_PCT:-15}"
usage() {
cat <<'EOF'
Usage: check-doc-drift.sh [--warn-only] [-h|--help]
Compares hand-written claims in README.md and DOCKER_HUB.md against the build
files they describe (Dockerfile.base, Dockerfile.variant).
--warn-only Report drift but exit 0 (advisory use, e.g. a local pre-push hook).
Environment:
SKIP_SIZE_CHECK=1 skip check 8 (published size claims vs Docker Hub)
SKIP_REF_CHECK=1 skip check 9 (refs moved since the last release are named)
SIZE_TOLERANCE_PCT check 8 tolerance, default 15 (see comment for its bounds)
Exit: 0 = in sync, 1 = drift, 2 = cannot run.
EOF
}
while [ $# -gt 0 ]; do
case "$1" in
--warn-only) WARN_ONLY=1; shift ;;
-h|--help) usage; exit 0 ;;
*) echo "::error::unknown argument: $1" >&2; usage >&2; exit 2 ;;
esac
done
for f in "$README" "$HUB" "$DF_VARIANT" "$DF_BASE"; do
if [ ! -f "$f" ]; then
echo "::error::$f not found (cwd $PWD). Cannot evaluate doc drift, so this is exit 2, not a pass."
exit 2
fi
done
# Read `ARG NAME=value` from a Dockerfile. Exit 2 when absent: if the ARG this
# gate is built around has been renamed, the gate is measuring nothing and must
# say so rather than silently comparing against an empty string.
read_arg() {
local file="$1" name="$2" value
value="$(sed -n "s/^ARG ${name}=\\(.*\\)\$/\\1/p" "$file" | head -1)"
if [ -z "$value" ]; then
echo "::error::ARG ${name} not found in ${file}. It was probably renamed;" >&2
echo "::error::update check-doc-drift.sh to match, because this gate is now blind." >&2
exit 2
fi
printf '%s' "$value"
}
# One row of README's "Version pins" table: `| pi | `0.85.1` | ... |`
read_pin_row() {
sed -n "s/^| $1 | \`\\([^\`]*\`*\\)\` |.*/\\1/p" "$README" | head -1
}
fail() {
FAILURES=$((FAILURES + 1))
echo "::error::$1"
}
ok() { printf ' OK %s\n' "$1"; }
# A check that could not be EVALUATED, as distinct from one that passed.
# Deliberately neither ok() nor fail(): printing it as OK would launder an
# unmeasured claim into a passing one (the exact habit this file exists to
# break), while failing on a third party's uptime would make every release
# hostage to Docker Hub's API. Loud, counted, and surfaced in the summary.
skip() { SKIPS=$((SKIPS + 1)); printf ' SKIP %s\n' "$1"; }
echo "Checking hand-maintained doc claims against the build files they describe."
echo
# ---------------------------------------------------------------------------
# 1-3. README's version-pin table vs the ARGs it names by name.
# ---------------------------------------------------------------------------
check_pin() {
local label="$1" documented="$2" actual="$3" where="$4"
if [ -z "$documented" ]; then
fail "README.md: no '| $label |' row found in the version-pin table. Either the
table was restructured (update this gate) or the row was dropped (restore it)."
return
fi
if [ "$documented" != "$actual" ]; then
fail "README.md version-pin table is stale for $label: says '$documented',
$where says '$actual'. Fix the table — it is the reviewable record of what
this repo deliberately freezes, so a wrong row defeats its only purpose."
return
fi
ok "README pin $label = $actual"
}
PI_ACTUAL="$(read_arg "$DF_VARIANT" PI_VERSION)"
ATELIER_ACTUAL="$(read_arg "$DF_VARIANT" PI_ATELIER_REF)"
MEMPALACE_ACTUAL="$(read_arg "$DF_BASE" MEMPALACE_VERSION)"
check_pin pi "$(read_pin_row pi)" "$PI_ACTUAL" "ARG PI_VERSION in $DF_VARIANT"
check_pin pi-atelier "$(read_pin_row pi-atelier)" "$ATELIER_ACTUAL" "ARG PI_ATELIER_REF in $DF_VARIANT"
check_pin mempalace "$(read_pin_row mempalace)" "$MEMPALACE_ACTUAL" "ARG MEMPALACE_VERSION in $DF_BASE"
# ---------------------------------------------------------------------------
# 4. DOCKER_HUB.md's Node claim vs ARG NODE_VERSION. This is the published page,
# so it is the one whose staleness reaches users.
# ---------------------------------------------------------------------------
NODE_ACTUAL="$(read_arg "$DF_BASE" NODE_VERSION)"
NODE_DOCUMENTED="$(sed -n 's/.*\*\*Node\.js\*\* v\([0-9][0-9]*\).*/\1/p' "$HUB" | head -1)"
if [ -z "$NODE_DOCUMENTED" ]; then
fail "$HUB: could not find a '**Node.js** vNN' claim. If the wording changed,
update this gate; do not leave the published page unverified."
elif [ "$NODE_DOCUMENTED" != "$NODE_ACTUAL" ]; then
fail "$HUB claims Node v$NODE_DOCUMENTED but ARG NODE_VERSION=$NODE_ACTUAL.
This file is PUBLISHED to Docker Hub by update-description on every tag,
and it is read from the TAG — so fix it before tagging, not after."
else
ok "$HUB Node claim = v$NODE_ACTUAL"
fi
# ---------------------------------------------------------------------------
# 5. Placeholders CI will not substitute. docker-publish.yml substitutes exactly
# {{PI_VERSION}} and then greps for leftovers of that ONE token, so any other
# {{...}} sails through the guard and is published literally.
# ---------------------------------------------------------------------------
UNKNOWN_PLACEHOLDERS="$(grep -o '{{[A-Za-z0-9_]*}}' "$HUB" | sort -u | grep -v '^{{PI_VERSION}}$' || true)"
if [ -n "$UNKNOWN_PLACEHOLDERS" ]; then
fail "$HUB contains placeholders CI does not substitute, which would be
published verbatim: $(echo "$UNKNOWN_PLACEHOLDERS" | tr '\n' ' ')
docker-publish.yml only fills {{PI_VERSION}}; add substitution there first."
else
ok "$HUB has no placeholders beyond {{PI_VERSION}}"
fi
# Match only the UPPER_SNAKE placeholder convention CI uses. A bare '{{' search
# is WRONG here, and the first version of this check proved it by failing on
# README.md:900 — `docker inspect --format '{{json .Config.Labels}}'`, a Go
# template in a legitimate example, not a placeholder. The gate was wrong, not
# the doc. Keep this anchored to [A-Z] so Go/Jinja/Handlebars examples pass.
README_PLACEHOLDERS="$(grep -o '{{[A-Z][A-Z0-9_]*}}' "$README" | sort -u || true)"
if [ -n "$README_PLACEHOLDERS" ]; then
fail "$README contains placeholder(s) nothing substitutes, so they would render
literally for every reader: $(echo "$README_PLACEHOLDERS" | tr '\n' ' ')
Only DOCKER_HUB.md gets substitution, and only for {{PI_VERSION}}."
else
ok "$README has no unsubstituted placeholders"
fi
# ---------------------------------------------------------------------------
# 6. Hub full_description length.
# ---------------------------------------------------------------------------
HUB_CHARS="$(wc -c < "$HUB" | tr -d ' ')"
if [ "$HUB_CHARS" -gt "$HUB_MAX_CHARS" ]; then
fail "$HUB is $HUB_CHARS chars, over Docker Hub's $HUB_MAX_CHARS-char
full_description limit. update-description would fail with a non-200 AFTER
the full build. Trim it — this file is the essentials-only page, and
README.md is the long form on purpose."
else
ok "$HUB is $HUB_CHARS chars (limit $HUB_MAX_CHARS)"
fi
# ---------------------------------------------------------------------------
# 7. Stale "Unreleased" pointers. "Unreleased" is a CHANGELOG-only concept; in
# a user-facing doc it is always a pointer that outlived what it pointed at.
# This class has now bitten five times, hence a gate rather than vigilance.
# ---------------------------------------------------------------------------
STALE_MARKERS="$(grep -n 'Unreleased' "$README" "$HUB" || true)"
if [ -n "$STALE_MARKERS" ]; then
fail "'Unreleased' appears in a user-facing doc, which is always a stale
pointer once the thing ships (it has happened five times here):
${STALE_MARKERS//$'\n'/$'\n' }
State the fact directly, or move it to CHANGELOG.md where 'Unreleased' means something."
else
ok "no stale 'Unreleased' pointers in $README or $HUB"
fi
# ---------------------------------------------------------------------------
# 8. Published size claims vs Docker Hub's MEASURED full_size.
#
# Why this exists: every other claim in these docs is checked against a file
# in this repo, so it cannot rot without someone editing the thing it
# describes. The size claims had no such anchor -- nothing in the repo states
# the image size -- so they quietly went 24% wrong across eight releases
# (DOCKER_HUB.md said ~1.1 GB; :latest measured 1.37 GB on 2026-09-14).
# DOCKER_HUB.md is POSTed to Docker Hub by update-description, so that number
# is the first thing a stranger reads about this image.
#
# Hub's `full_size` tracks the FIRST manifest entry (amd64 here), NOT the sum
# across architectures -- measured: v1.9.2 full_size=1.228 GB, amd64=1.228,
# arm64=1.211, sum=2.439. That matches the table's per-arch "Size
# (compressed)" column, which is why full_size is the right field.
#
# NOT COVERED, deliberately: README.md's ~3.2 GB figures are UNCOMPRESSED
# on-disk sizes, and the registry API exposes compressed sizes only (layer
# sizes in a manifest are compressed; the config blob carries no uncompressed
# totals). Measuring them needs a real pull, so they are out of scope here --
# do not read a green check 8 as covering them.
# ---------------------------------------------------------------------------
# Shared by checks 8 and 9: which Hub repo, and its tag list (one request).
# Derive the repo from the doc's own rows rather than hardcoding it, so a
# rename cannot leave these checks silently probing a repo nobody publishes to.
# shellcheck disable=SC2016 # single quotes are deliberate: this is a sed
# script, and its \( \) groups and \1 backreference must reach sed unexpanded.
HUB_REPO_PATH="$(sed -n 's/^| `\([^:`]*\):[^`]*`.*/\1/p' "$HUB" | head -1)"
HUB_TAGS_JSON=""
HAVE_NET_TOOLS=0
if command -v curl >/dev/null 2>&1 && command -v python3 >/dev/null 2>&1; then
HAVE_NET_TOOLS=1
if [ -n "$HUB_REPO_PATH" ] && \
{ [ "${SKIP_SIZE_CHECK:-0}" != "1" ] || [ "${SKIP_REF_CHECK:-0}" != "1" ]; }; then
HUB_TAGS_JSON="$(curl -sS -m 20 \
"https://hub.docker.com/v2/repositories/${HUB_REPO_PATH}/tags/?page_size=100" \
2>/dev/null || true)"
fi
fi
if [ "${SKIP_SIZE_CHECK:-0}" = "1" ]; then
skip "size claims -- SKIP_SIZE_CHECK=1 was set"
elif [ "$HAVE_NET_TOOLS" = 0 ]; then
skip "size claims -- need both curl and python3 to measure them"
else
if [ -z "$HUB_REPO_PATH" ]; then
skip "size claims -- found no \`repo:tag\` image rows in $HUB to check"
else
if [ -z "$HUB_TAGS_JSON" ]; then
skip "size claims -- Docker Hub API unreachable (offline?); NOT verified"
else
SIZE_RC=0
# NO `|| true` on the python invocation: an early draft had one, and it
# swallowed the exit code so a printed DRIFT line still exited 0 -- a gate
# that reports the defect and passes anyway. The outer `|| SIZE_RC=$?` is
# what keeps `set -e` happy while preserving the code.
SIZE_OUT="$(HUB_MD="$HUB" HUB_JSON="$HUB_TAGS_JSON" TOL="$SIZE_TOLERANCE_PCT" \
python3 <<'PYEOF'
import json, os, re, sys
try:
data = json.loads(os.environ["HUB_JSON"])
except (ValueError, KeyError) as exc:
print(" SKIP size claims -- Hub API returned unparseable JSON (%s)" % exc)
sys.exit(3)
# full_size == first manifest entry (amd64), which is the per-arch number the
# table's "Size (compressed)" column claims. Verified against .images[] sizes.
sizes = {
r["name"]: r["full_size"] / 1e9
for r in data.get("results", [])
if isinstance(r.get("full_size"), int) and r.get("name")
}
if not sizes:
print(" SKIP size claims -- Hub API returned no usable tags")
sys.exit(3)
tol = float(os.environ["TOL"])
row = re.compile(r"^\|\s*`([^`:]+):([^`]+)`\s*\|[^|]*\|\s*~?([0-9]+(?:\.[0-9]+)?)\s*GB\s*\|")
checked = drift = 0
with open(os.environ["HUB_MD"], encoding="utf-8") as fh:
for line in fh:
m = row.match(line)
if not m:
continue # rows saying "same", and every non-image row
_repo, tag, claimed = m.group(1), m.group(2), float(m.group(3))
if "X.Y.Z" in tag:
continue # placeholder row; the concrete tag is checked instead
# base-<hash> is content-addressed and immutable, so its size is
# base-latest's by construction -- probe the alias that always exists.
probe = "base-latest" if tag.startswith("base-") else tag
actual = sizes.get(probe)
if actual is None:
print(" SKIP size %s -- tag '%s' not present on Hub" % (tag, probe))
continue
checked += 1
off = abs(claimed - actual) / actual * 100
if off <= tol:
print(" OK size %s claims ~%.2f GB, Hub measures %.2f GB (%.0f%% off)"
% (tag, claimed, actual, off))
else:
drift += 1
print(" DRIFT size %s claims ~%.2f GB but Hub measures %.2f GB"
" (%.0f%% off, tolerance %.0f%%)" % (tag, claimed, actual, off, tol))
if checked == 0:
print(" SKIP size claims -- no checkable rows resolved to a published tag")
sys.exit(3)
sys.exit(1 if drift else 0)
PYEOF
)" || SIZE_RC=$?
printf '%s\n' "$SIZE_OUT"
case "$SIZE_RC" in
0) : ;;
3) SKIPS=$((SKIPS + 1)) ;;
*)
fail "a published size claim in $HUB has drifted from what Docker Hub
actually serves (see DRIFT above). This page is POSTed to Docker Hub by
update-description, so it is the first size a stranger sees. Re-measure and
update the table:
curl -sS 'https://hub.docker.com/v2/repositories/${HUB_REPO_PATH}/tags/?page_size=100' |
jq -r '.results[] | \"\\(.name) \\(.full_size/1e9)\"'"
;;
esac
fi
fi
fi
# ---------------------------------------------------------------------------
# 9. Everything the NEXT build would bake differently from the LAST PUBLISHED
# release must be named in the CHANGELOG text above that release's heading.
#
# Why this exists, measured 2026-09-19: pi-extensions 25c1265 (a new `task`
# tool and a hook that blocks certain `fork` calls -- a change to how every
# agent in the container delegates work) and mempalace-toolkit 817b3a8 (the
# feed's mine deadline had never reached the transport) both reached this
# image through floating `*_REF=main` ARGs. Neither produced a diff in this
# repo, so nothing here asked for a CHANGELOG entry, and neither had one
# until a reader asked. This is the same shape as check 8: a fact with no
# in-repo anchor rots. The hand practice that existed for it -- the
# "Dependency audit" table in each release's notes ("Baked in vN | Upstream
# now") -- is precisely a "someone remembers" mechanism, and it had lapsed.
#
# How it measures, with no docker/crane/token: the last published `vX.Y.Z`
# is the highest such tag in Hub's tag list (shared with check 8); its
# amd64 config blob is read through the anonymous registry API (token ->
# manifest index -> per-arch manifest -> config) and carries one
# `se.jordbo.pi-devbox.<name>-ref` label per component, each holding the
# SHA that build-args actually baked (resolve-versions in docker-publish.yml
# turns every ref into a SHA before `docker build`). "What the next build
# would bake" is resolved the way that job does it: a 40-hex ARG is itself,
# a tag or branch is `git ls-remote`d (peeled `^{}` first -- an annotated
# tag's un-dereferenced SHA is the tag object, a false alarm this repo has
# already fallen for once), pi-studio is the highest semver tag, and
# `PI_VERSION` / `MEMPALACE_VERSION` are compared as literals against the
# `pi-version` / `mempalace-version` labels (the latter set in Dockerfile.base
# and inherited; absent on releases before it shipped, which reports SKIP).
#
# The rule: baked == would-bake is OK with no mention required. If they
# differ, the text ABOVE the last published version's `## ` heading -- i.e.
# `## Unreleased` plus any not-yet-published `## vX.Y.Z` section, which is
# what the release commit turns Unreleased into -- must contain the
# would-bake value's 7-char SHA prefix (or, for pi-studio, the tag name; for
# pi, the version string). Naming the SHA, not just the repo, is the point:
# it is what the audit table always recorded, and it makes the failure
# message's compare URL a copy-paste away from knowing what moved.
#
# Every upstream commit therefore re-reds this gate until the CHANGELOG
# names the new head. That is the intended cost: the thing that gets baked
# is the thing that gets named, and a typo-fix upstream costs one edited
# SHA here. Read from the TAG like everything else in these docs -- the
# release commit renames Unreleased, so the pending text still covers it.
#
# SKIPs, each counted: SKIP_REF_CHECK=1; no curl/python3; Hub unreachable;
# the release's labels unreadable; one component's upstream unreachable
# (that component only). A published tag whose heading is MISSING from the
# CHANGELOG is a failure, not a skip: that is drift in its own right.
# ---------------------------------------------------------------------------
if [ "${SKIP_REF_CHECK:-0}" = "1" ]; then
skip "ref moves -- SKIP_REF_CHECK=1 was set"
elif [ "$HAVE_NET_TOOLS" = 0 ]; then
skip "ref moves -- need both curl and python3 to read the published labels"
elif ! command -v git >/dev/null 2>&1; then
skip "ref moves -- need git (ls-remote) to resolve what the next build would bake"
elif [ -z "$HUB_REPO_PATH" ]; then
skip "ref moves -- found no \`repo:tag\` image rows in $HUB to locate the published image"
elif [ -z "$HUB_TAGS_JSON" ]; then
skip "ref moves -- Docker Hub API unreachable (offline?); NOT verified"
else
# One plain top-level assignment per ARG, on purpose: read_arg exits 2 on a
# missing ARG, and under `set -e` that only propagates from a bare
# `VAR="$(...)"`. Nested inside a heredoc's $(...) the exit would be swallowed
# by `cat`, and a renamed ARG would leave this check comparing a label against
# an empty string and reporting the component "unchanged".
TOOLKIT_REPO="$(read_arg "$DF_VARIANT" PI_TOOLKIT_REPO)"; TOOLKIT_REF="$(read_arg "$DF_VARIANT" PI_TOOLKIT_REF)"
EXTENSIONS_REPO="$(read_arg "$DF_VARIANT" PI_EXTENSIONS_REPO)"; EXTENSIONS_REF="$(read_arg "$DF_VARIANT" PI_EXTENSIONS_REF)"
FORK_REPO="$(read_arg "$DF_VARIANT" PI_FORK_REPO)"; FORK_REF="$(read_arg "$DF_VARIANT" PI_FORK_REF)"
OBSMEM_REPO="$(read_arg "$DF_VARIANT" PI_OBSMEM_REPO)"; OBSMEM_REF="$(read_arg "$DF_VARIANT" PI_OBSMEM_REF)"
ATELIER_REPO="$(read_arg "$DF_VARIANT" PI_ATELIER_REPO)"
MPTK_REPO="$(read_arg "$DF_BASE" MEMPALACE_TOOLKIT_REPO)"; MPTK_REF="$(read_arg "$DF_BASE" MEMPALACE_TOOLKIT_REF)"
STUDIO_REPO="$(read_arg "$DF_VARIANT" PI_STUDIO_REPO)"
SKILLSET_SNAPSHOT="$(read_arg "$DF_VARIANT" SKILLSET_SNAPSHOT_REF)"
# name|kind|repo|ref -- one line per label the variant image carries.
# kinds: ref = branch/tag/SHA resolved like resolve-versions does;
# studio = highest semver tag of the repo (label lives on <tag>-studio);
# literal = the ARG value IS the baked value (a SHA pin, a version).
REF_COMPONENTS="pi-toolkit|ref|$TOOLKIT_REPO|$TOOLKIT_REF
pi-extensions|ref|$EXTENSIONS_REPO|$EXTENSIONS_REF
pi-fork|ref|$FORK_REPO|$FORK_REF
pi-obsmem|ref|$OBSMEM_REPO|$OBSMEM_REF
pi-atelier|ref|$ATELIER_REPO|$ATELIER_ACTUAL
mempalace-toolkit|ref|$MPTK_REPO|$MPTK_REF
pi-studio|studio|$STUDIO_REPO|
skillset-snapshot|literal||$SKILLSET_SNAPSHOT
pi-version|literal||$PI_ACTUAL
mempalace-version|literal||$MEMPALACE_ACTUAL"
REF_RC=0
# Same discipline as check 8: no `|| true` on the python, or a printed DRIFT
# exits 0. Per-component SKIP lines are counted afterwards by grep, so a run
# that evaluated eight components and could not reach the ninth reports one
# skip, not a green tick over the ninth.
REF_OUT="$(HUB_REPO="$HUB_REPO_PATH" HUB_JSON="$HUB_TAGS_JSON" CHANGELOG="CHANGELOG.md" \
COMPONENTS="$REF_COMPONENTS" python3 <<'PYEOF'
import json, os, re, subprocess, sys, urllib.request, urllib.parse
SHA40 = re.compile(r"^[0-9a-f]{40}$")
SEMVER = re.compile(r"^v?[0-9]+\.[0-9]+\.[0-9]+$")
LABEL = "se.jordbo.pi-devbox."
def ver_key(tag):
return tuple(int(x) for x in tag.lstrip("v").split("."))
def http_json(url, headers=None, timeout=30):
req = urllib.request.Request(url, headers=headers or {})
with urllib.request.urlopen(req, timeout=timeout) as resp:
return json.loads(resp.read().decode("utf-8"))
def labels_of(repo, tag):
"""Config labels of <repo>:<tag>'s amd64 image via the anonymous registry API."""
tok = http_json(
"https://auth.docker.io/token?service=registry.docker.io&scope="
+ urllib.parse.quote(f"repository:{repo}:pull", safe=":")
)["token"]
hdr = {
"Authorization": f"Bearer {tok}",
"Accept": ", ".join([
"application/vnd.oci.image.index.v1+json",
"application/vnd.docker.distribution.manifest.list.v2+json",
"application/vnd.oci.image.manifest.v1+json",
"application/vnd.docker.distribution.manifest.v2+json",
]),
}
base = f"https://registry-1.docker.io/v2/{repo}"
man = http_json(f"{base}/manifests/{tag}", hdr)
if "manifests" in man: # multi-arch index: pick linux/amd64, as check 8 does
cands = [m for m in man["manifests"]
if m.get("platform", {}).get("architecture") == "amd64"
and m.get("platform", {}).get("os") == "linux"]
if not cands:
raise RuntimeError("no linux/amd64 entry in the manifest index")
man = http_json(f"{base}/manifests/{cands[0]['digest']}", hdr)
cfg = http_json(f"{base}/blobs/{man['config']['digest']}", hdr)
return cfg.get("config", {}).get("Labels") or {}
def ls_remote(repo, *patterns):
# GIT_TERMINAL_PROMPT=0: a repo flipped private must fail fast as a SKIP,
# not sit waiting for a username on a CI runner until the job times out.
env = dict(os.environ, GIT_TERMINAL_PROMPT="0")
out = subprocess.run(["git", "ls-remote", repo, *patterns], env=env,
capture_output=True, text=True, timeout=60, check=True).stdout
return {line.split("\t")[1]: line.split("\t")[0] for line in out.splitlines() if "\t" in line}
def resolve_ref(repo, ref):
"""What docker-publish.yml's resolve-versions would pass as the build-arg."""
if SHA40.match(ref):
return ref, ref
refs = ls_remote(repo, f"refs/heads/{ref}", f"refs/tags/{ref}", f"refs/tags/{ref}^{{}}")
for key in (f"refs/tags/{ref}^{{}}", f"refs/heads/{ref}", f"refs/tags/{ref}"):
if key in refs:
return refs[key], ref
raise RuntimeError(f"'{ref}' is neither a branch nor a tag of {repo}")
def resolve_studio(repo):
refs = ls_remote(repo, "refs/tags/*")
tags = {k[len("refs/tags/"):]: v for k, v in refs.items()}
names = sorted((t for t in tags if SEMVER.match(t)), key=ver_key)
if not names:
raise RuntimeError(f"no semver tag at {repo}")
tag = names[-1]
return tags.get(tag + "^{}", tags[tag]), tag
def compare_url(repo, a, b):
root = repo[:-4] if repo.endswith(".git") else repo
return f"{root}/compare/{a}...{b}"
try:
hub = json.loads(os.environ["HUB_JSON"])
except (ValueError, KeyError) as exc:
print(" SKIP ref moves -- Hub API returned unparseable JSON (%s)" % exc)
sys.exit(3)
released = sorted((r["name"] for r in hub.get("results", [])
if isinstance(r.get("name"), str) and re.fullmatch(r"v[0-9]+\.[0-9]+\.[0-9]+", r["name"])),
key=ver_key)
if not released:
print(" SKIP ref moves -- Hub lists no published vX.Y.Z tag to compare against")
sys.exit(3)
last = released[-1]
repo = os.environ["HUB_REPO"]
# The text every not-yet-published change lives in: everything above the last
# published version's heading. Its absence is drift, not a skip.
text = open(os.environ["CHANGELOG"], encoding="utf-8").read()
# (\s|$) rather than \b: a word boundary would accept "## v1.9.2-rc1" or
# "## v1.9.2-typo" as v1.9.2's heading. Caught by the sabotage test, not review.
m = re.search(r"^## v?%s(\s|$)" % re.escape(last.lstrip("v")), text, re.M)
if not m:
print(" DRIFT ref moves -- %s is the last PUBLISHED tag on Hub but %s has no '## %s' heading"
% (last, os.environ["CHANGELOG"], last))
sys.exit(1)
pending = text[:m.start()].lower()
try:
labels = labels_of(repo, last)
except Exception as exc: # network, auth, shape -- all "could not measure"
print(" SKIP ref moves -- could not read %s:%s's labels from the registry (%s); NOT verified"
% (repo, last, exc))
sys.exit(3)
studio_labels = None
checked = drift = 0
problems = []
for line in os.environ["COMPONENTS"].splitlines():
if not line.strip():
continue
name, kind, url, ref = line.split("|", 3)
# <name>-ref labels hold SHAs; names that already end in -version are the
# label (pi-version, mempalace-version) -- a version string, compared literally.
key = LABEL + name if name.endswith("-version") else LABEL + name + "-ref"
try:
if kind == "studio":
if studio_labels is None:
studio_labels = labels_of(repo, last + "-studio")
baked = studio_labels.get(key)
else:
baked = labels.get(key)
except Exception as exc:
print(" SKIP %-18s -- could not read %s:%s-studio's labels (%s)" % (name, repo, last, exc))
continue
if not baked:
print(" SKIP %-18s -- %s carries no %s label" % (name, last, key))
continue
try:
if kind == "ref":
now, shown = resolve_ref(url, ref)
elif kind == "studio":
now, shown = resolve_studio(url)
else:
now, shown = ref, ref
except Exception as exc:
print(" SKIP %-18s -- could not resolve what the next build would bake (%s)" % (name, exc))
continue
checked += 1
is_sha = bool(SHA40.match(now))
short = (lambda s: s[:7] if SHA40.match(s) else s)
if baked == now:
print(" OK %-18s unchanged since %s (%s)" % (name, last, short(now)))
continue
names = [now[:7].lower()] if is_sha else [now.lower()]
if kind == "studio":
names.append(shown.lower())
if any(n in pending for n in names):
print(" OK %-18s %s -> %s since %s, named above the %s heading"
% (name, short(baked), short(now), last, last))
continue
drift += 1
hint = compare_url(url, baked, now) if (url and is_sha and SHA40.match(baked)) else ""
problems.append(" %-18s %s -> %s%s" % (name, short(baked), short(now), (" " + hint) if hint else ""))
print(" DRIFT %-18s %s -> %s since %s, NOT named above the %s heading"
% (name, short(baked), short(now), last, last))
if problems:
print(" Name each new value (7-char SHA prefix, or the tag/version) in CHANGELOG.md above '## %s':" % last)
print("\n".join(problems))
if checked == 0 and drift == 0:
print(" SKIP ref moves -- no component could be evaluated")
sys.exit(3)
sys.exit(1 if drift else 0)
PYEOF
)" || REF_RC=$?
printf '%s\n' "$REF_OUT"
REF_SKIPS="$(printf '%s\n' "$REF_OUT" | grep -c '^ SKIP ' || true)"
case "$REF_RC" in
0) SKIPS=$((SKIPS + REF_SKIPS)) ;;
3) SKIPS=$((SKIPS + 1)) ;;
*)
SKIPS=$((SKIPS + REF_SKIPS))
fail "a component the next build would bake differently from the last published
release is not named in CHANGELOG.md (see DRIFT above). These reach the image
through floating refs, so nothing else in this repo records that they moved;
the CHANGELOG entry is the only place a reader of the next tag can learn it.
Name the new SHA (7 chars is enough) where you describe the change -- the
compare URL above shows what moved."
;;
esac
fi
echo
if [ "$FAILURES" -eq 0 ]; then
if [ "$SKIPS" -gt 0 ]; then
echo "OK: every checked doc claim matches the build files" \
"($SKIPS check(s) SKIPPED and therefore NOT verified -- see SKIP above)."
else
echo "OK: every checked doc claim matches the build files."
fi
exit 0
fi
echo "::error::$FAILURES doc claim(s) have drifted from the build files."
echo
echo "Docs are read from the TAG, not from main: docker-publish.yml checks out"
echo "github.ref, so a fix pushed after tagging does not reach the release or the"
echo "Hub page. Update the docs BEFORE you tag."
if [ "$WARN_ONLY" -eq 1 ]; then
echo "(--warn-only: exiting 0 anyway)"
exit 0
fi
exit 1
+166
View File
@@ -0,0 +1,166 @@
#!/usr/bin/env bash
# check-skill-floor.sh — fail when the vendored pi-extensions skill snapshot in
# rootfs/ ("the floor") has drifted from the package repo it is a snapshot of.
#
# THE DEFECT THIS EXISTS TO CATCH, measured 2026-09-10.
# rootfs/usr/local/share/pi-devbox/skills/pi-extensions/ ships a vendored copy
# of the pi-extensions skill so the skill is ALWAYS present in the image.
# Dockerfile.variant then copies the freshly-cloned package copy OVER the served
# path at /usr/local/share/... — but it never writes back to the repo floor. So
# the floor only silently rots, and it had: 34284 B, untouched since fa04d20
# (2026-07-30), while the package copy was 38973 B. Four copies existed with
# three different sizes.
#
# Why that is worse than ordinary staleness: the floor is a FALLBACK. The copy
# step is guarded by `if [ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build
# where the package clone yields no skill/ keeps the vendored snapshot and still
# succeeds — green, with no manifest flag and no label saying which copy was
# served. The image would ship a July skill and nothing would say so. Keeping
# the floor fresh means that fallback is harmless instead of a silent regression.
#
# WHY A DIRECTORY HASH AND NOT `sha256sum SKILL.md`.
# The same pipeline Dockerfile.variant uses for skillset_snapshot_tree_sha256,
# and for the same documented reason: a file-only compare answers "did this one
# file change", not "is this the same skill". pi-extensions ships TWO files
# (SKILL.md + evaluate-extension-usage.py), so a sibling-file edit would pass a
# file-only check. If you change the pipeline here, change it there too.
#
# WHY GATING ON ANOTHER REPO IS PROPORTIONATE HERE, since that is normally a
# smell: this fires only when the package's skill/ DIRECTORY HASH changes, which
# is exactly and only when the floor has genuinely gone stale. pi-extensions
# commits that do not touch skill/ leave the hash alone and cannot turn this red.
# The repo is also anonymously clonable (verified 2026-09-10 with `git ls-remote`
# and no credentials), so this needs no secret and cannot break on token expiry.
#
# Exit codes — deliberately three, matching scripts/lint-shell.sh's philosophy
# that a gate which cannot run must not pass:
# 0 in sync (or the package legitimately has no skill/ at this ref)
# 1 DRIFT — the floor differs from the package
# 2 cannot run — no package copy could be obtained
set -euo pipefail
REPO_ROOT=$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)
FLOOR_DIR="${REPO_ROOT}/rootfs/usr/local/share/pi-devbox/skills/pi-extensions"
# Defaults mirror Dockerfile.variant's ARGs so this checks what the build builds.
PI_EXTENSIONS_REPO="${PI_EXTENSIONS_REPO:-https://gitea.jordbo.se/joakimp/pi-extensions.git}"
PI_EXTENSIONS_REF="${PI_EXTENSIONS_REF:-main}"
PACKAGE_DIR=""
WARN_ONLY=0
TMPDIR_CLONE=""
usage() {
cat <<'EOF'
Usage: scripts/check-skill-floor.sh [options]
--package-dir DIR Compare against an existing skill directory instead of
cloning. In a devbox container use /opt/pi-extensions/skill
for a fully offline run.
--warn-only Report drift but exit 0 (advisory use, e.g. a local hook).
-h, --help This text.
Environment: PI_EXTENSIONS_REPO, PI_EXTENSIONS_REF (default main) — both mirror
the Dockerfile.variant ARGs of the same name.
EOF
}
while [ $# -gt 0 ]; do
case "$1" in
--package-dir) PACKAGE_DIR="${2:-}"; shift 2 ;;
--warn-only) WARN_ONLY=1; shift ;;
-h|--help) usage; exit 0 ;;
*) echo "::error::unknown argument: $1" >&2; usage >&2; exit 2 ;;
esac
done
cleanup() {
if [ -n "$TMPDIR_CLONE" ]; then rm -rf "$TMPDIR_CLONE"; fi
}
trap cleanup EXIT
# Identical to Dockerfile.variant's tree_sha256(): relative paths + per-file
# sha256 over a sorted `find`, folded into one digest. Deterministic, never
# readdir order.
tree_sha256() {
( cd "$1" && find . -type f -print | LC_ALL=C sort | xargs -r sha256sum ) \
2>/dev/null | sha256sum | cut -d' ' -f1
}
if [ ! -d "$FLOOR_DIR" ]; then
echo "::error::floor directory is missing: ${FLOOR_DIR}"
echo "::error::rootfs/ is supposed to guarantee the skill is always in the image."
exit 2
fi
SOURCE_DESC=""
if [ -n "$PACKAGE_DIR" ]; then
if [ ! -d "$PACKAGE_DIR" ]; then
echo "::error::--package-dir does not exist: ${PACKAGE_DIR}"
exit 2
fi
SOURCE_DESC="local directory ${PACKAGE_DIR}"
else
command -v git >/dev/null 2>&1 || { echo "::error::git not found; cannot obtain the package copy."; exit 2; }
TMPDIR_CLONE=$(mktemp -d)
# Fetch the single ref shallowly. `git fetch <ref>` accepts a branch, a tag
# and (on Gitea) a reachable commit, which is why this is not `clone --branch`
# — CI resolves PI_EXTENSIONS_REF to a 40-hex SHA before the build.
if ! ( cd "$TMPDIR_CLONE" \
&& git init -q . \
&& git remote add origin "$PI_EXTENSIONS_REPO" \
&& git fetch -q --depth 1 origin "$PI_EXTENSIONS_REF" \
&& git checkout -q FETCH_HEAD ) 2>/dev/null; then
echo "::error::could not fetch ${PI_EXTENSIONS_REF} from ${PI_EXTENSIONS_REPO}"
echo "::error::Cannot determine whether the floor is stale, so this is exit 2, not a pass."
echo "::error::For an offline run, pass --package-dir /opt/pi-extensions/skill"
exit 2
fi
PACKAGE_SHA=$( cd "$TMPDIR_CLONE" && git rev-parse --short HEAD )
PACKAGE_DIR="${TMPDIR_CLONE}/skill"
SOURCE_DESC="${PI_EXTENSIONS_REPO} @ ${PI_EXTENSIONS_REF} (${PACKAGE_SHA})"
fi
# A ref with no skill/ is the documented fallback case: Dockerfile.variant keeps
# the vendored snapshot and the build succeeds. Nothing to compare, so this is
# not drift — but it IS the exact condition under which the floor ships, so say
# so loudly rather than printing a silent green tick.
if [ ! -d "$PACKAGE_DIR" ]; then
echo "::warning::package has no skill/ at this ref — the vendored floor is what will ship."
echo " source : ${SOURCE_DESC}"
echo " floor : $(tree_sha256 "$FLOOR_DIR")"
exit 0
fi
FLOOR_HASH=$(tree_sha256 "$FLOOR_DIR")
PKG_HASH=$(tree_sha256 "$PACKAGE_DIR")
if [ "$FLOOR_HASH" = "$PKG_HASH" ]; then
echo "OK: vendored pi-extensions floor matches the package."
echo " source : ${SOURCE_DESC}"
echo " tree_sha256: ${FLOOR_HASH}"
exit 0
fi
# `set -e` interacts badly with `[ … ] && x` as a bare statement, so both of
# these are explicit if-blocks rather than AND-lists.
LEVEL="error"
if [ "$WARN_ONLY" -eq 1 ]; then LEVEL="warning"; fi
echo "::${LEVEL}::vendored pi-extensions skill floor has DRIFTED from the package."
echo " source : ${SOURCE_DESC}"
echo " floor tree_sha256 : ${FLOOR_HASH}"
echo " pkg tree_sha256 : ${PKG_HASH}"
echo ""
echo " per-file differences:"
diff -rq "$FLOOR_DIR" "$PACKAGE_DIR" 2>&1 | sed 's/^/ /' || true
echo ""
echo " Remedy — re-sync the floor and commit it:"
echo " cp -a <pi-extensions>/skill/. ${FLOOR_DIR}/"
echo " git add ${FLOOR_DIR#"${REPO_ROOT}/"} && git commit"
echo ""
echo " NOTE this forces one full base rebuild: base_tag hashes Dockerfile.base"
echo " + rootfs/, and that rebuild is what re-bakes the refreshed floor."
if [ "$WARN_ONLY" -eq 1 ]; then exit 0; fi
exit 1
+97
View File
@@ -0,0 +1,97 @@
#!/usr/bin/env bash
# Shellcheck + syntax-check every shell script in this repo. Severity: error.
#
# SINGLE SOURCE OF TRUTH for two callers:
# .gitea/workflows/lint.yml — advisory, every branch push and PR
# .gitea/workflows/docker-publish.yml — the release GATE (lint-gate job)
# Extracted from lint.yml on 2026-09-08 rather than copied, because a second
# copy is exactly the drift this repo has been bitten by (see skillset's
# pi-extensions mirror, refreshed the same evening after sitting 9579 B behind).
#
# WHY THIS CHECK EXISTS AT ALL
# actionlint shellchecks workflow `run:` steps only. The repo's own scripts —
# entrypoint.sh, scripts/*.sh, and the extensionless tools under
# rootfs/usr/local/bin/ — were never shellchecked. A sibling repo with the same
# gap shipped a broken `echo "$json" | python3 <<'EOF' ... json.load(sys.stdin)`
# for two months: with no script argument python reads its SCRIPT from stdin,
# so the heredoc IS stdin and json.load hits EOF. shellcheck flags that at
# severity error (SC2259); nothing ever ran it.
#
# WHY THE RELEASE GATES ON IT (added 2026-09-08, the expensive way round)
# v1.8.14's first attempt failed after build-base had already spent ~46 min:
# scripts/smoke-test.sh had an apostrophe inside a single-quoted exec_test body
# ("the fleet\'s"), which CLOSES the string, so the body truncated and its tail
# ran on the CI runner instead of inside the image. shellcheck had already
# caught it as SC2289 at severity error — the lint job went red on the very
# push that introduced it and stayed red for 24 hours, unread. lint.yml
# deliberately does not run on tag pushes (sound: the tagged tree was linted on
# main, and a tag-ref lint run sorts above the publish run and makes a release
# look finished early). The gap was never "lint the tag" — it was that a tree
# whose lint FAILED could still be released. Hence a gate inside the publish
# workflow, ~40 s, ahead of everything expensive.
#
# SEVERITY CHOICE
# -S error is 0 findings across this repo when clean, so it is free to add.
# -S warning is NOT free here (20x SC2088 tilde-in-quotes in
# recreate-sanity-check.sh, plus assorted SC2016 — both intentional), and a
# noisy gate trains people to ignore it. Error-only, matching the
# SHELLCHECK_OPTS philosophy in lint.yml.
# Reproduce the count before editing it (the `$ ` prefix is load-bearing: a
# comment whose first word is "shellcheck" is parsed as a DIRECTIVE, and a
# malformed one is SC1072/SC1073 at severity error — this gate caught exactly
# that when the line was first written without it):
# $ shellcheck -S warning -f gcc scripts/*.sh rootfs/usr/local/bin/* \
# entrypoint*.sh hooks/* | grep -c SC2088
#
# Usage: bash scripts/lint-shell.sh [root] (default root: repo top level)
set -uo pipefail
root="${1:-}"
if [ -z "$root" ]; then
root="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
fi
cd "$root" || { echo "::error::cannot cd to $root"; exit 2; }
# A gate that cannot run must not pass. Without this, a machine (or a CI job
# whose install step was reordered away) without shellcheck would sail through
# printing nothing, which is the failure mode this whole file exists to prevent.
if ! command -v shellcheck >/dev/null 2>&1; then
echo "::error::shellcheck not found — the gate cannot run, so it must not pass" >&2
echo " install it (apt-get install -y shellcheck) or run this in CI" >&2
exit 2
fi
# Union of two signals, because either alone misses a real case: a shebang scan
# misses a sourced fragment with no shebang, and a *.sh glob misses the
# extensionless tools in rootfs/usr/local/bin/. Silent skipping is precisely the
# failure mode this gate exists to prevent, so err toward over-collecting.
# -print0/mapfile -d '' so a path containing a space cannot silently split.
mapfile -d '' -t all_files < <(find . -not -path './.git/*' -type f -print0)
sh_files=()
for f in "${all_files[@]}"; do
case "$f" in *.sh) sh_files+=("$f"); continue;; esac
if head -n1 "$f" 2>/dev/null | grep -qE '^#!.*\b(bash|sh)\b'; then
sh_files+=("$f")
fi
done
echo "Checking ${#sh_files[@]} shell file(s) with $(shellcheck --version | awk '/version:/{print $2}')"
# A green tick over an empty file set is not a check.
if [ "${#sh_files[@]}" -eq 0 ]; then
echo "::error::no shell files found — the shebang scan or the checkout is wrong"
exit 1
fi
rc=0
shellcheck -S error -f gcc "${sh_files[@]}" || rc=1
# bash -n catches a different class than shellcheck (unbalanced constructs it
# declines to parse), so both run and both count.
for f in "${sh_files[@]}"; do
bash -n "$f" || { echo "::error file=$f::bash -n failed"; rc=1; }
done
if [ "$rc" -eq 0 ]; then
echo "OK: ${#sh_files[@]} shell file(s) clean at severity error"
fi
exit "$rc"
+258 -17
View File
@@ -2,8 +2,8 @@
# Runtime post-recreate verification for pi-devbox.
#
# Verifies that after `docker compose up -d --force-recreate`:
# - The new image is actually live (pi version matches, when an expected
# version is supplied — see the version note below)
# - The new image is actually live (both the pi version and — when asked —
# the pi-devbox image release tag; see the two version notes below)
# - Persisted named volumes survived (~/.pi config, shell history, zoxide,
# nvim data, uv cache, ssh-local)
# - pi runtime wiring is intact: keybindings symlink, AGENTS.md symlink,
@@ -11,7 +11,8 @@
# pi-observational-memory / (studio variant) pi-studio package
# registrations in settings.json packages[]
# - Shell defaults re-seeded from /etc/skel-devbox
# - /tmp/sshcm exists with mode 700 (ssh ControlMaster dir)
# - ssh ControlMaster works: /tmp/sshcm exists 700 AND the ControlPath that
# ssh actually resolves (ssh -G) is a writable directory
# - /opt toolkits intact
# - Known expected-absences don't regress
#
@@ -25,13 +26,33 @@
# the pi-devbox repo (which a maintainer already has for CI builds). A plain
# `docker pull` consumer is not the audience and will not have this file.
#
# Version note: pi's version is resolved from `latest` at CI build time and is
# NOT pinned to a concrete value in Dockerfile.variant (ARG PI_VERSION=latest).
# So unlike opencode-devbox, this script cannot self-derive an expected version
# from the Dockerfile. Pass --expected-version to assert a match; without it the
# live pi version is reported as an informational WARN, not a failure.
# TWO DIFFERENT VERSIONS, TWO DIFFERENT FLAGS. This distinction has already
# cost a release day, so it is spelled out here and in AGENTS.md step 4:
#
# Usage: ./scripts/recreate-sanity-check.sh [--expected-version X.Y.Z] [--variant studio|plain]
# --expected-version the PI CODING AGENT version, e.g. 0.84.3
# (`pi --version`; pinned as ARG PI_VERSION in
# Dockerfile.variant, which CI reads as the source
# of truth)
# --expected-image-version the PI-DEVBOX IMAGE release tag, e.g. 1.8.9 or
# v1.8.9 (the `release_tag` baked into
# /etc/pi-devbox/build-manifest.json)
#
# Passing a release tag to --expected-version used to report
# "pi version mismatch: expected 1.8.8, got 0.84.3" — an accusation aimed at
# the wrong component, on the last gate of a release. Both flags now detect
# being handed the other one's value and say so instead.
#
# Neither flag is required. Both values are derivable from the image's own
# build manifest, so by default the script asserts the LIVE pi version against
# the version recorded at build time — which is not a tautology: a stale
# `pi` in the ~/.pi/npm-global volume can shadow the baked one, exactly the
# way a stale npm:pi-atelier can (see the packages[] check below). Pass the
# flags when you want an assertion against a value you name yourself, which
# is what a release checklist wants.
#
# Usage: ./scripts/recreate-sanity-check.sh [--expected-version X.Y.Z]
# [--expected-image-version X.Y.Z]
# [--variant studio|plain]
#
# Exit codes:
# 0 all checks passed
@@ -41,22 +62,61 @@
set -euo pipefail
EXPECTED_VERSION=""
EXPECTED_IMAGE_VERSION=""
VARIANT=""
REPO_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
MANIFEST=/etc/pi-devbox/build-manifest.json
# Parse arguments
usage() {
cat >&2 <<'EOF'
usage: recreate-sanity-check.sh [--expected-version X.Y.Z]
[--expected-image-version X.Y.Z]
[--variant studio|plain]
--expected-version pi coding agent version, e.g. 0.84.3 (`pi --version`)
--expected-image-version pi-devbox image release tag, e.g. 1.8.9 or v1.8.9
--variant studio|plain (auto-detected when omitted)
These are two different versions. Both are read from the image's own build
manifest when the corresponding flag is omitted.
EOF
}
# Parse arguments. Every flag takes a value, so reject a missing one rather
# than swallowing the next flag as if it were the value.
need_value() {
case "${2:-}" in
""|-*)
echo "$1 requires a value" >&2
usage
exit 2
;;
esac
}
while [[ $# -gt 0 ]]; do
case "$1" in
--expected-version)
need_value "$@"
EXPECTED_VERSION="$2"
shift 2
;;
--expected-image-version)
need_value "$@"
EXPECTED_IMAGE_VERSION="$2"
shift 2
;;
--variant)
need_value "$@"
VARIANT="$2"
shift 2
;;
--help|-h)
usage
exit 0
;;
*)
echo "usage: $0 [--expected-version X.Y.Z] [--variant studio|plain]" >&2
echo "unknown option: $1" >&2
usage
exit 2
;;
esac
@@ -67,6 +127,19 @@ pass() { echo " ✓ $1"; }
fail() { echo " ✗ $1" >&2; FAILED=$((FAILED + 1)); }
warn() { echo " ⚠ $1" >&2; }
# Read one top-level field from the build manifest, or print nothing. The
# manifest is the image's own ground truth (written at `docker build` time by
# Dockerfile.variant), so it needs no checkout and no network. Absent on an
# image built before it existed, hence every caller treats "" as unknown.
manifest_field() {
[ -f "$MANIFEST" ] || return 0
command -v jq >/dev/null 2>&1 || return 0
jq -r --arg k "$1" '.[$k] // empty' "$MANIFEST" 2>/dev/null || true
}
# Release tags are written with a leading v in the manifest and quoted without
# one in checklists; compare on the bare number so both spellings work.
strip_v() { printf '%s' "${1#v}"; }
# Auto-detect variant if not provided. The studio variant vendors pi-studio to
# /opt/pi-studio; the plain variant does not.
if [ -z "$VARIANT" ]; then
@@ -86,21 +159,59 @@ else
fi
echo
echo "-- pi version --"
MANIFEST_PI_VERSION=$(manifest_field pi_version)
MANIFEST_RELEASE_TAG=$(manifest_field release_tag)
echo "-- pi (coding agent) version --"
if ACTUAL_VERSION=$(pi --version 2>&1 | head -1); then
if [ -n "$EXPECTED_VERSION" ]; then
if [ "$ACTUAL_VERSION" = "$EXPECTED_VERSION" ]; then
pass "pi version $ACTUAL_VERSION"
if [ "$(strip_v "$EXPECTED_VERSION")" = "$(strip_v "$ACTUAL_VERSION")" ]; then
pass "pi version $ACTUAL_VERSION (matches --expected-version)"
elif [ -n "$MANIFEST_RELEASE_TAG" ] &&
[ "$(strip_v "$EXPECTED_VERSION")" = "$(strip_v "$MANIFEST_RELEASE_TAG")" ]; then
# Exact, not heuristic: the value handed over IS this image's release
# tag, so it cannot be a pi version anyone meant.
fail "--expected-version $EXPECTED_VERSION is the pi-devbox IMAGE version, not the pi version — use --expected-image-version $EXPECTED_VERSION (live pi is $ACTUAL_VERSION)"
else
fail "pi version mismatch: expected $EXPECTED_VERSION, got $ACTUAL_VERSION"
fail "pi version mismatch: expected $EXPECTED_VERSION, got $ACTUAL_VERSION (this flag asserts the pi coding agent version; for the image release tag use --expected-image-version)"
fi
elif [ -n "$MANIFEST_PI_VERSION" ]; then
# Not a tautology: the manifest records what pi reported at BUILD time,
# while `pi --version` resolves through PATH, which a stale npm-global
# volume install can shadow.
if [ "$MANIFEST_PI_VERSION" = "$ACTUAL_VERSION" ]; then
pass "pi version $ACTUAL_VERSION (matches this image's build manifest)"
else
fail "live pi $ACTUAL_VERSION != $MANIFEST_PI_VERSION recorded in $MANIFEST — a stale pi in the ~/.pi/npm-global volume is shadowing the baked one"
fi
else
warn "pi version $ACTUAL_VERSION (no --expected-version given; pi is built from 'latest', cannot self-derive — informational only)"
warn "pi version $ACTUAL_VERSION (no --expected-version and no build manifest to compare against — informational only)"
fi
else
fail "pi --version failed"
fi
echo
echo "-- pi-devbox image version --"
if [ -z "$MANIFEST_RELEASE_TAG" ]; then
if [ -n "$EXPECTED_IMAGE_VERSION" ]; then
fail "cannot verify --expected-image-version $EXPECTED_IMAGE_VERSION: no readable release_tag in $MANIFEST (image built before the manifest existed, or jq missing)"
else
warn "image release tag unknown (no readable $MANIFEST) — pi-devbox-version would say the same"
fi
elif [ -n "$EXPECTED_IMAGE_VERSION" ]; then
if [ "$(strip_v "$EXPECTED_IMAGE_VERSION")" = "$(strip_v "$MANIFEST_RELEASE_TAG")" ]; then
pass "image version $MANIFEST_RELEASE_TAG (matches --expected-image-version)"
elif [ -n "$MANIFEST_PI_VERSION" ] &&
[ "$(strip_v "$EXPECTED_IMAGE_VERSION")" = "$MANIFEST_PI_VERSION" ]; then
fail "--expected-image-version $EXPECTED_IMAGE_VERSION is the pi version, not the image release tag — use --expected-version $EXPECTED_IMAGE_VERSION (this image is $MANIFEST_RELEASE_TAG)"
else
fail "image version mismatch: expected $EXPECTED_IMAGE_VERSION, got $MANIFEST_RELEASE_TAG — the recreate did not pick up the intended image"
fi
else
warn "image version $MANIFEST_RELEASE_TAG (no --expected-image-version given — informational only)"
fi
echo
echo "-- Persisted named volumes (must survive --force-recreate) --"
@@ -264,6 +375,34 @@ if [ -f "$HOME/.pi/agent/settings.json" ]; then
fi
fi
# ── agent-browser must resolve to the image, not the config volume ────
# The same volume-shadowing hazard already asserted for pi (above) and
# pi-atelier (just now), for the third package it has bitten. This check
# belongs HERE rather than only in smoke-test.sh: a build-time container has an
# empty ~/.pi/npm-global, so smoke-test can never see the stale copy that a
# real recreate inherits. Measured instance: 0.27.0 from 2026-07-17 shadowed
# the image's 0.35.2 for ~7 weeks on mbp-m1-2020, silently supplying an older
# BUNDLED SKILL (3 skillsets vs 8) — the agent read the stale instructions
# without any version mismatch ever being surfaced.
AB_PATH=$(command -v agent-browser 2>/dev/null || true)
if [ -z "$AB_PATH" ]; then
warn "agent-browser not on PATH (expected in v1.6.0+ images; skipping shadow check)"
else
AB_REAL=$(readlink -f "$AB_PATH" 2>/dev/null || echo "$AB_PATH")
AB_VER=$(agent-browser --version 2>/dev/null | head -n1)
case "$AB_REAL" in
/usr/*)
pass "agent-browser resolves to the image copy (${AB_VER:-version unknown})"
;;
*)
fail "agent-browser resolves to $AB_REAL (${AB_VER:-version unknown}) — a ~/.pi/npm-global VOLUME copy is shadowing the image; the entrypoint retirement guard did not run or could not move it"
;;
esac
if [ -d "$HOME/.pi/npm-global/lib/node_modules/agent-browser" ]; then
fail "stale agent-browser still present in the ~/.pi/npm-global volume (entrypoint guard did not retire it)"
fi
fi
# ── pi <-> pi-atelier compatibility floor ─────────────────────────────
# atelier < 0.7.1 wraps pi's private TUI renderer in a way that recurses under
# pi >= 0.84: pi hangs at startup burning CPU, with no error message. atelier's
@@ -287,13 +426,115 @@ if [ -d /opt/pi-atelier ] && command -v jq >/dev/null 2>&1; then
fi
echo
echo "-- ssh ControlMaster dir --"
echo "-- ssh ControlMaster: socket dir + EFFECTIVE ControlPath --"
# TWO LAYERS, and the second is the one that has actually broken in the field.
#
# LAYER 1 (original check): /tmp/sshcm, the directory entrypoint-user.sh creates
# for the base image's system drop-in
# (/etc/ssh/ssh_config.d/00-devbox-controlmaster.conf).
#
# LAYER 2 (added 2026-09-15): the directory a config NAMES — which is not the
# same question, and asserting layer 1 is structurally blind to it. On
# emb-7kj4vr4g a durable ~/.pi/ssh/config pointed ControlPath at /tmp/ssh-cm
# (with a hyphen), a directory nothing in the image creates. EVERY ssh died
# unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory
# rc=255 with the remote command never running — while this script printed a
# green tick for layer 1, truthfully, about the wrong object.
#
# The same rc=255 has a second, independent cause already documented in prose in
# Dockerfile.base ("SSH client defaults" CAVEAT) and never verified anywhere: a
# per-host `ControlPath ~/.ssh/cm/%r@%h:%p` inherited from a bind-mounted
# READ-ONLY ~/.ssh. Measured to be the identical failure class:
# unix_listener: cannot bind to path ~/.ssh/cm/...: Read-only file system
# So do not guess which config wins — ask ssh. `ssh -G` applies real config
# precedence (first-obtained-value-wins, system drop-in, Include, -F override)
# and prints the fully expanded ControlPath. Require its parent to exist and be
# writable. Cost measured at 0.116 s for 48 hosts; -G never opens a connection.
if [ -d /tmp/sshcm ] && [ "$(stat -c %a /tmp/sshcm 2>/dev/null)" = "700" ]; then
pass "/tmp/sshcm exists with mode 700"
else
fail "/tmp/sshcm missing or not mode 700"
fi
# Probe one route. $1 = label, $2 = config to force with -F ("" = ssh's own
# default precedence), $3 = severity when a ControlPath dir is unusable.
#
# SEVERITY SPLIT IS DELIBERATE. The default route legitimately resolves into the
# read-only ~/.ssh on any host whose own config pins ControlPath there, and the
# supported workaround (`ssh -F ~/.ssh-local/config`) already exists — so that
# is a warn, not a fail. Failing it would paint this script red on every run of
# every device, and a check that fires benignly every time is one you learn to
# ignore. The sidecar route is the PRESCRIBED one, so there it is a hard fail.
_ssh_cm_probe() {
local label="$1" cfg="${2:-}" sev="${3:-fail}"
local h out cm cp dir n=0 shown
local bad=()
while IFS= read -r h; do
[ -n "$h" ] || continue
if [ -n "$cfg" ]; then
out=$(ssh -F "$cfg" -G "$h" 2>/dev/null) || continue
else
out=$(ssh -G "$h" 2>/dev/null) || continue
fi
cm=$(printf '%s\n' "$out" | awk '/^controlmaster /{print $2; exit}')
case "$cm" in '' | no | none | false) continue ;; esac
cp=$(printf '%s\n' "$out" | awk '/^controlpath /{print $2; exit}')
case "$cp" in '' | none) continue ;; esac
n=$((n + 1))
dir=$(dirname "$cp")
if [ ! -d "$dir" ] || [ ! -w "$dir" ]; then
bad+=("$h")
fi
done <<< "$SSH_CM_HOSTS"
# Cap the host list. A 50-name line is the noisy gate this repo already warns
# about in lint-shell.sh: unreadable output is ignored output. Six names plus
# a count is enough to identify the class and act.
if [ "${#bad[@]}" -gt 0 ]; then
shown="${bad[*]:0:6}"
if [ "${#bad[@]}" -gt 6 ]; then
shown="$shown (+$(( ${#bad[@]} - 6 )) more)"
fi
fi
if [ "$n" -eq 0 ]; then
warn "$label: no host resolves to ControlMaster on — effective ControlPath not exercised"
elif [ "${#bad[@]}" -eq 0 ]; then
pass "$label: $n ControlMaster host(s), every ControlPath dir exists and is writable"
elif [ "$sev" = "warn" ]; then
warn "$label: ${#bad[@]}/$n host(s) resolve ControlPath to a missing or unwritable dir [$shown] — expected when ~/.ssh/config pins ControlPath inside the read-only ~/.ssh; use 'ssh -F ~/.ssh-local/config' (see Dockerfile.base CAVEAT)"
else
fail "$label: ${#bad[@]}/$n host(s) resolve ControlPath to a missing or unwritable dir [$shown] — ssh dies rc=255 'unix_listener: cannot bind to path' and the remote command never runs"
fi
}
if command -v ssh >/dev/null 2>&1; then
# Concrete Host aliases only: patterns (*, ?) and negations (!) are not
# connectable targets, so `ssh -G` on them proves nothing.
_cm_cfgs=()
if [ -r "$HOME/.ssh/config" ]; then _cm_cfgs+=("$HOME/.ssh/config"); fi
if [ -r "$HOME/.ssh-local/config" ]; then _cm_cfgs+=("$HOME/.ssh-local/config"); fi
if [ "${#_cm_cfgs[@]}" -gt 0 ]; then
SSH_CM_HOSTS=$(awk 'tolower($1)=="host"{for(i=2;i<=NF;i++) if ($i !~ /[*?!]/) print $i}' \
"${_cm_cfgs[@]}" 2>/dev/null | sort -u)
else
SSH_CM_HOSTS=""
fi
if [ -z "$SSH_CM_HOSTS" ]; then
warn "no concrete Host aliases in ~/.ssh/config or ~/.ssh-local/config — effective ControlPath not verified"
else
_ssh_cm_probe "default ssh precedence" "" warn
if [ -r "$HOME/.ssh-local/config" ]; then
_ssh_cm_probe "ssh -F ~/.ssh-local/config" "$HOME/.ssh-local/config" fail
else
warn "~/.ssh-local/config absent — setup-lan-access.sh did not run; the prescribed multiplex route is unverified"
fi
fi
else
warn "ssh not on PATH — effective ControlPath not verified"
fi
echo
echo "-- Shell defaults re-seeded from /etc/skel-devbox --"
if [ -f "$HOME/.bash_aliases" ]; then
+273 -8
View File
@@ -5,6 +5,7 @@
#
# Verifies:
# - pi binary present and (if EXPECTED_PI_VERSION set) matches CI's resolved version
# - node MAJOR matches Dockerfile.base's ARG NODE_VERSION (if EXPECTED_NODE_MAJOR set)
# - mempalace core matches the audited pin (if EXPECTED_MEMPALACE_VERSION set)
# - new v1.0.0 base additions (pandoc, graphviz, imagemagick, yq, tealdeer)
# - typst PDF engine for pandoc (v1.4.0) — `pandoc --pdf-engine=typst`
@@ -27,6 +28,9 @@
# (human, --json, --quiet)
# - (studio variant only, auto-detected) pi-studio cloned + prebuilt
# client bundle present + registered via `pi install`
# - no foreign npm-11 platform packages (@esbuild, clipboard) beyond the host
# - no build-time npm cache (/root/.npm) shipped in the image
# - esbuild compiles + clipboard native loads at every install site
# - image size within threshold
set -euo pipefail
@@ -91,8 +95,31 @@ if [ -n "${EXPECTED_PI_VERSION:-}" ]; then
else
run "pi" "pi --version"
fi
run "node" "node --version"
# Until 2026-09-07 this was a bare `run "node" "node --version"`, which asserts
# only that the binary exists and exits 0 — the printed version was never
# compared to anything. A node major bump would therefore have passed this suite
# SILENTLY, while a reader skimming it would reasonably assume node regressions
# were covered. EXPECTED_NODE_MAJOR closes that: CI derives it from
# Dockerfile.base's ARG NODE_VERSION (the single source of truth), so this also
# catches a stale cached layer whose node does not match the declared ARG.
if [ -n "${EXPECTED_NODE_MAJOR:-}" ]; then
run_expect "node major matches Dockerfile ARG" "node --version" "v${EXPECTED_NODE_MAJOR}."
else
run "node" "node --version"
fi
run "git" "git --version"
# NOTE: the shellcheck binary is a GATE DEPENDENCY, not a convenience.
# scripts/lint-shell.sh is the release gate (the lint-gate job resolve-versions
# depends on) and it exits 2 when the binary is missing, by design — "a gate that
# cannot run must not pass". Measured on v1.8.14: it was absent from the image, so
# that gate could not be run by a developer in ANY container, only in CI.
# Asserted here so its absence fails a build instead of being discovered by a hook
# that then refuses every push (hooks/pre-push).
#
# This comment must not BEGIN with the tool's name: a line starting with
# `# shellcheck` is parsed as a DIRECTIVE, not a comment (SC1073/SC1072). The
# gate added in this same change caught that here, before the push.
run "shellcheck (lint gate dependency)" "shellcheck --version | grep -qE '^version: [0-9]'"
run "aws" "aws --version"
run "uv" "uv --version"
run "nvim" "nvim --version"
@@ -245,6 +272,8 @@ run "socat" "socat -V"
run "studio-expose helper" "test -x /usr/local/bin/studio-expose"
run "image-baked pi-devbox-environment skill" \
"test -f /usr/local/share/pi-devbox/skills/pi-devbox-environment/SKILL.md"
run "image-baked credential-incident-response skill" \
"test -f /usr/local/share/pi-devbox/skills/credential-incident-response/SKILL.md"
run "global-AGENTS append snippet present" \
"test -f /usr/local/share/pi-devbox/pi-global-AGENTS.append.md"
run "pi-devbox block merged into pi-global-AGENTS.md" \
@@ -289,8 +318,24 @@ run "pi-toolkit clone" "test -d /opt/pi-toolkit && git -C /opt/pi-toolkit rev
run "pi-extensions clone" "test -d /opt/pi-extensions && git -C /opt/pi-extensions rev-parse --short HEAD"
run "pi-fork clone + node_modules" \
"test -f /opt/pi-fork/package.json && test -d /opt/pi-fork/node_modules"
run "pi-observational-memory clone + node_modules" \
"test -f /opt/pi-observational-memory/package.json && test -d /opt/pi-observational-memory/node_modules"
# om is checked differently from pi-fork ON PURPOSE. It declares ZERO runtime
# dependencies: 8 devDependencies (omitted by --omit=dev) and 4 peerDependencies,
# which pi itself provides. npm 10 still materialised a node_modules for it, but
# that directory held exactly ONE file (.package-lock.json, 4 KB) and no nested
# package.json at all — 20 empty scope dirs. npm 11 stopped creating it, so the
# old `test -d node_modules` assertion went red on v1.9.0 while nothing about om
# had changed or broken. It was asserting an npm artefact, not a property of the
# shipped software. What actually has to hold is that the entry point pi loads
# exists, so assert THAT, straight out of the manifest pi reads
# (package.json -> pi.extensions), rather than a hardcoded path that could drift.
run "pi-observational-memory clone + declared pi entry point" \
"test -f /opt/pi-observational-memory/package.json && \
node -e 'const p=require(\"/opt/pi-observational-memory/package.json\"),f=require(\"fs\"),h=require(\"path\"); \
const l=(p.pi&&p.pi.extensions)||[]; \
if(!l.length){console.error(\"package.json declares no pi.extensions\");process.exit(1)} \
for(const e of l){const t=h.resolve(\"/opt/pi-observational-memory\",e); \
if(!f.existsSync(t)){console.error(\"declared entry missing: \"+t);process.exit(1)}} \
console.log(\"entries ok: \"+l.join(\",\"))'"
# ...and that the clone carries the AUTH FIX, not merely that it exists. om's
# pre-flight hasUsableAuth() check silently disabled `recall` for ~8 weeks once
# pi moved to request-time SigV4 signing and stopped exposing a static Bedrock
@@ -500,6 +545,50 @@ run "manifest skill fingerprint matches the baked snapshot" '
echo "manifest=[$m] actual=[$a]" >&2
[ -n "$m" ] && [ "$m" = "$a" ]
'
# ── Which pi-extensions skill copy shipped ──────────────────────────────
# Closes the silent-fallback hole. The refresh in Dockerfile.variant is guarded
# by `[ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates
# the co-located skill keeps the vendored floor and still succeeds GREEN, with
# nothing recording that a snapshot shipped instead of the package copy. Measured
# 2026-09-10: the floor had been stale since 2026-07-30, so that path would have
# shipped a six-week-old skill in silence. The floor is fresh now and gated by the
# skill-floor lint job, but "the fallback is currently harmless" is a fact with a
# shelf life, whereas "the image says which copy it got" keeps working.
#
# vendored-floor FAILS here rather than merely warning: these images track main,
# where the package has co-located skill/ since fa04d20, so a fallback means the
# clone did not resolve as intended and that is a defect to investigate. A fork
# deliberately pointing at a mirror without skill/ is the one case that should
# edit this assertion — which is the honest place for that decision to surface.
run "manifest names which pi-extensions skill copy shipped" '
j=/etc/pi-devbox/build-manifest.json
s=$(jq -r ".pi_extensions_skill_source // empty" $j)
h=$(jq -r ".pi_extensions_skill_tree_sha256 // empty" $j)
echo "source=[$s] tree_sha256=[$h]" >&2
printf "%s" "$h" | grep -qxE "[0-9a-f]{64}" || {
echo "pi_extensions_skill_tree_sha256 is not a 64-hex digest" >&2; exit 1; }
case "$s" in
package) ;;
vendored-floor)
echo "FALLBACK: clone had no skill/ at this ref, so the image ships the committed floor" >&2; exit 1 ;;
divergent)
echo "MIXED: served directory is part package and part floor" >&2; exit 1 ;;
*)
echo "pi_extensions_skill_source absent or unrecognised" >&2; exit 1 ;;
esac
'
# Same shape as the mempalace fingerprint check above, and for the same reason: a
# recorded hash that is never recomputed is a claim, not a measurement.
run "recorded pi-extensions skill hash matches the served bytes" '
j=/etc/pi-devbox/build-manifest.json
d=/usr/local/share/pi-devbox/skills/pi-extensions
m=$(jq -r ".pi_extensions_skill_tree_sha256 // empty" $j)
a=$( (cd "$d" && find . -type f -print | LC_ALL=C sort | xargs -r sha256sum) | sha256sum | cut -d" " -f1)
echo "manifest=[$m] actual=[$a]" >&2
[ -n "$m" ] && [ "$m" = "$a" ]
'
# OCI labels live in the image config, not the container fs — inspect them
# from the host docker rather than via `docker run`.
LBL=$(docker inspect --format '{{ index .Config.Labels "se.jordbo.pi-devbox.pi-extensions-ref" }}' "$IMAGE" 2>/dev/null || true)
@@ -508,6 +597,19 @@ if [ -n "$LBL" ] && [ "$LBL" != "<no value>" ]; then
else
printf " ❌ OCI label se.jordbo.pi-devbox.pi-extensions-ref missing or empty\n"; FAIL=$((FAIL+1))
fi
# mempalace-version is set in Dockerfile.base and INHERITED by the variant, so
# it states the pin of the base this image actually built on. It must equal the
# installed binary: the one way they diverge is a base built with
# INSTALL_MEMPALACE=false (label says 3.x, nothing installed) or an install that
# resolved to something other than the pin — both invisible to a label-only
# check. Same ground-truth rule as the manifest assertion above.
MP_LBL=$(docker inspect --format '{{ index .Config.Labels "se.jordbo.pi-devbox.mempalace-version" }}' "$IMAGE" 2>/dev/null || true)
MP_BIN=$(docker run --rm --entrypoint= "$IMAGE" sh -c 'mempalace --version 2>/dev/null | head -n1 | tr -d "\r"' 2>/dev/null || true); MP_BIN=${MP_BIN##* }
if [ -n "$MP_LBL" ] && [ "$MP_LBL" != "<no value>" ] && [ "$MP_LBL" = "$MP_BIN" ]; then
printf " ✅ OCI label se.jordbo.pi-devbox.mempalace-version=%s equals the installed core\n" "$MP_LBL"; PASS=$((PASS+1))
else
printf " ❌ OCI label se.jordbo.pi-devbox.mempalace-version=[%s] vs installed mempalace=[%s]\n" "$MP_LBL" "$MP_BIN"; FAIL=$((FAIL+1))
fi
# ── Runtime deployment (needs entrypoint to run) ──────────────────────
echo ""
@@ -594,11 +696,40 @@ exec_test "mempalace skill linked (fallback)" 'test -L $HOME/.agents/skills
# This assertion is kept because it is orthogonal and free: it pins content,
# not provenance, so it still catches a re-vendored snapshot whose ref was
# bumped correctly but whose bytes came from the wrong place.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Provenance is stamped for you" "$f" && ! grep -q "Attribute what you file yourself" "$f" && echo ok'
#
# v1.8.13: RE-PINNED on refresh a12fe5e -> e9e09d9, which is the whole point of
# the mechanism — the previous pair ("Provenance is stamped for you" present /
# "Attribute what you file yourself" absent) still passed against the NEW
# snapshot, so leaving it would have produced a canary that is green on both the
# old and the new bytes, i.e. blind to precisely the refresh it exists to
# witness. Same false-green family as the pre-v1.8.5 canary this comment warns
# about. The replacement pair was chosen by MEASURING direction against both
# files rather than by reading the diff: "Diaries self-heal; plain drawers do
# not" is new=1/old=0, "Agent diaries live in" is new=0/old=1 — so each string
# discriminates on its own and the pair still fails loudly in BOTH directions
# (forgotten bump AND re-vendored stale snapshot). Upstream content behind this
# refresh: the bare project-name wing convention and the <harness>@<device>
# added_by rule.
#
# Unreleased: RE-PINNED again on refresh e9e09d9 -> e9e45f7. The retired pair was
# still green against the new snapshot (the diaries section was untouched), so it
# was blind to this refresh for the same reason the v1.8.13 pair was blind to
# that one. The replacement pair is unusually strong because BOTH witnesses come
# out of the same upstream commit: skillset e9e45f7 ADDED the "Withdrawing an ask
# you sent" bullet and DELETED the sentence "there is nothing anyone can do about
# it from the other end" that the new bullet contradicts. Directions were
# MEASURED against both files, not read off the diff: "Withdrawing an ask you
# sent" is new=1/old=0, "nothing anyone can do about it from the other end" is
# new=0/old=1. A canary whose negative witness was removed by the very commit it
# pins fails loudly on the OLD bytes instead of merely failing to notice them,
# which is the property every previous pair here lacked. Upstream content:
# requester-side ask withdrawal became DEPLOYED behaviour once v1.9.1 baked
# mempalace-toolkit e68ee20 (>= e2b060a) through the floating MEMPALACE_TOOLKIT_REF.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Withdrawing an ask you sent" "$f" && ! grep -q "nothing anyone can do about it from the other end" "$f" && echo ok'
# Link TARGETS, not just link existence: with no skillset mounted (as here) the
# baked tree must be what resolves, for all three vendored skills.
# baked tree must be what resolves, for all four vendored skills.
exec_test "vendored skills resolve to the baked tree (no skillset mounted)" \
'for s in mempalace pi-extensions pi-devbox-environment; do
'for s in mempalace pi-extensions pi-devbox-environment credential-incident-response; do
case "$(readlink -f $HOME/.agents/skills/$s)" in
/usr/local/share/pi-devbox/skills/$s) ;;
*) echo "$s resolves to $(readlink -f $HOME/.agents/skills/$s)" >&2; exit 1 ;;
@@ -612,10 +743,18 @@ exec_test "vendored skills resolve to the baked tree (no skillset mounted)" \
exec_test "pi-devbox-version reports skill sources (all baked, no skillset here)" \
'out=$(pi-devbox-version)
echo "$out" | grep -q "skills:" || { echo "no skills section" >&2; exit 1; }
for s in mempalace pi-extensions pi-devbox-environment; do
echo "$out" | grep -qE "^ $s +baked$" \
for s in mempalace pi-extensions pi-devbox-environment credential-incident-response; do
echo "$out" | grep -qE "^ $s +baked( \([^)]*\))?$" \
|| { echo "$s not reported as baked" >&2; exit 1; }
done; echo ok'
# The optional " (...)" above is what pi-extensions now appends to say WHICH copy
# shipped — "baked (package copy)", or a loud FALLBACK/MIXED annotation. Without
# allowing it, adding that annotation turned this assertion red on v1.9.0 even
# though the state it reported was the correct one. The suffix is deliberately
# matched loosely rather than pinned to "(package copy)", because WHICH copy
# shipped is already asserted authoritatively above, against the manifest field
# and its measured tree hash, and duplicating that here in a regex would just
# create a second place to update whenever the wording changes.
# The boot banner must NOT carry the section: entrypoint-user.sh prints the
# version FIRST, before the baked links exist and long before the skillset
# deploy + reconcile run last, so anything it said about skill sources would be
@@ -716,10 +855,110 @@ exec_test "pi-atelier registered in packages[] (TUI sidebar)" \
exec_test "pi-atelier registered from /opt, not npm: (volume-shadowing guard)" \
'jq -e "((.packages // []) | any((type == \"string\") and endswith(\"/pi-atelier\"))) and (((.packages // []) | any(. == \"npm:pi-atelier\")) | not)" $HOME/.pi/agent/settings.json'
# agent-browser: the third package hit by ~/.pi/npm-global volume shadowing
# (after pi itself and pi-atelier). This build-time check is deliberately WEAK
# and says so: a `docker run` container has an EMPTY config volume, so it can
# only prove the image ships a sane copy and nothing in the image itself
# shadows it. The check that actually bites lives in
# recreate-sanity-check.sh, which runs where the volume is real — that is
# where a 7-week-old 0.27.0 was caught shadowing 0.35.2 on 2026-09-06.
# EXECUTION is ASSERTED here, not printed. Until 2026-09-07 the version was
# captured inside an echo with 2>/dev/null, so a binary that could not run at all
# still PASSED and simply printed version=[] -- the same failure class as the bare
# `node --version` two hundred lines up: a value displayed rather than compared.
#
# Why this exit code matters more than most: smoke runs `platforms: linux/amd64`
# on an x86 runner, i.e. NATIVE amd64, so this is the fleet's only recurring
# amd64 runtime proof for the linux-x64 ELF. No devbox can supply one -- every
# machine in the pi fleet is an Apple Silicon Mac (mbp-m1-2020; tor-ms22 = Mac
# Studio Mac13,1 M1 Max, verified 2026-08-17 by system_profiler; emb-7kj4vr4g =
# Apple Silicon, 4 routes 2026-09-07). Asking a device for that proof is asking
# for the impossible; CI already had it and was discarding it.
#
# KEEP PROSE OUT OF THE QUOTED BODY BELOW. On 2026-09-07 this explanation lived
# INSIDE the single-quoted argument and contained an apostrophe ("the fleet's").
# Inside '...' bash treats a backslash literally, so \' does not escape -- it
# CLOSES the string. The body silently truncated, the remaining lines were parsed
# by the RUNNER's shell instead of the container's, and `agent-browser --version`
# ran on a host that has no agent-browser: "line 770: command not found", release
# v1.8.14's smoke job failed after the base had already built. shellcheck caught
# it as SC2289 the same day and the red lint job went unread for 24h.
exec_test "agent-browser resolves under /usr (volume-shadowing guard, build-time half)" '
p=$(command -v agent-browser) || { echo "agent-browser not on PATH" >&2; exit 1; }
r=$(readlink -f "$p")
v=$(agent-browser --version) || { echo "agent-browser did not EXECUTE" >&2; exit 1; }
test -n "$v" || { echo "agent-browser --version produced no output" >&2; exit 1; }
echo "resolved=[$r] version=[$(printf %s "$v" | head -n1)]" >&2
case "$r" in /usr/*) ;; *) exit 1 ;; esac
test ! -d "$HOME/.pi/npm-global/lib/node_modules/agent-browser" || exit 1
echo ok
'
# pi-fork capability floor. `extensions: []` makes a fork child run with
# --no-extensions, which is the only MECHANICAL guarantee that a fork cannot
# file drawers or diary entries under the parent's identity — the mempalace
# bridge is an extension, so removing extensions removes the write path.
# Asserted because it is a security-shaped default that a settings merge or a
# hand-edit could silently drop, and its absence is invisible until a fork
# writes to the shared palace as you (measured twice: 2026-09-01, 2026-09-06).
# Deliberately compares to [] and not "is falsy": null means "load normal
# extensions", i.e. exactly the unguarded state this asserts against.
exec_test "pi-fork extensions floor is [] (forks cannot write to the palace)" \
'jq -e ".[\"pi-fork\"].extensions == []" $HOME/.pi/agent/settings.json'
# ── /tmp/sshcm directory created by entrypoint ────────────────────────
exec_test "/tmp/sshcm dir mode 700 (ssh ControlMaster)" \
'test -d /tmp/sshcm && [ "$(stat -c %a /tmp/sshcm)" = "700" ] && echo ok'
# ── Build-time leftovers (npm 11 bloat sentinels) ─────────────────────
# Both of these are worth a PASS/FAIL assertion rather than a size-gate
# diagnostic, because the size gate has ~225 MB of deliberate margin: v1.9.1
# shipped +131 MB of pure build residue and stayed green. These name the
# residue directly, so a regression is legible instead of merely "bigger".
echo ""
echo "── Build-time leftovers ──"
# npm 11 installs EVERY optional platform package of a native dependency, not
# just the host's (it ignores os/cpu, --os/--cpu and npmrc os=/cpu=). Two
# families are affected and pruned in Dockerfile.variant: @esbuild/<platform>
# and @mariozechner/clipboard-<triple>. Keep-set is the host arch only, plus
# clipboard's gnu AND musl (its napi loader picks between them at runtime).
# Runs as root because the image declares no USER; that is also what lets the
# cache assertion below read /root.
run "no foreign platform packages (npm 11 sentinel)" \
'arch=$(node -p process.arch); bad=$(find /usr/lib/node_modules /opt -type d \( -regex ".*/@esbuild/[^/]+" -o -regex ".*/@mariozechner/clipboard-[^/]+" \) ! -name "linux-$arch" ! -name "clipboard-linux-$arch-gnu" ! -name "clipboard-linux-$arch-musl" -prune -print 2>/dev/null); if [ -n "$bad" ]; then echo "foreign platform dirs shipped:" >&2; echo "$bad" >&2; du -sm $bad 2>/dev/null | sort -rn | head -5 >&2; exit 1; fi; echo ok'
# The build's own npm download cache is not free: it lands in the layer that
# created it. v1.9.1 shipped 145 MB of /root/.npm/_cacache (35 MB in v1.8.14)
# — the largest single item in its +131 MB residual, and invisible to the
# size-gate diagnostics because those only looked under node_modules and /opt.
# Nothing at runtime reads it: root's cache, while the container runs as
# `developer`. NOTE the assertion must run as root or a permission error on
# mode-700 /root would make `test ! -d` pass for the wrong reason.
run "no build-time npm cache shipped (/root/.npm)" \
'test "$(id -u)" = "0" || { echo "assertion needs root to read /root" >&2; exit 1; }; if [ -e /root/.npm ]; then echo "/root/.npm shipped: $(du -sm /root/.npm | cut -f1) MB" >&2; exit 1; fi; echo ok'
# The prune's risk is not "too big" but "removed something needed", and only a
# FUNCTIONAL check covers that. These load the natives from every install site
# found in the image, so they also scale to the studio variant's third site.
#
# NOTE THE PATH-QUALIFIED require(). The obvious form, `node -e
# 'require("esbuild")...'`, resolves by walking up from the CURRENT DIRECTORY —
# so it fails with MODULE_NOT_FOUND from /workspace on a perfectly good image,
# because esbuild lives nested inside the pi trees and global installs are not
# on node's require path (NODE_PATH is unset). That exact command was left in a
# runbook as "if this fails, revert the release", and it duly failed for the
# wrong reason on the first machine that ran it. A check must fail only for the
# thing it is checking.
run "esbuild works at every install site (prune removed weight, not function)" \
'sites=$(find /usr/lib/node_modules /opt -type d -path "*/node_modules/esbuild" -prune 2>/dev/null); if [ -z "$sites" ]; then echo "no esbuild install found at all" >&2; exit 1; fi; for d in $sites; do node -e "require(\"$d\").transformSync(\"const x:number=1\",{loader:\"ts\"})" || { echo "esbuild broken at $d" >&2; exit 1; }; done; echo ok'
# Clipboard is the family pruned second, and its napi-rs loader picks its native
# binding at require() time — so a successful load IS the proof that the kept
# platform package is the one this image needs.
run "clipboard native loads at every install site" \
'sites=$(find /usr/lib/node_modules /opt -type d -path "*/node_modules/@mariozechner/clipboard" -prune 2>/dev/null); if [ -z "$sites" ]; then echo "no @mariozechner/clipboard install found at all" >&2; exit 1; fi; for d in $sites; do node -e "var c=require(\"$d\"); if (typeof c.setText !== \"function\") { throw new Error(\"native binding missing\"); }" || { echo "clipboard native broken at $d" >&2; exit 1; }; done; echo ok'
# ── Image size ────────────────────────────────────────────────────────
echo ""
echo "── Image size ──"
@@ -747,6 +986,32 @@ elif [ "$SIZE_MB" -le "$SIZE_THRESHOLD_MB" ]; then
printf " ✅ size: %d MB (threshold %d MB)\n" "$SIZE_MB" "$SIZE_THRESHOLD_MB"; PASS=$((PASS+1))
else
printf " ❌ size: %d MB exceeds threshold %d MB\n" "$SIZE_MB" "$SIZE_THRESHOLD_MB"; FAIL=$((FAIL+1))
# A bare "too big" verdict cost a full CI-log dig plus a local npm bisect to
# attribute the v1.9.0 overshoot (+431 MB, which turned out to be npm 11
# installing 26 @esbuild platform binaries per pi-coding-agent copy). The
# container already knows where its bytes are, so make it say so: the biggest
# layers, and the biggest directories under the paths that historically grow.
# Same principle as the run() helper above — a red assertion should carry its
# own diagnostic rather than send the next reader spelunking.
echo " ── largest layers (docker history) ──"
docker history --format '{{.Size}}\t{{.CreatedBy}}' "$IMAGE" 2>/dev/null \
| grep -vE '^0B' | head -12 | sed 's/^/ /' | cut -c1-160
echo " ── largest directories in the image ──"
docker run --rm --entrypoint sh "$IMAGE" -c \
'du -sm /usr/lib/node_modules/* /opt/* /usr/local/share/ms-playwright 2>/dev/null | sort -rn | head -12' \
2>/dev/null | sed 's/^/ /' || echo " (could not inspect directories)"
echo " ── build caches that should not be in the image ──"
# v1.9.1's residual was 145 MB of npm cache under /root, and the du list
# above cannot see it: it enumerates node_modules and /opt only. A gate whose
# diagnostic looks only where the bytes were LAST time sends the next reader
# spelunking again, so name the cache paths explicitly.
docker run --rm --entrypoint sh "$IMAGE" -c \
'du -sm /root/.npm /root/.cache /tmp/node-compile-cache /home/developer/.npm 2>/dev/null | sort -rn' \
2>/dev/null | sed 's/^/ /' || true
echo " ── foreign platform dirs (npm 11 regression sentinel) ──"
docker run --rm --entrypoint sh "$IMAGE" -c \
'find /usr/lib/node_modules /opt -type d \( -regex ".*/@esbuild/[^/]+" -o -regex ".*/@mariozechner/clipboard-[^/]+" \) -printf "%f\n" 2>/dev/null | sort | uniq -c | sort -rn | head' \
2>/dev/null | sed 's/^/ /' || true
fi
# ── Summary ───────────────────────────────────────────────────────────