Compare commits

..

16 Commits

Author SHA1 Message Date
joakimp f5c53b8693 release: date the v1.9.2 section for the tag
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 27s
Publish Docker Image / base-decide (push) Successful in 13s
Publish Docker Image / build-base (push) Successful in 43m59s
Publish Docker Image / smoke (push) Successful in 5m29s
Publish Docker Image / smoke-studio (push) Successful in 5m50s
Publish Docker Image / build-variant (push) Successful in 17m29s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / update-description (push) Successful in 10s
Publish Docker Image / build-variant-studio (push) Successful in 21m50s
CI checks out the tag, so the CHANGELOG heading has to name the version
before the tag exists, not after. v1.9.1's tagged tree carried no
"## Unreleased" either -- same convention, made explicit here.

Contents of v1.9.2 (all measured, nothing inherited):
  - 1baba79  the +131 MB v1.9.1 residual: /root/.npm (110 MB) plus
             @mariozechner/clipboard-* foreign natives at BOTH install
             sites (global + /opt/pi-fork's nested pi-coding-agent copy)
  - 852f900  smoke sentinels that name the residue instead of letting
             the size gate's ~225 MB margin swallow it
  - dab989b  mempalace-toolkit: the owed-withdrawal suite gate
             (node --check never type-strips)
  - 9aaff26  vendored mempalace skill snapshot -> e9e45f7, which is what
             turns the currently-red canary green

Deliberately NOT in this tag: DOCKER_HUB.md's "~1.1 GB" size claim is
low (Hub says v1.9.1 = 1.37 GB, v1.8.14 = 1.23 GB). This release
*changes* the size, so no number is right both before and after it --
writing the predicted ~1.24 GB would be an unmeasured claim in a
user-visible page. Measure post-build, fix on the next tag, and extend
check-doc-drift.sh to gate size claims against Hub's full_size so the
number cannot rot silently again.
2026-09-14 17:55:16 +02:00
joakimp 735565b9be docs: retire the deployed-and-unproven label on isWithdrawn, and name dab989b
Lint / skill-floor (push) Successful in 8s
Lint / hadolint (push) Successful in 17s
Lint / doc-drift (push) Successful in 7s
Lint / actionlint (push) Successful in 20s
Two things, and the second was found by doing the first.

v1.9.1's notes closed with an explicit open item: isWithdrawn had 17 assertions
and 4 mutation kills but had never been exercised on a released image against the
live logstream, which is a different state from untested and a worse one to leave
unlabelled. It has now been exercised, from a container recreated onto v1.9.1, and
the label is retired with the measurement rather than with an assurance.

What the entry records is the instrument as much as the result, because the result
is only as good as the thing that produced it: the shipped extractor was lifted
out of the toolkit's own test script and used to cut the four predicates out of
the BAKED mempalace.ts, and deriveOwed's three queries were replayed with their
exact shipped parameters. No second copy of the logic was written, which is the
same discipline the suite itself is built on. Baseline owed was established by
three independent routes BEFORE any probe was planted, because without a recorded
baseline a withdrawal that appears to work proves nothing. Predictions were
written down before each measurement; the table reports both columns.

The per-predicate attribution is in the entry deliberately. "The count went back
to baseline" is a claim about arithmetic and would have been satisfied by a
pre-filter or by an unrelated reply; isAnswered=false with isWithdrawn=true on a
candidate that was still in the raw set is a claim about mechanism. The negative
controls carry the same weight as the positive one -- a rule that lets a third
party or a prose sentence empty someone's mailbox would have passed a
count-only check.

Also recorded: the rule fires on the real incident it was written for (seq 112),
which is worth more than any synthetic probe, and the one path still unwitnessed
(the extension's own in-process poll, which only fires when the agent is idle) is
named as unwitnessed rather than folded into the pass.

Second: reaching for the suite as corroboration showed it cannot run on this image
at all -- its own precondition gate rejects the source it exists to check, because
`node --check` does not type-strip on any node and v1.9.1 moved to node 24. All 17
assertions had been skipped since the bump. Fixed in toolkit dab989b, named here
per the rule this repo adopted two commits ago after the withdrawal fix shipped
unnamed in v1.9.1. The gate now also distinguishes "cannot run" from "does not
parse", since collapsing those is how the defect pointed readers at the wrong
file.

The rebuild cost is stated and, having been measured, is smaller than the earlier
draft of this entry claimed: 9aaff26 already moved base_tag by refreshing the
vendored skill snapshot under rootfs/, so the base rebuild was pending before this
fix existed and the toolkit pickup rides along with it.
2026-09-14 16:29:52 +02:00
joakimp 9aaff26e3a docs: the withdrawal fix shipped in v1.9.1 unnamed — say so, and re-pin the canary
Lint / hadolint (push) Successful in 11s
Lint / doc-drift (push) Successful in 12s
Lint / skill-floor (push) Successful in 12s
Lint / actionlint (push) Successful in 16s
v1.9.1 bakes mempalace-toolkit e68ee20, which contains e2b060a. Requester-side
ask withdrawal has therefore been LIVE on every v1.9.1 device since 2026-09-10,
while the v1.9.0 section of CHANGELOG.md still read "not yet pinned ... this
image still pins e45f6b4", and the mempalace skill still told every agent at
session start that a withdrawal is impossible.

Measured, two independent routes, expectation recorded before looking:

  - published image label, :v1.9.1-studio and :latest-studio (same digest):
      se.jordbo.pi-devbox.mempalace-toolkit-ref = e68ee2071ca3ad39...
  - ancestry: e2b060a is an ancestor of e68ee20
  - baked mempalace.ts sha 7c16fe14 != v1.8.14's dfca71e9 (different bytes)
  - grep -c isWithdrawn on this container's baked copy = 0 (v1.8.14, so
    tor-ms22 cannot exercise the behaviour it is documenting)

The first label read came back EMPTY, and that empty was a claim about the
request rather than the image: Hub redirects blob fetches to a CDN and curl
without -L returns 0 bytes at exit 0. Recorded in the notes, because a registry
audit reporting "no labels" is missing -L until proven otherwise.

Changes:

  - CHANGELOG Unreleased: the floating-ref mechanism, which is the reusable
    part. ARG MEMPALACE_TOOLKIT_REF=main + CI resolving it to a SHA at build
    time means a release absorbs whatever toolkit main holds, and "what
    behaviour did this image gain" is a question nobody is forced to answer.
    The fleet rule (name the pickup before tagging) was honoured for the
    feed-tick commits v1.9.1 names, and missed for one in the same range.
  - CHANGELOG v1.9.0: the stale caveat is ANNOTATED, not rewritten. The
    sentence is the evidence for how the drift happened; deleting it would
    destroy the only trace. Same discipline as fleet-ops d42412c.
  - CHANGELOG Unreleased: corrected the size entry's "needs no base rebuild"
    claim. True of that entry alone, false of the release now that the
    vendored skill snapshot -- a base_tag input -- moves with it (~67 min).
  - rootfs mempalace skill snapshot 44472 -> 46045 B, refreshed via
    scripts/vendor-mempalace-skill.sh so the bytes and SKILLSET_SNAPSHOT_REF
    (4d7c0ea -> e9e45f7) move together; --check confirms exact match.
  - scripts/smoke-test.sh canary RE-PINNED. The retired pair was still green
    against the new snapshot, i.e. blind to this refresh for the same reason
    the pre-v1.8.13 pair was blind to that one. The replacement is stronger
    than any predecessor here because BOTH witnesses come from the same
    upstream commit: e9e45f7 added "Withdrawing an ask you sent" and deleted
    "nothing anyone can do about it from the other end", the sentence the new
    bullet contradicts. Directions measured against both files, then the
    canary body EXECUTED against each: new -> rc=0 "ok", old -> rc=1 empty.
    A canary whose negative witness was removed by the commit it pins fails
    loudly on stale bytes instead of merely failing to notice them.

Gates: check-doc-drift OK, check-skill-floor OK, vendor --check OK,
hooks/pre-push OK (16 shell files clean at severity error).

Not fixable here: isWithdrawn is deployed and UNPROVEN. 17 assertions and 4
mutation kills, never once exercised on a released image against the live
logstream. Ask routed to a v1.9.1 device.
2026-09-14 15:07:26 +02:00
Joakim Persson 852f900b53 test(smoke): make CI prove the natives still work — the runbook check didn't
Lint / doc-drift (push) Successful in 6s
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 9s
Lint / actionlint (push) Successful in 18m11s
The post-boot check v1.9.1 left for the next machine was

  node -e 'require("esbuild").transformSync("const x:number=1",{loader:"ts"})'

with "if this fails, the prune removed something needed on arm64 -> revert to
v1.8.14 and fix" attached. Run from /workspace on the freshly recreated v1.9.1
image it fails with MODULE_NOT_FOUND — on a perfectly healthy image. require()
resolves by walking up from the CURRENT DIRECTORY, esbuild is nested inside the
two pi trees, and a global install is not on node's require path (NODE_PATH is
unset). The version that "verified the prune" last night only passed because the
shell happened to sit inside the tree. So the check was CWD-dependent, its red
meant nothing, and the remedy it prescribed was reverting a good release.

Path-qualified, both sites compile TS today. Rather than leave a fragile command
in a note for a human to paste, CI now owns it:

  - esbuild must transformSync at EVERY install site found in the image
  - @mariozechner/clipboard must load with its native binding attached at every
    site — for the clipboard prune that IS the proof, since napi-rs resolves the
    platform package at require() time

Sites are discovered with find, so the studio variant's third site is covered
without naming it.

Verified as a four-way matrix, not just a green run: both assertions pass
against the real image FROM /workspace (the cwd that broke the old command), and
both fail (exit 1, "The package \"@esbuild/linux-arm64\" could not be found, and
is needed by esbuild") against copies of the same packages with the host
platform binary removed. An assertion never observed failing is not evidence.

Also documents both failure shapes in AGENTS.md next to the size gate: this one,
and the mode-700 /root one where `test ! -d` passes on a permission error.
2026-09-11 11:06:18 +02:00
Joakim Persson 1baba79c96 fix(size): the v1.9.1 residual was npm's own cache, not the platform binaries
Lint / doc-drift (push) Successful in 6s
Lint / hadolint (push) Successful in 10s
Lint / skill-floor (push) Successful in 9s
Lint / actionlint (push) Successful in 4m26s
v1.9.1's @esbuild prune fixed the 431 MB size-gate failure but still shipped
+131 MB compressed over v1.8.14, nearly all of it in the pi/extensions install
layer (87 -> 206 MB). That leftover was filed as an open item with an explicit
hypothesis — the @mariozechner/clipboard-* family, same npm 11 behaviour, a
different package — and an explicit warning that the hypothesis was not a
measured cause. Measured now, after recreating onto v1.9.1, and the hypothesis
accounted for one sixth of it:

  +110 MB  /root/.npm/_cacache  (35.2 -> 145.3 MB)   the build's npm cache
  + 21 MB  clipboard foreign platform packages, both install sites
  = 131 MB  i.e. the whole delta, no unexplained remainder

Method, since there is no docker CLI inside the container: pulled both variant
layer blobs straight from the registry with a token + manifest + blob fetch and
listed the tarballs (29 789 vs 29 889 entries, 270.8 vs 401.9 MB uncompressed),
then aggregated per package. The file COUNT barely moved, which is what said
"few large files", not "npm installed more packages".

  - purge_build_caches: npm cache clean --force + rm -rf /root/.npm, in the SAME
    layer as the installs, in both the main RUN and the studio RUN. npm 11
    caches every platform tarball it downloads, including the ones the prune
    then deletes, so the cache grew faster than the tree. Nothing at runtime
    reads it: build is root, container is developer with its own cache in $HOME.
  - prune_foreign_esbuild -> prune_foreign_natives: now covers both MEASURED
    families. Clipboard keeps linux-$arch-gnu AND -musl because its napi-rs
    loader picks between them at runtime via its own isMusl() probe; the musl
    package is a 420-byte stub. The bare @mariozechner/clipboard wrapper has no
    hyphen suffix and cannot match the pattern.

Verified on arm64 before writing the glob — a widened rm -rf against a tree you
cannot inspect is the one change shape not to write blind, which is why the
order was update-then-patch. Exercised against a copy of the real trees with
foreign dirs fabricated back in (aix-ppc64, android-arm64, darwin-arm64,
win32-x64, linux-x64): all removed, host linux-arm64 kept at both sites, 21 MB
freed, require('@mariozechner/clipboard') still loads and exports all 18
functions, esbuild.transformSync still compiles TS at both sites. v1.9.1's own
arm64 validation also passed here; CI could only smoke amd64.

Two sentinel assertions, because the size gate did not catch this: it has
~225 MB of margin, so 131 MB of residue stayed green. Both were verified RED
against the running v1.9.1 image and GREEN against a pruned tree. The cache one
refuses to run as non-root: test ! -d /root/.npm on mode-700 /root would
otherwise pass for the wrong reason. Size-failure diagnostics now list cache
paths too — they previously enumerated only node_modules and /opt, where these
bytes were not.

Also fixed: -printf '%f\\n' reaches the shell with both backslashes (confirmed
from the published image's recorded created_by), so v1.9.1's progress line
printed a mangled "li ux-arm64" — find emitted a literal backslash and tr ate
the n out of the name. Single backslash now.

Deliberately not purged: /tmp/node-compile-cache (1.3 MB). The manifest RUN
calls pi --version again, so deleting it earlier only relocates those bytes into
that layer — today's manifest layer is 128 kB precisely because it finds the
cache warm.

Dockerfile.base untouched, so no base rebuild: this rides the next release.
2026-09-11 09:39:54 +02:00
Joakim Persson 42bd29d654 fix(ci): unblock the release — npm 11 esbuild bloat + two self-inflicted assertions
Lint / hadolint (push) Successful in 8s
Lint / skill-floor (push) Successful in 9s
Lint / doc-drift (push) Successful in 10s
Lint / actionlint (push) Successful in 19s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 14s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Successful in 5m4s
Publish Docker Image / smoke-studio (push) Successful in 13m17s
Publish Docker Image / build-variant (push) Successful in 33m49s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 14s
Publish Docker Image / build-variant-studio (push) Successful in 30m58s
v1.9.0 was tagged but never published: smoke failed 90-passed/3-failed and
build-variant needs smoke, so nothing reached the registry. All three are fixed.

1. Size, 431 MB over threshold. Node 24 brings npm 11, which installs EVERY
   @esbuild/<platform> optional binary instead of the matching one: 26 dirs,
   284 MB per pi-coding-agent copy. Measured on pi-fork: npm 10.9.8 -> 165 MB
   (exactly what v1.8.14 shipped), npm 11.19.0 -> 449 MB. npm 11 ignores the
   os/cpu constraints AND --os/--cpu AND an npmrc carrying them, all measured,
   so Dockerfile.variant prunes explicitly, keeping linux-$(node -p
   process.arch) so one line is right on both arches. Pruned in the SAME layer
   as each install, or the bytes survive in the earlier layer. Three sites:
   global pi, pi-fork, pi-studio. Verified esbuild still transforms TS after.

2. om's node_modules assertion tested an npm artefact. om has zero runtime deps;
   its node_modules held ONE file (.package-lock.json) and 20 empty scope dirs.
   npm 11 stopped creating it. Now asserts the entry point pi actually loads,
   read from package.json -> pi.extensions.

3. The skill-source annotation added in v1.9.0 ("baked (package copy)") broke
   the assertion matching "baked$". Pattern now allows an optional suffix; which
   copy shipped stays authoritatively asserted against the manifest + tree hash.

Also: a failed size check now prints the largest layers, largest directories and
an @esbuild sentinel, so this class attributes itself next time instead of
costing a CI dig plus a local npm bisect.

Threshold stays 3800 MB: it caught a real regression and raising it would have
thrown the signal away. No Dockerfile.base/rootfs change, so base-0fb1256c7f99
is reused and build-base is skipped.
2026-09-10 23:39:40 +02:00
Joakim Persson 3a44e81cad feat(ci): gate documentation drift, and make docs a pre-tag release step
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Five doc claims had rotted by v1.9.0, all the same shape: a value written once
by hand, in a file nothing verifies, about a number that lives elsewhere and
moved. README pin table wrong on all three rows; a "Planned" section describing
something already shipped; DOCKER_HUB.md claiming Node v22 against Node 24.

DOCKER_HUB.md is why this is a gate and not a resolution to be careful: it is
PUBLISHED (update-description POSTs it as Docker Hub full_description on every
tag), it had gone eight releases untouched, nothing generates it, and it is read
from the TAG -- so the stale page shipped with v1.9.0 regardless.

scripts/check-doc-drift.sh: seven checks, all repo-local (no network, token,
image, or sibling clone). Exit 0/1/2 matching lint-shell.sh; a renamed ARG is a
red 2, not a green tick. Wired as a fourth lint.yml job so "the docs lie" is its
own red name.

Verified with 15 controls, including two false-positive controls: the first
placeholder check flagged README.md:900, a Go template in a legitimate
`docker inspect --format` example. The gate was wrong, not the doc, so the
pattern is now anchored to the UPPER_SNAKE convention CI substitutes.

Not gated, deliberately: counts/sizes needing a running image (they belong in
smoke-test.sh -- a guessing gate is worse than none), and Dockerfile.base
BASE_REBUILD_DATE, because base_tag hashes that file content-wise and demanding
it be current would force a ~60 min rebuild on releases that touch no base
files. Free during a rebuild, expensive otherwise.

AGENTS.md step 3 rewritten around the mechanism: checkout@v4 with no ref: means
every job reads github.ref, the tag. Docs must be right BEFORE tagging.
2026-09-10 22:01:49 +02:00
Joakim Persson 35964abd01 docs(hub): the Docker Hub page claimed Node v22; v1.9.0 ships Node 24
DOCKER_HUB.md is hand-maintained (no generator: CI only substitutes
{{PI_VERSION}} and POSTs the file as Docker Hub full_description), and it was
last touched at v1.8.6 -- eight releases ago. The Node claim is now false as of
this release, and unlike README.md this file IS published, so a stale claim here
is user-visible rather than internal.

Verified countable claims rather than assuming: "7 user-facing extensions" is
correct (7 files in pi-extensions/extensions/). The "29 mempalace_* tools" claim
is suspect -- this session sees 45 -- but left alone because I cannot attribute
that count to the baked 3.9.0 server without measuring it.
2026-09-10 21:43:17 +02:00
Joakim Persson 6353d59e63 docs(readme): correct the three stale version pins and the shipped-as-planned claim
The version-pin table existed precisely to be the reviewable record of what is
deliberately frozen, and it was wrong on every row: pi 0.84.4 -> 0.85.1,
pi-atelier v0.10.0 -> v0.10.1, mempalace 3.8.0 -> 3.9.0. Verified by parsing the
table and comparing against the ARGs it names rather than by eye.

Also: the "Planned for an upcoming minor release" section listed typst PDF
export, which shipped long ago and even carried a self-contradicting
"(shipped in Unreleased/base)" marker -- the fourth instance of the stale
in-repo Unreleased-pointer class this CHANGELOG already documents. typst 0.15.1
confirmed live in the running image, so the item is now stated as current fact.

The pi-devbox-version sample was v1.5.0-era and structurally outdated: it
predates the palace line the surrounding prose advertises, the pi-atelier
component, and the whole skills: block. Replaced with real observed output
rather than hand-written text.

README has no CI coupling (no workflow or gate reads it; the Docker Hub page
comes from DOCKER_HUB.md), so this cannot affect the in-flight v1.9.0 build.
2026-09-10 21:33:07 +02:00
Joakim Persson 8f0960e134 release: v1.9.0
Lint / skill-floor (push) Successful in 8s
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / lint-gate (push) Successful in 22s
Publish Docker Image / resolve-versions (push) Successful in 15s
Publish Docker Image / base-decide (push) Successful in 8s
Publish Docker Image / build-base (push) Successful in 42m19s
Publish Docker Image / smoke-studio (push) Failing after 6m17s
Publish Docker Image / build-variant-studio (push) Has been skipped
Publish Docker Image / smoke (push) Failing after 9m7s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Rename the Unreleased changelog section to its release heading, matching the
established `## vX.Y.Z — YYYY-MM-DD` format (em dash), per AGENTS.md release
step 3.

Minor rather than patch: CHANGELOG.md scopes minor to "significant base
additions", and this release bumps the Node runtime under every baked JS tool
(pi, agent-browser, playwright, mempalace) from 22 to 24 — the first
NODE_VERSION change since it was introduced at v1.0.0 — alongside five new
base packages (shellcheck, bind9-dnsutils, ldap-utils, xxd, python3-yaml).
Package additions alone have been patch here (v1.8.12, v1.8.13); the runtime
major is what lifts this one.
2026-09-10 21:21:09 +02:00
Joakim Persson 1c905480e3 docs(changelog): note the mempalace feed-tick fix this image carries
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 10s
Lint / actionlint (push) Successful in 19s
The recurring "[mempalace ext] feed (tick) failed: mine timed out after 30000ms"
message is fixed in mempalace-toolkit (309980b + e68ee20) and this image is what
delivers it, since CI resolves MEMPALACE_TOOLKIT_REF to a commit SHA at build
time. Worth a changelog entry rather than leaving it implicit in a ref bump: it
is the most visible symptom operators on this fleet have been living with, and
the entry records that it was a genuine defect (overlapping mines on a
single-writer palace) rather than the cosmetic annoyance it was parked as.
2026-09-10 20:58:42 +02:00
Joakim Persson ff6fd1492a feat(manifest): record WHICH pi-extensions skill copy shipped
Closes the half deliberately left open by cac5e00's skill-floor gate, and the
more important half: "the floor is currently fresh" is a fact with a shelf
life, whereas "the image says which copy it got" keeps working.

The refresh in Dockerfile.variant is guarded by
`[ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates the
co-located skill keeps the vendored floor and still succeeds GREEN, with nothing
in the manifest, labels or logs separating that from a normal build. Afterwards
the two are indistinguishable by inspection -- same path, same filenames, same
permissions -- which is exactly how the floor went unnoticed from 2026-07-30 to
2026-09-10.

build-manifest.json gains pi_extensions_skill_source and
pi_extensions_skill_tree_sha256, MEASURED rather than passed as build-args, per
the ground-truth rule the surrounding block already follows -- and necessarily
so, since the outcome depends on the clone's contents and no ARG could express
it. Three values, because two would force a lie: package (served bytes equal
the clone's skill/), vendored-floor (clone had no skill/ at this ref), and
divergent (both exist but differ -- e.g. the clone ships SKILL.md but not
evaluate-extension-usage.py, so the served directory is a genuine MIX). No OCI
label mirrors these deliberately: LABEL cannot take a RUN-computed value, and a
label fed from an ARG would be the claim-not-measurement being removed here.

Two smoke assertions make the record a gate: the source must be named and be
`package` -- vendored-floor FAILS rather than warns, since these images track
main where the package has shipped skill/ since fa04d20, so a fallback means
the clone did not resolve as intended -- and the tree hash is recomputed over
the served directory, because a recorded hash never recompared is a claim.

pi-devbox-version annotates the line too: "baked (package copy)" normally, or a
yellow "(FALLBACK: vendored floor)". Its existing section reports which copy is
READ at runtime; this is the one fact decided at BUILD time and unrecoverable
later. Old images degrade cleanly -- field absent, jq // empty yields nothing,
line prints plain "baked" as before (verified against this v1.8.14 manifest).

Tested by running the exact logic against this container's real layout, with
the expected value written down before each: package (served == clone),
vendored-floor (clone path absent), divergent (clone lacking the .py while the
served dir has it), and null (empty served dir) -- all four as predicted. The
five pi-devbox-version render branches likewise, including the absent-field
case. Emitted JSON validated with jq for both the populated and null forms.

Gates green: lint-shell.sh (15 files), hadolint 2.15.1, actionlint 1.7.12,
check-base-hash.sh, check-skill-floor.sh, vendor-mempalace-skill.sh --check.
2026-09-10 20:30:16 +02:00
Joakim Persson edc7659add chore(deps): node 22->24, actionlint 1.7.12, hadolint 2.15.1, skillset ref
Audited every component the image obtains OUTSIDE debian/apt. Of ~23, the 19
that resolve `latest` at build time were already current or refresh themselves
on the next rebuild, and the hard pins for pi (0.85.1), mempalace (3.9.0) and
pi-atelier (v0.10.1) were already newest. Four needed a human.

NODE_VERSION 22 -> 24 (LTS "Krypton"). This was a latent defect rather than
housekeeping: agent-browser publishes engines.node ">=24.0.0", so the image sat
BELOW a declared requirement -- v1.8.14 shipped node 22.23.2 with agent-browser
0.37.1, so every build installed it with an npm EBADENGINE warning and ran the
baked browser automation outside its supported range. pi (">=22.19.0") and
playwright (">=20") are satisfied either way. Verified before bumping, since a
missing NodeSource suite breaks every arch at once: setup_24.x returns HTTP 200
and node_24.x advertises `Architectures: amd64 arm64 armhf x86_64`, covering
the arm64 fleet and the amd64 CI runners. Nothing else pinned the node major.

actionlint 1.7.7 -> 1.7.12 and hadolint 2.14.0 -> 2.15.1, each RUN AGAINST THIS
TREE at the new version before being pinned -- both clean, no new findings. A
linter bump is the one dependency update that can turn CI red on unchanged
code, so it is verified locally rather than discovered on a round trip.

SKILLSET_SNAPSHOT_REF e9e09d9 -> 4d7c0ea via scripts/vendor-mempalace-skill.sh,
never by hand: that script is the only thing permitted to write the ARG,
because a cp without a matching bump yields a manifest that confidently lies.
This proved PROVENANCE-ONLY -- the ref was 6 commits behind, but
skills/mempalace/SKILL.md is byte-identical at both (3675bfab), so the snapshot
was already correct and only its recorded origin was stale. No rootfs/ bytes
changed, the smoke-test phrase canary stays valid, and this ARG alone would not
force a base rebuild (the node bump does).

Two measurement traps worth recording, since both would have produced a wrong
answer: GitHub's releases/latest reports pi-atelier v0.10.0 as newest because
v0.10.1 is a TAG WITH NO RELEASE OBJECT -- the pin was already current, and
`git ls-remote --tags` is the instrument that shows it. And gitea-mcp is hosted
on gitea.com, not GitHub, so querying api.github.com returned nothing at all
rather than an error.

Verified with every gate this repo owns, all green, using the NEW linter pins:
lint-shell.sh (15 files), check-workflow-shell.sh, check-base-hash.sh,
actionlint 1.7.12, hadolint 2.15.1, check-skill-floor.sh, and
vendor-mempalace-skill.sh --check.
2026-09-10 20:17:32 +02:00
Joakim Persson cac5e00a31 feat(ci): gate the vendored pi-extensions skill floor, and bake python3-yaml
Follows ecfd2fc, which refreshed the stale floor by hand. A one-off refresh
fixes the symptom; this makes the drift impossible to reintroduce silently.

scripts/check-skill-floor.sh compares the repo floor
(rootfs/usr/local/share/pi-devbox/skills/pi-extensions/) against the package
repo it is a snapshot of, wired in as a new `skill-floor` job in lint.yml.

DIRECTORY hash, not `sha256sum SKILL.md`, using the same tree_sha256 pipeline
Dockerfile.variant uses for skillset_snapshot_tree_sha256 and for the reason
already documented there: a file-only compare answers "did this one file
change", not "is this the same skill". Verified by NEGATIVE CONTROL rather
than asserted -- with SKILL.md left byte-identical and only
evaluate-extension-usage.py edited, the directory check fails (rc=1) where a
file-only compare would have passed. Seven behaviour tests, each with its
expected rc written down before running: in-sync via local dir (0), in-sync
via anonymous remote clone (0), missing --package-dir (2), bad argument (2),
content drift (1), the sibling-file case (1), and --warn-only over drift (0).

Exit codes 0 in sync / 1 drift / 2 cannot-run, matching scripts/lint-shell.sh:
a gate that cannot run must not pass, so an unreachable package repo is a red
2 and never a green tick. A ref with no skill/ is NOT drift -- that is the
documented fallback -- but it emits ::warning:: because it is precisely the
condition under which the floor ships.

Gating on another repo is normally a smell. It is proportionate here because
the check can only fire when skill/ itself changed, which is exactly when the
floor has gone stale; pi-extensions commits that leave skill/ alone cannot
turn this red. It also needs no secret: pi-extensions is anonymously clonable
(verified with `git ls-remote` and no credentials), so it cannot start failing
when a token expires.

Also bakes python3-yaml (552 KB, zero extra deps) into Dockerfile.base. This
is the shellcheck story repeating exactly: scripts/check-workflow-shell.sh --
the guard against the Gitea sh/dash footgun that broke resolve-versions
(ed49b8d) and promote-base-latest (b7197e8) -- hard-exits with "python3 yaml
module missing", so a gate this repo already owns could not be run locally by
anyone. lint.yml installing it explicitly in CI was the evidence. Found while
wiring the job above: the guard could not be run before pushing.

CHANGELOG Unreleased updated for both this and ecfd2fc, including an explicit
note on what is NOT fixed -- the silent-fallback half still has no manifest
flag recording which copy was served.

Verified locally with every gate this repo owns, all green: lint-shell.sh (15
files clean), check-workflow-shell.sh, check-base-hash.sh, actionlint 1.7.7
(pinned, same version as CI), hadolint 2.14.0, and the new check itself.
2026-09-10 19:14:01 +02:00
Joakim Persson ecfd2fc2e5 feat: bake dig/ldapsearch/xxd and refresh the stale pi-extensions rootfs floor
Two changes that share one forced base rebuild, hence one commit.

1. THREE PACKAGES, each closing a capability gap measured during the
   gitea.egl.lan/FreeIPA work on 2026-09-09..10 rather than a preference:

   bind9-dnsutils (~6.1 MB measured) -- dig/host/nslookup were ALL absent,
   so the container could resolve names but had no way to interrogate a
   SPECIFIC nameserver. `getent hosts` only follows the resolver's default
   path, so diagnosing "gateway 172.16.88.1 NXDOMAINs the egl.lan zone
   while 10.20.253.1 is authoritative for it" had to be hand-rolled in
   python3. Split-horizon DNS is a recurring class of bug on this fleet.
   Note the package name: plain `dnsutils` is transitional in trixie.

   ldap-utils (1244 KB, pulls nothing extra) -- the fleet authenticates
   against FreeIPA, yet every LDAP probe had to be run by SSHing to an
   already-enrolled host. Simple binds only; GSSAPI would additionally
   need krb5-user + libsasl2-modules-gssapi-mit, deliberately not added
   as that is a Kerberos-client decision, not a tool.

   xxd (198 KB) -- convenience for verifying git-crypt blob magic in
   myconfigs; `od -c` from coreutils already does the same job.

   netcat-openbsd was in the original proposal and is deliberately NOT
   here: measured redundant, because socat is already baked and bash's
   /dev/tcp does reachability checks with zero packages (verified against
   gitea.egl.lan:3000). Recorded in the Dockerfile so the omission reads
   as a decision rather than an oversight.

2. ROOTFS FLOOR REFRESH: rootfs/.../pi-extensions/SKILL.md was 34284 B,
   unchanged since fa04d20 (2026-07-30), while the canonical package copy
   is 38973 B. Dockerfile.variant copies the fresh package copy over the
   SERVED path at build time but never writes back to this floor, so the
   floor is a silent fallback: if that build-time copy is ever absent it
   ships the July skill with no log line or manifest flag to say which
   version deployed. Refreshed from pi-extensions@c64c122, verified
   byte-identical to both the canonical and the runtime-served copies.

Why one commit: the base_tag hash folds in `cat Dockerfile.base` AND
`find rootfs -type f | xargs cat` (.gitea/workflows/docker-publish.yml),
so either change alone forces the same full base rebuild -- and that
rebuild is precisely what re-bakes rootfs/ as it then stands. Emulating
the workflow hash with a fixed toolkit ref: f3d6462c7416 -> fc4edda03c54.

Verified: scripts/check-base-hash.sh passes (no new ARG *_REF added), and
no shell scripts are touched so the lint-shell gate is unaffected. Sizes
and dependency fan-out measured via apt-get --no-install-recommends
--dry-run on Debian 13 trixie.
2026-09-10 18:56:11 +02:00
joakimp 15a3728ae9 feat: bake shellcheck and add a client-side pre-push lint gate
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 19s
v1.8.14 made shell lint a RELEASE gate (scripts/lint-shell.sh, shared by lint.yml
and the new lint-gate job that resolve-versions depends on), and that script
correctly exits 2 when shellcheck is absent -- "a gate that cannot run must not
pass". Measured on v1.8.14 on 2026-09-09 by three routes (command -v, dpkg -l, a
filesystem search): shellcheck was NOT IN THE IMAGE AT ALL. So the gate could not
be run by a developer in any container, only in CI, and the loop stayed
write-shell -> push -> wait for CI -> discover. That is the loop the gate was
added to shorten, after v1.8.14's first attempt burned ~46 min on a tree whose
lint had already been red for 24 hours.

shellcheck 0.10.0-1 added to the Dockerfile.base apt block: ~39 MB installed
(Installed-Size 40112 KB), measured to pull ZERO additional packages under
--no-install-recommends because libc6/libffi8/libgmp10 are already present.
NOTE this forces one full base rebuild -- base-decide hashes Dockerfile.base +
rootfs/, so unlike a scripts/ change it cannot reuse the existing base- layer.

hooks/pre-push is opt-in per clone (git config core.hooksPath hooks), bypassable
with --no-verify, and execs scripts/lint-shell.sh rather than reimplementing it
-- one copy, because a duplicated check that drifts is the failure this repo
keeps paying for. Matches the idiom skillset/ and myconfigs/ already use.

WHY THIS REPO HAD NO HOOKS, since it was reported as drift and is not: a peer
asked tor-ms22 for core.hooksPath per clone on the premise that unset meant the
gates were unverified there. Measured: pi-devbox unset, skillset hooks, myconfigs
common/hooks, pi-toolkit unset -- but `git ls-files | grep -i hook` is EMPTY in
both pi-devbox and pi-toolkit, so there was nothing to point at on any machine
and unset was the only correct value. This closes the real half for pi-devbox;
pi-toolkit still ships none.

Verified, expected result written down before each check:
  * refusal paths -- shellcheck absent => rc 2 with the remedy named; linter
    missing => rc 2. Never waved through on the assumption CI will catch it.
  * the hook is IN the scan set -- "Checking 14 shell file(s)" with it present,
    13 with it moved aside, so the extensionless file is found by the shebang
    half of the linter's two-signal union. This check exists because the first
    attempt was ambiguous: a planted `[ $UNSET_VAR = "x" ]` was not reported,
    which could equally have meant "not scanned" or "below -S error". It was the
    latter. A count that moves is unambiguous; a clean run is not.
  * it catches the REAL v1.8.14 defect -- planting `echo 'the fleet\'s thing'`
    in hooks/pre-push yields SC1073/SC1072 at severity error, rc=1.

And the gate earned its keep inside this commit: the first version of the
smoke-test assertion carried a comment beginning "# shellcheck is a GATE
DEPENDENCY", and a comment whose first word is the tool's name is parsed as a
DIRECTIVE, not a comment. The new gate failed it with SC1073/SC1072 before the
push -- same family as the v1.8.14 apostrophe, a line that reads as prose to a
human and as syntax to the parser.
2026-09-09 08:57:34 +02:00
14 changed files with 1814 additions and 35 deletions
+80 -2
View File
@@ -92,7 +92,7 @@ jobs:
- name: Install actionlint (pinned)
env:
ACTIONLINT_VERSION: 1.7.7
ACTIONLINT_VERSION: 1.7.12
run: |
curl -fsSL \
"https://github.com/rhysd/actionlint/releases/download/v${ACTIONLINT_VERSION}/actionlint_${ACTIONLINT_VERSION}_linux_amd64.tar.gz" \
@@ -127,7 +127,7 @@ jobs:
- name: Install hadolint (pinned)
env:
HADOLINT_VERSION: 2.14.0
HADOLINT_VERSION: 2.15.1
run: |
curl -fsSL \
"https://github.com/hadolint/hadolint/releases/download/v${HADOLINT_VERSION}/hadolint-Linux-x86_64" \
@@ -137,3 +137,81 @@ jobs:
- name: Run hadolint
run: hadolint Dockerfile.base Dockerfile.variant
skill-floor:
# Gate the VENDORED pi-extensions skill snapshot in rootfs/ against the
# package repo it is a snapshot of. Its own job rather than a step in
# `actionlint`, so "the floor is stale" is a distinct red name in the runs
# list instead of being buried in a lint job that is about something else.
#
# The gap it closes, measured 2026-09-10: the floor sat at 34284 B, untouched
# since fa04d20 (2026-07-30), while the package copy was 38973 B.
# Dockerfile.variant copies the fresh package copy over the SERVED path but
# never writes back to the floor, so nothing in the repo ever noticed. That
# matters because the floor is a FALLBACK: the copy is guarded by
# `if [ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone
# yields no skill/ ships the vendored snapshot and still goes green, with no
# manifest flag or label saying which copy was served.
#
# Gating on another repo is normally a smell; it is proportionate here
# because the check compares the skill DIRECTORY hash, so it can only fire
# when that directory actually changed — which is exactly when the floor has
# gone stale. pi-extensions commits that leave skill/ alone cannot turn this
# red. No secret is needed either: the repo is anonymously clonable (verified
# 2026-09-10 with `git ls-remote` and no credentials), so this cannot start
# failing when a token expires.
#
# Exit codes are 0 in sync / 1 drift / 2 cannot-run, matching
# scripts/lint-shell.sh: a gate that cannot run must not pass, so an
# unreachable package repo is a red 2 rather than a green tick.
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
steps:
- uses: actions/checkout@v4
- name: Vendored pi-extensions skill floor matches the package
run: bash scripts/check-skill-floor.sh
doc-drift:
# Gate hand-maintained doc claims against the build files they describe.
# Its own job for the same reason as skill-floor: "the docs lie" should be a
# distinct red name, not a line buried in a job about workflow syntax.
#
# The gap it closes, measured 2026-09-10 while preparing v1.9.0 — five
# claims had rotted, every one of them a fact written by hand in a file
# nothing verified:
# * README.md's "Version pins" table was wrong on ALL THREE rows (pi
# 0.84.4 vs 0.85.1, pi-atelier v0.10.0 vs v0.10.1, mempalace 3.8.0 vs
# 3.9.0) — and that table exists specifically to be the reviewable
# record of what the repo freezes on purpose, so a wrong row destroys
# the only thing it is for.
# * README.md listed already-shipped typst PDF export under "Planned for
# an upcoming minor release", marked "(shipped in Unreleased/base)".
# * DOCKER_HUB.md claimed "Node.js v22" while v1.9.0 ships Node 24.
#
# DOCKER_HUB.md is why this is a gate and not a habit. It is PUBLISHED —
# update-description POSTs it to Docker Hub as full_description on every tag
# — and it had gone eight releases (v1.8.6 -> v1.9.0) untouched. Nothing
# generates it and nothing checked it, so the only thing keeping it true was
# someone remembering. It is also read from the TAG, so a fix pushed to main
# after tagging never reaches the published page.
#
# Cheap and hermetic on purpose: every check compares a doc string against a
# value that exists in this repo, so no network, no token, no built image,
# and no sibling clone. Claims that genuinely need a running container (image
# sizes, the "N mempalace_* tools" count) are deliberately left out — a gate
# that cannot evaluate a claim honestly would have to guess, and a guessing
# gate is worse than none. Assert those in scripts/smoke-test.sh instead.
#
# Exit codes 0 in sync / 1 drift / 2 cannot-run, matching lint-shell.sh and
# check-skill-floor.sh. A renamed ARG makes the gate blind, so that is a red
# 2, not a green tick.
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
steps:
- uses: actions/checkout@v4
- name: Doc claims match the build files
run: bash scripts/check-doc-drift.sh
+63 -3
View File
@@ -92,7 +92,44 @@ re-brand of opencode-devbox's `pi-only` variant.
is hashed into `base_tag`, so it costs a base rebuild (~67 min); and if the
section the phrase canary names has changed, re-pin it in
`scripts/smoke-test.sh`.
3. Update `CHANGELOG.md` Unreleased → vX.Y.Z section.
3. **Update the docs this release makes stale — BEFORE you tag.** Rename
`CHANGELOG.md`'s `## Unreleased` to `## vX.Y.Z — YYYY-MM-DD` (em dash, as
every prior release heading uses), then run the gate:
```bash
bash scripts/check-doc-drift.sh # 0 in sync / 1 drift / 2 cannot run
```
It compares README.md's version-pin table against the ARGs it names, and
DOCKER_HUB.md's Node claim against `ARG NODE_VERSION`, plus Hub's
25 000-char limit, unsubstituted `{{PLACEHOLDERS}}`, and stale `Unreleased`
pointers in user-facing docs.
**Why before and not after:** `docker-publish.yml` runs `actions/checkout@v4`
with no `ref:`, so every job reads `github.ref` — the **tag**. A doc fix
pushed to `main` after tagging does not reach the release, and for
`DOCKER_HUB.md` it does not reach the published Hub page either, because
`update-description` POSTs that file as Docker Hub's `full_description` from
the tag's tree. Getting it in afterwards means re-pointing the tag, which is
its own hazard (v1.8.14 went `601fc98` → `361babd` and broke deploy
verification until `git fetch --tags --force`).
The gate is deliberately narrow — it only checks claims verifiable from files
in this repo. Still eyeball, because these are NOT gated:
- counts and sizes (`~1.1 GB`, "N `mempalace_*` tools", "7 extensions") —
they need a running image; assert them in `scripts/smoke-test.sh` instead
- feature prose that quietly became false, e.g. a "Planned for an upcoming
release" section describing something that already shipped
- `Dockerfile.base`'s `# BASE_REBUILD_DATE:` marker. Ungated on purpose:
`base_tag` hashes Dockerfile.base's content, comments included, so
demanding it be current would force a ~60 min base rebuild on a release
that touched no base files. **Fix it when the base is already rebuilding —
then it is free.**
Measured cost of skipping this, 2026-09-10 (v1.9.0): five stale claims, one
of them published. README's pin table was wrong on all three rows, and
DOCKER_HUB.md — untouched for eight releases — still said Node v22 while the
image shipped Node 24.
4. Verify `docker compose up` works locally with the current `latest` image
if you're upgrading users from a previous version. Then run the
**post-recreate sanity check** inside the running container to confirm
@@ -243,8 +280,10 @@ shipped the same image bytes); preventatively fixed for `PI_VERSION` +
image. Verifies binaries, repo clones, runtime deployment (waits for
keybindings + mempalace bridge + ≥4 extensions before sampling — fixes
the parallel-build-load race documented in opencode-devbox c6f9d11
2026-06-08), and image size threshold (3500 MB; revisit after a few
releases as actuals settle).
2026-06-08), build-time leftovers (see below), and image size threshold
(3800 MB in `SIZE_THRESHOLD_MB`; revisit after a few releases as actuals
settle — this doc said 3500 until 2026-09-11, after the bar had already
moved twice).
If smoke fails on size threshold but build is otherwise fine: bump
`SIZE_THRESHOLD_MB` in scripts/smoke-test.sh in a follow-up commit and
@@ -252,6 +291,27 @@ re-run. The threshold exists to catch *runaway* growth (an accidental
texlive bake-in, a forgotten chrome dependency), not to block ordinary
upstream bumps.
**The size gate is not a substitute for naming the residue.** It carries
~225 MB of deliberate margin, so v1.9.1 shipped +131 MB of pure build
residue — 110 MB of it npm's own download cache under `/root/.npm`, the
rest foreign platform packages — and stayed green. Four named assertions
now cover that ground: no foreign npm-11 platform packages beyond the
host arch (`@esbuild/*`, `@mariozechner/clipboard-*`), no `/root/.npm` in
the image, and — because the prune's real risk is *removing something
needed*, not size — esbuild must compile TS and clipboard must load its
native binding at **every** install site.
Two failure shapes to copy from those, both of which bit here:
- `test ! -d /root/.npm` on mode-700 `/root` passes for a **permission**
error, so the cache assertion refuses to run as non-root. Watch for
this in any assertion about a path you may not be allowed to read.
- `node -e 'require("esbuild")'` resolves by walking up from the CURRENT
DIRECTORY, so it fails with `MODULE_NOT_FOUND` from `/workspace` on a
perfectly healthy image (esbuild is nested inside the pi trees;
`NODE_PATH` is unset). Always path-qualify: `require("<abs>/esbuild")`.
A runbook shipped the bare form with "if this fails, revert the
release" attached, and it duly went red for the wrong reason.
## Build pipeline notes
- **Two-phase**: base + variant. Base is rebuilt only when
+653
View File
@@ -11,6 +11,659 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
---
## v1.9.2 — 2026-09-14
**The v1.9.1 residual is attributed and fixed: it was mostly npm's own download
cache, not the platform binaries it looked like.** v1.9.1 pruned the 26 foreign
`@esbuild/<platform>` directories that npm 11 installs, which fixed the 431 MB
size-gate failure — but the image still shipped **+131 MB compressed over
v1.8.14**, nearly all of it in the single pi/extensions install layer (87 MB →
206 MB). v1.9.1's notes recorded the leftover as an open item with an explicit
hypothesis (the `@mariozechner/clipboard-*` family, same npm 11 behaviour, a
different package) and an explicit warning that the hypothesis was **not** a
measured cause. It was measured on 2026-09-11, after recreating onto v1.9.1, and
the hypothesis accounted for only a sixth of it:
| item | v1.8.14 | v1.9.1 | delta |
|---|---|---|---|
| `/root/.npm/_cacache` — the build's npm download cache | 35.2 MB | 145.3 MB | **+110 MB** |
| `@mariozechner/clipboard-*` foreign platform packages (2 sites) | 0 MB | 21.1 MB | **+21 MB** |
| variant install layer, uncompressed total | 270.8 MB | 401.9 MB | +131 MB |
That is the whole delta with no unexplained remainder. Both items are now
deleted in the **same layer** that creates them, in both the main install RUN and
the studio RUN:
- **`purge_build_caches`** — `npm cache clean --force` plus `rm -rf /root/.npm`.
npm 11 caches every platform tarball it downloads, including the ones the
prune then deletes, so the cache grew far faster than the installed tree.
Nothing at runtime reads it: the build runs as root, the container runs as
`developer` with its own cache under `$HOME`.
- **`prune_foreign_esbuild` → `prune_foreign_natives`** — now covers both
measured families. For clipboard the keep-set is `clipboard-linux-$arch-gnu`
**and** `-musl`, because its napi-rs loader chooses between them at runtime
from its own `isMusl()` probe; the musl package is a 420-byte stub, so keeping
it is free insurance. The bare `@mariozechner/clipboard` wrapper has no
hyphen suffix and cannot match the pattern.
**Verified on arm64 before writing the patch, which is why the order was
update-then-patch:** a widened `rm -rf` glob is the worst possible change to
write against a tree you cannot inspect, and there is no docker CLI inside the
container — but after a recreate the container *is* the image. The prune was
exercised against a copy of the real trees with foreign directories fabricated
back in (aix-ppc64, android-arm64, darwin-arm64, win32-x64, linux-x64): all
removed, host `linux-arm64` kept at both sites, 21 MB freed, and
`require('@mariozechner/clipboard')` still loads and exports all 18 functions.
`esbuild.transformSync` still compiles TS at both sites. v1.9.1's own arm64
validation of the esbuild prune also passed here — CI could only smoke-test
amd64.
**Two sentinel assertions in `smoke-test.sh`, because the size gate did not
catch this.** The gate has ~225 MB of deliberate margin, so 131 MB of pure
build residue stayed green. There are now named PASS/FAIL checks for *foreign
platform packages beyond the host arch* and for *`/root/.npm` being shipped* —
the latter deliberately refuses to run as non-root, because `test ! -d
/root/.npm` on mode-700 `/root` would otherwise pass for the wrong reason. The
size-failure diagnostics now also list cache paths: they previously enumerated
only `node_modules` and `/opt`, where these bytes were not.
**Also fixed: the prune's own progress line was mangled.** `-printf '%f\\n'`
reaches the shell with both backslashes (confirmed from the published image's
recorded `created_by`), so `find` emitted a literal backslash and `tr` then ate
the `n` out of the name — v1.9.1 printed `esbuild platform dirs kept: li
ux-arm64`. Single backslash now.
Deliberately **not** changed: `/tmp/node-compile-cache` (1.3 MB). The manifest
RUN at the end of `Dockerfile.variant` calls `pi --version` again, so deleting
it earlier only relocates those bytes into that layer — today's manifest layer
is 128 kB precisely because it finds the cache warm.
**Two functional assertions as well, after the runbook command left for the
next machine failed for the wrong reason.** v1.9.1's open item prescribed
`node -e 'require("esbuild").transformSync(...)'` as the post-boot check, with
"if this fails, the prune removed something needed → revert to v1.8.14". Run
from `/workspace` it fails with `MODULE_NOT_FOUND` on a perfectly good image:
`require` resolves by walking up from the current directory, esbuild lives
nested inside the two pi trees, and global installs are not on node's require
path (`NODE_PATH` is unset). The check that verified the prune last time only
passed because the shell happened to be inside the tree. Smoke now does it
properly and CI owns it: for every install site found in the image (so the
studio variant's third site is covered automatically), esbuild must compile TS
and `@mariozechner/clipboard` must load with its native binding attached — the
latter is the real proof for the clipboard prune, since napi-rs resolves the
platform package at `require()` time. Both were verified as a four-way matrix:
green on the real image *from `/workspace`*, and red against copies of the same
packages with the host platform binary removed (`The package
"@esbuild/linux-arm64" could not be found`).
Touches `Dockerfile.variant` and `scripts/smoke-test.sh` only: `Dockerfile.base`
is unchanged, so this needs no base rebuild and should **ride the next release**
rather than burn a cycle of its own.
> **Superseded by the entry below:** that entry refreshes the vendored mempalace
> skill snapshot, which **is** hashed into `base_tag`. The release as a whole now
> costs a base rebuild (~67 min). The *size* work above still needs none of its
> own; the two simply travel together now.
**A behaviour change reached the fleet without any release naming it, and a
paragraph in these notes kept saying it had not.** v1.9.1 bakes
`mempalace-toolkit` **`e68ee20`**, which contains **`e2b060a`** — requester-side
ask withdrawal (`isWithdrawn`, RFC 003 §3.3 clause 4). So the behaviour has been
live on every v1.9.1 device since 2026-09-10, while the v1.9.0 section of this
file still read "not yet pinned … this image still pins `e45f6b4`" and the
mempalace skill still told every agent, at session start, that a withdrawal is
impossible: *"there is nothing anyone can do about it from the other end."*
The mechanism is the point, because it will do this again. `Dockerfile.variant`
carries `ARG MEMPALACE_TOOLKIT_REF=main` and `docker-publish.yml` resolves it to
a commit SHA at build time (`gitea_sha mempalace-toolkit`). A release therefore
absorbs *whatever toolkit `main` holds at that moment*, and "what behaviour did
this image gain?" is a question **nobody is structurally forced to answer**. This
fleet already has the rule — a floating ref that pulls a behaviour change into
the image must be named in the CHANGELOG *before* tagging. It was honoured for
the feed-tick fix, which v1.9.1 names explicitly (`309980b`, `e68ee20`), and
missed for the commit sitting in the same range.
Measured before being written, two independent routes, expectation recorded
first ("label should read ≥ `e68ee20`, since the build at 22:00Z postdates that
commit's 18:58Z"):
| route | result |
|---|---|
| Docker Hub config-blob label, `:v1.9.1-studio` and `:latest-studio` (same digest) | `se.jordbo.pi-devbox.mempalace-toolkit-ref = e68ee2071ca3ad39…` |
| `git merge-base --is-ancestor e2b060a e68ee20` | ancestor — the fix is inside the baked ref |
| baked `extensions/pi/mempalace.ts` sha256 vs v1.8.14's | `7c16fe14…` vs `dfca71e9…` — different bytes, so not the pre-fix file |
| `grep -c isWithdrawn` on this v1.8.14 container's baked copy | `0` — confirms the split, and that tor-ms22 cannot exercise it |
The first attempt at that label read **empty**, and the empty result was a claim
about the request, not the image: Docker Hub redirects blob fetches to a CDN and
`curl` without `-L` returns 0 bytes with exit 0. A registry audit that reports
"no labels" should be assumed to be missing `-L` until proven otherwise.
So the skill is updated rather than deferred (skillset **`e9e45f7`**, vendored
here with `scripts/vendor-mempalace-skill.sh`, `44472 → 46045 B`). Two things
were deliberate:
- **The "silence is not an answer" rule keeps its teeth.** An ask still stays
owed until the *recipient's* terminal event; what is new is a release by the
**asker**, explicitly marked. Stated that way round on purpose — the wrong
reading of this change is "withdrawals happen, so I need not reply".
- **The precondition ships with the rule**, because this skill is read on images
that lack the behaviour (tor-ms22, right now):
`grep -c isWithdrawn /opt/mempalace-toolkit/extensions/pi/mempalace.ts`, where
`0` means the withdrawal will not reach the recipient's mailbox. Same shape as
the provenance bullet's live-bridge check.
**The snapshot canary is re-pinned, and this time it fails on the old bytes
instead of merely failing to notice them.** The retired pair (`"Diaries
self-heal…"` present / `"Agent diaries live in"` absent) was still green against
the new snapshot, so it was blind to this refresh exactly as the pre-v1.8.13 pair
was blind to that one. The replacement is stronger than any predecessor here
because **both witnesses come from the same upstream commit**: `e9e45f7` added
`"Withdrawing an ask you sent"` and deleted `"nothing anyone can do about it from
the other end"`, the sentence the new bullet contradicts. Directions were
measured against both files rather than read off the diff (`new=1/old=0` and
`new=0/old=1`), then the canary body was **executed** against each: new → `rc=0
ok`, old → `rc=1` empty.
**`isWithdrawn` is no longer deployed-and-unproven — and the suite that pins it
had been dark since the node 24 bump.** The rule was exercised on the released
image against the live logstream on 2026-09-14, from a container recreated onto
v1.9.1 (born 13:21:41Z, confirmed by entrypoint-written mtimes and docker-written
`/etc` files agreeing to the second; `/proc/uptime` and `ps -o lstart=` were not
used, per their retraction). Baked toolkit `e68ee20`, `grep -c isWithdrawn` = 3,
mempalace.ts sha256 `7c16fe14…`, and the file pi actually loads verified to be
that same path and hash rather than a stale copy.
The instrument matters as much as the result: the shipped extractor was lifted
out of `scripts/test-owed-withdrawal.sh` and used to cut `TERMINAL_STATUS`,
`isStrictlyAfter`, `isAnswered` and `isWithdrawn` out of the **baked**
mempalace.ts by brace matching, then `deriveOwed`'s three queries were replayed
with their exact shipped parameters over the same `/mcp` transport the extension
uses. Shipped bytes, live data, no second copy of the logic. Baseline owed = 1,
agreed by three independent routes (the extension's own wake-up card; a hand
derivation of 23 candidates; the shipped predicates). Every expectation was
recorded before its measurement:
| probe | predicted | observed |
|---|---|---|
| positive control planted | owed 1 → 2 | 2 |
| requester withdraws it | back to 1 | 1, and `isAnswered=false isWithdrawn=true` |
| **third party** retracts someone else's ask | no effect | still owed |
| requester withdraws in **prose**, no marker | no effect | still owed |
| cleanup by ordinary replies | owed == baseline | 1, then 0 |
The count returning to baseline is only arithmetic; the per-predicate verdict is
what makes it a statement about mechanism. Probe A remained a raw candidate
throughout and no terminal reply of ours existed on its correlation, so its
removal is attributable to the withdrawal rule alone. Final attribution: A
cleared by `isWithdrawn` only, the two negative controls cleared by `isAnswered`
only — each probe retired by the predicate it was built to exercise. Both marker
spellings now have live witnesses (`metadata.withdraws` naming the ask's event id,
and naming its correlation). Unplanned and worth more than the probes: against
live data the rule also fires on the incident it was written for — seq 112 is
reported withdrawn-by-requester, i.e. mbp-m1-2020's seq 119 withdrawal now works,
so the 41h false obligation cannot recur. Full record in the coordination log at
`project/pi-devbox` seq 140; harness preserved as artifact
`art_20260914T141704_a7a9a294a4bb`.
Not proven, and deliberately not claimed: the extension's **own in-process
mailbox poll** surfacing a withdrawn ask. That poll fires at `agent_settled` when
the agent is idle; two probes were left owed across the rest of the session to
give it a window and it did not fire. Calling that confirmed would be a claim
about the session's patience, not about the code.
**Toolkit pickup `dab989b` — the owed-set suite could not run on this image, and
its own gate was why.** Reaching for the suite as corroboration exposed a second
defect: its precondition line `node --experimental-strip-types --check "$SRC"`
exits 2 on an unmodified mempalace.ts under node 24, and with `set -euo pipefail`
that skipped all 17 assertions and both regression guards. The cause is narrower
than the error suggests — it names the inline type-import, but `node --check` does
not type-strip at all: a file containing only `const x: number = 1` fails
identically, while executing the same file works. So the gate could never
validate TypeScript on any node; v1.9.1's bump from v22.23.2 to v24.21.0 is the
most likely trigger, though with no node 22 on the box that half stays labelled
inference rather than measurement. It failed **closed** — loud exit 2, never
vacuously green — which is the good direction, and the reason it went unnoticed is
that nothing in CI runs this suite.
That matters more than a red test, because this suite is the only thing that makes
`isWithdrawn`'s failure mode visible: a wrong rule there does not throw and does
not log, it makes a real unanswered ask vanish from a mailbox forever. `dab989b`
strips first and syntax-checks the emitted JS, and splits the exit codes so that
`3` means the gate cannot run while `2` means the source does not parse —
collapsing those is how the defect disguised itself as "mempalace.ts does not
parse" while mempalace.ts was fine. Verified in six directions with expectations
written first: clean source `rc=0` 17/17; malformed TypeScript `rc=2`; stripper
made unavailable `rc=3` without the misleading message; and the four mutation
kills back at e2b060a's counts of 3/1/2/1, so sensitivity is restored rather than
asserted.
> **Cost, and it is smaller than it looks:** `MEMPALACE_TOOLKIT_REF` is resolved
> by CI to the head of the toolkit's `main` and folded into the `base_tag` hash —
> verified at `docker-publish.yml:126-128`, whose own comment gives the reason
> ("otherwise a toolkit-only fix never lands"). So `dab989b` moves `base_tag` and
> the next tag rebuilds the base, with nothing to remember to trigger. But it adds
> no rebuild that was not already owed: `9aaff26` refreshed the vendored skill
> snapshot under `rootfs/`, which is also hashed into `base_tag`, so a base
> rebuild has been pending since before this fix existed. The toolkit pickup rides
> along with it, and the same rebuild is what finally bakes skillset `e9e45f7` and
> turns the snapshot canary above green against the image's own floor.
---
## v1.9.1 — 2026-09-10
**v1.9.0 was tagged but never published: its own smoke gate stopped it, and it
was right to.** `build-base` succeeded, then `smoke` failed 90-passed/3-failed,
and because `build-variant` needs `smoke`, both variants, `promote-base-latest`
and `update-description` were skipped. No image reached the registry, so
`latest` still pointed at v1.8.14. v1.9.1 carries everything listed under v1.9.0
below, plus the three fixes here. Two of the three failures were self-inflicted
by v1.9.0's own changes, and the third was a real regression that the Node bump
dragged in — which is the case for keeping the gate strict.
**Failure 1 — the image was 431 MB over its size threshold, and npm 11 was the
cause.** Node 22 → 24 brings npm 10 → 11, and npm 11 installs **every**
`@esbuild/<platform>` optional binary rather than only the one matching the host:
26 platform directories covering aix-ppc64, android, darwin, freebsd, netbsd,
openbsd, win32, s390x, riscv64 and more, none of which this image can execute.
Measured on pi-fork's dependency tree, same repo and same command:
| npm | packages | `node_modules` |
| --- | --- | --- |
| 10.9.8 | 136 | **165 MB** |
| 11.19.0 | 169 | **449 MB** |
The 165 MB figure reproduces exactly what v1.8.14 shipped, which is what
identified npm rather than the image as the variable. esbuild declares those
binaries with `os`/`cpu` constraints, but npm 11 ignores them — and also ignores
`--os`/`--cpu` flags and an `.npmrc` carrying `os=`/`cpu=` (all three measured,
all three still produced 26 directories). So `Dockerfile.variant` now prunes
explicitly, keeping only `linux-$(node -p process.arch)` so one line is correct
on amd64 and arm64. Verified this removes dead weight and not function: after
pruning, `esbuild.transformSync` still compiles TypeScript. The prune runs in the
**same layer** as each `npm install` — deleting in a later `RUN` would leave the
bytes in the earlier layer and shrink the image by nothing. Three sites are
covered: the global pi install, pi-fork, and pi-studio (which pulls its own
pi-coding-agent copy), for roughly 548 MB recovered in the non-studio variant and
822 MB in studio. The threshold stays at 3800 MB deliberately: it caught a real
regression, and raising it to accommodate one would have discarded the signal.
**Failure 2 — the om `node_modules` assertion was checking an npm artefact, not
the software.** `pi-observational-memory` declares **zero** runtime dependencies:
8 devDependencies (omitted by `--omit=dev`) and 4 peerDependencies, which pi
itself provides. npm 10 still materialised a `node_modules` for it, but that
directory contained exactly **one file** (`.package-lock.json`, 4 KB) and no
nested `package.json` — 20 empty scope directories. npm 11 stopped creating it,
so `test -d node_modules` went red while nothing about om had changed or broken.
The assertion now checks what must actually hold — that the entry point pi loads
exists — read out of the manifest pi itself reads (`package.json` →
`pi.extensions`) rather than a hardcoded path that could drift. pi-fork keeps its
`node_modules` check, because pi-fork has real dependencies where the directory's
absence would mean something.
**Failure 3 — the skill-source annotation broke the assertion that reads it.**
v1.9.0 taught `pi-devbox-version` to say *which* pi-extensions copy shipped
(`baked (package copy)`, or a loud FALLBACK/MIXED marker). The smoke assertion
matched `^ $s +baked$`, anchored at the end, so the annotation failed it even
though the state reported was correct. The pattern now allows an optional
` (...)` suffix, matched loosely on purpose: *which* copy shipped is already
asserted authoritatively against the manifest field and its measured tree hash,
and re-encoding that wording in a second regex would just add a second place to
update. The lesson recorded rather than the fix alone: the display branches were
tested in an isolated harness that passed, but the assertion **consuming** them
was never run — harness-passes-therefore-consumer-passes was an assumption.
**A red size assertion now carries its own diagnostic.** Attributing the 431 MB
took a full CI-log dig plus a local npm bisect, while the container knew where
its bytes were the whole time. On failure the check now prints the largest
layers, the largest directories, and a count of `@esbuild` platform directories
as a sentinel for this exact regression recurring — the same principle the `run()`
helper already applies to every other assertion.
No `Dockerfile.base` or `rootfs/` change, so the base fingerprint is untouched
and `base-decide` reuses `base-0fb1256c7f99` built during the v1.9.0 attempt.
**A gate for documentation drift, because five claims rotted in one release and
one of them was published.** Preparing v1.9.0 turned up a cluster of stale
facts, all the same shape — a value written once by hand, in a file nothing
verifies, about a number that lives somewhere else and moved:
- `README.md`'s "Version pins" table was wrong on **all three rows**: pi
`0.84.4` vs `ARG PI_VERSION=0.85.1`, pi-atelier `v0.10.0` vs `v0.10.1`,
mempalace `3.8.0` vs `3.9.0`. That table is the worst possible place for this,
because it exists *specifically* to be the reviewable record of what the repo
freezes deliberately — so a wrong row destroys the only thing it is for.
- `README.md` listed already-shipped typst PDF export under "Planned for an
upcoming minor release", carrying the self-contradicting marker "(shipped in
Unreleased/base)". The **fourth** instance of the stale-`Unreleased`-pointer
class this changelog already documented three of.
- `DOCKER_HUB.md` claimed "Node.js v22" while v1.9.0 ships Node 24.
The last one is why this became a gate rather than a resolution to be careful.
`DOCKER_HUB.md` is **published**: `update-description` POSTs it to Docker Hub as
`full_description` on every tag. It had gone **eight releases** (v1.8.6 →
v1.9.0) without a touch. Nothing generates it — CI only substitutes
`{{PI_VERSION}}` — and nothing checked it, so the sole mechanism keeping it true
was whoever remembered. Worse, it is read from the **tag**, so the stale page
published with v1.9.0 anyway and the fix could only ride the next release.
**New: `scripts/check-doc-drift.sh` + a `doc-drift` job in `lint.yml`.** Seven
checks, all comparing a doc string to a value that exists in this repo, so it
needs no network, no token, no built image, and no sibling clone:
- README's three pin-table rows vs the ARGs they name *by name*
- `DOCKER_HUB.md`'s Node claim vs `ARG NODE_VERSION`
- placeholders CI will not substitute — the publish step greps for leftovers of
`{{PI_VERSION}}` only, so any *second* token sails through and publishes
literally
- `DOCKER_HUB.md` under Docker Hub's 25 000-char `full_description` limit
(previously discoverable only as a non-200 *after* the full build)
- `Unreleased` appearing in a user-facing doc, which is always a pointer that
outlived what it pointed at
Exit codes match `lint-shell.sh` and `check-skill-floor.sh`: `0` in sync, `1`
drift, `2` cannot run — a renamed ARG makes the gate blind, which is a red `2`,
never a green tick. Verified with **15 controls**: every check fails when its
claim is broken, the real v1.9.0 Node bug is caught, and two false-positive
controls pass — the first version of the placeholder check wrongly flagged
`README.md:900`'s `docker inspect --format '{{json .Config.Labels}}'`, a Go
template in a legitimate example, so the pattern is now anchored to the
UPPER_SNAKE convention CI actually substitutes. **The gate was wrong, not the
doc** — which is the whole reason a gate gets negative controls.
Deliberately **not** gated, and the reasons matter more than the list:
- Counts and sizes (`~1.1 GB`, "N `mempalace_*` tools", "7 extensions") need a
running image. A gate that cannot evaluate a claim honestly would have to
guess, and a guessing gate is worse than none — assert these in
`scripts/smoke-test.sh`, where a real image exists.
- `Dockerfile.base`'s `# BASE_REBUILD_DATE:` marker, itself stale (2026-07-13,
three base rebuilds ago). `base_tag` hashes Dockerfile.base's *content*,
comments included, so demanding it be current would force a ~60 min base
rebuild on a release that touched no base files at all. It is free to fix
while the base is *already* rebuilding, and expensive at any other moment.
That cost asymmetry is now written into the release checklist rather than
enforced.
**Release checklist step 3 rewritten** (`AGENTS.md`) around the mechanism that
made this expensive: `docker-publish.yml` runs `actions/checkout@v4` with no
`ref:`, so every job reads `github.ref` — the tag. Docs must be correct *before*
tagging; afterwards the only routes are re-pointing the tag (its own hazard —
v1.8.14 went `601fc98` → `361babd` and broke deploy verification until
`git fetch --tags --force`) or waiting for the next release. The step now also
names what the gate cannot see, so "gate is green" is not mistaken for "docs are
true". The same reflex went into the `ci-release-watcher` skill, as the first
correctness rule — it is the only one that expires once the tag exists.
Also fixed in passing: README's `pi-devbox-version` sample was v1.5.0-era and
structurally outdated (it predated the `palace:` line the surrounding prose
advertises, the `pi-atelier` component, and the whole `skills:` block). Replaced
with real observed output rather than hand-written text. `DOCKER_HUB.md`'s "7
user-facing extensions" was **verified correct**; its "29 `mempalace_*` tools"
is stale (a live client shows 45) but left alone rather than corrected on a
guess, since that count cannot be attributed to the baked 3.9.0 server without
measuring it.
## v1.9.0 — 2026-09-10 (tagged, never published — superseded by v1.9.1)
> This tag exists in git but no image was ever pushed for it: `smoke` failed
> three assertions and skipped every downstream job. Everything below ships in
> **v1.9.1**, whose entry explains the three failures and their fixes. Kept as
> its own section rather than folded away, because the tag is real and someone
> will eventually find it and wonder why Docker Hub has no v1.9.0.
**`shellcheck` is now in the image, because the release gate it depends on could
not be run by anyone.** v1.8.14 made shell lint a release gate: `scripts/lint-shell.sh`
became the single source of truth for `lint.yml` and a new `lint-gate` job that
`resolve-versions` depends on, and it deliberately exits 2 when `shellcheck` is
absent — *a gate that cannot run must not pass*. Measured on v1.8.14 on
2026-09-09, by three routes (`command -v`, `dpkg -l`, a filesystem search):
**`shellcheck` was not in the image at all.** So `bash scripts/lint-shell.sh`
exited 2 in every devbox container, and the only place the gate could ever run
was CI. The developer loop was therefore write-shell → push → wait for CI →
discover — which is the loop the gate was added to shorten, after v1.8.14's first
attempt burned ~46 minutes on a tree whose lint had already been red for 24
hours. Added to the `apt-get` block in `Dockerfile.base`: `shellcheck 0.10.0-1`,
~39 MB installed (`Installed-Size` 40112 KB), and measured to pull **zero**
additional packages under `--no-install-recommends` because its three deps
(`libc6`, `libffi8`, `libgmp10`) are already present. **This forces one full base
rebuild** — `base-decide` hashes `Dockerfile.base` + `rootfs/`, so unlike a
`scripts/` change it cannot reuse the existing `base-` layer.
**A client-side pre-push lint gate: `hooks/pre-push`.** Opt-in per clone with
`git config core.hooksPath hooks`, bypass with `git push --no-verify`, matching
the idiom the `skillset` and `myconfigs` repos already use. It is a thin wrapper
that `exec`s `scripts/lint-shell.sh` — the same script CI runs, one copy, because
a duplicated check that drifts is the failure this repo keeps paying for (the
`pi-extensions` skill mirror sat 9579 B behind for weeks; the shell-lint logic
was extracted to one file for exactly this reason).
> **Why this repo had no hooks at all, which is worth stating because it was
> reported as drift and is not.** A fleet peer asked `tor-ms22` to report
> `git config core.hooksPath` per clone on the premise that an unset value meant
> "no secret-scan and no shell-lint hook locally", leaving the drift/secret gates
> unverified. Measured: `pi-devbox` **unset**, `skillset` `hooks`, `myconfigs`
> `common/hooks`, `pi-toolkit` **unset**. But `git ls-files | grep -i hook` is
> **empty** in both `pi-devbox` and `pi-toolkit` — neither repo tracked a single
> hook file, so there was nothing for `core.hooksPath` to point at on any machine
> and unset was the only correct value. The two repos that do ship hooks were
> already wired correctly. This entry closes the real half of that gap for
> `pi-devbox`; `pi-toolkit` still ships none.
**The hook is verified to catch the defect that motivated it, not merely to
exist.** Three measurements, each with the expected result written down first:
- **Refusal paths.** With `shellcheck` absent (the state of every container built
before this change) the hook exits **2** and names the remedy; with
`scripts/lint-shell.sh` missing it also exits **2**. It never waves a push
through on the assumption that CI will catch it.
- **The hook is actually in the scan set.** `lint-shell.sh` reports `Checking 14
shell file(s)` with `hooks/pre-push` present and **13** with it moved aside — so
the extensionless file is discovered by the shebang half of the linter's
two-signal union, rather than being silently skipped. This check exists because
the first attempt at it was ambiguous: a planted `[ $UNSET_VAR = "x" ]` was not
reported, which could equally have meant "file not scanned" or "defect below
`-S error`". It was the latter. A count that moves is unambiguous; a clean run
is not.
- **It catches the real v1.8.14 defect.** Planting the exact failing shape — an
apostrophe inside a single-quoted string, `echo 'the fleet\'s thing'` — in
`hooks/pre-push` produces `SC1073`/`SC1072` at severity **error** and `rc=1`.
That is the defect that closed a string, truncated an `exec_test` body, sent its
tail to the runner's shell, and cost a 46-minute build.
**The gate earned its keep inside the commit that added it.** The first version of
the `smoke-test.sh` assertion above carried a comment beginning `# shellcheck is a
GATE DEPENDENCY…`. A comment whose first word is the tool's name is parsed as a
**shellcheck directive**, not a comment, so the new gate immediately failed with
`SC1073`/`SC1072` at severity error — on the change that introduced it. Same family
as the v1.8.14 apostrophe: a line that reads as prose to a human and as syntax to
the parser. Before this change that defect would have been discovered in CI.
**Also queued, not yet pinned:** the `mempalace-toolkit` owed-set derivation now
honours a requester withdrawing its *own* ask (`isWithdrawn`, RFC 003 §3.3
clause 4, with `scripts/test-owed-withdrawal.sh`) — toolkit commit **`e2b060a`**,
which is the minimum revision for the behaviour. This image still pins `e45f6b4`.
Until an image bakes `e2b060a` or later, a sender must assume its withdrawal has
no effect on the recipient's mailbox — measured cost of the gap: a withdrawn
v1.8.13 rollout ask was still being reported as owed on `tor-ms22` 41 hours later,
for a release that device never installed. The same commit also anchors the
derivation's `mine` query at the newest end (`order: "desc"`); with the previous
default `asc` + `limit: 100`, a device passing 100 authored events would have its
recent replies fall out of the join window and see answered asks resurface.
> **Corrected 2026-09-14 (`pi@tor-ms22`, the device in the measured cost above).**
> "This image still pins `e45f6b4`" and "until an image bakes `e2b060a` or later"
> were true when written on 2026-09-09 and are **false for the running fleet**.
> **v1.9.1 bakes `e68ee20`**, a descendant of `e2b060a`, so requester-side
> withdrawal is LIVE wherever v1.9.1 runs. Measured from the published image's own
> label (`se.jordbo.pi-devbox.mempalace-toolkit-ref`) rather than from these
> notes, and cross-checked by ancestry and by the baked `mempalace.ts` sha
> differing from v1.8.14's. Nothing here was mis-stated on purpose:
> `ARG MEMPALACE_TOOLKIT_REF=main` is resolved to a commit SHA by CI at build
> time, so the release absorbed the commit without anybody having to name it,
> while this paragraph went on asserting it had not. Left standing rather than
> rewritten — the sentence is the evidence for how the drift happened. See
> Unreleased.
**Four small packages, each chosen from a gap that was measured rather than
imagined.** All four were picked by looking back at a real session — the
`gitea.egl.lan`/FreeIPA debugging of 2026-09-09..10 — and asking which absences
actually cost time, not which tools sound useful. `bind9-dnsutils` (~6.1 MB, 10
packages): `dig`, `host` **and** `nslookup` were all absent, so the container
could resolve names but had no way to interrogate a *specific* nameserver —
`getent hosts` only follows the resolver's default path, so diagnosing "gateway
`172.16.88.1` NXDOMAINs the `egl.lan` zone while `10.20.253.1` is authoritative
for it" had to be hand-rolled in `python3`. Note the package name: plain
`dnsutils` is transitional in trixie. `ldap-utils` (1244 KB, **zero** extra deps):
the fleet authenticates against FreeIPA, yet every LDAP probe had to be run by
SSHing to an already-enrolled host; this gives simple binds only, since GSSAPI
would additionally need `krb5-user` + `libsasl2-modules-gssapi-mit`, which is a
Kerberos-client decision rather than a tool. `xxd` (198 KB) is frank convenience
— `od -c` already does the job. `python3-yaml` (552 KB, zero extra deps) is the
shellcheck story repeating exactly: `scripts/check-workflow-shell.sh`, the guard
against the Gitea `sh`/dash footgun that broke `resolve-versions` (`ed49b8d`) and
`promote-base-latest` (`b7197e8`), hard-exits with "python3 yaml module missing"
without it — and `lint.yml` installing it explicitly in CI was the evidence the
image lacked it. **`netcat-openbsd` was proposed and deliberately rejected**:
measured redundant, because `socat` is already baked and bash's `/dev/tcp` does
reachability checks with zero packages. The reason is recorded in
`Dockerfile.base` so the omission reads as a decision rather than an oversight.
**The vendored `pi-extensions` skill floor was 41 days stale, and is now gated so
it cannot silently rot again.** `rootfs/usr/local/share/pi-devbox/skills/pi-extensions/`
sat at 34284 B, untouched since `fa04d20` (2026-07-30), while the package copy
was 38973 B — four copies of one skill existed across the fleet with three
different sizes. `Dockerfile.variant` copies the freshly-cloned package copy over
the **served** path but never writes back to the repo floor, so nothing in the
repo ever noticed. That is worse than ordinary staleness because the floor is a
**fallback**: the copy is guarded by `if [ -f /opt/pi-extensions/skill/SKILL.md ]`,
so a build whose clone yields no `skill/` keeps the vendored snapshot and still
goes **green**, with no manifest flag and no label recording which copy was
served — the image would ship a July skill and nothing would say so. The floor is
refreshed here from `pi-extensions@c64c122`, and the new `skill-floor` job in
`lint.yml` runs `scripts/check-skill-floor.sh` to keep it that way.
The check compares the **directory** hash, using the same `tree_sha256` pipeline
`Dockerfile.variant` uses for `skillset_snapshot_tree_sha256` and for the same
documented reason: a `sha256sum SKILL.md` answers "did this one file change", not
"is this the same skill", and `pi-extensions` ships two files. That is not
hypothetical — it was **verified by negative control**: with `SKILL.md` left
byte-identical and only `evaluate-extension-usage.py` edited, the directory check
correctly fails while a file-only compare would have passed. Exit codes are `0`
in sync / `1` drift / `2` cannot-run, matching `scripts/lint-shell.sh`, so an
unreachable package repo is a red `2` rather than a green tick. Gating on another
repo is normally a smell; it is proportionate here because the check can only
fire when `skill/` itself changed — which is exactly when the floor has gone
stale — and it needs no secret, since `pi-extensions` is anonymously clonable
(verified with `git ls-remote` and no credentials).
> **What this does *not* fix, stated so nobody reads more into it than is there.**
> The floor is now fresh and guarded, but the *silent-fallback* half remains:
> if the build-time copy is ever absent, the build still succeeds with no
> manifest flag or OCI label recording that the vendored snapshot was served
> instead of the package copy. The durable fix for that is a manifest field
> alongside the existing `skillset_snapshot_tree_sha256`, which this change does
> not add.
**Four pinned dependencies bumped, after an audit of everything the image gets
from outside apt.** The audit itself is the useful part: of ~23 externally-managed
components, the 19 that resolve `latest` at build time were already current or
refresh themselves on the next rebuild, and the hard pins for `pi` (0.85.1),
`mempalace` (3.9.0) and `pi-atelier` (v0.10.1) were all already the newest
available. Only four needed a human.
**`NODE_VERSION` 22 → 24 (LTS "Krypton") — this one was a latent defect, not
housekeeping.** `agent-browser` publishes `engines.node ">=24.0.0"`, so the image
was *below a declared requirement*: v1.8.14 shipped node 22.23.2 with
`agent-browser` 0.37.1, meaning every build installed it with an npm `EBADENGINE`
warning and then ran the baked browser automation outside its supported range.
The other two npm consumers are satisfied either way — `pi` declares `>=22.19.0`,
`playwright` `>=20`. Verified before bumping, because a missing NodeSource suite
would break every architecture at once: `setup_24.x` returns HTTP 200 and the
`node_24.x` suite advertises `Architectures: amd64 arm64 armhf x86_64`, covering
both the arm64 fleet and the amd64 CI runners. Nothing else in the repo pinned the
node major.
**`actionlint` 1.7.7 → 1.7.12 and `hadolint` 2.14.0 → 2.15.1**, each run against
the current tree at the new version *before* being pinned — both clean, no new
findings. That ordering is the point: a linter bump is the one dependency update
that can turn CI red on unchanged code, so discovering it locally costs a minute
and discovering it in CI costs a round trip.
**`SKILLSET_SNAPSHOT_REF` `e9e09d9` → `4d7c0ea`**, via
`scripts/vendor-mempalace-skill.sh` rather than by hand, because that script is
the only thing that may write the ARG — a `cp` without a matching bump produces a
manifest that confidently lies. This turned out to be **provenance-only**: the
recorded ref was 6 commits behind, but `skills/mempalace/SKILL.md` is byte-identical
at both (`3675bfab…`), so the vendored snapshot was already correct and only its
recorded origin was stale. Consequently no `rootfs/` bytes changed, the
smoke-test phrase canary stays valid, and this ARG alone would not have forced a
base rebuild — the node bump does that anyway.
**The silent-fallback hole is closed: the image now records WHICH `pi-extensions`
skill copy it shipped.** This was the half deliberately left open by the
`skill-floor` gate above, and it is the more important half, because "the floor is
currently fresh" is a fact with a shelf life while "the image says which copy it
got" keeps working. The refresh step in `Dockerfile.variant` is guarded by
`if [ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates the
co-located skill kept the vendored floor and still succeeded **green**, with
nothing in the manifest, the labels or the logs distinguishing that from a normal
build. The two outcomes are indistinguishable by inspection afterwards — same
path, same filenames, same permissions — which is exactly how the floor went
unnoticed from 2026-07-30 to 2026-09-10.
`build-manifest.json` gains `pi_extensions_skill_source` and
`pi_extensions_skill_tree_sha256`, both **measured rather than passed in as
build-args**, per the ground-truth rule the rest of that block already follows —
and necessarily so here, since the outcome depends on the clone's contents and no
ARG could express it. Three values, because two would force a lie:
`package` (served bytes equal the clone's `skill/`), `vendored-floor` (the clone
had no `skill/` at this ref, so the fallback shipped), and `divergent` — both
exist but differ, e.g. the clone ships `SKILL.md` but not
`evaluate-extension-usage.py`, leaving the served directory a genuine **mix** of
package and floor. No OCI label mirrors these, deliberately: `LABEL` cannot take a
value computed in a `RUN`, and a label fed from an ARG would be precisely the
claim-not-measurement this change exists to remove.
Two `scripts/smoke-test.sh` assertions turn the record into a gate: one that the
source is named and is `package` — `vendored-floor` **fails** rather than warns,
since these images track `main` where the package has co-located `skill/` since
`fa04d20`, so a fallback means the clone did not resolve as intended — and one
that recomputes the tree hash over the served directory, because a recorded hash
that is never recompared is a claim rather than a measurement. `pi-devbox-version`
also annotates the line: `pi-extensions baked (package copy)` on the normal path,
and a yellow `(FALLBACK: vendored floor — clone had no skill/)` otherwise. Its
existing skill section reports which copy is being **read** at runtime; this is
the one fact that is decided at **build** time and cannot be recovered later.
Older images degrade cleanly — the field is absent, `jq // empty` yields nothing,
and the line prints plain `baked` exactly as before.
**This image also carries a real fix for the recurring
`[mempalace ext] feed (tick) failed: mine timed out after 30000ms` message** that
has been appearing in the pi TUI across the fleet since August
(`mempalace-toolkit` `309980b` + `e68ee20`, picked up because CI resolves
`MEMPALACE_TOOLKIT_REF` to a commit SHA at build time). It was parked as cosmetic
on 2026-08-27 and it was not cosmetic: `lastFeedAt` was recorded only after a
*successful* wait, but the extension's `Promise.race` abandons only the **wait**
and cannot cancel the mine, so a timeout left the 10-minute debounce clock stale
— and with `feedInFlight` already cleared, **both** guards stood open and every
following settled turn started another mine on top of the one still running.
Overlapping writers on a single-writer palace, each making the next slower and
the next timeout likelier, which is why the message appeared many times per
session instead of at most once per debounce window. Simulated over ten minutes
of settled turns with a 60s mine: **16 mines launched, 15 of them overlapping**
before; **2 and 0** after. Nothing was ever lost — the transcript is staged
before the mine and `mine --mode convos` is idempotent — so this was wasted work
and a misleading error, not data loss. The deadline also rose from 30s to 5
minutes: the mine is the slowest call the extension makes (30–60s normally) yet
carried the tightest deadline, 4x tighter than the `prepare` before it and 10x
tighter than the init handshake. On a healthy fleet the message should now be
absent; if it appears it is informative — a mine exceeding five minutes.
---
## v1.8.14 — 2026-09-08
> **First release attempt failed; fixed in this same entry.** The `smoke` and
+1 -1
View File
@@ -94,7 +94,7 @@ The entrypoint deploys/registers all of these on first container start. Re-runni
uv run --with jupyterlab jupyter lab --no-browser --port 8888
uv run --with marimo marimo edit
```
- **Node.js** v22 + npm (used by pi itself)
- **Node.js** v24 LTS + npm (used by pi itself)
- **Rust** — `rustup-init` is on PATH; install toolchains on demand
- **Go** — opt-in via `--build-arg INSTALL_GO=true` if rebuilding from source
+90 -1
View File
@@ -101,6 +101,80 @@ ENV DEBIAN_FRONTEND=noninteractive
# container: `ss` lands at /usr/bin/ss, `ip` at /usr/sbin/ip
# (both already on the developer PATH), and `portcheck --all`
# then correctly identifies the socat listener on 8765.
# shellcheck — shell linter. Added 2026-09-09 to close a CAPABILITY gap, not
# a style preference. `scripts/lint-shell.sh` is the release
# GATE (the `lint-gate` job that `resolve-versions` depends
# on), and it correctly refuses to pass when shellcheck is
# missing — "a gate that cannot run must not pass". Measured on
# v1.8.14: shellcheck was absent from this image by all three
# routes (PATH, dpkg, filesystem), so `bash
# scripts/lint-shell.sh` exited 2 in EVERY devbox container and
# no developer could run the release gate locally at all. The
# loop was therefore write-shell → push → wait for CI → discover,
# which is the loop the gate was added to shorten: v1.8.14's
# first attempt burned ~46 min on a tree whose lint had already
# been red for 24 h. This is also what makes a client-side
# pre-push hook possible (see hooks/pre-push); without the
# binary that hook would refuse every push. ~39 MB installed
# (Installed-Size 40112 KB, shellcheck 0.10.0-1) and measured
# to pull ZERO additional packages under
# --no-install-recommends: its deps (libc6, libffi8, libgmp10)
# are already present. NOTE this file feeds the base-decide
# hash (Dockerfile.base + rootfs/), so adding it forces one
# full base rebuild.
# bind9-dnsutils — `dig` and `nslookup`. Added 2026-09-10 to close a
# DIAGNOSTIC gap measured during the gitea.egl.lan/FreeIPA
# work: the container could resolve names but had NO way to
# ask a SPECIFIC nameserver anything. `getent hosts` only
# follows the resolver's default path, so the whole "gateway
# 172.16.88.1 returns NXDOMAIN for the egl.lan zone while
# 10.20.253.1 is authoritative for it" diagnosis had to be
# hand-rolled in python3 — dig, host AND nslookup were all
# absent. `dig @10.20.253.1 freeipa-4.egl.lan` is the
# one-liner that replaces it, and split-horizon DNS is a
# recurring class of bug on this fleet, not a one-off. NOTE
# the package to name is bind9-dnsutils: plain `dnsutils` is
# a transitional package in trixie. ~6.1 MB total (6210 KB
# measured): bind9-dnsutils 721 KB + bind9-host 161 KB +
# bind9-libs 3804 KB plus 7 small libs (libfstrm0,
# libjson-c5, liblmdb0, libmaxminddb0, libprotobuf-c1,
# liburcu8t64, libuv1t64) under --no-install-recommends.
# ldap-utils — `ldapsearch`/`ldapmodify`. Added 2026-09-10. This fleet
# authenticates against FreeIPA (EGL.LAN), and every LDAP
# probe during the Gitea auth work had to be run by SSHing to
# an already-enrolled host because the container had no LDAP
# client at all. 1244 KB and pulls NOTHING extra under
# --no-install-recommends — its deps (libldap, libsasl2) are
# already present. CAVEAT: this gives SIMPLE binds only,
# which is what Gitea itself uses and what most probes need.
# GSSAPI binds (`ldapsearch -Y GSSAPI`) additionally require
# krb5-user + libsasl2-modules-gssapi-mit, deliberately NOT
# added here — that is a Kerberos-client decision with
# /etc/krb5.conf implications, not just a tool.
# xxd — hex dump. 198 KB, no extra deps. Convenience, and honestly
# marginal: `od -c` from coreutils is always present and does
# the same job. Earned its place because verifying that
# git-crypt actually encrypted a staged blob (the \0GITCRYPT\0
# magic) is a recurring check in myconfigs and xxd is the
# muscle-memory command for it.
# NOT added — netcat-openbsd (133 KB): measured redundant on
# 2026-09-10, because socat is already baked above AND bash's
# /dev/tcp does reachability checks with zero packages
# (verified against gitea.egl.lan:3000). Recorded here so the
# omission reads as a decision rather than an oversight.
# python3-yaml — PyYAML. Added 2026-09-10 for precisely the same reason as
# shellcheck above: a gate this repo ALREADY OWNS could not be
# run locally by anyone. scripts/check-workflow-shell.sh — the
# guard that catches the "bash-only syntax under Gitea's default
# sh/dash shell" footgun that broke resolve-versions (ed49b8d)
# and promote-base-latest (b7197e8) — hard-exits with "ERROR:
# python3 yaml module missing" without it. lint.yml installs it
# explicitly in CI (`shellcheck python3-yaml`), which is itself
# the evidence that the image lacked it. Measured 2026-09-10
# while wiring the skill-floor job: the guard could not be run
# before pushing — the same write → push → wait-for-CI loop that
# shellcheck was baked to shorten. 552 KB, and pulls ZERO extra
# packages under --no-install-recommends.
RUN apt-get update && \
apt-get upgrade -y --no-install-recommends && \
apt-get install -y --no-install-recommends \
@@ -120,6 +194,7 @@ RUN apt-get update && \
make \
patch \
diffutils \
shellcheck \
git-crypt \
age \
file \
@@ -141,6 +216,10 @@ RUN apt-get update && \
kitty-terminfo \
ncurses-term \
iproute2 \
bind9-dnsutils \
ldap-utils \
xxd \
python3-yaml \
&& ln -s /usr/bin/fdfind /usr/local/bin/fd \
&& apt-get clean \
&& rm -rf /var/lib/apt/lists/*
@@ -554,7 +633,17 @@ ENV COLORTERM=truecolor
ENV PATH="/home/developer/.local/bin:/home/developer/.cargo/bin:${PATH}"
# ── Node.js (required for pi + MCP servers + tldr) ──
ARG NODE_VERSION=22
# 24 (LTS "Krypton"), raised from 22 on 2026-09-10 because the image was BELOW a
# DECLARED requirement, not merely behind the newest release: `agent-browser`
# publishes engines.node ">=24.0.0", so every build on 22 installed it with an npm
# EBADENGINE warning and then ran it outside its supported range — measured on
# v1.8.14, which shipped node 22.23.2 with agent-browser 0.37.1. The other two npm
# consumers are satisfied either way: pi declares ">=22.19.0" and playwright
# ">=20". Verified before bumping, because a missing NodeSource suite would break
# the build for every arch at once: deb.nodesource.com/setup_24.x returns HTTP 200
# and the node_24.x suite advertises `Architectures: amd64 arm64 armhf x86_64`, so
# both the arm64 fleet and the amd64 CI runners resolve.
ARG NODE_VERSION=24
RUN curl -fsSL --retry 5 --retry-delay 5 --retry-all-errors https://deb.nodesource.com/setup_${NODE_VERSION}.x | bash - && \
apt-get install -y --no-install-recommends nodejs && \
rm -rf /var/lib/apt/lists/*
+130 -1
View File
@@ -196,6 +196,71 @@ RUN set -e && \
done; \
return 1; \
} && \
# prune_foreign_natives: npm 11 (shipped with Node 24) installs EVERY optional
# platform package of a native dependency, not just the one matching the host.
# TWO families are affected in this image, and BOTH have been measured — add a
# family here only after measuring it, never by widening the pattern on a hunch:
#
# @esbuild/<platform> 26 dirs, 284 MB (found first, v1.9.1)
# @mariozechner/clipboard-<triple> 11 dirs, 12 MB per site, 10 MB foreign
#
# esbuild declares those with os/cpu constraints, but npm 11 ignores them and
# ALSO ignores --os/--cpu and an npmrc carrying os=/cpu= (all three measured).
# So prune explicitly, keeping only the host platform, computed from
# `node -p process.arch` so one line stays correct on amd64 and arm64.
# Measured on pi-fork's tree: npm 10.9.8 -> 165 MB, npm 11.19.0 -> 449 MB,
# and the 165 MB figure reproduces what v1.9.0's predecessor actually shipped.
#
# WHY THE CLIPBOARD FAMILY WAS ADDED (2026-09-11): v1.9.1 pruned @esbuild only
# and still shipped +131 MB compressed over v1.8.14. That residual was
# attributed by listing the PUBLISHED arm64 layer tarballs straight from the
# registry (there is no docker CLI inside the container, so `docker history`
# was not available): +110 MB /root/.npm/_cacache (purged below) and +21 MB of
# clipboard platform packages across the two install sites — 131 MB total, so
# the delta is now fully accounted for with no unexplained remainder.
#
# Keeping linux-$arch-{gnu,musl} is deliberate: clipboard's napi-rs loader
# tries ./<name>.node then the platform package, per platform in try/catch, and
# chooses gnu vs musl at runtime from its own isMusl() probe — so both host-arch
# branches must survive. The musl package is a 420-byte stub, i.e. free. The
# bare wrapper `@mariozechner/clipboard` has no hyphen suffix and therefore
# cannot match the regex below. Verified on arm64 against a copy of the real
# tree before this was written: after pruning to those two,
# require('@mariozechner/clipboard') still loads and exports all 18 functions.
# esbuild likewise still compiles TS via transformSync at both install sites.
# This removes dead weight, not function.
#
# MUST run in the SAME layer as the npm installs above: deleting in a later RUN
# leaves the bytes in this layer and shrinks the image by nothing.
# NOTE the single backslash in -printf '%f\n': Docker passes '\\n' through
# verbatim, so v1.9.1's doubled version printed a mangled "li ux-arm64"
# (find emitted a literal backslash, then `tr` translated the n out of the
# name). Confirmed from the published image's own recorded created_by.
prune_foreign_natives() { \
arch="$(node -p process.arch)"; \
find /usr/lib/node_modules /opt -type d -regex '.*/@esbuild/[^/]+' \
! -name "linux-$arch" -prune -exec rm -rf {} + ; \
find /usr/lib/node_modules /opt -type d -regex '.*/@mariozechner/clipboard-[^/]+' \
! -name "clipboard-linux-$arch-gnu" ! -name "clipboard-linux-$arch-musl" \
-prune -exec rm -rf {} + ; \
echo "native platform dirs kept: $(find /usr/lib/node_modules /opt -type d \( -regex '.*/@esbuild/[^/]+' -o -regex '.*/@mariozechner/clipboard-[^/]+' \) -printf '%f\n' 2>/dev/null | sort | uniq -c | tr '\n' ' ')"; \
} && \
# purge_build_caches: the build's own download caches are NOT free — they land
# in whichever layer created them. Measured on the published v1.9.1 arm64
# variant layer: root/.npm/_cacache was 145.2 MB of a 401.9 MB layer (35.2 MB
# in v1.8.14), the single biggest item in the +131 MB residual, because npm 11
# caches every platform tarball it fetched — including the ones just pruned.
# Nothing at runtime reads it: the build runs as root, the container runs as
# `developer` with its own cache under $HOME (and $HOME/.pi is a volume).
# DELIBERATELY NOT purged here: /tmp/node-compile-cache (1.3 MB, written by
# `pi --version` below). The manifest RUN at the end of this file calls
# `pi --version` again, so deleting it here only relocates those bytes into
# that layer instead of removing them from the image — measured, not assumed:
# today the manifest layer is 128 kB precisely because it finds the cache warm.
purge_build_caches() { \
npm cache clean --force >/dev/null 2>&1 || true; \
rm -rf /root/.npm; \
} && \
if [ "${PI_VERSION}" = "latest" ]; then \
NPM_CONFIG_PREFIX=/usr npm install -g @earendil-works/pi-coding-agent ; \
else \
@@ -209,6 +274,8 @@ RUN set -e && \
git_fetch_ref "${PI_ATELIER_REPO}" "${PI_ATELIER_REF}" /opt/pi-atelier && \
(cd /opt/pi-fork && npm install --omit=dev --no-audit --no-fund) && \
(cd /opt/pi-observational-memory && npm install --omit=dev --no-audit --no-fund) && \
prune_foreign_natives && \
purge_build_caches && \
echo "pi-toolkit at $(cd /opt/pi-toolkit && git rev-parse --short HEAD)" && \
echo "pi-extensions at $(cd /opt/pi-extensions && git rev-parse --short HEAD)" && \
echo "pi-fork at $(cd /opt/pi-fork && git rev-parse --short HEAD)" && \
@@ -309,6 +376,26 @@ ARG PI_STUDIO_REF=main
ARG PI_STUDIO_VERSION=none
RUN if [ "${INSTALL_STUDIO}" = "true" ]; then \
set -e; \
# Same prune + cache purge as the main install RUN — see the comments there.
# They have to be redefined because shell functions do not survive across
# layers, and they have to run in THIS layer because pi-studio's npm install
# happens here: deleting in a later RUN would leave the bytes in this layer
# and shrink nothing. pi-studio pulls its own pi-coding-agent copy, so it is
# a third ~274 MB site on top of the two in the non-studio variant — and its
# npm install refills /root/.npm, which the main RUN emptied in ITS layer.
prune_foreign_natives() { \
arch="$(node -p process.arch)"; \
find /usr/lib/node_modules /opt -type d -regex '.*/@esbuild/[^/]+' \
! -name "linux-$arch" -prune -exec rm -rf {} + ; \
find /usr/lib/node_modules /opt -type d -regex '.*/@mariozechner/clipboard-[^/]+' \
! -name "clipboard-linux-$arch-gnu" ! -name "clipboard-linux-$arch-musl" \
-prune -exec rm -rf {} + ; \
echo "native platform dirs kept: $(find /usr/lib/node_modules /opt -type d \( -regex '.*/@esbuild/[^/]+' -o -regex '.*/@mariozechner/clipboard-[^/]+' \) -printf '%f\n' 2>/dev/null | sort | uniq -c | tr '\n' ' ')"; \
}; \
purge_build_caches() { \
npm cache clean --force >/dev/null 2>&1 || true; \
rm -rf /root/.npm; \
}; \
rm -rf /opt/pi-studio && mkdir -p /opt/pi-studio && \
git -C /opt/pi-studio init -q && \
git -C /opt/pi-studio remote add origin "${PI_STUDIO_REPO}" && \
@@ -320,6 +407,8 @@ RUN if [ "${INSTALL_STUDIO}" = "true" ]; then \
done; \
[ "$ok" = "1" ] && \
(cd /opt/pi-studio && npm install --omit=dev --no-audit --no-fund) && \
prune_foreign_natives && \
purge_build_caches && \
echo "pi-studio at $(cd /opt/pi-studio && git rev-parse --short HEAD)"; \
fi
@@ -392,7 +481,7 @@ ARG MEMPALACE_TOOLKIT_REF=main
# no ~67-minute base rebuild. (scripts/check-base-hash.sh scans only
# Dockerfile.base, so no folding into the base hash is required — nor would
# it be correct, since this ARG changes nothing about the base's contents.)
ARG SKILLSET_SNAPSHOT_REF=e9e09d95f92670536a199fc986dfa24d787f18d1
ARG SKILLSET_SNAPSHOT_REF=e9e45f7acdde490c3b5d24ce5f508bff8785c2c7
# Dockerfile.base sets description="pi-devbox — base image (variant-independent)"
# and every variant INHERITS it, so both published images used to advertise
@@ -473,6 +562,44 @@ RUN set -e; \
if [ -d "$_snap_dir" ] && [ -n "$(find "$_snap_dir" -type f -print -quit)" ]; then \
SKILL_SNAP="\"$(tree_sha256 "$_snap_dir")\""; \
fi; \
# ── WHICH pi-extensions skill copy actually shipped ──
# Closes the silent-fallback hole. The refresh step above is guarded by
# `[ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates
# the co-located skill (or a fork pointing at a mirror without it) keeps the
# vendored floor and still succeeds — GREEN, with nothing anywhere recording
# that a snapshot shipped instead of the package copy. Measured 2026-09-10:
# the floor had been stale since 2026-07-30, so that fallback would have
# shipped a six-week-old skill silently. The floor is fresh now and gated by
# the skill-floor CI job, but "the fallback is currently harmless" is not the
# same as "you can tell which copy you got", and only the second survives.
#
# MEASURED, never claimed, per the ground-truth rule above: the branch
# condition is re-derived from the same test the refresh step used, and the
# served bytes are then compared against the clone. A build-arg could not
# express this at all, since the outcome depends on the clone's contents.
# package served bytes == the clone's skill/ (the normal path)
# vendored-floor the clone has no skill/ at this ref (fallback shipped)
# divergent both exist but differ — e.g. the clone ships SKILL.md but
# not evaluate-extension-usage.py, so the served directory is
# a MIX of package and floor. Worth its own value: it is the
# one state neither of the other two names honestly.
# No OCI label mirrors this, deliberately: LABEL cannot take a value computed
# in a RUN, and a label fed from an ARG would be exactly the claim-not-
# measurement this block exists to avoid.
_px_dir=/usr/local/share/pi-devbox/skills/pi-extensions; \
PIEXT_SRC='null'; PIEXT_HASH='null'; \
if [ -d "$_px_dir" ] && [ -n "$(find "$_px_dir" -type f -print -quit)" ]; then \
PIEXT_HASH="\"$(tree_sha256 "$_px_dir")\""; \
if [ -f /opt/pi-extensions/skill/SKILL.md ]; then \
if [ "$(tree_sha256 "$_px_dir")" = "$(tree_sha256 /opt/pi-extensions/skill)" ]; then \
PIEXT_SRC='"package"'; \
else \
PIEXT_SRC='"divergent"'; \
fi; \
else \
PIEXT_SRC='"vendored-floor"'; \
fi; \
fi; \
{ \
echo '{'; \
echo " \"release_tag\": \"${RELEASE_TAG}\","; \
@@ -493,6 +620,8 @@ RUN set -e; \
# vendored skill directory, not one file — see tree_sha256() above.
echo " \"skillset_snapshot_ref\": \"${SKILLSET_SNAPSHOT_REF}\","; \
echo " \"skillset_snapshot_tree_sha256\": ${SKILL_SNAP},"; \
echo " \"pi_extensions_skill_source\": ${PIEXT_SRC},"; \
echo " \"pi_extensions_skill_tree_sha256\": ${PIEXT_HASH},"; \
echo " \"components\": {"; \
echo " \"pi-toolkit\": \"$(rev /opt/pi-toolkit)\","; \
echo " \"pi-extensions\": \"$(rev /opt/pi-extensions)\","; \
+24 -19
View File
@@ -175,12 +175,10 @@ Currently published:
| `joakimp/pi-devbox:latest-studio` | `latest` + [pi-studio](https://github.com/omaclaren/pi-studio) (browser prompt editor, KaTeX/Mermaid preview, tmux-backed literate REPLs) | ~3.25 GB |
| `joakimp/pi-devbox:vX.Y.Z-studio` | pinned-version studio equivalent | ~3.25 GB |
Planned for an upcoming minor release:
- *(shipped in Unreleased/base)* **PDF export from Studio/pandoc** now works:
the base image ships **`typst`** as the PDF engine (`pandoc --pdf-engine=typst`),
a single ~30 MB static binary — no separate `-tex` variant needed.
`texlive-xetex` stays the higher-fidelity fallback (install on demand).
Both variants ship **`typst`** as the pandoc PDF engine
(`pandoc --pdf-engine=typst`), a single ~30 MB static binary, so PDF export from
Studio/pandoc works out of the box — no separate `-tex` variant needed.
`texlive-xetex` stays the higher-fidelity fallback (install on demand).
## Using pi-studio (`-studio` variant)
@@ -919,16 +917,23 @@ through `jq` yourself:
```console
$ pi-devbox-version
pi-devbox v1.5.0
built: 2026-07-13T17:53:16Z (source d68674d11e06)
pi: 0.80.6
pi-devbox v1.8.14
built: 2026-09-08T21:54:07Z (source 361babd4fd61)
pi: 0.85.1
palace: 3.9.0
components:
pi-toolkit: 9a8f6faeaa08
pi-extensions: 61c98e004e3d
pi-fork: 4a09af4ef527
pi-observational-memory: 27a5195eaf90
mempalace-toolkit: 96699f2a1781
pi-studio: 2ef38ef31cea
pi-toolkit: adfb553f5c8a
pi-extensions: 2610545c83bb
pi-fork: e69725c39603
pi-observational-memory: ce9fc982b3a2
pi-atelier: 734258bbcb62
mempalace-toolkit: e45f6b430181
pi-studio: e04fc7aa3275
skills:
credential-incident-response baked
mempalace live /workspace/skillset @ 4d7c0ea (identical to baked snapshot)
pi-devbox-environment baked
pi-extensions baked
```
It also flags **live drift** — if `pi --version` no longer matches what was
@@ -1093,7 +1098,7 @@ persisted volumes survived, and pi runtime wiring is intact:
```bash
./scripts/recreate-sanity-check.sh # auto-detects variant
./scripts/recreate-sanity-check.sh --expected-image-version 1.8.9 # assert the pi-devbox release tag
./scripts/recreate-sanity-check.sh --expected-version 0.84.4 # assert the pi coding agent version
./scripts/recreate-sanity-check.sh --expected-version 0.85.1 # assert the pi coding agent version
```
Those are **two different versions**, and the flags are not interchangeable:
@@ -1132,9 +1137,9 @@ resolved to `latest` at build time:
| Component | Pin | Where |
|---|---|---|
| pi | `0.84.4` | `ARG PI_VERSION` — `Dockerfile.variant` |
| pi-atelier | `v0.10.0` | `ARG PI_ATELIER_REF` — `Dockerfile.variant` |
| mempalace | `3.8.0` | `ARG MEMPALACE_VERSION` — `Dockerfile.base` |
| pi | `0.85.1` | `ARG PI_VERSION` — `Dockerfile.variant` |
| pi-atelier | `v0.10.1` | `ARG PI_ATELIER_REF` — `Dockerfile.variant` |
| mempalace | `3.9.0` | `ARG MEMPALACE_VERSION` — `Dockerfile.base` |
The objective is **not** to freeze versions. Bumping is routine — usually one
line plus a changelog note. The objective is that adopting a new upstream
Executable
+64
View File
@@ -0,0 +1,64 @@
#!/usr/bin/env bash
# Pre-push gate for pi-devbox: shellcheck every shell script before it leaves
# this clone. Thin wrapper — all logic lives in scripts/lint-shell.sh, which is
# the SAME script the CI release gate runs. One copy, not two: a duplicated
# check that drifts is the failure this repo keeps paying for.
#
# Install per clone: git config core.hooksPath hooks
# Bypass this gate: git push --no-verify (a guard, not a wall)
#
# WHY THIS HOOK EXISTS
# v1.8.14's first release attempt died at scripts/smoke-test.sh:770 after
# build-base had already spent ~46 minutes. shellcheck had ALREADY caught the
# defect — SC2289 at severity error, on the very push that introduced it — and
# the lint job stayed red for 24 hours, unread, across three runs. The fix at
# the time was to gate the release on the same script (the `lint-gate` job).
# This hook is the cheaper end of that: the same finding, before the push,
# in seconds rather than after a 40 s CI gate or a 46 min build.
#
# WHY IT COULD NOT EXIST UNTIL NOW
# Measured on v1.8.14 (2026-09-09): shellcheck was absent from the devbox
# image by all three routes — PATH, dpkg and a filesystem search. So
# lint-shell.sh exited 2 in every container, and a hook calling it would have
# refused EVERY push rather than gating anything. `shellcheck` was added to
# Dockerfile.base in the same change that added this file; on an image built
# before that, enable this hook and you will simply be told the gate cannot
# run. That is the correct behaviour, but it is not a working hook — so do not
# set core.hooksPath on a container older than the release that bakes it.
#
# NOTE ON SCOPE: this lints the WORKING TREE, not the exact commit range being
# pushed. That is deliberate and matches what the CI gate does to the tagged
# tree. It means a defect you have staged-but-not-committed is also reported,
# which is noisy in the safe direction.
set -euo pipefail
HOOK_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
REPO_ROOT="$(cd "$HOOK_DIR/.." && pwd)"
LINTER="$REPO_ROOT/scripts/lint-shell.sh"
tag="[lint-shell]"
# Same rule the gate itself applies, applied one level up: a missing check is
# not a pass. If the script is gone, the push is refused rather than waved
# through on the assumption that CI will catch it.
if [ ! -r "$LINTER" ]; then
echo "$tag refusing the push: $LINTER is missing, so the gate cannot" >&2
echo "$tag run. A gate that cannot run must not pass." >&2
exit 2
fi
# Point the message at the actual remedy when the binary is absent, because the
# linter's own message ("install it or run this in CI") is written for a CI
# runner and is misleading inside a container the developer cannot apt-install
# into persistently.
if ! command -v shellcheck >/dev/null 2>&1; then
echo "$tag refusing the push: shellcheck is not installed, so the gate" >&2
echo "$tag cannot run. A gate that cannot run must not pass." >&2
echo "$tag" >&2
echo "$tag This container predates the image that bakes shellcheck." >&2
echo "$tag Either recreate onto an image that has it, or unset the hook:" >&2
echo "$tag git config --unset core.hooksPath" >&2
echo "$tag To push this once without the gate: git push --no-verify" >&2
exit 2
fi
exec bash "$LINTER" "$REPO_ROOT"
+27 -1
View File
@@ -157,6 +157,11 @@ if [ "$SHOW_SKILLS" = "yes" ] && [ -d "$BAKED_SKILLS" ] && [ -d "$SKILLS_DIR" ];
# is not hypothetical.
snap_ref=$(jq -r '.skillset_snapshot_ref // empty' "$MANIFEST")
snap_sha=$(jq -r '.skillset_snapshot_tree_sha256 // empty' "$MANIFEST")
# Which pi-extensions copy the BUILD baked. Distinct from everything else in
# this section, which reports which copy is being READ at runtime: for
# pi-extensions the baked tree is itself one of two possible copies, and that
# choice was made at build time and is not recoverable by inspection.
px_src=$(jq -r '.pi_extensions_skill_source // empty' "$MANIFEST")
# Same pipeline Dockerfile.variant uses to measure the baked directory at
# build time: relative paths in `find | sort` order, each hashed, the whole
# listing folded into one sha256. Keep the two definitions identical — they
@@ -187,7 +192,28 @@ if [ "$SHOW_SKILLS" = "yes" ] && [ -d "$BAKED_SKILLS" ] && [ -d "$SKILLS_DIR" ];
_target=$(readlink -f "$_link" 2>/dev/null || echo "$_link")
case "$_target" in
"$BAKED_SKILLS"/*|"$BAKED_SKILLS")
printf ' %-22s baked\n' "$_name"
# "baked" alone used to be the whole story. For pi-extensions it is not:
# the baked tree holds EITHER the package copy that Dockerfile.variant
# lays over the snapshot, OR the vendored floor, when the clone had no
# skill/ at that ref. The two are indistinguishable by inspection — same
# path, same filenames, same permissions — so the build records which one
# it used and this reports it. Without this line a six-week-stale
# fallback skill looks exactly like a current one, which is precisely how
# the floor went unnoticed from 2026-07-30 to 2026-09-10.
if [ "$_name" = "pi-extensions" ] && [ -n "$px_src" ]; then
case "$px_src" in
package)
printf ' %-22s baked (package copy)\n' "$_name" ;;
vendored-floor)
printf ' %-22s baked \033[33m(FALLBACK: vendored floor — clone had no skill/)\033[0m\n' "$_name" ;;
divergent)
printf ' %-22s baked \033[33m(MIXED: part package, part floor)\033[0m\n' "$_name" ;;
*)
printf ' %-22s baked\n' "$_name" ;;
esac
else
printf ' %-22s baked\n' "$_name"
fi
continue
;;
esac
@@ -535,7 +535,10 @@ Two consequences worth internalising:
- **"Seen, not doing it" is a legitimate ack** — `status="blocked"` or
`"superseded"` plus the reason. Silence is not, and it is not merely rude:
with no terminal event of yours to join to, the ask stays in the owed set
indefinitely and there is nothing anyone can do about it from the other end.
indefinitely. The *original requester* — and nobody else — can release it from
the other end, but only by saying so explicitly: see **Withdrawing an ask you
sent** below. That is a release by the asker, not an escape for the answerer.
While the ask still stands, only *your* terminal event clears it.
- **Nothing expires, and it should not.** An `open` with no terminal reply is
still live by definition, and the finished threads are valuable history. If
content is genuinely perishable ("do not push to main for the next hour"), say
@@ -570,6 +573,24 @@ Two consequences worth internalising:
be matched to it at all.
- **Corrections are new events, never edits.** Say explicitly what you retract
and name the id — drawer or event — that carried the withdrawn claim.
- **Withdrawing an ask you sent: state it, never imply it.** Your release only
counts when the terminal event (a) comes from the same `from_agent` that sent
the ask, (b) is directed at that recipient exactly — never `*`, so a broadcast
can neither oblige nor release, (c) carries a terminal status (`claimed` and
`ready` are not terminal and do not release anything), (d) is strictly after
the ask, (e) joins it via `ack_of` or the same `correlation_id`, **and (f)
names that ask in `metadata.withdraws` or `metadata.closes`.** Prose in the
body does not count, and neither does a bare terminal event on the
correlation: inferring release from *any* terminal would let your own
bookkeeping silently delete a real obligation, so the release must be stated.
Needs toolkit ≥ `e2b060a` (image ≥ `v1.9.1`) — check with
`grep -c isWithdrawn /opt/mempalace-toolkit/extensions/pi/mempalace.ts` and
read `0` as "my withdrawal will have no effect on their mailbox". Measured
cost of getting it wrong: a `v1.8.13` rollout ask was withdrawn by its sender,
who recorded it as done; the recipient's derivation never saw the release and
still reported the ask owed **41 hours later**, for a release that device
never installed — and the asymmetry was invisible from the sender's side
(RFC 003 §3.3 clause 4).
- **Put a retraction where the reader will look.** A *directed open ask* reaches a
live agent's mailbox; a **terminal-status event reaches no mailbox at all**, and
a *drawer* is what a future semantic search finds. If you filed advice as a
@@ -1,7 +1,7 @@
---
name: pi-extensions
description: >-
Use the pi extensions (pi-fork, pi-observational-memory, ssh-controlmaster) effectively in the pi coding agent harness. Load this skill only when running inside pi (detection - `fork` and `recall` are present in your tool list, or `pi --ssh` was used to start the session). pi-fork dispatches focused subtasks to forked agents at fast/balanced/deep effort tiers; pi-observational-memory compacts long sessions into recallable observations + reflections; ssh-controlmaster rewires pi's read/write/edit/bash tools to execute on a remote host over a multiplexed SSH connection. This skill covers tier selection, task design, boundary discipline, when to use recall, and remote-pi mechanics.
Use the pi extensions (pi-fork, pi-observational-memory, ssh-controlmaster) effectively in the pi coding agent harness. Load this skill only when running inside pi (detection - `fork` and `recall` are present in your tool list, or `pi --ssh` was used to start the session). pi-fork dispatches focused subtasks to forked agents at fast/balanced/deep effort tiers; pi-observational-memory compacts long sessions into recallable observations + reflections; ssh-controlmaster rewires pi's read/write/edit/bash tools to execute on a remote host over a multiplexed SSH connection. Also covers the context ladder L0-L4 and when to reach for the separate `pi-task` CLI instead of `fork` - isolated child, immutable spec, machine-checked envelope, write-boundary diff. This skill covers tier selection, task design, boundary discipline, when to use recall, and remote-pi mechanics.
---
# Pi Extensions: pi-fork, pi-observational-memory, ssh-controlmaster
@@ -161,6 +161,68 @@ The "three" things it completed were exactly the main thread's pending todos, vi
- Distrust **quantities** and **provenance claims** in fork prose specifically ("all N sessions", "shipped with the image", "as expected") — those are the slots confabulation fills.
- The fact that the fork was "right anyway" is not the same as the fork having followed instructions.
### The context ladder — and the second dispatch mechanism (`pi-task`)
Everything above describes a child that inherits everything. That is not a fixed
cost of delegation — **how much context a child gets is a choice**, and `fork`
sits at one extreme of it. Five rungs:
| rung | what the child sees | mechanism | built? |
|---|---|---|---|
| **L0** | nothing but the goal | `pi-task` default: fresh `--session-id pitask-<id>-<stamp>` in a private `--session-dir` | yes |
| **L1** | goal + **names** of files/commands to read itself | `pi-task` spec `context.files` / `context.commands` (`bin/pi-task:154,157`) | yes |
| **L2** | goal + an **excerpt the parent curated** | `pi-task` spec `context.facts`, pasted verbatim (`bin/pi-task:151`) | yes |
| **L3** | a **truncated tail** of the parent branch | *nothing implements this* — would need a new spec key plus `--session <trimmed snapshot>` | **no** |
| **L4** | the **entire** parent branch | `fork(task=…)` — `getHeader()+getBranch()`, no offset or limit anywhere in the call chain | yes |
**`pi-task` is a CLI, not an extension — it will never appear in your tool list.**
Invoke it with `bash`: `/opt/pi-toolkit/bin/pi-task run <spec.json>` (source at
`/workspace/pi-toolkit/bin/pi-task`, `schema` subcommand prints the spec fields).
It reads an immutable JSON spec, and "inherit the session" is not expressible in
that schema — the isolation is structural, not a request.
**Choose the lowest rung that can do the job:**
- **`fork` (L4)** when the subtask only makes sense against this conversation,
when you want several independent opinions in parallel from one message, or for
read-only exploration whose detail you will discard. Everything in "Boundary
discipline" above applies in full.
- **`pi-task` (L0–L2)** when the brief contains a **prohibition** (the inherited
transcript is exactly what overrides those), when you want a **pass/fail**
result instead of prose, when you need an **audit trail**, or when writes
outside an authorised set must be caught.
- **Neither** for trivial work, iterative work (both are one-shot), or judgement
that needs context only you have.
**What `pi-task` gets you that no brief can.** The envelope must parse or the run
FAILED, however fluent the prose. `roots[]` is the WATCHED set and
`write_allowed` the CHANGEABLE subset, diffed before and after with git
`--porcelain --ignored`. That `--ignored` flag is load-bearing: in the T4 test the
child obeyed its brief perfectly and still tripped the detector, because
`py_compile` wrote `__pycache__` into a watched-but-not-writable root — a
gitignored path that plain `--porcelain` reports as clean. Note the structural
point that test exposed: under `read_only: true` a write is *defiance*, so a
well-behaved child never produces a delta and the detector is never exercised.
Splitting WATCHED from WRITABLE is what lets an **obedient** child reveal a
violation, which is the realistic hazard.
**What it does not fix.** `--no-extensions` removes extensions, not the core
`read`/`write`/`edit`/`bash` tools — exactly as described above — so the boundary
diff is post-hoc **detection, not prevention**. And a fresh L0 context removes the
*narrative* failures (parent voice, invented continuity) without removing
confabulation: given an under-specified spec built on a false premise, the child
still filled the `deliverable` slot with a confident shape. The envelope's own
structure creates that pressure. Verify decisive claims from the filesystem
regardless of which rung you used.
**Trap — the capability floor is inverted from intuition.** `runner.ts:188` reads
`if (extensions !== null) args.push("--no-extensions")`. So `pi-fork.extensions:
[]` passes the flag and the floor is **on**; setting it to `null` — documented in
`settings.json` as the way to "restore normal extension loading" — passes nothing
and the floor is **off**, restoring palace writes inside every fork child.
Changing `[]` to `null` as a tidy-up re-arms what was deliberately disarmed.
`pi-task` hardcodes the flag and cannot drift this way.
### Anti-patterns
- **Forking trivial work.** A fork has overhead. If the task takes < 30 seconds in your main thread, just do it.
@@ -230,7 +292,7 @@ When entries conflict, **the most recent observation reflects the latest known s
## Quick Reference
```
fork(task=..., effort=fast|balanced|deep)
fork(task=..., effort=fast|balanced|deep) # L4: child inherits your WHOLE branch
- state decision authority explicitly
- pass verified context up front
- specify deliverable shape
@@ -240,6 +302,13 @@ fork(task=..., effort=fast|balanced|deep)
- write-capable? demand "What I did NOT do", then verify from git/fs, not the report
- prohibition in the brief => not a `fast` task
bash: /opt/pi-toolkit/bin/pi-task run <spec> # L0-L2: isolated child, NOT a tool
- schema | selftest | run [--dry-run]
- context.facts (pasted) / .files (names only) / .commands
- roots[] = WATCHED, write_allowed[] = CHANGEABLE subset
- envelope must parse or the run FAILED
- audit + cost: ~/.pi/agent/pi-task/<stamp>-<id>/result.json
recall(id=<12-char-hex>)
- only when stakes justify the cost
- id must already be visible in your context
+246
View File
@@ -0,0 +1,246 @@
#!/usr/bin/env bash
# check-doc-drift.sh — fail when a hand-maintained doc claim contradicts the
# build files it describes.
#
# THE DEFECT CLASS THIS EXISTS TO CATCH, measured 2026-09-10 while preparing
# v1.9.0. Five separate claims had rotted, all of them the same shape: a fact
# written once by hand, in a file nothing verifies, about a value that lives
# somewhere else and moved.
#
# 1..3. README.md's "Version pins" table was wrong on EVERY row — pi `0.84.4`
# vs ARG PI_VERSION=0.85.1, pi-atelier `v0.10.0` vs v0.10.1, mempalace
# `3.8.0` vs 3.9.0. That table is the worst possible place for this: it
# exists precisely to be the reviewable record of what is deliberately
# frozen, so when it lies, the review it enables is worthless.
# 4. README.md carried a "Planned for an upcoming minor release" section
# listing typst PDF export, which had ALREADY SHIPPED, tagged with a
# self-contradicting "(shipped in Unreleased/base)" marker. The
# CHANGELOG had already documented three earlier instances of exactly
# this stale-"Unreleased"-pointer class (see its v1.8.7 notes).
# 5. DOCKER_HUB.md claimed "Node.js v22" while this release ships Node 24.
# This one is the reason the gate exists at all: DOCKER_HUB.md is
# PUBLISHED. `update-description` in docker-publish.yml POSTs it to Hub
# as full_description on every tag, so unlike README.md — which no
# workflow or gate reads — a stale claim here is what users see.
#
# WHY A GATE AND NOT "REMEMBER TO CHECK". DOCKER_HUB.md had gone eight releases
# (v1.8.6 → v1.9.0) without a touch. Nothing generates it and nothing verifies
# it; the only mechanism keeping it true was whoever remembered. That is the
# same failure mode check-skill-floor.sh was written for, and the same fix:
# convert "someone remembers" into "CI refuses".
#
# WHY THESE FIVE CHECKS AND NOT MORE. Every check here compares a doc string to
# a value that EXISTS IN THIS REPO, so it can never be wrong about the world and
# needs no network, no token, and no built image. Claims that require a running
# container to verify (image sizes, the "N mempalace_* tools" count) are
# deliberately NOT gated: a check that cannot be evaluated honestly at lint time
# would either be skipped or guessed, and a guessing gate is worse than none.
# If you want those, assert them in scripts/smoke-test.sh where a real image is
# available.
#
# DELIBERATELY NOT GATED: Dockerfile.base's `# BASE_REBUILD_DATE:` comment, which
# is also stale (2026-07-13, three base rebuilds ago). base_tag is a hash of
# Dockerfile.base's CONTENT plus rootfs/, comments included, so a gate that
# demanded that comment be current would force a ~60 min base rebuild on any
# release that touched no base files at all. Fix it when you are already
# rebuilding the base — then it is free. This is a real cost asymmetry, not
# laziness.
#
# EXIT CODES (same contract as lint-shell.sh and check-skill-floor.sh):
# 0 every checked claim matches
# 1 at least one claim has drifted
# 2 cannot run (a file or ARG this gate reads is missing/unparseable)
# A gate that cannot run must not pass, so a missing input is 2, never 0.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
cd "$REPO_ROOT"
README="README.md"
HUB="DOCKER_HUB.md"
DF_VARIANT="Dockerfile.variant"
DF_BASE="Dockerfile.base"
# Docker Hub rejects a full_description longer than this. docker-publish.yml has
# no size check of its own; it only notices via a non-200 from the API, i.e.
# after paying the whole build. Catching it here makes it a 2-second failure.
HUB_MAX_CHARS=25000
WARN_ONLY=0
FAILURES=0
usage() {
cat <<'EOF'
Usage: check-doc-drift.sh [--warn-only] [-h|--help]
Compares hand-written claims in README.md and DOCKER_HUB.md against the build
files they describe (Dockerfile.base, Dockerfile.variant).
--warn-only Report drift but exit 0 (advisory use, e.g. a local pre-push hook).
Exit: 0 = in sync, 1 = drift, 2 = cannot run.
EOF
}
while [ $# -gt 0 ]; do
case "$1" in
--warn-only) WARN_ONLY=1; shift ;;
-h|--help) usage; exit 0 ;;
*) echo "::error::unknown argument: $1" >&2; usage >&2; exit 2 ;;
esac
done
for f in "$README" "$HUB" "$DF_VARIANT" "$DF_BASE"; do
if [ ! -f "$f" ]; then
echo "::error::$f not found (cwd $PWD). Cannot evaluate doc drift, so this is exit 2, not a pass."
exit 2
fi
done
# Read `ARG NAME=value` from a Dockerfile. Exit 2 when absent: if the ARG this
# gate is built around has been renamed, the gate is measuring nothing and must
# say so rather than silently comparing against an empty string.
read_arg() {
local file="$1" name="$2" value
value="$(sed -n "s/^ARG ${name}=\\(.*\\)\$/\\1/p" "$file" | head -1)"
if [ -z "$value" ]; then
echo "::error::ARG ${name} not found in ${file}. It was probably renamed;" >&2
echo "::error::update check-doc-drift.sh to match, because this gate is now blind." >&2
exit 2
fi
printf '%s' "$value"
}
# One row of README's "Version pins" table: `| pi | `0.85.1` | ... |`
read_pin_row() {
sed -n "s/^| $1 | \`\\([^\`]*\`*\\)\` |.*/\\1/p" "$README" | head -1
}
fail() {
FAILURES=$((FAILURES + 1))
echo "::error::$1"
}
ok() { printf ' OK %s\n' "$1"; }
echo "Checking hand-maintained doc claims against the build files they describe."
echo
# ---------------------------------------------------------------------------
# 1-3. README's version-pin table vs the ARGs it names by name.
# ---------------------------------------------------------------------------
check_pin() {
local label="$1" documented="$2" actual="$3" where="$4"
if [ -z "$documented" ]; then
fail "README.md: no '| $label |' row found in the version-pin table. Either the
table was restructured (update this gate) or the row was dropped (restore it)."
return
fi
if [ "$documented" != "$actual" ]; then
fail "README.md version-pin table is stale for $label: says '$documented',
$where says '$actual'. Fix the table — it is the reviewable record of what
this repo deliberately freezes, so a wrong row defeats its only purpose."
return
fi
ok "README pin $label = $actual"
}
PI_ACTUAL="$(read_arg "$DF_VARIANT" PI_VERSION)"
ATELIER_ACTUAL="$(read_arg "$DF_VARIANT" PI_ATELIER_REF)"
MEMPALACE_ACTUAL="$(read_arg "$DF_BASE" MEMPALACE_VERSION)"
check_pin pi "$(read_pin_row pi)" "$PI_ACTUAL" "ARG PI_VERSION in $DF_VARIANT"
check_pin pi-atelier "$(read_pin_row pi-atelier)" "$ATELIER_ACTUAL" "ARG PI_ATELIER_REF in $DF_VARIANT"
check_pin mempalace "$(read_pin_row mempalace)" "$MEMPALACE_ACTUAL" "ARG MEMPALACE_VERSION in $DF_BASE"
# ---------------------------------------------------------------------------
# 4. DOCKER_HUB.md's Node claim vs ARG NODE_VERSION. This is the published page,
# so it is the one whose staleness reaches users.
# ---------------------------------------------------------------------------
NODE_ACTUAL="$(read_arg "$DF_BASE" NODE_VERSION)"
NODE_DOCUMENTED="$(sed -n 's/.*\*\*Node\.js\*\* v\([0-9][0-9]*\).*/\1/p' "$HUB" | head -1)"
if [ -z "$NODE_DOCUMENTED" ]; then
fail "$HUB: could not find a '**Node.js** vNN' claim. If the wording changed,
update this gate; do not leave the published page unverified."
elif [ "$NODE_DOCUMENTED" != "$NODE_ACTUAL" ]; then
fail "$HUB claims Node v$NODE_DOCUMENTED but ARG NODE_VERSION=$NODE_ACTUAL.
This file is PUBLISHED to Docker Hub by update-description on every tag,
and it is read from the TAG — so fix it before tagging, not after."
else
ok "$HUB Node claim = v$NODE_ACTUAL"
fi
# ---------------------------------------------------------------------------
# 5. Placeholders CI will not substitute. docker-publish.yml substitutes exactly
# {{PI_VERSION}} and then greps for leftovers of that ONE token, so any other
# {{...}} sails through the guard and is published literally.
# ---------------------------------------------------------------------------
UNKNOWN_PLACEHOLDERS="$(grep -o '{{[A-Za-z0-9_]*}}' "$HUB" | sort -u | grep -v '^{{PI_VERSION}}$' || true)"
if [ -n "$UNKNOWN_PLACEHOLDERS" ]; then
fail "$HUB contains placeholders CI does not substitute, which would be
published verbatim: $(echo "$UNKNOWN_PLACEHOLDERS" | tr '\n' ' ')
docker-publish.yml only fills {{PI_VERSION}}; add substitution there first."
else
ok "$HUB has no placeholders beyond {{PI_VERSION}}"
fi
# Match only the UPPER_SNAKE placeholder convention CI uses. A bare '{{' search
# is WRONG here, and the first version of this check proved it by failing on
# README.md:900 — `docker inspect --format '{{json .Config.Labels}}'`, a Go
# template in a legitimate example, not a placeholder. The gate was wrong, not
# the doc. Keep this anchored to [A-Z] so Go/Jinja/Handlebars examples pass.
README_PLACEHOLDERS="$(grep -o '{{[A-Z][A-Z0-9_]*}}' "$README" | sort -u || true)"
if [ -n "$README_PLACEHOLDERS" ]; then
fail "$README contains placeholder(s) nothing substitutes, so they would render
literally for every reader: $(echo "$README_PLACEHOLDERS" | tr '\n' ' ')
Only DOCKER_HUB.md gets substitution, and only for {{PI_VERSION}}."
else
ok "$README has no unsubstituted placeholders"
fi
# ---------------------------------------------------------------------------
# 6. Hub full_description length.
# ---------------------------------------------------------------------------
HUB_CHARS="$(wc -c < "$HUB" | tr -d ' ')"
if [ "$HUB_CHARS" -gt "$HUB_MAX_CHARS" ]; then
fail "$HUB is $HUB_CHARS chars, over Docker Hub's $HUB_MAX_CHARS-char
full_description limit. update-description would fail with a non-200 AFTER
the full build. Trim it — this file is the essentials-only page, and
README.md is the long form on purpose."
else
ok "$HUB is $HUB_CHARS chars (limit $HUB_MAX_CHARS)"
fi
# ---------------------------------------------------------------------------
# 7. Stale "Unreleased" pointers. "Unreleased" is a CHANGELOG-only concept; in
# a user-facing doc it is always a pointer that outlived what it pointed at.
# This class has now bitten five times, hence a gate rather than vigilance.
# ---------------------------------------------------------------------------
STALE_MARKERS="$(grep -n 'Unreleased' "$README" "$HUB" || true)"
if [ -n "$STALE_MARKERS" ]; then
fail "'Unreleased' appears in a user-facing doc, which is always a stale
pointer once the thing ships (it has happened five times here):
${STALE_MARKERS//$'\n'/$'\n' }
State the fact directly, or move it to CHANGELOG.md where 'Unreleased' means something."
else
ok "no stale 'Unreleased' pointers in $README or $HUB"
fi
echo
if [ "$FAILURES" -eq 0 ]; then
echo "OK: every checked doc claim matches the build files."
exit 0
fi
echo "::error::$FAILURES doc claim(s) have drifted from the build files."
echo
echo "Docs are read from the TAG, not from main: docker-publish.yml checks out"
echo "github.ref, so a fix pushed after tagging does not reach the release or the"
echo "Hub page. Update the docs BEFORE you tag."
if [ "$WARN_ONLY" -eq 1 ]; then
echo "(--warn-only: exiting 0 anyway)"
exit 0
fi
exit 1
+166
View File
@@ -0,0 +1,166 @@
#!/usr/bin/env bash
# check-skill-floor.sh — fail when the vendored pi-extensions skill snapshot in
# rootfs/ ("the floor") has drifted from the package repo it is a snapshot of.
#
# THE DEFECT THIS EXISTS TO CATCH, measured 2026-09-10.
# rootfs/usr/local/share/pi-devbox/skills/pi-extensions/ ships a vendored copy
# of the pi-extensions skill so the skill is ALWAYS present in the image.
# Dockerfile.variant then copies the freshly-cloned package copy OVER the served
# path at /usr/local/share/... — but it never writes back to the repo floor. So
# the floor only silently rots, and it had: 34284 B, untouched since fa04d20
# (2026-07-30), while the package copy was 38973 B. Four copies existed with
# three different sizes.
#
# Why that is worse than ordinary staleness: the floor is a FALLBACK. The copy
# step is guarded by `if [ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build
# where the package clone yields no skill/ keeps the vendored snapshot and still
# succeeds — green, with no manifest flag and no label saying which copy was
# served. The image would ship a July skill and nothing would say so. Keeping
# the floor fresh means that fallback is harmless instead of a silent regression.
#
# WHY A DIRECTORY HASH AND NOT `sha256sum SKILL.md`.
# The same pipeline Dockerfile.variant uses for skillset_snapshot_tree_sha256,
# and for the same documented reason: a file-only compare answers "did this one
# file change", not "is this the same skill". pi-extensions ships TWO files
# (SKILL.md + evaluate-extension-usage.py), so a sibling-file edit would pass a
# file-only check. If you change the pipeline here, change it there too.
#
# WHY GATING ON ANOTHER REPO IS PROPORTIONATE HERE, since that is normally a
# smell: this fires only when the package's skill/ DIRECTORY HASH changes, which
# is exactly and only when the floor has genuinely gone stale. pi-extensions
# commits that do not touch skill/ leave the hash alone and cannot turn this red.
# The repo is also anonymously clonable (verified 2026-09-10 with `git ls-remote`
# and no credentials), so this needs no secret and cannot break on token expiry.
#
# Exit codes — deliberately three, matching scripts/lint-shell.sh's philosophy
# that a gate which cannot run must not pass:
# 0 in sync (or the package legitimately has no skill/ at this ref)
# 1 DRIFT — the floor differs from the package
# 2 cannot run — no package copy could be obtained
set -euo pipefail
REPO_ROOT=$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)
FLOOR_DIR="${REPO_ROOT}/rootfs/usr/local/share/pi-devbox/skills/pi-extensions"
# Defaults mirror Dockerfile.variant's ARGs so this checks what the build builds.
PI_EXTENSIONS_REPO="${PI_EXTENSIONS_REPO:-https://gitea.jordbo.se/joakimp/pi-extensions.git}"
PI_EXTENSIONS_REF="${PI_EXTENSIONS_REF:-main}"
PACKAGE_DIR=""
WARN_ONLY=0
TMPDIR_CLONE=""
usage() {
cat <<'EOF'
Usage: scripts/check-skill-floor.sh [options]
--package-dir DIR Compare against an existing skill directory instead of
cloning. In a devbox container use /opt/pi-extensions/skill
for a fully offline run.
--warn-only Report drift but exit 0 (advisory use, e.g. a local hook).
-h, --help This text.
Environment: PI_EXTENSIONS_REPO, PI_EXTENSIONS_REF (default main) — both mirror
the Dockerfile.variant ARGs of the same name.
EOF
}
while [ $# -gt 0 ]; do
case "$1" in
--package-dir) PACKAGE_DIR="${2:-}"; shift 2 ;;
--warn-only) WARN_ONLY=1; shift ;;
-h|--help) usage; exit 0 ;;
*) echo "::error::unknown argument: $1" >&2; usage >&2; exit 2 ;;
esac
done
cleanup() {
if [ -n "$TMPDIR_CLONE" ]; then rm -rf "$TMPDIR_CLONE"; fi
}
trap cleanup EXIT
# Identical to Dockerfile.variant's tree_sha256(): relative paths + per-file
# sha256 over a sorted `find`, folded into one digest. Deterministic, never
# readdir order.
tree_sha256() {
( cd "$1" && find . -type f -print | LC_ALL=C sort | xargs -r sha256sum ) \
2>/dev/null | sha256sum | cut -d' ' -f1
}
if [ ! -d "$FLOOR_DIR" ]; then
echo "::error::floor directory is missing: ${FLOOR_DIR}"
echo "::error::rootfs/ is supposed to guarantee the skill is always in the image."
exit 2
fi
SOURCE_DESC=""
if [ -n "$PACKAGE_DIR" ]; then
if [ ! -d "$PACKAGE_DIR" ]; then
echo "::error::--package-dir does not exist: ${PACKAGE_DIR}"
exit 2
fi
SOURCE_DESC="local directory ${PACKAGE_DIR}"
else
command -v git >/dev/null 2>&1 || { echo "::error::git not found; cannot obtain the package copy."; exit 2; }
TMPDIR_CLONE=$(mktemp -d)
# Fetch the single ref shallowly. `git fetch <ref>` accepts a branch, a tag
# and (on Gitea) a reachable commit, which is why this is not `clone --branch`
# — CI resolves PI_EXTENSIONS_REF to a 40-hex SHA before the build.
if ! ( cd "$TMPDIR_CLONE" \
&& git init -q . \
&& git remote add origin "$PI_EXTENSIONS_REPO" \
&& git fetch -q --depth 1 origin "$PI_EXTENSIONS_REF" \
&& git checkout -q FETCH_HEAD ) 2>/dev/null; then
echo "::error::could not fetch ${PI_EXTENSIONS_REF} from ${PI_EXTENSIONS_REPO}"
echo "::error::Cannot determine whether the floor is stale, so this is exit 2, not a pass."
echo "::error::For an offline run, pass --package-dir /opt/pi-extensions/skill"
exit 2
fi
PACKAGE_SHA=$( cd "$TMPDIR_CLONE" && git rev-parse --short HEAD )
PACKAGE_DIR="${TMPDIR_CLONE}/skill"
SOURCE_DESC="${PI_EXTENSIONS_REPO} @ ${PI_EXTENSIONS_REF} (${PACKAGE_SHA})"
fi
# A ref with no skill/ is the documented fallback case: Dockerfile.variant keeps
# the vendored snapshot and the build succeeds. Nothing to compare, so this is
# not drift — but it IS the exact condition under which the floor ships, so say
# so loudly rather than printing a silent green tick.
if [ ! -d "$PACKAGE_DIR" ]; then
echo "::warning::package has no skill/ at this ref — the vendored floor is what will ship."
echo " source : ${SOURCE_DESC}"
echo " floor : $(tree_sha256 "$FLOOR_DIR")"
exit 0
fi
FLOOR_HASH=$(tree_sha256 "$FLOOR_DIR")
PKG_HASH=$(tree_sha256 "$PACKAGE_DIR")
if [ "$FLOOR_HASH" = "$PKG_HASH" ]; then
echo "OK: vendored pi-extensions floor matches the package."
echo " source : ${SOURCE_DESC}"
echo " tree_sha256: ${FLOOR_HASH}"
exit 0
fi
# `set -e` interacts badly with `[ … ] && x` as a bare statement, so both of
# these are explicit if-blocks rather than AND-lists.
LEVEL="error"
if [ "$WARN_ONLY" -eq 1 ]; then LEVEL="warning"; fi
echo "::${LEVEL}::vendored pi-extensions skill floor has DRIFTED from the package."
echo " source : ${SOURCE_DESC}"
echo " floor tree_sha256 : ${FLOOR_HASH}"
echo " pkg tree_sha256 : ${PKG_HASH}"
echo ""
echo " per-file differences:"
diff -rq "$FLOOR_DIR" "$PACKAGE_DIR" 2>&1 | sed 's/^/ /' || true
echo ""
echo " Remedy — re-sync the floor and commit it:"
echo " cp -a <pi-extensions>/skill/. ${FLOOR_DIR}/"
echo " git add ${FLOOR_DIR#"${REPO_ROOT}/"} && git commit"
echo ""
echo " NOTE this forces one full base rebuild: base_tag hashes Dockerfile.base"
echo " + rootfs/, and that rebuild is what re-bakes the refreshed floor."
if [ "$WARN_ONLY" -eq 1 ]; then exit 0; fi
exit 1
+177 -4
View File
@@ -28,6 +28,9 @@
# (human, --json, --quiet)
# - (studio variant only, auto-detected) pi-studio cloned + prebuilt
# client bundle present + registered via `pi install`
# - no foreign npm-11 platform packages (@esbuild, clipboard) beyond the host
# - no build-time npm cache (/root/.npm) shipped in the image
# - esbuild compiles + clipboard native loads at every install site
# - image size within threshold
set -euo pipefail
@@ -105,6 +108,18 @@ else
run "node" "node --version"
fi
run "git" "git --version"
# NOTE: the shellcheck binary is a GATE DEPENDENCY, not a convenience.
# scripts/lint-shell.sh is the release gate (the lint-gate job resolve-versions
# depends on) and it exits 2 when the binary is missing, by design — "a gate that
# cannot run must not pass". Measured on v1.8.14: it was absent from the image, so
# that gate could not be run by a developer in ANY container, only in CI.
# Asserted here so its absence fails a build instead of being discovered by a hook
# that then refuses every push (hooks/pre-push).
#
# This comment must not BEGIN with the tool's name: a line starting with
# `# shellcheck` is parsed as a DIRECTIVE, not a comment (SC1073/SC1072). The
# gate added in this same change caught that here, before the push.
run "shellcheck (lint gate dependency)" "shellcheck --version | grep -qE '^version: [0-9]'"
run "aws" "aws --version"
run "uv" "uv --version"
run "nvim" "nvim --version"
@@ -303,8 +318,24 @@ run "pi-toolkit clone" "test -d /opt/pi-toolkit && git -C /opt/pi-toolkit rev
run "pi-extensions clone" "test -d /opt/pi-extensions && git -C /opt/pi-extensions rev-parse --short HEAD"
run "pi-fork clone + node_modules" \
"test -f /opt/pi-fork/package.json && test -d /opt/pi-fork/node_modules"
run "pi-observational-memory clone + node_modules" \
"test -f /opt/pi-observational-memory/package.json && test -d /opt/pi-observational-memory/node_modules"
# om is checked differently from pi-fork ON PURPOSE. It declares ZERO runtime
# dependencies: 8 devDependencies (omitted by --omit=dev) and 4 peerDependencies,
# which pi itself provides. npm 10 still materialised a node_modules for it, but
# that directory held exactly ONE file (.package-lock.json, 4 KB) and no nested
# package.json at all — 20 empty scope dirs. npm 11 stopped creating it, so the
# old `test -d node_modules` assertion went red on v1.9.0 while nothing about om
# had changed or broken. It was asserting an npm artefact, not a property of the
# shipped software. What actually has to hold is that the entry point pi loads
# exists, so assert THAT, straight out of the manifest pi reads
# (package.json -> pi.extensions), rather than a hardcoded path that could drift.
run "pi-observational-memory clone + declared pi entry point" \
"test -f /opt/pi-observational-memory/package.json && \
node -e 'const p=require(\"/opt/pi-observational-memory/package.json\"),f=require(\"fs\"),h=require(\"path\"); \
const l=(p.pi&&p.pi.extensions)||[]; \
if(!l.length){console.error(\"package.json declares no pi.extensions\");process.exit(1)} \
for(const e of l){const t=h.resolve(\"/opt/pi-observational-memory\",e); \
if(!f.existsSync(t)){console.error(\"declared entry missing: \"+t);process.exit(1)}} \
console.log(\"entries ok: \"+l.join(\",\"))'"
# ...and that the clone carries the AUTH FIX, not merely that it exists. om's
# pre-flight hasUsableAuth() check silently disabled `recall` for ~8 weeks once
# pi moved to request-time SigV4 signing and stopped exposing a static Bedrock
@@ -514,6 +545,50 @@ run "manifest skill fingerprint matches the baked snapshot" '
echo "manifest=[$m] actual=[$a]" >&2
[ -n "$m" ] && [ "$m" = "$a" ]
'
# ── Which pi-extensions skill copy shipped ──────────────────────────────
# Closes the silent-fallback hole. The refresh in Dockerfile.variant is guarded
# by `[ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates
# the co-located skill keeps the vendored floor and still succeeds GREEN, with
# nothing recording that a snapshot shipped instead of the package copy. Measured
# 2026-09-10: the floor had been stale since 2026-07-30, so that path would have
# shipped a six-week-old skill in silence. The floor is fresh now and gated by the
# skill-floor lint job, but "the fallback is currently harmless" is a fact with a
# shelf life, whereas "the image says which copy it got" keeps working.
#
# vendored-floor FAILS here rather than merely warning: these images track main,
# where the package has co-located skill/ since fa04d20, so a fallback means the
# clone did not resolve as intended and that is a defect to investigate. A fork
# deliberately pointing at a mirror without skill/ is the one case that should
# edit this assertion — which is the honest place for that decision to surface.
run "manifest names which pi-extensions skill copy shipped" '
j=/etc/pi-devbox/build-manifest.json
s=$(jq -r ".pi_extensions_skill_source // empty" $j)
h=$(jq -r ".pi_extensions_skill_tree_sha256 // empty" $j)
echo "source=[$s] tree_sha256=[$h]" >&2
printf "%s" "$h" | grep -qxE "[0-9a-f]{64}" || {
echo "pi_extensions_skill_tree_sha256 is not a 64-hex digest" >&2; exit 1; }
case "$s" in
package) ;;
vendored-floor)
echo "FALLBACK: clone had no skill/ at this ref, so the image ships the committed floor" >&2; exit 1 ;;
divergent)
echo "MIXED: served directory is part package and part floor" >&2; exit 1 ;;
*)
echo "pi_extensions_skill_source absent or unrecognised" >&2; exit 1 ;;
esac
'
# Same shape as the mempalace fingerprint check above, and for the same reason: a
# recorded hash that is never recomputed is a claim, not a measurement.
run "recorded pi-extensions skill hash matches the served bytes" '
j=/etc/pi-devbox/build-manifest.json
d=/usr/local/share/pi-devbox/skills/pi-extensions
m=$(jq -r ".pi_extensions_skill_tree_sha256 // empty" $j)
a=$( (cd "$d" && find . -type f -print | LC_ALL=C sort | xargs -r sha256sum) | sha256sum | cut -d" " -f1)
echo "manifest=[$m] actual=[$a]" >&2
[ -n "$m" ] && [ "$m" = "$a" ]
'
# OCI labels live in the image config, not the container fs — inspect them
# from the host docker rather than via `docker run`.
LBL=$(docker inspect --format '{{ index .Config.Labels "se.jordbo.pi-devbox.pi-extensions-ref" }}' "$IMAGE" 2>/dev/null || true)
@@ -622,7 +697,22 @@ exec_test "mempalace skill linked (fallback)" 'test -L $HOME/.agents/skills
# (forgotten bump AND re-vendored stale snapshot). Upstream content behind this
# refresh: the bare project-name wing convention and the <harness>@<device>
# added_by rule.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Diaries self-heal; plain drawers do not" "$f" && ! grep -q "Agent diaries live in" "$f" && echo ok'
#
# Unreleased: RE-PINNED again on refresh e9e09d9 -> e9e45f7. The retired pair was
# still green against the new snapshot (the diaries section was untouched), so it
# was blind to this refresh for the same reason the v1.8.13 pair was blind to
# that one. The replacement pair is unusually strong because BOTH witnesses come
# out of the same upstream commit: skillset e9e45f7 ADDED the "Withdrawing an ask
# you sent" bullet and DELETED the sentence "there is nothing anyone can do about
# it from the other end" that the new bullet contradicts. Directions were
# MEASURED against both files, not read off the diff: "Withdrawing an ask you
# sent" is new=1/old=0, "nothing anyone can do about it from the other end" is
# new=0/old=1. A canary whose negative witness was removed by the very commit it
# pins fails loudly on the OLD bytes instead of merely failing to notice them,
# which is the property every previous pair here lacked. Upstream content:
# requester-side ask withdrawal became DEPLOYED behaviour once v1.9.1 baked
# mempalace-toolkit e68ee20 (>= e2b060a) through the floating MEMPALACE_TOOLKIT_REF.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Withdrawing an ask you sent" "$f" && ! grep -q "nothing anyone can do about it from the other end" "$f" && echo ok'
# Link TARGETS, not just link existence: with no skillset mounted (as here) the
# baked tree must be what resolves, for all four vendored skills.
exec_test "vendored skills resolve to the baked tree (no skillset mounted)" \
@@ -641,9 +731,17 @@ exec_test "pi-devbox-version reports skill sources (all baked, no skillset here)
'out=$(pi-devbox-version)
echo "$out" | grep -q "skills:" || { echo "no skills section" >&2; exit 1; }
for s in mempalace pi-extensions pi-devbox-environment credential-incident-response; do
echo "$out" | grep -qE "^ $s +baked$" \
echo "$out" | grep -qE "^ $s +baked( \([^)]*\))?$" \
|| { echo "$s not reported as baked" >&2; exit 1; }
done; echo ok'
# The optional " (...)" above is what pi-extensions now appends to say WHICH copy
# shipped — "baked (package copy)", or a loud FALLBACK/MIXED annotation. Without
# allowing it, adding that annotation turned this assertion red on v1.9.0 even
# though the state it reported was the correct one. The suffix is deliberately
# matched loosely rather than pinned to "(package copy)", because WHICH copy
# shipped is already asserted authoritatively above, against the manifest field
# and its measured tree hash, and duplicating that here in a regex would just
# create a second place to update whenever the wording changes.
# The boot banner must NOT carry the section: entrypoint-user.sh prints the
# version FIRST, before the baked links exist and long before the skillset
# deploy + reconcile run last, so anything it said about skill sources would be
@@ -799,6 +897,55 @@ exec_test "pi-fork extensions floor is [] (forks cannot write to the palace)" \
exec_test "/tmp/sshcm dir mode 700 (ssh ControlMaster)" \
'test -d /tmp/sshcm && [ "$(stat -c %a /tmp/sshcm)" = "700" ] && echo ok'
# ── Build-time leftovers (npm 11 bloat sentinels) ─────────────────────
# Both of these are worth a PASS/FAIL assertion rather than a size-gate
# diagnostic, because the size gate has ~225 MB of deliberate margin: v1.9.1
# shipped +131 MB of pure build residue and stayed green. These name the
# residue directly, so a regression is legible instead of merely "bigger".
echo ""
echo "── Build-time leftovers ──"
# npm 11 installs EVERY optional platform package of a native dependency, not
# just the host's (it ignores os/cpu, --os/--cpu and npmrc os=/cpu=). Two
# families are affected and pruned in Dockerfile.variant: @esbuild/<platform>
# and @mariozechner/clipboard-<triple>. Keep-set is the host arch only, plus
# clipboard's gnu AND musl (its napi loader picks between them at runtime).
# Runs as root because the image declares no USER; that is also what lets the
# cache assertion below read /root.
run "no foreign platform packages (npm 11 sentinel)" \
'arch=$(node -p process.arch); bad=$(find /usr/lib/node_modules /opt -type d \( -regex ".*/@esbuild/[^/]+" -o -regex ".*/@mariozechner/clipboard-[^/]+" \) ! -name "linux-$arch" ! -name "clipboard-linux-$arch-gnu" ! -name "clipboard-linux-$arch-musl" -prune -print 2>/dev/null); if [ -n "$bad" ]; then echo "foreign platform dirs shipped:" >&2; echo "$bad" >&2; du -sm $bad 2>/dev/null | sort -rn | head -5 >&2; exit 1; fi; echo ok'
# The build's own npm download cache is not free: it lands in the layer that
# created it. v1.9.1 shipped 145 MB of /root/.npm/_cacache (35 MB in v1.8.14)
# — the largest single item in its +131 MB residual, and invisible to the
# size-gate diagnostics because those only looked under node_modules and /opt.
# Nothing at runtime reads it: root's cache, while the container runs as
# `developer`. NOTE the assertion must run as root or a permission error on
# mode-700 /root would make `test ! -d` pass for the wrong reason.
run "no build-time npm cache shipped (/root/.npm)" \
'test "$(id -u)" = "0" || { echo "assertion needs root to read /root" >&2; exit 1; }; if [ -e /root/.npm ]; then echo "/root/.npm shipped: $(du -sm /root/.npm | cut -f1) MB" >&2; exit 1; fi; echo ok'
# The prune's risk is not "too big" but "removed something needed", and only a
# FUNCTIONAL check covers that. These load the natives from every install site
# found in the image, so they also scale to the studio variant's third site.
#
# NOTE THE PATH-QUALIFIED require(). The obvious form, `node -e
# 'require("esbuild")...'`, resolves by walking up from the CURRENT DIRECTORY —
# so it fails with MODULE_NOT_FOUND from /workspace on a perfectly good image,
# because esbuild lives nested inside the pi trees and global installs are not
# on node's require path (NODE_PATH is unset). That exact command was left in a
# runbook as "if this fails, revert the release", and it duly failed for the
# wrong reason on the first machine that ran it. A check must fail only for the
# thing it is checking.
run "esbuild works at every install site (prune removed weight, not function)" \
'sites=$(find /usr/lib/node_modules /opt -type d -path "*/node_modules/esbuild" -prune 2>/dev/null); if [ -z "$sites" ]; then echo "no esbuild install found at all" >&2; exit 1; fi; for d in $sites; do node -e "require(\"$d\").transformSync(\"const x:number=1\",{loader:\"ts\"})" || { echo "esbuild broken at $d" >&2; exit 1; }; done; echo ok'
# Clipboard is the family pruned second, and its napi-rs loader picks its native
# binding at require() time — so a successful load IS the proof that the kept
# platform package is the one this image needs.
run "clipboard native loads at every install site" \
'sites=$(find /usr/lib/node_modules /opt -type d -path "*/node_modules/@mariozechner/clipboard" -prune 2>/dev/null); if [ -z "$sites" ]; then echo "no @mariozechner/clipboard install found at all" >&2; exit 1; fi; for d in $sites; do node -e "var c=require(\"$d\"); if (typeof c.setText !== \"function\") { throw new Error(\"native binding missing\"); }" || { echo "clipboard native broken at $d" >&2; exit 1; }; done; echo ok'
# ── Image size ────────────────────────────────────────────────────────
echo ""
echo "── Image size ──"
@@ -826,6 +973,32 @@ elif [ "$SIZE_MB" -le "$SIZE_THRESHOLD_MB" ]; then
printf " ✅ size: %d MB (threshold %d MB)\n" "$SIZE_MB" "$SIZE_THRESHOLD_MB"; PASS=$((PASS+1))
else
printf " ❌ size: %d MB exceeds threshold %d MB\n" "$SIZE_MB" "$SIZE_THRESHOLD_MB"; FAIL=$((FAIL+1))
# A bare "too big" verdict cost a full CI-log dig plus a local npm bisect to
# attribute the v1.9.0 overshoot (+431 MB, which turned out to be npm 11
# installing 26 @esbuild platform binaries per pi-coding-agent copy). The
# container already knows where its bytes are, so make it say so: the biggest
# layers, and the biggest directories under the paths that historically grow.
# Same principle as the run() helper above — a red assertion should carry its
# own diagnostic rather than send the next reader spelunking.
echo " ── largest layers (docker history) ──"
docker history --format '{{.Size}}\t{{.CreatedBy}}' "$IMAGE" 2>/dev/null \
| grep -vE '^0B' | head -12 | sed 's/^/ /' | cut -c1-160
echo " ── largest directories in the image ──"
docker run --rm --entrypoint sh "$IMAGE" -c \
'du -sm /usr/lib/node_modules/* /opt/* /usr/local/share/ms-playwright 2>/dev/null | sort -rn | head -12' \
2>/dev/null | sed 's/^/ /' || echo " (could not inspect directories)"
echo " ── build caches that should not be in the image ──"
# v1.9.1's residual was 145 MB of npm cache under /root, and the du list
# above cannot see it: it enumerates node_modules and /opt only. A gate whose
# diagnostic looks only where the bytes were LAST time sends the next reader
# spelunking again, so name the cache paths explicitly.
docker run --rm --entrypoint sh "$IMAGE" -c \
'du -sm /root/.npm /root/.cache /tmp/node-compile-cache /home/developer/.npm 2>/dev/null | sort -rn' \
2>/dev/null | sed 's/^/ /' || true
echo " ── foreign platform dirs (npm 11 regression sentinel) ──"
docker run --rm --entrypoint sh "$IMAGE" -c \
'find /usr/lib/node_modules /opt -type d \( -regex ".*/@esbuild/[^/]+" -o -regex ".*/@mariozechner/clipboard-[^/]+" \) -printf "%f\n" 2>/dev/null | sort | uniq -c | sort -rn | head' \
2>/dev/null | sed 's/^/ /' || true
fi
# ── Summary ───────────────────────────────────────────────────────────