Compare commits

...

13 Commits

Author SHA1 Message Date
Joakim Persson 1c905480e3 docs(changelog): note the mempalace feed-tick fix this image carries
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 10s
Lint / actionlint (push) Successful in 19s
The recurring "[mempalace ext] feed (tick) failed: mine timed out after 30000ms"
message is fixed in mempalace-toolkit (309980b + e68ee20) and this image is what
delivers it, since CI resolves MEMPALACE_TOOLKIT_REF to a commit SHA at build
time. Worth a changelog entry rather than leaving it implicit in a ref bump: it
is the most visible symptom operators on this fleet have been living with, and
the entry records that it was a genuine defect (overlapping mines on a
single-writer palace) rather than the cosmetic annoyance it was parked as.
2026-09-10 20:58:42 +02:00
Joakim Persson ff6fd1492a feat(manifest): record WHICH pi-extensions skill copy shipped
Closes the half deliberately left open by cac5e00's skill-floor gate, and the
more important half: "the floor is currently fresh" is a fact with a shelf
life, whereas "the image says which copy it got" keeps working.

The refresh in Dockerfile.variant is guarded by
`[ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates the
co-located skill keeps the vendored floor and still succeeds GREEN, with nothing
in the manifest, labels or logs separating that from a normal build. Afterwards
the two are indistinguishable by inspection -- same path, same filenames, same
permissions -- which is exactly how the floor went unnoticed from 2026-07-30 to
2026-09-10.

build-manifest.json gains pi_extensions_skill_source and
pi_extensions_skill_tree_sha256, MEASURED rather than passed as build-args, per
the ground-truth rule the surrounding block already follows -- and necessarily
so, since the outcome depends on the clone's contents and no ARG could express
it. Three values, because two would force a lie: package (served bytes equal
the clone's skill/), vendored-floor (clone had no skill/ at this ref), and
divergent (both exist but differ -- e.g. the clone ships SKILL.md but not
evaluate-extension-usage.py, so the served directory is a genuine MIX). No OCI
label mirrors these deliberately: LABEL cannot take a RUN-computed value, and a
label fed from an ARG would be the claim-not-measurement being removed here.

Two smoke assertions make the record a gate: the source must be named and be
`package` -- vendored-floor FAILS rather than warns, since these images track
main where the package has shipped skill/ since fa04d20, so a fallback means
the clone did not resolve as intended -- and the tree hash is recomputed over
the served directory, because a recorded hash never recompared is a claim.

pi-devbox-version annotates the line too: "baked (package copy)" normally, or a
yellow "(FALLBACK: vendored floor)". Its existing section reports which copy is
READ at runtime; this is the one fact decided at BUILD time and unrecoverable
later. Old images degrade cleanly -- field absent, jq // empty yields nothing,
line prints plain "baked" as before (verified against this v1.8.14 manifest).

Tested by running the exact logic against this container's real layout, with
the expected value written down before each: package (served == clone),
vendored-floor (clone path absent), divergent (clone lacking the .py while the
served dir has it), and null (empty served dir) -- all four as predicted. The
five pi-devbox-version render branches likewise, including the absent-field
case. Emitted JSON validated with jq for both the populated and null forms.

Gates green: lint-shell.sh (15 files), hadolint 2.15.1, actionlint 1.7.12,
check-base-hash.sh, check-skill-floor.sh, vendor-mempalace-skill.sh --check.
2026-09-10 20:30:16 +02:00
Joakim Persson edc7659add chore(deps): node 22->24, actionlint 1.7.12, hadolint 2.15.1, skillset ref
Audited every component the image obtains OUTSIDE debian/apt. Of ~23, the 19
that resolve `latest` at build time were already current or refresh themselves
on the next rebuild, and the hard pins for pi (0.85.1), mempalace (3.9.0) and
pi-atelier (v0.10.1) were already newest. Four needed a human.

NODE_VERSION 22 -> 24 (LTS "Krypton"). This was a latent defect rather than
housekeeping: agent-browser publishes engines.node ">=24.0.0", so the image sat
BELOW a declared requirement -- v1.8.14 shipped node 22.23.2 with agent-browser
0.37.1, so every build installed it with an npm EBADENGINE warning and ran the
baked browser automation outside its supported range. pi (">=22.19.0") and
playwright (">=20") are satisfied either way. Verified before bumping, since a
missing NodeSource suite breaks every arch at once: setup_24.x returns HTTP 200
and node_24.x advertises `Architectures: amd64 arm64 armhf x86_64`, covering
the arm64 fleet and the amd64 CI runners. Nothing else pinned the node major.

actionlint 1.7.7 -> 1.7.12 and hadolint 2.14.0 -> 2.15.1, each RUN AGAINST THIS
TREE at the new version before being pinned -- both clean, no new findings. A
linter bump is the one dependency update that can turn CI red on unchanged
code, so it is verified locally rather than discovered on a round trip.

SKILLSET_SNAPSHOT_REF e9e09d9 -> 4d7c0ea via scripts/vendor-mempalace-skill.sh,
never by hand: that script is the only thing permitted to write the ARG,
because a cp without a matching bump yields a manifest that confidently lies.
This proved PROVENANCE-ONLY -- the ref was 6 commits behind, but
skills/mempalace/SKILL.md is byte-identical at both (3675bfab), so the snapshot
was already correct and only its recorded origin was stale. No rootfs/ bytes
changed, the smoke-test phrase canary stays valid, and this ARG alone would not
force a base rebuild (the node bump does).

Two measurement traps worth recording, since both would have produced a wrong
answer: GitHub's releases/latest reports pi-atelier v0.10.0 as newest because
v0.10.1 is a TAG WITH NO RELEASE OBJECT -- the pin was already current, and
`git ls-remote --tags` is the instrument that shows it. And gitea-mcp is hosted
on gitea.com, not GitHub, so querying api.github.com returned nothing at all
rather than an error.

Verified with every gate this repo owns, all green, using the NEW linter pins:
lint-shell.sh (15 files), check-workflow-shell.sh, check-base-hash.sh,
actionlint 1.7.12, hadolint 2.15.1, check-skill-floor.sh, and
vendor-mempalace-skill.sh --check.
2026-09-10 20:17:32 +02:00
Joakim Persson cac5e00a31 feat(ci): gate the vendored pi-extensions skill floor, and bake python3-yaml
Follows ecfd2fc, which refreshed the stale floor by hand. A one-off refresh
fixes the symptom; this makes the drift impossible to reintroduce silently.

scripts/check-skill-floor.sh compares the repo floor
(rootfs/usr/local/share/pi-devbox/skills/pi-extensions/) against the package
repo it is a snapshot of, wired in as a new `skill-floor` job in lint.yml.

DIRECTORY hash, not `sha256sum SKILL.md`, using the same tree_sha256 pipeline
Dockerfile.variant uses for skillset_snapshot_tree_sha256 and for the reason
already documented there: a file-only compare answers "did this one file
change", not "is this the same skill". Verified by NEGATIVE CONTROL rather
than asserted -- with SKILL.md left byte-identical and only
evaluate-extension-usage.py edited, the directory check fails (rc=1) where a
file-only compare would have passed. Seven behaviour tests, each with its
expected rc written down before running: in-sync via local dir (0), in-sync
via anonymous remote clone (0), missing --package-dir (2), bad argument (2),
content drift (1), the sibling-file case (1), and --warn-only over drift (0).

Exit codes 0 in sync / 1 drift / 2 cannot-run, matching scripts/lint-shell.sh:
a gate that cannot run must not pass, so an unreachable package repo is a red
2 and never a green tick. A ref with no skill/ is NOT drift -- that is the
documented fallback -- but it emits ::warning:: because it is precisely the
condition under which the floor ships.

Gating on another repo is normally a smell. It is proportionate here because
the check can only fire when skill/ itself changed, which is exactly when the
floor has gone stale; pi-extensions commits that leave skill/ alone cannot
turn this red. It also needs no secret: pi-extensions is anonymously clonable
(verified with `git ls-remote` and no credentials), so it cannot start failing
when a token expires.

Also bakes python3-yaml (552 KB, zero extra deps) into Dockerfile.base. This
is the shellcheck story repeating exactly: scripts/check-workflow-shell.sh --
the guard against the Gitea sh/dash footgun that broke resolve-versions
(ed49b8d) and promote-base-latest (b7197e8) -- hard-exits with "python3 yaml
module missing", so a gate this repo already owns could not be run locally by
anyone. lint.yml installing it explicitly in CI was the evidence. Found while
wiring the job above: the guard could not be run before pushing.

CHANGELOG Unreleased updated for both this and ecfd2fc, including an explicit
note on what is NOT fixed -- the silent-fallback half still has no manifest
flag recording which copy was served.

Verified locally with every gate this repo owns, all green: lint-shell.sh (15
files clean), check-workflow-shell.sh, check-base-hash.sh, actionlint 1.7.7
(pinned, same version as CI), hadolint 2.14.0, and the new check itself.
2026-09-10 19:14:01 +02:00
Joakim Persson ecfd2fc2e5 feat: bake dig/ldapsearch/xxd and refresh the stale pi-extensions rootfs floor
Two changes that share one forced base rebuild, hence one commit.

1. THREE PACKAGES, each closing a capability gap measured during the
   gitea.egl.lan/FreeIPA work on 2026-09-09..10 rather than a preference:

   bind9-dnsutils (~6.1 MB measured) -- dig/host/nslookup were ALL absent,
   so the container could resolve names but had no way to interrogate a
   SPECIFIC nameserver. `getent hosts` only follows the resolver's default
   path, so diagnosing "gateway 172.16.88.1 NXDOMAINs the egl.lan zone
   while 10.20.253.1 is authoritative for it" had to be hand-rolled in
   python3. Split-horizon DNS is a recurring class of bug on this fleet.
   Note the package name: plain `dnsutils` is transitional in trixie.

   ldap-utils (1244 KB, pulls nothing extra) -- the fleet authenticates
   against FreeIPA, yet every LDAP probe had to be run by SSHing to an
   already-enrolled host. Simple binds only; GSSAPI would additionally
   need krb5-user + libsasl2-modules-gssapi-mit, deliberately not added
   as that is a Kerberos-client decision, not a tool.

   xxd (198 KB) -- convenience for verifying git-crypt blob magic in
   myconfigs; `od -c` from coreutils already does the same job.

   netcat-openbsd was in the original proposal and is deliberately NOT
   here: measured redundant, because socat is already baked and bash's
   /dev/tcp does reachability checks with zero packages (verified against
   gitea.egl.lan:3000). Recorded in the Dockerfile so the omission reads
   as a decision rather than an oversight.

2. ROOTFS FLOOR REFRESH: rootfs/.../pi-extensions/SKILL.md was 34284 B,
   unchanged since fa04d20 (2026-07-30), while the canonical package copy
   is 38973 B. Dockerfile.variant copies the fresh package copy over the
   SERVED path at build time but never writes back to this floor, so the
   floor is a silent fallback: if that build-time copy is ever absent it
   ships the July skill with no log line or manifest flag to say which
   version deployed. Refreshed from pi-extensions@c64c122, verified
   byte-identical to both the canonical and the runtime-served copies.

Why one commit: the base_tag hash folds in `cat Dockerfile.base` AND
`find rootfs -type f | xargs cat` (.gitea/workflows/docker-publish.yml),
so either change alone forces the same full base rebuild -- and that
rebuild is precisely what re-bakes rootfs/ as it then stands. Emulating
the workflow hash with a fixed toolkit ref: f3d6462c7416 -> fc4edda03c54.

Verified: scripts/check-base-hash.sh passes (no new ARG *_REF added), and
no shell scripts are touched so the lint-shell gate is unaffected. Sizes
and dependency fan-out measured via apt-get --no-install-recommends
--dry-run on Debian 13 trixie.
2026-09-10 18:56:11 +02:00
joakimp 15a3728ae9 feat: bake shellcheck and add a client-side pre-push lint gate
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 19s
v1.8.14 made shell lint a RELEASE gate (scripts/lint-shell.sh, shared by lint.yml
and the new lint-gate job that resolve-versions depends on), and that script
correctly exits 2 when shellcheck is absent -- "a gate that cannot run must not
pass". Measured on v1.8.14 on 2026-09-09 by three routes (command -v, dpkg -l, a
filesystem search): shellcheck was NOT IN THE IMAGE AT ALL. So the gate could not
be run by a developer in any container, only in CI, and the loop stayed
write-shell -> push -> wait for CI -> discover. That is the loop the gate was
added to shorten, after v1.8.14's first attempt burned ~46 min on a tree whose
lint had already been red for 24 hours.

shellcheck 0.10.0-1 added to the Dockerfile.base apt block: ~39 MB installed
(Installed-Size 40112 KB), measured to pull ZERO additional packages under
--no-install-recommends because libc6/libffi8/libgmp10 are already present.
NOTE this forces one full base rebuild -- base-decide hashes Dockerfile.base +
rootfs/, so unlike a scripts/ change it cannot reuse the existing base- layer.

hooks/pre-push is opt-in per clone (git config core.hooksPath hooks), bypassable
with --no-verify, and execs scripts/lint-shell.sh rather than reimplementing it
-- one copy, because a duplicated check that drifts is the failure this repo
keeps paying for. Matches the idiom skillset/ and myconfigs/ already use.

WHY THIS REPO HAD NO HOOKS, since it was reported as drift and is not: a peer
asked tor-ms22 for core.hooksPath per clone on the premise that unset meant the
gates were unverified there. Measured: pi-devbox unset, skillset hooks, myconfigs
common/hooks, pi-toolkit unset -- but `git ls-files | grep -i hook` is EMPTY in
both pi-devbox and pi-toolkit, so there was nothing to point at on any machine
and unset was the only correct value. This closes the real half for pi-devbox;
pi-toolkit still ships none.

Verified, expected result written down before each check:
  * refusal paths -- shellcheck absent => rc 2 with the remedy named; linter
    missing => rc 2. Never waved through on the assumption CI will catch it.
  * the hook is IN the scan set -- "Checking 14 shell file(s)" with it present,
    13 with it moved aside, so the extensionless file is found by the shebang
    half of the linter's two-signal union. This check exists because the first
    attempt was ambiguous: a planted `[ $UNSET_VAR = "x" ]` was not reported,
    which could equally have meant "not scanned" or "below -S error". It was the
    latter. A count that moves is unambiguous; a clean run is not.
  * it catches the REAL v1.8.14 defect -- planting `echo 'the fleet\'s thing'`
    in hooks/pre-push yields SC1073/SC1072 at severity error, rc=1.

And the gate earned its keep inside this commit: the first version of the
smoke-test assertion carried a comment beginning "# shellcheck is a GATE
DEPENDENCY", and a comment whose first word is the tool's name is parsed as a
DIRECTIVE, not a comment. The new gate failed it with SC1073/SC1072 before the
push -- same family as the v1.8.14 apostrophe, a line that reads as prose to a
human and as syntax to the parser.
2026-09-09 08:57:34 +02:00
joakimp 361babd4fd ci: gate the release on shell lint, from one shared script
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Publish Docker Image / smoke (push) Successful in 7m39s
Publish Docker Image / build-variant-studio (push) Successful in 17m15s
Publish Docker Image / build-variant (push) Successful in 17m58s
Publish Docker Image / update-description (push) Successful in 8s
Publish Docker Image / promote-base-latest (push) Successful in 12s
v1.8.14's first attempt spent ~46 minutes building a base image for a tree whose
own lint had been failing for 24 hours. shellcheck had already flagged the
defect (SC2289, severity error) on the push that introduced it; the lint
workflow went red at run 186 and nobody read it.

lint.yml deliberately skips tag pushes and its reasoning is sound -- the tagged
tree was already linted on main, and a tag-ref lint run sorts above the publish
run, making a release look finished before anything ships. The missing invariant
was never "lint the tag". It was "do not RELEASE a tree whose lint failed", and
only a job inside the publish workflow can enforce that.

So: extract the shell-lint logic from lint.yml into scripts/lint-shell.sh and
call it from both places, then add a lint-gate job that resolve-versions depends
on. resolve-versions is the graph root, so gating it gates everything. Cost is
~40 s at the front of a release; the alternative already cost fifty minutes.

Extracted rather than copied on purpose. A second copy of a check is the drift
this repo keeps paying for -- the same evening produced a skillset mirror that
had sat 9579 B behind its upstream through two consecutive edits.

The script adds one behaviour the inline version lacked: if shellcheck is not
installed it exits 2 rather than silently finding nothing, inheriting the
existing "a gate that cannot run must not pass" rule from hooks/pre-commit in
the skillset repo. Without that, reordering the install step away would turn the
gate into a green tick over zero checks.

Verified locally with a stubbed shellcheck (the real binary is not in the
devbox), five cases, each with its expectation stated first: absent shellcheck
-> rc=2; stub pass -> rc=0 and a non-zero file count; stub fail -> rc=1; a
deliberately unterminated `if` planted in scripts/ -> rc=1 via the bash -n half,
naming the file; removal -> rc=0 again. Discovery cross-checks against CI's own
number: the inline version reported 12 files, the extracted one reports 13, the
difference being lint-shell.sh itself. YAML re-parsed (10 jobs, was 9) with an
assertion that resolve-versions needs lint-gate, and the repo's
check-workflow-shell.sh guard still passes.
2026-09-08 23:41:44 +02:00
joakimp 70e675afee fix(smoke): keep prose out of the single-quoted exec_test body
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 22s
The agent-browser execution guard added on 2026-09-07 carried its explanation
INSIDE the single-quoted script body, and the explanation contained an
apostrophe ("the fleet\'s only recurring amd64 runtime proof"). Inside '...'
bash treats a backslash as literal, so \' does not escape the quote -- it CLOSES
the string. The body truncated at that point and the remaining lines were parsed
by the calling shell.

Consequences, both measured rather than inferred:
  - exec_test received 12 arguments instead of 2 (verified two-sided: the fixed
    tree yields argc=2, HEAD yields argc=12).
  - the leaked `v=$(agent-browser --version)` ran on the CI RUNNER instead of
    inside the image. The runner has no agent-browser, so smoke and
    smoke-studio both failed with "line 770: command not found" after
    build-base had already spent ~46 minutes. Every downstream job was skipped.
  - the truncated body still passed inside the container and printed its green
    tick first, so the log shows a PASS immediately followed by the failure --
    the tick was real, it just no longer covered the assertion.

The prose now sits above the exec_test call, where an apostrophe cannot
terminate anything, and a comment at that spot records why it must stay there.

Not a new failure class: shellcheck flagged it as SC2289 at severity error the
same day, so the lint job has been red since run 186 (2026-09-07 21:21) and was
not read. The gate did its job; nobody looked.
2026-09-08 23:31:48 +02:00
joakimp 601fc98a49 docs(changelog): release v1.8.14
Lint / hadolint (push) Successful in 9s
Publish Docker Image / resolve-versions (push) Successful in 14s
Lint / actionlint (push) Failing after 24s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 50m32s
Publish Docker Image / smoke-studio (push) Failing after 5m13s
Publish Docker Image / build-variant-studio (push) Has been skipped
Publish Docker Image / smoke (push) Failing after 7m40s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Converts the Unreleased section and records what this build carries beyond it:
the mempalace-toolkit bump that makes closing replies reach the mailbox
(deriveClosed, 21023e7 -> e45f6b4), and the L0-L4 subtask documentation landing
via pi-toolkit adfb553 + pi-extensions c64c122.

Notes the mechanism that makes the toolkit fix land at all — the resolved
toolkit SHA is folded into the content-addressed base tag, so the toolkit
moving forces a base rebuild rather than waiting for one — and the consequence
for the amd64 item already in this section: v1.8.13's base was cached, so
Dockerfile.base:607's agent-browser assertion never ran. This base is not
cached, so the native-amd64 proof is finally collected instead of discarded.
2026-09-08 22:09:31 +02:00
joakimp 7e0e66997d docs(env): name MEMPALACE_MAILBOX_NOTIFY — auto-detect cannot work in a container
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Failing after 16s
Unset means the mailbox is silent outside the pi TUI, and the reason is
structural: docker exec does not forward KITTY_WINDOW_ID/TERM_PROGRAM, so
'desktop' detection always falls through to OSC 777, which Kitty does not
implement — the notification then silently does nothing, the worst failure for a
feature whose only job is to break a silence. Documents the four modes, and that
MEMPALACE_MAILBOX_POLL_MS is a FLOOR BETWEEN activity-coupled polls rather than a
wall-clock interval (an idle session polls zero times) — the exact expectation
mismatch reported today.
2026-09-07 21:43:01 +02:00
joakimp 6bd8b79d3a test(smoke): assert agent-browser EXECUTES — it was the discarded amd64 proof
Lint / hadolint (push) Successful in 11s
Lint / actionlint (push) Failing after 23s
Second instance of the same bug class as the node line, in the same file, found
the same way. The agent-browser guard captured the version inside an echo with
2>/dev/null:

  echo "resolved=[$r] version=[$(agent-browser --version 2>/dev/null|head -n1)]" >&2

so the exit code was discarded and a binary that could not execute at all still
PASSED, printing version=[]. Verified two-sided: a stub exiting 127 passes the old
form and is caught by the new one.

Why this exit code matters more than most: smoke runs platforms: linux/amd64 on an
x86 runner, i.e. NATIVE amd64, so this line is the fleet's only recurring amd64
runtime proof for agent-browser's linux-x64 ELF.

NO DEVBOX CAN EVER SUPPLY THAT PROOF. Every machine in the pi fleet is an Apple
Silicon Mac: mbp-m1-2020; tor-ms22 = Mac Studio Mac13,1 M1 Max (fleet-ops
hosts/tor-ms22.md, verified 2026-08-17 with system_profiler); emb-7kj4vr4g =
Apple Silicon, verified 4 routes 2026-09-07. The open "amd64 runtime proof still
needed" ask sent to two devices was asking for the impossible, and emb's reply
naming tor-ms22 as "the only remaining candidate" is wrong for the same reason.
CI had the answer all along and was throwing it away.

Dockerfile.base:607 DOES assert it (`agent-browser --version && \`), but only when
the base rebuilds, and v1.8.13's base was cached — so smoke is where the recurring
gate belongs.
2026-09-07 21:21:16 +02:00
joakimp fabf1274aa docs(changelog): Unreleased section for the smoke node assertion + agent-browser correction
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 19s
Summarises what changed since v1.8.13: the node-major assertion (a bump would
have passed the suite silently), the two-sided verification of the derivation,
and the v1.8.13 agent-browser 0.35.2 -> 0.36.0 correction. No image content
changes; NODE_VERSION still 22.
2026-09-07 21:17:06 +02:00
joakimp 5972a2c535 test+docs: assert the node major in smoke, and correct v1.8.13's agent-browser version
Two findings from a delegated read-only audit of this repo, both verified from the
filesystem before patching.

1. No test asserted the node major, so a node-24 bump would have passed the smoke
   suite SILENTLY. scripts/smoke-test.sh:94 was a bare `run "node" "node --version"`
   — exit-0 and non-empty output only, the printed version compared to nothing —
   while the line above it uses run_expect against $EXPECTED_PI_VERSION for pi. A
   reader skimming the suite would reasonably assume node regressions were covered.
   Worse, this is where the "node v22.23.2 verified" line in the v1.8.13 recreate
   notes came from: printed output, not an assertion.

   Now gated on EXPECTED_NODE_MAJOR, which CI derives from Dockerfile.base's ARG
   NODE_VERSION — the single source of truth (Dockerfile.base:557 is the ONLY hard
   pin in the repo; Dockerfile.variant has no node install at all). That also
   catches a stale cached layer whose node disagrees with the declared ARG.
   Unset => previous behaviour, so this is backward compatible.

   Verified two-sided rather than assumed: the sed derivation yields 22 (empty
   would have silently disabled the assertion, reintroducing the bug); grep -Fq
   "v22." matches v22.23.2; "v24." does NOT match, so a wrong major is caught; and
   "v2." does not prefix-collide. Workflow YAML re-parsed after editing (9 jobs).

2. The v1.8.13 entry claimed "the image's own 0.35.2" for agent-browser. The image
   ships 0.36.0: /usr/lib/node_modules/agent-browser/package.json says version
   0.36.0, engines.node >=24.0.0, and no 0.35.2 exists anywhere in the image. The
   claim was also internally incoherent, contrasting 0.36.0 against a version that
   is not present. Corrected in place with a visible note, since the entry is
   already released. The reasoning survives untouched: the engines floor really is
   vestigial, because /usr/bin/agent-browser is a prebuilt aarch64 ELF invoked
   directly and never through node — which is why 0.36.0 runs fine on 22.23.2.
2026-09-07 21:05:24 +02:00
12 changed files with 1168 additions and 39 deletions
+29
View File
@@ -87,6 +87,35 @@ SSH_KEY_PATH=~/.ssh
# MEMPALACE_PI_REMOTE_PATH=/data/feed
# MEMPALACE_PI_DEVICE=
# ── Mailbox notification: MUST BE NAMED, auto-detect CANNOT work here ──
# The mempalace extension polls the logstream for fleet asks addressed to this
# device and queues them into the next turn. That part needs no config. The
# NOTIFICATION that tells the human it happened does, and unset means SILENT
# outside the pi TUI.
#
# Why there is no working default: terminal identity lives in env vars set by
# the emulator (KITTY_WINDOW_ID, TERM_PROGRAM) and `docker exec` does NOT
# forward them — inside the container pi sees only TERM=xterm-256color no matter
# what is rendering it. So "desktop" auto-detection always falls through to
# OSC 777, which Kitty does not implement, and the notification silently does
# nothing: the worst outcome for a feature whose only job is to break a silence.
# Naming the protocol is what makes it fire.
#
# kitty OSC 99 desktop notification (correct for Kitty, incl. over SSH)
# osc777 OSC 777 (tmux/iTerm2/foot and others)
# desktop OSC 99 if KITTY_WINDOW_ID is visible, else OSC 777 — inside a
# container that means effectively always OSC 777, so prefer naming
# 0 / off suppress entirely (in-TUI notify still shows)
# MEMPALACE_MAILBOX_NOTIFY=kitty
#
# Cadence, if the delivery ever feels late: the poll is coupled to session
# activity (it runs when the agent settles), NOT to a wall clock.
# MEMPALACE_MAILBOX_POLL_MS is therefore a FLOOR BETWEEN POLLS (default 300000),
# not a promise of one every 5 minutes — an idle session polls zero times, and
# session start does the first look.
# MEMPALACE_MAILBOX_POLL_MS=300000
# MEMPALACE_MAILBOX_RESURFACE_MS=3600000
# ── LAN access from the container (host-OS-agnostic) ─────────────────
# On VM-backed hosts (macOS OrbStack / Docker Desktop) the container can't
# reach the host's directly-attached LAN peers by default. The entrypoint
+46 -2
View File
@@ -157,7 +157,41 @@ jobs:
# buildcache silently reuses the layer from whatever pi version was
# current when the cache was first populated. Same class of bug as
# pi-devbox v0.74.0..v0.75.5 (fixed in v0.75.5b 2026-05-23).
# ── release gate ──────────────────────────────────────────────
# Refuse to spend a base build on a tree whose own shell scripts do not lint.
#
# v1.8.14's first attempt is why this exists. smoke and smoke-studio both failed
# at scripts/smoke-test.sh:770 AFTER build-base had already spent ~46 minutes,
# on a defect shellcheck had flagged as SC2289 (severity error) a day earlier:
# the lint workflow went red on the very push that introduced it (run 186) and
# stayed red for runs 187 and 188, unread.
#
# lint.yml deliberately does not run on tag pushes, and its reasoning is sound
# (the tagged tree was already linted on main; a tag-ref lint run sorts above
# the publish run and makes a release look finished before anything ships). The
# missing invariant was never "lint the tag" -- it was "do not RELEASE a tree
# whose lint failed", and only a job inside THIS workflow can enforce that.
#
# ~40 s, ahead of everything expensive, and it runs scripts/lint-shell.sh --
# the same file lint.yml calls, not a second copy that drifts.
lint-gate:
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
steps:
- uses: actions/checkout@v4
- name: Install shellcheck
run: |
apt-get update
apt-get install -y --no-install-recommends shellcheck
- name: "Shellcheck + syntax-check repository scripts (severity: error)"
run: bash scripts/lint-shell.sh
resolve-versions:
# Gated: a defective tree must not reach a 46-minute base build.
needs: [lint-gate]
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
@@ -543,7 +577,12 @@ jobs:
env:
EXPECTED_PI_VERSION: ${{ needs.resolve-versions.outputs.pi_version }}
EXPECTED_MEMPALACE_VERSION: ${{ needs.resolve-versions.outputs.mempalace_version }}
run: bash scripts/smoke-test.sh pi-devbox:smoke
run: |
# Single source of truth for the node major is Dockerfile.base's ARG.
# Asserting the BUILT image matches it also catches a stale cached layer.
EXPECTED_NODE_MAJOR=$(sed -n 's/^ARG NODE_VERSION=\([0-9][0-9]*\).*/\1/p' Dockerfile.base)
export EXPECTED_NODE_MAJOR
bash scripts/smoke-test.sh pi-devbox:smoke
# ── Phase 3b: amd64 smoke for the studio variant ────────────────────
# Additive + independent of the core `smoke` job: gates ONLY
@@ -606,7 +645,12 @@ jobs:
env:
EXPECTED_PI_VERSION: ${{ needs.resolve-versions.outputs.pi_version }}
EXPECTED_MEMPALACE_VERSION: ${{ needs.resolve-versions.outputs.mempalace_version }}
run: bash scripts/smoke-test.sh pi-devbox:smoke-studio
run: |
# Single source of truth for the node major is Dockerfile.base's ARG.
# Asserting the BUILT image matches it also catches a stale cached layer.
EXPECTED_NODE_MAJOR=$(sed -n 's/^ARG NODE_VERSION=\([0-9][0-9]*\).*/\1/p' Dockerfile.base)
export EXPECTED_NODE_MAJOR
bash scripts/smoke-test.sh pi-devbox:smoke-studio
# ── Phase 4: multi-arch publish ─────────────────────────────────────
build-variant:
+42 -27
View File
@@ -75,31 +75,11 @@ jobs:
# are shell scripts with no extension. -print0/mapfile -d '' so a path
# with a space cannot silently split, and the file count is asserted
# non-zero — a green tick over an empty file set is not a check.
run: |
# Union of two signals, because either alone misses a real case:
# a shebang scan misses a sourced fragment with no shebang, and a
# *.sh glob misses the extensionless tools in rootfs/usr/local/bin/.
# Silent skipping is precisely the failure mode this gate exists to
# prevent, so err toward over-collecting.
mapfile -d '' -t all_files < <(find . -not -path './.git/*' -type f -print0)
sh_files=()
for f in "${all_files[@]}"; do
case "$f" in *.sh) sh_files+=("$f"); continue;; esac
if head -n1 "$f" 2>/dev/null | grep -qE '^#!.*\b(bash|sh)\b'; then
sh_files+=("$f")
fi
done
echo "Checking ${#sh_files[@]} shell file(s)"
if [ "${#sh_files[@]}" -eq 0 ]; then
echo "::error::no shell files found — the shebang scan or the checkout is wrong"
exit 1
fi
shellcheck -S error -f gcc "${sh_files[@]}"
rc=0
for f in "${sh_files[@]}"; do
bash -n "$f" || { echo "::error file=$f::bash -n failed"; rc=1; }
done
exit "$rc"
#
# The implementation moved to scripts/lint-shell.sh on 2026-09-08 so the
# release gate in docker-publish.yml runs the SAME code rather than a
# second copy that drifts. Edit the script, not a copy of it.
run: bash scripts/lint-shell.sh
- name: Gitea shell guard (catches the actionlint blind spot)
# actionlint models GitHub Actions, where the default run shell is
@@ -112,7 +92,7 @@ jobs:
- name: Install actionlint (pinned)
env:
ACTIONLINT_VERSION: 1.7.7
ACTIONLINT_VERSION: 1.7.12
run: |
curl -fsSL \
"https://github.com/rhysd/actionlint/releases/download/v${ACTIONLINT_VERSION}/actionlint_${ACTIONLINT_VERSION}_linux_amd64.tar.gz" \
@@ -147,7 +127,7 @@ jobs:
- name: Install hadolint (pinned)
env:
HADOLINT_VERSION: 2.14.0
HADOLINT_VERSION: 2.15.1
run: |
curl -fsSL \
"https://github.com/hadolint/hadolint/releases/download/v${HADOLINT_VERSION}/hadolint-Linux-x86_64" \
@@ -157,3 +137,38 @@ jobs:
- name: Run hadolint
run: hadolint Dockerfile.base Dockerfile.variant
skill-floor:
# Gate the VENDORED pi-extensions skill snapshot in rootfs/ against the
# package repo it is a snapshot of. Its own job rather than a step in
# `actionlint`, so "the floor is stale" is a distinct red name in the runs
# list instead of being buried in a lint job that is about something else.
#
# The gap it closes, measured 2026-09-10: the floor sat at 34284 B, untouched
# since fa04d20 (2026-07-30), while the package copy was 38973 B.
# Dockerfile.variant copies the fresh package copy over the SERVED path but
# never writes back to the floor, so nothing in the repo ever noticed. That
# matters because the floor is a FALLBACK: the copy is guarded by
# `if [ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone
# yields no skill/ ships the vendored snapshot and still goes green, with no
# manifest flag or label saying which copy was served.
#
# Gating on another repo is normally a smell; it is proportionate here
# because the check compares the skill DIRECTORY hash, so it can only fire
# when that directory actually changed — which is exactly when the floor has
# gone stale. pi-extensions commits that leave skill/ alone cannot turn this
# red. No secret is needed either: the repo is anonymously clonable (verified
# 2026-09-10 with `git ls-remote` and no credentials), so this cannot start
# failing when a token expires.
#
# Exit codes are 0 in sync / 1 drift / 2 cannot-run, matching
# scripts/lint-shell.sh: a gate that cannot run must not pass, so an
# unreachable package repo is a red 2 rather than a green tick.
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
steps:
- uses: actions/checkout@v4
- name: Vendored pi-extensions skill floor matches the package
run: bash scripts/check-skill-floor.sh
+408 -3
View File
@@ -11,6 +11,405 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
---
## Unreleased
**`shellcheck` is now in the image, because the release gate it depends on could
not be run by anyone.** v1.8.14 made shell lint a release gate: `scripts/lint-shell.sh`
became the single source of truth for `lint.yml` and a new `lint-gate` job that
`resolve-versions` depends on, and it deliberately exits 2 when `shellcheck` is
absent — *a gate that cannot run must not pass*. Measured on v1.8.14 on
2026-09-09, by three routes (`command -v`, `dpkg -l`, a filesystem search):
**`shellcheck` was not in the image at all.** So `bash scripts/lint-shell.sh`
exited 2 in every devbox container, and the only place the gate could ever run
was CI. The developer loop was therefore write-shell → push → wait for CI →
discover — which is the loop the gate was added to shorten, after v1.8.14's first
attempt burned ~46 minutes on a tree whose lint had already been red for 24
hours. Added to the `apt-get` block in `Dockerfile.base`: `shellcheck 0.10.0-1`,
~39 MB installed (`Installed-Size` 40112 KB), and measured to pull **zero**
additional packages under `--no-install-recommends` because its three deps
(`libc6`, `libffi8`, `libgmp10`) are already present. **This forces one full base
rebuild** — `base-decide` hashes `Dockerfile.base` + `rootfs/`, so unlike a
`scripts/` change it cannot reuse the existing `base-` layer.
**A client-side pre-push lint gate: `hooks/pre-push`.** Opt-in per clone with
`git config core.hooksPath hooks`, bypass with `git push --no-verify`, matching
the idiom the `skillset` and `myconfigs` repos already use. It is a thin wrapper
that `exec`s `scripts/lint-shell.sh` — the same script CI runs, one copy, because
a duplicated check that drifts is the failure this repo keeps paying for (the
`pi-extensions` skill mirror sat 9579 B behind for weeks; the shell-lint logic
was extracted to one file for exactly this reason).
> **Why this repo had no hooks at all, which is worth stating because it was
> reported as drift and is not.** A fleet peer asked `tor-ms22` to report
> `git config core.hooksPath` per clone on the premise that an unset value meant
> "no secret-scan and no shell-lint hook locally", leaving the drift/secret gates
> unverified. Measured: `pi-devbox` **unset**, `skillset` `hooks`, `myconfigs`
> `common/hooks`, `pi-toolkit` **unset**. But `git ls-files | grep -i hook` is
> **empty** in both `pi-devbox` and `pi-toolkit` — neither repo tracked a single
> hook file, so there was nothing for `core.hooksPath` to point at on any machine
> and unset was the only correct value. The two repos that do ship hooks were
> already wired correctly. This entry closes the real half of that gap for
> `pi-devbox`; `pi-toolkit` still ships none.
**The hook is verified to catch the defect that motivated it, not merely to
exist.** Three measurements, each with the expected result written down first:
- **Refusal paths.** With `shellcheck` absent (the state of every container built
before this change) the hook exits **2** and names the remedy; with
`scripts/lint-shell.sh` missing it also exits **2**. It never waves a push
through on the assumption that CI will catch it.
- **The hook is actually in the scan set.** `lint-shell.sh` reports `Checking 14
shell file(s)` with `hooks/pre-push` present and **13** with it moved aside — so
the extensionless file is discovered by the shebang half of the linter's
two-signal union, rather than being silently skipped. This check exists because
the first attempt at it was ambiguous: a planted `[ $UNSET_VAR = "x" ]` was not
reported, which could equally have meant "file not scanned" or "defect below
`-S error`". It was the latter. A count that moves is unambiguous; a clean run
is not.
- **It catches the real v1.8.14 defect.** Planting the exact failing shape — an
apostrophe inside a single-quoted string, `echo 'the fleet\'s thing'` — in
`hooks/pre-push` produces `SC1073`/`SC1072` at severity **error** and `rc=1`.
That is the defect that closed a string, truncated an `exec_test` body, sent its
tail to the runner's shell, and cost a 46-minute build.
**The gate earned its keep inside the commit that added it.** The first version of
the `smoke-test.sh` assertion above carried a comment beginning `# shellcheck is a
GATE DEPENDENCY…`. A comment whose first word is the tool's name is parsed as a
**shellcheck directive**, not a comment, so the new gate immediately failed with
`SC1073`/`SC1072` at severity error — on the change that introduced it. Same family
as the v1.8.14 apostrophe: a line that reads as prose to a human and as syntax to
the parser. Before this change that defect would have been discovered in CI.
**Also queued, not yet pinned:** the `mempalace-toolkit` owed-set derivation now
honours a requester withdrawing its *own* ask (`isWithdrawn`, RFC 003 §3.3
clause 4, with `scripts/test-owed-withdrawal.sh`) — toolkit commit **`e2b060a`**,
which is the minimum revision for the behaviour. This image still pins `e45f6b4`.
Until an image bakes `e2b060a` or later, a sender must assume its withdrawal has
no effect on the recipient's mailbox — measured cost of the gap: a withdrawn
v1.8.13 rollout ask was still being reported as owed on `tor-ms22` 41 hours later,
for a release that device never installed. The same commit also anchors the
derivation's `mine` query at the newest end (`order: "desc"`); with the previous
default `asc` + `limit: 100`, a device passing 100 authored events would have its
recent replies fall out of the join window and see answered asks resurface.
**Four small packages, each chosen from a gap that was measured rather than
imagined.** All four were picked by looking back at a real session — the
`gitea.egl.lan`/FreeIPA debugging of 2026-09-09..10 — and asking which absences
actually cost time, not which tools sound useful. `bind9-dnsutils` (~6.1 MB, 10
packages): `dig`, `host` **and** `nslookup` were all absent, so the container
could resolve names but had no way to interrogate a *specific* nameserver —
`getent hosts` only follows the resolver's default path, so diagnosing "gateway
`172.16.88.1` NXDOMAINs the `egl.lan` zone while `10.20.253.1` is authoritative
for it" had to be hand-rolled in `python3`. Note the package name: plain
`dnsutils` is transitional in trixie. `ldap-utils` (1244 KB, **zero** extra deps):
the fleet authenticates against FreeIPA, yet every LDAP probe had to be run by
SSHing to an already-enrolled host; this gives simple binds only, since GSSAPI
would additionally need `krb5-user` + `libsasl2-modules-gssapi-mit`, which is a
Kerberos-client decision rather than a tool. `xxd` (198 KB) is frank convenience
— `od -c` already does the job. `python3-yaml` (552 KB, zero extra deps) is the
shellcheck story repeating exactly: `scripts/check-workflow-shell.sh`, the guard
against the Gitea `sh`/dash footgun that broke `resolve-versions` (`ed49b8d`) and
`promote-base-latest` (`b7197e8`), hard-exits with "python3 yaml module missing"
without it — and `lint.yml` installing it explicitly in CI was the evidence the
image lacked it. **`netcat-openbsd` was proposed and deliberately rejected**:
measured redundant, because `socat` is already baked and bash's `/dev/tcp` does
reachability checks with zero packages. The reason is recorded in
`Dockerfile.base` so the omission reads as a decision rather than an oversight.
**The vendored `pi-extensions` skill floor was 41 days stale, and is now gated so
it cannot silently rot again.** `rootfs/usr/local/share/pi-devbox/skills/pi-extensions/`
sat at 34284 B, untouched since `fa04d20` (2026-07-30), while the package copy
was 38973 B — four copies of one skill existed across the fleet with three
different sizes. `Dockerfile.variant` copies the freshly-cloned package copy over
the **served** path but never writes back to the repo floor, so nothing in the
repo ever noticed. That is worse than ordinary staleness because the floor is a
**fallback**: the copy is guarded by `if [ -f /opt/pi-extensions/skill/SKILL.md ]`,
so a build whose clone yields no `skill/` keeps the vendored snapshot and still
goes **green**, with no manifest flag and no label recording which copy was
served — the image would ship a July skill and nothing would say so. The floor is
refreshed here from `pi-extensions@c64c122`, and the new `skill-floor` job in
`lint.yml` runs `scripts/check-skill-floor.sh` to keep it that way.
The check compares the **directory** hash, using the same `tree_sha256` pipeline
`Dockerfile.variant` uses for `skillset_snapshot_tree_sha256` and for the same
documented reason: a `sha256sum SKILL.md` answers "did this one file change", not
"is this the same skill", and `pi-extensions` ships two files. That is not
hypothetical — it was **verified by negative control**: with `SKILL.md` left
byte-identical and only `evaluate-extension-usage.py` edited, the directory check
correctly fails while a file-only compare would have passed. Exit codes are `0`
in sync / `1` drift / `2` cannot-run, matching `scripts/lint-shell.sh`, so an
unreachable package repo is a red `2` rather than a green tick. Gating on another
repo is normally a smell; it is proportionate here because the check can only
fire when `skill/` itself changed — which is exactly when the floor has gone
stale — and it needs no secret, since `pi-extensions` is anonymously clonable
(verified with `git ls-remote` and no credentials).
> **What this does *not* fix, stated so nobody reads more into it than is there.**
> The floor is now fresh and guarded, but the *silent-fallback* half remains:
> if the build-time copy is ever absent, the build still succeeds with no
> manifest flag or OCI label recording that the vendored snapshot was served
> instead of the package copy. The durable fix for that is a manifest field
> alongside the existing `skillset_snapshot_tree_sha256`, which this change does
> not add.
**Four pinned dependencies bumped, after an audit of everything the image gets
from outside apt.** The audit itself is the useful part: of ~23 externally-managed
components, the 19 that resolve `latest` at build time were already current or
refresh themselves on the next rebuild, and the hard pins for `pi` (0.85.1),
`mempalace` (3.9.0) and `pi-atelier` (v0.10.1) were all already the newest
available. Only four needed a human.
**`NODE_VERSION` 22 → 24 (LTS "Krypton") — this one was a latent defect, not
housekeeping.** `agent-browser` publishes `engines.node ">=24.0.0"`, so the image
was *below a declared requirement*: v1.8.14 shipped node 22.23.2 with
`agent-browser` 0.37.1, meaning every build installed it with an npm `EBADENGINE`
warning and then ran the baked browser automation outside its supported range.
The other two npm consumers are satisfied either way — `pi` declares `>=22.19.0`,
`playwright` `>=20`. Verified before bumping, because a missing NodeSource suite
would break every architecture at once: `setup_24.x` returns HTTP 200 and the
`node_24.x` suite advertises `Architectures: amd64 arm64 armhf x86_64`, covering
both the arm64 fleet and the amd64 CI runners. Nothing else in the repo pinned the
node major.
**`actionlint` 1.7.7 → 1.7.12 and `hadolint` 2.14.0 → 2.15.1**, each run against
the current tree at the new version *before* being pinned — both clean, no new
findings. That ordering is the point: a linter bump is the one dependency update
that can turn CI red on unchanged code, so discovering it locally costs a minute
and discovering it in CI costs a round trip.
**`SKILLSET_SNAPSHOT_REF` `e9e09d9` → `4d7c0ea`**, via
`scripts/vendor-mempalace-skill.sh` rather than by hand, because that script is
the only thing that may write the ARG — a `cp` without a matching bump produces a
manifest that confidently lies. This turned out to be **provenance-only**: the
recorded ref was 6 commits behind, but `skills/mempalace/SKILL.md` is byte-identical
at both (`3675bfab…`), so the vendored snapshot was already correct and only its
recorded origin was stale. Consequently no `rootfs/` bytes changed, the
smoke-test phrase canary stays valid, and this ARG alone would not have forced a
base rebuild — the node bump does that anyway.
**The silent-fallback hole is closed: the image now records WHICH `pi-extensions`
skill copy it shipped.** This was the half deliberately left open by the
`skill-floor` gate above, and it is the more important half, because "the floor is
currently fresh" is a fact with a shelf life while "the image says which copy it
got" keeps working. The refresh step in `Dockerfile.variant` is guarded by
`if [ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates the
co-located skill kept the vendored floor and still succeeded **green**, with
nothing in the manifest, the labels or the logs distinguishing that from a normal
build. The two outcomes are indistinguishable by inspection afterwards — same
path, same filenames, same permissions — which is exactly how the floor went
unnoticed from 2026-07-30 to 2026-09-10.
`build-manifest.json` gains `pi_extensions_skill_source` and
`pi_extensions_skill_tree_sha256`, both **measured rather than passed in as
build-args**, per the ground-truth rule the rest of that block already follows —
and necessarily so here, since the outcome depends on the clone's contents and no
ARG could express it. Three values, because two would force a lie:
`package` (served bytes equal the clone's `skill/`), `vendored-floor` (the clone
had no `skill/` at this ref, so the fallback shipped), and `divergent` — both
exist but differ, e.g. the clone ships `SKILL.md` but not
`evaluate-extension-usage.py`, leaving the served directory a genuine **mix** of
package and floor. No OCI label mirrors these, deliberately: `LABEL` cannot take a
value computed in a `RUN`, and a label fed from an ARG would be precisely the
claim-not-measurement this change exists to remove.
Two `scripts/smoke-test.sh` assertions turn the record into a gate: one that the
source is named and is `package` — `vendored-floor` **fails** rather than warns,
since these images track `main` where the package has co-located `skill/` since
`fa04d20`, so a fallback means the clone did not resolve as intended — and one
that recomputes the tree hash over the served directory, because a recorded hash
that is never recompared is a claim rather than a measurement. `pi-devbox-version`
also annotates the line: `pi-extensions baked (package copy)` on the normal path,
and a yellow `(FALLBACK: vendored floor — clone had no skill/)` otherwise. Its
existing skill section reports which copy is being **read** at runtime; this is
the one fact that is decided at **build** time and cannot be recovered later.
Older images degrade cleanly — the field is absent, `jq // empty` yields nothing,
and the line prints plain `baked` exactly as before.
**This image also carries a real fix for the recurring
`[mempalace ext] feed (tick) failed: mine timed out after 30000ms` message** that
has been appearing in the pi TUI across the fleet since August
(`mempalace-toolkit` `309980b` + `e68ee20`, picked up because CI resolves
`MEMPALACE_TOOLKIT_REF` to a commit SHA at build time). It was parked as cosmetic
on 2026-08-27 and it was not cosmetic: `lastFeedAt` was recorded only after a
*successful* wait, but the extension's `Promise.race` abandons only the **wait**
and cannot cancel the mine, so a timeout left the 10-minute debounce clock stale
— and with `feedInFlight` already cleared, **both** guards stood open and every
following settled turn started another mine on top of the one still running.
Overlapping writers on a single-writer palace, each making the next slower and
the next timeout likelier, which is why the message appeared many times per
session instead of at most once per debounce window. Simulated over ten minutes
of settled turns with a 60s mine: **16 mines launched, 15 of them overlapping**
before; **2 and 0** after. Nothing was ever lost — the transcript is staged
before the mine and `mine --mode convos` is idempotent — so this was wasted work
and a misleading error, not data loss. The deadline also rose from 30s to 5
minutes: the mine is the slowest call the extension makes (30–60s normally) yet
carried the tightest deadline, 4x tighter than the `prepare` before it and 10x
tighter than the init handshake. On a healthy fleet the message should now be
absent; if it appears it is informative — a mine exceeding five minutes.
---
## v1.8.14 — 2026-09-08
> **First release attempt failed; fixed in this same entry.** The `smoke` and
> `smoke-studio` jobs both failed at `scripts/smoke-test.sh:770` with
> `agent-browser: command not found`, after `build-base` had already succeeded
> (~46 min spent). Root cause was in the agent-browser execution guard added the
> day before: the explanatory comment inside the **single-quoted** `exec_test`
> body contained an apostrophe (`the fleet\'s`). Inside `'...'` bash treats a
> backslash literally, so `\'` does not escape — it **closes the string**. The
> body silently truncated (measured: `exec_test` received **12** arguments
> instead of 2), and the remaining lines, including the `agent-browser --version`
> assertion, were parsed by the **runner's** shell instead of executing inside
> the image — and the runner has no agent-browser. The prose now lives above the
> call, where an apostrophe is harmless.
>
> **The lint job had already caught this, and it went unread for 24 hours.**
> `shellcheck` flagged it as `SC2289` at severity *error*, so the `actionlint`
> job went red at run 186 on 2026-09-07 21:21 — the exact push that introduced
> the guard — and stayed red for runs 187 and 188. `lint.yml` deliberately
> excludes tag pushes (documented: the tagged tree was already linted on main,
> and a tag-ref lint run would sort above the publish run), which is sound; the
> broken assumption was different, namely that a tree whose lint FAILED would not
> then be released. `docker-publish.yml` has no dependency on lint, so it built
> for 50 minutes on a tree known to be defective.
>
> **Fixed, then gated.** The prose moved above the `exec_test` call so an
> apostrophe cannot terminate anything, and the shell-lint logic moved out of
> `lint.yml` into **`scripts/lint-shell.sh`** — now called by both `lint.yml` and
> a new `lint-gate` job here that `resolve-versions` depends on. A release with a
> lint error refuses in ~40 s instead of failing after fifty minutes. One copy,
> not two: a duplicated check that drifts is the failure this repo keeps paying
> for. The script also refuses to pass when `shellcheck` is absent, inheriting
> the existing principle that a gate which cannot run must not pass.
>
> **`v1.8.14` was re-pointed** from `601fc98` to the fix commit. Nothing had
> consumed the original tag — no `v1.8.14` image was ever published, only the
> content-addressed `base-a365dd24de21`. `scripts/` does not feed the base hash,
> so the re-run reuses that base and skips the 46-minute rebuild.
**A test that was quietly checking nothing, and a version number that was wrong.**
Both found by delegating a read-only audit of this repo to a headless worker
(`pi-toolkit` `bin/pi-task`) and then spot-checking its pointers from the
filesystem — 5 of 5 held, and it also corrected a false premise planted in its
own brief.
**The node major is now asserted, not merely printed.**
`scripts/smoke-test.sh` ran `run "node" "node --version"`, which asserts only
that the binary exists and exits 0 — the printed version was compared to
nothing. The line above it has always used `run_expect` against
`$EXPECTED_PI_VERSION` for `pi`, so the suite *looked* like it covered node.
**A node major bump would have passed the whole smoke suite silently.** Worse,
this is where the "node v22.23.2 verified" line in the v1.8.13 recreate notes
came from: printed output, not an assertion — an expectation stated up front and
then falsified by the check.
Now gated on `EXPECTED_NODE_MAJOR`, which CI derives from `Dockerfile.base`'s
`ARG NODE_VERSION` — the single source of truth, and the *only* hard node pin in
the repo (`Dockerfile.variant` has no node install at all, so the two Dockerfiles
cannot disagree). That also catches a stale cached layer whose node disagrees
with the declared ARG. Unset ⇒ previous behaviour, so nothing breaks for anyone
running the suite by hand.
Verified two-sided, because a silent failure here reintroduces the exact bug it
fixes: the `sed` derivation yields `22` (an empty result would disable the
assertion silently); `grep -Fq "v22."` matches `v22.23.2`; `"v24."` does **not**
match, so a wrong major is caught; `"v2."` does not prefix-collide. The workflow
YAML was re-parsed after editing (9 jobs).
**v1.8.13's agent-browser version was wrong.** That entry said "the image's own
0.35.2". The image ships **0.36.0** — `/usr/lib/node_modules/agent-browser` at
0.36.0 with `engines.node >=24.0.0`, and no 0.35.2 exists anywhere in the image.
The sentence was also internally incoherent, contrasting 0.36.0 against a version
that is not present. Corrected in place with a visible note, since that entry is
already released. **The reasoning survives untouched**: the engines floor really
is vestigial, because `/usr/bin/agent-browser` is a prebuilt aarch64 ELF invoked
directly and never through node — which is exactly why 0.36.0 runs fine on
22.23.2, consistent with the runtime proof collected on 2026-09-07 and with the
retraction of the earlier false "0.36.0 requires node >= 24" alert.
No image content changes: `NODE_VERSION` still 22, no pins moved. This is a test
and a docs correction only.
**The same bug class, twice in one file — and the second one was throwing away a
proof the fleet cannot obtain any other way.** `scripts/smoke-test.sh`'s
agent-browser guard captured the version *inside an `echo`, with `2>/dev/null`*:
```sh
echo "resolved=[$r] version=[$(agent-browser --version 2>/dev/null | head -n1)]" >&2
```
The exit code was discarded, so a binary that could not execute at all still
**passed**, printing `version=[]`. Verified two-sided: a stub exiting 127 passes
the old form and is caught by the new one.
Why that exit code matters more than most: smoke runs `platforms: linux/amd64` on
an x86 runner, i.e. **native amd64**, making this line the fleet's only recurring
amd64 runtime proof for agent-browser's `linux-x64` ELF. **No devbox can ever
supply one** — every machine in the pi fleet is an Apple Silicon Mac
(`mbp-m1-2020`; `tor-ms22` = Mac Studio `Mac13,1` M1 Max, verified 2026-08-17 by
`system_profiler`; `emb-7kj4vr4g` = Apple Silicon, verified 4 ways 2026-09-07).
The "amd64 runtime proof still needed" item that was sent to two devices was
therefore asking for the impossible, while CI already had the answer and was
discarding it. `Dockerfile.base:607` does assert it (`agent-browser --version &&`),
but only when the base actually rebuilds — and v1.8.13's base was cached.
**The mailbox now announces replies that CLOSE your own asks.** `mempalace-toolkit`
`21023e7` → `e45f6b4`, which adds `deriveClosed()` alongside `deriveOwed()`. The old
path queried `status: open` and joined for a reply, which by construction can only
surface asks *you owe someone else*; a terminal reply carries `status: applied`
(or `blocked`/`failed`), so **the answer to your own question was structurally
invisible** — the one notification a human actually wants. Measured: `emb-7kj4vr4g`
closed the v1.8.13 rollout ask at 18:31Z with `status=applied`, the operator
reasonably expected to hear about it, and the mailbox stayed silent while being
correct by its own definition. Nine closed correlations were sitting unannounced.
Shares the 1-hour resurface floor, so a close is announced once and is news rather
than a nag.
This lands **because the base rebuilds**, which is worth stating explicitly: the
CI-resolved `mempalace-toolkit` SHA is folded into the content-addressed base tag
(`base-decide`), precisely so a toolkit-only fix cannot silently fail to land
behind an unchanged `Dockerfile.base`. The floating `main` ref was left alone on
purpose — the toolkit moving *forces* the rebuild rather than waiting for one.
**A consequence worth noting for the amd64 item above: this release actually
collects that proof.** v1.8.13's base was cached, which is why
`Dockerfile.base:607`'s `agent-browser --version &&` never ran. v1.8.14's base is
not cached, so both that assertion and the new `EXPECTED_NODE_MAJOR` gate execute
on a native `linux/amd64` runner. The fleet's first *kept* amd64 runtime proof for
the `linux-x64` ELF should be an artefact of this build rather than something
asked of a device that cannot supply it.
**Subtask delegation is documented — including the rung nobody built.** The image
picks these up through their resolved refs (`pi-toolkit` `adfb553`,
`pi-extensions` `c64c122`, the latter also refreshing the baked fallback skill):
- an operator-facing decision guide in pi-toolkit's `README.md`, built on the
L0–L4 context ladder — how much of the parent session a child can see is the
axis that explains nearly every observed good and bad behaviour;
- the canonical `pi-extensions` skill gains the same ladder next to *Boundary
discipline*, which until now diagnosed why an inherited transcript defeats a
brief without offering any alternative to "don't fork that";
- one bullet in the global `AGENTS.md`, so the choice is visible without loading
a skill, and naming `pi-task` as a **CLI** — an agent hunting for a `pi_task`
tool finds none and concludes it is unavailable.
What the ladder records: **L0/L1/L2 exist** in `pi-task` (`context.facts` /
`.files` / `.commands`), **L4** is `fork`'s only behaviour (`getHeader()` +
`getBranch()`, no offset or limit anywhere in the call chain), and **L3** — a
truncated branch — **is not implemented by anything**, which is now written down
instead of being a design idea somebody remembers.
Also recorded, found while writing the above: `pi-fork/src/runner.ts:188` reads
`if (extensions !== null) args.push("--no-extensions")`. So `extensions: []` turns
the capability floor **on** and `null` turns it **off** — and `null` is the
documented way to "restore normal extension loading", so tidying `[]` to `null`
as a no-op re-arms palace writes inside every fork child. `pi-task` hardcodes the
flag and cannot drift this way. Documented in three places because the edit that
triggers it looks harmless.
---
## v1.8.13 — 2026-09-06
**Version audit + three pins moved, one deliberately not moved.** `pi`
@@ -56,9 +455,15 @@ identically to the 0.84.4 control, so "alive" could be distinguished from
safe: all five prebuilt native addons in pi use NAPI (ABI-stable, no
NODE_MODULE_VERSION lock, no binding.gyp), nothing in the image declares a node
CEILING, and the install is one token (`setup_${NODE_VERSION}.x`). agent-browser
0.36.0 declares `engines.node >=24`, but that field is vestigial for the
artifact actually shipped: the image's own 0.35.2 declares the same floor and
runs fine on 22.23.2 as a prebuilt aarch64 ELF. The reason to wait is
0.36.0 declares `engines.node >=24.0.0`, but that field is vestigial for the
artifact actually shipped: `/usr/bin/agent-browser` is the prebuilt aarch64 ELF
`bin/agent-browser-linux-arm64`, invoked directly and never through node, so npm's
engines floor is never enforced at runtime — verified running under 22.23.2 in
this image. (Corrected 2026-09-07: this paragraph originally said "the image's own
0.35.2 declares the same floor". That was wrong and incoherent — it contrasted
0.36.0 against a 0.35.2 that does not exist in the image. There is exactly one
agent-browser present, `/usr/lib/node_modules/agent-browser` at 0.36.0. The
argument is unaffected; only the version was wrong.) The reason to wait is
attribution, not compatibility — this release already moves pi a minor,
mempalace a minor and bakes a Studio RC, so adding a node major would leave four
suspects if the image misbehaves. Worth doing as its own release with the smoke
+90 -1
View File
@@ -101,6 +101,80 @@ ENV DEBIAN_FRONTEND=noninteractive
# container: `ss` lands at /usr/bin/ss, `ip` at /usr/sbin/ip
# (both already on the developer PATH), and `portcheck --all`
# then correctly identifies the socat listener on 8765.
# shellcheck — shell linter. Added 2026-09-09 to close a CAPABILITY gap, not
# a style preference. `scripts/lint-shell.sh` is the release
# GATE (the `lint-gate` job that `resolve-versions` depends
# on), and it correctly refuses to pass when shellcheck is
# missing — "a gate that cannot run must not pass". Measured on
# v1.8.14: shellcheck was absent from this image by all three
# routes (PATH, dpkg, filesystem), so `bash
# scripts/lint-shell.sh` exited 2 in EVERY devbox container and
# no developer could run the release gate locally at all. The
# loop was therefore write-shell → push → wait for CI → discover,
# which is the loop the gate was added to shorten: v1.8.14's
# first attempt burned ~46 min on a tree whose lint had already
# been red for 24 h. This is also what makes a client-side
# pre-push hook possible (see hooks/pre-push); without the
# binary that hook would refuse every push. ~39 MB installed
# (Installed-Size 40112 KB, shellcheck 0.10.0-1) and measured
# to pull ZERO additional packages under
# --no-install-recommends: its deps (libc6, libffi8, libgmp10)
# are already present. NOTE this file feeds the base-decide
# hash (Dockerfile.base + rootfs/), so adding it forces one
# full base rebuild.
# bind9-dnsutils — `dig` and `nslookup`. Added 2026-09-10 to close a
# DIAGNOSTIC gap measured during the gitea.egl.lan/FreeIPA
# work: the container could resolve names but had NO way to
# ask a SPECIFIC nameserver anything. `getent hosts` only
# follows the resolver's default path, so the whole "gateway
# 172.16.88.1 returns NXDOMAIN for the egl.lan zone while
# 10.20.253.1 is authoritative for it" diagnosis had to be
# hand-rolled in python3 — dig, host AND nslookup were all
# absent. `dig @10.20.253.1 freeipa-4.egl.lan` is the
# one-liner that replaces it, and split-horizon DNS is a
# recurring class of bug on this fleet, not a one-off. NOTE
# the package to name is bind9-dnsutils: plain `dnsutils` is
# a transitional package in trixie. ~6.1 MB total (6210 KB
# measured): bind9-dnsutils 721 KB + bind9-host 161 KB +
# bind9-libs 3804 KB plus 7 small libs (libfstrm0,
# libjson-c5, liblmdb0, libmaxminddb0, libprotobuf-c1,
# liburcu8t64, libuv1t64) under --no-install-recommends.
# ldap-utils — `ldapsearch`/`ldapmodify`. Added 2026-09-10. This fleet
# authenticates against FreeIPA (EGL.LAN), and every LDAP
# probe during the Gitea auth work had to be run by SSHing to
# an already-enrolled host because the container had no LDAP
# client at all. 1244 KB and pulls NOTHING extra under
# --no-install-recommends — its deps (libldap, libsasl2) are
# already present. CAVEAT: this gives SIMPLE binds only,
# which is what Gitea itself uses and what most probes need.
# GSSAPI binds (`ldapsearch -Y GSSAPI`) additionally require
# krb5-user + libsasl2-modules-gssapi-mit, deliberately NOT
# added here — that is a Kerberos-client decision with
# /etc/krb5.conf implications, not just a tool.
# xxd — hex dump. 198 KB, no extra deps. Convenience, and honestly
# marginal: `od -c` from coreutils is always present and does
# the same job. Earned its place because verifying that
# git-crypt actually encrypted a staged blob (the \0GITCRYPT\0
# magic) is a recurring check in myconfigs and xxd is the
# muscle-memory command for it.
# NOT added — netcat-openbsd (133 KB): measured redundant on
# 2026-09-10, because socat is already baked above AND bash's
# /dev/tcp does reachability checks with zero packages
# (verified against gitea.egl.lan:3000). Recorded here so the
# omission reads as a decision rather than an oversight.
# python3-yaml — PyYAML. Added 2026-09-10 for precisely the same reason as
# shellcheck above: a gate this repo ALREADY OWNS could not be
# run locally by anyone. scripts/check-workflow-shell.sh — the
# guard that catches the "bash-only syntax under Gitea's default
# sh/dash shell" footgun that broke resolve-versions (ed49b8d)
# and promote-base-latest (b7197e8) — hard-exits with "ERROR:
# python3 yaml module missing" without it. lint.yml installs it
# explicitly in CI (`shellcheck python3-yaml`), which is itself
# the evidence that the image lacked it. Measured 2026-09-10
# while wiring the skill-floor job: the guard could not be run
# before pushing — the same write → push → wait-for-CI loop that
# shellcheck was baked to shorten. 552 KB, and pulls ZERO extra
# packages under --no-install-recommends.
RUN apt-get update && \
apt-get upgrade -y --no-install-recommends && \
apt-get install -y --no-install-recommends \
@@ -120,6 +194,7 @@ RUN apt-get update && \
make \
patch \
diffutils \
shellcheck \
git-crypt \
age \
file \
@@ -141,6 +216,10 @@ RUN apt-get update && \
kitty-terminfo \
ncurses-term \
iproute2 \
bind9-dnsutils \
ldap-utils \
xxd \
python3-yaml \
&& ln -s /usr/bin/fdfind /usr/local/bin/fd \
&& apt-get clean \
&& rm -rf /var/lib/apt/lists/*
@@ -554,7 +633,17 @@ ENV COLORTERM=truecolor
ENV PATH="/home/developer/.local/bin:/home/developer/.cargo/bin:${PATH}"
# ── Node.js (required for pi + MCP servers + tldr) ──
ARG NODE_VERSION=22
# 24 (LTS "Krypton"), raised from 22 on 2026-09-10 because the image was BELOW a
# DECLARED requirement, not merely behind the newest release: `agent-browser`
# publishes engines.node ">=24.0.0", so every build on 22 installed it with an npm
# EBADENGINE warning and then ran it outside its supported range — measured on
# v1.8.14, which shipped node 22.23.2 with agent-browser 0.37.1. The other two npm
# consumers are satisfied either way: pi declares ">=22.19.0" and playwright
# ">=20". Verified before bumping, because a missing NodeSource suite would break
# the build for every arch at once: deb.nodesource.com/setup_24.x returns HTTP 200
# and the node_24.x suite advertises `Architectures: amd64 arm64 armhf x86_64`, so
# both the arm64 fleet and the amd64 CI runners resolve.
ARG NODE_VERSION=24
RUN curl -fsSL --retry 5 --retry-delay 5 --retry-all-errors https://deb.nodesource.com/setup_${NODE_VERSION}.x | bash - && \
apt-get install -y --no-install-recommends nodejs && \
rm -rf /var/lib/apt/lists/*
+41 -1
View File
@@ -392,7 +392,7 @@ ARG MEMPALACE_TOOLKIT_REF=main
# no ~67-minute base rebuild. (scripts/check-base-hash.sh scans only
# Dockerfile.base, so no folding into the base hash is required — nor would
# it be correct, since this ARG changes nothing about the base's contents.)
ARG SKILLSET_SNAPSHOT_REF=e9e09d95f92670536a199fc986dfa24d787f18d1
ARG SKILLSET_SNAPSHOT_REF=4d7c0ea9caeb3a1d6d9b04cf34f3fca5f9df4985
# Dockerfile.base sets description="pi-devbox — base image (variant-independent)"
# and every variant INHERITS it, so both published images used to advertise
@@ -473,6 +473,44 @@ RUN set -e; \
if [ -d "$_snap_dir" ] && [ -n "$(find "$_snap_dir" -type f -print -quit)" ]; then \
SKILL_SNAP="\"$(tree_sha256 "$_snap_dir")\""; \
fi; \
# ── WHICH pi-extensions skill copy actually shipped ──
# Closes the silent-fallback hole. The refresh step above is guarded by
# `[ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates
# the co-located skill (or a fork pointing at a mirror without it) keeps the
# vendored floor and still succeeds — GREEN, with nothing anywhere recording
# that a snapshot shipped instead of the package copy. Measured 2026-09-10:
# the floor had been stale since 2026-07-30, so that fallback would have
# shipped a six-week-old skill silently. The floor is fresh now and gated by
# the skill-floor CI job, but "the fallback is currently harmless" is not the
# same as "you can tell which copy you got", and only the second survives.
#
# MEASURED, never claimed, per the ground-truth rule above: the branch
# condition is re-derived from the same test the refresh step used, and the
# served bytes are then compared against the clone. A build-arg could not
# express this at all, since the outcome depends on the clone's contents.
# package served bytes == the clone's skill/ (the normal path)
# vendored-floor the clone has no skill/ at this ref (fallback shipped)
# divergent both exist but differ — e.g. the clone ships SKILL.md but
# not evaluate-extension-usage.py, so the served directory is
# a MIX of package and floor. Worth its own value: it is the
# one state neither of the other two names honestly.
# No OCI label mirrors this, deliberately: LABEL cannot take a value computed
# in a RUN, and a label fed from an ARG would be exactly the claim-not-
# measurement this block exists to avoid.
_px_dir=/usr/local/share/pi-devbox/skills/pi-extensions; \
PIEXT_SRC='null'; PIEXT_HASH='null'; \
if [ -d "$_px_dir" ] && [ -n "$(find "$_px_dir" -type f -print -quit)" ]; then \
PIEXT_HASH="\"$(tree_sha256 "$_px_dir")\""; \
if [ -f /opt/pi-extensions/skill/SKILL.md ]; then \
if [ "$(tree_sha256 "$_px_dir")" = "$(tree_sha256 /opt/pi-extensions/skill)" ]; then \
PIEXT_SRC='"package"'; \
else \
PIEXT_SRC='"divergent"'; \
fi; \
else \
PIEXT_SRC='"vendored-floor"'; \
fi; \
fi; \
{ \
echo '{'; \
echo " \"release_tag\": \"${RELEASE_TAG}\","; \
@@ -493,6 +531,8 @@ RUN set -e; \
# vendored skill directory, not one file — see tree_sha256() above.
echo " \"skillset_snapshot_ref\": \"${SKILLSET_SNAPSHOT_REF}\","; \
echo " \"skillset_snapshot_tree_sha256\": ${SKILL_SNAP},"; \
echo " \"pi_extensions_skill_source\": ${PIEXT_SRC},"; \
echo " \"pi_extensions_skill_tree_sha256\": ${PIEXT_HASH},"; \
echo " \"components\": {"; \
echo " \"pi-toolkit\": \"$(rev /opt/pi-toolkit)\","; \
echo " \"pi-extensions\": \"$(rev /opt/pi-extensions)\","; \
Executable
+64
View File
@@ -0,0 +1,64 @@
#!/usr/bin/env bash
# Pre-push gate for pi-devbox: shellcheck every shell script before it leaves
# this clone. Thin wrapper — all logic lives in scripts/lint-shell.sh, which is
# the SAME script the CI release gate runs. One copy, not two: a duplicated
# check that drifts is the failure this repo keeps paying for.
#
# Install per clone: git config core.hooksPath hooks
# Bypass this gate: git push --no-verify (a guard, not a wall)
#
# WHY THIS HOOK EXISTS
# v1.8.14's first release attempt died at scripts/smoke-test.sh:770 after
# build-base had already spent ~46 minutes. shellcheck had ALREADY caught the
# defect — SC2289 at severity error, on the very push that introduced it — and
# the lint job stayed red for 24 hours, unread, across three runs. The fix at
# the time was to gate the release on the same script (the `lint-gate` job).
# This hook is the cheaper end of that: the same finding, before the push,
# in seconds rather than after a 40 s CI gate or a 46 min build.
#
# WHY IT COULD NOT EXIST UNTIL NOW
# Measured on v1.8.14 (2026-09-09): shellcheck was absent from the devbox
# image by all three routes — PATH, dpkg and a filesystem search. So
# lint-shell.sh exited 2 in every container, and a hook calling it would have
# refused EVERY push rather than gating anything. `shellcheck` was added to
# Dockerfile.base in the same change that added this file; on an image built
# before that, enable this hook and you will simply be told the gate cannot
# run. That is the correct behaviour, but it is not a working hook — so do not
# set core.hooksPath on a container older than the release that bakes it.
#
# NOTE ON SCOPE: this lints the WORKING TREE, not the exact commit range being
# pushed. That is deliberate and matches what the CI gate does to the tagged
# tree. It means a defect you have staged-but-not-committed is also reported,
# which is noisy in the safe direction.
set -euo pipefail
HOOK_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
REPO_ROOT="$(cd "$HOOK_DIR/.." && pwd)"
LINTER="$REPO_ROOT/scripts/lint-shell.sh"
tag="[lint-shell]"
# Same rule the gate itself applies, applied one level up: a missing check is
# not a pass. If the script is gone, the push is refused rather than waved
# through on the assumption that CI will catch it.
if [ ! -r "$LINTER" ]; then
echo "$tag refusing the push: $LINTER is missing, so the gate cannot" >&2
echo "$tag run. A gate that cannot run must not pass." >&2
exit 2
fi
# Point the message at the actual remedy when the binary is absent, because the
# linter's own message ("install it or run this in CI") is written for a CI
# runner and is misleading inside a container the developer cannot apt-install
# into persistently.
if ! command -v shellcheck >/dev/null 2>&1; then
echo "$tag refusing the push: shellcheck is not installed, so the gate" >&2
echo "$tag cannot run. A gate that cannot run must not pass." >&2
echo "$tag" >&2
echo "$tag This container predates the image that bakes shellcheck." >&2
echo "$tag Either recreate onto an image that has it, or unset the hook:" >&2
echo "$tag git config --unset core.hooksPath" >&2
echo "$tag To push this once without the gate: git push --no-verify" >&2
exit 2
fi
exec bash "$LINTER" "$REPO_ROOT"
+26
View File
@@ -157,6 +157,11 @@ if [ "$SHOW_SKILLS" = "yes" ] && [ -d "$BAKED_SKILLS" ] && [ -d "$SKILLS_DIR" ];
# is not hypothetical.
snap_ref=$(jq -r '.skillset_snapshot_ref // empty' "$MANIFEST")
snap_sha=$(jq -r '.skillset_snapshot_tree_sha256 // empty' "$MANIFEST")
# Which pi-extensions copy the BUILD baked. Distinct from everything else in
# this section, which reports which copy is being READ at runtime: for
# pi-extensions the baked tree is itself one of two possible copies, and that
# choice was made at build time and is not recoverable by inspection.
px_src=$(jq -r '.pi_extensions_skill_source // empty' "$MANIFEST")
# Same pipeline Dockerfile.variant uses to measure the baked directory at
# build time: relative paths in `find | sort` order, each hashed, the whole
# listing folded into one sha256. Keep the two definitions identical — they
@@ -187,7 +192,28 @@ if [ "$SHOW_SKILLS" = "yes" ] && [ -d "$BAKED_SKILLS" ] && [ -d "$SKILLS_DIR" ];
_target=$(readlink -f "$_link" 2>/dev/null || echo "$_link")
case "$_target" in
"$BAKED_SKILLS"/*|"$BAKED_SKILLS")
# "baked" alone used to be the whole story. For pi-extensions it is not:
# the baked tree holds EITHER the package copy that Dockerfile.variant
# lays over the snapshot, OR the vendored floor, when the clone had no
# skill/ at that ref. The two are indistinguishable by inspection — same
# path, same filenames, same permissions — so the build records which one
# it used and this reports it. Without this line a six-week-stale
# fallback skill looks exactly like a current one, which is precisely how
# the floor went unnoticed from 2026-07-30 to 2026-09-10.
if [ "$_name" = "pi-extensions" ] && [ -n "$px_src" ]; then
case "$px_src" in
package)
printf ' %-22s baked (package copy)\n' "$_name" ;;
vendored-floor)
printf ' %-22s baked \033[33m(FALLBACK: vendored floor — clone had no skill/)\033[0m\n' "$_name" ;;
divergent)
printf ' %-22s baked \033[33m(MIXED: part package, part floor)\033[0m\n' "$_name" ;;
*)
printf ' %-22s baked\n' "$_name" ;;
esac
else
printf ' %-22s baked\n' "$_name"
fi
continue
;;
esac
@@ -1,7 +1,7 @@
---
name: pi-extensions
description: >-
Use the pi extensions (pi-fork, pi-observational-memory, ssh-controlmaster) effectively in the pi coding agent harness. Load this skill only when running inside pi (detection - `fork` and `recall` are present in your tool list, or `pi --ssh` was used to start the session). pi-fork dispatches focused subtasks to forked agents at fast/balanced/deep effort tiers; pi-observational-memory compacts long sessions into recallable observations + reflections; ssh-controlmaster rewires pi's read/write/edit/bash tools to execute on a remote host over a multiplexed SSH connection. This skill covers tier selection, task design, boundary discipline, when to use recall, and remote-pi mechanics.
Use the pi extensions (pi-fork, pi-observational-memory, ssh-controlmaster) effectively in the pi coding agent harness. Load this skill only when running inside pi (detection - `fork` and `recall` are present in your tool list, or `pi --ssh` was used to start the session). pi-fork dispatches focused subtasks to forked agents at fast/balanced/deep effort tiers; pi-observational-memory compacts long sessions into recallable observations + reflections; ssh-controlmaster rewires pi's read/write/edit/bash tools to execute on a remote host over a multiplexed SSH connection. Also covers the context ladder L0-L4 and when to reach for the separate `pi-task` CLI instead of `fork` - isolated child, immutable spec, machine-checked envelope, write-boundary diff. This skill covers tier selection, task design, boundary discipline, when to use recall, and remote-pi mechanics.
---
# Pi Extensions: pi-fork, pi-observational-memory, ssh-controlmaster
@@ -161,6 +161,68 @@ The "three" things it completed were exactly the main thread's pending todos, vi
- Distrust **quantities** and **provenance claims** in fork prose specifically ("all N sessions", "shipped with the image", "as expected") — those are the slots confabulation fills.
- The fact that the fork was "right anyway" is not the same as the fork having followed instructions.
### The context ladder — and the second dispatch mechanism (`pi-task`)
Everything above describes a child that inherits everything. That is not a fixed
cost of delegation — **how much context a child gets is a choice**, and `fork`
sits at one extreme of it. Five rungs:
| rung | what the child sees | mechanism | built? |
|---|---|---|---|
| **L0** | nothing but the goal | `pi-task` default: fresh `--session-id pitask-<id>-<stamp>` in a private `--session-dir` | yes |
| **L1** | goal + **names** of files/commands to read itself | `pi-task` spec `context.files` / `context.commands` (`bin/pi-task:154,157`) | yes |
| **L2** | goal + an **excerpt the parent curated** | `pi-task` spec `context.facts`, pasted verbatim (`bin/pi-task:151`) | yes |
| **L3** | a **truncated tail** of the parent branch | *nothing implements this* — would need a new spec key plus `--session <trimmed snapshot>` | **no** |
| **L4** | the **entire** parent branch | `fork(task=…)` — `getHeader()+getBranch()`, no offset or limit anywhere in the call chain | yes |
**`pi-task` is a CLI, not an extension — it will never appear in your tool list.**
Invoke it with `bash`: `/opt/pi-toolkit/bin/pi-task run <spec.json>` (source at
`/workspace/pi-toolkit/bin/pi-task`, `schema` subcommand prints the spec fields).
It reads an immutable JSON spec, and "inherit the session" is not expressible in
that schema — the isolation is structural, not a request.
**Choose the lowest rung that can do the job:**
- **`fork` (L4)** when the subtask only makes sense against this conversation,
when you want several independent opinions in parallel from one message, or for
read-only exploration whose detail you will discard. Everything in "Boundary
discipline" above applies in full.
- **`pi-task` (L0–L2)** when the brief contains a **prohibition** (the inherited
transcript is exactly what overrides those), when you want a **pass/fail**
result instead of prose, when you need an **audit trail**, or when writes
outside an authorised set must be caught.
- **Neither** for trivial work, iterative work (both are one-shot), or judgement
that needs context only you have.
**What `pi-task` gets you that no brief can.** The envelope must parse or the run
FAILED, however fluent the prose. `roots[]` is the WATCHED set and
`write_allowed` the CHANGEABLE subset, diffed before and after with git
`--porcelain --ignored`. That `--ignored` flag is load-bearing: in the T4 test the
child obeyed its brief perfectly and still tripped the detector, because
`py_compile` wrote `__pycache__` into a watched-but-not-writable root — a
gitignored path that plain `--porcelain` reports as clean. Note the structural
point that test exposed: under `read_only: true` a write is *defiance*, so a
well-behaved child never produces a delta and the detector is never exercised.
Splitting WATCHED from WRITABLE is what lets an **obedient** child reveal a
violation, which is the realistic hazard.
**What it does not fix.** `--no-extensions` removes extensions, not the core
`read`/`write`/`edit`/`bash` tools — exactly as described above — so the boundary
diff is post-hoc **detection, not prevention**. And a fresh L0 context removes the
*narrative* failures (parent voice, invented continuity) without removing
confabulation: given an under-specified spec built on a false premise, the child
still filled the `deliverable` slot with a confident shape. The envelope's own
structure creates that pressure. Verify decisive claims from the filesystem
regardless of which rung you used.
**Trap — the capability floor is inverted from intuition.** `runner.ts:188` reads
`if (extensions !== null) args.push("--no-extensions")`. So `pi-fork.extensions:
[]` passes the flag and the floor is **on**; setting it to `null` — documented in
`settings.json` as the way to "restore normal extension loading" — passes nothing
and the floor is **off**, restoring palace writes inside every fork child.
Changing `[]` to `null` as a tidy-up re-arms what was deliberately disarmed.
`pi-task` hardcodes the flag and cannot drift this way.
### Anti-patterns
- **Forking trivial work.** A fork has overhead. If the task takes < 30 seconds in your main thread, just do it.
@@ -230,7 +292,7 @@ When entries conflict, **the most recent observation reflects the latest known s
## Quick Reference
```
fork(task=..., effort=fast|balanced|deep)
fork(task=..., effort=fast|balanced|deep) # L4: child inherits your WHOLE branch
- state decision authority explicitly
- pass verified context up front
- specify deliverable shape
@@ -240,6 +302,13 @@ fork(task=..., effort=fast|balanced|deep)
- write-capable? demand "What I did NOT do", then verify from git/fs, not the report
- prohibition in the brief => not a `fast` task
bash: /opt/pi-toolkit/bin/pi-task run <spec> # L0-L2: isolated child, NOT a tool
- schema | selftest | run [--dry-run]
- context.facts (pasted) / .files (names only) / .commands
- roots[] = WATCHED, write_allowed[] = CHANGEABLE subset
- envelope must parse or the run FAILED
- audit + cost: ~/.pi/agent/pi-task/<stamp>-<id>/result.json
recall(id=<12-char-hex>)
- only when stakes justify the cost
- id must already be visible in your context
+166
View File
@@ -0,0 +1,166 @@
#!/usr/bin/env bash
# check-skill-floor.sh — fail when the vendored pi-extensions skill snapshot in
# rootfs/ ("the floor") has drifted from the package repo it is a snapshot of.
#
# THE DEFECT THIS EXISTS TO CATCH, measured 2026-09-10.
# rootfs/usr/local/share/pi-devbox/skills/pi-extensions/ ships a vendored copy
# of the pi-extensions skill so the skill is ALWAYS present in the image.
# Dockerfile.variant then copies the freshly-cloned package copy OVER the served
# path at /usr/local/share/... — but it never writes back to the repo floor. So
# the floor only silently rots, and it had: 34284 B, untouched since fa04d20
# (2026-07-30), while the package copy was 38973 B. Four copies existed with
# three different sizes.
#
# Why that is worse than ordinary staleness: the floor is a FALLBACK. The copy
# step is guarded by `if [ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build
# where the package clone yields no skill/ keeps the vendored snapshot and still
# succeeds — green, with no manifest flag and no label saying which copy was
# served. The image would ship a July skill and nothing would say so. Keeping
# the floor fresh means that fallback is harmless instead of a silent regression.
#
# WHY A DIRECTORY HASH AND NOT `sha256sum SKILL.md`.
# The same pipeline Dockerfile.variant uses for skillset_snapshot_tree_sha256,
# and for the same documented reason: a file-only compare answers "did this one
# file change", not "is this the same skill". pi-extensions ships TWO files
# (SKILL.md + evaluate-extension-usage.py), so a sibling-file edit would pass a
# file-only check. If you change the pipeline here, change it there too.
#
# WHY GATING ON ANOTHER REPO IS PROPORTIONATE HERE, since that is normally a
# smell: this fires only when the package's skill/ DIRECTORY HASH changes, which
# is exactly and only when the floor has genuinely gone stale. pi-extensions
# commits that do not touch skill/ leave the hash alone and cannot turn this red.
# The repo is also anonymously clonable (verified 2026-09-10 with `git ls-remote`
# and no credentials), so this needs no secret and cannot break on token expiry.
#
# Exit codes — deliberately three, matching scripts/lint-shell.sh's philosophy
# that a gate which cannot run must not pass:
# 0 in sync (or the package legitimately has no skill/ at this ref)
# 1 DRIFT — the floor differs from the package
# 2 cannot run — no package copy could be obtained
set -euo pipefail
REPO_ROOT=$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)
FLOOR_DIR="${REPO_ROOT}/rootfs/usr/local/share/pi-devbox/skills/pi-extensions"
# Defaults mirror Dockerfile.variant's ARGs so this checks what the build builds.
PI_EXTENSIONS_REPO="${PI_EXTENSIONS_REPO:-https://gitea.jordbo.se/joakimp/pi-extensions.git}"
PI_EXTENSIONS_REF="${PI_EXTENSIONS_REF:-main}"
PACKAGE_DIR=""
WARN_ONLY=0
TMPDIR_CLONE=""
usage() {
cat <<'EOF'
Usage: scripts/check-skill-floor.sh [options]
--package-dir DIR Compare against an existing skill directory instead of
cloning. In a devbox container use /opt/pi-extensions/skill
for a fully offline run.
--warn-only Report drift but exit 0 (advisory use, e.g. a local hook).
-h, --help This text.
Environment: PI_EXTENSIONS_REPO, PI_EXTENSIONS_REF (default main) — both mirror
the Dockerfile.variant ARGs of the same name.
EOF
}
while [ $# -gt 0 ]; do
case "$1" in
--package-dir) PACKAGE_DIR="${2:-}"; shift 2 ;;
--warn-only) WARN_ONLY=1; shift ;;
-h|--help) usage; exit 0 ;;
*) echo "::error::unknown argument: $1" >&2; usage >&2; exit 2 ;;
esac
done
cleanup() {
if [ -n "$TMPDIR_CLONE" ]; then rm -rf "$TMPDIR_CLONE"; fi
}
trap cleanup EXIT
# Identical to Dockerfile.variant's tree_sha256(): relative paths + per-file
# sha256 over a sorted `find`, folded into one digest. Deterministic, never
# readdir order.
tree_sha256() {
( cd "$1" && find . -type f -print | LC_ALL=C sort | xargs -r sha256sum ) \
2>/dev/null | sha256sum | cut -d' ' -f1
}
if [ ! -d "$FLOOR_DIR" ]; then
echo "::error::floor directory is missing: ${FLOOR_DIR}"
echo "::error::rootfs/ is supposed to guarantee the skill is always in the image."
exit 2
fi
SOURCE_DESC=""
if [ -n "$PACKAGE_DIR" ]; then
if [ ! -d "$PACKAGE_DIR" ]; then
echo "::error::--package-dir does not exist: ${PACKAGE_DIR}"
exit 2
fi
SOURCE_DESC="local directory ${PACKAGE_DIR}"
else
command -v git >/dev/null 2>&1 || { echo "::error::git not found; cannot obtain the package copy."; exit 2; }
TMPDIR_CLONE=$(mktemp -d)
# Fetch the single ref shallowly. `git fetch <ref>` accepts a branch, a tag
# and (on Gitea) a reachable commit, which is why this is not `clone --branch`
# — CI resolves PI_EXTENSIONS_REF to a 40-hex SHA before the build.
if ! ( cd "$TMPDIR_CLONE" \
&& git init -q . \
&& git remote add origin "$PI_EXTENSIONS_REPO" \
&& git fetch -q --depth 1 origin "$PI_EXTENSIONS_REF" \
&& git checkout -q FETCH_HEAD ) 2>/dev/null; then
echo "::error::could not fetch ${PI_EXTENSIONS_REF} from ${PI_EXTENSIONS_REPO}"
echo "::error::Cannot determine whether the floor is stale, so this is exit 2, not a pass."
echo "::error::For an offline run, pass --package-dir /opt/pi-extensions/skill"
exit 2
fi
PACKAGE_SHA=$( cd "$TMPDIR_CLONE" && git rev-parse --short HEAD )
PACKAGE_DIR="${TMPDIR_CLONE}/skill"
SOURCE_DESC="${PI_EXTENSIONS_REPO} @ ${PI_EXTENSIONS_REF} (${PACKAGE_SHA})"
fi
# A ref with no skill/ is the documented fallback case: Dockerfile.variant keeps
# the vendored snapshot and the build succeeds. Nothing to compare, so this is
# not drift — but it IS the exact condition under which the floor ships, so say
# so loudly rather than printing a silent green tick.
if [ ! -d "$PACKAGE_DIR" ]; then
echo "::warning::package has no skill/ at this ref — the vendored floor is what will ship."
echo " source : ${SOURCE_DESC}"
echo " floor : $(tree_sha256 "$FLOOR_DIR")"
exit 0
fi
FLOOR_HASH=$(tree_sha256 "$FLOOR_DIR")
PKG_HASH=$(tree_sha256 "$PACKAGE_DIR")
if [ "$FLOOR_HASH" = "$PKG_HASH" ]; then
echo "OK: vendored pi-extensions floor matches the package."
echo " source : ${SOURCE_DESC}"
echo " tree_sha256: ${FLOOR_HASH}"
exit 0
fi
# `set -e` interacts badly with `[ … ] && x` as a bare statement, so both of
# these are explicit if-blocks rather than AND-lists.
LEVEL="error"
if [ "$WARN_ONLY" -eq 1 ]; then LEVEL="warning"; fi
echo "::${LEVEL}::vendored pi-extensions skill floor has DRIFTED from the package."
echo " source : ${SOURCE_DESC}"
echo " floor tree_sha256 : ${FLOOR_HASH}"
echo " pkg tree_sha256 : ${PKG_HASH}"
echo ""
echo " per-file differences:"
diff -rq "$FLOOR_DIR" "$PACKAGE_DIR" 2>&1 | sed 's/^/ /' || true
echo ""
echo " Remedy — re-sync the floor and commit it:"
echo " cp -a <pi-extensions>/skill/. ${FLOOR_DIR}/"
echo " git add ${FLOOR_DIR#"${REPO_ROOT}/"} && git commit"
echo ""
echo " NOTE this forces one full base rebuild: base_tag hashes Dockerfile.base"
echo " + rootfs/, and that rebuild is what re-bakes the refreshed floor."
if [ "$WARN_ONLY" -eq 1 ]; then exit 0; fi
exit 1
+91
View File
@@ -0,0 +1,91 @@
#!/usr/bin/env bash
# Shellcheck + syntax-check every shell script in this repo. Severity: error.
#
# SINGLE SOURCE OF TRUTH for two callers:
# .gitea/workflows/lint.yml — advisory, every branch push and PR
# .gitea/workflows/docker-publish.yml — the release GATE (lint-gate job)
# Extracted from lint.yml on 2026-09-08 rather than copied, because a second
# copy is exactly the drift this repo has been bitten by (see skillset's
# pi-extensions mirror, refreshed the same evening after sitting 9579 B behind).
#
# WHY THIS CHECK EXISTS AT ALL
# actionlint shellchecks workflow `run:` steps only. The repo's own scripts —
# entrypoint.sh, scripts/*.sh, and the extensionless tools under
# rootfs/usr/local/bin/ — were never shellchecked. A sibling repo with the same
# gap shipped a broken `echo "$json" | python3 <<'EOF' ... json.load(sys.stdin)`
# for two months: with no script argument python reads its SCRIPT from stdin,
# so the heredoc IS stdin and json.load hits EOF. shellcheck flags that at
# severity error (SC2259); nothing ever ran it.
#
# WHY THE RELEASE GATES ON IT (added 2026-09-08, the expensive way round)
# v1.8.14's first attempt failed after build-base had already spent ~46 min:
# scripts/smoke-test.sh had an apostrophe inside a single-quoted exec_test body
# ("the fleet\'s"), which CLOSES the string, so the body truncated and its tail
# ran on the CI runner instead of inside the image. shellcheck had already
# caught it as SC2289 at severity error — the lint job went red on the very
# push that introduced it and stayed red for 24 hours, unread. lint.yml
# deliberately does not run on tag pushes (sound: the tagged tree was linted on
# main, and a tag-ref lint run sorts above the publish run and makes a release
# look finished early). The gap was never "lint the tag" — it was that a tree
# whose lint FAILED could still be released. Hence a gate inside the publish
# workflow, ~40 s, ahead of everything expensive.
#
# SEVERITY CHOICE
# -S error is 0 findings across this repo when clean, so it is free to add.
# -S warning is NOT free here (19x SC2088 tilde-in-quotes in
# recreate-sanity-check.sh, plus assorted SC2016 — both intentional), and a
# noisy gate trains people to ignore it. Error-only, matching the
# SHELLCHECK_OPTS philosophy in lint.yml.
#
# Usage: bash scripts/lint-shell.sh [root] (default root: repo top level)
set -uo pipefail
root="${1:-}"
if [ -z "$root" ]; then
root="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
fi
cd "$root" || { echo "::error::cannot cd to $root"; exit 2; }
# A gate that cannot run must not pass. Without this, a machine (or a CI job
# whose install step was reordered away) without shellcheck would sail through
# printing nothing, which is the failure mode this whole file exists to prevent.
if ! command -v shellcheck >/dev/null 2>&1; then
echo "::error::shellcheck not found — the gate cannot run, so it must not pass" >&2
echo " install it (apt-get install -y shellcheck) or run this in CI" >&2
exit 2
fi
# Union of two signals, because either alone misses a real case: a shebang scan
# misses a sourced fragment with no shebang, and a *.sh glob misses the
# extensionless tools in rootfs/usr/local/bin/. Silent skipping is precisely the
# failure mode this gate exists to prevent, so err toward over-collecting.
# -print0/mapfile -d '' so a path containing a space cannot silently split.
mapfile -d '' -t all_files < <(find . -not -path './.git/*' -type f -print0)
sh_files=()
for f in "${all_files[@]}"; do
case "$f" in *.sh) sh_files+=("$f"); continue;; esac
if head -n1 "$f" 2>/dev/null | grep -qE '^#!.*\b(bash|sh)\b'; then
sh_files+=("$f")
fi
done
echo "Checking ${#sh_files[@]} shell file(s) with $(shellcheck --version | awk '/version:/{print $2}')"
# A green tick over an empty file set is not a check.
if [ "${#sh_files[@]}" -eq 0 ]; then
echo "::error::no shell files found — the shebang scan or the checkout is wrong"
exit 1
fi
rc=0
shellcheck -S error -f gcc "${sh_files[@]}" || rc=1
# bash -n catches a different class than shellcheck (unbalanced constructs it
# declines to parse), so both run and both count.
for f in "${sh_files[@]}"; do
bash -n "$f" || { echo "::error file=$f::bash -n failed"; rc=1; }
done
if [ "$rc" -eq 0 ]; then
echo "OK: ${#sh_files[@]} shell file(s) clean at severity error"
fi
exit "$rc"
+93 -2
View File
@@ -5,6 +5,7 @@
#
# Verifies:
# - pi binary present and (if EXPECTED_PI_VERSION set) matches CI's resolved version
# - node MAJOR matches Dockerfile.base's ARG NODE_VERSION (if EXPECTED_NODE_MAJOR set)
# - mempalace core matches the audited pin (if EXPECTED_MEMPALACE_VERSION set)
# - new v1.0.0 base additions (pandoc, graphviz, imagemagick, yq, tealdeer)
# - typst PDF engine for pandoc (v1.4.0) — `pandoc --pdf-engine=typst`
@@ -91,8 +92,31 @@ if [ -n "${EXPECTED_PI_VERSION:-}" ]; then
else
run "pi" "pi --version"
fi
run "node" "node --version"
# Until 2026-09-07 this was a bare `run "node" "node --version"`, which asserts
# only that the binary exists and exits 0 — the printed version was never
# compared to anything. A node major bump would therefore have passed this suite
# SILENTLY, while a reader skimming it would reasonably assume node regressions
# were covered. EXPECTED_NODE_MAJOR closes that: CI derives it from
# Dockerfile.base's ARG NODE_VERSION (the single source of truth), so this also
# catches a stale cached layer whose node does not match the declared ARG.
if [ -n "${EXPECTED_NODE_MAJOR:-}" ]; then
run_expect "node major matches Dockerfile ARG" "node --version" "v${EXPECTED_NODE_MAJOR}."
else
run "node" "node --version"
fi
run "git" "git --version"
# NOTE: the shellcheck binary is a GATE DEPENDENCY, not a convenience.
# scripts/lint-shell.sh is the release gate (the lint-gate job resolve-versions
# depends on) and it exits 2 when the binary is missing, by design — "a gate that
# cannot run must not pass". Measured on v1.8.14: it was absent from the image, so
# that gate could not be run by a developer in ANY container, only in CI.
# Asserted here so its absence fails a build instead of being discovered by a hook
# that then refuses every push (hooks/pre-push).
#
# This comment must not BEGIN with the tool's name: a line starting with
# `# shellcheck` is parsed as a DIRECTIVE, not a comment (SC1073/SC1072). The
# gate added in this same change caught that here, before the push.
run "shellcheck (lint gate dependency)" "shellcheck --version | grep -qE '^version: [0-9]'"
run "aws" "aws --version"
run "uv" "uv --version"
run "nvim" "nvim --version"
@@ -502,6 +526,50 @@ run "manifest skill fingerprint matches the baked snapshot" '
echo "manifest=[$m] actual=[$a]" >&2
[ -n "$m" ] && [ "$m" = "$a" ]
'
# ── Which pi-extensions skill copy shipped ──────────────────────────────
# Closes the silent-fallback hole. The refresh in Dockerfile.variant is guarded
# by `[ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates
# the co-located skill keeps the vendored floor and still succeeds GREEN, with
# nothing recording that a snapshot shipped instead of the package copy. Measured
# 2026-09-10: the floor had been stale since 2026-07-30, so that path would have
# shipped a six-week-old skill in silence. The floor is fresh now and gated by the
# skill-floor lint job, but "the fallback is currently harmless" is a fact with a
# shelf life, whereas "the image says which copy it got" keeps working.
#
# vendored-floor FAILS here rather than merely warning: these images track main,
# where the package has co-located skill/ since fa04d20, so a fallback means the
# clone did not resolve as intended and that is a defect to investigate. A fork
# deliberately pointing at a mirror without skill/ is the one case that should
# edit this assertion — which is the honest place for that decision to surface.
run "manifest names which pi-extensions skill copy shipped" '
j=/etc/pi-devbox/build-manifest.json
s=$(jq -r ".pi_extensions_skill_source // empty" $j)
h=$(jq -r ".pi_extensions_skill_tree_sha256 // empty" $j)
echo "source=[$s] tree_sha256=[$h]" >&2
printf "%s" "$h" | grep -qxE "[0-9a-f]{64}" || {
echo "pi_extensions_skill_tree_sha256 is not a 64-hex digest" >&2; exit 1; }
case "$s" in
package) ;;
vendored-floor)
echo "FALLBACK: clone had no skill/ at this ref, so the image ships the committed floor" >&2; exit 1 ;;
divergent)
echo "MIXED: served directory is part package and part floor" >&2; exit 1 ;;
*)
echo "pi_extensions_skill_source absent or unrecognised" >&2; exit 1 ;;
esac
'
# Same shape as the mempalace fingerprint check above, and for the same reason: a
# recorded hash that is never recomputed is a claim, not a measurement.
run "recorded pi-extensions skill hash matches the served bytes" '
j=/etc/pi-devbox/build-manifest.json
d=/usr/local/share/pi-devbox/skills/pi-extensions
m=$(jq -r ".pi_extensions_skill_tree_sha256 // empty" $j)
a=$( (cd "$d" && find . -type f -print | LC_ALL=C sort | xargs -r sha256sum) | sha256sum | cut -d" " -f1)
echo "manifest=[$m] actual=[$a]" >&2
[ -n "$m" ] && [ "$m" = "$a" ]
'
# OCI labels live in the image config, not the container fs — inspect them
# from the host docker rather than via `docker run`.
LBL=$(docker inspect --format '{{ index .Config.Labels "se.jordbo.pi-devbox.pi-extensions-ref" }}' "$IMAGE" 2>/dev/null || true)
@@ -739,10 +807,33 @@ exec_test "pi-atelier registered from /opt, not npm: (volume-shadowing guard)" \
# shadows it. The check that actually bites lives in
# recreate-sanity-check.sh, which runs where the volume is real — that is
# where a 7-week-old 0.27.0 was caught shadowing 0.35.2 on 2026-09-06.
# EXECUTION is ASSERTED here, not printed. Until 2026-09-07 the version was
# captured inside an echo with 2>/dev/null, so a binary that could not run at all
# still PASSED and simply printed version=[] -- the same failure class as the bare
# `node --version` two hundred lines up: a value displayed rather than compared.
#
# Why this exit code matters more than most: smoke runs `platforms: linux/amd64`
# on an x86 runner, i.e. NATIVE amd64, so this is the fleet's only recurring
# amd64 runtime proof for the linux-x64 ELF. No devbox can supply one -- every
# machine in the pi fleet is an Apple Silicon Mac (mbp-m1-2020; tor-ms22 = Mac
# Studio Mac13,1 M1 Max, verified 2026-08-17 by system_profiler; emb-7kj4vr4g =
# Apple Silicon, 4 routes 2026-09-07). Asking a device for that proof is asking
# for the impossible; CI already had it and was discarding it.
#
# KEEP PROSE OUT OF THE QUOTED BODY BELOW. On 2026-09-07 this explanation lived
# INSIDE the single-quoted argument and contained an apostrophe ("the fleet's").
# Inside '...' bash treats a backslash literally, so \' does not escape -- it
# CLOSES the string. The body silently truncated, the remaining lines were parsed
# by the RUNNER's shell instead of the container's, and `agent-browser --version`
# ran on a host that has no agent-browser: "line 770: command not found", release
# v1.8.14's smoke job failed after the base had already built. shellcheck caught
# it as SC2289 the same day and the red lint job went unread for 24h.
exec_test "agent-browser resolves under /usr (volume-shadowing guard, build-time half)" '
p=$(command -v agent-browser) || { echo "agent-browser not on PATH" >&2; exit 1; }
r=$(readlink -f "$p")
echo "resolved=[$r] version=[$(agent-browser --version 2>/dev/null | head -n1)]" >&2
v=$(agent-browser --version) || { echo "agent-browser did not EXECUTE" >&2; exit 1; }
test -n "$v" || { echo "agent-browser --version produced no output" >&2; exit 1; }
echo "resolved=[$r] version=[$(printf %s "$v" | head -n1)]" >&2
case "$r" in /usr/*) ;; *) exit 1 ;; esac
test ! -d "$HOME/.pi/npm-global/lib/node_modules/agent-browser" || exit 1
echo ok