Compare commits

...

42 Commits

Author SHA1 Message Date
Joakim Persson 12f99c49e3 docs(v1.10.0): adopt pi's fullscreen default and document the way back
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 15s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 14s
Publish Docker Image / lint-gate (push) Successful in 21s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 55m32s
Publish Docker Image / smoke (push) Failing after 8m50s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Publish Docker Image / smoke-studio (push) Failing after 9m55s
Publish Docker Image / build-variant-studio (push) Has been skipped
Decision: keep pi 1.0.0's default (tuiMode "fullscreen"). No tuiMode is baked,
so the image follows upstream rather than pinning the fleet to either mode -
but "we inherited a changed default" is only acceptable if the revert is
written down, so it now is.

README gains a "Terminal UI mode" section ahead of the pi-atelier section,
because the two interact. It gives all three scopes, taken from pi 1.0.0's own
docs/settings.md and docs/cli.md rather than from the changelog prose:

  - permanent: "tuiMode": "regular" in ~/.pi/agent/settings.json
  - one session: pi --tui-mode regular   (the flag is real: cli.md:221)
  - one project: the same key in .pi/settings.json, which overrides the agent
    directory

It also documents the four related settings (fullscreenExitOutput,
fullscreenScrollbar, fullscreenCopyOnSelect, fullscreenWheelScrollLines) and
two things specific to this image:

  - tmux: fullscreen uses the alternate screen, so tmux copy-mode shows the
    pane history AROUND pi, not pi's transcript. That is the concrete reason
    someone here would want "regular" back.
  - fullscreenWheelScrollLines "auto" behaves differently over SSH (caps fast
    wheel spins at 6 lines/event) than in a local macOS terminal (1 line) -
    worth naming because this container is normally driven over SSH.
  - pi-atelier works in BOTH modes (upstream has handled regular vs fullscreen
    renderers separately since pi 0.84), so reverting costs nothing. Stated so
    nobody assumes the sidebar is the price of the old scrollback.

The python3 merge snippet in that section was RUN before being documented:
against a realistic settings.json under an overridden HOME, twice, confirming
it is idempotent and preserves sibling keys including nested objects and the
_comment fields the seeded file uses. Documented code that has never been
executed is a guess.

DOCKER_HUB.md gets a short version of the same note, because CI reads it from
the TAG and POSTs it to Docker Hub as full_description - a behaviour change
this visible should not require reading the repo to undo.

Gates: doc-drift 23 OK / 0 DRIFT / 0 SKIP / 0 FAIL; base-hash, workflow-shell,
skill-floor, lint-shell all rc=0.
2026-10-02 09:00:51 +02:00
Joakim Persson 93eb3dbca7 docs(v1.10.0): name sqlite3/bc/dc/column in the Hub description tool list
Lint / skill-floor (push) Successful in 7s
Lint / actionlint (push) Successful in 18s
Lint / hadolint (push) Successful in 15s
Lint / doc-drift (push) Successful in 12s
DOCKER_HUB.md is POSTed to Docker Hub as full_description by the
update-description job, and CI reads it from the TAG, not from main. Its
"Modern CLI tooling" list said "Data: jq, yq" while this release adds sqlite3,
bc, dc and column to the base - so the public page would have shipped
incomplete for the whole v1.10.0 cycle and the fix could not land until the
next tag.

Caught by the pre-tag doc sweep rather than by a gate: check-doc-drift.sh
verifies version CLAIMS (pins, Node major) against the build files, but it
cannot know that a hand-written feature list grew stale, because nothing
declares that list's contents. Same failure shape as the v1.9.0 Node 22/24
mismatch that shipped eight releases running.

README carries no equivalent list, so there is nothing to mirror.
2026-10-02 08:20:01 +02:00
Joakim Persson e026f6f65e release(v1.10.0): pi 1.0.0 and pi-atelier v0.13.0, on measured evidence
Lint / hadolint (push) Successful in 10s
Lint / skill-floor (push) Successful in 12s
Lint / actionlint (push) Successful in 19s
Lint / doc-drift (push) Successful in 19s
Renames the unpublished v1.9.5 section to v1.10.0 and adopts the two bumps
that section had deliberately deferred. v1.9.5 was never tagged or published,
so nothing shipped under that name.

A minor, not a patch. v1.9.5 was numbered under the "patch - pi version bumps"
rule, but this release carries a pi MAJOR, a three-minor pi-atelier jump, four
new base packages and a new mailbox feature. The release carrying a 1.0.0
should not be the one numbered as a patch.

pi 0.85.1 -> 1.0.0. v1.9.5 held at 0.87.1 because obsmem's only compatibility
work names Pi 0.87 and no upstream issue mentions 0.99 or 1.0. That reasoning
had a hole worth naming: "no issue mentions 1.0" is an ABSENCE OF A STATEMENT,
not a measurement. So 1.0.0 was measured, against the published npm tarballs
for 0.87.1 and 1.0.0 unpacked side by side:

  - 1.0.0 has NO "### Breaking Changes" section at all. Those belong to 0.87.0,
    0.86.0, 0.84.3, 0.84.0, 0.83.0, 0.80.8, 0.80.7, 0.75.0 - none to anything
    between 0.88 and 1.0.0. The major is a milestone (fullscreen default,
    leaner codemode), not an API break.
  - finishTurn is still in 1.0.0's dist, so the pinned obsmem SHA keeps the API
    it migrated to.
  - All 8 pi.* APIs our extensions call exist in 1.0.0's dist: registerTool,
    registerCommand, registerFlag, getFlag, on, exec, sendMessage,
    sendUserMessage.
  - engines.node >=22.19.0 on both; the image ships 24.x.

The 0.99.2 change that looked fatal and is not: from 0.99.2 the DEFAULT MCP
exposure is `codemode`, so such tools are "neither declared to the model nor
listed" and must be found with searchTools(). That would gut the MemPalace
protocol if MemPalace were a builtin-MCP server. It is not - mempalace.ts and
mcp-loader.ts each run their own MCP client and register tools via
pi.registerTool(), which is why they are named mempalace_search and not
mcp__mempalace__search. Second route: settings.json has neither an mcpServers
block nor an mcp block. Filed upstream under "Changed", not "Breaking", so it
would have been easy to meet the hard way.

pi-atelier v0.10.3 (ed3837b) -> v0.13.0 (34d26f1), closing the gap v1.9.5
flagged as the pi bump's residual risk. peerDependencies are unchanged at pi
>=0.84.0 across v0.10.3/v0.12.1/v0.13.0 - a FLOOR, so not evidence of
anything. What was checked instead: v0.13.0 carries its OWN breaking change
(the entry point no longer exports the internal registry and layout helpers),
which reaches nothing of ours - grepping pi-atelier and registerSidebarPanel
across mempalace-toolkit, pi-devbox, pi-extensions, pi-fork,
pi-observational-memory and pi-studio returns ZERO matches in all six. v0.12.0
removed the showSessionActions setting (0 references here) and moved the
Control Center shortcut Alt+A -> F6 (0 references here, so no doc goes stale).
showSidebarAgent/showSidebarTodos, the only atelier keys README names, are
both still in v0.13.0's README. TuiMainScreen and renderLayoutFrame both still
exist in pi 1.0.0's dist.

PI_ATELIER_VERSION is bumped with PI_ATELIER_REF; it is a separate ARG and the
image label would otherwise have lied.

v0.13.0 is an ANNOTATED tag: `ls-remote --tags` reports the tag object
(dd06971), not the commit (34d26f1). check-doc-drift.sh resolves the commit, so
that is what the CHANGELOG names. The gate caught the wrong SHA on the first
attempt - noted in the CHANGELOG for future bumps.

NOT PROVEN, and stated so acceptance does not mistake it for cleared: grepping
dist shows the SYMBOLS survive, not that their SIGNATURES are unchanged -
necessary, not sufficient. 1.0.0 is four days old and neither obsmem nor
atelier has a commit naming it. Acceptance must prove the obsmem workers CAP
TURNS (peerDeps are *, so a mismatch is silent) and that the atelier sidebar
PAINTS.

User-visible behaviour change, deliberately NOT overridden: 1.0.0 makes the TUI
fullscreen by default, replacing the terminal's normal scrollback.
tuiMode: "regular" restores the old behaviour. No baked default is set, so the
image inherits upstream's choice rather than silently pinning the fleet.

Gates: check-doc-drift 23 OK / 0 DRIFT / 0 SKIP / 0 FAIL; base-hash,
workflow-shell, skill-floor, lint-shell all rc=0.
2026-10-02 08:16:37 +02:00
Joakim Persson 9d0b3dec0b release(v1.9.5): pi 0.87.1 with pi-obsmem pinned to the merged finishTurn fix
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 19s
Lint / doc-drift (push) Successful in 12s
v1.9.4 held pi at 0.85.1 because 0.87.0 REMOVED `shouldStopAfterTurn`, which
pi-observational-memory 3.1.4 still used in all three workers. Its peerDeps are
`*`, so nothing refuses at install time and the breakage is silent at runtime --
turn caps ignored, workers losing their specialised prompts. PR #83 fixes that
and merged 2026-09-23.

Re-measured 2026-10-01 across src/agents/{observer,reflector,dropper}/agent.ts,
deliberately reusing the v1.9.4 audit's counting method so the numbers compare:

  e7d77dc (3.1.4, baked in v1.9.4): shouldStopAfterTurn 3, finishTurn 0, systemPrompt 3
  731c3d4 (pinned here)           : shouldStopAfterTurn 0, finishTurn 3, systemPrompt 0

All three workers migrated, and the 0.86.0 AgentContext.systemPrompt reads are
gone too. The coupling is asymmetric and that is why these move in ONE commit:
3.1.4 + 0.87.1 silently ignores turn caps, and 731c3d4 + 0.85.1 breaks the
workers outright, because finishTurn does not exist before 0.87.0.

Pinned to a SHA rather than waiting for a tag, departing from the v1.9.4
instruction to wait for a release: #83 is merged but the newest obsmem tag is
still 3.1.4, cut 2026-09-20, BEFORE the merge. Upstream tags slowly and moves
master often (e7d77dc -> 1529e14 -> 731c3d4 in nine days), so waiting means
holding pi indefinitely. A pinned SHA keeps the property `master` lacks:
rebuilding this tag later produces the same image.

The 40-char form is load-bearing, not pedantry. check-doc-drift.sh recognises a
literal SHA only through a 40-char match, so a 7-char pin would fall through to
its branch-or-tag lookup, fail to resolve, and downgrade pi-obsmem's drift check
to a silent SKIP -- a pin that reads correctly and is no longer verified. Full
gate run after this change: 23 OK, 0 DRIFT, 0 SKIP, 0 FAIL.

0.99.0/0.99.1/0.99.2 and 1.0.0 all exist upstream and are deliberately skipped:
obsmem's only compatibility work names Pi 0.87 (2b1dc1c) and a repo-wide
issue/PR search for 0.99 or 1.0 returns zero matches. 1.0.0 is its own round.

pi-atelier stays at v0.10.3 and is the residual risk. Its peerDeps declare pi
>=0.84.0 -- a floor, satisfied -- on both v0.10.3 and the current v0.12.1, so it
spans this bump. But atelier hooks pi TUI internals that a declared floor does
not protect, and an under-declared floor is exactly what failed to warn anyone
at pi 0.84. Acceptance must confirm the sidebar PAINTS, using the two-sided
check from 0.84.4/0.85.1 that distinguishes "loaded" from "silently absent".

Also here:
- check-doc-drift.sh gains a pi-obsmem pin check. The ref-move check covers the
  same component, but for a pinned SHA it can only answer "upstream did not
  move", never "the table still says what we bake".
- scripts/lint-shell.sh mode 100644 -> 100755. Pre-existing since f25efa0 and
  the only non-executable script in scripts/; it was latent because both CI
  steps call it as `bash scripts/lint-shell.sh`, but it failed rc=126 "bad
  interpreter" when invoked directly. Same dropped-exec-bit signature recorded
  on 2026-09-22, found the same way: by RUNNING it, not by reading a diff.
- Unreleased section renamed to `## v1.9.5 — 2026-10-02`, satisfying the
  release-gate rule that a tag's CHANGELOG must name its own version.
2026-10-02 00:23:28 +02:00
Joakim Persson cb8969ef2c changelog: adopt pi-studio v0.9.61 and mempalace-toolkit 975ab92
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 18s
Lint / doc-drift (push) Failing after 17s
Lint / skill-floor (push) Has been cancelled
Both are floating refs that moved since v1.9.4, so the next tag would bake them
either way; the gate's complaint was that nothing in the repo said so. Naming
them is the fix the gate actually asks for.

mempalace-toolkit 2167a1b -> 975ab92 adds mailbox dormancy: an ask may declare a
`dormant_unless` predicate and is withheld from the ANNOUNCED owed-set while
every condition still matches its baseline. It is backward compatible by
construction -- an ask without the key behaves exactly as before -- and
fail-visible: any predicate that cannot be evaluated announces the ask rather
than hiding it, because the dangerous failure is work that disappears, not a
spurious nag.

pi-studio v0.9.60 -> v0.9.61 is adopted as-is. NOT pinned, deliberately, and the
reason is worth recording: check-doc-drift.sh gives pi-studio kind `studio`,
which resolves the highest semver TAG of the upstream repo and ignores
PI_STUDIO_REF entirely (resolve_studio, ~line 551). Setting
PI_STUDIO_REF=v0.9.61 would therefore not be tracked by the gate -- the moment
upstream tags v0.9.62 the gate would report DRIFT against a value we no longer
bake, i.e. a false positive by construction. Pinning pi-studio is a coherent
thing to want, but it requires moving it from kind `studio` to kind `ref` in the
gate at the same time, which changes release-gate semantics and is a separate
decision from adopting this release.

Still drifting, and left drifting on purpose: pi-obsmem e7d77dc -> 731c3d4.
PI_OBSMEM_REF=master is CI-resolved at build time, master is 4 commits ahead of
the 3.1.4 tag, and the newest tag remains 3.1.4 -- so PR #83 ("Fix Pi 0.87
memory worker compatibility", merged 2026-09-23) is still reachable only from
master. Whether to bake an unreleased master or stay on 3.1.4 and hold pi is a
pin decision with a silent failure mode behind it (obsmem's peerDeps are all
`*`), so it is not being made in a changelog commit.
2026-10-02 00:01:24 +02:00
Joakim Persson 31f36182bf feat(base): add sqlite3, bc, dc and bsdextrautils
Lint / skill-floor (push) Successful in 10s
Lint / doc-drift (push) Failing after 13s
Lint / actionlint (push) Successful in 20s
Lint / hadolint (push) Successful in 21s
sqlite3 is the only one of the four that was genuinely missing rather than
merely absent. MemPalace keeps both the palace and the logstream as SQLite
files (palace/chroma.sqlite3, logstream.sqlite3) and mempalace_status reports a
sqlite_integrity block, so every integrity or forensics check on this fleet has
so far gone through a python3 -c one-liner because the CLI was not in the image.

bc and dc are low value on their own and the CHANGELOG says so plainly: awk,
python3 and perl are all already baked and each is strictly more capable. They
are here because copy-pasted shell snippets assume bc exists. What actually went
wrong on 2026-09-28 was NOT bc's absence: printf '%6.2f' was handed the empty
output of the missing bc and rendered it as a confident "0.00 days" for a figure
that was really 4.81 days. No package fixes that failure mode -- only not taking
a formatted number on trust does.

bsdextrautils ships /usr/bin/column.

Cost MEASURED rather than estimated: 1311 KB total (587 + 236 + 149 + 339) and
ZERO transitive packages. libsqlite3-0, libreadline8t64, zlib1g, libsmartcols1
and libtinfo6 each already report "install ok installed", so nothing new is
pulled under --no-install-recommends.

Package names were verified with apt-cache and dpkg -S on a real trixie host
instead of being inferred, which caught two traps that would each have produced
either a build failure or a silently missing binary: dc is a SEPARATE binary
package from bc on Debian, and column ships in bsdextrautils, not in the
pre-bullseye bsdmainutils where it used to live.

Declined at the same time, recorded so the omission reads as a decision rather
than an oversight: datamash and xsv/csvkit. Python's stdlib csv module handled a
real 7-file Excel-export concatenation that day -- UTF-8 BOM, no trailing
newlines, and bare CR/LF inside quoted fields -- correctly and without them.

Dockerfile.base is in the base hash, so the next tag rebuilds the base (~64 min)
regardless of what else it carries.
2026-10-01 23:31:03 +02:00
joakimp 2278b22ba7 fix(ssh): wire git core.sshCommand to the writable sidecar, and assert it
Lint / hadolint (push) Successful in 13s
Lint / skill-floor (push) Successful in 14s
Lint / doc-drift (push) Successful in 18s
Lint / actionlint (push) Successful in 29s
~/.ssh is commonly bind-mounted READ-ONLY from the host, so a per-host
`ControlPath ~/.ssh/cm/%r@%h:%p` — the standard CGNAT multiplexing recipe, and
correct on the host — resolves inside an unwritable dir in the container. Every
push dies `unix_listener: cannot bind to path ...: Read-only file system`,
behind git's misleading "make sure you have the correct access rights".

setup-lan-access.sh already writes the fix: ~/.ssh-local/config overrides
ControlPath into the writable ~/.ssh-local/cm BEFORE `Include ~/.ssh/config`,
so -F repairs the socket path and keeps every per-host User/Port/IdentityFile.
entrypoint-user.sh now points git at it, guarded on the sidecar existing —
setup-lan-access.sh writes none on native Linux Docker, where -F at a missing
file would break every git-over-ssh call instead of fixing one. An existing
core.sshCommand is left alone (first-wins, as for the three git settings above).

Why code and not another doc line: the remedy was already in the global
AGENTS.md, in pi-devbox-environment SKILL.md §3, in 24 MemPalace drawers from
three devices, and printed verbatim by recreate-sanity-check.sh — and an agent
that had run that script two hours earlier still hit the failure and reinvented
a /tmp/sshcm workaround. A fifth copy was not the missing piece.

Assertions, each where it can actually pass:
- smoke-test.sh: two STATIC greps (wiring line + its [ -r ] guard). `run` uses
  --entrypoint="", so asserting the runtime value there would repeat the v1.8.0
  mistake of an assertion that cannot pass, unvalidated until the next tag.
- smoke-test.sh runtime phase: a BICONDITIONAL — sidecar present => must route
  through it; absent => must be unset. The absent arm is the one CI exercises
  (native Linux runner), so "is set" would have failed CI for a correct image.
- recreate-sanity-check.sh: the runtime assertion, plus an explicit fail for the
  inverted state (set while the sidecar is missing). The permanent "default ssh
  precedence" warning keeps its severity but now states that it is structural and
  can never reach zero, and whether git is wired, unwired, or has no sidecar.

All five arms exercised against the real script before commit; that caught a
defect in the first draft, which reported "git IS wired ... unaffected" about a
state where the sidecar was gone and every git-over-ssh call failed.

Host ~/.ssh/config needs no change: the same line is right on the host and
unusable through a read-only mount, so the fix belongs in the container layer.
2026-09-22 23:08:59 +02:00
joakimp 37fcbfcf04 fix: three v1.9.4-acceptance findings (installer WARN, init sentinel, hash collation)
Lint / hadolint (push) Successful in 10s
Lint / skill-floor (push) Successful in 9s
Lint / actionlint (push) Successful in 20s
Lint / doc-drift (push) Successful in 18s
Unreleased; no pin moves except pi-extensions master 25c1265 -> 143a214.

1. entrypoint-user.sh: first-run sentinel is now config.json, which is what
   `mempalace init` writes. The old test on palace/ (created by mining, not
   init) re-fired "Initializing MemPalace" on every boot of a container that
   never mined locally: v1.9.4 acceptance measured 1 line after one boot, 2
   after a restart. Idempotent, so harmless; the comment lied. This file is a
   base-hash input, so the next tag rebuilds the base (326fb7c03949 predicted
   with the pinned sort below; f40c4b7b103d before).

2. Dockerfile.variant: `git config --system --add safe.directory` for the
   seven root-owned /opt clones (listed, not `*`). Since git 2.35.2 any git
   command in a repo owned by another user fails "dubious ownership"; as
   `developer` that made `git -C /opt/pi-atelier rev-parse` print nothing
   (one false FAIL in the v1.9.4 acceptance) and is the root cause of the
   per-boot "WARN: pi-extensions install.sh failed" (fixed at the source in
   pi-extensions 143a214; this is the belt to that suspender). Verified live
   in a v1.9.3 container: add -> rev-parse prints 25c1265, unset -> fatal
   again; /etc/gitconfig restored to its 5 lines afterwards.

3. docker-publish.yml: `find -print0 | LC_ALL=C sort -z` in base-decide.
   sort collates per locale; identical rootfs hashed to base-f40c4b7b103d
   under C/C.UTF-8 (== run 695) and base-d8df62216a81 under sv_SE/en_US.UTF-8.
   The runner exports LANG=C.UTF-8, so the hash was stable by accident. Pin is
   hash-neutral: pinned sort + HEAD inputs reproduces f40c4b7b103d exactly.
   Per-command prefix only; nobody's locale changes.

CHANGELOG: new `## Unreleased` above v1.9.4 naming pi-extensions 143a214
(check 9 rc=0), the three fixes, and the synlig 3.10.0 hub upgrade that
v1.9.4 listed as still open. One invented URL org (gwpl) caught before commit;
upstream is elpapi42, as Dockerfile.variant:151 says.

Gates: check-doc-drift rc=0, check-base-hash rc=0, lint-shell rc=0 (16 files),
hadolint 2.15.1 rc=0, YAML parses (10 jobs), bash -n rc=0.
2026-09-22 17:20:10 +02:00
joakimp 16fddebd43 docs(v1.9.4): the client/server skew is narrower than the release commit said
Lint / hadolint (push) Successful in 8s
Lint / skill-floor (push) Successful in 11s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 21s
Publish Docker Image / resolve-versions (push) Successful in 15s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-base (push) Successful in 1h4m42s
Publish Docker Image / smoke (push) Successful in 6m4s
Publish Docker Image / smoke-studio (push) Successful in 9m18s
Publish Docker Image / build-variant (push) Successful in 19m20s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / build-variant-studio (push) Successful in 24m15s
e3b38cd named "the feeder against the 3.9.0 hub" as the one path to exercise
before tagging, on the reasoning that 3.10.0's CLI writes now follow the daemon
write-routing policy. That was a reading of the mempalace changelog, not a
measurement of this image's feeder. Measured now:

  mempalace-pi-session in remote mode (MEMPALACE_REMOTE_URL set, which is how
  every fleet devbox runs) stages transcripts locally in python, rsyncs them to
  the hub host, and calls the hub's own `mempalace_mine` MCP tool over HTTP
  (run_remote_mine). The local `mempalace` CLI is required only in local mode
  (line 475) and invoked only on the local branch (line 1029). With a PATH shim
  logging every `mempalace` invocation, a 47-session `--dry-run` from this
  container logged ZERO calls; the shim's positive control logged one. The pi
  extension likewise speaks HTTP to the hub and spawns no local mempalace-mcp.

So the 3.10.0 CLI never runs against the 3.9.0 hub from this image. In remote
mode the client pin touches first-run `mempalace init` and the on-disk layout,
nothing else -- which is exactly what the ENV MEMPALACE_CONFIG_DIR change is
for. Both the Dockerfile.base comment and the CHANGELOG paragraph now say so,
and the "Still open" item becomes the thing that is actually unmeasured: a
first-boot acceptance of the built image from an empty ~/.mempalace volume.

Gates re-run on this tree: check-doc-drift.sh rc=0, lint-shell.sh rc=0,
hadolint 2.15.1 rc=0. lint.yml for e3b38cd: run 693, completed, success.
2026-09-22 14:51:22 +02:00
joakimp e3b38cdb0b release(v1.9.4): mempalace 3.10.0 with a pinned palace root, pi-atelier v0.10.3, pi held at 0.85.1
Lint / skill-floor (push) Successful in 8s
Lint / hadolint (push) Successful in 10s
Lint / doc-drift (push) Successful in 13s
Lint / actionlint (push) Successful in 22s
Three pins move or are deliberately held; the CHANGELOG's Unreleased section
becomes the v1.9.4 entry in this same commit, because since ee6cb9e the tag
build runs check-doc-drift.sh check 9 and every would-bake value must be named
above the last published heading.

mempalace 3.9.0 -> 3.10.0 (Dockerfile.base). Deferred at v1.9.3 for two
"Upgrade notes" items; both re-measured against the 3.10.0 wheel:

  - New installs resolve ~/.config/mempalace. The order is $MEMPALACE_CONFIG_DIR,
    then ~/.mempalace IF it holds config.json / people_map.json /
    palace/chroma.sqlite3, then XDG. An EMPTY ~/.mempalace fails that test, and
    an empty ~/.mempalace is what a freshly mounted devbox-palace volume (or
    entrypoint.sh's mkdir on a volume-less container) looks like at first boot.
    Measured with a fresh $HOME and `uvx mempalace==3.10.0 init`: config.json
    landed in ~/.config/mempalace, ~/.mempalace stayed empty, and
    entrypoint-user.sh's `[ ! -d ~/.mempalace/palace ]` would fire every start.
    Same run with MEMPALACE_CONFIG_DIR set: everything in ~/.mempalace.
    => ENV MEMPALACE_CONFIG_DIR=/home/${USER_NAME}/.mempalace, placed in the
    non-root-user section where USER_NAME is in scope, and entrypoint-user.sh
    reads PALACE_DIR from the same variable (old path as fallback). Existing
    volumes were safe either way; this is for first boots. MEMPALACE_PALACE_PATH
    is still honoured (config.py:927), so smoke-test.sh's stage test holds.
  - event_list defaults newest-first without a cursor. Server-side: the
    extension talks to synlig over MEMPALACE_REMOTE_URL, so it lands when the
    hub upgrades. mempalace-toolkit 2167a1b (floating main) makes every
    cursor-less call say order:"desc" so the mailbox reads the same window on
    either server version.

Corrected a stale sequencing comment while there: synlig serves the palace as
a `uv tool` under systemd (mempalace-serve.service; measured 2026-09-22:
mempalace 3.9.0, python 3.12.13, chromadb 1.5.9, no palace container), not via
docker-compose.mempalace.yml as the comment had said for two releases.

pi-atelier v0.10.1 -> v0.10.3 (Dockerfile.variant). Fixes only (both tags
released 2026-09-22); package.json at v0.10.3 still has zero runtime deps and no
build script, peerDeps still >=0.84.0. Annotated tags: the CHANGELOG names the
peeled SHA ed3837b, which is what resolve-versions bakes and check 9 compares.

pi HELD at 0.85.1 while 0.86.1 / 0.87.0 exist. 0.87.0 removed the
shouldStopAfterTurn agent option; pi-observational-memory 3.1.4 still uses it
and AgentContext.systemPrompt in observer/reflector/dropper (measured: 1 hit
per worker in src/agents on both the baked 3.1.3 tree and upstream master;
finishTurn 0). peerDeps are `*`, so it would break at runtime, not install.
Upstream #82, fix PR #83 open and unmerged. Bump pi and pi-obsmem together
once #83 ships. The other extensions are clean against the 0.86/0.87 lists.

Verified locally the way CI does: check-doc-drift.sh rc=0 with check 9
reporting all four moved components as named above the v1.9.3 heading;
lint-shell.sh rc=0; hadolint 2.15.1 (CI's pin, with .hadolint.yaml) rc=0 on
both Dockerfiles; PyPI 3.10.0 present, not yanked. Two claims caught by
re-measurement before commit and corrected in the text: a pi issue number
that does not exist on the 0.87.0 changelog line, and "8 shouldStopAfterTurn
hits" that was 3 (one per worker) in src/.

Still open before the tag: run the feeder from the built image against the
3.9.0 hub (dry-run first); recreate acceptance must show ~/.mempalace/palace
still passing and no ~/.config/mempalace appearing. After acceptance: upgrade
synlig's tool to 3.10.0.
2026-09-22 14:26:10 +02:00
joakimp ee6cb9e62a ci(release): a tag must not publish a component no CHANGELOG entry names
Lint / skill-floor (push) Successful in 9s
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 18s
Lint / doc-drift (push) Successful in 16s
lint-gate already existed to enforce "do not RELEASE a tree whose lint failed",
because lint.yml does not run on tag pushes. scripts/check-doc-drift.sh had the
same gap and it was never extended to cover it: check 9 ran only in lint.yml, so
no tag build has ever evaluated it. Add it to lint-gate, which resolve-versions
already needs, so it fails in ~8 s ahead of the 46-minute base build.

For this check the gap is strictly worse than it is for shellcheck. Shellcheck
judges the tree, so green on main is still green at the tag -- the bytes did not
move. Check 9 judges the tree against upstream NOW, and the floating refs it
watches move with no commit here at all, so a green reading on main carries no
information about tag time. v1.9.3 is the worked example: pi-observational-memory
moved cba0334 -> e7d77dc the day AFTER the tag, nothing went red, and it
surfaced only because someone ran the gate by hand.

Measured, not assumed:
- Same file lint.yml calls (one reference in each workflow), not a second copy.
- Works on a CI-shaped checkout: cloned --depth 1 --no-tags (0 tags, 1 commit),
  rc=0. It needs no local tags because `last` comes from the Hub tags API, not
  `git tag`, so the plain actions/checkout@v4 above is sufficient.
- Has teeth: deleting the Unreleased section from that clone gives rc=1 and
  names the component; restoring it gives rc=0.
- ~8 s (7.8-8.3 s measured), vs ~1 s for lint-shell.sh.

Residual, accepted: the gate resolves the refs seconds before resolve-versions
resolves them again, so an upstream push inside that window still slips past.
Check 9 on the next release names it then.
2026-09-21 23:23:29 +02:00
joakimp 0d324f1855 changelog: name pi-obsmem 3.1.4 (e7d77dc) as the next rebuild's implicit adoption
Lint / skill-floor (push) Successful in 13s
Lint / hadolint (push) Successful in 15s
Lint / doc-drift (push) Successful in 14s
Lint / actionlint (push) Successful in 23s
check-doc-drift.sh was rc=1: pi-observational-memory moved cba0334 (3.1.3) ->
e7d77dc (3.1.4) upstream on 2026-09-20, the day after the v1.9.3 tag, and
nothing named it above the v1.9.3 heading.

No code change is needed or wanted here. PI_OBSMEM_REF defaults to master in
Dockerfile.variant and CI's resolve-versions turns that into a SHA at build
time, so the next rebuild adopts this whether or not anyone acts -- which is
precisely why it has to be written down. There is no pin in this repo to bump,
so the CHANGELOG is the only record a reader of the next tag would have.

v1.9.3 itself has no defect: its manifest, image labels and baked
/opt/pi-observational-memory clone all agree on cba0334. The drift is
post-tag, so a green drift reading taken on 2026-09-19 was correct when taken.

Measured, not assumed: the range cba0334...e7d77dc is 4 commits whose only
substance is one line in src/agents/worker-stream.ts (preserve registry
receiver when resolving stream, PR #78 / issue #77) plus a 19-line regression
test; the rest is the version bump and release merge. peerDependencies are
unchanged (all four @earendil-works/* peers still *) and there is no engines
block, so no pi or node floor to clear.

Gate now rc=0: "OK pi-obsmem cba0334 -> e7d77dc since v1.9.3, named above the
v1.9.3 heading".
2026-09-21 23:15:11 +02:00
Joakim Persson 6c13f43ac8 release: date the v1.9.3 section for the tag
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 16s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 41m49s
Publish Docker Image / smoke (push) Successful in 8m45s
Publish Docker Image / smoke-studio (push) Successful in 9m8s
Publish Docker Image / build-variant (push) Successful in 17m42s
Publish Docker Image / update-description (push) Successful in 11s
Publish Docker Image / promote-base-latest (push) Successful in 26s
Publish Docker Image / build-variant-studio (push) Successful in 21m28s
CI checks out the tag, so the CHANGELOG heading has to name the version
before the tag exists (same convention as v1.9.1/v1.9.2). check-doc-drift
check 9 was sabotage-tested for exactly this rename and still passes:
everything above the published v1.9.2 heading is the pending text.

Contents of v1.9.3 (all measured, nothing inherited):
  - f25efa0  recreate-sanity-check resolves the ControlPath ssh actually
             uses (ssh -G); the ~/.pi/ssh/config hyphen typo deleted
  - b8d818e  Hub size claims corrected; check 8 gates them against full_size
  - c7d369f  promote-base-latest's conditional re-tag measured (run 669)
  - 50153e6  skill floor -> pi-extensions 25c1265 (task tool + fork-gate)
  - ea89605  changelog for the two floating-ref changes: pi-extensions
             25c1265 and mempalace-toolkit 817b3a8 (mine deadline)
  - cb6d9e5  check 9: what the next build bakes differently must be named
             in the CHANGELOG; first run found pi-obsmem 7b397f4->cba0334
  - b9057fd  LABEL se.jordbo.pi-devbox.mempalace-version (inherited)
  - 960aada  dependency audit 2026-09-19; mempalace 3.10.0 deferred

Floating refs this tag will bake, named per check 9: pi-toolkit 9c87ee8,
pi-extensions 25c1265, mempalace-toolkit 817b3a8, pi-observational-memory
cba0334 (3.1.3). Base REBUILDS (Dockerfile.base + rootfs/ changed), so the
BASE_REBUILD_DATE marker is refreshed now, when it is free -- it had been
stale since 2026-07-13 through two rebuilds.

Deliberately NOT in this tag: mempalace 3.10.0 (see the audit section for
the three measured preconditions).
2026-09-19 17:45:50 +02:00
Joakim Persson 960aada769 changelog: dependency audit 2026-09-19 — mempalace 3.10.0 measured and deliberately deferred
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 17s
Lint / skill-floor (push) Successful in 8s
Lint / doc-drift (push) Successful in 13s
Every row by direct command: npm/PyPI/GitHub release APIs, git ls-remote,
the v1.9.2 config labels via the registry API, and binaries in a running
v1.9.2 container for the floating tools. Only pin with a newer upstream is
mempalace 3.9.0 -> 3.10.0, and it is NOT a small bump: new installs move to
~/.config/mempalace (entrypoint-user.sh tests ~/.mempalace/palace and
entrypoint.sh persists ~/.mempalace, so a fresh container would initialise
outside the volume), and mempalace_event_list flips to newest-first by
default (the toolkit's deriveClosed passes no order, deriveOwed on 2 of 3
calls — server-default-dependent). Deferred to its own release with the
three measured preconditions listed.
2026-09-19 17:38:02 +02:00
Joakim Persson b9057fdc8c Dockerfile.base: LABEL se.jordbo.pi-devbox.mempalace-version, inherited by both variants
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 10s
Lint / doc-drift (push) Successful in 16s
Lint / actionlint (push) Successful in 21s
Closes the blind spot check 9 (cb6d9e5) named: no label recorded the palace
pin, so a MEMPALACE_VERSION bump — the one component whose skew against the
shared central palace is fleet-wide — could ship without a CHANGELOG line.

In Dockerfile.base, not Dockerfile.variant, deliberately: the value sits next
to the ARG that defines it (a copy in the variant is one more pin able to
drift); labels are inherited by every image built FROM the base, so no
build-arg to plumb through the variant's four call sites; and inheritance
means the label states the pin of the base the image ACTUALLY built on,
which is the question when base-decide cache-hits an older base. Both
mechanisms measured on the published v1.9.2 config blob rather than assumed:
maintainer + image.source (set only in Dockerfile.base) are present on the
variant image, and pi-version=0.85.1 is an ARG expanded inside a LABEL.

Intent, like every se.jordbo.pi-devbox.* label; the manifest's
mempalace_version (read from the installed binary) stays the ground truth,
and smoke-test.sh now asserts label == installed core — the one way they
diverge is a base built with INSTALL_MEMPALACE=false or an off-pin install,
both invisible to a label-only check.

check-doc-drift check 9 gains the component (literal, against
ARG MEMPALACE_VERSION in Dockerfile.base); the label-key rule generalises to
"names ending in -version are the label itself". Until a release carries the
label it reports a counted SKIP, not OK — measured: "v1.9.2 carries no
se.jordbo.pi-devbox.mempalace-version label", summary says 1 SKIPPED.

Costs nothing extra: this Unreleased already forces a base rebuild
(50153e6 rootfs/ skill floor). check-base-hash unchanged (no new *_REF).
2026-09-19 17:27:34 +02:00
Joakim Persson cb6d9e5dd0 check-doc-drift: check 9 — what the next build bakes differently must be named in the CHANGELOG
Lint / hadolint (push) Successful in 8s
Lint / skill-floor (push) Successful in 10s
Lint / doc-drift (push) Successful in 14s
Lint / actionlint (push) Successful in 19s
The two most recent Unreleased entries (pi-extensions 25c1265: task tool +
fork-gate; mempalace-toolkit 817b3a8: the mine deadline that never reached
the transport) reached this image through floating *_REF=main ARGs. Neither
produced a diff here, so nothing asked for a CHANGELOG line, and neither had
one until a reader asked. Same class as check 8 — a claim with no in-repo
anchor rots — and the hand practice that covered it (the per-release
"Dependency audit" table, Baked-in vs Upstream-now) is a someone-remembers
mechanism that had lapsed.

Method, no docker/crane/token: last published vX.Y.Z from Hub's tag list
(one request, now shared with check 8); its amd64 config blob via the
anonymous registry API carries one se.jordbo.pi-devbox.<name>-ref label per
component with the SHA the build-args baked. "Would bake" is resolved the way
resolve-versions does it: 40-hex ARG is itself, branch/tag via git ls-remote
with the peeled ^{} form preferred, pi-studio = highest semver tag (label on
<tag>-studio), PI_VERSION as a literal against the pi-version label. Nine
components, ~7.5 s. Unchanged: OK. Moved: the new value's 7-char prefix (tag
for studio, version for pi) must appear above the last published version's
"## " heading — Unreleased plus any not-yet-published section, which is what
the release commit turns Unreleased into, so the tag build passes on the same
text. A published tag with no heading is a failure, not a skip.

First run found a move nobody had recorded: pi-observational-memory 7b397f4
-> cba0334 (6 commits, 3.1.1 -> 3.1.3; memory workers now routed by the
model's exact provider instead of the first registered provider with a
matching api — upstream #70). Named in the CHANGELOG with the honest scope:
no expected effect on the shipped bedrock-only configuration, unmeasured.

Sabotage-tested, expectation written before each run:
  obsmem SHA removed from CHANGELOG ............ rc=1  DRIFT
  Unreleased renamed to "## v1.9.3 — …" ........ rc=0  (release-commit shape)
  "## v1.9.2" heading mangled to v1.9.2-typo .... rc=1  — this one FAILED first:
      \b accepted "v1.9.2-typo" as v1.9.2's heading; now (\s|$)
  bare "## v1.9.2" (no date) ................... rc=0
  ARG PI_OBSMEM_REPO renamed ................... rc=2  (blind gate must not pass;
      read_arg calls are bare top-level assignments on purpose — nested in a
      heredoc's $(...) the exit 2 is swallowed by cat)
  offline (proxy to a closed port) ............. rc=0, 2 SKIPs counted
  SKIP_REF_CHECK=1 ............................. rc=0, 1 SKIP counted
  restored ..................................... rc=0, files byte-identical
GIT_TERMINAL_PROMPT=0 on ls-remote: a repo flipped private fails in 0.3 s as
a SKIP instead of waiting for a username until the job times out.

Header, lint.yml job comment and AGENTS.md all still said this gate "needs
no network" — false since check 8 (2026-09-14). Now: two classes, hermetic
1-7 and published-state 8-9 that SKIP loudly and counted when offline.
Considered and not added: compose <-> .env.example variable cross-check (the
four mismatches are commented-out lines, the mempalace-server compose file's
own variables, and entrypoint-consumed ones — it would fire on nothing
wrong); "documented tag exists on Hub" (check 8 already SKIPs by name; a
hard fail would misreport the tag-to-publish window).
2026-09-19 17:15:12 +02:00
Joakim Persson ea896054df changelog: task tool + fork-gate (pi-extensions 25c1265), and the mine deadline that never reached the transport (mempalace-toolkit 817b3a8)
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 18s
Lint / doc-drift (push) Successful in 6s
Both land in the image through floating refs (PI_EXTENSIONS_REF=main,
MEMPALACE_TOOLKIT_REF=main), so neither shows up as a diff in this repo — the
changelog is the only place a reader of a pi-devbox tag learns that the
container's delegation and feed behaviour changed. Recorded under Unreleased
with the measurements: 5/5 real fork briefs redirected, live pi -p proof for
gate and tool, 2-of-6 → 6-of-6 on the toolkit's new timeout test.

Dockerfile.base: the "Stall protection" comment listed MEMPALACE_MCP_TIMEOUT_MS
as THE tool-call deadline; it now says the feed's mine carries its own.
Comment-only, but Dockerfile.base is hashed into base_tag; the rebuild was
already forced by 50153e6 (rootfs/ skill floor).
2026-09-19 17:00:59 +02:00
Joakim Persson 50153e65b7 skill floor: refresh vendored pi-extensions skill to pi-extensions@25c1265 (task tool + fork-gate)
Lint / hadolint (push) Successful in 11s
Lint / doc-drift (push) Successful in 10s
Lint / skill-floor (push) Successful in 10s
Lint / actionlint (push) Successful in 21s
check-skill-floor.sh: OK, tree_sha256 9b85a633… matches the package at main (25c1265).
Forces one base rebuild (base_tag hashes rootfs/), which is also what bakes
extensions/task.ts and fork-gate.ts into /opt/pi-extensions via the floating
PI_EXTENSIONS_REF=main; install.sh symlinks every extensions/*.ts on start.
2026-09-19 16:42:06 +02:00
Joakim Persson f25efa074d fix(sanity): verify the ControlPath ssh RESOLVES, not the dir the image creates
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 13s
Lint / doc-drift (push) Successful in 12s
Lint / actionlint (push) Successful in 24s
recreate-sanity-check.sh asserted `/tmp/sshcm exists with mode 700` and printed
a green tick while every ssh in the container died rc=255. It was right about
what it checked: the breakage was in the directory a CONFIG named, not the one
the image creates, and the old check could not see the disagreement between them.

Found on the v1.9.2 first boot on emb-7kj4vr4g. A durable ~/.pi/ssh/config,
hand-written into the ~/.pi named volume by the previous session so it would
survive the recreate, declared `ControlPath /tmp/ssh-cm/%C` — with a hyphen.
Nothing here creates that path; the canonical spelling is /tmp/sshcm, identical
in Dockerfile.base, entrypoint-user.sh, recreate-sanity-check.sh and
smoke-test.sh. ControlMaster auto with an unusable ControlPath does not degrade
to an unmultiplexed connection — it fails hard:

    unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory

rc=255, remote command never runs.

Now: resolve the ControlPath ssh itself would use via `ssh -G` and require its
parent to exist and be writable. -G applies real config precedence (first-value-
wins, the system drop-in, Include, an -F override), so it answers "which rule
captured this host" instead of re-implementing the guess, and it never opens a
connection — 0.116s for 48 hosts.

This also puts a check under a caveat documented in prose in Dockerfile.base
("SSH client defaults") and verified nowhere: a per-host ControlPath under a
read-only bind-mounted ~/.ssh gives the identical failure, "Read-only file
system". On the machine this was built on that is 16 of 50 hosts — freeipa-1..6,
gitea.egl.lan, runner-1..3, tor-ms22 — never reported by anything before.

Severity split is deliberate. Default ssh precedence legitimately lands in the
read-only ~/.ssh for any host whose own config pins it there, and the supported
workaround (ssh -F ~/.ssh-local/config, generated every start by
setup-lan-access.sh) already exists, so that is a warn. Failing it would paint
the script red on every run of every device, and a check that fires benignly
every time is one you learn to ignore — the same reasoning that keeps
lint-shell.sh at -S error. The sidecar route is prescribed, so there it is a
hard fail. Host lists cap at six names plus a count: unreadable output is
ignored output.

Teeth proven both directions, each sabotage confirmed by diff BEFORE the result
was believed: ControlPath -> nonexistent dir => rc=1; -> existing-but-read-only
dir => rc=1; restored => rc=0 with the sidecar byte-identical.

No config was added to fix the original problem — the fix was to DELETE
~/.pi/ssh/config, a third hand-maintained copy of what setup-lan-access.sh
already generates from version control on every container start, with fewer
features and one typo. The host-owned ~/.ssh/config is left alone on purpose:
those ~/.ssh/cm paths are correct on the host, where ~/.ssh is writable.

Also lint-shell.sh: the SC2088 count in the severity-choice rationale said 19
and is now 20, with the reproduce command recorded. That command needs a `$ `
prefix — a comment whose first word is "shellcheck" is parsed as a directive,
and the malformed one tripped SC1072/SC1073 at severity error. The gate caught
it on the very commit that introduced it.
2026-09-15 22:59:43 +02:00
joakimp c7d369f28d docs(changelog): measure promote-base-latest's conditional re-tag
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 20s
Closes the open question from v1.9.2's release verification. crane copy DOES
bump Hub's last_updated when it runs (run 669: digests differed, copy ran
17:10:04->17:10:08, base-latest.last_updated moved to 17:10). The
identical-digest path stays unproven by construction -- the job prints
'base-latest already current; nothing to promote.' and never copies -- so a
freshness assertion on the alias is conditional on base-decide's need_build,
not a blanket upgrade. Operational rule lives in the ci-release-watcher skill
(skillset c8034af); the pipeline property is recorded here.
2026-09-14 20:49:23 +02:00
joakimp b8d818ed99 fix(docs): correct the Hub size claims, and gate them so they cannot rot again
Lint / hadolint (push) Successful in 9s
Lint / doc-drift (push) Successful in 13s
Lint / skill-floor (push) Successful in 15s
Lint / actionlint (push) Successful in 18s
DOCKER_HUB.md claimed ~1.1 GB for :latest while Docker Hub served 1.37 GB --
20% wrong, drifting quietly across eight releases, and POSTed to Docker Hub by
update-description every time. It is the first number a stranger reads about
this image.

Root cause is structural, not carelessness: every OTHER claim check-doc-drift.sh
guards is anchored to a build file (a pin, an ARG, a placeholder), so it cannot
rot without someone editing the thing it describes. Nothing in this repo states
the image size, so nothing measured it.

Corrected against measured full_size after v1.9.2 published (amd64/arm64):
  :latest         ~1.1  -> ~1.23 GB   (1.228 / 1.211)
  :latest-studio  ~1.15 -> ~1.25 GB   (1.255 / 1.238)
  :base-latest    ~1.0  -> ~1.17 GB   (1.167 / 1.151)
full_size tracks the FIRST manifest entry (amd64), not the sum across arches --
measured: full_size=1.228, amd64=1.228, arm64=1.211, sum=2.439 -- which is what
the table's per-arch "Size (compressed)" column claims.

New check 8 gates those claims against Hub. Fails on drift; SKIPS LOUDLY via a
new skip() helper (counted, named in the summary) when curl/python3 are absent,
the API is unreachable, or SKIP_SIZE_CHECK=1. A skip is deliberately neither OK
nor a failure: printing an unverified claim as OK is the habit this file exists
to break, and failing on Docker Hub's uptime would hold releases hostage to a
third party. Not wired into hooks/pre-push, so pushes stay offline.

Two bugs caught by writing the expected exit code down before running the check:

  1. The tolerance would have missed its own motivating case. The percentage is
     computed against the MEASURED size, but the 20% first chosen came from the
     claim-relative figure; the real drift was |1.1-1.37|/1.37 = 19.7% and would
     have passed. Now 15%, inside a window with both bounds measured: above the
     largest legitimate skew (11.4%, a claim describing the published release
     while the next tag changes the size) and below the rot it must catch.

  2. A `|| true` on the python invocation made the gate FAIL OPEN -- it printed
     DRIFT and exited 0. Removed; the outer `|| SIZE_RC=$?` satisfies set -e
     without swallowing the code. Verified two-sided: 19% drift exits 1 at the
     default tolerance and 0 at SIZE_TOLERANCE_PCT=25, so the threshold does the
     work rather than the ordering.

NOT covered, and the script says so where a reader will see it: README.md's
~3.2 GB figures are UNCOMPRESSED and the registry exposes compressed sizes only,
so measuring them needs a real pull. A green check 8 says nothing about them.

Verified: check-doc-drift.sh rc=0 on the corrected tree, rc=1 on injected drift,
rc=0 with --warn-only, SKIP+rc=0 on the offline path (proxy to a dead port);
shellcheck clean at default severity (the sed single-quote SC2016 carries a
disable with its reason); lint-shell.sh 16 files clean.
2026-09-14 20:33:40 +02:00
joakimp f5c53b8693 release: date the v1.9.2 section for the tag
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 27s
Publish Docker Image / base-decide (push) Successful in 13s
Publish Docker Image / build-base (push) Successful in 43m59s
Publish Docker Image / smoke (push) Successful in 5m29s
Publish Docker Image / smoke-studio (push) Successful in 5m50s
Publish Docker Image / build-variant (push) Successful in 17m29s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / update-description (push) Successful in 10s
Publish Docker Image / build-variant-studio (push) Successful in 21m50s
CI checks out the tag, so the CHANGELOG heading has to name the version
before the tag exists, not after. v1.9.1's tagged tree carried no
"## Unreleased" either -- same convention, made explicit here.

Contents of v1.9.2 (all measured, nothing inherited):
  - 1baba79  the +131 MB v1.9.1 residual: /root/.npm (110 MB) plus
             @mariozechner/clipboard-* foreign natives at BOTH install
             sites (global + /opt/pi-fork's nested pi-coding-agent copy)
  - 852f900  smoke sentinels that name the residue instead of letting
             the size gate's ~225 MB margin swallow it
  - dab989b  mempalace-toolkit: the owed-withdrawal suite gate
             (node --check never type-strips)
  - 9aaff26  vendored mempalace skill snapshot -> e9e45f7, which is what
             turns the currently-red canary green

Deliberately NOT in this tag: DOCKER_HUB.md's "~1.1 GB" size claim is
low (Hub says v1.9.1 = 1.37 GB, v1.8.14 = 1.23 GB). This release
*changes* the size, so no number is right both before and after it --
writing the predicted ~1.24 GB would be an unmeasured claim in a
user-visible page. Measure post-build, fix on the next tag, and extend
check-doc-drift.sh to gate size claims against Hub's full_size so the
number cannot rot silently again.
2026-09-14 17:55:16 +02:00
joakimp 735565b9be docs: retire the deployed-and-unproven label on isWithdrawn, and name dab989b
Lint / skill-floor (push) Successful in 8s
Lint / hadolint (push) Successful in 17s
Lint / doc-drift (push) Successful in 7s
Lint / actionlint (push) Successful in 20s
Two things, and the second was found by doing the first.

v1.9.1's notes closed with an explicit open item: isWithdrawn had 17 assertions
and 4 mutation kills but had never been exercised on a released image against the
live logstream, which is a different state from untested and a worse one to leave
unlabelled. It has now been exercised, from a container recreated onto v1.9.1, and
the label is retired with the measurement rather than with an assurance.

What the entry records is the instrument as much as the result, because the result
is only as good as the thing that produced it: the shipped extractor was lifted
out of the toolkit's own test script and used to cut the four predicates out of
the BAKED mempalace.ts, and deriveOwed's three queries were replayed with their
exact shipped parameters. No second copy of the logic was written, which is the
same discipline the suite itself is built on. Baseline owed was established by
three independent routes BEFORE any probe was planted, because without a recorded
baseline a withdrawal that appears to work proves nothing. Predictions were
written down before each measurement; the table reports both columns.

The per-predicate attribution is in the entry deliberately. "The count went back
to baseline" is a claim about arithmetic and would have been satisfied by a
pre-filter or by an unrelated reply; isAnswered=false with isWithdrawn=true on a
candidate that was still in the raw set is a claim about mechanism. The negative
controls carry the same weight as the positive one -- a rule that lets a third
party or a prose sentence empty someone's mailbox would have passed a
count-only check.

Also recorded: the rule fires on the real incident it was written for (seq 112),
which is worth more than any synthetic probe, and the one path still unwitnessed
(the extension's own in-process poll, which only fires when the agent is idle) is
named as unwitnessed rather than folded into the pass.

Second: reaching for the suite as corroboration showed it cannot run on this image
at all -- its own precondition gate rejects the source it exists to check, because
`node --check` does not type-strip on any node and v1.9.1 moved to node 24. All 17
assertions had been skipped since the bump. Fixed in toolkit dab989b, named here
per the rule this repo adopted two commits ago after the withdrawal fix shipped
unnamed in v1.9.1. The gate now also distinguishes "cannot run" from "does not
parse", since collapsing those is how the defect pointed readers at the wrong
file.

The rebuild cost is stated and, having been measured, is smaller than the earlier
draft of this entry claimed: 9aaff26 already moved base_tag by refreshing the
vendored skill snapshot under rootfs/, so the base rebuild was pending before this
fix existed and the toolkit pickup rides along with it.
2026-09-14 16:29:52 +02:00
joakimp 9aaff26e3a docs: the withdrawal fix shipped in v1.9.1 unnamed — say so, and re-pin the canary
Lint / hadolint (push) Successful in 11s
Lint / doc-drift (push) Successful in 12s
Lint / skill-floor (push) Successful in 12s
Lint / actionlint (push) Successful in 16s
v1.9.1 bakes mempalace-toolkit e68ee20, which contains e2b060a. Requester-side
ask withdrawal has therefore been LIVE on every v1.9.1 device since 2026-09-10,
while the v1.9.0 section of CHANGELOG.md still read "not yet pinned ... this
image still pins e45f6b4", and the mempalace skill still told every agent at
session start that a withdrawal is impossible.

Measured, two independent routes, expectation recorded before looking:

  - published image label, :v1.9.1-studio and :latest-studio (same digest):
      se.jordbo.pi-devbox.mempalace-toolkit-ref = e68ee2071ca3ad39...
  - ancestry: e2b060a is an ancestor of e68ee20
  - baked mempalace.ts sha 7c16fe14 != v1.8.14's dfca71e9 (different bytes)
  - grep -c isWithdrawn on this container's baked copy = 0 (v1.8.14, so
    tor-ms22 cannot exercise the behaviour it is documenting)

The first label read came back EMPTY, and that empty was a claim about the
request rather than the image: Hub redirects blob fetches to a CDN and curl
without -L returns 0 bytes at exit 0. Recorded in the notes, because a registry
audit reporting "no labels" is missing -L until proven otherwise.

Changes:

  - CHANGELOG Unreleased: the floating-ref mechanism, which is the reusable
    part. ARG MEMPALACE_TOOLKIT_REF=main + CI resolving it to a SHA at build
    time means a release absorbs whatever toolkit main holds, and "what
    behaviour did this image gain" is a question nobody is forced to answer.
    The fleet rule (name the pickup before tagging) was honoured for the
    feed-tick commits v1.9.1 names, and missed for one in the same range.
  - CHANGELOG v1.9.0: the stale caveat is ANNOTATED, not rewritten. The
    sentence is the evidence for how the drift happened; deleting it would
    destroy the only trace. Same discipline as fleet-ops d42412c.
  - CHANGELOG Unreleased: corrected the size entry's "needs no base rebuild"
    claim. True of that entry alone, false of the release now that the
    vendored skill snapshot -- a base_tag input -- moves with it (~67 min).
  - rootfs mempalace skill snapshot 44472 -> 46045 B, refreshed via
    scripts/vendor-mempalace-skill.sh so the bytes and SKILLSET_SNAPSHOT_REF
    (4d7c0ea -> e9e45f7) move together; --check confirms exact match.
  - scripts/smoke-test.sh canary RE-PINNED. The retired pair was still green
    against the new snapshot, i.e. blind to this refresh for the same reason
    the pre-v1.8.13 pair was blind to that one. The replacement is stronger
    than any predecessor here because BOTH witnesses come from the same
    upstream commit: e9e45f7 added "Withdrawing an ask you sent" and deleted
    "nothing anyone can do about it from the other end", the sentence the new
    bullet contradicts. Directions measured against both files, then the
    canary body EXECUTED against each: new -> rc=0 "ok", old -> rc=1 empty.
    A canary whose negative witness was removed by the commit it pins fails
    loudly on stale bytes instead of merely failing to notice them.

Gates: check-doc-drift OK, check-skill-floor OK, vendor --check OK,
hooks/pre-push OK (16 shell files clean at severity error).

Not fixable here: isWithdrawn is deployed and UNPROVEN. 17 assertions and 4
mutation kills, never once exercised on a released image against the live
logstream. Ask routed to a v1.9.1 device.
2026-09-14 15:07:26 +02:00
Joakim Persson 852f900b53 test(smoke): make CI prove the natives still work — the runbook check didn't
Lint / doc-drift (push) Successful in 6s
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 9s
Lint / actionlint (push) Successful in 18m11s
The post-boot check v1.9.1 left for the next machine was

  node -e 'require("esbuild").transformSync("const x:number=1",{loader:"ts"})'

with "if this fails, the prune removed something needed on arm64 -> revert to
v1.8.14 and fix" attached. Run from /workspace on the freshly recreated v1.9.1
image it fails with MODULE_NOT_FOUND — on a perfectly healthy image. require()
resolves by walking up from the CURRENT DIRECTORY, esbuild is nested inside the
two pi trees, and a global install is not on node's require path (NODE_PATH is
unset). The version that "verified the prune" last night only passed because the
shell happened to sit inside the tree. So the check was CWD-dependent, its red
meant nothing, and the remedy it prescribed was reverting a good release.

Path-qualified, both sites compile TS today. Rather than leave a fragile command
in a note for a human to paste, CI now owns it:

  - esbuild must transformSync at EVERY install site found in the image
  - @mariozechner/clipboard must load with its native binding attached at every
    site — for the clipboard prune that IS the proof, since napi-rs resolves the
    platform package at require() time

Sites are discovered with find, so the studio variant's third site is covered
without naming it.

Verified as a four-way matrix, not just a green run: both assertions pass
against the real image FROM /workspace (the cwd that broke the old command), and
both fail (exit 1, "The package \"@esbuild/linux-arm64\" could not be found, and
is needed by esbuild") against copies of the same packages with the host
platform binary removed. An assertion never observed failing is not evidence.

Also documents both failure shapes in AGENTS.md next to the size gate: this one,
and the mode-700 /root one where `test ! -d` passes on a permission error.
2026-09-11 11:06:18 +02:00
Joakim Persson 1baba79c96 fix(size): the v1.9.1 residual was npm's own cache, not the platform binaries
Lint / doc-drift (push) Successful in 6s
Lint / hadolint (push) Successful in 10s
Lint / skill-floor (push) Successful in 9s
Lint / actionlint (push) Successful in 4m26s
v1.9.1's @esbuild prune fixed the 431 MB size-gate failure but still shipped
+131 MB compressed over v1.8.14, nearly all of it in the pi/extensions install
layer (87 -> 206 MB). That leftover was filed as an open item with an explicit
hypothesis — the @mariozechner/clipboard-* family, same npm 11 behaviour, a
different package — and an explicit warning that the hypothesis was not a
measured cause. Measured now, after recreating onto v1.9.1, and the hypothesis
accounted for one sixth of it:

  +110 MB  /root/.npm/_cacache  (35.2 -> 145.3 MB)   the build's npm cache
  + 21 MB  clipboard foreign platform packages, both install sites
  = 131 MB  i.e. the whole delta, no unexplained remainder

Method, since there is no docker CLI inside the container: pulled both variant
layer blobs straight from the registry with a token + manifest + blob fetch and
listed the tarballs (29 789 vs 29 889 entries, 270.8 vs 401.9 MB uncompressed),
then aggregated per package. The file COUNT barely moved, which is what said
"few large files", not "npm installed more packages".

  - purge_build_caches: npm cache clean --force + rm -rf /root/.npm, in the SAME
    layer as the installs, in both the main RUN and the studio RUN. npm 11
    caches every platform tarball it downloads, including the ones the prune
    then deletes, so the cache grew faster than the tree. Nothing at runtime
    reads it: build is root, container is developer with its own cache in $HOME.
  - prune_foreign_esbuild -> prune_foreign_natives: now covers both MEASURED
    families. Clipboard keeps linux-$arch-gnu AND -musl because its napi-rs
    loader picks between them at runtime via its own isMusl() probe; the musl
    package is a 420-byte stub. The bare @mariozechner/clipboard wrapper has no
    hyphen suffix and cannot match the pattern.

Verified on arm64 before writing the glob — a widened rm -rf against a tree you
cannot inspect is the one change shape not to write blind, which is why the
order was update-then-patch. Exercised against a copy of the real trees with
foreign dirs fabricated back in (aix-ppc64, android-arm64, darwin-arm64,
win32-x64, linux-x64): all removed, host linux-arm64 kept at both sites, 21 MB
freed, require('@mariozechner/clipboard') still loads and exports all 18
functions, esbuild.transformSync still compiles TS at both sites. v1.9.1's own
arm64 validation also passed here; CI could only smoke amd64.

Two sentinel assertions, because the size gate did not catch this: it has
~225 MB of margin, so 131 MB of residue stayed green. Both were verified RED
against the running v1.9.1 image and GREEN against a pruned tree. The cache one
refuses to run as non-root: test ! -d /root/.npm on mode-700 /root would
otherwise pass for the wrong reason. Size-failure diagnostics now list cache
paths too — they previously enumerated only node_modules and /opt, where these
bytes were not.

Also fixed: -printf '%f\\n' reaches the shell with both backslashes (confirmed
from the published image's recorded created_by), so v1.9.1's progress line
printed a mangled "li ux-arm64" — find emitted a literal backslash and tr ate
the n out of the name. Single backslash now.

Deliberately not purged: /tmp/node-compile-cache (1.3 MB). The manifest RUN
calls pi --version again, so deleting it earlier only relocates those bytes into
that layer — today's manifest layer is 128 kB precisely because it finds the
cache warm.

Dockerfile.base untouched, so no base rebuild: this rides the next release.
2026-09-11 09:39:54 +02:00
Joakim Persson 42bd29d654 fix(ci): unblock the release — npm 11 esbuild bloat + two self-inflicted assertions
Lint / hadolint (push) Successful in 8s
Lint / skill-floor (push) Successful in 9s
Lint / doc-drift (push) Successful in 10s
Lint / actionlint (push) Successful in 19s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 14s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Successful in 5m4s
Publish Docker Image / smoke-studio (push) Successful in 13m17s
Publish Docker Image / build-variant (push) Successful in 33m49s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 14s
Publish Docker Image / build-variant-studio (push) Successful in 30m58s
v1.9.0 was tagged but never published: smoke failed 90-passed/3-failed and
build-variant needs smoke, so nothing reached the registry. All three are fixed.

1. Size, 431 MB over threshold. Node 24 brings npm 11, which installs EVERY
   @esbuild/<platform> optional binary instead of the matching one: 26 dirs,
   284 MB per pi-coding-agent copy. Measured on pi-fork: npm 10.9.8 -> 165 MB
   (exactly what v1.8.14 shipped), npm 11.19.0 -> 449 MB. npm 11 ignores the
   os/cpu constraints AND --os/--cpu AND an npmrc carrying them, all measured,
   so Dockerfile.variant prunes explicitly, keeping linux-$(node -p
   process.arch) so one line is right on both arches. Pruned in the SAME layer
   as each install, or the bytes survive in the earlier layer. Three sites:
   global pi, pi-fork, pi-studio. Verified esbuild still transforms TS after.

2. om's node_modules assertion tested an npm artefact. om has zero runtime deps;
   its node_modules held ONE file (.package-lock.json) and 20 empty scope dirs.
   npm 11 stopped creating it. Now asserts the entry point pi actually loads,
   read from package.json -> pi.extensions.

3. The skill-source annotation added in v1.9.0 ("baked (package copy)") broke
   the assertion matching "baked$". Pattern now allows an optional suffix; which
   copy shipped stays authoritatively asserted against the manifest + tree hash.

Also: a failed size check now prints the largest layers, largest directories and
an @esbuild sentinel, so this class attributes itself next time instead of
costing a CI dig plus a local npm bisect.

Threshold stays 3800 MB: it caught a real regression and raising it would have
thrown the signal away. No Dockerfile.base/rootfs change, so base-0fb1256c7f99
is reused and build-base is skipped.
2026-09-10 23:39:40 +02:00
Joakim Persson 3a44e81cad feat(ci): gate documentation drift, and make docs a pre-tag release step
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Five doc claims had rotted by v1.9.0, all the same shape: a value written once
by hand, in a file nothing verifies, about a number that lives elsewhere and
moved. README pin table wrong on all three rows; a "Planned" section describing
something already shipped; DOCKER_HUB.md claiming Node v22 against Node 24.

DOCKER_HUB.md is why this is a gate and not a resolution to be careful: it is
PUBLISHED (update-description POSTs it as Docker Hub full_description on every
tag), it had gone eight releases untouched, nothing generates it, and it is read
from the TAG -- so the stale page shipped with v1.9.0 regardless.

scripts/check-doc-drift.sh: seven checks, all repo-local (no network, token,
image, or sibling clone). Exit 0/1/2 matching lint-shell.sh; a renamed ARG is a
red 2, not a green tick. Wired as a fourth lint.yml job so "the docs lie" is its
own red name.

Verified with 15 controls, including two false-positive controls: the first
placeholder check flagged README.md:900, a Go template in a legitimate
`docker inspect --format` example. The gate was wrong, not the doc, so the
pattern is now anchored to the UPPER_SNAKE convention CI substitutes.

Not gated, deliberately: counts/sizes needing a running image (they belong in
smoke-test.sh -- a guessing gate is worse than none), and Dockerfile.base
BASE_REBUILD_DATE, because base_tag hashes that file content-wise and demanding
it be current would force a ~60 min rebuild on releases that touch no base
files. Free during a rebuild, expensive otherwise.

AGENTS.md step 3 rewritten around the mechanism: checkout@v4 with no ref: means
every job reads github.ref, the tag. Docs must be right BEFORE tagging.
2026-09-10 22:01:49 +02:00
Joakim Persson 35964abd01 docs(hub): the Docker Hub page claimed Node v22; v1.9.0 ships Node 24
DOCKER_HUB.md is hand-maintained (no generator: CI only substitutes
{{PI_VERSION}} and POSTs the file as Docker Hub full_description), and it was
last touched at v1.8.6 -- eight releases ago. The Node claim is now false as of
this release, and unlike README.md this file IS published, so a stale claim here
is user-visible rather than internal.

Verified countable claims rather than assuming: "7 user-facing extensions" is
correct (7 files in pi-extensions/extensions/). The "29 mempalace_* tools" claim
is suspect -- this session sees 45 -- but left alone because I cannot attribute
that count to the baked 3.9.0 server without measuring it.
2026-09-10 21:43:17 +02:00
Joakim Persson 6353d59e63 docs(readme): correct the three stale version pins and the shipped-as-planned claim
The version-pin table existed precisely to be the reviewable record of what is
deliberately frozen, and it was wrong on every row: pi 0.84.4 -> 0.85.1,
pi-atelier v0.10.0 -> v0.10.1, mempalace 3.8.0 -> 3.9.0. Verified by parsing the
table and comparing against the ARGs it names rather than by eye.

Also: the "Planned for an upcoming minor release" section listed typst PDF
export, which shipped long ago and even carried a self-contradicting
"(shipped in Unreleased/base)" marker -- the fourth instance of the stale
in-repo Unreleased-pointer class this CHANGELOG already documents. typst 0.15.1
confirmed live in the running image, so the item is now stated as current fact.

The pi-devbox-version sample was v1.5.0-era and structurally outdated: it
predates the palace line the surrounding prose advertises, the pi-atelier
component, and the whole skills: block. Replaced with real observed output
rather than hand-written text.

README has no CI coupling (no workflow or gate reads it; the Docker Hub page
comes from DOCKER_HUB.md), so this cannot affect the in-flight v1.9.0 build.
2026-09-10 21:33:07 +02:00
Joakim Persson 8f0960e134 release: v1.9.0
Lint / skill-floor (push) Successful in 8s
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / lint-gate (push) Successful in 22s
Publish Docker Image / resolve-versions (push) Successful in 15s
Publish Docker Image / base-decide (push) Successful in 8s
Publish Docker Image / build-base (push) Successful in 42m19s
Publish Docker Image / smoke-studio (push) Failing after 6m17s
Publish Docker Image / build-variant-studio (push) Has been skipped
Publish Docker Image / smoke (push) Failing after 9m7s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Rename the Unreleased changelog section to its release heading, matching the
established `## vX.Y.Z — YYYY-MM-DD` format (em dash), per AGENTS.md release
step 3.

Minor rather than patch: CHANGELOG.md scopes minor to "significant base
additions", and this release bumps the Node runtime under every baked JS tool
(pi, agent-browser, playwright, mempalace) from 22 to 24 — the first
NODE_VERSION change since it was introduced at v1.0.0 — alongside five new
base packages (shellcheck, bind9-dnsutils, ldap-utils, xxd, python3-yaml).
Package additions alone have been patch here (v1.8.12, v1.8.13); the runtime
major is what lifts this one.
2026-09-10 21:21:09 +02:00
Joakim Persson 1c905480e3 docs(changelog): note the mempalace feed-tick fix this image carries
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 10s
Lint / actionlint (push) Successful in 19s
The recurring "[mempalace ext] feed (tick) failed: mine timed out after 30000ms"
message is fixed in mempalace-toolkit (309980b + e68ee20) and this image is what
delivers it, since CI resolves MEMPALACE_TOOLKIT_REF to a commit SHA at build
time. Worth a changelog entry rather than leaving it implicit in a ref bump: it
is the most visible symptom operators on this fleet have been living with, and
the entry records that it was a genuine defect (overlapping mines on a
single-writer palace) rather than the cosmetic annoyance it was parked as.
2026-09-10 20:58:42 +02:00
Joakim Persson ff6fd1492a feat(manifest): record WHICH pi-extensions skill copy shipped
Closes the half deliberately left open by cac5e00's skill-floor gate, and the
more important half: "the floor is currently fresh" is a fact with a shelf
life, whereas "the image says which copy it got" keeps working.

The refresh in Dockerfile.variant is guarded by
`[ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates the
co-located skill keeps the vendored floor and still succeeds GREEN, with nothing
in the manifest, labels or logs separating that from a normal build. Afterwards
the two are indistinguishable by inspection -- same path, same filenames, same
permissions -- which is exactly how the floor went unnoticed from 2026-07-30 to
2026-09-10.

build-manifest.json gains pi_extensions_skill_source and
pi_extensions_skill_tree_sha256, MEASURED rather than passed as build-args, per
the ground-truth rule the surrounding block already follows -- and necessarily
so, since the outcome depends on the clone's contents and no ARG could express
it. Three values, because two would force a lie: package (served bytes equal
the clone's skill/), vendored-floor (clone had no skill/ at this ref), and
divergent (both exist but differ -- e.g. the clone ships SKILL.md but not
evaluate-extension-usage.py, so the served directory is a genuine MIX). No OCI
label mirrors these deliberately: LABEL cannot take a RUN-computed value, and a
label fed from an ARG would be the claim-not-measurement being removed here.

Two smoke assertions make the record a gate: the source must be named and be
`package` -- vendored-floor FAILS rather than warns, since these images track
main where the package has shipped skill/ since fa04d20, so a fallback means
the clone did not resolve as intended -- and the tree hash is recomputed over
the served directory, because a recorded hash never recompared is a claim.

pi-devbox-version annotates the line too: "baked (package copy)" normally, or a
yellow "(FALLBACK: vendored floor)". Its existing section reports which copy is
READ at runtime; this is the one fact decided at BUILD time and unrecoverable
later. Old images degrade cleanly -- field absent, jq // empty yields nothing,
line prints plain "baked" as before (verified against this v1.8.14 manifest).

Tested by running the exact logic against this container's real layout, with
the expected value written down before each: package (served == clone),
vendored-floor (clone path absent), divergent (clone lacking the .py while the
served dir has it), and null (empty served dir) -- all four as predicted. The
five pi-devbox-version render branches likewise, including the absent-field
case. Emitted JSON validated with jq for both the populated and null forms.

Gates green: lint-shell.sh (15 files), hadolint 2.15.1, actionlint 1.7.12,
check-base-hash.sh, check-skill-floor.sh, vendor-mempalace-skill.sh --check.
2026-09-10 20:30:16 +02:00
Joakim Persson edc7659add chore(deps): node 22->24, actionlint 1.7.12, hadolint 2.15.1, skillset ref
Audited every component the image obtains OUTSIDE debian/apt. Of ~23, the 19
that resolve `latest` at build time were already current or refresh themselves
on the next rebuild, and the hard pins for pi (0.85.1), mempalace (3.9.0) and
pi-atelier (v0.10.1) were already newest. Four needed a human.

NODE_VERSION 22 -> 24 (LTS "Krypton"). This was a latent defect rather than
housekeeping: agent-browser publishes engines.node ">=24.0.0", so the image sat
BELOW a declared requirement -- v1.8.14 shipped node 22.23.2 with agent-browser
0.37.1, so every build installed it with an npm EBADENGINE warning and ran the
baked browser automation outside its supported range. pi (">=22.19.0") and
playwright (">=20") are satisfied either way. Verified before bumping, since a
missing NodeSource suite breaks every arch at once: setup_24.x returns HTTP 200
and node_24.x advertises `Architectures: amd64 arm64 armhf x86_64`, covering
the arm64 fleet and the amd64 CI runners. Nothing else pinned the node major.

actionlint 1.7.7 -> 1.7.12 and hadolint 2.14.0 -> 2.15.1, each RUN AGAINST THIS
TREE at the new version before being pinned -- both clean, no new findings. A
linter bump is the one dependency update that can turn CI red on unchanged
code, so it is verified locally rather than discovered on a round trip.

SKILLSET_SNAPSHOT_REF e9e09d9 -> 4d7c0ea via scripts/vendor-mempalace-skill.sh,
never by hand: that script is the only thing permitted to write the ARG,
because a cp without a matching bump yields a manifest that confidently lies.
This proved PROVENANCE-ONLY -- the ref was 6 commits behind, but
skills/mempalace/SKILL.md is byte-identical at both (3675bfab), so the snapshot
was already correct and only its recorded origin was stale. No rootfs/ bytes
changed, the smoke-test phrase canary stays valid, and this ARG alone would not
force a base rebuild (the node bump does).

Two measurement traps worth recording, since both would have produced a wrong
answer: GitHub's releases/latest reports pi-atelier v0.10.0 as newest because
v0.10.1 is a TAG WITH NO RELEASE OBJECT -- the pin was already current, and
`git ls-remote --tags` is the instrument that shows it. And gitea-mcp is hosted
on gitea.com, not GitHub, so querying api.github.com returned nothing at all
rather than an error.

Verified with every gate this repo owns, all green, using the NEW linter pins:
lint-shell.sh (15 files), check-workflow-shell.sh, check-base-hash.sh,
actionlint 1.7.12, hadolint 2.15.1, check-skill-floor.sh, and
vendor-mempalace-skill.sh --check.
2026-09-10 20:17:32 +02:00
Joakim Persson cac5e00a31 feat(ci): gate the vendored pi-extensions skill floor, and bake python3-yaml
Follows ecfd2fc, which refreshed the stale floor by hand. A one-off refresh
fixes the symptom; this makes the drift impossible to reintroduce silently.

scripts/check-skill-floor.sh compares the repo floor
(rootfs/usr/local/share/pi-devbox/skills/pi-extensions/) against the package
repo it is a snapshot of, wired in as a new `skill-floor` job in lint.yml.

DIRECTORY hash, not `sha256sum SKILL.md`, using the same tree_sha256 pipeline
Dockerfile.variant uses for skillset_snapshot_tree_sha256 and for the reason
already documented there: a file-only compare answers "did this one file
change", not "is this the same skill". Verified by NEGATIVE CONTROL rather
than asserted -- with SKILL.md left byte-identical and only
evaluate-extension-usage.py edited, the directory check fails (rc=1) where a
file-only compare would have passed. Seven behaviour tests, each with its
expected rc written down before running: in-sync via local dir (0), in-sync
via anonymous remote clone (0), missing --package-dir (2), bad argument (2),
content drift (1), the sibling-file case (1), and --warn-only over drift (0).

Exit codes 0 in sync / 1 drift / 2 cannot-run, matching scripts/lint-shell.sh:
a gate that cannot run must not pass, so an unreachable package repo is a red
2 and never a green tick. A ref with no skill/ is NOT drift -- that is the
documented fallback -- but it emits ::warning:: because it is precisely the
condition under which the floor ships.

Gating on another repo is normally a smell. It is proportionate here because
the check can only fire when skill/ itself changed, which is exactly when the
floor has gone stale; pi-extensions commits that leave skill/ alone cannot
turn this red. It also needs no secret: pi-extensions is anonymously clonable
(verified with `git ls-remote` and no credentials), so it cannot start failing
when a token expires.

Also bakes python3-yaml (552 KB, zero extra deps) into Dockerfile.base. This
is the shellcheck story repeating exactly: scripts/check-workflow-shell.sh --
the guard against the Gitea sh/dash footgun that broke resolve-versions
(ed49b8d) and promote-base-latest (b7197e8) -- hard-exits with "python3 yaml
module missing", so a gate this repo already owns could not be run locally by
anyone. lint.yml installing it explicitly in CI was the evidence. Found while
wiring the job above: the guard could not be run before pushing.

CHANGELOG Unreleased updated for both this and ecfd2fc, including an explicit
note on what is NOT fixed -- the silent-fallback half still has no manifest
flag recording which copy was served.

Verified locally with every gate this repo owns, all green: lint-shell.sh (15
files clean), check-workflow-shell.sh, check-base-hash.sh, actionlint 1.7.7
(pinned, same version as CI), hadolint 2.14.0, and the new check itself.
2026-09-10 19:14:01 +02:00
Joakim Persson ecfd2fc2e5 feat: bake dig/ldapsearch/xxd and refresh the stale pi-extensions rootfs floor
Two changes that share one forced base rebuild, hence one commit.

1. THREE PACKAGES, each closing a capability gap measured during the
   gitea.egl.lan/FreeIPA work on 2026-09-09..10 rather than a preference:

   bind9-dnsutils (~6.1 MB measured) -- dig/host/nslookup were ALL absent,
   so the container could resolve names but had no way to interrogate a
   SPECIFIC nameserver. `getent hosts` only follows the resolver's default
   path, so diagnosing "gateway 172.16.88.1 NXDOMAINs the egl.lan zone
   while 10.20.253.1 is authoritative for it" had to be hand-rolled in
   python3. Split-horizon DNS is a recurring class of bug on this fleet.
   Note the package name: plain `dnsutils` is transitional in trixie.

   ldap-utils (1244 KB, pulls nothing extra) -- the fleet authenticates
   against FreeIPA, yet every LDAP probe had to be run by SSHing to an
   already-enrolled host. Simple binds only; GSSAPI would additionally
   need krb5-user + libsasl2-modules-gssapi-mit, deliberately not added
   as that is a Kerberos-client decision, not a tool.

   xxd (198 KB) -- convenience for verifying git-crypt blob magic in
   myconfigs; `od -c` from coreutils already does the same job.

   netcat-openbsd was in the original proposal and is deliberately NOT
   here: measured redundant, because socat is already baked and bash's
   /dev/tcp does reachability checks with zero packages (verified against
   gitea.egl.lan:3000). Recorded in the Dockerfile so the omission reads
   as a decision rather than an oversight.

2. ROOTFS FLOOR REFRESH: rootfs/.../pi-extensions/SKILL.md was 34284 B,
   unchanged since fa04d20 (2026-07-30), while the canonical package copy
   is 38973 B. Dockerfile.variant copies the fresh package copy over the
   SERVED path at build time but never writes back to this floor, so the
   floor is a silent fallback: if that build-time copy is ever absent it
   ships the July skill with no log line or manifest flag to say which
   version deployed. Refreshed from pi-extensions@c64c122, verified
   byte-identical to both the canonical and the runtime-served copies.

Why one commit: the base_tag hash folds in `cat Dockerfile.base` AND
`find rootfs -type f | xargs cat` (.gitea/workflows/docker-publish.yml),
so either change alone forces the same full base rebuild -- and that
rebuild is precisely what re-bakes rootfs/ as it then stands. Emulating
the workflow hash with a fixed toolkit ref: f3d6462c7416 -> fc4edda03c54.

Verified: scripts/check-base-hash.sh passes (no new ARG *_REF added), and
no shell scripts are touched so the lint-shell gate is unaffected. Sizes
and dependency fan-out measured via apt-get --no-install-recommends
--dry-run on Debian 13 trixie.
2026-09-10 18:56:11 +02:00
joakimp 15a3728ae9 feat: bake shellcheck and add a client-side pre-push lint gate
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 19s
v1.8.14 made shell lint a RELEASE gate (scripts/lint-shell.sh, shared by lint.yml
and the new lint-gate job that resolve-versions depends on), and that script
correctly exits 2 when shellcheck is absent -- "a gate that cannot run must not
pass". Measured on v1.8.14 on 2026-09-09 by three routes (command -v, dpkg -l, a
filesystem search): shellcheck was NOT IN THE IMAGE AT ALL. So the gate could not
be run by a developer in any container, only in CI, and the loop stayed
write-shell -> push -> wait for CI -> discover. That is the loop the gate was
added to shorten, after v1.8.14's first attempt burned ~46 min on a tree whose
lint had already been red for 24 hours.

shellcheck 0.10.0-1 added to the Dockerfile.base apt block: ~39 MB installed
(Installed-Size 40112 KB), measured to pull ZERO additional packages under
--no-install-recommends because libc6/libffi8/libgmp10 are already present.
NOTE this forces one full base rebuild -- base-decide hashes Dockerfile.base +
rootfs/, so unlike a scripts/ change it cannot reuse the existing base- layer.

hooks/pre-push is opt-in per clone (git config core.hooksPath hooks), bypassable
with --no-verify, and execs scripts/lint-shell.sh rather than reimplementing it
-- one copy, because a duplicated check that drifts is the failure this repo
keeps paying for. Matches the idiom skillset/ and myconfigs/ already use.

WHY THIS REPO HAD NO HOOKS, since it was reported as drift and is not: a peer
asked tor-ms22 for core.hooksPath per clone on the premise that unset meant the
gates were unverified there. Measured: pi-devbox unset, skillset hooks, myconfigs
common/hooks, pi-toolkit unset -- but `git ls-files | grep -i hook` is EMPTY in
both pi-devbox and pi-toolkit, so there was nothing to point at on any machine
and unset was the only correct value. This closes the real half for pi-devbox;
pi-toolkit still ships none.

Verified, expected result written down before each check:
  * refusal paths -- shellcheck absent => rc 2 with the remedy named; linter
    missing => rc 2. Never waved through on the assumption CI will catch it.
  * the hook is IN the scan set -- "Checking 14 shell file(s)" with it present,
    13 with it moved aside, so the extensionless file is found by the shebang
    half of the linter's two-signal union. This check exists because the first
    attempt was ambiguous: a planted `[ $UNSET_VAR = "x" ]` was not reported,
    which could equally have meant "not scanned" or "below -S error". It was the
    latter. A count that moves is unambiguous; a clean run is not.
  * it catches the REAL v1.8.14 defect -- planting `echo 'the fleet\'s thing'`
    in hooks/pre-push yields SC1073/SC1072 at severity error, rc=1.

And the gate earned its keep inside this commit: the first version of the
smoke-test assertion carried a comment beginning "# shellcheck is a GATE
DEPENDENCY", and a comment whose first word is the tool's name is parsed as a
DIRECTIVE, not a comment. The new gate failed it with SC1073/SC1072 before the
push -- same family as the v1.8.14 apostrophe, a line that reads as prose to a
human and as syntax to the parser.
2026-09-09 08:57:34 +02:00
joakimp 361babd4fd ci: gate the release on shell lint, from one shared script
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Publish Docker Image / smoke (push) Successful in 7m39s
Publish Docker Image / build-variant-studio (push) Successful in 17m15s
Publish Docker Image / build-variant (push) Successful in 17m58s
Publish Docker Image / update-description (push) Successful in 8s
Publish Docker Image / promote-base-latest (push) Successful in 12s
v1.8.14's first attempt spent ~46 minutes building a base image for a tree whose
own lint had been failing for 24 hours. shellcheck had already flagged the
defect (SC2289, severity error) on the push that introduced it; the lint
workflow went red at run 186 and nobody read it.

lint.yml deliberately skips tag pushes and its reasoning is sound -- the tagged
tree was already linted on main, and a tag-ref lint run sorts above the publish
run, making a release look finished before anything ships. The missing invariant
was never "lint the tag". It was "do not RELEASE a tree whose lint failed", and
only a job inside the publish workflow can enforce that.

So: extract the shell-lint logic from lint.yml into scripts/lint-shell.sh and
call it from both places, then add a lint-gate job that resolve-versions depends
on. resolve-versions is the graph root, so gating it gates everything. Cost is
~40 s at the front of a release; the alternative already cost fifty minutes.

Extracted rather than copied on purpose. A second copy of a check is the drift
this repo keeps paying for -- the same evening produced a skillset mirror that
had sat 9579 B behind its upstream through two consecutive edits.

The script adds one behaviour the inline version lacked: if shellcheck is not
installed it exits 2 rather than silently finding nothing, inheriting the
existing "a gate that cannot run must not pass" rule from hooks/pre-commit in
the skillset repo. Without that, reordering the install step away would turn the
gate into a green tick over zero checks.

Verified locally with a stubbed shellcheck (the real binary is not in the
devbox), five cases, each with its expectation stated first: absent shellcheck
-> rc=2; stub pass -> rc=0 and a non-zero file count; stub fail -> rc=1; a
deliberately unterminated `if` planted in scripts/ -> rc=1 via the bash -n half,
naming the file; removal -> rc=0 again. Discovery cross-checks against CI's own
number: the inline version reported 12 files, the extracted one reports 13, the
difference being lint-shell.sh itself. YAML re-parsed (10 jobs, was 9) with an
assertion that resolve-versions needs lint-gate, and the repo's
check-workflow-shell.sh guard still passes.
2026-09-08 23:41:44 +02:00
joakimp 70e675afee fix(smoke): keep prose out of the single-quoted exec_test body
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 22s
The agent-browser execution guard added on 2026-09-07 carried its explanation
INSIDE the single-quoted script body, and the explanation contained an
apostrophe ("the fleet\'s only recurring amd64 runtime proof"). Inside '...'
bash treats a backslash as literal, so \' does not escape the quote -- it CLOSES
the string. The body truncated at that point and the remaining lines were parsed
by the calling shell.

Consequences, both measured rather than inferred:
  - exec_test received 12 arguments instead of 2 (verified two-sided: the fixed
    tree yields argc=2, HEAD yields argc=12).
  - the leaked `v=$(agent-browser --version)` ran on the CI RUNNER instead of
    inside the image. The runner has no agent-browser, so smoke and
    smoke-studio both failed with "line 770: command not found" after
    build-base had already spent ~46 minutes. Every downstream job was skipped.
  - the truncated body still passed inside the container and printed its green
    tick first, so the log shows a PASS immediately followed by the failure --
    the tick was real, it just no longer covered the assertion.

The prose now sits above the exec_test call, where an apostrophe cannot
terminate anything, and a comment at that spot records why it must stay there.

Not a new failure class: shellcheck flagged it as SC2289 at severity error the
same day, so the lint job has been red since run 186 (2026-09-07 21:21) and was
not read. The gate did its job; nobody looked.
2026-09-08 23:31:48 +02:00
joakimp 601fc98a49 docs(changelog): release v1.8.14
Lint / hadolint (push) Successful in 9s
Publish Docker Image / resolve-versions (push) Successful in 14s
Lint / actionlint (push) Failing after 24s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 50m32s
Publish Docker Image / smoke-studio (push) Failing after 5m13s
Publish Docker Image / build-variant-studio (push) Has been skipped
Publish Docker Image / smoke (push) Failing after 7m40s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Converts the Unreleased section and records what this build carries beyond it:
the mempalace-toolkit bump that makes closing replies reach the mailbox
(deriveClosed, 21023e7 -> e45f6b4), and the L0-L4 subtask documentation landing
via pi-toolkit adfb553 + pi-extensions c64c122.

Notes the mechanism that makes the toolkit fix land at all — the resolved
toolkit SHA is folded into the content-addressed base tag, so the toolkit
moving forces a base rebuild rather than waiting for one — and the consequence
for the amd64 item already in this section: v1.8.13's base was cached, so
Dockerfile.base:607's agent-browser assertion never ran. This base is not
cached, so the native-amd64 proof is finally collected instead of discarded.
2026-09-08 22:09:31 +02:00
joakimp 7e0e66997d docs(env): name MEMPALACE_MAILBOX_NOTIFY — auto-detect cannot work in a container
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Failing after 16s
Unset means the mailbox is silent outside the pi TUI, and the reason is
structural: docker exec does not forward KITTY_WINDOW_ID/TERM_PROGRAM, so
'desktop' detection always falls through to OSC 777, which Kitty does not
implement — the notification then silently does nothing, the worst failure for a
feature whose only job is to break a silence. Documents the four modes, and that
MEMPALACE_MAILBOX_POLL_MS is a FLOOR BETWEEN activity-coupled polls rather than a
wall-clock interval (an idle session polls zero times) — the exact expectation
mismatch reported today.
2026-09-07 21:43:01 +02:00
joakimp 6bd8b79d3a test(smoke): assert agent-browser EXECUTES — it was the discarded amd64 proof
Lint / hadolint (push) Successful in 11s
Lint / actionlint (push) Failing after 23s
Second instance of the same bug class as the node line, in the same file, found
the same way. The agent-browser guard captured the version inside an echo with
2>/dev/null:

  echo "resolved=[$r] version=[$(agent-browser --version 2>/dev/null|head -n1)]" >&2

so the exit code was discarded and a binary that could not execute at all still
PASSED, printing version=[]. Verified two-sided: a stub exiting 127 passes the old
form and is caught by the new one.

Why this exit code matters more than most: smoke runs platforms: linux/amd64 on an
x86 runner, i.e. NATIVE amd64, so this line is the fleet's only recurring amd64
runtime proof for agent-browser's linux-x64 ELF.

NO DEVBOX CAN EVER SUPPLY THAT PROOF. Every machine in the pi fleet is an Apple
Silicon Mac: mbp-m1-2020; tor-ms22 = Mac Studio Mac13,1 M1 Max (fleet-ops
hosts/tor-ms22.md, verified 2026-08-17 with system_profiler); emb-7kj4vr4g =
Apple Silicon, verified 4 routes 2026-09-07. The open "amd64 runtime proof still
needed" ask sent to two devices was asking for the impossible, and emb's reply
naming tor-ms22 as "the only remaining candidate" is wrong for the same reason.
CI had the answer all along and was throwing it away.

Dockerfile.base:607 DOES assert it (`agent-browser --version && \`), but only when
the base rebuilds, and v1.8.13's base was cached — so smoke is where the recurring
gate belongs.
2026-09-07 21:21:16 +02:00
19 changed files with 4197 additions and 98 deletions
+29
View File
@@ -87,6 +87,35 @@ SSH_KEY_PATH=~/.ssh
# MEMPALACE_PI_REMOTE_PATH=/data/feed
# MEMPALACE_PI_DEVICE=
# ── Mailbox notification: MUST BE NAMED, auto-detect CANNOT work here ──
# The mempalace extension polls the logstream for fleet asks addressed to this
# device and queues them into the next turn. That part needs no config. The
# NOTIFICATION that tells the human it happened does, and unset means SILENT
# outside the pi TUI.
#
# Why there is no working default: terminal identity lives in env vars set by
# the emulator (KITTY_WINDOW_ID, TERM_PROGRAM) and `docker exec` does NOT
# forward them — inside the container pi sees only TERM=xterm-256color no matter
# what is rendering it. So "desktop" auto-detection always falls through to
# OSC 777, which Kitty does not implement, and the notification silently does
# nothing: the worst outcome for a feature whose only job is to break a silence.
# Naming the protocol is what makes it fire.
#
# kitty OSC 99 desktop notification (correct for Kitty, incl. over SSH)
# osc777 OSC 777 (tmux/iTerm2/foot and others)
# desktop OSC 99 if KITTY_WINDOW_ID is visible, else OSC 777 — inside a
# container that means effectively always OSC 777, so prefer naming
# 0 / off suppress entirely (in-TUI notify still shows)
# MEMPALACE_MAILBOX_NOTIFY=kitty
#
# Cadence, if the delivery ever feels late: the poll is coupled to session
# activity (it runs when the agent settles), NOT to a wall clock.
# MEMPALACE_MAILBOX_POLL_MS is therefore a FLOOR BETWEEN POLLS (default 300000),
# not a promise of one every 5 minutes — an idle session polls zero times, and
# session start does the first look.
# MEMPALACE_MAILBOX_POLL_MS=300000
# MEMPALACE_MAILBOX_RESURFACE_MS=3600000
# ── LAN access from the container (host-OS-agnostic) ─────────────────
# On VM-backed hosts (macOS OrbStack / Docker Desktop) the container can't
# reach the host's directly-attached LAN peers by default. The entrypoint
+69 -2
View File
@@ -108,7 +108,14 @@ jobs:
id: compute
run: |
# Hash inputs that determine the base image's contents.
# Order is fixed via `find -print0 | sort -z` for reproducibility.
# Order is fixed via `find -print0 | LC_ALL=C sort -z`. The LC_ALL=C is
# load-bearing: sort collates per locale, and a dictionary locale
# (sv_SE/en_US.UTF-8) orders rootfs differently from byte order, giving
# a different hash for identical content (measured 2026-09-22:
# base-f40c4b7b103d under C/C.UTF-8 vs base-d8df62216a81 under
# sv_SE.UTF-8). The runner ships LANG=C.UTF-8 today, so this pin
# changes nothing now; it stops the hash depending on that accident.
# Predicting base_tag locally MUST use the same prefix on this sort.
# Junk filters: __pycache__/*.pyc and macOS metadata are gitignored
# locally but still picked up by `find rootfs -type f` on a clean CI
# checkout. Exclude them defensively.
@@ -120,7 +127,7 @@ jobs:
! -name '*.pyc' \
! -name '.DS_Store' \
! -name '._*' \
-print0 2>/dev/null | sort -z | xargs -0 cat 2>/dev/null
-print0 2>/dev/null | LC_ALL=C sort -z | xargs -0 cat 2>/dev/null
cat entrypoint.sh entrypoint-user.sh
# mempalace-toolkit is cloned in Dockerfile.base at a ref CI
# resolves to a SHA; fold it in so base_tag changes when the
@@ -157,7 +164,67 @@ jobs:
# buildcache silently reuses the layer from whatever pi version was
# current when the cache was first populated. Same class of bug as
# pi-devbox v0.74.0..v0.75.5 (fixed in v0.75.5b 2026-05-23).
# ── release gate ──────────────────────────────────────────────
# Refuse to spend a base build on a tree whose own shell scripts do not lint.
#
# v1.8.14's first attempt is why this exists. smoke and smoke-studio both failed
# at scripts/smoke-test.sh:770 AFTER build-base had already spent ~46 minutes,
# on a defect shellcheck had flagged as SC2289 (severity error) a day earlier:
# the lint workflow went red on the very push that introduced it (run 186) and
# stayed red for runs 187 and 188, unread.
#
# lint.yml deliberately does not run on tag pushes, and its reasoning is sound
# (the tagged tree was already linted on main; a tag-ref lint run sorts above
# the publish run and makes a release look finished before anything ships). The
# missing invariant was never "lint the tag" -- it was "do not RELEASE a tree
# whose lint failed", and only a job inside THIS workflow can enforce that.
#
# ~40 s, ahead of everything expensive, and it runs scripts/lint-shell.sh --
# the same file lint.yml calls, not a second copy that drifts.
#
# scripts/check-doc-drift.sh is here for the same reason and closes the same
# gap -- and for it the gap is strictly worse. Shellcheck judges the TREE:
# green on main is still green at the tag, because the bytes did not move.
# Check 9 judges the tree against UPSTREAM NOW, and the floating refs it
# watches (PI_OBSMEM_REF=master and friends, which resolve-versions below turns
# into SHAs) move with no commit in this repo at all -- so a green reading on
# main carries no information about tag time, and that window is exactly where
# releases live. Worked example: pi-observational-memory moved cba0334 ->
# e7d77dc the day AFTER v1.9.3 was tagged. Nothing went red; it surfaced only
# because someone ran the gate by hand. Without this step a tag can publish a
# component that no CHANGELOG entry names, and the floating ref means no other
# file in the repo would record it either.
#
# Adds ~8 s. No token and no built image: it takes the last published vX.Y.Z
# from the Hub tags API, that release's baked labels from the anonymous
# registry API, and `git ls-remote`s each upstream -- so a plain checkout is
# enough, with no tags or history to fetch. Offline it SKIPs loudly and
# counted rather than passing, so an outage degrades it to a visible skip
# instead of a false green. Residual, accepted: it resolves the refs seconds
# before resolve-versions resolves them again, so an upstream push landing
# inside that window still slips through -- and check 9 on the NEXT release
# would then name it.
lint-gate:
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
steps:
- uses: actions/checkout@v4
- name: Install shellcheck
run: |
apt-get update
apt-get install -y --no-install-recommends shellcheck
- name: "Shellcheck + syntax-check repository scripts (severity: error)"
run: bash scripts/lint-shell.sh
- name: Components the next build would bake differently are named in the CHANGELOG
run: bash scripts/check-doc-drift.sh
resolve-versions:
# Gated: a defective tree must not reach a 46-minute base build.
needs: [lint-gate]
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
+92 -27
View File
@@ -75,31 +75,11 @@ jobs:
# are shell scripts with no extension. -print0/mapfile -d '' so a path
# with a space cannot silently split, and the file count is asserted
# non-zero — a green tick over an empty file set is not a check.
run: |
# Union of two signals, because either alone misses a real case:
# a shebang scan misses a sourced fragment with no shebang, and a
# *.sh glob misses the extensionless tools in rootfs/usr/local/bin/.
# Silent skipping is precisely the failure mode this gate exists to
# prevent, so err toward over-collecting.
mapfile -d '' -t all_files < <(find . -not -path './.git/*' -type f -print0)
sh_files=()
for f in "${all_files[@]}"; do
case "$f" in *.sh) sh_files+=("$f"); continue;; esac
if head -n1 "$f" 2>/dev/null | grep -qE '^#!.*\b(bash|sh)\b'; then
sh_files+=("$f")
fi
done
echo "Checking ${#sh_files[@]} shell file(s)"
if [ "${#sh_files[@]}" -eq 0 ]; then
echo "::error::no shell files found — the shebang scan or the checkout is wrong"
exit 1
fi
shellcheck -S error -f gcc "${sh_files[@]}"
rc=0
for f in "${sh_files[@]}"; do
bash -n "$f" || { echo "::error file=$f::bash -n failed"; rc=1; }
done
exit "$rc"
#
# The implementation moved to scripts/lint-shell.sh on 2026-09-08 so the
# release gate in docker-publish.yml runs the SAME code rather than a
# second copy that drifts. Edit the script, not a copy of it.
run: bash scripts/lint-shell.sh
- name: Gitea shell guard (catches the actionlint blind spot)
# actionlint models GitHub Actions, where the default run shell is
@@ -112,7 +92,7 @@ jobs:
- name: Install actionlint (pinned)
env:
ACTIONLINT_VERSION: 1.7.7
ACTIONLINT_VERSION: 1.7.12
run: |
curl -fsSL \
"https://github.com/rhysd/actionlint/releases/download/v${ACTIONLINT_VERSION}/actionlint_${ACTIONLINT_VERSION}_linux_amd64.tar.gz" \
@@ -147,7 +127,7 @@ jobs:
- name: Install hadolint (pinned)
env:
HADOLINT_VERSION: 2.14.0
HADOLINT_VERSION: 2.15.1
run: |
curl -fsSL \
"https://github.com/hadolint/hadolint/releases/download/v${HADOLINT_VERSION}/hadolint-Linux-x86_64" \
@@ -157,3 +137,88 @@ jobs:
- name: Run hadolint
run: hadolint Dockerfile.base Dockerfile.variant
skill-floor:
# Gate the VENDORED pi-extensions skill snapshot in rootfs/ against the
# package repo it is a snapshot of. Its own job rather than a step in
# `actionlint`, so "the floor is stale" is a distinct red name in the runs
# list instead of being buried in a lint job that is about something else.
#
# The gap it closes, measured 2026-09-10: the floor sat at 34284 B, untouched
# since fa04d20 (2026-07-30), while the package copy was 38973 B.
# Dockerfile.variant copies the fresh package copy over the SERVED path but
# never writes back to the floor, so nothing in the repo ever noticed. That
# matters because the floor is a FALLBACK: the copy is guarded by
# `if [ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone
# yields no skill/ ships the vendored snapshot and still goes green, with no
# manifest flag or label saying which copy was served.
#
# Gating on another repo is normally a smell; it is proportionate here
# because the check compares the skill DIRECTORY hash, so it can only fire
# when that directory actually changed — which is exactly when the floor has
# gone stale. pi-extensions commits that leave skill/ alone cannot turn this
# red. No secret is needed either: the repo is anonymously clonable (verified
# 2026-09-10 with `git ls-remote` and no credentials), so this cannot start
# failing when a token expires.
#
# Exit codes are 0 in sync / 1 drift / 2 cannot-run, matching
# scripts/lint-shell.sh: a gate that cannot run must not pass, so an
# unreachable package repo is a red 2 rather than a green tick.
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
steps:
- uses: actions/checkout@v4
- name: Vendored pi-extensions skill floor matches the package
run: bash scripts/check-skill-floor.sh
doc-drift:
# Gate hand-maintained doc claims against the build files they describe.
# Its own job for the same reason as skill-floor: "the docs lie" should be a
# distinct red name, not a line buried in a job about workflow syntax.
#
# The gap it closes, measured 2026-09-10 while preparing v1.9.0 — five
# claims had rotted, every one of them a fact written by hand in a file
# nothing verified:
# * README.md's "Version pins" table was wrong on ALL THREE rows (pi
# 0.84.4 vs 0.85.1, pi-atelier v0.10.0 vs v0.10.1, mempalace 3.8.0 vs
# 3.9.0) — and that table exists specifically to be the reviewable
# record of what the repo freezes on purpose, so a wrong row destroys
# the only thing it is for.
# * README.md listed already-shipped typst PDF export under "Planned for
# an upcoming minor release", marked "(shipped in Unreleased/base)".
# * DOCKER_HUB.md claimed "Node.js v22" while v1.9.0 ships Node 24.
#
# DOCKER_HUB.md is why this is a gate and not a habit. It is PUBLISHED —
# update-description POSTs it to Docker Hub as full_description on every tag
# — and it had gone eight releases (v1.8.6 -> v1.9.0) untouched. Nothing
# generates it and nothing checked it, so the only thing keeping it true was
# someone remembering. It is also read from the TAG, so a fix pushed to main
# after tagging never reaches the published page.
#
# Two classes of check. 1-7 are hermetic: each compares a doc string against
# a value that exists in this repo — no network, no token, no built image.
# 8-9 compare against what is PUBLISHED, because those claims have no
# in-repo anchor and rotted for exactly that reason: 8 reads Docker Hub's
# measured sizes; 9 reads the ref labels baked into the last released image
# (anonymous registry API, no docker/crane) and `git ls-remote`s each
# floating upstream, then requires every component the next build would
# bake differently to be NAMED in the CHANGELOG above that release's
# heading. Both SKIP loudly and counted when offline — a skip is neither OK
# nor a failure. Claims that genuinely need a running container (the "N
# mempalace_* tools" count, uncompressed sizes) are still left out — a gate
# that cannot evaluate a claim honestly would have to guess, and a guessing
# gate is worse than none. Assert those in scripts/smoke-test.sh instead.
#
# Exit codes 0 in sync / 1 drift / 2 cannot-run, matching lint-shell.sh and
# check-skill-floor.sh. A renamed ARG makes the gate blind, so that is a red
# 2, not a green tick.
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
steps:
- uses: actions/checkout@v4
- name: Doc claims match the build files
run: bash scripts/check-doc-drift.sh
+75 -3
View File
@@ -92,7 +92,56 @@ re-brand of opencode-devbox's `pi-only` variant.
is hashed into `base_tag`, so it costs a base rebuild (~67 min); and if the
section the phrase canary names has changed, re-pin it in
`scripts/smoke-test.sh`.
3. Update `CHANGELOG.md` Unreleased → vX.Y.Z section.
3. **Update the docs this release makes stale — BEFORE you tag.** Rename
`CHANGELOG.md`'s `## Unreleased` to `## vX.Y.Z — YYYY-MM-DD` (em dash, as
every prior release heading uses), then run the gate:
```bash
bash scripts/check-doc-drift.sh # 0 in sync / 1 drift / 2 cannot run
```
It compares README.md's version-pin table against the ARGs it names, and
DOCKER_HUB.md's Node claim against `ARG NODE_VERSION`, plus Hub's
25 000-char limit, unsubstituted `{{PLACEHOLDERS}}`, and stale `Unreleased`
pointers in user-facing docs. With network it also checks DOCKER_HUB.md's
size claims against Hub's measured sizes (check 8) and — check 9 — that
**every component the next build would bake differently from the last
published release is named in the CHANGELOG** above that release's heading:
it reads the `se.jordbo.pi-devbox.*-ref` labels off the published image and
`git ls-remote`s each floating `*_REF`. A red check 9 means an upstream
(pi-toolkit, pi-extensions, mempalace-toolkit, pi-fork,
pi-observational-memory, pi-studio) moved and no entry names the new SHA;
the failure prints the compare URL; a `PI_VERSION` or `MEMPALACE_VERSION`
bump is caught the same way via the `pi-version` / `mempalace-version`
labels. Name the 7-char SHA (or version) where you describe
the change — that is what the old "Dependency audit" tables recorded by
hand, now required.
**Why before and not after:** `docker-publish.yml` runs `actions/checkout@v4`
with no `ref:`, so every job reads `github.ref` — the **tag**. A doc fix
pushed to `main` after tagging does not reach the release, and for
`DOCKER_HUB.md` it does not reach the published Hub page either, because
`update-description` POSTs that file as Docker Hub's `full_description` from
the tag's tree. Getting it in afterwards means re-pointing the tag, which is
its own hazard (v1.8.14 went `601fc98` → `361babd` and broke deploy
verification until `git fetch --tags --force`).
The gate is deliberately narrow — it only checks claims verifiable from files
in this repo. Still eyeball, because these are NOT gated:
- counts and sizes (`~1.1 GB`, "N `mempalace_*` tools", "7 extensions") —
they need a running image; assert them in `scripts/smoke-test.sh` instead
- feature prose that quietly became false, e.g. a "Planned for an upcoming
release" section describing something that already shipped
- `Dockerfile.base`'s `# BASE_REBUILD_DATE:` marker. Ungated on purpose:
`base_tag` hashes Dockerfile.base's content, comments included, so
demanding it be current would force a ~60 min base rebuild on a release
that touched no base files. **Fix it when the base is already rebuilding —
then it is free.**
Measured cost of skipping this, 2026-09-10 (v1.9.0): five stale claims, one
of them published. README's pin table was wrong on all three rows, and
DOCKER_HUB.md — untouched for eight releases — still said Node v22 while the
image shipped Node 24.
4. Verify `docker compose up` works locally with the current `latest` image
if you're upgrading users from a previous version. Then run the
**post-recreate sanity check** inside the running container to confirm
@@ -243,8 +292,10 @@ shipped the same image bytes); preventatively fixed for `PI_VERSION` +
image. Verifies binaries, repo clones, runtime deployment (waits for
keybindings + mempalace bridge + ≥4 extensions before sampling — fixes
the parallel-build-load race documented in opencode-devbox c6f9d11
2026-06-08), and image size threshold (3500 MB; revisit after a few
releases as actuals settle).
2026-06-08), build-time leftovers (see below), and image size threshold
(3800 MB in `SIZE_THRESHOLD_MB`; revisit after a few releases as actuals
settle — this doc said 3500 until 2026-09-11, after the bar had already
moved twice).
If smoke fails on size threshold but build is otherwise fine: bump
`SIZE_THRESHOLD_MB` in scripts/smoke-test.sh in a follow-up commit and
@@ -252,6 +303,27 @@ re-run. The threshold exists to catch *runaway* growth (an accidental
texlive bake-in, a forgotten chrome dependency), not to block ordinary
upstream bumps.
**The size gate is not a substitute for naming the residue.** It carries
~225 MB of deliberate margin, so v1.9.1 shipped +131 MB of pure build
residue — 110 MB of it npm's own download cache under `/root/.npm`, the
rest foreign platform packages — and stayed green. Four named assertions
now cover that ground: no foreign npm-11 platform packages beyond the
host arch (`@esbuild/*`, `@mariozechner/clipboard-*`), no `/root/.npm` in
the image, and — because the prune's real risk is *removing something
needed*, not size — esbuild must compile TS and clipboard must load its
native binding at **every** install site.
Two failure shapes to copy from those, both of which bit here:
- `test ! -d /root/.npm` on mode-700 `/root` passes for a **permission**
error, so the cache assertion refuses to run as non-root. Watch for
this in any assertion about a path you may not be allowed to read.
- `node -e 'require("esbuild")'` resolves by walking up from the CURRENT
DIRECTORY, so it fails with `MODULE_NOT_FOUND` from `/workspace` on a
perfectly healthy image (esbuild is nested inside the pi trees;
`NODE_PATH` is unset). Always path-qualify: `require("<abs>/esbuild")`.
A runbook shipped the bare form with "if this fails, revert the
release" attached, and it duly went red for the wrong reason.
## Build pipeline notes
- **Two-phase**: base + variant. Base is rebuilt only when
+1647 -1
View File
File diff suppressed because it is too large Load Diff
+8 -6
View File
@@ -8,12 +8,12 @@ A self-contained Docker container for the [pi coding-agent](https://github.com/e
| Tag | Architectures | Size (compressed) | What you get |
|---|---|---|---|
| `joakimp/pi-devbox:latest` | amd64, arm64 | ~1.1 GB | Self-contained: base + pi `{{PI_VERSION}}` + companions |
| `joakimp/pi-devbox:latest` | amd64, arm64 | ~1.23 GB | Self-contained: base + pi `{{PI_VERSION}}` + companions |
| `joakimp/pi-devbox:vX.Y.Z` | amd64, arm64 | same | Pinned semver release |
| `joakimp/pi-devbox:latest-studio` | amd64, arm64 | ~1.15 GB | `latest` + [pi-studio](https://github.com/omaclaren/pi-studio): browser prompt editor, KaTeX/Mermaid preview, tmux-backed literate REPLs |
| `joakimp/pi-devbox:latest-studio` | amd64, arm64 | ~1.25 GB | `latest` + [pi-studio](https://github.com/omaclaren/pi-studio): browser prompt editor, KaTeX/Mermaid preview, tmux-backed literate REPLs |
| `joakimp/pi-devbox:vX.Y.Z-studio` | amd64, arm64 | same | Pinned semver studio release |
| `joakimp/pi-devbox:base-latest` | amd64, arm64 | ~1.0 GB | Base layer alias (internal building block; pull `:latest` instead) |
| `joakimp/pi-devbox:base-<hash>` | amd64, arm64 | ~1.0 GB | Content-addressed base; immutable. Stable parent for variant rebuilds. |
| `joakimp/pi-devbox:base-latest` | amd64, arm64 | ~1.17 GB | Base layer alias (internal building block; pull `:latest` instead) |
| `joakimp/pi-devbox:base-<hash>` | amd64, arm64 | ~1.17 GB | Content-addressed base; immutable. Stable parent for variant rebuilds. |
> **pi-studio (`-studio` tags):** launch with `/studio --no-browser --port 8765` inside a pi session. The server binds `127.0.0.1` **inside the container**, so reach it via host networking or a loopback bridge (and `ssh -L` for a remote host; mosh needs a parallel `ssh -L`). Full recipe: [README → Using pi-studio](https://gitea.jordbo.se/joakimp/pi-devbox#using-pi-studio--studio-variant).
@@ -56,6 +56,8 @@ Full setup guide — authentication for each provider (Anthropic, OpenAI, Gemini
The entrypoint deploys/registers all of these on first container start. Re-running is idempotent and preserves user edits.
**Terminal UI mode — fullscreen by default.** pi 1.0.0 made the TUI fullscreen, and this image adopts upstream's default. Fullscreen uses the terminal's alternate screen, so the transcript no longer lands in your terminal's (or tmux's) native scrollback. To get the previous behaviour back, set `"tuiMode": "regular"` in `~/.pi/agent/settings.json`, or pass `pi --tui-mode regular` for a single session. The bundled pi-atelier sidebar works in both modes. See the README for the related `fullscreenExitOutput` / `fullscreenScrollbar` / `fullscreenCopyOnSelect` / `fullscreenWheelScrollLines` settings — the last one behaves differently over SSH, which is how this container is usually driven.
### MemPalace (persistent agent memory)
- **MemPalace** + MCP server — semantic search over conversation history, knowledge graph, diary; queryable via 29 `mempalace_*` tools inside pi
@@ -80,7 +82,7 @@ The entrypoint deploys/registers all of these on first container start. Re-runni
- **Editor**: neovim (system-wide `termguicolors` default; bring your own config/plugins), tmux (configured for 0-indexed sessions)
- **Search/nav**: ripgrep, fd, fzf, zoxide
- **Display**: bat, eza, htop, tree
- **Data**: jq, yq
- **Data**: jq, yq, sqlite3, bc/dc, column
- **Help**: tldr (tealdeer — Rust port; run `tldr --update` once to populate cache)
- **Git**: git-lfs, git-crypt, gitleaks (for pre-commit secret scanning)
- **Build**: gcc, g++, make, patch
@@ -94,7 +96,7 @@ The entrypoint deploys/registers all of these on first container start. Re-runni
uv run --with jupyterlab jupyter lab --no-browser --port 8888
uv run --with marimo marimo edit
```
- **Node.js** v22 + npm (used by pi itself)
- **Node.js** v24 LTS + npm (used by pi itself)
- **Rust** — `rustup-init` is on PATH; install toolchains on demand
- **Go** — opt-in via `--build-arg INSTALL_GO=true` if rebuilding from source
+218 -12
View File
@@ -14,7 +14,7 @@
# content-addressed over this file, so any byte change invalidates the
# cache. Recommended cadence: once per release for security updates.
#
# BASE_REBUILD_DATE: 2026-07-13 (Unreleased — agent-browser CLI + Playwright Chromium for headless browser automation; prior: typst PDF engine + xz-utils + pandoc typst-template default-font patch)
# BASE_REBUILD_DATE: 2026-09-22 (v1.9.4 — mempalace 3.9.0 -> 3.10.0 with ENV MEMPALACE_CONFIG_DIR pinning the layout, mempalace-toolkit 2167a1b explicit event_list order; previous marker 2026-09-19 / v1.9.3)
#
# ── Lineage note ─────────────────────────────────────────────────────
# Adapted from opencode-devbox/Dockerfile.base (commit before v1.16.2).
@@ -101,6 +101,116 @@ ENV DEBIAN_FRONTEND=noninteractive
# container: `ss` lands at /usr/bin/ss, `ip` at /usr/sbin/ip
# (both already on the developer PATH), and `portcheck --all`
# then correctly identifies the socat listener on 8765.
# shellcheck — shell linter. Added 2026-09-09 to close a CAPABILITY gap, not
# a style preference. `scripts/lint-shell.sh` is the release
# GATE (the `lint-gate` job that `resolve-versions` depends
# on), and it correctly refuses to pass when shellcheck is
# missing — "a gate that cannot run must not pass". Measured on
# v1.8.14: shellcheck was absent from this image by all three
# routes (PATH, dpkg, filesystem), so `bash
# scripts/lint-shell.sh` exited 2 in EVERY devbox container and
# no developer could run the release gate locally at all. The
# loop was therefore write-shell → push → wait for CI → discover,
# which is the loop the gate was added to shorten: v1.8.14's
# first attempt burned ~46 min on a tree whose lint had already
# been red for 24 h. This is also what makes a client-side
# pre-push hook possible (see hooks/pre-push); without the
# binary that hook would refuse every push. ~39 MB installed
# (Installed-Size 40112 KB, shellcheck 0.10.0-1) and measured
# to pull ZERO additional packages under
# --no-install-recommends: its deps (libc6, libffi8, libgmp10)
# are already present. NOTE this file feeds the base-decide
# hash (Dockerfile.base + rootfs/), so adding it forces one
# full base rebuild.
# bind9-dnsutils — `dig` and `nslookup`. Added 2026-09-10 to close a
# DIAGNOSTIC gap measured during the gitea.egl.lan/FreeIPA
# work: the container could resolve names but had NO way to
# ask a SPECIFIC nameserver anything. `getent hosts` only
# follows the resolver's default path, so the whole "gateway
# 172.16.88.1 returns NXDOMAIN for the egl.lan zone while
# 10.20.253.1 is authoritative for it" diagnosis had to be
# hand-rolled in python3 — dig, host AND nslookup were all
# absent. `dig @10.20.253.1 freeipa-4.egl.lan` is the
# one-liner that replaces it, and split-horizon DNS is a
# recurring class of bug on this fleet, not a one-off. NOTE
# the package to name is bind9-dnsutils: plain `dnsutils` is
# a transitional package in trixie. ~6.1 MB total (6210 KB
# measured): bind9-dnsutils 721 KB + bind9-host 161 KB +
# bind9-libs 3804 KB plus 7 small libs (libfstrm0,
# libjson-c5, liblmdb0, libmaxminddb0, libprotobuf-c1,
# liburcu8t64, libuv1t64) under --no-install-recommends.
# ldap-utils — `ldapsearch`/`ldapmodify`. Added 2026-09-10. This fleet
# authenticates against FreeIPA (EGL.LAN), and every LDAP
# probe during the Gitea auth work had to be run by SSHing to
# an already-enrolled host because the container had no LDAP
# client at all. 1244 KB and pulls NOTHING extra under
# --no-install-recommends — its deps (libldap, libsasl2) are
# already present. CAVEAT: this gives SIMPLE binds only,
# which is what Gitea itself uses and what most probes need.
# GSSAPI binds (`ldapsearch -Y GSSAPI`) additionally require
# krb5-user + libsasl2-modules-gssapi-mit, deliberately NOT
# added here — that is a Kerberos-client decision with
# /etc/krb5.conf implications, not just a tool.
# xxd — hex dump. 198 KB, no extra deps. Convenience, and honestly
# marginal: `od -c` from coreutils is always present and does
# the same job. Earned its place because verifying that
# git-crypt actually encrypted a staged blob (the \0GITCRYPT\0
# magic) is a recurring check in myconfigs and xxd is the
# muscle-memory command for it.
# NOT added — netcat-openbsd (133 KB): measured redundant on
# 2026-09-10, because socat is already baked above AND bash's
# /dev/tcp does reachability checks with zero packages
# (verified against gitea.egl.lan:3000). Recorded here so the
# omission reads as a decision rather than an oversight.
# sqlite3 — the `sqlite3` CLI. Added 2026-09-28 at ALC's request. 587 KB.
# This is the one that was genuinely missing rather than merely
# absent: MemPalace keeps BOTH the palace and the logstream as
# SQLite files (palace/chroma.sqlite3, logstream.sqlite3),
# mempalace_status reports a sqlite_integrity block, and every
# integrity or forensics check on this fleet has so far been
# done through a python3 -c one-liner because the CLI did not
# exist in the image.
# bc, dc — arbitrary-precision calculators. 236 KB + 149 KB. SEPARATE
# binary packages on Debian (dc split out of bc before
# bookworm), so both must be named — verified with apt-cache on
# a trixie host, not assumed. Added 2026-09-28. HONEST
# RATIONALE: low value on their own, because awk, python3 and
# perl are all already baked and each is strictly more capable.
# They are here because copy-pasted shell snippets assume bc
# exists. Note what actually went wrong on 2026-09-28, since it
# was NOT bc's absence: `printf '%6.2f'` was handed the empty
# output of the missing bc and rendered it as a confident
# "0.00 days" for a figure that was really 4.81 days. A missing
# tool that formats as a plausible number is worse than one
# that fails loudly, and no package fixes that — only not
# taking a formatted value on trust does.
# bsdextrautils— ships /usr/bin/column. 339 KB. Added 2026-09-28. Confirmed
# with `dpkg -S /usr/bin/column` on trixie rather than inferred
# from the pre-bullseye bsdmainutils name, which is where column
# used to live and is the obvious way to get this wrong.
# MEASURED COST OF ALL FOUR: 1311 KB total and ZERO transitive
# packages — libsqlite3-0, libreadline8t64, zlib1g,
# libsmartcols1 and libtinfo6 are each already present in the
# image (checked with dpkg-query, all five report
# "install ok installed"), so under --no-install-recommends
# nothing new is pulled.
# NOT added — datamash, xsv/csvkit: considered 2026-09-28 and
# declined. Python's stdlib csv module handled a real 7-file
# Excel-export concatenation that day (BOM, CRLF, embedded
# newlines inside quoted fields) correctly and without them.
# python3-yaml — PyYAML. Added 2026-09-10 for precisely the same reason as
# shellcheck above: a gate this repo ALREADY OWNS could not be
# run locally by anyone. scripts/check-workflow-shell.sh — the
# guard that catches the "bash-only syntax under Gitea's default
# sh/dash shell" footgun that broke resolve-versions (ed49b8d)
# and promote-base-latest (b7197e8) — hard-exits with "ERROR:
# python3 yaml module missing" without it. lint.yml installs it
# explicitly in CI (`shellcheck python3-yaml`), which is itself
# the evidence that the image lacked it. Measured 2026-09-10
# while wiring the skill-floor job: the guard could not be run
# before pushing — the same write → push → wait-for-CI loop that
# shellcheck was baked to shorten. 552 KB, and pulls ZERO extra
# packages under --no-install-recommends.
RUN apt-get update && \
apt-get upgrade -y --no-install-recommends && \
apt-get install -y --no-install-recommends \
@@ -120,6 +230,7 @@ RUN apt-get update && \
make \
patch \
diffutils \
shellcheck \
git-crypt \
age \
file \
@@ -141,6 +252,14 @@ RUN apt-get update && \
kitty-terminfo \
ncurses-term \
iproute2 \
bind9-dnsutils \
ldap-utils \
xxd \
python3-yaml \
sqlite3 \
bc \
dc \
bsdextrautils \
&& ln -s /usr/bin/fdfind /usr/local/bin/fd \
&& apt-get clean \
&& rm -rf /var/lib/apt/lists/*
@@ -373,7 +492,10 @@ RUN ARCH=$(case "${TARGETARCH}" in amd64) echo "x86_64" ;; arm64) echo "aarch64"
# uninterruptibly. A stall-kill is no longer a permanent latch either: the
# next tool call respawns the server with capped exponential backoff (the
# budget resets on any successful response). Tunables:
# MEMPALACE_MCP_TIMEOUT_MS (default 60000), MEMPALACE_MCP_INIT_TIMEOUT_MS
# MEMPALACE_MCP_TIMEOUT_MS (default 60000; the feed's `mempalace_mine` carries
# its own longer MEMPALACE_FEED_MINE_TIMEOUT_MS, default 300000, since toolkit
# 817b3a8 — before that the 60 s deadline cut every honest mine off),
# MEMPALACE_MCP_INIT_TIMEOUT_MS
# (default 300000 — generous so a genuine first cold-open isn't killed),
# MEMPALACE_MCP_MAX_RESPAWNS (default 2; 0 disables self-heal),
# MEMPALACE_MCP_RESPAWN_BACKOFF_MS (default 1000); timeouts of 0 disable.
@@ -450,14 +572,22 @@ ARG INSTALL_MEMPALACE=true
# the part that should stay manual.
#
# Deployment sequencing note for whoever ships this bump: synlig (the shared
# central palace host) serves mempalace 3.8.0 SERVER-SIDE via
# docker-compose.mempalace.yml, which reuses this same devbox image. (Measured
# 2026-09-06 over ssh: synlig's UV_TOOL_DIR mempalace entry last changed
# 2026-08-25 15:33 — this comment previously said 3.7.1, which was stale.)
# Bumping this ARG changes only the CLIENT version baked into pi-devbox
# images: it introduces client/server skew until synlig's compose stack is
# separately rebuilt/redeployed with the new pin. Not something to code around
# here — just sequence the redeploy.
# central palace host) serves mempalace SERVER-SIDE as a `uv tool` install run
# by the systemd unit `mempalace-serve.service` (`python -m mempalace.mcp_server
# --transport http`), NOT via docker-compose.mempalace.yml — that compose file
# exists in this repo but is not what runs there. (Measured 2026-09-22 over
# ssh: `uv tool list` -> mempalace v3.9.0, python 3.12.13, chromadb 1.5.9;
# `docker ps` matched no palace container. This comment previously said the
# compose stack served 3.8.0, which was stale on both counts.) Bumping this ARG
# changes only the CLIENT version baked into pi-devbox images: it introduces
# client/server skew until synlig's tool is upgraded (`uv tool upgrade
# mempalace` + restart the unit). Not something to code around here — just
# sequence the upgrade. And note which side OWNS what: MCP tool semantics
# (event_list ordering, kg_timeline pagination, search result fields) come
# from the SERVER the extension talks to over MEMPALACE_REMOTE_URL, so they
# change when synlig upgrades; only the local CLI (`mempalace init` at first
# run, the mempalace-pi-session feeder) and the on-disk layout under
# ~/.mempalace change when THIS pin does.
#
# v1.8.13: 3.8.0 -> 3.9.0. Audited: no Breaking/Removed changelog headings.
# Adopted mainly for #2281 (`mempalace_mine` accepts a single conversation
@@ -469,7 +599,61 @@ ARG INSTALL_MEMPALACE=true
# (release awareness, `task create`/`task launch` MCP tools) are SERVER-side,
# so they stay dark until synlig is redeployed — a client bump alone cannot
# light them up.
ARG MEMPALACE_VERSION=3.9.0
#
# v1.9.4: 3.9.0 -> 3.10.0 (PyPI 2026-09-15). Deferred at v1.9.3 for two
# "Upgrade notes" items; both re-measured against the 3.10.0 wheel, one needed
# an adaptation:
# - "New installs keep config and palace under ~/.config/mempalace". The
# resolution order is $MEMPALACE_CONFIG_DIR, then ~/.mempalace IF it holds
# config.json / people_map.json / palace/chroma.sqlite3, then XDG. An
# EMPTY ~/.mempalace does not count — and an empty ~/.mempalace is exactly
# what a freshly mounted devbox-palace volume (or entrypoint.sh's mkdir on
# a volume-less container) looks like at first boot. Measured with a fresh
# $HOME: `mempalace init` wrote to ~/.config/mempalace, outside the
# persisted path, and entrypoint-user.sh's first-run test
# `[ ! -d ~/.mempalace/palace ]` would stay true on every start. With
# MEMPALACE_CONFIG_DIR set, everything landed in ~/.mempalace. Hence the
# ENV MEMPALACE_CONFIG_DIR below (in the non-root-user section, where
# ${USER_NAME} is in scope): first in the resolution order, so the
# heuristic never runs and the image's layout contract no longer depends
# on it. Existing volumes were safe either way (config.json is a legacy
# marker); the ENV is for first boots. palace_path still defaults to
# <config_dir>/palace and MEMPALACE_PALACE_PATH is still honoured (config.py
# :927), so scripts/smoke-test.sh's stage-path test keeps its meaning.
# - "MCP event listing returns the newest events first when no cursor is
# given". SERVER-side (see above), so it lands when synlig upgrades, not
# here. mempalace-toolkit 2167a1b made every cursor-less event_list call
# in the pi extension say `order: "desc"` explicitly, so the mailbox reads
# the same window against either server version.
# Also in the notes, neither reaching this image: `mempalace rules` dropped
# `--agent` (no caller in pi-devbox, mempalace-toolkit, skillset or myconfigs);
# `get_collection()` refuses unknown collection names (library callers only).
# MCP tool-schema review, as always: no tool removed or renamed; additive
# fields on search results (filed_at / content_date provenance), `limit` /
# `offset` on kg_timeline, `last_modified` on drawers. Skew while synlig stays
# on 3.9.0 is narrower than it looks: the pi extension speaks HTTP to the hub
# (no local mempalace-mcp is spawned), and the feeder in remote mode stages
# locally in python, rsyncs, and calls the hub's own `mempalace_mine` tool.
# Measured 2026-09-22 with a PATH shim in front of `mempalace`: a 47-session
# `mempalace-pi-session --dry-run` made ZERO local CLI calls (the shim's
# positive control logged one). So 3.10.0's new CLI write-routing policy never
# runs against the hub from this image; the client pin touches first-run
# `mempalace init` and the on-disk layout, nothing else in remote mode.
ARG MEMPALACE_VERSION=3.10.0
# Recorded as a label HERE, not in Dockerfile.variant, for three reasons: the
# value lives next to the ARG that defines it (a second copy in the variant
# would be one more pin able to drift, which is the class check-doc-drift.sh
# exists to catch); labels are inherited by every image built FROM this one, so
# both variants carry it with no build-arg to plumb through four call sites;
# and inheritance means the label states the pin of the base the variant
# ACTUALLY built on — which is the question when base-decide cache-hits an
# older base. Like every se.jordbo.pi-devbox.* label this records INTENT; the
# ground truth is /etc/pi-devbox/build-manifest.json's mempalace_version, read
# from the installed binary, and scripts/smoke-test.sh asserts the two agree.
# check-doc-drift.sh check 9 reads this off the last published image so that a
# pin bump must be named in the CHANGELOG — until this label ships, that
# component reports SKIP (label absent on the published release), not OK.
LABEL se.jordbo.pi-devbox.mempalace-version="${MEMPALACE_VERSION}"
ENV UV_TOOL_DIR=/opt/uv-tools
ENV UV_TOOL_BIN_DIR=/usr/local/bin
RUN if [ "${INSTALL_MEMPALACE}" = "true" ]; then \
@@ -554,7 +738,17 @@ ENV COLORTERM=truecolor
ENV PATH="/home/developer/.local/bin:/home/developer/.cargo/bin:${PATH}"
# ── Node.js (required for pi + MCP servers + tldr) ──
ARG NODE_VERSION=22
# 24 (LTS "Krypton"), raised from 22 on 2026-09-10 because the image was BELOW a
# DECLARED requirement, not merely behind the newest release: `agent-browser`
# publishes engines.node ">=24.0.0", so every build on 22 installed it with an npm
# EBADENGINE warning and then ran it outside its supported range — measured on
# v1.8.14, which shipped node 22.23.2 with agent-browser 0.37.1. The other two npm
# consumers are satisfied either way: pi declares ">=22.19.0" and playwright
# ">=20". Verified before bumping, because a missing NodeSource suite would break
# the build for every arch at once: deb.nodesource.com/setup_24.x returns HTTP 200
# and the node_24.x suite advertises `Architectures: amd64 arm64 armhf x86_64`, so
# both the arm64 fleet and the amd64 CI runners resolve.
ARG NODE_VERSION=24
RUN curl -fsSL --retry 5 --retry-delay 5 --retry-all-errors https://deb.nodesource.com/setup_${NODE_VERSION}.x | bash - && \
apt-get install -y --no-install-recommends nodejs && \
rm -rf /var/lib/apt/lists/*
@@ -726,6 +920,18 @@ print('chromadb embedding model warmed: all-MiniLM-L6-v2')" && \
ENV NPM_CONFIG_PREFIX=/home/${USER_NAME}/.pi/npm-global
ENV PATH="/home/${USER_NAME}/.pi/npm-global/bin:${PATH}"
# ── MemPalace config/palace root: pin it, do not let a heuristic pick it ──
# mempalace >= 3.10.0 resolves its config dir as $MEMPALACE_CONFIG_DIR, then
# ~/.mempalace ONLY if it already holds a config/palace, else ~/.config/mempalace
# (XDG). An empty ~/.mempalace — a fresh devbox-palace volume, or entrypoint.sh's
# mkdir on a volume-less container — fails that test, so a first boot would put
# the palace outside the persisted path and re-run first-run init forever. This
# ENV is first in the order, so the layout is what entrypoint.sh (mkdir),
# entrypoint-user.sh (first-run test), the feeder's <palace-root>/pi-stage and
# scripts/recreate-sanity-check.sh all already assume. Rationale and the
# measurement live with ARG MEMPALACE_VERSION above; keep the two in step.
ENV MEMPALACE_CONFIG_DIR=/home/${USER_NAME}/.mempalace
# ── Shell defaults (bash history, aliases, readline) ─────────────────
RUN mkdir -p /etc/skel-devbox
COPY rootfs/home/developer/.bash_aliases /etc/skel-devbox/.bash_aliases
+278 -5
View File
@@ -113,7 +113,91 @@ ARG USER_NAME=developer
# signature is ~5s of sustained CPU. Two-sided check: the atelier sidebar
# painted ACTIVITY+WORKSPACE identically to the 0.84.4 control, so the test
# could distinguish "loaded" from "silently absent".
ARG PI_VERSION=0.85.1
#
# v1.9.4: HELD at 0.85.1 while 0.86.1 and 0.87.0 exist upstream. 0.87.0's
# changelog: "Removed the inherited `shouldStopAfterTurn` agent option. Use
# `finishTurn` and return `{ action: "end" }` instead" (no issue number on
# that line); 0.86.0 moved provider stream inputs to `TranscriptContext`, with
# system prompts read from `context.messages`. pi-observational-memory 3.1.4
# (the `master` ref baked below) still uses both the old option and
# `AgentContext.systemPrompt` in its observer, reflector and dropper workers —
# measured 2026-09-22 in src/agents/*/agent.ts, both the baked 3.1.3 tree and
# upstream master 3.1.4: `shouldStopAfterTurn` 1 per worker (3), `finishTurn`
# 0, `systemPrompt` 1 per worker. Its peerDependencies are `*`,
# so nothing at install time would refuse; it would break at runtime (turn
# caps ignored, workers losing their specialised prompts). Upstream tracks it
# as pi-observational-memory #82 with fix PR #83 (opened 2026-09-21, mergeable,
# not merged at this writing). Bump pi and pi-obsmem TOGETHER once #83 has
# shipped in a release. The other extensions were checked against the 0.86.0
# and 0.87.0 breaking lists and are clean: ssh-controlmaster's `user_bash`
# handler already returns `undefined | { operations }` (0.86.0 fail-closed
# contract) and its registerTool calls spread the built-in tools so they
# carry parameter schemas (#9300). 0.86.1 as an intermediate is untested and
# not worth the pty matrix for a stop that #83 will make moot.
#
# v1.9.5: 0.85.1 -> 0.87.1, and pi-obsmem moves off `master` to a pinned SHA in
# the SAME commit, because neither is safe alone. 0.87.0 REMOVED
# `shouldStopAfterTurn`, which 3.1.4 still used; the pinned tip uses
# `finishTurn`, which does not exist before 0.87.0. So 3.1.4 + 0.87.1 silently
# ignores turn caps, and the new tip + 0.85.1 breaks the workers outright — the
# pair only works together, exactly as the v1.9.4 note predicted.
# MEASURED 2026-10-01 across src/agents/{observer,reflector,dropper}/agent.ts,
# deliberately the same counting method as the v1.9.4 audit above so the two
# numbers are comparable:
# 3.1.4 (e7d77dc): shouldStopAfterTurn 3, finishTurn 0, systemPrompt 3
# 731c3d4 (pinned) : shouldStopAfterTurn 0, finishTurn 3, systemPrompt 0
# i.e. all three workers migrated, and the 0.86.0 `AgentContext.systemPrompt`
# reads are gone too. #82/#83 merged 2026-09-23.
# WHY A SHA AND NOT A TAG, departing from the v1.9.4 instruction to wait for a
# release: #83 is merged, but the newest obsmem tag is STILL 3.1.4, cut
# 2026-09-20 — before the merge. Upstream tags slowly while moving master often
# (e7d77dc -> 1529e14 -> 731c3d4 in the nine days to 2026-10-01), so waiting for
# a tag means holding pi indefinitely. A full SHA keeps the one property
# `master` does not have: rebuilding this tag later produces the SAME image.
# v1.9.5 SKIPPED 0.99.x and 1.0.0 for lack of evidence. v1.10.0 went and got the
# evidence instead of waiting for obsmem to mention a version, because "no issue
# names 1.0" is an absence of a statement, not a measurement.
#
# v1.10.0: 0.87.1 -> 1.0.0. MEASURED 2026-10-02 against the published npm
# tarballs for 0.87.1 and 1.0.0, unpacked side by side:
# 1. NO `### Breaking Changes` SECTION IN 1.0.0 AT ALL. The changelog's
# breaking sections belong to 0.87.0, 0.86.0, 0.84.3, 0.84.0, 0.83.0,
# 0.80.8, 0.80.7 and 0.75.0 — none to 0.88+..1.0.0. The major is a
# milestone (fullscreen default, leaner codemode), not an API break. The
# last break that touched us was 0.87.0's `shouldStopAfterTurn` removal,
# which v1.9.5 already absorbed.
# 2. `finishTurn` IS STILL IN 1.0.0's dist, so the obsmem SHA pinned below
# keeps the API it migrated to. (`shouldStopAfterTurn`: absent from both
# 0.87.1 and 1.0.0, as expected after its 0.87.0 removal.)
# 3. EVERY pi.* API our extensions call exists in 1.0.0's dist — all 8 of
# registerTool, registerCommand, registerFlag, getFlag, on, exec,
# sendMessage, sendUserMessage, extracted from mempalace.ts, the
# pi-extensions tree and obsmem's src/.
# 4. engines.node is `>=22.19.0` on both; the image ships 24.x.
# THE 0.99.2 CHANGE THAT LOOKED FATAL AND IS NOT: from 0.99.2 the DEFAULT MCP
# `exposure` is `codemode`, i.e. such tools are "neither declared to the model
# nor listed" and must be found with `searchTools()`. That would gut the
# MemPalace protocol if MemPalace were a builtin-MCP server. It is not: both
# mempalace.ts and mcp-loader.ts run their OWN MCP client and surface tools via
# `pi.registerTool()`, which is why they are named `mempalace_search` and not
# `mcp__mempalace__search`. Confirmed on a second route — settings.json has
# neither an `mcpServers` block (builtin, exposure-governed) nor an `mcp` block.
# Extension-registered tools are declared like built-ins, so `exposure` cannot
# reach them.
# WHAT IS NOT PROVEN, stated so acceptance does not mistake this for cleared:
# grepping dist shows the SYMBOLS survive, not that their SIGNATURES are
# unchanged — necessary, not sufficient. 1.0.0 is also four days of upstream old
# (published 2026-10-01T19:15Z) and neither obsmem nor atelier has a commit
# naming it. Acceptance must prove the obsmem workers CAP TURNS (peerDeps are
# `*`, so a mismatch is silent) and that the atelier sidebar PAINTS.
# USER-VISIBLE BEHAVIOUR CHANGE, decided rather than inherited: 1.0.0 makes the
# TUI fullscreen by default, which replaces the terminal's normal scrollback.
# Upstream's default is ADOPTED on purpose and no `tuiMode` is baked here, so
# the image follows pi instead of pinning the fleet to either mode. The revert
# is documented (README -> "Terminal UI mode"): `"tuiMode": "regular"` in
# ~/.pi/agent/settings.json, `pi --tui-mode regular` for one session, or the
# same key in a project's .pi/settings.json.
ARG PI_VERSION=1.0.0
ARG PI_TOOLKIT_REF=main
ARG PI_EXTENSIONS_REF=main
# Repo URLs default to the canonical gitea origin but are overridable so a
@@ -128,7 +212,14 @@ ARG PI_EXTENSIONS_REPO=https://gitea.jordbo.se/joakimp/pi-extensions.git
ARG PI_FORK_REPO=https://github.com/elpapi42/pi-fork.git
ARG PI_FORK_REF=master
ARG PI_OBSMEM_REPO=https://github.com/elpapi42/pi-observational-memory.git
ARG PI_OBSMEM_REF=master
# PINNED to a full 40-char SHA as of v1.9.5, not `master` — see the PI_VERSION
# note above for the measurement and the reasoning. Moves together with
# PI_VERSION by necessity, not by convention. The width matters: 731c3d4 is the
# same commit, but check-doc-drift.sh recognises a literal SHA only via a
# 40-char match (SHA40), so a short pin would fall through to its
# branch-or-tag lookup, fail, and downgrade that component's drift check to a
# silent SKIP — a pin that reads fine and is no longer verified.
ARG PI_OBSMEM_REF=731c3d49288580f4d79cabdbfbc0d16b34db0f41
# pi-atelier (TUI sidebar: ordered panels, split-pane, themes) is PINNED TO A
# TAG, which CI resolves to that tag's commit SHA — same treatment as
# pi-studio, for reproducibility plus cache-busting.
@@ -170,9 +261,46 @@ ARG PI_ATELIER_REPO=https://github.com/michaelmjhhhh/pi-atelier.git
# old and new pin. Included because it was already exercised: the pty matrix
# for PI_VERSION above ran atelier v0.10.1 against pi 0.85.1 and painted the
# sidebar identically to v0.10.0.
ARG PI_ATELIER_REF=v0.10.1
#
# v1.9.4: v0.10.1 -> v0.10.3 (both v0.10.2 and v0.10.3 released 2026-09-22).
# Fixes only per the release notes: sidebar text/borders preserved beside
# inline images (#53), transcript images hidden while capturing overlays are
# open, Workspace Pulse skips redundant HEAD/diff when nothing tracked changed
# (#61), sidebar height from row counts (#59), git/usage scans suspended while
# disabled, Display Revert + Undo ordering. Checked before bumping: package.json
# at v0.10.3 still declares zero runtime dependencies and no build script (so
# the no-`npm install` reasoning above holds) and peerDependencies are still
# pi >=0.84.0, so it still spans the pinned 0.85.1. 13 commits v0.10.1..v0.10.3,
# all under src/ tests/ docs/ scripts/ plus metadata; no entry-point move.
#
# v1.10.0: v0.10.3 -> v0.13.0, closing the two-minor gap v1.9.5 flagged as the
# residual risk of the pi bump. peerDependencies are UNCHANGED at pi >=0.84.0
# across v0.10.3, v0.12.1 and v0.13.0 — still a FLOOR, so still not evidence of
# anything; the floor is satisfied by 1.0.0 either way. What was actually
# checked, 2026-10-02:
# - v0.13.0 CARRIES ITS OWN BREAKING CHANGE, unrelated to pi: "The package
# entry point now exports only the sidebar contribution protocol
# (`registerSidebarPanel`, guards, size limits, event types). The internal
# registry and layout helpers are no longer exported." Harmless HERE only
# because nothing of ours imports them: a grep for `pi-atelier` and
# `registerSidebarPanel` across mempalace-toolkit, pi-devbox,
# pi-extensions, pi-fork, pi-observational-memory and pi-studio returns
# ZERO matches in all six. atelier is a leaf here — it registers its own
# sidebar and no one consumes its API.
# - Settings churn in the gap, checked against our own tree: v0.12.0 REMOVED
# the `showSessionActions` setting (0 references here) and migrated the
# Control Center shortcut Alt+A -> F6 (0 references here, so no doc goes
# stale). `showSidebarAgent` / `showSidebarTodos`, the only atelier keys
# README names, are both still documented in v0.13.0's README.
# - The TUI internals atelier patches (`TuiMainScreen`, `renderLayoutFrame`)
# are both still present in pi 1.0.0's dist, and atelier's changelog shows
# it has handled fullscreen vs regular renderers explicitly since 0.84.
# - Residual: v0.13.0 is SAME-DAY upstream (2026-10-02). Acceptance proves
# the sidebar PAINTS with the two-sided check from 0.84.4/0.85.1 that can
# tell "loaded" from "silently absent" — "no crash" is not the test.
ARG PI_ATELIER_REF=v0.13.0
# Human-readable tag PI_ATELIER_REF was resolved from; recorded as a label.
ARG PI_ATELIER_VERSION=v0.10.1
ARG PI_ATELIER_VERSION=v0.13.0
RUN set -e && \
# git_fetch_ref: clone-equivalent helper that accepts EITHER a branch name
@@ -196,6 +324,71 @@ RUN set -e && \
done; \
return 1; \
} && \
# prune_foreign_natives: npm 11 (shipped with Node 24) installs EVERY optional
# platform package of a native dependency, not just the one matching the host.
# TWO families are affected in this image, and BOTH have been measured — add a
# family here only after measuring it, never by widening the pattern on a hunch:
#
# @esbuild/<platform> 26 dirs, 284 MB (found first, v1.9.1)
# @mariozechner/clipboard-<triple> 11 dirs, 12 MB per site, 10 MB foreign
#
# esbuild declares those with os/cpu constraints, but npm 11 ignores them and
# ALSO ignores --os/--cpu and an npmrc carrying os=/cpu= (all three measured).
# So prune explicitly, keeping only the host platform, computed from
# `node -p process.arch` so one line stays correct on amd64 and arm64.
# Measured on pi-fork's tree: npm 10.9.8 -> 165 MB, npm 11.19.0 -> 449 MB,
# and the 165 MB figure reproduces what v1.9.0's predecessor actually shipped.
#
# WHY THE CLIPBOARD FAMILY WAS ADDED (2026-09-11): v1.9.1 pruned @esbuild only
# and still shipped +131 MB compressed over v1.8.14. That residual was
# attributed by listing the PUBLISHED arm64 layer tarballs straight from the
# registry (there is no docker CLI inside the container, so `docker history`
# was not available): +110 MB /root/.npm/_cacache (purged below) and +21 MB of
# clipboard platform packages across the two install sites — 131 MB total, so
# the delta is now fully accounted for with no unexplained remainder.
#
# Keeping linux-$arch-{gnu,musl} is deliberate: clipboard's napi-rs loader
# tries ./<name>.node then the platform package, per platform in try/catch, and
# chooses gnu vs musl at runtime from its own isMusl() probe — so both host-arch
# branches must survive. The musl package is a 420-byte stub, i.e. free. The
# bare wrapper `@mariozechner/clipboard` has no hyphen suffix and therefore
# cannot match the regex below. Verified on arm64 against a copy of the real
# tree before this was written: after pruning to those two,
# require('@mariozechner/clipboard') still loads and exports all 18 functions.
# esbuild likewise still compiles TS via transformSync at both install sites.
# This removes dead weight, not function.
#
# MUST run in the SAME layer as the npm installs above: deleting in a later RUN
# leaves the bytes in this layer and shrinks the image by nothing.
# NOTE the single backslash in -printf '%f\n': Docker passes '\\n' through
# verbatim, so v1.9.1's doubled version printed a mangled "li ux-arm64"
# (find emitted a literal backslash, then `tr` translated the n out of the
# name). Confirmed from the published image's own recorded created_by.
prune_foreign_natives() { \
arch="$(node -p process.arch)"; \
find /usr/lib/node_modules /opt -type d -regex '.*/@esbuild/[^/]+' \
! -name "linux-$arch" -prune -exec rm -rf {} + ; \
find /usr/lib/node_modules /opt -type d -regex '.*/@mariozechner/clipboard-[^/]+' \
! -name "clipboard-linux-$arch-gnu" ! -name "clipboard-linux-$arch-musl" \
-prune -exec rm -rf {} + ; \
echo "native platform dirs kept: $(find /usr/lib/node_modules /opt -type d \( -regex '.*/@esbuild/[^/]+' -o -regex '.*/@mariozechner/clipboard-[^/]+' \) -printf '%f\n' 2>/dev/null | sort | uniq -c | tr '\n' ' ')"; \
} && \
# purge_build_caches: the build's own download caches are NOT free — they land
# in whichever layer created them. Measured on the published v1.9.1 arm64
# variant layer: root/.npm/_cacache was 145.2 MB of a 401.9 MB layer (35.2 MB
# in v1.8.14), the single biggest item in the +131 MB residual, because npm 11
# caches every platform tarball it fetched — including the ones just pruned.
# Nothing at runtime reads it: the build runs as root, the container runs as
# `developer` with its own cache under $HOME (and $HOME/.pi is a volume).
# DELIBERATELY NOT purged here: /tmp/node-compile-cache (1.3 MB, written by
# `pi --version` below). The manifest RUN at the end of this file calls
# `pi --version` again, so deleting it here only relocates those bytes into
# that layer instead of removing them from the image — measured, not assumed:
# today the manifest layer is 128 kB precisely because it finds the cache warm.
purge_build_caches() { \
npm cache clean --force >/dev/null 2>&1 || true; \
rm -rf /root/.npm; \
} && \
if [ "${PI_VERSION}" = "latest" ]; then \
NPM_CONFIG_PREFIX=/usr npm install -g @earendil-works/pi-coding-agent ; \
else \
@@ -209,12 +402,30 @@ RUN set -e && \
git_fetch_ref "${PI_ATELIER_REPO}" "${PI_ATELIER_REF}" /opt/pi-atelier && \
(cd /opt/pi-fork && npm install --omit=dev --no-audit --no-fund) && \
(cd /opt/pi-observational-memory && npm install --omit=dev --no-audit --no-fund) && \
prune_foreign_natives && \
purge_build_caches && \
echo "pi-toolkit at $(cd /opt/pi-toolkit && git rev-parse --short HEAD)" && \
echo "pi-extensions at $(cd /opt/pi-extensions && git rev-parse --short HEAD)" && \
echo "pi-fork at $(cd /opt/pi-fork && git rev-parse --short HEAD)" && \
echo "pi-observational-memory at $(cd /opt/pi-observational-memory && git rev-parse --short HEAD)" && \
echo "pi-atelier at $(cd /opt/pi-atelier && git rev-parse --short HEAD) (${PI_ATELIER_VERSION})"
# ── git: let the unprivileged user read the root-owned /opt clones ──────────
# The clones above (and /opt/mempalace-toolkit from the base, /opt/pi-studio in
# the studio variant) are root-owned; the container runs as `developer`. Since
# git 2.35.2 (CVE-2022-24765) any git command in a repo owned by another user
# fails with "dubious ownership" — so `git -C /opt/pi-extensions rev-parse`
# returns nothing, `install.sh` used to abort on it, and acceptance checks that
# read a baked ref via git silently measured "" (v1.9.4 first-boot run: one
# false FAIL from exactly this). Listing the paths (not `*`) keeps the check
# meaningful for everything else, e.g. the virtiofs-mounted /workspace.
# Entries for paths absent in a variant (pi-studio) are inert.
RUN for d in pi-toolkit pi-extensions pi-fork pi-observational-memory pi-atelier \
mempalace-toolkit pi-studio; do \
git config --system --add safe.directory "/opt/${d}"; \
done && \
git config --system --get-all safe.directory
# ── Image-baked skill refresh: pi-extensions (Option 1 over Option 2) ──
# rootfs ships a VENDORED snapshot of the pi-extensions skill at
# /usr/local/share/pi-devbox/skills/pi-extensions/ (the "floor" — guarantees the
@@ -309,6 +520,26 @@ ARG PI_STUDIO_REF=main
ARG PI_STUDIO_VERSION=none
RUN if [ "${INSTALL_STUDIO}" = "true" ]; then \
set -e; \
# Same prune + cache purge as the main install RUN — see the comments there.
# They have to be redefined because shell functions do not survive across
# layers, and they have to run in THIS layer because pi-studio's npm install
# happens here: deleting in a later RUN would leave the bytes in this layer
# and shrink nothing. pi-studio pulls its own pi-coding-agent copy, so it is
# a third ~274 MB site on top of the two in the non-studio variant — and its
# npm install refills /root/.npm, which the main RUN emptied in ITS layer.
prune_foreign_natives() { \
arch="$(node -p process.arch)"; \
find /usr/lib/node_modules /opt -type d -regex '.*/@esbuild/[^/]+' \
! -name "linux-$arch" -prune -exec rm -rf {} + ; \
find /usr/lib/node_modules /opt -type d -regex '.*/@mariozechner/clipboard-[^/]+' \
! -name "clipboard-linux-$arch-gnu" ! -name "clipboard-linux-$arch-musl" \
-prune -exec rm -rf {} + ; \
echo "native platform dirs kept: $(find /usr/lib/node_modules /opt -type d \( -regex '.*/@esbuild/[^/]+' -o -regex '.*/@mariozechner/clipboard-[^/]+' \) -printf '%f\n' 2>/dev/null | sort | uniq -c | tr '\n' ' ')"; \
}; \
purge_build_caches() { \
npm cache clean --force >/dev/null 2>&1 || true; \
rm -rf /root/.npm; \
}; \
rm -rf /opt/pi-studio && mkdir -p /opt/pi-studio && \
git -C /opt/pi-studio init -q && \
git -C /opt/pi-studio remote add origin "${PI_STUDIO_REPO}" && \
@@ -320,6 +551,8 @@ RUN if [ "${INSTALL_STUDIO}" = "true" ]; then \
done; \
[ "$ok" = "1" ] && \
(cd /opt/pi-studio && npm install --omit=dev --no-audit --no-fund) && \
prune_foreign_natives && \
purge_build_caches && \
echo "pi-studio at $(cd /opt/pi-studio && git rev-parse --short HEAD)"; \
fi
@@ -392,7 +625,7 @@ ARG MEMPALACE_TOOLKIT_REF=main
# no ~67-minute base rebuild. (scripts/check-base-hash.sh scans only
# Dockerfile.base, so no folding into the base hash is required — nor would
# it be correct, since this ARG changes nothing about the base's contents.)
ARG SKILLSET_SNAPSHOT_REF=e9e09d95f92670536a199fc986dfa24d787f18d1
ARG SKILLSET_SNAPSHOT_REF=e9e45f7acdde490c3b5d24ce5f508bff8785c2c7
# Dockerfile.base sets description="pi-devbox — base image (variant-independent)"
# and every variant INHERITS it, so both published images used to advertise
@@ -473,6 +706,44 @@ RUN set -e; \
if [ -d "$_snap_dir" ] && [ -n "$(find "$_snap_dir" -type f -print -quit)" ]; then \
SKILL_SNAP="\"$(tree_sha256 "$_snap_dir")\""; \
fi; \
# ── WHICH pi-extensions skill copy actually shipped ──
# Closes the silent-fallback hole. The refresh step above is guarded by
# `[ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates
# the co-located skill (or a fork pointing at a mirror without it) keeps the
# vendored floor and still succeeds — GREEN, with nothing anywhere recording
# that a snapshot shipped instead of the package copy. Measured 2026-09-10:
# the floor had been stale since 2026-07-30, so that fallback would have
# shipped a six-week-old skill silently. The floor is fresh now and gated by
# the skill-floor CI job, but "the fallback is currently harmless" is not the
# same as "you can tell which copy you got", and only the second survives.
#
# MEASURED, never claimed, per the ground-truth rule above: the branch
# condition is re-derived from the same test the refresh step used, and the
# served bytes are then compared against the clone. A build-arg could not
# express this at all, since the outcome depends on the clone's contents.
# package served bytes == the clone's skill/ (the normal path)
# vendored-floor the clone has no skill/ at this ref (fallback shipped)
# divergent both exist but differ — e.g. the clone ships SKILL.md but
# not evaluate-extension-usage.py, so the served directory is
# a MIX of package and floor. Worth its own value: it is the
# one state neither of the other two names honestly.
# No OCI label mirrors this, deliberately: LABEL cannot take a value computed
# in a RUN, and a label fed from an ARG would be exactly the claim-not-
# measurement this block exists to avoid.
_px_dir=/usr/local/share/pi-devbox/skills/pi-extensions; \
PIEXT_SRC='null'; PIEXT_HASH='null'; \
if [ -d "$_px_dir" ] && [ -n "$(find "$_px_dir" -type f -print -quit)" ]; then \
PIEXT_HASH="\"$(tree_sha256 "$_px_dir")\""; \
if [ -f /opt/pi-extensions/skill/SKILL.md ]; then \
if [ "$(tree_sha256 "$_px_dir")" = "$(tree_sha256 /opt/pi-extensions/skill)" ]; then \
PIEXT_SRC='"package"'; \
else \
PIEXT_SRC='"divergent"'; \
fi; \
else \
PIEXT_SRC='"vendored-floor"'; \
fi; \
fi; \
{ \
echo '{'; \
echo " \"release_tag\": \"${RELEASE_TAG}\","; \
@@ -493,6 +764,8 @@ RUN set -e; \
# vendored skill directory, not one file — see tree_sha256() above.
echo " \"skillset_snapshot_ref\": \"${SKILLSET_SNAPSHOT_REF}\","; \
echo " \"skillset_snapshot_tree_sha256\": ${SKILL_SNAP},"; \
echo " \"pi_extensions_skill_source\": ${PIEXT_SRC},"; \
echo " \"pi_extensions_skill_tree_sha256\": ${PIEXT_HASH},"; \
echo " \"components\": {"; \
echo " \"pi-toolkit\": \"$(rev /opt/pi-toolkit)\","; \
echo " \"pi-extensions\": \"$(rev /opt/pi-extensions)\","; \
+91 -22
View File
@@ -175,12 +175,10 @@ Currently published:
| `joakimp/pi-devbox:latest-studio` | `latest` + [pi-studio](https://github.com/omaclaren/pi-studio) (browser prompt editor, KaTeX/Mermaid preview, tmux-backed literate REPLs) | ~3.25 GB |
| `joakimp/pi-devbox:vX.Y.Z-studio` | pinned-version studio equivalent | ~3.25 GB |
Planned for an upcoming minor release:
- *(shipped in Unreleased/base)* **PDF export from Studio/pandoc** now works:
the base image ships **`typst`** as the PDF engine (`pandoc --pdf-engine=typst`),
a single ~30 MB static binary — no separate `-tex` variant needed.
`texlive-xetex` stays the higher-fidelity fallback (install on demand).
Both variants ship **`typst`** as the pandoc PDF engine
(`pandoc --pdf-engine=typst`), a single ~30 MB static binary, so PDF export from
Studio/pandoc works out of the box — no separate `-tex` variant needed.
`texlive-xetex` stays the higher-fidelity fallback (install on demand).
## Using pi-studio (`-studio` variant)
@@ -360,6 +358,61 @@ DOT syntax errors instead of crashing. Then in Studio: open the PNG (or a
`.md` that embeds it) and hit **refresh-from-disk** after each edit.
Note: SVG is **not** in Studio's local-image-link allowlist — use PNG.
## Terminal UI mode: fullscreen is the default (since v1.10.0)
pi **1.0.0** changed the default terminal UI mode to **fullscreen**, and this
image adopts upstream's default rather than overriding it. Fullscreen draws into
the terminal's alternate screen, so pi's transcript no longer accumulates in your
terminal's native scrollback — you scroll inside pi instead, and on exit pi
prints the transcript (`fullscreenExitOutput`).
That is a real behaviour change if you were used to the old mode, so here is the
way back. Nothing in the image is pinned, so all three routes below work:
| Scope | How |
|---|---|
| **Permanently, for every session** | add `"tuiMode": "regular"` to `~/.pi/agent/settings.json` |
| **One session** | `pi --tui-mode regular` |
| **One project only** | add `"tuiMode": "regular"` to that project's `.pi/settings.json` — project settings override the agent directory |
The settings file is plain JSON and `tuiMode` is top-level, so the minimal
permanent change is:
```bash
# merge the key without disturbing the rest of the file (python3 is always present)
python3 - <<'PY'
import json, pathlib
p = pathlib.Path.home() / ".pi/agent/settings.json"
d = json.loads(p.read_text()) if p.exists() else {}
d["tuiMode"] = "regular" # "fullscreen" is pi's default
p.write_text(json.dumps(d, indent=2) + "\n")
PY
```
If you stay on fullscreen, four related settings are worth knowing — all
documented in pi's own `docs/settings.md` under *Terminal and display*:
- `fullscreenExitOutput` — `"transcript"` (default) or `"resume-hint"`: what pi
leaves behind in the terminal when fullscreen exits.
- `fullscreenScrollbar` — `"auto"` (default), `"always"`, `"hidden"`.
- `fullscreenCopyOnSelect` — `true` by default; selecting text copies it.
- `fullscreenWheelScrollLines` — `"auto"` by default. Relevant **here** in
particular: this container is normally driven over SSH, and over SSH `"auto"`
accelerates fast wheel spins to at most 6 lines per event (local macOS
terminals accelerate on their own, so there it moves one line). `Alt`+wheel
moves five times as far.
Two notes specific to this image:
- **tmux.** Fullscreen uses the alternate screen, so `tmux` copy-mode scrollback
shows the pane's history *around* pi, not pi's transcript. Scroll within pi,
or use `"tuiMode": "regular"` if you rely on tmux copy-mode to search the
conversation.
- **pi-atelier.** The bundled sidebar works in both modes — upstream has handled
regular and fullscreen renderers separately since pi 0.84, including
fullscreen divider dragging and keeping sidebar text out of fullscreen
selection — so switching back to `regular` does not cost you the sidebar.
## Using pi-atelier (TUI sidebar)
`pi-atelier` is bundled in **both** variants (vendored at `/opt/pi-atelier`,
@@ -574,7 +627,13 @@ ChromaDB ONNX embedding model so first-time semantic search is
instant.
The palace data lives at `~/.mempalace/palace` on the host
(bind-mounted into the container). This means:
(bind-mounted into the container). The image pins that layout with
`ENV MEMPALACE_CONFIG_DIR=/home/developer/.mempalace` (`Dockerfile.base`):
mempalace ≥ 3.10.0 would otherwise treat an *empty* `~/.mempalace` — a freshly
mounted volume at first boot — as "no install here" and put a new palace under
`~/.config/mempalace`, outside anything the compose files persist. With the
variable set, first in mempalace's resolution order, the location is a contract
rather than a heuristic. This means:
- A pi running on the host and a pi running inside this container see
the same palace.
@@ -903,8 +962,10 @@ docker inspect --format '{{json .Config.Labels}}' joakimp/pi-devbox:latest | jq
```
`org.opencontainers.image.{version,revision,created}` plus
`se.jordbo.pi-devbox.*-ref` record the intended pi version and companion
refs. The on-disk `/etc/pi-devbox/build-manifest.json` records **ground
`se.jordbo.pi-devbox.*-ref` and `se.jordbo.pi-devbox.*-version` record the
intended pi and mempalace versions and companion refs (`mempalace-version` is
set in `Dockerfile.base` and inherited, so it names the pin of the base the
image actually built on). The on-disk `/etc/pi-devbox/build-manifest.json` records **ground
truth** — the actual checked-out commit of each `/opt` clone, the live
`pi --version`, and (from v1.8.6) the live `mempalace --version` of the
installed palace core — so a tag is reconstructable after CI logs rotate:
@@ -919,16 +980,23 @@ through `jq` yourself:
```console
$ pi-devbox-version
pi-devbox v1.5.0
built: 2026-07-13T17:53:16Z (source d68674d11e06)
pi: 0.80.6
pi-devbox v1.8.14
built: 2026-09-08T21:54:07Z (source 361babd4fd61)
pi: 0.85.1
palace: 3.9.0
components:
pi-toolkit: 9a8f6faeaa08
pi-extensions: 61c98e004e3d
pi-fork: 4a09af4ef527
pi-observational-memory: 27a5195eaf90
mempalace-toolkit: 96699f2a1781
pi-studio: 2ef38ef31cea
pi-toolkit: adfb553f5c8a
pi-extensions: 2610545c83bb
pi-fork: e69725c39603
pi-observational-memory: ce9fc982b3a2
pi-atelier: 734258bbcb62
mempalace-toolkit: e45f6b430181
pi-studio: e04fc7aa3275
skills:
credential-incident-response baked
mempalace live /workspace/skillset @ 4d7c0ea (identical to baked snapshot)
pi-devbox-environment baked
pi-extensions baked
```
It also flags **live drift** — if `pi --version` no longer matches what was
@@ -1093,7 +1161,7 @@ persisted volumes survived, and pi runtime wiring is intact:
```bash
./scripts/recreate-sanity-check.sh # auto-detects variant
./scripts/recreate-sanity-check.sh --expected-image-version 1.8.9 # assert the pi-devbox release tag
./scripts/recreate-sanity-check.sh --expected-version 0.84.4 # assert the pi coding agent version
./scripts/recreate-sanity-check.sh --expected-version 1.0.0 # assert the pi coding agent version
```
Those are **two different versions**, and the flags are not interchangeable:
@@ -1132,9 +1200,10 @@ resolved to `latest` at build time:
| Component | Pin | Where |
|---|---|---|
| pi | `0.84.4` | `ARG PI_VERSION` — `Dockerfile.variant` |
| pi-atelier | `v0.10.0` | `ARG PI_ATELIER_REF` — `Dockerfile.variant` |
| mempalace | `3.8.0` | `ARG MEMPALACE_VERSION` — `Dockerfile.base` |
| pi | `1.0.0` | `ARG PI_VERSION` — `Dockerfile.variant` |
| pi-obsmem | `731c3d49288580f4d79cabdbfbc0d16b34db0f41` | `ARG PI_OBSMEM_REF` — `Dockerfile.variant` |
| pi-atelier | `v0.13.0` | `ARG PI_ATELIER_REF` — `Dockerfile.variant` |
| mempalace | `3.10.0` | `ARG MEMPALACE_VERSION` — `Dockerfile.base` |
The objective is **not** to freeze versions. Bumping is routine — usually one
line plus a changelog note. The objective is that adopting a new upstream
+44 -2
View File
@@ -100,8 +100,21 @@ fi
# existing data. `--yes` auto-accepts detected entities so the init is
# non-interactive.
if command -v mempalace &>/dev/null && [ -d /workspace ]; then
PALACE_DIR="${HOME}/.mempalace"
if [ ! -d "$PALACE_DIR/palace" ]; then
# Read the root from the same variable mempalace itself reads (set as an
# image ENV in Dockerfile.base since mempalace 3.10.0 started resolving
# ~/.config/mempalace for an EMPTY ~/.mempalace). The fallback keeps the
# historical location for anyone running this script with the ENV unset;
# the point of naming the variable here is that this test and mempalace's
# own resolution can no longer disagree about where the palace lives — a
# disagreement that would make this branch fire on every start.
PALACE_DIR="${MEMPALACE_CONFIG_DIR:-${HOME}/.mempalace}"
# Sentinel = config.json, because that is what `mempalace init` writes.
# It does NOT create palace/ — mining does — so the earlier test on palace/
# re-fired on every start of a container that had never mined locally
# (v1.9.4 acceptance: 1 "Initializing" line after one boot, 2 after a
# restart). Harmless (init is idempotent) but the log lied. Populated
# volumes (config.json present) skip either way.
if [ ! -f "$PALACE_DIR/config.json" ]; then
echo "Initializing MemPalace for workspace (non-interactive)..."
# </dev/null: mempalace init has an interactive "Mine this directory
# now? [Y/n]" prompt that --yes does not auto-answer in all paths.
@@ -306,6 +319,35 @@ fi
if [ -f "$HOME/.gitignore_global" ] && ! git config --global core.excludesFile &>/dev/null; then
git config --global core.excludesFile "$HOME/.gitignore_global"
fi
# Route git-over-ssh through the WRITABLE ssh sidecar. ~/.ssh is commonly
# bind-mounted read-only from the host, and a per-host
# ControlPath ~/.ssh/cm/%r@%h:%p
# inherited from that config (the standard CGNAT multiplexing recipe) kills every
# push with
# unix_listener: cannot bind to path ~/.ssh/cm/...: Read-only file system
# hidden behind git's misleading "Please make sure you have the correct access
# rights", which sends the reader hunting for a key problem that does not exist.
# setup-lan-access.sh (run near the top of this script) already wrote
# ~/.ssh-local/config, whose leading `Host *` block overrides ControlPath into
# the writable ~/.ssh-local/cm and only THEN `Include`s the user's own config —
# so -F repairs the socket path while keeping every per-host User/Port/
# IdentityFile. Wiring it here means no caller has to know any of that.
#
# WHY THIS IS NOT LEFT TO DOCUMENTATION: measured 2026-09-22 on tor-ms22, an
# agent with the remedy in its system prompt, in a loaded skill, in 24 palace
# drawers, AND printed verbatim by recreate-sanity-check.sh two hours earlier
# still hit this and reinvented a /tmp/sshcm workaround. The knowledge was
# available four times over, so a fifth copy is not the fix — removing the need
# to know is.
#
# The [ -r ] guard is load-bearing, not decoration: setup-lan-access.sh only
# writes the sidecar on VM-backed hosts (OrbStack / Docker Desktop). On native
# Linux Docker there is none, and pointing -F at a missing file would break EVERY
# git-over-ssh operation instead of fixing one. Respect a value the user already
# set — same first-wins convention as the three settings above.
if [ -r "$HOME/.ssh-local/config" ] && ! git config --global core.sshCommand &>/dev/null; then
git config --global core.sshCommand "ssh -F $HOME/.ssh-local/config"
fi
# ── pi: deploy toolkit + extensions + mempalace bridge ─────────────
# pi is always installed in pi-devbox; no INSTALL_PI guard needed.
Executable
+64
View File
@@ -0,0 +1,64 @@
#!/usr/bin/env bash
# Pre-push gate for pi-devbox: shellcheck every shell script before it leaves
# this clone. Thin wrapper — all logic lives in scripts/lint-shell.sh, which is
# the SAME script the CI release gate runs. One copy, not two: a duplicated
# check that drifts is the failure this repo keeps paying for.
#
# Install per clone: git config core.hooksPath hooks
# Bypass this gate: git push --no-verify (a guard, not a wall)
#
# WHY THIS HOOK EXISTS
# v1.8.14's first release attempt died at scripts/smoke-test.sh:770 after
# build-base had already spent ~46 minutes. shellcheck had ALREADY caught the
# defect — SC2289 at severity error, on the very push that introduced it — and
# the lint job stayed red for 24 hours, unread, across three runs. The fix at
# the time was to gate the release on the same script (the `lint-gate` job).
# This hook is the cheaper end of that: the same finding, before the push,
# in seconds rather than after a 40 s CI gate or a 46 min build.
#
# WHY IT COULD NOT EXIST UNTIL NOW
# Measured on v1.8.14 (2026-09-09): shellcheck was absent from the devbox
# image by all three routes — PATH, dpkg and a filesystem search. So
# lint-shell.sh exited 2 in every container, and a hook calling it would have
# refused EVERY push rather than gating anything. `shellcheck` was added to
# Dockerfile.base in the same change that added this file; on an image built
# before that, enable this hook and you will simply be told the gate cannot
# run. That is the correct behaviour, but it is not a working hook — so do not
# set core.hooksPath on a container older than the release that bakes it.
#
# NOTE ON SCOPE: this lints the WORKING TREE, not the exact commit range being
# pushed. That is deliberate and matches what the CI gate does to the tagged
# tree. It means a defect you have staged-but-not-committed is also reported,
# which is noisy in the safe direction.
set -euo pipefail
HOOK_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
REPO_ROOT="$(cd "$HOOK_DIR/.." && pwd)"
LINTER="$REPO_ROOT/scripts/lint-shell.sh"
tag="[lint-shell]"
# Same rule the gate itself applies, applied one level up: a missing check is
# not a pass. If the script is gone, the push is refused rather than waved
# through on the assumption that CI will catch it.
if [ ! -r "$LINTER" ]; then
echo "$tag refusing the push: $LINTER is missing, so the gate cannot" >&2
echo "$tag run. A gate that cannot run must not pass." >&2
exit 2
fi
# Point the message at the actual remedy when the binary is absent, because the
# linter's own message ("install it or run this in CI") is written for a CI
# runner and is misleading inside a container the developer cannot apt-install
# into persistently.
if ! command -v shellcheck >/dev/null 2>&1; then
echo "$tag refusing the push: shellcheck is not installed, so the gate" >&2
echo "$tag cannot run. A gate that cannot run must not pass." >&2
echo "$tag" >&2
echo "$tag This container predates the image that bakes shellcheck." >&2
echo "$tag Either recreate onto an image that has it, or unset the hook:" >&2
echo "$tag git config --unset core.hooksPath" >&2
echo "$tag To push this once without the gate: git push --no-verify" >&2
exit 2
fi
exec bash "$LINTER" "$REPO_ROOT"
+27 -1
View File
@@ -157,6 +157,11 @@ if [ "$SHOW_SKILLS" = "yes" ] && [ -d "$BAKED_SKILLS" ] && [ -d "$SKILLS_DIR" ];
# is not hypothetical.
snap_ref=$(jq -r '.skillset_snapshot_ref // empty' "$MANIFEST")
snap_sha=$(jq -r '.skillset_snapshot_tree_sha256 // empty' "$MANIFEST")
# Which pi-extensions copy the BUILD baked. Distinct from everything else in
# this section, which reports which copy is being READ at runtime: for
# pi-extensions the baked tree is itself one of two possible copies, and that
# choice was made at build time and is not recoverable by inspection.
px_src=$(jq -r '.pi_extensions_skill_source // empty' "$MANIFEST")
# Same pipeline Dockerfile.variant uses to measure the baked directory at
# build time: relative paths in `find | sort` order, each hashed, the whole
# listing folded into one sha256. Keep the two definitions identical — they
@@ -187,7 +192,28 @@ if [ "$SHOW_SKILLS" = "yes" ] && [ -d "$BAKED_SKILLS" ] && [ -d "$SKILLS_DIR" ];
_target=$(readlink -f "$_link" 2>/dev/null || echo "$_link")
case "$_target" in
"$BAKED_SKILLS"/*|"$BAKED_SKILLS")
printf ' %-22s baked\n' "$_name"
# "baked" alone used to be the whole story. For pi-extensions it is not:
# the baked tree holds EITHER the package copy that Dockerfile.variant
# lays over the snapshot, OR the vendored floor, when the clone had no
# skill/ at that ref. The two are indistinguishable by inspection — same
# path, same filenames, same permissions — so the build records which one
# it used and this reports it. Without this line a six-week-stale
# fallback skill looks exactly like a current one, which is precisely how
# the floor went unnoticed from 2026-07-30 to 2026-09-10.
if [ "$_name" = "pi-extensions" ] && [ -n "$px_src" ]; then
case "$px_src" in
package)
printf ' %-22s baked (package copy)\n' "$_name" ;;
vendored-floor)
printf ' %-22s baked \033[33m(FALLBACK: vendored floor — clone had no skill/)\033[0m\n' "$_name" ;;
divergent)
printf ' %-22s baked \033[33m(MIXED: part package, part floor)\033[0m\n' "$_name" ;;
*)
printf ' %-22s baked\n' "$_name" ;;
esac
else
printf ' %-22s baked\n' "$_name"
fi
continue
;;
esac
@@ -535,7 +535,10 @@ Two consequences worth internalising:
- **"Seen, not doing it" is a legitimate ack** — `status="blocked"` or
`"superseded"` plus the reason. Silence is not, and it is not merely rude:
with no terminal event of yours to join to, the ask stays in the owed set
indefinitely and there is nothing anyone can do about it from the other end.
indefinitely. The *original requester* — and nobody else — can release it from
the other end, but only by saying so explicitly: see **Withdrawing an ask you
sent** below. That is a release by the asker, not an escape for the answerer.
While the ask still stands, only *your* terminal event clears it.
- **Nothing expires, and it should not.** An `open` with no terminal reply is
still live by definition, and the finished threads are valuable history. If
content is genuinely perishable ("do not push to main for the next hour"), say
@@ -570,6 +573,24 @@ Two consequences worth internalising:
be matched to it at all.
- **Corrections are new events, never edits.** Say explicitly what you retract
and name the id — drawer or event — that carried the withdrawn claim.
- **Withdrawing an ask you sent: state it, never imply it.** Your release only
counts when the terminal event (a) comes from the same `from_agent` that sent
the ask, (b) is directed at that recipient exactly — never `*`, so a broadcast
can neither oblige nor release, (c) carries a terminal status (`claimed` and
`ready` are not terminal and do not release anything), (d) is strictly after
the ask, (e) joins it via `ack_of` or the same `correlation_id`, **and (f)
names that ask in `metadata.withdraws` or `metadata.closes`.** Prose in the
body does not count, and neither does a bare terminal event on the
correlation: inferring release from *any* terminal would let your own
bookkeeping silently delete a real obligation, so the release must be stated.
Needs toolkit ≥ `e2b060a` (image ≥ `v1.9.1`) — check with
`grep -c isWithdrawn /opt/mempalace-toolkit/extensions/pi/mempalace.ts` and
read `0` as "my withdrawal will have no effect on their mailbox". Measured
cost of getting it wrong: a `v1.8.13` rollout ask was withdrawn by its sender,
who recorded it as done; the recipient's derivation never saw the release and
still reported the ask owed **41 hours later**, for a release that device
never installed — and the asymmetry was invisible from the sender's side
(RFC 003 §3.3 clause 4).
- **Put a retraction where the reader will look.** A *directed open ask* reaches a
live agent's mailbox; a **terminal-status event reaches no mailbox at all**, and
a *drawer* is what a future semantic search finds. If you filed advice as a
@@ -1,17 +1,17 @@
---
name: pi-extensions
description: >-
Use the pi extensions (pi-fork, pi-observational-memory, ssh-controlmaster) effectively in the pi coding agent harness. Load this skill only when running inside pi (detection - `fork` and `recall` are present in your tool list, or `pi --ssh` was used to start the session). pi-fork dispatches focused subtasks to forked agents at fast/balanced/deep effort tiers; pi-observational-memory compacts long sessions into recallable observations + reflections; ssh-controlmaster rewires pi's read/write/edit/bash tools to execute on a remote host over a multiplexed SSH connection. This skill covers tier selection, task design, boundary discipline, when to use recall, and remote-pi mechanics.
Use the pi extensions (pi-fork, pi-observational-memory, ssh-controlmaster) effectively in the pi coding agent harness. Load this skill only when running inside pi (detection - `fork` and `recall` are present in your tool list, or `pi --ssh` was used to start the session). pi-fork dispatches focused subtasks to forked agents at fast/balanced/deep effort tiers; pi-observational-memory compacts long sessions into recallable observations + reflections; ssh-controlmaster rewires pi's read/write/edit/bash tools to execute on a remote host over a multiplexed SSH connection. Also covers the context ladder L0-L4 and the `task` tool (pi-task: isolated child, immutable spec, machine-checked envelope, write-boundary diff) that is the DEFAULT for delegated work, with `fork` reserved for read-only exploration and parallel opinions, plus the `fork-gate` hook that enforces the split. This skill covers rung and tier selection, task design, boundary discipline, when to use recall, and remote-pi mechanics.
---
# Pi Extensions: pi-fork, pi-observational-memory, ssh-controlmaster
# Pi Extensions: pi-fork + task/fork-gate, pi-observational-memory, ssh-controlmaster
## When to Load This Skill
Load only when **both** of these are true:
1. You are running inside the **pi coding agent harness** (not Claude Code, not opencode, not any other harness).
2. The `fork` and/or `recall` tools appear in your available tool list, **or** the session was started with `pi --ssh ...`.
2. The `fork`, `task` and/or `recall` tools appear in your available tool list, **or** the session was started with `pi --ssh ...`.
If you do not see those tools, this skill does not apply — skip it. Other harnesses do not have these extensions and the patterns below will not work there.
@@ -71,7 +71,78 @@ ssh-controlmaster is orthogonal but composes cleanly: when pi is operating remot
---
## Part 1: pi-fork
## Part 1: delegating work — `task` and `fork`
### Decide the rung BEFORE the brief (read this first)
Two tools run a child agent. They differ in one thing, and it decides the
quality of what comes back: **what the child sees.**
| tool | child sees | rung | gives you | use for |
|---|---|---|---|---|
| **`task`** (pi-extensions `task.ts`, wraps `pi-task`) | **only your spec** — goal, named files, curated facts | L0–L2 | immutable spec, PASS/FAIL envelope, per-root boundary diff, audit dir, budgets | **any delegated work that changes files or must obey a rule** — the default |
| **`fork`** (pi-fork) | **your entire branch**, brief appended last | L4 | prose report, effort tiers, parallel dispatch from one message | read-only exploration that needs this conversation; N independent opinions |
**Pre-flight before any `fork(...)` — one *yes* makes it a `task`:**
1. Does the brief say *do not / only / never / must not*?
2. Will the child write, edit, commit or push anything?
3. Do I want a PASS/FAIL I can check, rather than prose?
Why the text alone did not work (and why this is now enforced): the rule above
lived in this skill and in the global AGENTS.md for months and was still
violated by agents that had just read it — five fork briefs in one session on
2026-09-17, all carrying "do not", one of which returned confident verbatim
quotes that did not exist. `fork` is a *tool*: its self-recommending
description ("implementation, testing, review…") is in the tool list every turn
and survives compaction; this skill is gone after the first compaction, and
`pi-task` was a CLI to be remembered. Two structural fixes shipped 2026-09-19:
- **`task` is a tool** (`extensions/task.ts`), so both rungs sit in the tool
list with the decision rule in their descriptions, and `promptGuidelines`
puts the rule in the system prompt where compaction cannot remove it. It
rejects, before spending a model run, the two spec errors that make a
violation certain (see "roots" below) and serialises overlapping tasks.
- **`fork-gate`** (`extensions/fork-gate.ts`) is a `tool_call` hook that BLOCKS
a fork whose brief contains a prohibition, a write boundary or a clause-initial
file-changing imperative, and returns the `task(...)` call to make instead.
It matches wording, not intent: a genuinely read-only exploration brief that
trips it is rephrased, and one that cannot be rephrased without its
prohibition needed `task` all along. `PI_FORK_GATE=off` logs instead of
blocking; `/ext` disables it entirely.
Minimal call:
```
task(id="slug", goal="…verbatim; the child has NO other context…",
deliverable="…exact shape wanted…", effort="fast|balanced|deep",
read_only=false,
roots=["/abs/repo/docs", "/abs/repo/src"], # WATCHED, each diffed alone
write_allowed=["/abs/repo/docs"], # exact subset of roots
facts=["verified fact"], files=["/abs/path/to/read"])
```
**Roots — the two errors the tool refuses up front.** Every root is diffed on
its own and a delta is allowed only if that *exact root string* is in
`write_allowed`. So (a) `write_allowed` must be a subset of `roots`, not a
subdirectory of one, and (b) a writable root must not lie inside a watched-only
root — the parent's porcelain would change and register a violation every time.
List the writable part as its own root and leave the enclosing repo out. This is
the shape the 2026-09-17 migration tasks used (sibling roots, `write_allowed`
naming four of them) and it passed cleanly.
**Overlap.** Sibling `task` calls whose roots overlap would see each other's
writes as violations; the tool runs them one after another automatically. Do
not rely on that for ordering *semantics* — if B needs A's output, call B after
A returns.
**What isolation does not fix.** L0 removes the *narrative* failures (parent
voice, invented continuity, ignored prohibitions). It does not remove
confabulation: an under-specified spec still gets a confident deliverable. The
report prints the evidence pointers under a "SPOT-CHECK THESE" heading for a
reason.
Everything below about tiers, brief design and boundary discipline applies to
**both** tools — a `task` spec is a brief too.
### Effort tier mapping
@@ -85,15 +156,15 @@ Configured in `~/.pi/agent/settings.json` under `pi-fork.effortProfiles`. The co
**Rule of thumb:** start at `balanced` unless you have a specific reason to go up or down. Going too cheap on a deep task wastes a fork; going too expensive on a mechanical task is just slow.
### When to fork vs. do it yourself
### When to delegate vs. do it yourself
Fork when **any** of:
Having chosen the rung above, delegate (either tool) when **any** of:
- The task requires reading many files whose contents you don't need to keep in your main context afterwards (the fork returns a dense summary; raw file contents stay in the fork's context and are discarded).
- You want to run multiple analyses in **parallel** (especially: comparing N options, where independent reasoning is itself a signal — see "parallel forks" below).
- The task is well-scoped enough to specify completely up front and well-bounded enough that returning a dense report is more useful than continuing the dialogue.
- You are about to do something that would burn a lot of tokens on tool calls (long file reads, many bash invocations) whose output you will mostly discard.
Don't fork when:
Don't delegate when:
- The work fits in your current context budget without crowding out what comes next.
- The task is exploratory and you'll need to iterate based on what you find (forking turns iteration into round-trips with full task-spec rewrites).
- You need to make decisions during the work that depend on context only the main thread has.
@@ -161,6 +232,69 @@ The "three" things it completed were exactly the main thread's pending todos, vi
- Distrust **quantities** and **provenance claims** in fork prose specifically ("all N sessions", "shipped with the image", "as expected") — those are the slots confabulation fills.
- The fact that the fork was "right anyway" is not the same as the fork having followed instructions.
### The context ladder — and the second dispatch mechanism (`pi-task`)
Everything above describes a child that inherits everything. That is not a fixed
cost of delegation — **how much context a child gets is a choice**, and `fork`
sits at one extreme of it. Five rungs:
| rung | what the child sees | mechanism | built? |
|---|---|---|---|
| **L0** | nothing but the goal | `pi-task` default: fresh `--session-id pitask-<id>-<stamp>` in a private `--session-dir` | yes |
| **L1** | goal + **names** of files/commands to read itself | `pi-task` spec `context.files` / `context.commands` (`bin/pi-task:154,157`) | yes |
| **L2** | goal + an **excerpt the parent curated** | `pi-task` spec `context.facts`, pasted verbatim (`bin/pi-task:151`) | yes |
| **L3** | a **truncated tail** of the parent branch | *nothing implements this* — would need a new spec key plus `--session <trimmed snapshot>` | **no** |
| **L4** | the **entire** parent branch | `fork(task=…)` — `getHeader()+getBranch()`, no offset or limit anywhere in the call chain | yes |
**`pi-task` is a CLI (`/opt/pi-toolkit/bin/pi-task`, `schema` prints the spec
fields); the `task` tool from pi-extensions wraps it** so it appears in your tool
list next to `fork`. If the tool is absent, invoke the CLI with `bash`:
`/opt/pi-toolkit/bin/pi-task run <spec.json>`. Either way it reads an immutable
JSON spec, and "inherit the session" is not expressible in that schema — the
isolation is structural, not a request.
**Choose the lowest rung that can do the job:**
- **`fork` (L4)** when the subtask only makes sense against this conversation,
when you want several independent opinions in parallel from one message, or for
read-only exploration whose detail you will discard. Everything in "Boundary
discipline" above applies in full.
- **`pi-task` (L0–L2)** when the brief contains a **prohibition** (the inherited
transcript is exactly what overrides those), when you want a **pass/fail**
result instead of prose, when you need an **audit trail**, or when writes
outside an authorised set must be caught.
- **Neither** for trivial work, iterative work (both are one-shot), or judgement
that needs context only you have.
**What `pi-task` gets you that no brief can.** The envelope must parse or the run
FAILED, however fluent the prose. `roots[]` is the WATCHED set and
`write_allowed` the CHANGEABLE subset, diffed before and after with git
`--porcelain --ignored`. That `--ignored` flag is load-bearing: in the T4 test the
child obeyed its brief perfectly and still tripped the detector, because
`py_compile` wrote `__pycache__` into a watched-but-not-writable root — a
gitignored path that plain `--porcelain` reports as clean. Note the structural
point that test exposed: under `read_only: true` a write is *defiance*, so a
well-behaved child never produces a delta and the detector is never exercised.
Splitting WATCHED from WRITABLE is what lets an **obedient** child reveal a
violation, which is the realistic hazard.
**What it does not fix.** `--no-extensions` removes extensions, not the core
`read`/`write`/`edit`/`bash` tools — exactly as described above — so the boundary
diff is post-hoc **detection, not prevention**. And a fresh L0 context removes the
*narrative* failures (parent voice, invented continuity) without removing
confabulation: given an under-specified spec built on a false premise, the child
still filled the `deliverable` slot with a confident shape. The envelope's own
structure creates that pressure. Verify decisive claims from the filesystem
regardless of which rung you used.
**Trap — the capability floor is inverted from intuition.** `runner.ts:188` reads
`if (extensions !== null) args.push("--no-extensions")`. So `pi-fork.extensions:
[]` passes the flag and the floor is **on**; setting it to `null` — documented in
`settings.json` as the way to "restore normal extension loading" — passes nothing
and the floor is **off**, restoring palace writes inside every fork child.
Changing `[]` to `null` as a tidy-up re-arms what was deliberately disarmed.
`pi-task` hardcodes the flag and cannot drift this way.
### Anti-patterns
- **Forking trivial work.** A fork has overhead. If the task takes < 30 seconds in your main thread, just do it.
@@ -230,7 +364,17 @@ When entries conflict, **the most recent observation reflects the latest known s
## Quick Reference
```
fork(task=..., effort=fast|balanced|deep)
task(id, goal, deliverable, effort, read_only, roots, write_allowed, facts, files, commands, wall_s, usd)
- L0-L2: isolated child sees ONLY the spec — DEFAULT for work that writes or has rules
- roots[] = WATCHED (each diffed alone); write_allowed[] = exact subset of roots,
never nested inside a watched-only root (the tool rejects both errors up front)
- envelope must parse or the run FAILED; spot-check evidence pointers
- overlapping-root tasks are serialised; audit: ~/.pi/agent/pi-task/<stamp>-<id>/result.json
- CLI fallback: bash /opt/pi-toolkit/bin/pi-task run <spec.json> (schema | selftest | run --dry-run)
fork(task=..., effort=fast|balanced|deep) # L4: child inherits your WHOLE branch
- ONLY for read-only exploration needing this conversation, or N parallel opinions
- fork-gate BLOCKS briefs with do-not/only/never, write boundaries, or "edit/commit/fix …"
- state decision authority explicitly
- pass verified context up front
- specify deliverable shape
@@ -248,8 +392,9 @@ recall(id=<12-char-hex>)
```
~/.pi/agent/settings.json
pi-fork.effortProfiles — model + thinking-depth per tier
pi-fork.effortProfiles — model + thinking-depth per tier (used by BOTH fork and task)
pi-fork.defaultEffort — usually "balanced"
env PI_FORK_GATE=off — fork-gate logs instead of blocking (default: block)
observational-memory.* — token thresholds, model, agentMaxTurns
observational-memory.debugLog: true — opt-in NDJSON telemetry at
~/.pi/agent/observational-memory/debug/<session>.ndjson (off by default)
+704
View File
@@ -0,0 +1,704 @@
#!/usr/bin/env bash
# check-doc-drift.sh — fail when a hand-maintained doc claim contradicts the
# build files it describes.
#
# THE DEFECT CLASS THIS EXISTS TO CATCH, measured 2026-09-10 while preparing
# v1.9.0. Five separate claims had rotted, all of them the same shape: a fact
# written once by hand, in a file nothing verifies, about a value that lives
# somewhere else and moved.
#
# 1..3. README.md's "Version pins" table was wrong on EVERY row — pi `0.84.4`
# vs ARG PI_VERSION=0.85.1, pi-atelier `v0.10.0` vs v0.10.1, mempalace
# `3.8.0` vs 3.9.0. That table is the worst possible place for this: it
# exists precisely to be the reviewable record of what is deliberately
# frozen, so when it lies, the review it enables is worthless.
# 4. README.md carried a "Planned for an upcoming minor release" section
# listing typst PDF export, which had ALREADY SHIPPED, tagged with a
# self-contradicting "(shipped in Unreleased/base)" marker. The
# CHANGELOG had already documented three earlier instances of exactly
# this stale-"Unreleased"-pointer class (see its v1.8.7 notes).
# 5. DOCKER_HUB.md claimed "Node.js v22" while this release ships Node 24.
# This one is the reason the gate exists at all: DOCKER_HUB.md is
# PUBLISHED. `update-description` in docker-publish.yml POSTs it to Hub
# as full_description on every tag, so unlike README.md — which no
# workflow or gate reads — a stale claim here is what users see.
#
# WHY A GATE AND NOT "REMEMBER TO CHECK". DOCKER_HUB.md had gone eight releases
# (v1.8.6 → v1.9.0) without a touch. Nothing generates it and nothing verifies
# it; the only mechanism keeping it true was whoever remembered. That is the
# same failure mode check-skill-floor.sh was written for, and the same fix:
# convert "someone remembers" into "CI refuses".
#
# TWO CLASSES OF CHECK, DELIBERATELY. Checks 1-7 compare a doc string to a
# value that EXISTS IN THIS REPO, so they can never be wrong about the world and
# need no network, no token, and no built image. Checks 8-9 compare against what
# is PUBLISHED (Docker Hub's measured sizes; the ref labels baked into the last
# released image), because those claims have no in-repo anchor at all and had
# rotted for exactly that reason. They need the network and therefore SKIP,
# loudly and counted, when it is absent -- a skip is neither OK nor a failure,
# because printing an unverified claim as OK is the habit this file exists to
# break, while failing on a third party's uptime would make every release
# hostage to it. Claims that need a RUNNING CONTAINER (the "N mempalace_* tools"
# count, uncompressed on-disk sizes) are still not gated here; assert them in
# scripts/smoke-test.sh where a real image is available.
#
# DELIBERATELY NOT GATED: Dockerfile.base's `# BASE_REBUILD_DATE:` comment, which
# is also stale (2026-07-13, three base rebuilds ago). base_tag is a hash of
# Dockerfile.base's CONTENT plus rootfs/, comments included, so a gate that
# demanded that comment be current would force a ~60 min base rebuild on any
# release that touched no base files at all. Fix it when you are already
# rebuilding the base — then it is free. This is a real cost asymmetry, not
# laziness.
#
# EXIT CODES (same contract as lint-shell.sh and check-skill-floor.sh):
# 0 every checked claim matches
# 1 at least one claim has drifted
# 2 cannot run (a file or ARG this gate reads is missing/unparseable)
# A gate that cannot run must not pass, so a missing input is 2, never 0.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
cd "$REPO_ROOT"
README="README.md"
HUB="DOCKER_HUB.md"
DF_VARIANT="Dockerfile.variant"
DF_BASE="Dockerfile.base"
# Docker Hub rejects a full_description longer than this. docker-publish.yml has
# no size check of its own; it only notices via a non-200 from the API, i.e.
# after paying the whole build. Catching it here makes it a 2-second failure.
HUB_MAX_CHARS=25000
WARN_ONLY=0
FAILURES=0
SKIPS=0
# Tolerance for the published size claims (check 8), as a percentage OF THE
# MEASURED SIZE. The denominator matters: against the claim instead, the same
# drift reads as a different number, and an early draft of this gate took 20%
# from the claim-relative figure and would therefore have MISSED its own
# motivating case. Both bounds are measured, not guessed:
# - the rot that motivated this check: claimed 1.1 GB vs measured 1.37 GB
# = 19.7% off, so the threshold must sit BELOW that or the gate is theatre.
# - the largest legitimate skew, i.e. a claim describing the currently-published
# release while the next tag changes the size: v1.9.1's 1.37 GB against
# v1.9.2's measured 1.23 GB = 11.4% off, so the threshold must sit ABOVE that
# or every size-changing release trips it.
# 15% sits in that 11.4%-19.7% window. Widen it only with a measured reason, and
# re-derive both bounds if you do.
SIZE_TOLERANCE_PCT="${SIZE_TOLERANCE_PCT:-15}"
usage() {
cat <<'EOF'
Usage: check-doc-drift.sh [--warn-only] [-h|--help]
Compares hand-written claims in README.md and DOCKER_HUB.md against the build
files they describe (Dockerfile.base, Dockerfile.variant).
--warn-only Report drift but exit 0 (advisory use, e.g. a local pre-push hook).
Environment:
SKIP_SIZE_CHECK=1 skip check 8 (published size claims vs Docker Hub)
SKIP_REF_CHECK=1 skip check 9 (refs moved since the last release are named)
SIZE_TOLERANCE_PCT check 8 tolerance, default 15 (see comment for its bounds)
Exit: 0 = in sync, 1 = drift, 2 = cannot run.
EOF
}
while [ $# -gt 0 ]; do
case "$1" in
--warn-only) WARN_ONLY=1; shift ;;
-h|--help) usage; exit 0 ;;
*) echo "::error::unknown argument: $1" >&2; usage >&2; exit 2 ;;
esac
done
for f in "$README" "$HUB" "$DF_VARIANT" "$DF_BASE"; do
if [ ! -f "$f" ]; then
echo "::error::$f not found (cwd $PWD). Cannot evaluate doc drift, so this is exit 2, not a pass."
exit 2
fi
done
# Read `ARG NAME=value` from a Dockerfile. Exit 2 when absent: if the ARG this
# gate is built around has been renamed, the gate is measuring nothing and must
# say so rather than silently comparing against an empty string.
read_arg() {
local file="$1" name="$2" value
value="$(sed -n "s/^ARG ${name}=\\(.*\\)\$/\\1/p" "$file" | head -1)"
if [ -z "$value" ]; then
echo "::error::ARG ${name} not found in ${file}. It was probably renamed;" >&2
echo "::error::update check-doc-drift.sh to match, because this gate is now blind." >&2
exit 2
fi
printf '%s' "$value"
}
# One row of README's "Version pins" table: `| pi | `0.85.1` | ... |`
read_pin_row() {
sed -n "s/^| $1 | \`\\([^\`]*\`*\\)\` |.*/\\1/p" "$README" | head -1
}
fail() {
FAILURES=$((FAILURES + 1))
echo "::error::$1"
}
ok() { printf ' OK %s\n' "$1"; }
# A check that could not be EVALUATED, as distinct from one that passed.
# Deliberately neither ok() nor fail(): printing it as OK would launder an
# unmeasured claim into a passing one (the exact habit this file exists to
# break), while failing on a third party's uptime would make every release
# hostage to Docker Hub's API. Loud, counted, and surfaced in the summary.
skip() { SKIPS=$((SKIPS + 1)); printf ' SKIP %s\n' "$1"; }
echo "Checking hand-maintained doc claims against the build files they describe."
echo
# ---------------------------------------------------------------------------
# 1-3. README's version-pin table vs the ARGs it names by name.
# ---------------------------------------------------------------------------
check_pin() {
local label="$1" documented="$2" actual="$3" where="$4"
if [ -z "$documented" ]; then
fail "README.md: no '| $label |' row found in the version-pin table. Either the
table was restructured (update this gate) or the row was dropped (restore it)."
return
fi
if [ "$documented" != "$actual" ]; then
fail "README.md version-pin table is stale for $label: says '$documented',
$where says '$actual'. Fix the table — it is the reviewable record of what
this repo deliberately freezes, so a wrong row defeats its only purpose."
return
fi
ok "README pin $label = $actual"
}
PI_ACTUAL="$(read_arg "$DF_VARIANT" PI_VERSION)"
ATELIER_ACTUAL="$(read_arg "$DF_VARIANT" PI_ATELIER_REF)"
MEMPALACE_ACTUAL="$(read_arg "$DF_BASE" MEMPALACE_VERSION)"
# pi-obsmem became a PIN in v1.9.5 (was the floating `master`), so it joins the
# reviewable table. It is also covered by the ref-move check below, but that one
# can only ever report "unchanged" for a pinned SHA -- it answers "did upstream
# move?", never "does the table still say what we bake?", which is this check.
OBSMEM_PIN_ACTUAL="$(read_arg "$DF_VARIANT" PI_OBSMEM_REF)"
check_pin pi "$(read_pin_row pi)" "$PI_ACTUAL" "ARG PI_VERSION in $DF_VARIANT"
check_pin pi-obsmem "$(read_pin_row pi-obsmem)" "$OBSMEM_PIN_ACTUAL" "ARG PI_OBSMEM_REF in $DF_VARIANT"
check_pin pi-atelier "$(read_pin_row pi-atelier)" "$ATELIER_ACTUAL" "ARG PI_ATELIER_REF in $DF_VARIANT"
check_pin mempalace "$(read_pin_row mempalace)" "$MEMPALACE_ACTUAL" "ARG MEMPALACE_VERSION in $DF_BASE"
# ---------------------------------------------------------------------------
# 4. DOCKER_HUB.md's Node claim vs ARG NODE_VERSION. This is the published page,
# so it is the one whose staleness reaches users.
# ---------------------------------------------------------------------------
NODE_ACTUAL="$(read_arg "$DF_BASE" NODE_VERSION)"
NODE_DOCUMENTED="$(sed -n 's/.*\*\*Node\.js\*\* v\([0-9][0-9]*\).*/\1/p' "$HUB" | head -1)"
if [ -z "$NODE_DOCUMENTED" ]; then
fail "$HUB: could not find a '**Node.js** vNN' claim. If the wording changed,
update this gate; do not leave the published page unverified."
elif [ "$NODE_DOCUMENTED" != "$NODE_ACTUAL" ]; then
fail "$HUB claims Node v$NODE_DOCUMENTED but ARG NODE_VERSION=$NODE_ACTUAL.
This file is PUBLISHED to Docker Hub by update-description on every tag,
and it is read from the TAG — so fix it before tagging, not after."
else
ok "$HUB Node claim = v$NODE_ACTUAL"
fi
# ---------------------------------------------------------------------------
# 5. Placeholders CI will not substitute. docker-publish.yml substitutes exactly
# {{PI_VERSION}} and then greps for leftovers of that ONE token, so any other
# {{...}} sails through the guard and is published literally.
# ---------------------------------------------------------------------------
UNKNOWN_PLACEHOLDERS="$(grep -o '{{[A-Za-z0-9_]*}}' "$HUB" | sort -u | grep -v '^{{PI_VERSION}}$' || true)"
if [ -n "$UNKNOWN_PLACEHOLDERS" ]; then
fail "$HUB contains placeholders CI does not substitute, which would be
published verbatim: $(echo "$UNKNOWN_PLACEHOLDERS" | tr '\n' ' ')
docker-publish.yml only fills {{PI_VERSION}}; add substitution there first."
else
ok "$HUB has no placeholders beyond {{PI_VERSION}}"
fi
# Match only the UPPER_SNAKE placeholder convention CI uses. A bare '{{' search
# is WRONG here, and the first version of this check proved it by failing on
# README.md:900 — `docker inspect --format '{{json .Config.Labels}}'`, a Go
# template in a legitimate example, not a placeholder. The gate was wrong, not
# the doc. Keep this anchored to [A-Z] so Go/Jinja/Handlebars examples pass.
README_PLACEHOLDERS="$(grep -o '{{[A-Z][A-Z0-9_]*}}' "$README" | sort -u || true)"
if [ -n "$README_PLACEHOLDERS" ]; then
fail "$README contains placeholder(s) nothing substitutes, so they would render
literally for every reader: $(echo "$README_PLACEHOLDERS" | tr '\n' ' ')
Only DOCKER_HUB.md gets substitution, and only for {{PI_VERSION}}."
else
ok "$README has no unsubstituted placeholders"
fi
# ---------------------------------------------------------------------------
# 6. Hub full_description length.
# ---------------------------------------------------------------------------
HUB_CHARS="$(wc -c < "$HUB" | tr -d ' ')"
if [ "$HUB_CHARS" -gt "$HUB_MAX_CHARS" ]; then
fail "$HUB is $HUB_CHARS chars, over Docker Hub's $HUB_MAX_CHARS-char
full_description limit. update-description would fail with a non-200 AFTER
the full build. Trim it — this file is the essentials-only page, and
README.md is the long form on purpose."
else
ok "$HUB is $HUB_CHARS chars (limit $HUB_MAX_CHARS)"
fi
# ---------------------------------------------------------------------------
# 7. Stale "Unreleased" pointers. "Unreleased" is a CHANGELOG-only concept; in
# a user-facing doc it is always a pointer that outlived what it pointed at.
# This class has now bitten five times, hence a gate rather than vigilance.
# ---------------------------------------------------------------------------
STALE_MARKERS="$(grep -n 'Unreleased' "$README" "$HUB" || true)"
if [ -n "$STALE_MARKERS" ]; then
fail "'Unreleased' appears in a user-facing doc, which is always a stale
pointer once the thing ships (it has happened five times here):
${STALE_MARKERS//$'\n'/$'\n' }
State the fact directly, or move it to CHANGELOG.md where 'Unreleased' means something."
else
ok "no stale 'Unreleased' pointers in $README or $HUB"
fi
# ---------------------------------------------------------------------------
# 8. Published size claims vs Docker Hub's MEASURED full_size.
#
# Why this exists: every other claim in these docs is checked against a file
# in this repo, so it cannot rot without someone editing the thing it
# describes. The size claims had no such anchor -- nothing in the repo states
# the image size -- so they quietly went 24% wrong across eight releases
# (DOCKER_HUB.md said ~1.1 GB; :latest measured 1.37 GB on 2026-09-14).
# DOCKER_HUB.md is POSTed to Docker Hub by update-description, so that number
# is the first thing a stranger reads about this image.
#
# Hub's `full_size` tracks the FIRST manifest entry (amd64 here), NOT the sum
# across architectures -- measured: v1.9.2 full_size=1.228 GB, amd64=1.228,
# arm64=1.211, sum=2.439. That matches the table's per-arch "Size
# (compressed)" column, which is why full_size is the right field.
#
# NOT COVERED, deliberately: README.md's ~3.2 GB figures are UNCOMPRESSED
# on-disk sizes, and the registry API exposes compressed sizes only (layer
# sizes in a manifest are compressed; the config blob carries no uncompressed
# totals). Measuring them needs a real pull, so they are out of scope here --
# do not read a green check 8 as covering them.
# ---------------------------------------------------------------------------
# Shared by checks 8 and 9: which Hub repo, and its tag list (one request).
# Derive the repo from the doc's own rows rather than hardcoding it, so a
# rename cannot leave these checks silently probing a repo nobody publishes to.
# shellcheck disable=SC2016 # single quotes are deliberate: this is a sed
# script, and its \( \) groups and \1 backreference must reach sed unexpanded.
HUB_REPO_PATH="$(sed -n 's/^| `\([^:`]*\):[^`]*`.*/\1/p' "$HUB" | head -1)"
HUB_TAGS_JSON=""
HAVE_NET_TOOLS=0
if command -v curl >/dev/null 2>&1 && command -v python3 >/dev/null 2>&1; then
HAVE_NET_TOOLS=1
if [ -n "$HUB_REPO_PATH" ] && \
{ [ "${SKIP_SIZE_CHECK:-0}" != "1" ] || [ "${SKIP_REF_CHECK:-0}" != "1" ]; }; then
HUB_TAGS_JSON="$(curl -sS -m 20 \
"https://hub.docker.com/v2/repositories/${HUB_REPO_PATH}/tags/?page_size=100" \
2>/dev/null || true)"
fi
fi
if [ "${SKIP_SIZE_CHECK:-0}" = "1" ]; then
skip "size claims -- SKIP_SIZE_CHECK=1 was set"
elif [ "$HAVE_NET_TOOLS" = 0 ]; then
skip "size claims -- need both curl and python3 to measure them"
else
if [ -z "$HUB_REPO_PATH" ]; then
skip "size claims -- found no \`repo:tag\` image rows in $HUB to check"
else
if [ -z "$HUB_TAGS_JSON" ]; then
skip "size claims -- Docker Hub API unreachable (offline?); NOT verified"
else
SIZE_RC=0
# NO `|| true` on the python invocation: an early draft had one, and it
# swallowed the exit code so a printed DRIFT line still exited 0 -- a gate
# that reports the defect and passes anyway. The outer `|| SIZE_RC=$?` is
# what keeps `set -e` happy while preserving the code.
SIZE_OUT="$(HUB_MD="$HUB" HUB_JSON="$HUB_TAGS_JSON" TOL="$SIZE_TOLERANCE_PCT" \
python3 <<'PYEOF'
import json, os, re, sys
try:
data = json.loads(os.environ["HUB_JSON"])
except (ValueError, KeyError) as exc:
print(" SKIP size claims -- Hub API returned unparseable JSON (%s)" % exc)
sys.exit(3)
# full_size == first manifest entry (amd64), which is the per-arch number the
# table's "Size (compressed)" column claims. Verified against .images[] sizes.
sizes = {
r["name"]: r["full_size"] / 1e9
for r in data.get("results", [])
if isinstance(r.get("full_size"), int) and r.get("name")
}
if not sizes:
print(" SKIP size claims -- Hub API returned no usable tags")
sys.exit(3)
tol = float(os.environ["TOL"])
row = re.compile(r"^\|\s*`([^`:]+):([^`]+)`\s*\|[^|]*\|\s*~?([0-9]+(?:\.[0-9]+)?)\s*GB\s*\|")
checked = drift = 0
with open(os.environ["HUB_MD"], encoding="utf-8") as fh:
for line in fh:
m = row.match(line)
if not m:
continue # rows saying "same", and every non-image row
_repo, tag, claimed = m.group(1), m.group(2), float(m.group(3))
if "X.Y.Z" in tag:
continue # placeholder row; the concrete tag is checked instead
# base-<hash> is content-addressed and immutable, so its size is
# base-latest's by construction -- probe the alias that always exists.
probe = "base-latest" if tag.startswith("base-") else tag
actual = sizes.get(probe)
if actual is None:
print(" SKIP size %s -- tag '%s' not present on Hub" % (tag, probe))
continue
checked += 1
off = abs(claimed - actual) / actual * 100
if off <= tol:
print(" OK size %s claims ~%.2f GB, Hub measures %.2f GB (%.0f%% off)"
% (tag, claimed, actual, off))
else:
drift += 1
print(" DRIFT size %s claims ~%.2f GB but Hub measures %.2f GB"
" (%.0f%% off, tolerance %.0f%%)" % (tag, claimed, actual, off, tol))
if checked == 0:
print(" SKIP size claims -- no checkable rows resolved to a published tag")
sys.exit(3)
sys.exit(1 if drift else 0)
PYEOF
)" || SIZE_RC=$?
printf '%s\n' "$SIZE_OUT"
case "$SIZE_RC" in
0) : ;;
3) SKIPS=$((SKIPS + 1)) ;;
*)
fail "a published size claim in $HUB has drifted from what Docker Hub
actually serves (see DRIFT above). This page is POSTed to Docker Hub by
update-description, so it is the first size a stranger sees. Re-measure and
update the table:
curl -sS 'https://hub.docker.com/v2/repositories/${HUB_REPO_PATH}/tags/?page_size=100' |
jq -r '.results[] | \"\\(.name) \\(.full_size/1e9)\"'"
;;
esac
fi
fi
fi
# ---------------------------------------------------------------------------
# 9. Everything the NEXT build would bake differently from the LAST PUBLISHED
# release must be named in the CHANGELOG text above that release's heading.
#
# Why this exists, measured 2026-09-19: pi-extensions 25c1265 (a new `task`
# tool and a hook that blocks certain `fork` calls -- a change to how every
# agent in the container delegates work) and mempalace-toolkit 817b3a8 (the
# feed's mine deadline had never reached the transport) both reached this
# image through floating `*_REF=main` ARGs. Neither produced a diff in this
# repo, so nothing here asked for a CHANGELOG entry, and neither had one
# until a reader asked. This is the same shape as check 8: a fact with no
# in-repo anchor rots. The hand practice that existed for it -- the
# "Dependency audit" table in each release's notes ("Baked in vN | Upstream
# now") -- is precisely a "someone remembers" mechanism, and it had lapsed.
#
# How it measures, with no docker/crane/token: the last published `vX.Y.Z`
# is the highest such tag in Hub's tag list (shared with check 8); its
# amd64 config blob is read through the anonymous registry API (token ->
# manifest index -> per-arch manifest -> config) and carries one
# `se.jordbo.pi-devbox.<name>-ref` label per component, each holding the
# SHA that build-args actually baked (resolve-versions in docker-publish.yml
# turns every ref into a SHA before `docker build`). "What the next build
# would bake" is resolved the way that job does it: a 40-hex ARG is itself,
# a tag or branch is `git ls-remote`d (peeled `^{}` first -- an annotated
# tag's un-dereferenced SHA is the tag object, a false alarm this repo has
# already fallen for once), pi-studio is the highest semver tag, and
# `PI_VERSION` / `MEMPALACE_VERSION` are compared as literals against the
# `pi-version` / `mempalace-version` labels (the latter set in Dockerfile.base
# and inherited; absent on releases before it shipped, which reports SKIP).
#
# The rule: baked == would-bake is OK with no mention required. If they
# differ, the text ABOVE the last published version's `## ` heading -- i.e.
# `## Unreleased` plus any not-yet-published `## vX.Y.Z` section, which is
# what the release commit turns Unreleased into -- must contain the
# would-bake value's 7-char SHA prefix (or, for pi-studio, the tag name; for
# pi, the version string). Naming the SHA, not just the repo, is the point:
# it is what the audit table always recorded, and it makes the failure
# message's compare URL a copy-paste away from knowing what moved.
#
# Every upstream commit therefore re-reds this gate until the CHANGELOG
# names the new head. That is the intended cost: the thing that gets baked
# is the thing that gets named, and a typo-fix upstream costs one edited
# SHA here. Read from the TAG like everything else in these docs -- the
# release commit renames Unreleased, so the pending text still covers it.
#
# SKIPs, each counted: SKIP_REF_CHECK=1; no curl/python3; Hub unreachable;
# the release's labels unreadable; one component's upstream unreachable
# (that component only). A published tag whose heading is MISSING from the
# CHANGELOG is a failure, not a skip: that is drift in its own right.
# ---------------------------------------------------------------------------
if [ "${SKIP_REF_CHECK:-0}" = "1" ]; then
skip "ref moves -- SKIP_REF_CHECK=1 was set"
elif [ "$HAVE_NET_TOOLS" = 0 ]; then
skip "ref moves -- need both curl and python3 to read the published labels"
elif ! command -v git >/dev/null 2>&1; then
skip "ref moves -- need git (ls-remote) to resolve what the next build would bake"
elif [ -z "$HUB_REPO_PATH" ]; then
skip "ref moves -- found no \`repo:tag\` image rows in $HUB to locate the published image"
elif [ -z "$HUB_TAGS_JSON" ]; then
skip "ref moves -- Docker Hub API unreachable (offline?); NOT verified"
else
# One plain top-level assignment per ARG, on purpose: read_arg exits 2 on a
# missing ARG, and under `set -e` that only propagates from a bare
# `VAR="$(...)"`. Nested inside a heredoc's $(...) the exit would be swallowed
# by `cat`, and a renamed ARG would leave this check comparing a label against
# an empty string and reporting the component "unchanged".
TOOLKIT_REPO="$(read_arg "$DF_VARIANT" PI_TOOLKIT_REPO)"; TOOLKIT_REF="$(read_arg "$DF_VARIANT" PI_TOOLKIT_REF)"
EXTENSIONS_REPO="$(read_arg "$DF_VARIANT" PI_EXTENSIONS_REPO)"; EXTENSIONS_REF="$(read_arg "$DF_VARIANT" PI_EXTENSIONS_REF)"
FORK_REPO="$(read_arg "$DF_VARIANT" PI_FORK_REPO)"; FORK_REF="$(read_arg "$DF_VARIANT" PI_FORK_REF)"
OBSMEM_REPO="$(read_arg "$DF_VARIANT" PI_OBSMEM_REPO)"; OBSMEM_REF="$(read_arg "$DF_VARIANT" PI_OBSMEM_REF)"
ATELIER_REPO="$(read_arg "$DF_VARIANT" PI_ATELIER_REPO)"
MPTK_REPO="$(read_arg "$DF_BASE" MEMPALACE_TOOLKIT_REPO)"; MPTK_REF="$(read_arg "$DF_BASE" MEMPALACE_TOOLKIT_REF)"
STUDIO_REPO="$(read_arg "$DF_VARIANT" PI_STUDIO_REPO)"
SKILLSET_SNAPSHOT="$(read_arg "$DF_VARIANT" SKILLSET_SNAPSHOT_REF)"
# name|kind|repo|ref -- one line per label the variant image carries.
# kinds: ref = branch/tag/SHA resolved like resolve-versions does;
# studio = highest semver tag of the repo (label lives on <tag>-studio);
# literal = the ARG value IS the baked value (a SHA pin, a version).
REF_COMPONENTS="pi-toolkit|ref|$TOOLKIT_REPO|$TOOLKIT_REF
pi-extensions|ref|$EXTENSIONS_REPO|$EXTENSIONS_REF
pi-fork|ref|$FORK_REPO|$FORK_REF
pi-obsmem|ref|$OBSMEM_REPO|$OBSMEM_REF
pi-atelier|ref|$ATELIER_REPO|$ATELIER_ACTUAL
mempalace-toolkit|ref|$MPTK_REPO|$MPTK_REF
pi-studio|studio|$STUDIO_REPO|
skillset-snapshot|literal||$SKILLSET_SNAPSHOT
pi-version|literal||$PI_ACTUAL
mempalace-version|literal||$MEMPALACE_ACTUAL"
REF_RC=0
# Same discipline as check 8: no `|| true` on the python, or a printed DRIFT
# exits 0. Per-component SKIP lines are counted afterwards by grep, so a run
# that evaluated eight components and could not reach the ninth reports one
# skip, not a green tick over the ninth.
REF_OUT="$(HUB_REPO="$HUB_REPO_PATH" HUB_JSON="$HUB_TAGS_JSON" CHANGELOG="CHANGELOG.md" \
COMPONENTS="$REF_COMPONENTS" python3 <<'PYEOF'
import json, os, re, subprocess, sys, urllib.request, urllib.parse
SHA40 = re.compile(r"^[0-9a-f]{40}$")
SEMVER = re.compile(r"^v?[0-9]+\.[0-9]+\.[0-9]+$")
LABEL = "se.jordbo.pi-devbox."
def ver_key(tag):
return tuple(int(x) for x in tag.lstrip("v").split("."))
def http_json(url, headers=None, timeout=30):
req = urllib.request.Request(url, headers=headers or {})
with urllib.request.urlopen(req, timeout=timeout) as resp:
return json.loads(resp.read().decode("utf-8"))
def labels_of(repo, tag):
"""Config labels of <repo>:<tag>'s amd64 image via the anonymous registry API."""
tok = http_json(
"https://auth.docker.io/token?service=registry.docker.io&scope="
+ urllib.parse.quote(f"repository:{repo}:pull", safe=":")
)["token"]
hdr = {
"Authorization": f"Bearer {tok}",
"Accept": ", ".join([
"application/vnd.oci.image.index.v1+json",
"application/vnd.docker.distribution.manifest.list.v2+json",
"application/vnd.oci.image.manifest.v1+json",
"application/vnd.docker.distribution.manifest.v2+json",
]),
}
base = f"https://registry-1.docker.io/v2/{repo}"
man = http_json(f"{base}/manifests/{tag}", hdr)
if "manifests" in man: # multi-arch index: pick linux/amd64, as check 8 does
cands = [m for m in man["manifests"]
if m.get("platform", {}).get("architecture") == "amd64"
and m.get("platform", {}).get("os") == "linux"]
if not cands:
raise RuntimeError("no linux/amd64 entry in the manifest index")
man = http_json(f"{base}/manifests/{cands[0]['digest']}", hdr)
cfg = http_json(f"{base}/blobs/{man['config']['digest']}", hdr)
return cfg.get("config", {}).get("Labels") or {}
def ls_remote(repo, *patterns):
# GIT_TERMINAL_PROMPT=0: a repo flipped private must fail fast as a SKIP,
# not sit waiting for a username on a CI runner until the job times out.
env = dict(os.environ, GIT_TERMINAL_PROMPT="0")
out = subprocess.run(["git", "ls-remote", repo, *patterns], env=env,
capture_output=True, text=True, timeout=60, check=True).stdout
return {line.split("\t")[1]: line.split("\t")[0] for line in out.splitlines() if "\t" in line}
def resolve_ref(repo, ref):
"""What docker-publish.yml's resolve-versions would pass as the build-arg."""
if SHA40.match(ref):
return ref, ref
refs = ls_remote(repo, f"refs/heads/{ref}", f"refs/tags/{ref}", f"refs/tags/{ref}^{{}}")
for key in (f"refs/tags/{ref}^{{}}", f"refs/heads/{ref}", f"refs/tags/{ref}"):
if key in refs:
return refs[key], ref
raise RuntimeError(f"'{ref}' is neither a branch nor a tag of {repo}")
def resolve_studio(repo):
refs = ls_remote(repo, "refs/tags/*")
tags = {k[len("refs/tags/"):]: v for k, v in refs.items()}
names = sorted((t for t in tags if SEMVER.match(t)), key=ver_key)
if not names:
raise RuntimeError(f"no semver tag at {repo}")
tag = names[-1]
return tags.get(tag + "^{}", tags[tag]), tag
def compare_url(repo, a, b):
root = repo[:-4] if repo.endswith(".git") else repo
return f"{root}/compare/{a}...{b}"
try:
hub = json.loads(os.environ["HUB_JSON"])
except (ValueError, KeyError) as exc:
print(" SKIP ref moves -- Hub API returned unparseable JSON (%s)" % exc)
sys.exit(3)
released = sorted((r["name"] for r in hub.get("results", [])
if isinstance(r.get("name"), str) and re.fullmatch(r"v[0-9]+\.[0-9]+\.[0-9]+", r["name"])),
key=ver_key)
if not released:
print(" SKIP ref moves -- Hub lists no published vX.Y.Z tag to compare against")
sys.exit(3)
last = released[-1]
repo = os.environ["HUB_REPO"]
# The text every not-yet-published change lives in: everything above the last
# published version's heading. Its absence is drift, not a skip.
text = open(os.environ["CHANGELOG"], encoding="utf-8").read()
# (\s|$) rather than \b: a word boundary would accept "## v1.9.2-rc1" or
# "## v1.9.2-typo" as v1.9.2's heading. Caught by the sabotage test, not review.
m = re.search(r"^## v?%s(\s|$)" % re.escape(last.lstrip("v")), text, re.M)
if not m:
print(" DRIFT ref moves -- %s is the last PUBLISHED tag on Hub but %s has no '## %s' heading"
% (last, os.environ["CHANGELOG"], last))
sys.exit(1)
pending = text[:m.start()].lower()
try:
labels = labels_of(repo, last)
except Exception as exc: # network, auth, shape -- all "could not measure"
print(" SKIP ref moves -- could not read %s:%s's labels from the registry (%s); NOT verified"
% (repo, last, exc))
sys.exit(3)
studio_labels = None
checked = drift = 0
problems = []
for line in os.environ["COMPONENTS"].splitlines():
if not line.strip():
continue
name, kind, url, ref = line.split("|", 3)
# <name>-ref labels hold SHAs; names that already end in -version are the
# label (pi-version, mempalace-version) -- a version string, compared literally.
key = LABEL + name if name.endswith("-version") else LABEL + name + "-ref"
try:
if kind == "studio":
if studio_labels is None:
studio_labels = labels_of(repo, last + "-studio")
baked = studio_labels.get(key)
else:
baked = labels.get(key)
except Exception as exc:
print(" SKIP %-18s -- could not read %s:%s-studio's labels (%s)" % (name, repo, last, exc))
continue
if not baked:
print(" SKIP %-18s -- %s carries no %s label" % (name, last, key))
continue
try:
if kind == "ref":
now, shown = resolve_ref(url, ref)
elif kind == "studio":
now, shown = resolve_studio(url)
else:
now, shown = ref, ref
except Exception as exc:
print(" SKIP %-18s -- could not resolve what the next build would bake (%s)" % (name, exc))
continue
checked += 1
is_sha = bool(SHA40.match(now))
short = (lambda s: s[:7] if SHA40.match(s) else s)
if baked == now:
print(" OK %-18s unchanged since %s (%s)" % (name, last, short(now)))
continue
names = [now[:7].lower()] if is_sha else [now.lower()]
if kind == "studio":
names.append(shown.lower())
if any(n in pending for n in names):
print(" OK %-18s %s -> %s since %s, named above the %s heading"
% (name, short(baked), short(now), last, last))
continue
drift += 1
hint = compare_url(url, baked, now) if (url and is_sha and SHA40.match(baked)) else ""
problems.append(" %-18s %s -> %s%s" % (name, short(baked), short(now), (" " + hint) if hint else ""))
print(" DRIFT %-18s %s -> %s since %s, NOT named above the %s heading"
% (name, short(baked), short(now), last, last))
if problems:
print(" Name each new value (7-char SHA prefix, or the tag/version) in CHANGELOG.md above '## %s':" % last)
print("\n".join(problems))
if checked == 0 and drift == 0:
print(" SKIP ref moves -- no component could be evaluated")
sys.exit(3)
sys.exit(1 if drift else 0)
PYEOF
)" || REF_RC=$?
printf '%s\n' "$REF_OUT"
REF_SKIPS="$(printf '%s\n' "$REF_OUT" | grep -c '^ SKIP ' || true)"
case "$REF_RC" in
0) SKIPS=$((SKIPS + REF_SKIPS)) ;;
3) SKIPS=$((SKIPS + 1)) ;;
*)
SKIPS=$((SKIPS + REF_SKIPS))
fail "a component the next build would bake differently from the last published
release is not named in CHANGELOG.md (see DRIFT above). These reach the image
through floating refs, so nothing else in this repo records that they moved;
the CHANGELOG entry is the only place a reader of the next tag can learn it.
Name the new SHA (7 chars is enough) where you describe the change -- the
compare URL above shows what moved."
;;
esac
fi
echo
if [ "$FAILURES" -eq 0 ]; then
if [ "$SKIPS" -gt 0 ]; then
echo "OK: every checked doc claim matches the build files" \
"($SKIPS check(s) SKIPPED and therefore NOT verified -- see SKIP above)."
else
echo "OK: every checked doc claim matches the build files."
fi
exit 0
fi
echo "::error::$FAILURES doc claim(s) have drifted from the build files."
echo
echo "Docs are read from the TAG, not from main: docker-publish.yml checks out"
echo "github.ref, so a fix pushed after tagging does not reach the release or the"
echo "Hub page. Update the docs BEFORE you tag."
if [ "$WARN_ONLY" -eq 1 ]; then
echo "(--warn-only: exiting 0 anyway)"
exit 0
fi
exit 1
+166
View File
@@ -0,0 +1,166 @@
#!/usr/bin/env bash
# check-skill-floor.sh — fail when the vendored pi-extensions skill snapshot in
# rootfs/ ("the floor") has drifted from the package repo it is a snapshot of.
#
# THE DEFECT THIS EXISTS TO CATCH, measured 2026-09-10.
# rootfs/usr/local/share/pi-devbox/skills/pi-extensions/ ships a vendored copy
# of the pi-extensions skill so the skill is ALWAYS present in the image.
# Dockerfile.variant then copies the freshly-cloned package copy OVER the served
# path at /usr/local/share/... — but it never writes back to the repo floor. So
# the floor only silently rots, and it had: 34284 B, untouched since fa04d20
# (2026-07-30), while the package copy was 38973 B. Four copies existed with
# three different sizes.
#
# Why that is worse than ordinary staleness: the floor is a FALLBACK. The copy
# step is guarded by `if [ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build
# where the package clone yields no skill/ keeps the vendored snapshot and still
# succeeds — green, with no manifest flag and no label saying which copy was
# served. The image would ship a July skill and nothing would say so. Keeping
# the floor fresh means that fallback is harmless instead of a silent regression.
#
# WHY A DIRECTORY HASH AND NOT `sha256sum SKILL.md`.
# The same pipeline Dockerfile.variant uses for skillset_snapshot_tree_sha256,
# and for the same documented reason: a file-only compare answers "did this one
# file change", not "is this the same skill". pi-extensions ships TWO files
# (SKILL.md + evaluate-extension-usage.py), so a sibling-file edit would pass a
# file-only check. If you change the pipeline here, change it there too.
#
# WHY GATING ON ANOTHER REPO IS PROPORTIONATE HERE, since that is normally a
# smell: this fires only when the package's skill/ DIRECTORY HASH changes, which
# is exactly and only when the floor has genuinely gone stale. pi-extensions
# commits that do not touch skill/ leave the hash alone and cannot turn this red.
# The repo is also anonymously clonable (verified 2026-09-10 with `git ls-remote`
# and no credentials), so this needs no secret and cannot break on token expiry.
#
# Exit codes — deliberately three, matching scripts/lint-shell.sh's philosophy
# that a gate which cannot run must not pass:
# 0 in sync (or the package legitimately has no skill/ at this ref)
# 1 DRIFT — the floor differs from the package
# 2 cannot run — no package copy could be obtained
set -euo pipefail
REPO_ROOT=$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)
FLOOR_DIR="${REPO_ROOT}/rootfs/usr/local/share/pi-devbox/skills/pi-extensions"
# Defaults mirror Dockerfile.variant's ARGs so this checks what the build builds.
PI_EXTENSIONS_REPO="${PI_EXTENSIONS_REPO:-https://gitea.jordbo.se/joakimp/pi-extensions.git}"
PI_EXTENSIONS_REF="${PI_EXTENSIONS_REF:-main}"
PACKAGE_DIR=""
WARN_ONLY=0
TMPDIR_CLONE=""
usage() {
cat <<'EOF'
Usage: scripts/check-skill-floor.sh [options]
--package-dir DIR Compare against an existing skill directory instead of
cloning. In a devbox container use /opt/pi-extensions/skill
for a fully offline run.
--warn-only Report drift but exit 0 (advisory use, e.g. a local hook).
-h, --help This text.
Environment: PI_EXTENSIONS_REPO, PI_EXTENSIONS_REF (default main) — both mirror
the Dockerfile.variant ARGs of the same name.
EOF
}
while [ $# -gt 0 ]; do
case "$1" in
--package-dir) PACKAGE_DIR="${2:-}"; shift 2 ;;
--warn-only) WARN_ONLY=1; shift ;;
-h|--help) usage; exit 0 ;;
*) echo "::error::unknown argument: $1" >&2; usage >&2; exit 2 ;;
esac
done
cleanup() {
if [ -n "$TMPDIR_CLONE" ]; then rm -rf "$TMPDIR_CLONE"; fi
}
trap cleanup EXIT
# Identical to Dockerfile.variant's tree_sha256(): relative paths + per-file
# sha256 over a sorted `find`, folded into one digest. Deterministic, never
# readdir order.
tree_sha256() {
( cd "$1" && find . -type f -print | LC_ALL=C sort | xargs -r sha256sum ) \
2>/dev/null | sha256sum | cut -d' ' -f1
}
if [ ! -d "$FLOOR_DIR" ]; then
echo "::error::floor directory is missing: ${FLOOR_DIR}"
echo "::error::rootfs/ is supposed to guarantee the skill is always in the image."
exit 2
fi
SOURCE_DESC=""
if [ -n "$PACKAGE_DIR" ]; then
if [ ! -d "$PACKAGE_DIR" ]; then
echo "::error::--package-dir does not exist: ${PACKAGE_DIR}"
exit 2
fi
SOURCE_DESC="local directory ${PACKAGE_DIR}"
else
command -v git >/dev/null 2>&1 || { echo "::error::git not found; cannot obtain the package copy."; exit 2; }
TMPDIR_CLONE=$(mktemp -d)
# Fetch the single ref shallowly. `git fetch <ref>` accepts a branch, a tag
# and (on Gitea) a reachable commit, which is why this is not `clone --branch`
# — CI resolves PI_EXTENSIONS_REF to a 40-hex SHA before the build.
if ! ( cd "$TMPDIR_CLONE" \
&& git init -q . \
&& git remote add origin "$PI_EXTENSIONS_REPO" \
&& git fetch -q --depth 1 origin "$PI_EXTENSIONS_REF" \
&& git checkout -q FETCH_HEAD ) 2>/dev/null; then
echo "::error::could not fetch ${PI_EXTENSIONS_REF} from ${PI_EXTENSIONS_REPO}"
echo "::error::Cannot determine whether the floor is stale, so this is exit 2, not a pass."
echo "::error::For an offline run, pass --package-dir /opt/pi-extensions/skill"
exit 2
fi
PACKAGE_SHA=$( cd "$TMPDIR_CLONE" && git rev-parse --short HEAD )
PACKAGE_DIR="${TMPDIR_CLONE}/skill"
SOURCE_DESC="${PI_EXTENSIONS_REPO} @ ${PI_EXTENSIONS_REF} (${PACKAGE_SHA})"
fi
# A ref with no skill/ is the documented fallback case: Dockerfile.variant keeps
# the vendored snapshot and the build succeeds. Nothing to compare, so this is
# not drift — but it IS the exact condition under which the floor ships, so say
# so loudly rather than printing a silent green tick.
if [ ! -d "$PACKAGE_DIR" ]; then
echo "::warning::package has no skill/ at this ref — the vendored floor is what will ship."
echo " source : ${SOURCE_DESC}"
echo " floor : $(tree_sha256 "$FLOOR_DIR")"
exit 0
fi
FLOOR_HASH=$(tree_sha256 "$FLOOR_DIR")
PKG_HASH=$(tree_sha256 "$PACKAGE_DIR")
if [ "$FLOOR_HASH" = "$PKG_HASH" ]; then
echo "OK: vendored pi-extensions floor matches the package."
echo " source : ${SOURCE_DESC}"
echo " tree_sha256: ${FLOOR_HASH}"
exit 0
fi
# `set -e` interacts badly with `[ … ] && x` as a bare statement, so both of
# these are explicit if-blocks rather than AND-lists.
LEVEL="error"
if [ "$WARN_ONLY" -eq 1 ]; then LEVEL="warning"; fi
echo "::${LEVEL}::vendored pi-extensions skill floor has DRIFTED from the package."
echo " source : ${SOURCE_DESC}"
echo " floor tree_sha256 : ${FLOOR_HASH}"
echo " pkg tree_sha256 : ${PKG_HASH}"
echo ""
echo " per-file differences:"
diff -rq "$FLOOR_DIR" "$PACKAGE_DIR" 2>&1 | sed 's/^/ /' || true
echo ""
echo " Remedy — re-sync the floor and commit it:"
echo " cp -a <pi-extensions>/skill/. ${FLOOR_DIR}/"
echo " git add ${FLOOR_DIR#"${REPO_ROOT}/"} && git commit"
echo ""
echo " NOTE this forces one full base rebuild: base_tag hashes Dockerfile.base"
echo " + rootfs/, and that rebuild is what re-bakes the refreshed floor."
if [ "$WARN_ONLY" -eq 1 ]; then exit 0; fi
exit 1
+97
View File
@@ -0,0 +1,97 @@
#!/usr/bin/env bash
# Shellcheck + syntax-check every shell script in this repo. Severity: error.
#
# SINGLE SOURCE OF TRUTH for two callers:
# .gitea/workflows/lint.yml — advisory, every branch push and PR
# .gitea/workflows/docker-publish.yml — the release GATE (lint-gate job)
# Extracted from lint.yml on 2026-09-08 rather than copied, because a second
# copy is exactly the drift this repo has been bitten by (see skillset's
# pi-extensions mirror, refreshed the same evening after sitting 9579 B behind).
#
# WHY THIS CHECK EXISTS AT ALL
# actionlint shellchecks workflow `run:` steps only. The repo's own scripts —
# entrypoint.sh, scripts/*.sh, and the extensionless tools under
# rootfs/usr/local/bin/ — were never shellchecked. A sibling repo with the same
# gap shipped a broken `echo "$json" | python3 <<'EOF' ... json.load(sys.stdin)`
# for two months: with no script argument python reads its SCRIPT from stdin,
# so the heredoc IS stdin and json.load hits EOF. shellcheck flags that at
# severity error (SC2259); nothing ever ran it.
#
# WHY THE RELEASE GATES ON IT (added 2026-09-08, the expensive way round)
# v1.8.14's first attempt failed after build-base had already spent ~46 min:
# scripts/smoke-test.sh had an apostrophe inside a single-quoted exec_test body
# ("the fleet\'s"), which CLOSES the string, so the body truncated and its tail
# ran on the CI runner instead of inside the image. shellcheck had already
# caught it as SC2289 at severity error — the lint job went red on the very
# push that introduced it and stayed red for 24 hours, unread. lint.yml
# deliberately does not run on tag pushes (sound: the tagged tree was linted on
# main, and a tag-ref lint run sorts above the publish run and makes a release
# look finished early). The gap was never "lint the tag" — it was that a tree
# whose lint FAILED could still be released. Hence a gate inside the publish
# workflow, ~40 s, ahead of everything expensive.
#
# SEVERITY CHOICE
# -S error is 0 findings across this repo when clean, so it is free to add.
# -S warning is NOT free here (20x SC2088 tilde-in-quotes in
# recreate-sanity-check.sh, plus assorted SC2016 — both intentional), and a
# noisy gate trains people to ignore it. Error-only, matching the
# SHELLCHECK_OPTS philosophy in lint.yml.
# Reproduce the count before editing it (the `$ ` prefix is load-bearing: a
# comment whose first word is "shellcheck" is parsed as a DIRECTIVE, and a
# malformed one is SC1072/SC1073 at severity error — this gate caught exactly
# that when the line was first written without it):
# $ shellcheck -S warning -f gcc scripts/*.sh rootfs/usr/local/bin/* \
# entrypoint*.sh hooks/* | grep -c SC2088
#
# Usage: bash scripts/lint-shell.sh [root] (default root: repo top level)
set -uo pipefail
root="${1:-}"
if [ -z "$root" ]; then
root="$(git rev-parse --show-toplevel 2>/dev/null || pwd)"
fi
cd "$root" || { echo "::error::cannot cd to $root"; exit 2; }
# A gate that cannot run must not pass. Without this, a machine (or a CI job
# whose install step was reordered away) without shellcheck would sail through
# printing nothing, which is the failure mode this whole file exists to prevent.
if ! command -v shellcheck >/dev/null 2>&1; then
echo "::error::shellcheck not found — the gate cannot run, so it must not pass" >&2
echo " install it (apt-get install -y shellcheck) or run this in CI" >&2
exit 2
fi
# Union of two signals, because either alone misses a real case: a shebang scan
# misses a sourced fragment with no shebang, and a *.sh glob misses the
# extensionless tools in rootfs/usr/local/bin/. Silent skipping is precisely the
# failure mode this gate exists to prevent, so err toward over-collecting.
# -print0/mapfile -d '' so a path containing a space cannot silently split.
mapfile -d '' -t all_files < <(find . -not -path './.git/*' -type f -print0)
sh_files=()
for f in "${all_files[@]}"; do
case "$f" in *.sh) sh_files+=("$f"); continue;; esac
if head -n1 "$f" 2>/dev/null | grep -qE '^#!.*\b(bash|sh)\b'; then
sh_files+=("$f")
fi
done
echo "Checking ${#sh_files[@]} shell file(s) with $(shellcheck --version | awk '/version:/{print $2}')"
# A green tick over an empty file set is not a check.
if [ "${#sh_files[@]}" -eq 0 ]; then
echo "::error::no shell files found — the shebang scan or the checkout is wrong"
exit 1
fi
rc=0
shellcheck -S error -f gcc "${sh_files[@]}" || rc=1
# bash -n catches a different class than shellcheck (unbalanced constructs it
# declines to parse), so both run and both count.
for f in "${sh_files[@]}"; do
bash -n "$f" || { echo "::error file=$f::bash -n failed"; rc=1; }
done
if [ "$rc" -eq 0 ]; then
echo "OK: ${#sh_files[@]} shell file(s) clean at severity error"
fi
exit "$rc"
+165 -2
View File
@@ -11,7 +11,8 @@
# pi-observational-memory / (studio variant) pi-studio package
# registrations in settings.json packages[]
# - Shell defaults re-seeded from /etc/skel-devbox
# - /tmp/sshcm exists with mode 700 (ssh ControlMaster dir)
# - ssh ControlMaster works: /tmp/sshcm exists 700 AND the ControlPath that
# ssh actually resolves (ssh -G) is a writable directory
# - /opt toolkits intact
# - Known expected-absences don't regress
#
@@ -425,13 +426,175 @@ if [ -d /opt/pi-atelier ] && command -v jq >/dev/null 2>&1; then
fi
echo
echo "-- ssh ControlMaster dir --"
echo "-- ssh ControlMaster: socket dir + EFFECTIVE ControlPath --"
# TWO LAYERS, and the second is the one that has actually broken in the field.
#
# LAYER 1 (original check): /tmp/sshcm, the directory entrypoint-user.sh creates
# for the base image's system drop-in
# (/etc/ssh/ssh_config.d/00-devbox-controlmaster.conf).
#
# LAYER 2 (added 2026-09-15): the directory a config NAMES — which is not the
# same question, and asserting layer 1 is structurally blind to it. On
# emb-7kj4vr4g a durable ~/.pi/ssh/config pointed ControlPath at /tmp/ssh-cm
# (with a hyphen), a directory nothing in the image creates. EVERY ssh died
# unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory
# rc=255 with the remote command never running — while this script printed a
# green tick for layer 1, truthfully, about the wrong object.
#
# The same rc=255 has a second, independent cause already documented in prose in
# Dockerfile.base ("SSH client defaults" CAVEAT) and never verified anywhere: a
# per-host `ControlPath ~/.ssh/cm/%r@%h:%p` inherited from a bind-mounted
# READ-ONLY ~/.ssh. Measured to be the identical failure class:
# unix_listener: cannot bind to path ~/.ssh/cm/...: Read-only file system
# So do not guess which config wins — ask ssh. `ssh -G` applies real config
# precedence (first-obtained-value-wins, system drop-in, Include, -F override)
# and prints the fully expanded ControlPath. Require its parent to exist and be
# writable. Cost measured at 0.116 s for 48 hosts; -G never opens a connection.
if [ -d /tmp/sshcm ] && [ "$(stat -c %a /tmp/sshcm 2>/dev/null)" = "700" ]; then
pass "/tmp/sshcm exists with mode 700"
else
fail "/tmp/sshcm missing or not mode 700"
fi
# Probe one route. $1 = label, $2 = config to force with -F ("" = ssh's own
# default precedence), $3 = severity when a ControlPath dir is unusable.
#
# SEVERITY SPLIT IS DELIBERATE. The default route legitimately resolves into the
# read-only ~/.ssh on any host whose own config pins ControlPath there, and the
# supported workaround (`ssh -F ~/.ssh-local/config`) already exists — so that
# is a warn, not a fail. Failing it would paint this script red on every run of
# every device, and a check that fires benignly every time is one you learn to
# ignore. The sidecar route is the PRESCRIBED one, so there it is a hard fail.
_ssh_cm_probe() {
local label="$1" cfg="${2:-}" sev="${3:-fail}"
local h out cm cp dir n=0 shown
local bad=()
while IFS= read -r h; do
[ -n "$h" ] || continue
if [ -n "$cfg" ]; then
out=$(ssh -F "$cfg" -G "$h" 2>/dev/null) || continue
else
out=$(ssh -G "$h" 2>/dev/null) || continue
fi
cm=$(printf '%s\n' "$out" | awk '/^controlmaster /{print $2; exit}')
case "$cm" in '' | no | none | false) continue ;; esac
cp=$(printf '%s\n' "$out" | awk '/^controlpath /{print $2; exit}')
case "$cp" in '' | none) continue ;; esac
n=$((n + 1))
dir=$(dirname "$cp")
if [ ! -d "$dir" ] || [ ! -w "$dir" ]; then
bad+=("$h")
fi
done <<< "$SSH_CM_HOSTS"
# Cap the host list. A 50-name line is the noisy gate this repo already warns
# about in lint-shell.sh: unreadable output is ignored output. Six names plus
# a count is enough to identify the class and act.
if [ "${#bad[@]}" -gt 0 ]; then
shown="${bad[*]:0:6}"
if [ "${#bad[@]}" -gt 6 ]; then
shown="$shown (+$(( ${#bad[@]} - 6 )) more)"
fi
fi
if [ "$n" -eq 0 ]; then
warn "$label: no host resolves to ControlMaster on — effective ControlPath not exercised"
elif [ "${#bad[@]}" -eq 0 ]; then
pass "$label: $n ControlMaster host(s), every ControlPath dir exists and is writable"
elif [ "$sev" = "warn" ]; then
warn "$label: ${#bad[@]}/$n host(s) resolve ControlPath to a missing or unwritable dir [$shown] — STRUCTURAL and permanent while ~/.ssh/config pins ControlPath inside the read-only ~/.ssh, so this line can never reach zero and is not a to-do; $_git_ssh_note (see Dockerfile.base CAVEAT)"
else
fail "$label: ${#bad[@]}/$n host(s) resolve ControlPath to a missing or unwritable dir [$shown] — ssh dies rc=255 'unix_listener: cannot bind to path' and the remote command never runs"
fi
}
if command -v ssh >/dev/null 2>&1; then
# Whether git-over-ssh already routes through the sidecar decides how much the
# permanent default-route warning below actually matters, so state it IN that
# message rather than leaving each reader to work it out. Asserted properly as
# its own pass/fail in the next section.
#
# THREE states, not two, and the [ -r ] test is why: core.sshCommand NAMING the
# sidecar does not mean the sidecar EXISTS. Without that test, the arm where the
# path is wired but the file is gone printed "git IS wired ... unaffected" about
# a state in which every single git-over-ssh call fails. Found by exercising all
# five arms of this check rather than only the healthy one.
if [ ! -r "$HOME/.ssh-local/config" ]; then
_git_ssh_note="there is no ssh sidecar on this host, so nothing for -F to point at and these hosts cannot multiplex at all — see the git-over-ssh check below"
elif command -v git >/dev/null 2>&1 &&
git config --global --get core.sshCommand 2>/dev/null |
grep -qF -- "-F $HOME/.ssh-local/config"; then
_git_ssh_note="git IS wired to the sidecar (core.sshCommand), so git push/fetch is unaffected; bare 'ssh' to these hosts still needs 'ssh -F ~/.ssh-local/config'"
else
_git_ssh_note="git is NOT wired to the sidecar (see the git-over-ssh check below), so both git and bare 'ssh' need 'ssh -F ~/.ssh-local/config'"
fi
# Concrete Host aliases only: patterns (*, ?) and negations (!) are not
# connectable targets, so `ssh -G` on them proves nothing.
_cm_cfgs=()
if [ -r "$HOME/.ssh/config" ]; then _cm_cfgs+=("$HOME/.ssh/config"); fi
if [ -r "$HOME/.ssh-local/config" ]; then _cm_cfgs+=("$HOME/.ssh-local/config"); fi
if [ "${#_cm_cfgs[@]}" -gt 0 ]; then
SSH_CM_HOSTS=$(awk 'tolower($1)=="host"{for(i=2;i<=NF;i++) if ($i !~ /[*?!]/) print $i}' \
"${_cm_cfgs[@]}" 2>/dev/null | sort -u)
else
SSH_CM_HOSTS=""
fi
if [ -z "$SSH_CM_HOSTS" ]; then
warn "no concrete Host aliases in ~/.ssh/config or ~/.ssh-local/config — effective ControlPath not verified"
else
_ssh_cm_probe "default ssh precedence" "" warn
if [ -r "$HOME/.ssh-local/config" ]; then
_ssh_cm_probe "ssh -F ~/.ssh-local/config" "$HOME/.ssh-local/config" fail
else
warn "~/.ssh-local/config absent — setup-lan-access.sh did not run; the prescribed multiplex route is unverified"
fi
fi
else
warn "ssh not on PATH — effective ControlPath not verified"
fi
echo
echo "-- git-over-ssh routed through the writable ssh sidecar --"
# The one question in this area that is BINARY, FIXABLE, and therefore worth a
# check that can reach zero and stay there.
#
# The ControlPath warning above cannot: while a bind-mounted ~/.ssh/config pins
# ControlPath inside the read-only ~/.ssh, the default route will ALWAYS resolve
# to an unwritable dir, so that line fires benignly on every run of every device
# forever — and the script's own comment above says why that is dangerous: it is
# a warning you learn to skip. MEASURED 2026-09-22, exactly that: an agent ran
# this script, recorded "two by-design warnings", then two hours later hit
# `unix_listener: cannot bind ... Read-only file system` on `git push`, failed to
# connect it to the warning it had already read, and reinvented a /tmp/sshcm
# workaround — while the remedy string sat inside the dismissed warning.
# entrypoint-user.sh now wires core.sshCommand to the sidecar so nobody has to
# know; this section asserts the wiring actually happened, which is the part a
# reader can act on.
if ! command -v git >/dev/null 2>&1; then
warn "git not on PATH — sidecar wiring not verified"
elif [ ! -r "$HOME/.ssh-local/config" ]; then
# No sidecar (native Linux Docker, where setup-lan-access.sh writes none). The
# correct state is UNSET: -F pointing at a missing file breaks every
# git-over-ssh call, which is worse than the problem being solved.
if [ -z "$(git config --global --get core.sshCommand 2>/dev/null || true)" ]; then
pass "no ssh sidecar on this host and core.sshCommand correctly unset (guard holds)"
else
fail "core.sshCommand is set but ~/.ssh-local/config does not exist — every git-over-ssh call dies on a missing -F file; entrypoint-user.sh's [ -r ] guard did not hold"
fi
else
_git_ssh_cmd=$(git config --global --get core.sshCommand 2>/dev/null || true)
case "$_git_ssh_cmd" in
*"-F $HOME/.ssh-local/config"*)
pass "git core.sshCommand routes through the sidecar ($_git_ssh_cmd)" ;;
'')
fail "a sidecar exists but git core.sshCommand is unset — 'git push' to any host whose config pins ControlPath inside the read-only ~/.ssh dies rc=255 behind git's misleading 'correct access rights'. entrypoint-user.sh should set it; expected on images built before that wiring landed, where the fix is: git config --global core.sshCommand \"ssh -F \$HOME/.ssh-local/config\"" ;;
*)
warn "git core.sshCommand set to something else and left alone (first-wins, deliberate): $_git_ssh_cmd" ;;
esac
fi
echo
echo "-- Shell defaults re-seeded from /etc/skel-devbox --"
if [ -f "$HOME/.bash_aliases" ]; then
+247 -5
View File
@@ -28,6 +28,9 @@
# (human, --json, --quiet)
# - (studio variant only, auto-detected) pi-studio cloned + prebuilt
# client bundle present + registered via `pi install`
# - no foreign npm-11 platform packages (@esbuild, clipboard) beyond the host
# - no build-time npm cache (/root/.npm) shipped in the image
# - esbuild compiles + clipboard native loads at every install site
# - image size within threshold
set -euo pipefail
@@ -105,6 +108,18 @@ else
run "node" "node --version"
fi
run "git" "git --version"
# NOTE: the shellcheck binary is a GATE DEPENDENCY, not a convenience.
# scripts/lint-shell.sh is the release gate (the lint-gate job resolve-versions
# depends on) and it exits 2 when the binary is missing, by design — "a gate that
# cannot run must not pass". Measured on v1.8.14: it was absent from the image, so
# that gate could not be run by a developer in ANY container, only in CI.
# Asserted here so its absence fails a build instead of being discovered by a hook
# that then refuses every push (hooks/pre-push).
#
# This comment must not BEGIN with the tool's name: a line starting with
# `# shellcheck` is parsed as a DIRECTIVE, not a comment (SC1073/SC1072). The
# gate added in this same change caught that here, before the push.
run "shellcheck (lint gate dependency)" "shellcheck --version | grep -qE '^version: [0-9]'"
run "aws" "aws --version"
run "uv" "uv --version"
run "nvim" "nvim --version"
@@ -235,6 +250,22 @@ run_expect "remote-palace-without-inbox skip is announced, not silent" \
"MemPalace catch-up skipped"
run "...and the skip notice names the variable that fixes it" \
"grep -A6 'MemPalace catch-up skipped' /usr/local/bin/entrypoint-user.sh | grep -q 'MEMPALACE_PI_SSH_TARGET'"
# git-over-ssh must be wired to the writable ssh sidecar, because ~/.ssh is
# commonly bind-mounted READ-ONLY and a per-host `ControlPath ~/.ssh/cm/...`
# inherited from it kills every push with `unix_listener: cannot bind ...:
# Read-only file system` behind git's misleading "correct access rights".
#
# STATIC assertion, deliberately. `run` executes `docker run --entrypoint=""`, so
# entrypoint-user.sh never runs here and `git config --global core.sshCommand` is
# necessarily unset — asserting the VALUE in this harness would repeat the v1.8.0
# mistake (an assertion that cannot pass, unvalidated until the next tag). What
# IS checkable at build time is that the wiring code shipped in the image. The
# runtime value is asserted below in the Runtime deployment section, where the
# real entrypoint chain has run.
run "entrypoint wires git core.sshCommand to the ssh sidecar" \
"grep -q 'core.sshCommand \"ssh -F' /usr/local/bin/entrypoint-user.sh"
run "...and guards it on the sidecar existing (native Linux has none)" \
"grep -B2 'core.sshCommand \"ssh -F' /usr/local/bin/entrypoint-user.sh | grep -q '\\[ -r \"\$HOME/.ssh-local/config\" \\]'"
# A remote mine that FAILS must not report success. MCP answers a hard tool
# failure with HTTP 200 and the tool's own JSON escaped inside
# result.content[].text, so the feeder's old `'\"error\"' in body` check could
@@ -303,8 +334,24 @@ run "pi-toolkit clone" "test -d /opt/pi-toolkit && git -C /opt/pi-toolkit rev
run "pi-extensions clone" "test -d /opt/pi-extensions && git -C /opt/pi-extensions rev-parse --short HEAD"
run "pi-fork clone + node_modules" \
"test -f /opt/pi-fork/package.json && test -d /opt/pi-fork/node_modules"
run "pi-observational-memory clone + node_modules" \
"test -f /opt/pi-observational-memory/package.json && test -d /opt/pi-observational-memory/node_modules"
# om is checked differently from pi-fork ON PURPOSE. It declares ZERO runtime
# dependencies: 8 devDependencies (omitted by --omit=dev) and 4 peerDependencies,
# which pi itself provides. npm 10 still materialised a node_modules for it, but
# that directory held exactly ONE file (.package-lock.json, 4 KB) and no nested
# package.json at all — 20 empty scope dirs. npm 11 stopped creating it, so the
# old `test -d node_modules` assertion went red on v1.9.0 while nothing about om
# had changed or broken. It was asserting an npm artefact, not a property of the
# shipped software. What actually has to hold is that the entry point pi loads
# exists, so assert THAT, straight out of the manifest pi reads
# (package.json -> pi.extensions), rather than a hardcoded path that could drift.
run "pi-observational-memory clone + declared pi entry point" \
"test -f /opt/pi-observational-memory/package.json && \
node -e 'const p=require(\"/opt/pi-observational-memory/package.json\"),f=require(\"fs\"),h=require(\"path\"); \
const l=(p.pi&&p.pi.extensions)||[]; \
if(!l.length){console.error(\"package.json declares no pi.extensions\");process.exit(1)} \
for(const e of l){const t=h.resolve(\"/opt/pi-observational-memory\",e); \
if(!f.existsSync(t)){console.error(\"declared entry missing: \"+t);process.exit(1)}} \
console.log(\"entries ok: \"+l.join(\",\"))'"
# ...and that the clone carries the AUTH FIX, not merely that it exists. om's
# pre-flight hasUsableAuth() check silently disabled `recall` for ~8 weeks once
# pi moved to request-time SigV4 signing and stopped exposing a static Bedrock
@@ -514,6 +561,50 @@ run "manifest skill fingerprint matches the baked snapshot" '
echo "manifest=[$m] actual=[$a]" >&2
[ -n "$m" ] && [ "$m" = "$a" ]
'
# ── Which pi-extensions skill copy shipped ──────────────────────────────
# Closes the silent-fallback hole. The refresh in Dockerfile.variant is guarded
# by `[ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates
# the co-located skill keeps the vendored floor and still succeeds GREEN, with
# nothing recording that a snapshot shipped instead of the package copy. Measured
# 2026-09-10: the floor had been stale since 2026-07-30, so that path would have
# shipped a six-week-old skill in silence. The floor is fresh now and gated by the
# skill-floor lint job, but "the fallback is currently harmless" is a fact with a
# shelf life, whereas "the image says which copy it got" keeps working.
#
# vendored-floor FAILS here rather than merely warning: these images track main,
# where the package has co-located skill/ since fa04d20, so a fallback means the
# clone did not resolve as intended and that is a defect to investigate. A fork
# deliberately pointing at a mirror without skill/ is the one case that should
# edit this assertion — which is the honest place for that decision to surface.
run "manifest names which pi-extensions skill copy shipped" '
j=/etc/pi-devbox/build-manifest.json
s=$(jq -r ".pi_extensions_skill_source // empty" $j)
h=$(jq -r ".pi_extensions_skill_tree_sha256 // empty" $j)
echo "source=[$s] tree_sha256=[$h]" >&2
printf "%s" "$h" | grep -qxE "[0-9a-f]{64}" || {
echo "pi_extensions_skill_tree_sha256 is not a 64-hex digest" >&2; exit 1; }
case "$s" in
package) ;;
vendored-floor)
echo "FALLBACK: clone had no skill/ at this ref, so the image ships the committed floor" >&2; exit 1 ;;
divergent)
echo "MIXED: served directory is part package and part floor" >&2; exit 1 ;;
*)
echo "pi_extensions_skill_source absent or unrecognised" >&2; exit 1 ;;
esac
'
# Same shape as the mempalace fingerprint check above, and for the same reason: a
# recorded hash that is never recomputed is a claim, not a measurement.
run "recorded pi-extensions skill hash matches the served bytes" '
j=/etc/pi-devbox/build-manifest.json
d=/usr/local/share/pi-devbox/skills/pi-extensions
m=$(jq -r ".pi_extensions_skill_tree_sha256 // empty" $j)
a=$( (cd "$d" && find . -type f -print | LC_ALL=C sort | xargs -r sha256sum) | sha256sum | cut -d" " -f1)
echo "manifest=[$m] actual=[$a]" >&2
[ -n "$m" ] && [ "$m" = "$a" ]
'
# OCI labels live in the image config, not the container fs — inspect them
# from the host docker rather than via `docker run`.
LBL=$(docker inspect --format '{{ index .Config.Labels "se.jordbo.pi-devbox.pi-extensions-ref" }}' "$IMAGE" 2>/dev/null || true)
@@ -522,6 +613,19 @@ if [ -n "$LBL" ] && [ "$LBL" != "<no value>" ]; then
else
printf " ❌ OCI label se.jordbo.pi-devbox.pi-extensions-ref missing or empty\n"; FAIL=$((FAIL+1))
fi
# mempalace-version is set in Dockerfile.base and INHERITED by the variant, so
# it states the pin of the base this image actually built on. It must equal the
# installed binary: the one way they diverge is a base built with
# INSTALL_MEMPALACE=false (label says 3.x, nothing installed) or an install that
# resolved to something other than the pin — both invisible to a label-only
# check. Same ground-truth rule as the manifest assertion above.
MP_LBL=$(docker inspect --format '{{ index .Config.Labels "se.jordbo.pi-devbox.mempalace-version" }}' "$IMAGE" 2>/dev/null || true)
MP_BIN=$(docker run --rm --entrypoint= "$IMAGE" sh -c 'mempalace --version 2>/dev/null | head -n1 | tr -d "\r"' 2>/dev/null || true); MP_BIN=${MP_BIN##* }
if [ -n "$MP_LBL" ] && [ "$MP_LBL" != "<no value>" ] && [ "$MP_LBL" = "$MP_BIN" ]; then
printf " ✅ OCI label se.jordbo.pi-devbox.mempalace-version=%s equals the installed core\n" "$MP_LBL"; PASS=$((PASS+1))
else
printf " ❌ OCI label se.jordbo.pi-devbox.mempalace-version=[%s] vs installed mempalace=[%s]\n" "$MP_LBL" "$MP_BIN"; FAIL=$((FAIL+1))
fi
# ── Runtime deployment (needs entrypoint to run) ──────────────────────
echo ""
@@ -571,6 +675,23 @@ exec_test "settings.json bootstrapped" 'test -f $HOME/.pi/agent/sett
exec_test "pi-devbox-environment skill linked" 'test -L $HOME/.agents/skills/pi-devbox-environment && test -f $HOME/.agents/skills/pi-devbox-environment/SKILL.md && echo ok'
exec_test "pi-extensions skill linked (fallback)" 'test -L $HOME/.agents/skills/pi-extensions && test -f $HOME/.agents/skills/pi-extensions/SKILL.md && echo ok'
exec_test "mempalace skill linked (fallback)" 'test -L $HOME/.agents/skills/mempalace && test -f $HOME/.agents/skills/mempalace/SKILL.md && echo ok'
# git-over-ssh sidecar wiring, asserted as a BICONDITIONAL rather than "is set".
# setup-lan-access.sh writes ~/.ssh-local/config only on VM-backed hosts, and a
# CI runner is native Linux Docker — so "core.sshCommand is set" would fail here
# for a correct image, which is precisely the v1.8.0 trap (an assertion whose
# environment was never checked, unvalidated until the next tag). Both arms are
# real: sidecar present => must route through it; sidecar absent => must be UNSET,
# because -F pointing at a missing file breaks every git-over-ssh call and is
# worse than the problem being fixed. This arm is the one CI actually exercises,
# so CI validates the guard; the other is covered by the static greps above and
# by scripts/recreate-sanity-check.sh on a real device.
exec_test "git core.sshCommand matches sidecar presence" '
cmd=$(git config --global --get core.sshCommand 2>/dev/null || true)
if [ -r "$HOME/.ssh-local/config" ]; then
case "$cmd" in *"-F $HOME/.ssh-local/config"*) echo "wired: $cmd" ;; *) exit 1 ;; esac
else
[ -z "$cmd" ] || exit 1; echo "no sidecar, correctly unset"
fi'
# The vendored mempalace snapshot is refreshed MANUALLY per release (see
# rootfs/usr/local/share/pi-devbox/skills/VENDORED.md). Through v1.8.4 it also
# silently SHADOWED the live skillset copy, so staleness was invisible — and the
@@ -622,7 +743,22 @@ exec_test "mempalace skill linked (fallback)" 'test -L $HOME/.agents/skills
# (forgotten bump AND re-vendored stale snapshot). Upstream content behind this
# refresh: the bare project-name wing convention and the <harness>@<device>
# added_by rule.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Diaries self-heal; plain drawers do not" "$f" && ! grep -q "Agent diaries live in" "$f" && echo ok'
#
# Unreleased: RE-PINNED again on refresh e9e09d9 -> e9e45f7. The retired pair was
# still green against the new snapshot (the diaries section was untouched), so it
# was blind to this refresh for the same reason the v1.8.13 pair was blind to
# that one. The replacement pair is unusually strong because BOTH witnesses come
# out of the same upstream commit: skillset e9e45f7 ADDED the "Withdrawing an ask
# you sent" bullet and DELETED the sentence "there is nothing anyone can do about
# it from the other end" that the new bullet contradicts. Directions were
# MEASURED against both files, not read off the diff: "Withdrawing an ask you
# sent" is new=1/old=0, "nothing anyone can do about it from the other end" is
# new=0/old=1. A canary whose negative witness was removed by the very commit it
# pins fails loudly on the OLD bytes instead of merely failing to notice them,
# which is the property every previous pair here lacked. Upstream content:
# requester-side ask withdrawal became DEPLOYED behaviour once v1.9.1 baked
# mempalace-toolkit e68ee20 (>= e2b060a) through the floating MEMPALACE_TOOLKIT_REF.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Withdrawing an ask you sent" "$f" && ! grep -q "nothing anyone can do about it from the other end" "$f" && echo ok'
# Link TARGETS, not just link existence: with no skillset mounted (as here) the
# baked tree must be what resolves, for all four vendored skills.
exec_test "vendored skills resolve to the baked tree (no skillset mounted)" \
@@ -641,9 +777,17 @@ exec_test "pi-devbox-version reports skill sources (all baked, no skillset here)
'out=$(pi-devbox-version)
echo "$out" | grep -q "skills:" || { echo "no skills section" >&2; exit 1; }
for s in mempalace pi-extensions pi-devbox-environment credential-incident-response; do
echo "$out" | grep -qE "^ $s +baked$" \
echo "$out" | grep -qE "^ $s +baked( \([^)]*\))?$" \
|| { echo "$s not reported as baked" >&2; exit 1; }
done; echo ok'
# The optional " (...)" above is what pi-extensions now appends to say WHICH copy
# shipped — "baked (package copy)", or a loud FALLBACK/MIXED annotation. Without
# allowing it, adding that annotation turned this assertion red on v1.9.0 even
# though the state it reported was the correct one. The suffix is deliberately
# matched loosely rather than pinned to "(package copy)", because WHICH copy
# shipped is already asserted authoritatively above, against the manifest field
# and its measured tree hash, and duplicating that here in a regex would just
# create a second place to update whenever the wording changes.
# The boot banner must NOT carry the section: entrypoint-user.sh prints the
# version FIRST, before the baked links exist and long before the skillset
# deploy + reconcile run last, so anything it said about skill sources would be
@@ -751,10 +895,33 @@ exec_test "pi-atelier registered from /opt, not npm: (volume-shadowing guard)" \
# shadows it. The check that actually bites lives in
# recreate-sanity-check.sh, which runs where the volume is real — that is
# where a 7-week-old 0.27.0 was caught shadowing 0.35.2 on 2026-09-06.
# EXECUTION is ASSERTED here, not printed. Until 2026-09-07 the version was
# captured inside an echo with 2>/dev/null, so a binary that could not run at all
# still PASSED and simply printed version=[] -- the same failure class as the bare
# `node --version` two hundred lines up: a value displayed rather than compared.
#
# Why this exit code matters more than most: smoke runs `platforms: linux/amd64`
# on an x86 runner, i.e. NATIVE amd64, so this is the fleet's only recurring
# amd64 runtime proof for the linux-x64 ELF. No devbox can supply one -- every
# machine in the pi fleet is an Apple Silicon Mac (mbp-m1-2020; tor-ms22 = Mac
# Studio Mac13,1 M1 Max, verified 2026-08-17 by system_profiler; emb-7kj4vr4g =
# Apple Silicon, 4 routes 2026-09-07). Asking a device for that proof is asking
# for the impossible; CI already had it and was discarding it.
#
# KEEP PROSE OUT OF THE QUOTED BODY BELOW. On 2026-09-07 this explanation lived
# INSIDE the single-quoted argument and contained an apostrophe ("the fleet's").
# Inside '...' bash treats a backslash literally, so \' does not escape -- it
# CLOSES the string. The body silently truncated, the remaining lines were parsed
# by the RUNNER's shell instead of the container's, and `agent-browser --version`
# ran on a host that has no agent-browser: "line 770: command not found", release
# v1.8.14's smoke job failed after the base had already built. shellcheck caught
# it as SC2289 the same day and the red lint job went unread for 24h.
exec_test "agent-browser resolves under /usr (volume-shadowing guard, build-time half)" '
p=$(command -v agent-browser) || { echo "agent-browser not on PATH" >&2; exit 1; }
r=$(readlink -f "$p")
echo "resolved=[$r] version=[$(agent-browser --version 2>/dev/null | head -n1)]" >&2
v=$(agent-browser --version) || { echo "agent-browser did not EXECUTE" >&2; exit 1; }
test -n "$v" || { echo "agent-browser --version produced no output" >&2; exit 1; }
echo "resolved=[$r] version=[$(printf %s "$v" | head -n1)]" >&2
case "$r" in /usr/*) ;; *) exit 1 ;; esac
test ! -d "$HOME/.pi/npm-global/lib/node_modules/agent-browser" || exit 1
echo ok
@@ -776,6 +943,55 @@ exec_test "pi-fork extensions floor is [] (forks cannot write to the palace)" \
exec_test "/tmp/sshcm dir mode 700 (ssh ControlMaster)" \
'test -d /tmp/sshcm && [ "$(stat -c %a /tmp/sshcm)" = "700" ] && echo ok'
# ── Build-time leftovers (npm 11 bloat sentinels) ─────────────────────
# Both of these are worth a PASS/FAIL assertion rather than a size-gate
# diagnostic, because the size gate has ~225 MB of deliberate margin: v1.9.1
# shipped +131 MB of pure build residue and stayed green. These name the
# residue directly, so a regression is legible instead of merely "bigger".
echo ""
echo "── Build-time leftovers ──"
# npm 11 installs EVERY optional platform package of a native dependency, not
# just the host's (it ignores os/cpu, --os/--cpu and npmrc os=/cpu=). Two
# families are affected and pruned in Dockerfile.variant: @esbuild/<platform>
# and @mariozechner/clipboard-<triple>. Keep-set is the host arch only, plus
# clipboard's gnu AND musl (its napi loader picks between them at runtime).
# Runs as root because the image declares no USER; that is also what lets the
# cache assertion below read /root.
run "no foreign platform packages (npm 11 sentinel)" \
'arch=$(node -p process.arch); bad=$(find /usr/lib/node_modules /opt -type d \( -regex ".*/@esbuild/[^/]+" -o -regex ".*/@mariozechner/clipboard-[^/]+" \) ! -name "linux-$arch" ! -name "clipboard-linux-$arch-gnu" ! -name "clipboard-linux-$arch-musl" -prune -print 2>/dev/null); if [ -n "$bad" ]; then echo "foreign platform dirs shipped:" >&2; echo "$bad" >&2; du -sm $bad 2>/dev/null | sort -rn | head -5 >&2; exit 1; fi; echo ok'
# The build's own npm download cache is not free: it lands in the layer that
# created it. v1.9.1 shipped 145 MB of /root/.npm/_cacache (35 MB in v1.8.14)
# — the largest single item in its +131 MB residual, and invisible to the
# size-gate diagnostics because those only looked under node_modules and /opt.
# Nothing at runtime reads it: root's cache, while the container runs as
# `developer`. NOTE the assertion must run as root or a permission error on
# mode-700 /root would make `test ! -d` pass for the wrong reason.
run "no build-time npm cache shipped (/root/.npm)" \
'test "$(id -u)" = "0" || { echo "assertion needs root to read /root" >&2; exit 1; }; if [ -e /root/.npm ]; then echo "/root/.npm shipped: $(du -sm /root/.npm | cut -f1) MB" >&2; exit 1; fi; echo ok'
# The prune's risk is not "too big" but "removed something needed", and only a
# FUNCTIONAL check covers that. These load the natives from every install site
# found in the image, so they also scale to the studio variant's third site.
#
# NOTE THE PATH-QUALIFIED require(). The obvious form, `node -e
# 'require("esbuild")...'`, resolves by walking up from the CURRENT DIRECTORY —
# so it fails with MODULE_NOT_FOUND from /workspace on a perfectly good image,
# because esbuild lives nested inside the pi trees and global installs are not
# on node's require path (NODE_PATH is unset). That exact command was left in a
# runbook as "if this fails, revert the release", and it duly failed for the
# wrong reason on the first machine that ran it. A check must fail only for the
# thing it is checking.
run "esbuild works at every install site (prune removed weight, not function)" \
'sites=$(find /usr/lib/node_modules /opt -type d -path "*/node_modules/esbuild" -prune 2>/dev/null); if [ -z "$sites" ]; then echo "no esbuild install found at all" >&2; exit 1; fi; for d in $sites; do node -e "require(\"$d\").transformSync(\"const x:number=1\",{loader:\"ts\"})" || { echo "esbuild broken at $d" >&2; exit 1; }; done; echo ok'
# Clipboard is the family pruned second, and its napi-rs loader picks its native
# binding at require() time — so a successful load IS the proof that the kept
# platform package is the one this image needs.
run "clipboard native loads at every install site" \
'sites=$(find /usr/lib/node_modules /opt -type d -path "*/node_modules/@mariozechner/clipboard" -prune 2>/dev/null); if [ -z "$sites" ]; then echo "no @mariozechner/clipboard install found at all" >&2; exit 1; fi; for d in $sites; do node -e "var c=require(\"$d\"); if (typeof c.setText !== \"function\") { throw new Error(\"native binding missing\"); }" || { echo "clipboard native broken at $d" >&2; exit 1; }; done; echo ok'
# ── Image size ────────────────────────────────────────────────────────
echo ""
echo "── Image size ──"
@@ -803,6 +1019,32 @@ elif [ "$SIZE_MB" -le "$SIZE_THRESHOLD_MB" ]; then
printf " ✅ size: %d MB (threshold %d MB)\n" "$SIZE_MB" "$SIZE_THRESHOLD_MB"; PASS=$((PASS+1))
else
printf " ❌ size: %d MB exceeds threshold %d MB\n" "$SIZE_MB" "$SIZE_THRESHOLD_MB"; FAIL=$((FAIL+1))
# A bare "too big" verdict cost a full CI-log dig plus a local npm bisect to
# attribute the v1.9.0 overshoot (+431 MB, which turned out to be npm 11
# installing 26 @esbuild platform binaries per pi-coding-agent copy). The
# container already knows where its bytes are, so make it say so: the biggest
# layers, and the biggest directories under the paths that historically grow.
# Same principle as the run() helper above — a red assertion should carry its
# own diagnostic rather than send the next reader spelunking.
echo " ── largest layers (docker history) ──"
docker history --format '{{.Size}}\t{{.CreatedBy}}' "$IMAGE" 2>/dev/null \
| grep -vE '^0B' | head -12 | sed 's/^/ /' | cut -c1-160
echo " ── largest directories in the image ──"
docker run --rm --entrypoint sh "$IMAGE" -c \
'du -sm /usr/lib/node_modules/* /opt/* /usr/local/share/ms-playwright 2>/dev/null | sort -rn | head -12' \
2>/dev/null | sed 's/^/ /' || echo " (could not inspect directories)"
echo " ── build caches that should not be in the image ──"
# v1.9.1's residual was 145 MB of npm cache under /root, and the du list
# above cannot see it: it enumerates node_modules and /opt only. A gate whose
# diagnostic looks only where the bytes were LAST time sends the next reader
# spelunking again, so name the cache paths explicitly.
docker run --rm --entrypoint sh "$IMAGE" -c \
'du -sm /root/.npm /root/.cache /tmp/node-compile-cache /home/developer/.npm 2>/dev/null | sort -rn' \
2>/dev/null | sed 's/^/ /' || true
echo " ── foreign platform dirs (npm 11 regression sentinel) ──"
docker run --rm --entrypoint sh "$IMAGE" -c \
'find /usr/lib/node_modules /opt -type d \( -regex ".*/@esbuild/[^/]+" -o -regex ".*/@mariozechner/clipboard-[^/]+" \) -printf "%f\n" 2>/dev/null | sort | uniq -c | sort -rn | head' \
2>/dev/null | sed 's/^/ /' || true
fi
# ── Summary ───────────────────────────────────────────────────────────