Compare commits

..

19 Commits

Author SHA1 Message Date
Joakim Persson 6c13f43ac8 release: date the v1.9.3 section for the tag
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 16s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 41m49s
Publish Docker Image / smoke (push) Successful in 8m45s
Publish Docker Image / smoke-studio (push) Successful in 9m8s
Publish Docker Image / build-variant (push) Successful in 17m42s
Publish Docker Image / update-description (push) Successful in 11s
Publish Docker Image / promote-base-latest (push) Successful in 26s
Publish Docker Image / build-variant-studio (push) Successful in 21m28s
CI checks out the tag, so the CHANGELOG heading has to name the version
before the tag exists (same convention as v1.9.1/v1.9.2). check-doc-drift
check 9 was sabotage-tested for exactly this rename and still passes:
everything above the published v1.9.2 heading is the pending text.

Contents of v1.9.3 (all measured, nothing inherited):
  - f25efa0  recreate-sanity-check resolves the ControlPath ssh actually
             uses (ssh -G); the ~/.pi/ssh/config hyphen typo deleted
  - b8d818e  Hub size claims corrected; check 8 gates them against full_size
  - c7d369f  promote-base-latest's conditional re-tag measured (run 669)
  - 50153e6  skill floor -> pi-extensions 25c1265 (task tool + fork-gate)
  - ea89605  changelog for the two floating-ref changes: pi-extensions
             25c1265 and mempalace-toolkit 817b3a8 (mine deadline)
  - cb6d9e5  check 9: what the next build bakes differently must be named
             in the CHANGELOG; first run found pi-obsmem 7b397f4->cba0334
  - b9057fd  LABEL se.jordbo.pi-devbox.mempalace-version (inherited)
  - 960aada  dependency audit 2026-09-19; mempalace 3.10.0 deferred

Floating refs this tag will bake, named per check 9: pi-toolkit 9c87ee8,
pi-extensions 25c1265, mempalace-toolkit 817b3a8, pi-observational-memory
cba0334 (3.1.3). Base REBUILDS (Dockerfile.base + rootfs/ changed), so the
BASE_REBUILD_DATE marker is refreshed now, when it is free -- it had been
stale since 2026-07-13 through two rebuilds.

Deliberately NOT in this tag: mempalace 3.10.0 (see the audit section for
the three measured preconditions).
2026-09-19 17:45:50 +02:00
Joakim Persson 960aada769 changelog: dependency audit 2026-09-19 — mempalace 3.10.0 measured and deliberately deferred
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 17s
Lint / skill-floor (push) Successful in 8s
Lint / doc-drift (push) Successful in 13s
Every row by direct command: npm/PyPI/GitHub release APIs, git ls-remote,
the v1.9.2 config labels via the registry API, and binaries in a running
v1.9.2 container for the floating tools. Only pin with a newer upstream is
mempalace 3.9.0 -> 3.10.0, and it is NOT a small bump: new installs move to
~/.config/mempalace (entrypoint-user.sh tests ~/.mempalace/palace and
entrypoint.sh persists ~/.mempalace, so a fresh container would initialise
outside the volume), and mempalace_event_list flips to newest-first by
default (the toolkit's deriveClosed passes no order, deriveOwed on 2 of 3
calls — server-default-dependent). Deferred to its own release with the
three measured preconditions listed.
2026-09-19 17:38:02 +02:00
Joakim Persson b9057fdc8c Dockerfile.base: LABEL se.jordbo.pi-devbox.mempalace-version, inherited by both variants
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 10s
Lint / doc-drift (push) Successful in 16s
Lint / actionlint (push) Successful in 21s
Closes the blind spot check 9 (cb6d9e5) named: no label recorded the palace
pin, so a MEMPALACE_VERSION bump — the one component whose skew against the
shared central palace is fleet-wide — could ship without a CHANGELOG line.

In Dockerfile.base, not Dockerfile.variant, deliberately: the value sits next
to the ARG that defines it (a copy in the variant is one more pin able to
drift); labels are inherited by every image built FROM the base, so no
build-arg to plumb through the variant's four call sites; and inheritance
means the label states the pin of the base the image ACTUALLY built on,
which is the question when base-decide cache-hits an older base. Both
mechanisms measured on the published v1.9.2 config blob rather than assumed:
maintainer + image.source (set only in Dockerfile.base) are present on the
variant image, and pi-version=0.85.1 is an ARG expanded inside a LABEL.

Intent, like every se.jordbo.pi-devbox.* label; the manifest's
mempalace_version (read from the installed binary) stays the ground truth,
and smoke-test.sh now asserts label == installed core — the one way they
diverge is a base built with INSTALL_MEMPALACE=false or an off-pin install,
both invisible to a label-only check.

check-doc-drift check 9 gains the component (literal, against
ARG MEMPALACE_VERSION in Dockerfile.base); the label-key rule generalises to
"names ending in -version are the label itself". Until a release carries the
label it reports a counted SKIP, not OK — measured: "v1.9.2 carries no
se.jordbo.pi-devbox.mempalace-version label", summary says 1 SKIPPED.

Costs nothing extra: this Unreleased already forces a base rebuild
(50153e6 rootfs/ skill floor). check-base-hash unchanged (no new *_REF).
2026-09-19 17:27:34 +02:00
Joakim Persson cb6d9e5dd0 check-doc-drift: check 9 — what the next build bakes differently must be named in the CHANGELOG
Lint / hadolint (push) Successful in 8s
Lint / skill-floor (push) Successful in 10s
Lint / doc-drift (push) Successful in 14s
Lint / actionlint (push) Successful in 19s
The two most recent Unreleased entries (pi-extensions 25c1265: task tool +
fork-gate; mempalace-toolkit 817b3a8: the mine deadline that never reached
the transport) reached this image through floating *_REF=main ARGs. Neither
produced a diff here, so nothing asked for a CHANGELOG line, and neither had
one until a reader asked. Same class as check 8 — a claim with no in-repo
anchor rots — and the hand practice that covered it (the per-release
"Dependency audit" table, Baked-in vs Upstream-now) is a someone-remembers
mechanism that had lapsed.

Method, no docker/crane/token: last published vX.Y.Z from Hub's tag list
(one request, now shared with check 8); its amd64 config blob via the
anonymous registry API carries one se.jordbo.pi-devbox.<name>-ref label per
component with the SHA the build-args baked. "Would bake" is resolved the way
resolve-versions does it: 40-hex ARG is itself, branch/tag via git ls-remote
with the peeled ^{} form preferred, pi-studio = highest semver tag (label on
<tag>-studio), PI_VERSION as a literal against the pi-version label. Nine
components, ~7.5 s. Unchanged: OK. Moved: the new value's 7-char prefix (tag
for studio, version for pi) must appear above the last published version's
"## " heading — Unreleased plus any not-yet-published section, which is what
the release commit turns Unreleased into, so the tag build passes on the same
text. A published tag with no heading is a failure, not a skip.

First run found a move nobody had recorded: pi-observational-memory 7b397f4
-> cba0334 (6 commits, 3.1.1 -> 3.1.3; memory workers now routed by the
model's exact provider instead of the first registered provider with a
matching api — upstream #70). Named in the CHANGELOG with the honest scope:
no expected effect on the shipped bedrock-only configuration, unmeasured.

Sabotage-tested, expectation written before each run:
  obsmem SHA removed from CHANGELOG ............ rc=1  DRIFT
  Unreleased renamed to "## v1.9.3 — …" ........ rc=0  (release-commit shape)
  "## v1.9.2" heading mangled to v1.9.2-typo .... rc=1  — this one FAILED first:
      \b accepted "v1.9.2-typo" as v1.9.2's heading; now (\s|$)
  bare "## v1.9.2" (no date) ................... rc=0
  ARG PI_OBSMEM_REPO renamed ................... rc=2  (blind gate must not pass;
      read_arg calls are bare top-level assignments on purpose — nested in a
      heredoc's $(...) the exit 2 is swallowed by cat)
  offline (proxy to a closed port) ............. rc=0, 2 SKIPs counted
  SKIP_REF_CHECK=1 ............................. rc=0, 1 SKIP counted
  restored ..................................... rc=0, files byte-identical
GIT_TERMINAL_PROMPT=0 on ls-remote: a repo flipped private fails in 0.3 s as
a SKIP instead of waiting for a username until the job times out.

Header, lint.yml job comment and AGENTS.md all still said this gate "needs
no network" — false since check 8 (2026-09-14). Now: two classes, hermetic
1-7 and published-state 8-9 that SKIP loudly and counted when offline.
Considered and not added: compose <-> .env.example variable cross-check (the
four mismatches are commented-out lines, the mempalace-server compose file's
own variables, and entrypoint-consumed ones — it would fire on nothing
wrong); "documented tag exists on Hub" (check 8 already SKIPs by name; a
hard fail would misreport the tag-to-publish window).
2026-09-19 17:15:12 +02:00
Joakim Persson ea896054df changelog: task tool + fork-gate (pi-extensions 25c1265), and the mine deadline that never reached the transport (mempalace-toolkit 817b3a8)
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 18s
Lint / doc-drift (push) Successful in 6s
Both land in the image through floating refs (PI_EXTENSIONS_REF=main,
MEMPALACE_TOOLKIT_REF=main), so neither shows up as a diff in this repo — the
changelog is the only place a reader of a pi-devbox tag learns that the
container's delegation and feed behaviour changed. Recorded under Unreleased
with the measurements: 5/5 real fork briefs redirected, live pi -p proof for
gate and tool, 2-of-6 → 6-of-6 on the toolkit's new timeout test.

Dockerfile.base: the "Stall protection" comment listed MEMPALACE_MCP_TIMEOUT_MS
as THE tool-call deadline; it now says the feed's mine carries its own.
Comment-only, but Dockerfile.base is hashed into base_tag; the rebuild was
already forced by 50153e6 (rootfs/ skill floor).
2026-09-19 17:00:59 +02:00
Joakim Persson 50153e65b7 skill floor: refresh vendored pi-extensions skill to pi-extensions@25c1265 (task tool + fork-gate)
Lint / hadolint (push) Successful in 11s
Lint / doc-drift (push) Successful in 10s
Lint / skill-floor (push) Successful in 10s
Lint / actionlint (push) Successful in 21s
check-skill-floor.sh: OK, tree_sha256 9b85a633… matches the package at main (25c1265).
Forces one base rebuild (base_tag hashes rootfs/), which is also what bakes
extensions/task.ts and fork-gate.ts into /opt/pi-extensions via the floating
PI_EXTENSIONS_REF=main; install.sh symlinks every extensions/*.ts on start.
2026-09-19 16:42:06 +02:00
Joakim Persson f25efa074d fix(sanity): verify the ControlPath ssh RESOLVES, not the dir the image creates
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 13s
Lint / doc-drift (push) Successful in 12s
Lint / actionlint (push) Successful in 24s
recreate-sanity-check.sh asserted `/tmp/sshcm exists with mode 700` and printed
a green tick while every ssh in the container died rc=255. It was right about
what it checked: the breakage was in the directory a CONFIG named, not the one
the image creates, and the old check could not see the disagreement between them.

Found on the v1.9.2 first boot on emb-7kj4vr4g. A durable ~/.pi/ssh/config,
hand-written into the ~/.pi named volume by the previous session so it would
survive the recreate, declared `ControlPath /tmp/ssh-cm/%C` — with a hyphen.
Nothing here creates that path; the canonical spelling is /tmp/sshcm, identical
in Dockerfile.base, entrypoint-user.sh, recreate-sanity-check.sh and
smoke-test.sh. ControlMaster auto with an unusable ControlPath does not degrade
to an unmultiplexed connection — it fails hard:

    unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory

rc=255, remote command never runs.

Now: resolve the ControlPath ssh itself would use via `ssh -G` and require its
parent to exist and be writable. -G applies real config precedence (first-value-
wins, the system drop-in, Include, an -F override), so it answers "which rule
captured this host" instead of re-implementing the guess, and it never opens a
connection — 0.116s for 48 hosts.

This also puts a check under a caveat documented in prose in Dockerfile.base
("SSH client defaults") and verified nowhere: a per-host ControlPath under a
read-only bind-mounted ~/.ssh gives the identical failure, "Read-only file
system". On the machine this was built on that is 16 of 50 hosts — freeipa-1..6,
gitea.egl.lan, runner-1..3, tor-ms22 — never reported by anything before.

Severity split is deliberate. Default ssh precedence legitimately lands in the
read-only ~/.ssh for any host whose own config pins it there, and the supported
workaround (ssh -F ~/.ssh-local/config, generated every start by
setup-lan-access.sh) already exists, so that is a warn. Failing it would paint
the script red on every run of every device, and a check that fires benignly
every time is one you learn to ignore — the same reasoning that keeps
lint-shell.sh at -S error. The sidecar route is prescribed, so there it is a
hard fail. Host lists cap at six names plus a count: unreadable output is
ignored output.

Teeth proven both directions, each sabotage confirmed by diff BEFORE the result
was believed: ControlPath -> nonexistent dir => rc=1; -> existing-but-read-only
dir => rc=1; restored => rc=0 with the sidecar byte-identical.

No config was added to fix the original problem — the fix was to DELETE
~/.pi/ssh/config, a third hand-maintained copy of what setup-lan-access.sh
already generates from version control on every container start, with fewer
features and one typo. The host-owned ~/.ssh/config is left alone on purpose:
those ~/.ssh/cm paths are correct on the host, where ~/.ssh is writable.

Also lint-shell.sh: the SC2088 count in the severity-choice rationale said 19
and is now 20, with the reproduce command recorded. That command needs a `$ `
prefix — a comment whose first word is "shellcheck" is parsed as a directive,
and the malformed one tripped SC1072/SC1073 at severity error. The gate caught
it on the very commit that introduced it.
2026-09-15 22:59:43 +02:00
joakimp c7d369f28d docs(changelog): measure promote-base-latest's conditional re-tag
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 20s
Closes the open question from v1.9.2's release verification. crane copy DOES
bump Hub's last_updated when it runs (run 669: digests differed, copy ran
17:10:04->17:10:08, base-latest.last_updated moved to 17:10). The
identical-digest path stays unproven by construction -- the job prints
'base-latest already current; nothing to promote.' and never copies -- so a
freshness assertion on the alias is conditional on base-decide's need_build,
not a blanket upgrade. Operational rule lives in the ci-release-watcher skill
(skillset c8034af); the pipeline property is recorded here.
2026-09-14 20:49:23 +02:00
joakimp b8d818ed99 fix(docs): correct the Hub size claims, and gate them so they cannot rot again
Lint / hadolint (push) Successful in 9s
Lint / doc-drift (push) Successful in 13s
Lint / skill-floor (push) Successful in 15s
Lint / actionlint (push) Successful in 18s
DOCKER_HUB.md claimed ~1.1 GB for :latest while Docker Hub served 1.37 GB --
20% wrong, drifting quietly across eight releases, and POSTed to Docker Hub by
update-description every time. It is the first number a stranger reads about
this image.

Root cause is structural, not carelessness: every OTHER claim check-doc-drift.sh
guards is anchored to a build file (a pin, an ARG, a placeholder), so it cannot
rot without someone editing the thing it describes. Nothing in this repo states
the image size, so nothing measured it.

Corrected against measured full_size after v1.9.2 published (amd64/arm64):
  :latest         ~1.1  -> ~1.23 GB   (1.228 / 1.211)
  :latest-studio  ~1.15 -> ~1.25 GB   (1.255 / 1.238)
  :base-latest    ~1.0  -> ~1.17 GB   (1.167 / 1.151)
full_size tracks the FIRST manifest entry (amd64), not the sum across arches --
measured: full_size=1.228, amd64=1.228, arm64=1.211, sum=2.439 -- which is what
the table's per-arch "Size (compressed)" column claims.

New check 8 gates those claims against Hub. Fails on drift; SKIPS LOUDLY via a
new skip() helper (counted, named in the summary) when curl/python3 are absent,
the API is unreachable, or SKIP_SIZE_CHECK=1. A skip is deliberately neither OK
nor a failure: printing an unverified claim as OK is the habit this file exists
to break, and failing on Docker Hub's uptime would hold releases hostage to a
third party. Not wired into hooks/pre-push, so pushes stay offline.

Two bugs caught by writing the expected exit code down before running the check:

  1. The tolerance would have missed its own motivating case. The percentage is
     computed against the MEASURED size, but the 20% first chosen came from the
     claim-relative figure; the real drift was |1.1-1.37|/1.37 = 19.7% and would
     have passed. Now 15%, inside a window with both bounds measured: above the
     largest legitimate skew (11.4%, a claim describing the published release
     while the next tag changes the size) and below the rot it must catch.

  2. A `|| true` on the python invocation made the gate FAIL OPEN -- it printed
     DRIFT and exited 0. Removed; the outer `|| SIZE_RC=$?` satisfies set -e
     without swallowing the code. Verified two-sided: 19% drift exits 1 at the
     default tolerance and 0 at SIZE_TOLERANCE_PCT=25, so the threshold does the
     work rather than the ordering.

NOT covered, and the script says so where a reader will see it: README.md's
~3.2 GB figures are UNCOMPRESSED and the registry exposes compressed sizes only,
so measuring them needs a real pull. A green check 8 says nothing about them.

Verified: check-doc-drift.sh rc=0 on the corrected tree, rc=1 on injected drift,
rc=0 with --warn-only, SKIP+rc=0 on the offline path (proxy to a dead port);
shellcheck clean at default severity (the sed single-quote SC2016 carries a
disable with its reason); lint-shell.sh 16 files clean.
2026-09-14 20:33:40 +02:00
joakimp f5c53b8693 release: date the v1.9.2 section for the tag
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 27s
Publish Docker Image / base-decide (push) Successful in 13s
Publish Docker Image / build-base (push) Successful in 43m59s
Publish Docker Image / smoke (push) Successful in 5m29s
Publish Docker Image / smoke-studio (push) Successful in 5m50s
Publish Docker Image / build-variant (push) Successful in 17m29s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / update-description (push) Successful in 10s
Publish Docker Image / build-variant-studio (push) Successful in 21m50s
CI checks out the tag, so the CHANGELOG heading has to name the version
before the tag exists, not after. v1.9.1's tagged tree carried no
"## Unreleased" either -- same convention, made explicit here.

Contents of v1.9.2 (all measured, nothing inherited):
  - 1baba79  the +131 MB v1.9.1 residual: /root/.npm (110 MB) plus
             @mariozechner/clipboard-* foreign natives at BOTH install
             sites (global + /opt/pi-fork's nested pi-coding-agent copy)
  - 852f900  smoke sentinels that name the residue instead of letting
             the size gate's ~225 MB margin swallow it
  - dab989b  mempalace-toolkit: the owed-withdrawal suite gate
             (node --check never type-strips)
  - 9aaff26  vendored mempalace skill snapshot -> e9e45f7, which is what
             turns the currently-red canary green

Deliberately NOT in this tag: DOCKER_HUB.md's "~1.1 GB" size claim is
low (Hub says v1.9.1 = 1.37 GB, v1.8.14 = 1.23 GB). This release
*changes* the size, so no number is right both before and after it --
writing the predicted ~1.24 GB would be an unmeasured claim in a
user-visible page. Measure post-build, fix on the next tag, and extend
check-doc-drift.sh to gate size claims against Hub's full_size so the
number cannot rot silently again.
2026-09-14 17:55:16 +02:00
joakimp 735565b9be docs: retire the deployed-and-unproven label on isWithdrawn, and name dab989b
Lint / skill-floor (push) Successful in 8s
Lint / hadolint (push) Successful in 17s
Lint / doc-drift (push) Successful in 7s
Lint / actionlint (push) Successful in 20s
Two things, and the second was found by doing the first.

v1.9.1's notes closed with an explicit open item: isWithdrawn had 17 assertions
and 4 mutation kills but had never been exercised on a released image against the
live logstream, which is a different state from untested and a worse one to leave
unlabelled. It has now been exercised, from a container recreated onto v1.9.1, and
the label is retired with the measurement rather than with an assurance.

What the entry records is the instrument as much as the result, because the result
is only as good as the thing that produced it: the shipped extractor was lifted
out of the toolkit's own test script and used to cut the four predicates out of
the BAKED mempalace.ts, and deriveOwed's three queries were replayed with their
exact shipped parameters. No second copy of the logic was written, which is the
same discipline the suite itself is built on. Baseline owed was established by
three independent routes BEFORE any probe was planted, because without a recorded
baseline a withdrawal that appears to work proves nothing. Predictions were
written down before each measurement; the table reports both columns.

The per-predicate attribution is in the entry deliberately. "The count went back
to baseline" is a claim about arithmetic and would have been satisfied by a
pre-filter or by an unrelated reply; isAnswered=false with isWithdrawn=true on a
candidate that was still in the raw set is a claim about mechanism. The negative
controls carry the same weight as the positive one -- a rule that lets a third
party or a prose sentence empty someone's mailbox would have passed a
count-only check.

Also recorded: the rule fires on the real incident it was written for (seq 112),
which is worth more than any synthetic probe, and the one path still unwitnessed
(the extension's own in-process poll, which only fires when the agent is idle) is
named as unwitnessed rather than folded into the pass.

Second: reaching for the suite as corroboration showed it cannot run on this image
at all -- its own precondition gate rejects the source it exists to check, because
`node --check` does not type-strip on any node and v1.9.1 moved to node 24. All 17
assertions had been skipped since the bump. Fixed in toolkit dab989b, named here
per the rule this repo adopted two commits ago after the withdrawal fix shipped
unnamed in v1.9.1. The gate now also distinguishes "cannot run" from "does not
parse", since collapsing those is how the defect pointed readers at the wrong
file.

The rebuild cost is stated and, having been measured, is smaller than the earlier
draft of this entry claimed: 9aaff26 already moved base_tag by refreshing the
vendored skill snapshot under rootfs/, so the base rebuild was pending before this
fix existed and the toolkit pickup rides along with it.
2026-09-14 16:29:52 +02:00
joakimp 9aaff26e3a docs: the withdrawal fix shipped in v1.9.1 unnamed — say so, and re-pin the canary
Lint / hadolint (push) Successful in 11s
Lint / doc-drift (push) Successful in 12s
Lint / skill-floor (push) Successful in 12s
Lint / actionlint (push) Successful in 16s
v1.9.1 bakes mempalace-toolkit e68ee20, which contains e2b060a. Requester-side
ask withdrawal has therefore been LIVE on every v1.9.1 device since 2026-09-10,
while the v1.9.0 section of CHANGELOG.md still read "not yet pinned ... this
image still pins e45f6b4", and the mempalace skill still told every agent at
session start that a withdrawal is impossible.

Measured, two independent routes, expectation recorded before looking:

  - published image label, :v1.9.1-studio and :latest-studio (same digest):
      se.jordbo.pi-devbox.mempalace-toolkit-ref = e68ee2071ca3ad39...
  - ancestry: e2b060a is an ancestor of e68ee20
  - baked mempalace.ts sha 7c16fe14 != v1.8.14's dfca71e9 (different bytes)
  - grep -c isWithdrawn on this container's baked copy = 0 (v1.8.14, so
    tor-ms22 cannot exercise the behaviour it is documenting)

The first label read came back EMPTY, and that empty was a claim about the
request rather than the image: Hub redirects blob fetches to a CDN and curl
without -L returns 0 bytes at exit 0. Recorded in the notes, because a registry
audit reporting "no labels" is missing -L until proven otherwise.

Changes:

  - CHANGELOG Unreleased: the floating-ref mechanism, which is the reusable
    part. ARG MEMPALACE_TOOLKIT_REF=main + CI resolving it to a SHA at build
    time means a release absorbs whatever toolkit main holds, and "what
    behaviour did this image gain" is a question nobody is forced to answer.
    The fleet rule (name the pickup before tagging) was honoured for the
    feed-tick commits v1.9.1 names, and missed for one in the same range.
  - CHANGELOG v1.9.0: the stale caveat is ANNOTATED, not rewritten. The
    sentence is the evidence for how the drift happened; deleting it would
    destroy the only trace. Same discipline as fleet-ops d42412c.
  - CHANGELOG Unreleased: corrected the size entry's "needs no base rebuild"
    claim. True of that entry alone, false of the release now that the
    vendored skill snapshot -- a base_tag input -- moves with it (~67 min).
  - rootfs mempalace skill snapshot 44472 -> 46045 B, refreshed via
    scripts/vendor-mempalace-skill.sh so the bytes and SKILLSET_SNAPSHOT_REF
    (4d7c0ea -> e9e45f7) move together; --check confirms exact match.
  - scripts/smoke-test.sh canary RE-PINNED. The retired pair was still green
    against the new snapshot, i.e. blind to this refresh for the same reason
    the pre-v1.8.13 pair was blind to that one. The replacement is stronger
    than any predecessor here because BOTH witnesses come from the same
    upstream commit: e9e45f7 added "Withdrawing an ask you sent" and deleted
    "nothing anyone can do about it from the other end", the sentence the new
    bullet contradicts. Directions measured against both files, then the
    canary body EXECUTED against each: new -> rc=0 "ok", old -> rc=1 empty.
    A canary whose negative witness was removed by the commit it pins fails
    loudly on stale bytes instead of merely failing to notice them.

Gates: check-doc-drift OK, check-skill-floor OK, vendor --check OK,
hooks/pre-push OK (16 shell files clean at severity error).

Not fixable here: isWithdrawn is deployed and UNPROVEN. 17 assertions and 4
mutation kills, never once exercised on a released image against the live
logstream. Ask routed to a v1.9.1 device.
2026-09-14 15:07:26 +02:00
Joakim Persson 852f900b53 test(smoke): make CI prove the natives still work — the runbook check didn't
Lint / doc-drift (push) Successful in 6s
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 9s
Lint / actionlint (push) Successful in 18m11s
The post-boot check v1.9.1 left for the next machine was

  node -e 'require("esbuild").transformSync("const x:number=1",{loader:"ts"})'

with "if this fails, the prune removed something needed on arm64 -> revert to
v1.8.14 and fix" attached. Run from /workspace on the freshly recreated v1.9.1
image it fails with MODULE_NOT_FOUND — on a perfectly healthy image. require()
resolves by walking up from the CURRENT DIRECTORY, esbuild is nested inside the
two pi trees, and a global install is not on node's require path (NODE_PATH is
unset). The version that "verified the prune" last night only passed because the
shell happened to sit inside the tree. So the check was CWD-dependent, its red
meant nothing, and the remedy it prescribed was reverting a good release.

Path-qualified, both sites compile TS today. Rather than leave a fragile command
in a note for a human to paste, CI now owns it:

  - esbuild must transformSync at EVERY install site found in the image
  - @mariozechner/clipboard must load with its native binding attached at every
    site — for the clipboard prune that IS the proof, since napi-rs resolves the
    platform package at require() time

Sites are discovered with find, so the studio variant's third site is covered
without naming it.

Verified as a four-way matrix, not just a green run: both assertions pass
against the real image FROM /workspace (the cwd that broke the old command), and
both fail (exit 1, "The package \"@esbuild/linux-arm64\" could not be found, and
is needed by esbuild") against copies of the same packages with the host
platform binary removed. An assertion never observed failing is not evidence.

Also documents both failure shapes in AGENTS.md next to the size gate: this one,
and the mode-700 /root one where `test ! -d` passes on a permission error.
2026-09-11 11:06:18 +02:00
Joakim Persson 1baba79c96 fix(size): the v1.9.1 residual was npm's own cache, not the platform binaries
Lint / doc-drift (push) Successful in 6s
Lint / hadolint (push) Successful in 10s
Lint / skill-floor (push) Successful in 9s
Lint / actionlint (push) Successful in 4m26s
v1.9.1's @esbuild prune fixed the 431 MB size-gate failure but still shipped
+131 MB compressed over v1.8.14, nearly all of it in the pi/extensions install
layer (87 -> 206 MB). That leftover was filed as an open item with an explicit
hypothesis — the @mariozechner/clipboard-* family, same npm 11 behaviour, a
different package — and an explicit warning that the hypothesis was not a
measured cause. Measured now, after recreating onto v1.9.1, and the hypothesis
accounted for one sixth of it:

  +110 MB  /root/.npm/_cacache  (35.2 -> 145.3 MB)   the build's npm cache
  + 21 MB  clipboard foreign platform packages, both install sites
  = 131 MB  i.e. the whole delta, no unexplained remainder

Method, since there is no docker CLI inside the container: pulled both variant
layer blobs straight from the registry with a token + manifest + blob fetch and
listed the tarballs (29 789 vs 29 889 entries, 270.8 vs 401.9 MB uncompressed),
then aggregated per package. The file COUNT barely moved, which is what said
"few large files", not "npm installed more packages".

  - purge_build_caches: npm cache clean --force + rm -rf /root/.npm, in the SAME
    layer as the installs, in both the main RUN and the studio RUN. npm 11
    caches every platform tarball it downloads, including the ones the prune
    then deletes, so the cache grew faster than the tree. Nothing at runtime
    reads it: build is root, container is developer with its own cache in $HOME.
  - prune_foreign_esbuild -> prune_foreign_natives: now covers both MEASURED
    families. Clipboard keeps linux-$arch-gnu AND -musl because its napi-rs
    loader picks between them at runtime via its own isMusl() probe; the musl
    package is a 420-byte stub. The bare @mariozechner/clipboard wrapper has no
    hyphen suffix and cannot match the pattern.

Verified on arm64 before writing the glob — a widened rm -rf against a tree you
cannot inspect is the one change shape not to write blind, which is why the
order was update-then-patch. Exercised against a copy of the real trees with
foreign dirs fabricated back in (aix-ppc64, android-arm64, darwin-arm64,
win32-x64, linux-x64): all removed, host linux-arm64 kept at both sites, 21 MB
freed, require('@mariozechner/clipboard') still loads and exports all 18
functions, esbuild.transformSync still compiles TS at both sites. v1.9.1's own
arm64 validation also passed here; CI could only smoke amd64.

Two sentinel assertions, because the size gate did not catch this: it has
~225 MB of margin, so 131 MB of residue stayed green. Both were verified RED
against the running v1.9.1 image and GREEN against a pruned tree. The cache one
refuses to run as non-root: test ! -d /root/.npm on mode-700 /root would
otherwise pass for the wrong reason. Size-failure diagnostics now list cache
paths too — they previously enumerated only node_modules and /opt, where these
bytes were not.

Also fixed: -printf '%f\\n' reaches the shell with both backslashes (confirmed
from the published image's recorded created_by), so v1.9.1's progress line
printed a mangled "li ux-arm64" — find emitted a literal backslash and tr ate
the n out of the name. Single backslash now.

Deliberately not purged: /tmp/node-compile-cache (1.3 MB). The manifest RUN
calls pi --version again, so deleting it earlier only relocates those bytes into
that layer — today's manifest layer is 128 kB precisely because it finds the
cache warm.

Dockerfile.base untouched, so no base rebuild: this rides the next release.
2026-09-11 09:39:54 +02:00
Joakim Persson 42bd29d654 fix(ci): unblock the release — npm 11 esbuild bloat + two self-inflicted assertions
Lint / hadolint (push) Successful in 8s
Lint / skill-floor (push) Successful in 9s
Lint / doc-drift (push) Successful in 10s
Lint / actionlint (push) Successful in 19s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 14s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Successful in 5m4s
Publish Docker Image / smoke-studio (push) Successful in 13m17s
Publish Docker Image / build-variant (push) Successful in 33m49s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 14s
Publish Docker Image / build-variant-studio (push) Successful in 30m58s
v1.9.0 was tagged but never published: smoke failed 90-passed/3-failed and
build-variant needs smoke, so nothing reached the registry. All three are fixed.

1. Size, 431 MB over threshold. Node 24 brings npm 11, which installs EVERY
   @esbuild/<platform> optional binary instead of the matching one: 26 dirs,
   284 MB per pi-coding-agent copy. Measured on pi-fork: npm 10.9.8 -> 165 MB
   (exactly what v1.8.14 shipped), npm 11.19.0 -> 449 MB. npm 11 ignores the
   os/cpu constraints AND --os/--cpu AND an npmrc carrying them, all measured,
   so Dockerfile.variant prunes explicitly, keeping linux-$(node -p
   process.arch) so one line is right on both arches. Pruned in the SAME layer
   as each install, or the bytes survive in the earlier layer. Three sites:
   global pi, pi-fork, pi-studio. Verified esbuild still transforms TS after.

2. om's node_modules assertion tested an npm artefact. om has zero runtime deps;
   its node_modules held ONE file (.package-lock.json) and 20 empty scope dirs.
   npm 11 stopped creating it. Now asserts the entry point pi actually loads,
   read from package.json -> pi.extensions.

3. The skill-source annotation added in v1.9.0 ("baked (package copy)") broke
   the assertion matching "baked$". Pattern now allows an optional suffix; which
   copy shipped stays authoritatively asserted against the manifest + tree hash.

Also: a failed size check now prints the largest layers, largest directories and
an @esbuild sentinel, so this class attributes itself next time instead of
costing a CI dig plus a local npm bisect.

Threshold stays 3800 MB: it caught a real regression and raising it would have
thrown the signal away. No Dockerfile.base/rootfs change, so base-0fb1256c7f99
is reused and build-base is skipped.
2026-09-10 23:39:40 +02:00
Joakim Persson 3a44e81cad feat(ci): gate documentation drift, and make docs a pre-tag release step
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Five doc claims had rotted by v1.9.0, all the same shape: a value written once
by hand, in a file nothing verifies, about a number that lives elsewhere and
moved. README pin table wrong on all three rows; a "Planned" section describing
something already shipped; DOCKER_HUB.md claiming Node v22 against Node 24.

DOCKER_HUB.md is why this is a gate and not a resolution to be careful: it is
PUBLISHED (update-description POSTs it as Docker Hub full_description on every
tag), it had gone eight releases untouched, nothing generates it, and it is read
from the TAG -- so the stale page shipped with v1.9.0 regardless.

scripts/check-doc-drift.sh: seven checks, all repo-local (no network, token,
image, or sibling clone). Exit 0/1/2 matching lint-shell.sh; a renamed ARG is a
red 2, not a green tick. Wired as a fourth lint.yml job so "the docs lie" is its
own red name.

Verified with 15 controls, including two false-positive controls: the first
placeholder check flagged README.md:900, a Go template in a legitimate
`docker inspect --format` example. The gate was wrong, not the doc, so the
pattern is now anchored to the UPPER_SNAKE convention CI substitutes.

Not gated, deliberately: counts/sizes needing a running image (they belong in
smoke-test.sh -- a guessing gate is worse than none), and Dockerfile.base
BASE_REBUILD_DATE, because base_tag hashes that file content-wise and demanding
it be current would force a ~60 min rebuild on releases that touch no base
files. Free during a rebuild, expensive otherwise.

AGENTS.md step 3 rewritten around the mechanism: checkout@v4 with no ref: means
every job reads github.ref, the tag. Docs must be right BEFORE tagging.
2026-09-10 22:01:49 +02:00
Joakim Persson 35964abd01 docs(hub): the Docker Hub page claimed Node v22; v1.9.0 ships Node 24
DOCKER_HUB.md is hand-maintained (no generator: CI only substitutes
{{PI_VERSION}} and POSTs the file as Docker Hub full_description), and it was
last touched at v1.8.6 -- eight releases ago. The Node claim is now false as of
this release, and unlike README.md this file IS published, so a stale claim here
is user-visible rather than internal.

Verified countable claims rather than assuming: "7 user-facing extensions" is
correct (7 files in pi-extensions/extensions/). The "29 mempalace_* tools" claim
is suspect -- this session sees 45 -- but left alone because I cannot attribute
that count to the baked 3.9.0 server without measuring it.
2026-09-10 21:43:17 +02:00
Joakim Persson 6353d59e63 docs(readme): correct the three stale version pins and the shipped-as-planned claim
The version-pin table existed precisely to be the reviewable record of what is
deliberately frozen, and it was wrong on every row: pi 0.84.4 -> 0.85.1,
pi-atelier v0.10.0 -> v0.10.1, mempalace 3.8.0 -> 3.9.0. Verified by parsing the
table and comparing against the ARGs it names rather than by eye.

Also: the "Planned for an upcoming minor release" section listed typst PDF
export, which shipped long ago and even carried a self-contradicting
"(shipped in Unreleased/base)" marker -- the fourth instance of the stale
in-repo Unreleased-pointer class this CHANGELOG already documents. typst 0.15.1
confirmed live in the running image, so the item is now stated as current fact.

The pi-devbox-version sample was v1.5.0-era and structurally outdated: it
predates the palace line the surrounding prose advertises, the pi-atelier
component, and the whole skills: block. Replaced with real observed output
rather than hand-written text.

README has no CI coupling (no workflow or gate reads it; the Docker Hub page
comes from DOCKER_HUB.md), so this cannot affect the in-flight v1.9.0 build.
2026-09-10 21:33:07 +02:00
Joakim Persson 8f0960e134 release: v1.9.0
Lint / skill-floor (push) Successful in 8s
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / lint-gate (push) Successful in 22s
Publish Docker Image / resolve-versions (push) Successful in 15s
Publish Docker Image / base-decide (push) Successful in 8s
Publish Docker Image / build-base (push) Successful in 42m19s
Publish Docker Image / smoke-studio (push) Failing after 6m17s
Publish Docker Image / build-variant-studio (push) Has been skipped
Publish Docker Image / smoke (push) Failing after 9m7s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Rename the Unreleased changelog section to its release heading, matching the
established `## vX.Y.Z — YYYY-MM-DD` format (em dash), per AGENTS.md release
step 3.

Minor rather than patch: CHANGELOG.md scopes minor to "significant base
additions", and this release bumps the Node runtime under every baked JS tool
(pi, agent-browser, playwright, mempalace) from 22 to 24 — the first
NODE_VERSION change since it was introduced at v1.0.0 — alongside five new
base packages (shellcheck, bind9-dnsutils, ldap-utils, xxd, python3-yaml).
Package additions alone have been patch here (v1.8.12, v1.8.13); the runtime
major is what lifts this one.
2026-09-10 21:21:09 +02:00
13 changed files with 2122 additions and 61 deletions
+50
View File
@@ -172,3 +172,53 @@ jobs:
- name: Vendored pi-extensions skill floor matches the package
run: bash scripts/check-skill-floor.sh
doc-drift:
# Gate hand-maintained doc claims against the build files they describe.
# Its own job for the same reason as skill-floor: "the docs lie" should be a
# distinct red name, not a line buried in a job about workflow syntax.
#
# The gap it closes, measured 2026-09-10 while preparing v1.9.0 — five
# claims had rotted, every one of them a fact written by hand in a file
# nothing verified:
# * README.md's "Version pins" table was wrong on ALL THREE rows (pi
# 0.84.4 vs 0.85.1, pi-atelier v0.10.0 vs v0.10.1, mempalace 3.8.0 vs
# 3.9.0) — and that table exists specifically to be the reviewable
# record of what the repo freezes on purpose, so a wrong row destroys
# the only thing it is for.
# * README.md listed already-shipped typst PDF export under "Planned for
# an upcoming minor release", marked "(shipped in Unreleased/base)".
# * DOCKER_HUB.md claimed "Node.js v22" while v1.9.0 ships Node 24.
#
# DOCKER_HUB.md is why this is a gate and not a habit. It is PUBLISHED —
# update-description POSTs it to Docker Hub as full_description on every tag
# — and it had gone eight releases (v1.8.6 -> v1.9.0) untouched. Nothing
# generates it and nothing checked it, so the only thing keeping it true was
# someone remembering. It is also read from the TAG, so a fix pushed to main
# after tagging never reaches the published page.
#
# Two classes of check. 1-7 are hermetic: each compares a doc string against
# a value that exists in this repo — no network, no token, no built image.
# 8-9 compare against what is PUBLISHED, because those claims have no
# in-repo anchor and rotted for exactly that reason: 8 reads Docker Hub's
# measured sizes; 9 reads the ref labels baked into the last released image
# (anonymous registry API, no docker/crane) and `git ls-remote`s each
# floating upstream, then requires every component the next build would
# bake differently to be NAMED in the CHANGELOG above that release's
# heading. Both SKIP loudly and counted when offline — a skip is neither OK
# nor a failure. Claims that genuinely need a running container (the "N
# mempalace_* tools" count, uncompressed sizes) are still left out — a gate
# that cannot evaluate a claim honestly would have to guess, and a guessing
# gate is worse than none. Assert those in scripts/smoke-test.sh instead.
#
# Exit codes 0 in sync / 1 drift / 2 cannot-run, matching lint-shell.sh and
# check-skill-floor.sh. A renamed ARG makes the gate blind, so that is a red
# 2, not a green tick.
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
steps:
- uses: actions/checkout@v4
- name: Doc claims match the build files
run: bash scripts/check-doc-drift.sh
+75 -3
View File
@@ -92,7 +92,56 @@ re-brand of opencode-devbox's `pi-only` variant.
is hashed into `base_tag`, so it costs a base rebuild (~67 min); and if the
section the phrase canary names has changed, re-pin it in
`scripts/smoke-test.sh`.
3. Update `CHANGELOG.md` Unreleased → vX.Y.Z section.
3. **Update the docs this release makes stale — BEFORE you tag.** Rename
`CHANGELOG.md`'s `## Unreleased` to `## vX.Y.Z — YYYY-MM-DD` (em dash, as
every prior release heading uses), then run the gate:
```bash
bash scripts/check-doc-drift.sh # 0 in sync / 1 drift / 2 cannot run
```
It compares README.md's version-pin table against the ARGs it names, and
DOCKER_HUB.md's Node claim against `ARG NODE_VERSION`, plus Hub's
25 000-char limit, unsubstituted `{{PLACEHOLDERS}}`, and stale `Unreleased`
pointers in user-facing docs. With network it also checks DOCKER_HUB.md's
size claims against Hub's measured sizes (check 8) and — check 9 — that
**every component the next build would bake differently from the last
published release is named in the CHANGELOG** above that release's heading:
it reads the `se.jordbo.pi-devbox.*-ref` labels off the published image and
`git ls-remote`s each floating `*_REF`. A red check 9 means an upstream
(pi-toolkit, pi-extensions, mempalace-toolkit, pi-fork,
pi-observational-memory, pi-studio) moved and no entry names the new SHA;
the failure prints the compare URL; a `PI_VERSION` or `MEMPALACE_VERSION`
bump is caught the same way via the `pi-version` / `mempalace-version`
labels. Name the 7-char SHA (or version) where you describe
the change — that is what the old "Dependency audit" tables recorded by
hand, now required.
**Why before and not after:** `docker-publish.yml` runs `actions/checkout@v4`
with no `ref:`, so every job reads `github.ref` — the **tag**. A doc fix
pushed to `main` after tagging does not reach the release, and for
`DOCKER_HUB.md` it does not reach the published Hub page either, because
`update-description` POSTs that file as Docker Hub's `full_description` from
the tag's tree. Getting it in afterwards means re-pointing the tag, which is
its own hazard (v1.8.14 went `601fc98` → `361babd` and broke deploy
verification until `git fetch --tags --force`).
The gate is deliberately narrow — it only checks claims verifiable from files
in this repo. Still eyeball, because these are NOT gated:
- counts and sizes (`~1.1 GB`, "N `mempalace_*` tools", "7 extensions") —
they need a running image; assert them in `scripts/smoke-test.sh` instead
- feature prose that quietly became false, e.g. a "Planned for an upcoming
release" section describing something that already shipped
- `Dockerfile.base`'s `# BASE_REBUILD_DATE:` marker. Ungated on purpose:
`base_tag` hashes Dockerfile.base's content, comments included, so
demanding it be current would force a ~60 min base rebuild on a release
that touched no base files. **Fix it when the base is already rebuilding —
then it is free.**
Measured cost of skipping this, 2026-09-10 (v1.9.0): five stale claims, one
of them published. README's pin table was wrong on all three rows, and
DOCKER_HUB.md — untouched for eight releases — still said Node v22 while the
image shipped Node 24.
4. Verify `docker compose up` works locally with the current `latest` image
if you're upgrading users from a previous version. Then run the
**post-recreate sanity check** inside the running container to confirm
@@ -243,8 +292,10 @@ shipped the same image bytes); preventatively fixed for `PI_VERSION` +
image. Verifies binaries, repo clones, runtime deployment (waits for
keybindings + mempalace bridge + ≥4 extensions before sampling — fixes
the parallel-build-load race documented in opencode-devbox c6f9d11
2026-06-08), and image size threshold (3500 MB; revisit after a few
releases as actuals settle).
2026-06-08), build-time leftovers (see below), and image size threshold
(3800 MB in `SIZE_THRESHOLD_MB`; revisit after a few releases as actuals
settle — this doc said 3500 until 2026-09-11, after the bar had already
moved twice).
If smoke fails on size threshold but build is otherwise fine: bump
`SIZE_THRESHOLD_MB` in scripts/smoke-test.sh in a follow-up commit and
@@ -252,6 +303,27 @@ re-run. The threshold exists to catch *runaway* growth (an accidental
texlive bake-in, a forgotten chrome dependency), not to block ordinary
upstream bumps.
**The size gate is not a substitute for naming the residue.** It carries
~225 MB of deliberate margin, so v1.9.1 shipped +131 MB of pure build
residue — 110 MB of it npm's own download cache under `/root/.npm`, the
rest foreign platform packages — and stayed green. Four named assertions
now cover that ground: no foreign npm-11 platform packages beyond the
host arch (`@esbuild/*`, `@mariozechner/clipboard-*`), no `/root/.npm` in
the image, and — because the prune's real risk is *removing something
needed*, not size — esbuild must compile TS and clipboard must load its
native binding at **every** install site.
Two failure shapes to copy from those, both of which bit here:
- `test ! -d /root/.npm` on mode-700 `/root` passes for a **permission**
error, so the cache assertion refuses to run as non-root. Watch for
this in any assertion about a path you may not be allowed to read.
- `node -e 'require("esbuild")'` resolves by walking up from the CURRENT
DIRECTORY, so it fails with `MODULE_NOT_FOUND` from `/workspace` on a
perfectly healthy image (esbuild is nested inside the pi trees;
`NODE_PATH` is unset). Always path-qualify: `require("<abs>/esbuild")`.
A runbook shipped the bare form with "if this fails, revert the
release" attached, and it duly went red for the wrong reason.
## Build pipeline notes
- **Two-phase**: base + variant. Base is rebuilt only when
+793 -1
View File
@@ -11,7 +11,785 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
---
## Unreleased
## v1.9.3 — 2026-09-19
### Dependency audit (2026-09-19)
Every component checked against upstream by direct command, not assumed. "Baked"
is v1.9.2's published amd64 config labels (read through the registry API) or,
for the floating `*_VERSION=latest` tools, the binaries in a running v1.9.2
container.
| Component | Baked in v1.9.2 | Upstream now | Action |
|---|---|---|---|
| pi | `0.85.1` (pinned) | `0.85.1` is npm latest (2026-09-05) | none |
| pi-atelier | `v0.10.1` (pinned) | `v0.10.1` highest tag | none |
| **mempalace** | `3.9.0` (pinned) | **`3.10.0`** (2026-09-16) | **not adopted — see below** |
| skillset (mempalace fallback snapshot) | `e9e45f7` | `--check` OK: `skills/mempalace/SKILL.md` byte-identical at skillset `debc8f6` | none |
| **mempalace-toolkit** | `dab989b` | **`817b3a8`** | ships the mine-deadline fix (above) |
| **pi-toolkit** | `adfb553` | **`9c87ee8`** | ships the `task`-first AGENTS.md (above) |
| **pi-extensions** | `2610545` | **`25c1265`** | ships `task.ts` + `fork-gate.ts` (above) |
| pi-fork | `e69725c` | `e69725c` | none |
| **pi-observational-memory** | `7b397f4` (3.1.1) | **`cba0334`** (3.1.3) | adopted implicitly via `master` (above); peerDeps still `*`, no pi floor to clear |
| pi-studio (studio variant) | `e04fc7a` | `e04fc7a` highest semver tag | none |
| floating `*_VERSION=latest` tools | — | 14 of 16 already at latest; `uv` `0.12.13`→`0.12.17` (four patch releases, none with a Breaking section), `agent-browser` `0.37.1`→`0.38.1` (minor: `screenshot --if-changed`, `snapshot --delta`, persistent refs; `.1` is a recording-timing fix) | adopted implicitly by the rebuild; named here per this repo's floating-ref rule |
| node | major pin `24`, installed `v24.21.0` | `v24.21.0` newest 24.x | none |
**mempalace 3.10.0 is deliberately not in this release.** It is not a small
bump: a Rust exact-vector engine, four modules split into packages, and three
agent-facing contract changes, two of which touch this image directly —
- **New installs put config and palace under `~/.config/mempalace`.**
`entrypoint-user.sh` decides "first run" by `[ ! -d ~/.mempalace/palace ]`
and `entrypoint.sh` provisions `~/.mempalace` as the persisted path. On
3.10.0 a fresh container would initialise into `~/.config/mempalace` —
outside the volume — and the entrypoint's test would stay true on every
start. Whether `MEMPALACE_HOME`/config resolution honours the old path on a
*pre-existing* `~/.mempalace` is stated ("unchanged") but unmeasured here.
- **`mempalace_event_list` defaults to newest-first when no cursor or `order`
is given.** The toolkit's `deriveOwed` passes `order: "desc"` on two of its
three queries and `deriveClosed` on neither, so their windows are
server-default-dependent. That is a server-side skew (the hub is synlig's
stack, whose default image is `joakimp/pi-devbox:latest` — so this pin *is*
the server's next version), and the right fix is in the toolkit: pass
`order` explicitly on every `event_list` call so the derivation is the same
on 3.9 and 3.10 servers. Filed as follow-up; not a blocker for this tag.
- `mempalace rules` dropped `--agent`; `get_collection()` refuses unknown
names — no caller of either in this repo or the toolkit (grepped).
Adopting it needs its own release: the entrypoint path test, a measured
upgrade of an existing `~/.mempalace`, the toolkit `order` hardening, and the
synlig redeploy sequencing the pin comment in `Dockerfile.base` describes.
---
**New check 9 in `scripts/check-doc-drift.sh`: anything the next build would bake
differently from the last *published* release must be named in the CHANGELOG text
above that release's heading.** The two entries below this one are why. The
`task` tool and `fork-gate` (pi-extensions `25c1265`) and the mine-deadline fix
(mempalace-toolkit `817b3a8`) both reach this image through floating
`*_REF=main` ARGs, so neither produced a diff in this repo, nothing here asked
for a CHANGELOG line, and neither had one until a reader asked. Same shape as
check 8: a claim with no in-repo anchor rots. The hand practice that existed for
it — the "Dependency audit" table in each release's notes, *Baked in vN* against
*Upstream now* — is a "someone remembers" mechanism, and it had lapsed.
How it measures, with no `docker`, `crane` or token: the last published
`vX.Y.Z` is the highest such tag in Hub's list (one request, shared with
check 8); that tag's amd64 config blob is read through the anonymous registry
API (token → index → per-arch manifest → config) and carries one
`se.jordbo.pi-devbox.<name>-ref` label per component holding the SHA the
build-args actually baked. "What the next build would bake" is resolved the way
`resolve-versions` does it — a 40-hex ARG is itself, a branch or tag is
`git ls-remote`d with the peeled `^{}` form preferred (the un-dereferenced SHA of
an annotated tag is the tag object; this repo has raised that false alarm once
already), pi-studio is the highest semver tag read from `<tag>-studio`'s labels,
and `PI_VERSION` is compared as a literal against the `pi-version` label. Nine
components, 7.5 s.
The rule: unchanged needs no mention. Moved requires the new value's 7-char SHA
prefix (tag name for pi-studio, version string for pi) somewhere above the last
published version's `## ` heading — `## Unreleased` plus any not-yet-published
`## vX.Y.Z`, which is what the release commit turns Unreleased into, so the tag
build passes on the same text — sabotage-tested: renaming `## Unreleased` to
`## v1.9.3 — …` stays green; mangling the published `## v1.9.2` heading goes
red, and that test caught a `\b` that would have accepted `v1.9.2-typo` as the
heading (now `(\s|$)`). Naming the SHA rather than the repo is
deliberate: it is what the audit table always recorded, and it makes the failure
message's compare URL one click from knowing what moved. Every upstream commit
re-reds the gate until the CHANGELOG names the new head; that is the intended
cost — **the thing that gets baked is the thing that gets named.** A published
tag with no CHANGELOG heading is a failure, not a skip.
**First run found a move nobody had recorded.** `pi-observational-memory`
`7b397f4 → cba0334` (6 upstream commits, 2026-09-14..16, 3.1.1 → 3.1.3): the
memory workers' `streamSimple` lookup used to iterate every extension-registered
provider and take the first whose `api` matched the model's, so two providers
sharing an API type could route the observer to the wrong one (upstream #70);
it now asks for the model's exact provider and keeps the `api` match as a
consistency check. Reaches this image on the next variant build via
`PI_OBSMEM_REF=master`. No behaviour change expected on the shipped
configuration — every profile here uses the built-in `amazon-bedrock` provider
and no extension registers one — but that is an expectation, not a
measurement; the code path only differs when an extension has called
`registerProvider`.
Header and `lint.yml` comment corrected alongside: both still said this gate
"needs no network", which check 8 made false on 2026-09-14. There are now two
classes — hermetic checks 1–7, and published-state checks 8–9 that SKIP loudly
and counted when offline. Considered and not added, with reasons in the script:
a `docker-compose.yml` ↔ `.env.example` variable cross-check (the four
mismatches are commented-out lines, the mempalace-server compose file's own
documented variables, and two entrypoint-consumed variables — a gate there
would fire on nothing wrong), and a "documented tag exists on Hub" check
(check 8 already SKIPs a missing tag by name, and a hard fail would
misreport the window between tagging and publish).
**Blind spot closed while it was free: `se.jordbo.pi-devbox.mempalace-version`.**
No label recorded the palace pin, so check 9 could not see a `MEMPALACE_VERSION`
bump — checks 1–3 keep README's pin table consistent, but nothing required a
CHANGELOG line for the one component whose skew against the shared central
palace is fleet-wide. The label is set in `Dockerfile.base` next to the `ARG`
that defines it and **inherited** by both variants: no second copy of the pin to
drift, no build-arg to plumb through four variant call sites, and it states the
pin of the base the image *actually* built on — the question that matters when
`base-decide` cache-hits an older base. Intent, like every label here; the
manifest's `mempalace_version` stays the ground truth, and `smoke-test.sh` now
asserts label == installed binary (the one way they diverge is a base built with
`INSTALL_MEMPALACE=false`, or an install that resolved off-pin). Until a release
carries the label, check 9 reports that component as a counted SKIP, not OK;
costs nothing extra because this Unreleased already forces a base rebuild
(`rootfs/` skill floor).
---
**The rule "use `pi-task`, not `fork`, for a brief that carries a prohibition" was
written in the global `AGENTS.md` and in the pi-extensions skill, and it lost to
the `fork` tool's own description anyway.** Measured on tor-ms22, 2026-09-17: five
of five fork briefs in one session carried "do not"; one returned verbatim quotes
that did not exist in the source, and four had disjoint write boundaries that
fork cannot enforce — re-run as `pi-task`, all four passed their envelope. The
image now ships the rule where the decision is *made*, not where it is read about:
- **`fork-gate.ts`** (pi-extensions `25c1265`): a `tool_call` hook that blocks a
`fork` whose brief contains a prohibition (*do not / never / only …*), a write
boundary (*only touch / read-only / stay within …*) or a clause-initial
file-changing imperative (*Edit …, Commit …, Fix …*). The block reason the model
reads **is** the `task(...)` call to make instead. It matches wording, not
intent, and says so. `PI_FORK_GATE=off` logs instead of blocking.
- **`task.ts`** (same commit): `pi-task` registered as the `task` tool, with the
decision rule in its description and in `promptGuidelines` — which pi appends
to the **system prompt**, the one place compaction cannot remove it from.
Before spending a model run it rejects the two spec errors that make a boundary
violation certain (`write_allowed` not an *exact subset* of `roots` — pi-task
keys deltas by root string; a writable root nested inside a watched-only root,
whose porcelain would change every time) and serialises sibling tasks whose
roots overlap (parallel siblings saw each other's writes as violations,
2026-09-17). A FAIL verdict is a *result*; only the CLI refusing to run is an
error.
- **`pi-global-AGENTS.md`** (pi-toolkit `9c87ee8`): the delegation section now
opens "`task` first, `fork` second" — the one discriminator (what the child
sees), a three-question pre-flight before any `fork(...)`, a copy-paste minimal
call with the roots contract. This file is in the system prompt; the skill is
not, which is why the rule lives here.
- **Skill floor refreshed** to `25c1265` (`check-skill-floor.sh` OK, tree
`9b85a633…`); mirror in skillset `debc8f6`.
Why prose failed, as mechanisms: `fork` is a **tool** — its self-recommending
description ("exploration, implementation, testing, review…") is in the tool
list on every turn and survives compaction; the skill is gone after the first
compaction; `pi-task` was a CLI to be remembered and reached through `bash` with a
hand-written JSON spec. The asymmetry *widens* in exactly the long sessions where
fork is worst, and fork deletes its temp dir on exit, so its failures were found
only by re-verifying the narrative. Third time on this fleet that a rule held in
prose and violated in practice was fixed by moving it into a hook
(`check-secrets`, `check-egl-only`, now delegation).
Evidence, two-sided: `test/fork-gate.test.mjs` pins 15 must-block briefs
(including the real shapes) and 10 must-pass (including *"Write a summary…"*,
*"Report which files were modified…"*, *"Give me an update…"* — verbs a naive
list misfires on); the shipped classifier over the five real briefs from the
motivating session redirects 5/5. Live in `pi -p`: the fork was intercepted
before any child spawned (no `/tmp/pi-fork-*` directory), `task` returned PASS in
4 s / $0.014 with an evidence pointer and audit dir, a `usd=0.000001` budget came
back as a FAIL *result* (`isError=false`), and a nested-root spec was rejected
with no audit dir created.
Nothing to wire in this repo: `PI_EXTENSIONS_REF=main` floats and
`pi-extensions/install.sh` symlinks every `extensions/*.ts` on container start, so
the next build ships both. A container running today has neither —
`~/.pi/agent/extensions/` links into `/opt/pi-extensions`, which is the baked ref.
---
**`[mempalace ext] feed (tick) failed: mempalace remote request 'tools/call'
failed: timed out after 60000ms` is the same event as the `mine timed out after
30000ms` message the 2026-09 toolkit fix addressed, one deadline further down —
and that fix was incomplete.** mempalace-toolkit `817b3a8` (2026-09-18; v1.9.2
baked `dab989b`; the floating `MEMPALACE_TOOLKIT_REF=main` picks it up on the next
build). `MEMPALACE_FEED_MINE_TIMEOUT_MS` had been raised to 300 000 but was only
*raced* against `client.callTool("mempalace_mine")`; `callTool()` had no way to
carry a deadline, so every mine went out under the transport's generic
per-request timeout — `MEMPALACE_MCP_TIMEOUT_MS`, 60 000, the value the "Stall
protection" comment in `Dockerfile.base` documents — which fired first on every
honest 60 s+ mine on the shared single-writer palace. The 300 s was unreachable.
Over HTTP nothing is lost (the mine continues server-side and is idempotent);
over stdio it was worse than noise — that transport **kills the child** on
timeout, so there the mine really was aborted at 60 s.
Fix: `callTool(name, args, { timeoutMs })` on both transports, the feed passes
its own deadline down, plain calls keep 60 s (a *query* that slow is wedged; the
race stays as the liveness guard for a transport with its timeout disabled).
`scripts/test-mcp-call-timeout.sh` cuts `RemoteMcpClient` out of the shipped
file, drives it against a local JSON-RPC server that delays `tools/call`, and
asserts three things — a plain call rejects at the generic deadline, the override
outlives it, the override is itself a deadline: 2 of 6 fail on `dab989b`, 6 of 6
pass on `817b3a8`. `test-owed-withdrawal.sh` (17) and `check-mcp-client-sync.sh`
stay green; sync token `v1` untouched, since nothing in the protocol changed.
Reading for the fleet: on a fixed build that message means a mine exceeded
*five* minutes — look at palace size or a competing writer, not at the timeout.
---
**`scripts/recreate-sanity-check.sh` asserted that `/tmp/sshcm` exists while every
`ssh` in the container was dying `rc=255`, and it was right to — it was checking
the directory the *image* creates, and the breakage was in the directory a
*config* named.** Found on the v1.9.2 first boot on `emb-7kj4vr4g`. A durable
`~/.pi/ssh/config` (hand-written into the `~/.pi` named volume by the previous
session, so it would survive the recreate) declared `ControlPath
/tmp/ssh-cm/%C` — with a hyphen. Nothing in this repo creates that path; the
canonical directory is `/tmp/sshcm`, spelled the same way in four places
(`Dockerfile.base`, `entrypoint-user.sh`, `recreate-sanity-check.sh`,
`smoke-test.sh`). Result:
```
unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory
```
`rc=255`, and the remote command never ran at all — `ControlMaster auto` with an
unusable `ControlPath` fails hard rather than falling back to an unmultiplexed
connection. The old check passed truthfully, about the wrong object. **Two
independent facts about the same subsystem can both be true while the subsystem
is dead; a check that asserts only one of them cannot see their disagreement.**
**New in `recreate-sanity-check.sh`: resolve the ControlPath `ssh` itself would
use, via `ssh -G`, and require its parent to exist and be writable.** `-G`
applies real config precedence — first-obtained-value-wins, the system drop-in,
`Include`, an `-F` override — so it answers "which rule captured this host"
instead of re-implementing the guess. It never opens a connection: measured
0.116 s for 48 hosts.
This also puts a check under a caveat that had been documented in prose in
`Dockerfile.base` ("SSH client defaults") and verified nowhere: a per-host
`ControlPath ~/.ssh/cm/%r@%h:%p` inherited from a **read-only** bind-mounted
`~/.ssh` produces the identical failure, `cannot bind … Read-only file system`.
On the machine where this was built that is **16 of 50 hosts** — `freeipa-1..6`,
`gitea.egl.lan`, `runner-1..3`, `tor-ms22` and more — none of which had ever been
reported by anything.
**The two routes get different severities, deliberately.** Default `ssh`
precedence legitimately lands in the read-only `~/.ssh` on any host whose own
config pins it there, and the supported workaround (`ssh -F
~/.ssh-local/config`, generated every container start by `setup-lan-access.sh`)
already exists — so that is a `warn`. Making it a failure would paint the script
red on every run of every device, and **a check that fires benignly every time is
one you learn to ignore**, which is the same reasoning that keeps `lint-shell.sh`
at `-S error`. The sidecar route is the prescribed one, so there an unusable
directory is a hard `fail`. Host lists are capped at six names plus a count for
the same reason: unreadable output is ignored output.
Teeth proven in both directions, with each sabotage confirmed by `diff` *before*
the result was believed — a vacuous sabotage that silently fails to apply
reports "gate passed" and is worse than no test:
| Sabotage | Class | Result |
|---|---|---|
| `ControlPath` → nonexistent dir | the original hyphen bug | `✗` `rc=1` |
| `ControlPath` → existing but read-only dir | the `Dockerfile.base` caveat | `✗` `rc=1` |
| restored | — | `✓` `rc=0`, sidecar byte-identical |
No config was added to this repo to fix the original problem, because the fix was
to **delete** the offending file: `~/.pi/ssh/config` was a third hand-maintained
copy of what `setup-lan-access.sh` already generates from version control on
every start, with fewer features (no `known_hosts` sidecar, no
`StrictHostKeyChecking accept-new`) and one typo. The host-owned `~/.ssh/config`
is also left alone on purpose — those `~/.ssh/cm` paths are correct *on the host*,
where `~/.ssh` is writable, and the container-side override is the right layer.
---
**The size numbers on the Docker Hub page were the only claim in these docs with
nothing in the repo to check them against, and they had gone 20% wrong across
eight releases.** Every other claim `scripts/check-doc-drift.sh` guards is
anchored to a build file — a pin, an `ARG`, a placeholder — so it cannot rot
without someone editing the thing it describes. Nothing in this repo states the
image size, so `DOCKER_HUB.md`'s `~1.1 GB` simply drifted while the image grew,
and `update-description` POSTed it to Docker Hub every release. It is the first
number a stranger reads about this image.
Corrected against Docker Hub's measured `full_size`, 2026-09-14 after v1.9.2
published (amd64 / arm64, compressed):
| Row | Claimed | Measured | Now says |
|---|---|---|---|
| `:latest` | ~1.1 GB | 1.228 / 1.211 | ~1.23 GB |
| `:latest-studio` | ~1.15 GB | 1.255 / 1.238 | ~1.25 GB |
| `:base-latest`, `:base-<hash>` | ~1.0 GB | 1.167 / 1.151 | ~1.17 GB |
`full_size` is the right field because it tracks the **first manifest entry**
(amd64), not the sum across architectures — measured on v1.9.2:
`full_size=1.228`, `amd64=1.228`, `arm64=1.211`, `sum=2.439`. That matches the
table's per-arch "Size (compressed)" column.
**New check 8 in `scripts/check-doc-drift.sh`: size claims vs Hub's measured
`full_size`,** so this class cannot rot silently again. It fails on drift beyond
tolerance, and **skips loudly** — a new `skip()` helper, counted and named in the
summary — when `curl`/`python3` are missing, the API is unreachable, or
`SKIP_SIZE_CHECK=1`. Skips are deliberately neither `OK` nor a failure: printing
an unverified claim as OK is the habit this file exists to break, while failing
on Docker Hub's uptime would make every release hostage to a third party. Not in
`hooks/pre-push` (that runs `lint-shell.sh` only), so pushes do not hit the
network.
Two bugs were caught while building it, both by writing the expected exit code
down *before* running the check:
- **The tolerance would have missed its own motivating case.** The percentage is
computed against the *measured* size, but the 20% first chosen came from the
claim-relative figure. The real drift was `|1.1 − 1.37| / 1.37 = 19.7%` — it
would have passed. Now 15%, sitting inside a window whose bounds are both
measured: above the largest legitimate skew (a claim describing the published
release while the next tag changes the size — v1.9.1's 1.37 against v1.9.2's
1.23 = 11.4%) and below the rot it exists to catch (19.7%).
- **A `|| true` on the python invocation made the gate fail open.** It printed
`DRIFT` and exited 0 — a gate that reports the defect and passes anyway.
Removed; the outer `|| SIZE_RC=$?` is what satisfies `set -e` without
swallowing the code. Verified two-sided afterwards: a 19% drift exits 1 at the
default tolerance and 0 at `SIZE_TOLERANCE_PCT=25`, so the threshold is doing
the work rather than the ordering.
**Explicitly NOT covered:** `README.md`'s `~3.2 GB` figures are *uncompressed*
on-disk sizes, and the registry exposes compressed sizes only (manifest layer
sizes are compressed; the config blob carries no uncompressed totals). Measuring
them needs a real pull, so they remain unverified — a green check 8 says nothing
about them, and the script says so where a reader will see it.
**`promote-base-latest`'s conditional re-tag is now measured, which closes an
open question from the v1.9.2 release verification.** The job compares digests
and re-tags `base-latest` only if stale, and that short-circuit determines
whether a release watcher may assert *freshness* on the alias or merely
*existence*.
Measured on run 669: the digests differed (`want sha256:8c452575…`,
`have sha256:f34ad201…`), the job logged
`Promoting base-latest -> …:base-b5f2d03baae2`, `crane copy` ran
`17:10:04 → 17:10:08`, and Docker Hub's `base-latest.last_updated` moved to
`17:10`. So **a `crane copy` that runs does bump Hub's timestamp** — a manifest
re-tag is a tag write.
The identical-digest case remains *unproven by construction*: when `base-latest`
already resolves to the new `base-<hash>` the job prints
`base-latest already current; nothing to promote.` and never copies, so the
timestamp legitimately stays put. A freshness assertion would then report a
correct release as stale. The rule is therefore conditional on
`base-decide`'s `need_build` — strengthen it to fresh-required only when
`Dockerfile.base`, `rootfs/` or a folded `*_REF` has moved, and leave it as an
existence check for the variant-only and docs-only releases that cache-hit the
base. Operational guidance lives with the tooling that consumes it
(`ci-release-watcher` skill, skillset `c8034af`); recorded here because it is a
property of *this* pipeline.
Worth restating alongside it, because v1.9.2 proved it the useful way: tag
freshness is not the authoritative evidence that a base rebuild baked what you
expected. The `se.jordbo.pi-devbox.*-ref` labels are — readable straight from the
registry with no `docker` or `crane` (token → manifest index → per-arch manifest
→ config blob), and worth validating against the *previous* tag first, since a
reader that cannot show the old SHA cannot prove the new one.
---
## v1.9.2 — 2026-09-14
**The v1.9.1 residual is attributed and fixed: it was mostly npm's own download
cache, not the platform binaries it looked like.** v1.9.1 pruned the 26 foreign
`@esbuild/<platform>` directories that npm 11 installs, which fixed the 431 MB
size-gate failure — but the image still shipped **+131 MB compressed over
v1.8.14**, nearly all of it in the single pi/extensions install layer (87 MB →
206 MB). v1.9.1's notes recorded the leftover as an open item with an explicit
hypothesis (the `@mariozechner/clipboard-*` family, same npm 11 behaviour, a
different package) and an explicit warning that the hypothesis was **not** a
measured cause. It was measured on 2026-09-11, after recreating onto v1.9.1, and
the hypothesis accounted for only a sixth of it:
| item | v1.8.14 | v1.9.1 | delta |
|---|---|---|---|
| `/root/.npm/_cacache` — the build's npm download cache | 35.2 MB | 145.3 MB | **+110 MB** |
| `@mariozechner/clipboard-*` foreign platform packages (2 sites) | 0 MB | 21.1 MB | **+21 MB** |
| variant install layer, uncompressed total | 270.8 MB | 401.9 MB | +131 MB |
That is the whole delta with no unexplained remainder. Both items are now
deleted in the **same layer** that creates them, in both the main install RUN and
the studio RUN:
- **`purge_build_caches`** — `npm cache clean --force` plus `rm -rf /root/.npm`.
npm 11 caches every platform tarball it downloads, including the ones the
prune then deletes, so the cache grew far faster than the installed tree.
Nothing at runtime reads it: the build runs as root, the container runs as
`developer` with its own cache under `$HOME`.
- **`prune_foreign_esbuild` → `prune_foreign_natives`** — now covers both
measured families. For clipboard the keep-set is `clipboard-linux-$arch-gnu`
**and** `-musl`, because its napi-rs loader chooses between them at runtime
from its own `isMusl()` probe; the musl package is a 420-byte stub, so keeping
it is free insurance. The bare `@mariozechner/clipboard` wrapper has no
hyphen suffix and cannot match the pattern.
**Verified on arm64 before writing the patch, which is why the order was
update-then-patch:** a widened `rm -rf` glob is the worst possible change to
write against a tree you cannot inspect, and there is no docker CLI inside the
container — but after a recreate the container *is* the image. The prune was
exercised against a copy of the real trees with foreign directories fabricated
back in (aix-ppc64, android-arm64, darwin-arm64, win32-x64, linux-x64): all
removed, host `linux-arm64` kept at both sites, 21 MB freed, and
`require('@mariozechner/clipboard')` still loads and exports all 18 functions.
`esbuild.transformSync` still compiles TS at both sites. v1.9.1's own arm64
validation of the esbuild prune also passed here — CI could only smoke-test
amd64.
**Two sentinel assertions in `smoke-test.sh`, because the size gate did not
catch this.** The gate has ~225 MB of deliberate margin, so 131 MB of pure
build residue stayed green. There are now named PASS/FAIL checks for *foreign
platform packages beyond the host arch* and for *`/root/.npm` being shipped* —
the latter deliberately refuses to run as non-root, because `test ! -d
/root/.npm` on mode-700 `/root` would otherwise pass for the wrong reason. The
size-failure diagnostics now also list cache paths: they previously enumerated
only `node_modules` and `/opt`, where these bytes were not.
**Also fixed: the prune's own progress line was mangled.** `-printf '%f\\n'`
reaches the shell with both backslashes (confirmed from the published image's
recorded `created_by`), so `find` emitted a literal backslash and `tr` then ate
the `n` out of the name — v1.9.1 printed `esbuild platform dirs kept: li
ux-arm64`. Single backslash now.
Deliberately **not** changed: `/tmp/node-compile-cache` (1.3 MB). The manifest
RUN at the end of `Dockerfile.variant` calls `pi --version` again, so deleting
it earlier only relocates those bytes into that layer — today's manifest layer
is 128 kB precisely because it finds the cache warm.
**Two functional assertions as well, after the runbook command left for the
next machine failed for the wrong reason.** v1.9.1's open item prescribed
`node -e 'require("esbuild").transformSync(...)'` as the post-boot check, with
"if this fails, the prune removed something needed → revert to v1.8.14". Run
from `/workspace` it fails with `MODULE_NOT_FOUND` on a perfectly good image:
`require` resolves by walking up from the current directory, esbuild lives
nested inside the two pi trees, and global installs are not on node's require
path (`NODE_PATH` is unset). The check that verified the prune last time only
passed because the shell happened to be inside the tree. Smoke now does it
properly and CI owns it: for every install site found in the image (so the
studio variant's third site is covered automatically), esbuild must compile TS
and `@mariozechner/clipboard` must load with its native binding attached — the
latter is the real proof for the clipboard prune, since napi-rs resolves the
platform package at `require()` time. Both were verified as a four-way matrix:
green on the real image *from `/workspace`*, and red against copies of the same
packages with the host platform binary removed (`The package
"@esbuild/linux-arm64" could not be found`).
Touches `Dockerfile.variant` and `scripts/smoke-test.sh` only: `Dockerfile.base`
is unchanged, so this needs no base rebuild and should **ride the next release**
rather than burn a cycle of its own.
> **Superseded by the entry below:** that entry refreshes the vendored mempalace
> skill snapshot, which **is** hashed into `base_tag`. The release as a whole now
> costs a base rebuild (~67 min). The *size* work above still needs none of its
> own; the two simply travel together now.
**A behaviour change reached the fleet without any release naming it, and a
paragraph in these notes kept saying it had not.** v1.9.1 bakes
`mempalace-toolkit` **`e68ee20`**, which contains **`e2b060a`** — requester-side
ask withdrawal (`isWithdrawn`, RFC 003 §3.3 clause 4). So the behaviour has been
live on every v1.9.1 device since 2026-09-10, while the v1.9.0 section of this
file still read "not yet pinned … this image still pins `e45f6b4`" and the
mempalace skill still told every agent, at session start, that a withdrawal is
impossible: *"there is nothing anyone can do about it from the other end."*
The mechanism is the point, because it will do this again. `Dockerfile.variant`
carries `ARG MEMPALACE_TOOLKIT_REF=main` and `docker-publish.yml` resolves it to
a commit SHA at build time (`gitea_sha mempalace-toolkit`). A release therefore
absorbs *whatever toolkit `main` holds at that moment*, and "what behaviour did
this image gain?" is a question **nobody is structurally forced to answer**. This
fleet already has the rule — a floating ref that pulls a behaviour change into
the image must be named in the CHANGELOG *before* tagging. It was honoured for
the feed-tick fix, which v1.9.1 names explicitly (`309980b`, `e68ee20`), and
missed for the commit sitting in the same range.
Measured before being written, two independent routes, expectation recorded
first ("label should read ≥ `e68ee20`, since the build at 22:00Z postdates that
commit's 18:58Z"):
| route | result |
|---|---|
| Docker Hub config-blob label, `:v1.9.1-studio` and `:latest-studio` (same digest) | `se.jordbo.pi-devbox.mempalace-toolkit-ref = e68ee2071ca3ad39…` |
| `git merge-base --is-ancestor e2b060a e68ee20` | ancestor — the fix is inside the baked ref |
| baked `extensions/pi/mempalace.ts` sha256 vs v1.8.14's | `7c16fe14…` vs `dfca71e9…` — different bytes, so not the pre-fix file |
| `grep -c isWithdrawn` on this v1.8.14 container's baked copy | `0` — confirms the split, and that tor-ms22 cannot exercise it |
The first attempt at that label read **empty**, and the empty result was a claim
about the request, not the image: Docker Hub redirects blob fetches to a CDN and
`curl` without `-L` returns 0 bytes with exit 0. A registry audit that reports
"no labels" should be assumed to be missing `-L` until proven otherwise.
So the skill is updated rather than deferred (skillset **`e9e45f7`**, vendored
here with `scripts/vendor-mempalace-skill.sh`, `44472 → 46045 B`). Two things
were deliberate:
- **The "silence is not an answer" rule keeps its teeth.** An ask still stays
owed until the *recipient's* terminal event; what is new is a release by the
**asker**, explicitly marked. Stated that way round on purpose — the wrong
reading of this change is "withdrawals happen, so I need not reply".
- **The precondition ships with the rule**, because this skill is read on images
that lack the behaviour (tor-ms22, right now):
`grep -c isWithdrawn /opt/mempalace-toolkit/extensions/pi/mempalace.ts`, where
`0` means the withdrawal will not reach the recipient's mailbox. Same shape as
the provenance bullet's live-bridge check.
**The snapshot canary is re-pinned, and this time it fails on the old bytes
instead of merely failing to notice them.** The retired pair (`"Diaries
self-heal…"` present / `"Agent diaries live in"` absent) was still green against
the new snapshot, so it was blind to this refresh exactly as the pre-v1.8.13 pair
was blind to that one. The replacement is stronger than any predecessor here
because **both witnesses come from the same upstream commit**: `e9e45f7` added
`"Withdrawing an ask you sent"` and deleted `"nothing anyone can do about it from
the other end"`, the sentence the new bullet contradicts. Directions were
measured against both files rather than read off the diff (`new=1/old=0` and
`new=0/old=1`), then the canary body was **executed** against each: new → `rc=0
ok`, old → `rc=1` empty.
**`isWithdrawn` is no longer deployed-and-unproven — and the suite that pins it
had been dark since the node 24 bump.** The rule was exercised on the released
image against the live logstream on 2026-09-14, from a container recreated onto
v1.9.1 (born 13:21:41Z, confirmed by entrypoint-written mtimes and docker-written
`/etc` files agreeing to the second; `/proc/uptime` and `ps -o lstart=` were not
used, per their retraction). Baked toolkit `e68ee20`, `grep -c isWithdrawn` = 3,
mempalace.ts sha256 `7c16fe14…`, and the file pi actually loads verified to be
that same path and hash rather than a stale copy.
The instrument matters as much as the result: the shipped extractor was lifted
out of `scripts/test-owed-withdrawal.sh` and used to cut `TERMINAL_STATUS`,
`isStrictlyAfter`, `isAnswered` and `isWithdrawn` out of the **baked**
mempalace.ts by brace matching, then `deriveOwed`'s three queries were replayed
with their exact shipped parameters over the same `/mcp` transport the extension
uses. Shipped bytes, live data, no second copy of the logic. Baseline owed = 1,
agreed by three independent routes (the extension's own wake-up card; a hand
derivation of 23 candidates; the shipped predicates). Every expectation was
recorded before its measurement:
| probe | predicted | observed |
|---|---|---|
| positive control planted | owed 1 → 2 | 2 |
| requester withdraws it | back to 1 | 1, and `isAnswered=false isWithdrawn=true` |
| **third party** retracts someone else's ask | no effect | still owed |
| requester withdraws in **prose**, no marker | no effect | still owed |
| cleanup by ordinary replies | owed == baseline | 1, then 0 |
The count returning to baseline is only arithmetic; the per-predicate verdict is
what makes it a statement about mechanism. Probe A remained a raw candidate
throughout and no terminal reply of ours existed on its correlation, so its
removal is attributable to the withdrawal rule alone. Final attribution: A
cleared by `isWithdrawn` only, the two negative controls cleared by `isAnswered`
only — each probe retired by the predicate it was built to exercise. Both marker
spellings now have live witnesses (`metadata.withdraws` naming the ask's event id,
and naming its correlation). Unplanned and worth more than the probes: against
live data the rule also fires on the incident it was written for — seq 112 is
reported withdrawn-by-requester, i.e. mbp-m1-2020's seq 119 withdrawal now works,
so the 41h false obligation cannot recur. Full record in the coordination log at
`project/pi-devbox` seq 140; harness preserved as artifact
`art_20260914T141704_a7a9a294a4bb`.
Not proven, and deliberately not claimed: the extension's **own in-process
mailbox poll** surfacing a withdrawn ask. That poll fires at `agent_settled` when
the agent is idle; two probes were left owed across the rest of the session to
give it a window and it did not fire. Calling that confirmed would be a claim
about the session's patience, not about the code.
**Toolkit pickup `dab989b` — the owed-set suite could not run on this image, and
its own gate was why.** Reaching for the suite as corroboration exposed a second
defect: its precondition line `node --experimental-strip-types --check "$SRC"`
exits 2 on an unmodified mempalace.ts under node 24, and with `set -euo pipefail`
that skipped all 17 assertions and both regression guards. The cause is narrower
than the error suggests — it names the inline type-import, but `node --check` does
not type-strip at all: a file containing only `const x: number = 1` fails
identically, while executing the same file works. So the gate could never
validate TypeScript on any node; v1.9.1's bump from v22.23.2 to v24.21.0 is the
most likely trigger, though with no node 22 on the box that half stays labelled
inference rather than measurement. It failed **closed** — loud exit 2, never
vacuously green — which is the good direction, and the reason it went unnoticed is
that nothing in CI runs this suite.
That matters more than a red test, because this suite is the only thing that makes
`isWithdrawn`'s failure mode visible: a wrong rule there does not throw and does
not log, it makes a real unanswered ask vanish from a mailbox forever. `dab989b`
strips first and syntax-checks the emitted JS, and splits the exit codes so that
`3` means the gate cannot run while `2` means the source does not parse —
collapsing those is how the defect disguised itself as "mempalace.ts does not
parse" while mempalace.ts was fine. Verified in six directions with expectations
written first: clean source `rc=0` 17/17; malformed TypeScript `rc=2`; stripper
made unavailable `rc=3` without the misleading message; and the four mutation
kills back at e2b060a's counts of 3/1/2/1, so sensitivity is restored rather than
asserted.
> **Cost, and it is smaller than it looks:** `MEMPALACE_TOOLKIT_REF` is resolved
> by CI to the head of the toolkit's `main` and folded into the `base_tag` hash —
> verified at `docker-publish.yml:126-128`, whose own comment gives the reason
> ("otherwise a toolkit-only fix never lands"). So `dab989b` moves `base_tag` and
> the next tag rebuilds the base, with nothing to remember to trigger. But it adds
> no rebuild that was not already owed: `9aaff26` refreshed the vendored skill
> snapshot under `rootfs/`, which is also hashed into `base_tag`, so a base
> rebuild has been pending since before this fix existed. The toolkit pickup rides
> along with it, and the same rebuild is what finally bakes skillset `e9e45f7` and
> turns the snapshot canary above green against the image's own floor.
---
## v1.9.1 — 2026-09-10
**v1.9.0 was tagged but never published: its own smoke gate stopped it, and it
was right to.** `build-base` succeeded, then `smoke` failed 90-passed/3-failed,
and because `build-variant` needs `smoke`, both variants, `promote-base-latest`
and `update-description` were skipped. No image reached the registry, so
`latest` still pointed at v1.8.14. v1.9.1 carries everything listed under v1.9.0
below, plus the three fixes here. Two of the three failures were self-inflicted
by v1.9.0's own changes, and the third was a real regression that the Node bump
dragged in — which is the case for keeping the gate strict.
**Failure 1 — the image was 431 MB over its size threshold, and npm 11 was the
cause.** Node 22 → 24 brings npm 10 → 11, and npm 11 installs **every**
`@esbuild/<platform>` optional binary rather than only the one matching the host:
26 platform directories covering aix-ppc64, android, darwin, freebsd, netbsd,
openbsd, win32, s390x, riscv64 and more, none of which this image can execute.
Measured on pi-fork's dependency tree, same repo and same command:
| npm | packages | `node_modules` |
| --- | --- | --- |
| 10.9.8 | 136 | **165 MB** |
| 11.19.0 | 169 | **449 MB** |
The 165 MB figure reproduces exactly what v1.8.14 shipped, which is what
identified npm rather than the image as the variable. esbuild declares those
binaries with `os`/`cpu` constraints, but npm 11 ignores them — and also ignores
`--os`/`--cpu` flags and an `.npmrc` carrying `os=`/`cpu=` (all three measured,
all three still produced 26 directories). So `Dockerfile.variant` now prunes
explicitly, keeping only `linux-$(node -p process.arch)` so one line is correct
on amd64 and arm64. Verified this removes dead weight and not function: after
pruning, `esbuild.transformSync` still compiles TypeScript. The prune runs in the
**same layer** as each `npm install` — deleting in a later `RUN` would leave the
bytes in the earlier layer and shrink the image by nothing. Three sites are
covered: the global pi install, pi-fork, and pi-studio (which pulls its own
pi-coding-agent copy), for roughly 548 MB recovered in the non-studio variant and
822 MB in studio. The threshold stays at 3800 MB deliberately: it caught a real
regression, and raising it to accommodate one would have discarded the signal.
**Failure 2 — the om `node_modules` assertion was checking an npm artefact, not
the software.** `pi-observational-memory` declares **zero** runtime dependencies:
8 devDependencies (omitted by `--omit=dev`) and 4 peerDependencies, which pi
itself provides. npm 10 still materialised a `node_modules` for it, but that
directory contained exactly **one file** (`.package-lock.json`, 4 KB) and no
nested `package.json` — 20 empty scope directories. npm 11 stopped creating it,
so `test -d node_modules` went red while nothing about om had changed or broken.
The assertion now checks what must actually hold — that the entry point pi loads
exists — read out of the manifest pi itself reads (`package.json` →
`pi.extensions`) rather than a hardcoded path that could drift. pi-fork keeps its
`node_modules` check, because pi-fork has real dependencies where the directory's
absence would mean something.
**Failure 3 — the skill-source annotation broke the assertion that reads it.**
v1.9.0 taught `pi-devbox-version` to say *which* pi-extensions copy shipped
(`baked (package copy)`, or a loud FALLBACK/MIXED marker). The smoke assertion
matched `^ $s +baked$`, anchored at the end, so the annotation failed it even
though the state reported was correct. The pattern now allows an optional
` (...)` suffix, matched loosely on purpose: *which* copy shipped is already
asserted authoritatively against the manifest field and its measured tree hash,
and re-encoding that wording in a second regex would just add a second place to
update. The lesson recorded rather than the fix alone: the display branches were
tested in an isolated harness that passed, but the assertion **consuming** them
was never run — harness-passes-therefore-consumer-passes was an assumption.
**A red size assertion now carries its own diagnostic.** Attributing the 431 MB
took a full CI-log dig plus a local npm bisect, while the container knew where
its bytes were the whole time. On failure the check now prints the largest
layers, the largest directories, and a count of `@esbuild` platform directories
as a sentinel for this exact regression recurring — the same principle the `run()`
helper already applies to every other assertion.
No `Dockerfile.base` or `rootfs/` change, so the base fingerprint is untouched
and `base-decide` reuses `base-0fb1256c7f99` built during the v1.9.0 attempt.
**A gate for documentation drift, because five claims rotted in one release and
one of them was published.** Preparing v1.9.0 turned up a cluster of stale
facts, all the same shape — a value written once by hand, in a file nothing
verifies, about a number that lives somewhere else and moved:
- `README.md`'s "Version pins" table was wrong on **all three rows**: pi
`0.84.4` vs `ARG PI_VERSION=0.85.1`, pi-atelier `v0.10.0` vs `v0.10.1`,
mempalace `3.8.0` vs `3.9.0`. That table is the worst possible place for this,
because it exists *specifically* to be the reviewable record of what the repo
freezes deliberately — so a wrong row destroys the only thing it is for.
- `README.md` listed already-shipped typst PDF export under "Planned for an
upcoming minor release", carrying the self-contradicting marker "(shipped in
Unreleased/base)". The **fourth** instance of the stale-`Unreleased`-pointer
class this changelog already documented three of.
- `DOCKER_HUB.md` claimed "Node.js v22" while v1.9.0 ships Node 24.
The last one is why this became a gate rather than a resolution to be careful.
`DOCKER_HUB.md` is **published**: `update-description` POSTs it to Docker Hub as
`full_description` on every tag. It had gone **eight releases** (v1.8.6 →
v1.9.0) without a touch. Nothing generates it — CI only substitutes
`{{PI_VERSION}}` — and nothing checked it, so the sole mechanism keeping it true
was whoever remembered. Worse, it is read from the **tag**, so the stale page
published with v1.9.0 anyway and the fix could only ride the next release.
**New: `scripts/check-doc-drift.sh` + a `doc-drift` job in `lint.yml`.** Seven
checks, all comparing a doc string to a value that exists in this repo, so it
needs no network, no token, no built image, and no sibling clone:
- README's three pin-table rows vs the ARGs they name *by name*
- `DOCKER_HUB.md`'s Node claim vs `ARG NODE_VERSION`
- placeholders CI will not substitute — the publish step greps for leftovers of
`{{PI_VERSION}}` only, so any *second* token sails through and publishes
literally
- `DOCKER_HUB.md` under Docker Hub's 25 000-char `full_description` limit
(previously discoverable only as a non-200 *after* the full build)
- `Unreleased` appearing in a user-facing doc, which is always a pointer that
outlived what it pointed at
Exit codes match `lint-shell.sh` and `check-skill-floor.sh`: `0` in sync, `1`
drift, `2` cannot run — a renamed ARG makes the gate blind, which is a red `2`,
never a green tick. Verified with **15 controls**: every check fails when its
claim is broken, the real v1.9.0 Node bug is caught, and two false-positive
controls pass — the first version of the placeholder check wrongly flagged
`README.md:900`'s `docker inspect --format '{{json .Config.Labels}}'`, a Go
template in a legitimate example, so the pattern is now anchored to the
UPPER_SNAKE convention CI actually substitutes. **The gate was wrong, not the
doc** — which is the whole reason a gate gets negative controls.
Deliberately **not** gated, and the reasons matter more than the list:
- Counts and sizes (`~1.1 GB`, "N `mempalace_*` tools", "7 extensions") need a
running image. A gate that cannot evaluate a claim honestly would have to
guess, and a guessing gate is worse than none — assert these in
`scripts/smoke-test.sh`, where a real image exists.
- `Dockerfile.base`'s `# BASE_REBUILD_DATE:` marker, itself stale (2026-07-13,
three base rebuilds ago). `base_tag` hashes Dockerfile.base's *content*,
comments included, so demanding it be current would force a ~60 min base
rebuild on a release that touched no base files at all. It is free to fix
while the base is *already* rebuilding, and expensive at any other moment.
That cost asymmetry is now written into the release checklist rather than
enforced.
**Release checklist step 3 rewritten** (`AGENTS.md`) around the mechanism that
made this expensive: `docker-publish.yml` runs `actions/checkout@v4` with no
`ref:`, so every job reads `github.ref` — the tag. Docs must be correct *before*
tagging; afterwards the only routes are re-pointing the tag (its own hazard —
v1.8.14 went `601fc98` → `361babd` and broke deploy verification until
`git fetch --tags --force`) or waiting for the next release. The step now also
names what the gate cannot see, so "gate is green" is not mistaken for "docs are
true". The same reflex went into the `ci-release-watcher` skill, as the first
correctness rule — it is the only one that expires once the tag exists.
Also fixed in passing: README's `pi-devbox-version` sample was v1.5.0-era and
structurally outdated (it predated the `palace:` line the surrounding prose
advertises, the `pi-atelier` component, and the whole `skills:` block). Replaced
with real observed output rather than hand-written text. `DOCKER_HUB.md`'s "7
user-facing extensions" was **verified correct**; its "29 `mempalace_*` tools"
is stale (a live client shows 45) but left alone rather than corrected on a
guess, since that count cannot be attributed to the baked 3.9.0 server without
measuring it.
## v1.9.0 — 2026-09-10 (tagged, never published — superseded by v1.9.1)
> This tag exists in git but no image was ever pushed for it: `smoke` failed
> three assertions and skipped every downstream job. Everything below ships in
> **v1.9.1**, whose entry explains the three failures and their fixes. Kept as
> its own section rather than folded away, because the tag is real and someone
> will eventually find it and wonder why Docker Hub has no v1.9.0.
**`shellcheck` is now in the image, because the release gate it depends on could
not be run by anyone.** v1.8.14 made shell lint a release gate: `scripts/lint-shell.sh`
@@ -92,6 +870,20 @@ derivation's `mine` query at the newest end (`order: "desc"`); with the previous
default `asc` + `limit: 100`, a device passing 100 authored events would have its
recent replies fall out of the join window and see answered asks resurface.
> **Corrected 2026-09-14 (`pi@tor-ms22`, the device in the measured cost above).**
> "This image still pins `e45f6b4`" and "until an image bakes `e2b060a` or later"
> were true when written on 2026-09-09 and are **false for the running fleet**.
> **v1.9.1 bakes `e68ee20`**, a descendant of `e2b060a`, so requester-side
> withdrawal is LIVE wherever v1.9.1 runs. Measured from the published image's own
> label (`se.jordbo.pi-devbox.mempalace-toolkit-ref`) rather than from these
> notes, and cross-checked by ancestry and by the baked `mempalace.ts` sha
> differing from v1.8.14's. Nothing here was mis-stated on purpose:
> `ARG MEMPALACE_TOOLKIT_REF=main` is resolved to a commit SHA by CI at build
> time, so the release absorbed the commit without anybody having to name it,
> while this paragraph went on asserting it had not. Left standing rather than
> rewritten — the sentence is the evidence for how the drift happened. See
> Unreleased.
**Four small packages, each chosen from a gap that was measured rather than
imagined.** All four were picked by looking back at a real session — the
`gitea.egl.lan`/FreeIPA debugging of 2026-09-09..10 — and asking which absences
+5 -5
View File
@@ -8,12 +8,12 @@ A self-contained Docker container for the [pi coding-agent](https://github.com/e
| Tag | Architectures | Size (compressed) | What you get |
|---|---|---|---|
| `joakimp/pi-devbox:latest` | amd64, arm64 | ~1.1 GB | Self-contained: base + pi `{{PI_VERSION}}` + companions |
| `joakimp/pi-devbox:latest` | amd64, arm64 | ~1.23 GB | Self-contained: base + pi `{{PI_VERSION}}` + companions |
| `joakimp/pi-devbox:vX.Y.Z` | amd64, arm64 | same | Pinned semver release |
| `joakimp/pi-devbox:latest-studio` | amd64, arm64 | ~1.15 GB | `latest` + [pi-studio](https://github.com/omaclaren/pi-studio): browser prompt editor, KaTeX/Mermaid preview, tmux-backed literate REPLs |
| `joakimp/pi-devbox:latest-studio` | amd64, arm64 | ~1.25 GB | `latest` + [pi-studio](https://github.com/omaclaren/pi-studio): browser prompt editor, KaTeX/Mermaid preview, tmux-backed literate REPLs |
| `joakimp/pi-devbox:vX.Y.Z-studio` | amd64, arm64 | same | Pinned semver studio release |
| `joakimp/pi-devbox:base-latest` | amd64, arm64 | ~1.0 GB | Base layer alias (internal building block; pull `:latest` instead) |
| `joakimp/pi-devbox:base-<hash>` | amd64, arm64 | ~1.0 GB | Content-addressed base; immutable. Stable parent for variant rebuilds. |
| `joakimp/pi-devbox:base-latest` | amd64, arm64 | ~1.17 GB | Base layer alias (internal building block; pull `:latest` instead) |
| `joakimp/pi-devbox:base-<hash>` | amd64, arm64 | ~1.17 GB | Content-addressed base; immutable. Stable parent for variant rebuilds. |
> **pi-studio (`-studio` tags):** launch with `/studio --no-browser --port 8765` inside a pi session. The server binds `127.0.0.1` **inside the container**, so reach it via host networking or a loopback bridge (and `ssh -L` for a remote host; mosh needs a parallel `ssh -L`). Full recipe: [README → Using pi-studio](https://gitea.jordbo.se/joakimp/pi-devbox#using-pi-studio--studio-variant).
@@ -94,7 +94,7 @@ The entrypoint deploys/registers all of these on first container start. Re-runni
uv run --with jupyterlab jupyter lab --no-browser --port 8888
uv run --with marimo marimo edit
```
- **Node.js** v22 + npm (used by pi itself)
- **Node.js** v24 LTS + npm (used by pi itself)
- **Rust** — `rustup-init` is on PATH; install toolchains on demand
- **Go** — opt-in via `--build-arg INSTALL_GO=true` if rebuilding from source
+19 -2
View File
@@ -14,7 +14,7 @@
# content-addressed over this file, so any byte change invalidates the
# cache. Recommended cadence: once per release for security updates.
#
# BASE_REBUILD_DATE: 2026-07-13 (Unreleased — agent-browser CLI + Playwright Chromium for headless browser automation; prior: typst PDF engine + xz-utils + pandoc typst-template default-font patch)
# BASE_REBUILD_DATE: 2026-09-19 (v1.9.3 — mempalace-version label, mempalace-toolkit 817b3a8 mine deadline, pi-extensions skill floor 25c1265; the marker had been stale since 2026-07-13 through the v1.9.0 Node 24 and v1.9.2 npm-residue rebuilds, updated now because the base is rebuilding anyway)
#
# ── Lineage note ─────────────────────────────────────────────────────
# Adapted from opencode-devbox/Dockerfile.base (commit before v1.16.2).
@@ -452,7 +452,10 @@ RUN ARCH=$(case "${TARGETARCH}" in amd64) echo "x86_64" ;; arm64) echo "aarch64"
# uninterruptibly. A stall-kill is no longer a permanent latch either: the
# next tool call respawns the server with capped exponential backoff (the
# budget resets on any successful response). Tunables:
# MEMPALACE_MCP_TIMEOUT_MS (default 60000), MEMPALACE_MCP_INIT_TIMEOUT_MS
# MEMPALACE_MCP_TIMEOUT_MS (default 60000; the feed's `mempalace_mine` carries
# its own longer MEMPALACE_FEED_MINE_TIMEOUT_MS, default 300000, since toolkit
# 817b3a8 — before that the 60 s deadline cut every honest mine off),
# MEMPALACE_MCP_INIT_TIMEOUT_MS
# (default 300000 — generous so a genuine first cold-open isn't killed),
# MEMPALACE_MCP_MAX_RESPAWNS (default 2; 0 disables self-heal),
# MEMPALACE_MCP_RESPAWN_BACKOFF_MS (default 1000); timeouts of 0 disable.
@@ -549,6 +552,20 @@ ARG INSTALL_MEMPALACE=true
# so they stay dark until synlig is redeployed — a client bump alone cannot
# light them up.
ARG MEMPALACE_VERSION=3.9.0
# Recorded as a label HERE, not in Dockerfile.variant, for three reasons: the
# value lives next to the ARG that defines it (a second copy in the variant
# would be one more pin able to drift, which is the class check-doc-drift.sh
# exists to catch); labels are inherited by every image built FROM this one, so
# both variants carry it with no build-arg to plumb through four call sites;
# and inheritance means the label states the pin of the base the variant
# ACTUALLY built on — which is the question when base-decide cache-hits an
# older base. Like every se.jordbo.pi-devbox.* label this records INTENT; the
# ground truth is /etc/pi-devbox/build-manifest.json's mempalace_version, read
# from the installed binary, and scripts/smoke-test.sh asserts the two agree.
# check-doc-drift.sh check 9 reads this off the last published image so that a
# pin bump must be named in the CHANGELOG — until this label ships, that
# component reports SKIP (label absent on the published release), not OK.
LABEL se.jordbo.pi-devbox.mempalace-version="${MEMPALACE_VERSION}"
ENV UV_TOOL_DIR=/opt/uv-tools
ENV UV_TOOL_BIN_DIR=/usr/local/bin
RUN if [ "${INSTALL_MEMPALACE}" = "true" ]; then \
+90 -1
View File
@@ -196,6 +196,71 @@ RUN set -e && \
done; \
return 1; \
} && \
# prune_foreign_natives: npm 11 (shipped with Node 24) installs EVERY optional
# platform package of a native dependency, not just the one matching the host.
# TWO families are affected in this image, and BOTH have been measured — add a
# family here only after measuring it, never by widening the pattern on a hunch:
#
# @esbuild/<platform> 26 dirs, 284 MB (found first, v1.9.1)
# @mariozechner/clipboard-<triple> 11 dirs, 12 MB per site, 10 MB foreign
#
# esbuild declares those with os/cpu constraints, but npm 11 ignores them and
# ALSO ignores --os/--cpu and an npmrc carrying os=/cpu= (all three measured).
# So prune explicitly, keeping only the host platform, computed from
# `node -p process.arch` so one line stays correct on amd64 and arm64.
# Measured on pi-fork's tree: npm 10.9.8 -> 165 MB, npm 11.19.0 -> 449 MB,
# and the 165 MB figure reproduces what v1.9.0's predecessor actually shipped.
#
# WHY THE CLIPBOARD FAMILY WAS ADDED (2026-09-11): v1.9.1 pruned @esbuild only
# and still shipped +131 MB compressed over v1.8.14. That residual was
# attributed by listing the PUBLISHED arm64 layer tarballs straight from the
# registry (there is no docker CLI inside the container, so `docker history`
# was not available): +110 MB /root/.npm/_cacache (purged below) and +21 MB of
# clipboard platform packages across the two install sites — 131 MB total, so
# the delta is now fully accounted for with no unexplained remainder.
#
# Keeping linux-$arch-{gnu,musl} is deliberate: clipboard's napi-rs loader
# tries ./<name>.node then the platform package, per platform in try/catch, and
# chooses gnu vs musl at runtime from its own isMusl() probe — so both host-arch
# branches must survive. The musl package is a 420-byte stub, i.e. free. The
# bare wrapper `@mariozechner/clipboard` has no hyphen suffix and therefore
# cannot match the regex below. Verified on arm64 against a copy of the real
# tree before this was written: after pruning to those two,
# require('@mariozechner/clipboard') still loads and exports all 18 functions.
# esbuild likewise still compiles TS via transformSync at both install sites.
# This removes dead weight, not function.
#
# MUST run in the SAME layer as the npm installs above: deleting in a later RUN
# leaves the bytes in this layer and shrinks the image by nothing.
# NOTE the single backslash in -printf '%f\n': Docker passes '\\n' through
# verbatim, so v1.9.1's doubled version printed a mangled "li ux-arm64"
# (find emitted a literal backslash, then `tr` translated the n out of the
# name). Confirmed from the published image's own recorded created_by.
prune_foreign_natives() { \
arch="$(node -p process.arch)"; \
find /usr/lib/node_modules /opt -type d -regex '.*/@esbuild/[^/]+' \
! -name "linux-$arch" -prune -exec rm -rf {} + ; \
find /usr/lib/node_modules /opt -type d -regex '.*/@mariozechner/clipboard-[^/]+' \
! -name "clipboard-linux-$arch-gnu" ! -name "clipboard-linux-$arch-musl" \
-prune -exec rm -rf {} + ; \
echo "native platform dirs kept: $(find /usr/lib/node_modules /opt -type d \( -regex '.*/@esbuild/[^/]+' -o -regex '.*/@mariozechner/clipboard-[^/]+' \) -printf '%f\n' 2>/dev/null | sort | uniq -c | tr '\n' ' ')"; \
} && \
# purge_build_caches: the build's own download caches are NOT free — they land
# in whichever layer created them. Measured on the published v1.9.1 arm64
# variant layer: root/.npm/_cacache was 145.2 MB of a 401.9 MB layer (35.2 MB
# in v1.8.14), the single biggest item in the +131 MB residual, because npm 11
# caches every platform tarball it fetched — including the ones just pruned.
# Nothing at runtime reads it: the build runs as root, the container runs as
# `developer` with its own cache under $HOME (and $HOME/.pi is a volume).
# DELIBERATELY NOT purged here: /tmp/node-compile-cache (1.3 MB, written by
# `pi --version` below). The manifest RUN at the end of this file calls
# `pi --version` again, so deleting it here only relocates those bytes into
# that layer instead of removing them from the image — measured, not assumed:
# today the manifest layer is 128 kB precisely because it finds the cache warm.
purge_build_caches() { \
npm cache clean --force >/dev/null 2>&1 || true; \
rm -rf /root/.npm; \
} && \
if [ "${PI_VERSION}" = "latest" ]; then \
NPM_CONFIG_PREFIX=/usr npm install -g @earendil-works/pi-coding-agent ; \
else \
@@ -209,6 +274,8 @@ RUN set -e && \
git_fetch_ref "${PI_ATELIER_REPO}" "${PI_ATELIER_REF}" /opt/pi-atelier && \
(cd /opt/pi-fork && npm install --omit=dev --no-audit --no-fund) && \
(cd /opt/pi-observational-memory && npm install --omit=dev --no-audit --no-fund) && \
prune_foreign_natives && \
purge_build_caches && \
echo "pi-toolkit at $(cd /opt/pi-toolkit && git rev-parse --short HEAD)" && \
echo "pi-extensions at $(cd /opt/pi-extensions && git rev-parse --short HEAD)" && \
echo "pi-fork at $(cd /opt/pi-fork && git rev-parse --short HEAD)" && \
@@ -309,6 +376,26 @@ ARG PI_STUDIO_REF=main
ARG PI_STUDIO_VERSION=none
RUN if [ "${INSTALL_STUDIO}" = "true" ]; then \
set -e; \
# Same prune + cache purge as the main install RUN — see the comments there.
# They have to be redefined because shell functions do not survive across
# layers, and they have to run in THIS layer because pi-studio's npm install
# happens here: deleting in a later RUN would leave the bytes in this layer
# and shrink nothing. pi-studio pulls its own pi-coding-agent copy, so it is
# a third ~274 MB site on top of the two in the non-studio variant — and its
# npm install refills /root/.npm, which the main RUN emptied in ITS layer.
prune_foreign_natives() { \
arch="$(node -p process.arch)"; \
find /usr/lib/node_modules /opt -type d -regex '.*/@esbuild/[^/]+' \
! -name "linux-$arch" -prune -exec rm -rf {} + ; \
find /usr/lib/node_modules /opt -type d -regex '.*/@mariozechner/clipboard-[^/]+' \
! -name "clipboard-linux-$arch-gnu" ! -name "clipboard-linux-$arch-musl" \
-prune -exec rm -rf {} + ; \
echo "native platform dirs kept: $(find /usr/lib/node_modules /opt -type d \( -regex '.*/@esbuild/[^/]+' -o -regex '.*/@mariozechner/clipboard-[^/]+' \) -printf '%f\n' 2>/dev/null | sort | uniq -c | tr '\n' ' ')"; \
}; \
purge_build_caches() { \
npm cache clean --force >/dev/null 2>&1 || true; \
rm -rf /root/.npm; \
}; \
rm -rf /opt/pi-studio && mkdir -p /opt/pi-studio && \
git -C /opt/pi-studio init -q && \
git -C /opt/pi-studio remote add origin "${PI_STUDIO_REPO}" && \
@@ -320,6 +407,8 @@ RUN if [ "${INSTALL_STUDIO}" = "true" ]; then \
done; \
[ "$ok" = "1" ] && \
(cd /opt/pi-studio && npm install --omit=dev --no-audit --no-fund) && \
prune_foreign_natives && \
purge_build_caches && \
echo "pi-studio at $(cd /opt/pi-studio && git rev-parse --short HEAD)"; \
fi
@@ -392,7 +481,7 @@ ARG MEMPALACE_TOOLKIT_REF=main
# no ~67-minute base rebuild. (scripts/check-base-hash.sh scans only
# Dockerfile.base, so no folding into the base hash is required — nor would
# it be correct, since this ARG changes nothing about the base's contents.)
ARG SKILLSET_SNAPSHOT_REF=4d7c0ea9caeb3a1d6d9b04cf34f3fca5f9df4985
ARG SKILLSET_SNAPSHOT_REF=e9e45f7acdde490c3b5d24ce5f508bff8785c2c7
# Dockerfile.base sets description="pi-devbox — base image (variant-independent)"
# and every variant INHERITS it, so both published images used to advertise
+28 -21
View File
@@ -175,12 +175,10 @@ Currently published:
| `joakimp/pi-devbox:latest-studio` | `latest` + [pi-studio](https://github.com/omaclaren/pi-studio) (browser prompt editor, KaTeX/Mermaid preview, tmux-backed literate REPLs) | ~3.25 GB |
| `joakimp/pi-devbox:vX.Y.Z-studio` | pinned-version studio equivalent | ~3.25 GB |
Planned for an upcoming minor release:
- *(shipped in Unreleased/base)* **PDF export from Studio/pandoc** now works:
the base image ships **`typst`** as the PDF engine (`pandoc --pdf-engine=typst`),
a single ~30 MB static binary — no separate `-tex` variant needed.
`texlive-xetex` stays the higher-fidelity fallback (install on demand).
Both variants ship **`typst`** as the pandoc PDF engine
(`pandoc --pdf-engine=typst`), a single ~30 MB static binary, so PDF export from
Studio/pandoc works out of the box — no separate `-tex` variant needed.
`texlive-xetex` stays the higher-fidelity fallback (install on demand).
## Using pi-studio (`-studio` variant)
@@ -903,8 +901,10 @@ docker inspect --format '{{json .Config.Labels}}' joakimp/pi-devbox:latest | jq
```
`org.opencontainers.image.{version,revision,created}` plus
`se.jordbo.pi-devbox.*-ref` record the intended pi version and companion
refs. The on-disk `/etc/pi-devbox/build-manifest.json` records **ground
`se.jordbo.pi-devbox.*-ref` and `se.jordbo.pi-devbox.*-version` record the
intended pi and mempalace versions and companion refs (`mempalace-version` is
set in `Dockerfile.base` and inherited, so it names the pin of the base the
image actually built on). The on-disk `/etc/pi-devbox/build-manifest.json` records **ground
truth** — the actual checked-out commit of each `/opt` clone, the live
`pi --version`, and (from v1.8.6) the live `mempalace --version` of the
installed palace core — so a tag is reconstructable after CI logs rotate:
@@ -919,16 +919,23 @@ through `jq` yourself:
```console
$ pi-devbox-version
pi-devbox v1.5.0
built: 2026-07-13T17:53:16Z (source d68674d11e06)
pi: 0.80.6
pi-devbox v1.8.14
built: 2026-09-08T21:54:07Z (source 361babd4fd61)
pi: 0.85.1
palace: 3.9.0
components:
pi-toolkit: 9a8f6faeaa08
pi-extensions: 61c98e004e3d
pi-fork: 4a09af4ef527
pi-observational-memory: 27a5195eaf90
mempalace-toolkit: 96699f2a1781
pi-studio: 2ef38ef31cea
pi-toolkit: adfb553f5c8a
pi-extensions: 2610545c83bb
pi-fork: e69725c39603
pi-observational-memory: ce9fc982b3a2
pi-atelier: 734258bbcb62
mempalace-toolkit: e45f6b430181
pi-studio: e04fc7aa3275
skills:
credential-incident-response baked
mempalace live /workspace/skillset @ 4d7c0ea (identical to baked snapshot)
pi-devbox-environment baked
pi-extensions baked
```
It also flags **live drift** — if `pi --version` no longer matches what was
@@ -1093,7 +1100,7 @@ persisted volumes survived, and pi runtime wiring is intact:
```bash
./scripts/recreate-sanity-check.sh # auto-detects variant
./scripts/recreate-sanity-check.sh --expected-image-version 1.8.9 # assert the pi-devbox release tag
./scripts/recreate-sanity-check.sh --expected-version 0.84.4 # assert the pi coding agent version
./scripts/recreate-sanity-check.sh --expected-version 0.85.1 # assert the pi coding agent version
```
Those are **two different versions**, and the flags are not interchangeable:
@@ -1132,9 +1139,9 @@ resolved to `latest` at build time:
| Component | Pin | Where |
|---|---|---|
| pi | `0.84.4` | `ARG PI_VERSION` — `Dockerfile.variant` |
| pi-atelier | `v0.10.0` | `ARG PI_ATELIER_REF` — `Dockerfile.variant` |
| mempalace | `3.8.0` | `ARG MEMPALACE_VERSION` — `Dockerfile.base` |
| pi | `0.85.1` | `ARG PI_VERSION` — `Dockerfile.variant` |
| pi-atelier | `v0.10.1` | `ARG PI_ATELIER_REF` — `Dockerfile.variant` |
| mempalace | `3.9.0` | `ARG MEMPALACE_VERSION` — `Dockerfile.base` |
The objective is **not** to freeze versions. Bumping is routine — usually one
line plus a changelog note. The objective is that adopting a new upstream
@@ -535,7 +535,10 @@ Two consequences worth internalising:
- **"Seen, not doing it" is a legitimate ack** — `status="blocked"` or
`"superseded"` plus the reason. Silence is not, and it is not merely rude:
with no terminal event of yours to join to, the ask stays in the owed set
indefinitely and there is nothing anyone can do about it from the other end.
indefinitely. The *original requester* — and nobody else — can release it from
the other end, but only by saying so explicitly: see **Withdrawing an ask you
sent** below. That is a release by the asker, not an escape for the answerer.
While the ask still stands, only *your* terminal event clears it.
- **Nothing expires, and it should not.** An `open` with no terminal reply is
still live by definition, and the finished threads are valuable history. If
content is genuinely perishable ("do not push to main for the next hour"), say
@@ -570,6 +573,24 @@ Two consequences worth internalising:
be matched to it at all.
- **Corrections are new events, never edits.** Say explicitly what you retract
and name the id — drawer or event — that carried the withdrawn claim.
- **Withdrawing an ask you sent: state it, never imply it.** Your release only
counts when the terminal event (a) comes from the same `from_agent` that sent
the ask, (b) is directed at that recipient exactly — never `*`, so a broadcast
can neither oblige nor release, (c) carries a terminal status (`claimed` and
`ready` are not terminal and do not release anything), (d) is strictly after
the ask, (e) joins it via `ack_of` or the same `correlation_id`, **and (f)
names that ask in `metadata.withdraws` or `metadata.closes`.** Prose in the
body does not count, and neither does a bare terminal event on the
correlation: inferring release from *any* terminal would let your own
bookkeeping silently delete a real obligation, so the release must be stated.
Needs toolkit ≥ `e2b060a` (image ≥ `v1.9.1`) — check with
`grep -c isWithdrawn /opt/mempalace-toolkit/extensions/pi/mempalace.ts` and
read `0` as "my withdrawal will have no effect on their mailbox". Measured
cost of getting it wrong: a `v1.8.13` rollout ask was withdrawn by its sender,
who recorded it as done; the recipient's derivation never saw the release and
still reported the ask owed **41 hours later**, for a release that device
never installed — and the asymmetry was invisible from the sender's side
(RFC 003 §3.3 clause 4).
- **Put a retraction where the reader will look.** A *directed open ask* reaches a
live agent's mailbox; a **terminal-status event reaches no mailbox at all**, and
a *drawer* is what a future semantic search finds. If you filed advice as a
@@ -1,17 +1,17 @@
---
name: pi-extensions
description: >-
Use the pi extensions (pi-fork, pi-observational-memory, ssh-controlmaster) effectively in the pi coding agent harness. Load this skill only when running inside pi (detection - `fork` and `recall` are present in your tool list, or `pi --ssh` was used to start the session). pi-fork dispatches focused subtasks to forked agents at fast/balanced/deep effort tiers; pi-observational-memory compacts long sessions into recallable observations + reflections; ssh-controlmaster rewires pi's read/write/edit/bash tools to execute on a remote host over a multiplexed SSH connection. Also covers the context ladder L0-L4 and when to reach for the separate `pi-task` CLI instead of `fork` - isolated child, immutable spec, machine-checked envelope, write-boundary diff. This skill covers tier selection, task design, boundary discipline, when to use recall, and remote-pi mechanics.
Use the pi extensions (pi-fork, pi-observational-memory, ssh-controlmaster) effectively in the pi coding agent harness. Load this skill only when running inside pi (detection - `fork` and `recall` are present in your tool list, or `pi --ssh` was used to start the session). pi-fork dispatches focused subtasks to forked agents at fast/balanced/deep effort tiers; pi-observational-memory compacts long sessions into recallable observations + reflections; ssh-controlmaster rewires pi's read/write/edit/bash tools to execute on a remote host over a multiplexed SSH connection. Also covers the context ladder L0-L4 and the `task` tool (pi-task: isolated child, immutable spec, machine-checked envelope, write-boundary diff) that is the DEFAULT for delegated work, with `fork` reserved for read-only exploration and parallel opinions, plus the `fork-gate` hook that enforces the split. This skill covers rung and tier selection, task design, boundary discipline, when to use recall, and remote-pi mechanics.
---
# Pi Extensions: pi-fork, pi-observational-memory, ssh-controlmaster
# Pi Extensions: pi-fork + task/fork-gate, pi-observational-memory, ssh-controlmaster
## When to Load This Skill
Load only when **both** of these are true:
1. You are running inside the **pi coding agent harness** (not Claude Code, not opencode, not any other harness).
2. The `fork` and/or `recall` tools appear in your available tool list, **or** the session was started with `pi --ssh ...`.
2. The `fork`, `task` and/or `recall` tools appear in your available tool list, **or** the session was started with `pi --ssh ...`.
If you do not see those tools, this skill does not apply — skip it. Other harnesses do not have these extensions and the patterns below will not work there.
@@ -71,7 +71,78 @@ ssh-controlmaster is orthogonal but composes cleanly: when pi is operating remot
---
## Part 1: pi-fork
## Part 1: delegating work — `task` and `fork`
### Decide the rung BEFORE the brief (read this first)
Two tools run a child agent. They differ in one thing, and it decides the
quality of what comes back: **what the child sees.**
| tool | child sees | rung | gives you | use for |
|---|---|---|---|---|
| **`task`** (pi-extensions `task.ts`, wraps `pi-task`) | **only your spec** — goal, named files, curated facts | L0–L2 | immutable spec, PASS/FAIL envelope, per-root boundary diff, audit dir, budgets | **any delegated work that changes files or must obey a rule** — the default |
| **`fork`** (pi-fork) | **your entire branch**, brief appended last | L4 | prose report, effort tiers, parallel dispatch from one message | read-only exploration that needs this conversation; N independent opinions |
**Pre-flight before any `fork(...)` — one *yes* makes it a `task`:**
1. Does the brief say *do not / only / never / must not*?
2. Will the child write, edit, commit or push anything?
3. Do I want a PASS/FAIL I can check, rather than prose?
Why the text alone did not work (and why this is now enforced): the rule above
lived in this skill and in the global AGENTS.md for months and was still
violated by agents that had just read it — five fork briefs in one session on
2026-09-17, all carrying "do not", one of which returned confident verbatim
quotes that did not exist. `fork` is a *tool*: its self-recommending
description ("implementation, testing, review…") is in the tool list every turn
and survives compaction; this skill is gone after the first compaction, and
`pi-task` was a CLI to be remembered. Two structural fixes shipped 2026-09-19:
- **`task` is a tool** (`extensions/task.ts`), so both rungs sit in the tool
list with the decision rule in their descriptions, and `promptGuidelines`
puts the rule in the system prompt where compaction cannot remove it. It
rejects, before spending a model run, the two spec errors that make a
violation certain (see "roots" below) and serialises overlapping tasks.
- **`fork-gate`** (`extensions/fork-gate.ts`) is a `tool_call` hook that BLOCKS
a fork whose brief contains a prohibition, a write boundary or a clause-initial
file-changing imperative, and returns the `task(...)` call to make instead.
It matches wording, not intent: a genuinely read-only exploration brief that
trips it is rephrased, and one that cannot be rephrased without its
prohibition needed `task` all along. `PI_FORK_GATE=off` logs instead of
blocking; `/ext` disables it entirely.
Minimal call:
```
task(id="slug", goal="…verbatim; the child has NO other context…",
deliverable="…exact shape wanted…", effort="fast|balanced|deep",
read_only=false,
roots=["/abs/repo/docs", "/abs/repo/src"], # WATCHED, each diffed alone
write_allowed=["/abs/repo/docs"], # exact subset of roots
facts=["verified fact"], files=["/abs/path/to/read"])
```
**Roots — the two errors the tool refuses up front.** Every root is diffed on
its own and a delta is allowed only if that *exact root string* is in
`write_allowed`. So (a) `write_allowed` must be a subset of `roots`, not a
subdirectory of one, and (b) a writable root must not lie inside a watched-only
root — the parent's porcelain would change and register a violation every time.
List the writable part as its own root and leave the enclosing repo out. This is
the shape the 2026-09-17 migration tasks used (sibling roots, `write_allowed`
naming four of them) and it passed cleanly.
**Overlap.** Sibling `task` calls whose roots overlap would see each other's
writes as violations; the tool runs them one after another automatically. Do
not rely on that for ordering *semantics* — if B needs A's output, call B after
A returns.
**What isolation does not fix.** L0 removes the *narrative* failures (parent
voice, invented continuity, ignored prohibitions). It does not remove
confabulation: an under-specified spec still gets a confident deliverable. The
report prints the evidence pointers under a "SPOT-CHECK THESE" heading for a
reason.
Everything below about tiers, brief design and boundary discipline applies to
**both** tools — a `task` spec is a brief too.
### Effort tier mapping
@@ -85,15 +156,15 @@ Configured in `~/.pi/agent/settings.json` under `pi-fork.effortProfiles`. The co
**Rule of thumb:** start at `balanced` unless you have a specific reason to go up or down. Going too cheap on a deep task wastes a fork; going too expensive on a mechanical task is just slow.
### When to fork vs. do it yourself
### When to delegate vs. do it yourself
Fork when **any** of:
Having chosen the rung above, delegate (either tool) when **any** of:
- The task requires reading many files whose contents you don't need to keep in your main context afterwards (the fork returns a dense summary; raw file contents stay in the fork's context and are discarded).
- You want to run multiple analyses in **parallel** (especially: comparing N options, where independent reasoning is itself a signal — see "parallel forks" below).
- The task is well-scoped enough to specify completely up front and well-bounded enough that returning a dense report is more useful than continuing the dialogue.
- You are about to do something that would burn a lot of tokens on tool calls (long file reads, many bash invocations) whose output you will mostly discard.
Don't fork when:
Don't delegate when:
- The work fits in your current context budget without crowding out what comes next.
- The task is exploratory and you'll need to iterate based on what you find (forking turns iteration into round-trips with full task-spec rewrites).
- You need to make decisions during the work that depend on context only the main thread has.
@@ -175,11 +246,12 @@ sits at one extreme of it. Five rungs:
| **L3** | a **truncated tail** of the parent branch | *nothing implements this* — would need a new spec key plus `--session <trimmed snapshot>` | **no** |
| **L4** | the **entire** parent branch | `fork(task=…)` — `getHeader()+getBranch()`, no offset or limit anywhere in the call chain | yes |
**`pi-task` is a CLI, not an extension — it will never appear in your tool list.**
Invoke it with `bash`: `/opt/pi-toolkit/bin/pi-task run <spec.json>` (source at
`/workspace/pi-toolkit/bin/pi-task`, `schema` subcommand prints the spec fields).
It reads an immutable JSON spec, and "inherit the session" is not expressible in
that schema — the isolation is structural, not a request.
**`pi-task` is a CLI (`/opt/pi-toolkit/bin/pi-task`, `schema` prints the spec
fields); the `task` tool from pi-extensions wraps it** so it appears in your tool
list next to `fork`. If the tool is absent, invoke the CLI with `bash`:
`/opt/pi-toolkit/bin/pi-task run <spec.json>`. Either way it reads an immutable
JSON spec, and "inherit the session" is not expressible in that schema — the
isolation is structural, not a request.
**Choose the lowest rung that can do the job:**
@@ -292,7 +364,17 @@ When entries conflict, **the most recent observation reflects the latest known s
## Quick Reference
```
task(id, goal, deliverable, effort, read_only, roots, write_allowed, facts, files, commands, wall_s, usd)
- L0-L2: isolated child sees ONLY the spec — DEFAULT for work that writes or has rules
- roots[] = WATCHED (each diffed alone); write_allowed[] = exact subset of roots,
never nested inside a watched-only root (the tool rejects both errors up front)
- envelope must parse or the run FAILED; spot-check evidence pointers
- overlapping-root tasks are serialised; audit: ~/.pi/agent/pi-task/<stamp>-<id>/result.json
- CLI fallback: bash /opt/pi-toolkit/bin/pi-task run <spec.json> (schema | selftest | run --dry-run)
fork(task=..., effort=fast|balanced|deep) # L4: child inherits your WHOLE branch
- ONLY for read-only exploration needing this conversation, or N parallel opinions
- fork-gate BLOCKS briefs with do-not/only/never, write boundaries, or "edit/commit/fix …"
- state decision authority explicitly
- pass verified context up front
- specify deliverable shape
@@ -302,13 +384,6 @@ fork(task=..., effort=fast|balanced|deep) # L4: child inherits your WHOLE b
- write-capable? demand "What I did NOT do", then verify from git/fs, not the report
- prohibition in the brief => not a `fast` task
bash: /opt/pi-toolkit/bin/pi-task run <spec> # L0-L2: isolated child, NOT a tool
- schema | selftest | run [--dry-run]
- context.facts (pasted) / .files (names only) / .commands
- roots[] = WATCHED, write_allowed[] = CHANGEABLE subset
- envelope must parse or the run FAILED
- audit + cost: ~/.pi/agent/pi-task/<stamp>-<id>/result.json
recall(id=<12-char-hex>)
- only when stakes justify the cost
- id must already be visible in your context
@@ -317,8 +392,9 @@ recall(id=<12-char-hex>)
```
~/.pi/agent/settings.json
pi-fork.effortProfiles — model + thinking-depth per tier
pi-fork.effortProfiles — model + thinking-depth per tier (used by BOTH fork and task)
pi-fork.defaultEffort — usually "balanced"
env PI_FORK_GATE=off — fork-gate logs instead of blocking (default: block)
observational-memory.* — token thresholds, model, agentMaxTurns
observational-memory.debugLog: true — opt-in NDJSON telemetry at
~/.pi/agent/observational-memory/debug/<session>.ndjson (off by default)
+698
View File
@@ -0,0 +1,698 @@
#!/usr/bin/env bash
# check-doc-drift.sh — fail when a hand-maintained doc claim contradicts the
# build files it describes.
#
# THE DEFECT CLASS THIS EXISTS TO CATCH, measured 2026-09-10 while preparing
# v1.9.0. Five separate claims had rotted, all of them the same shape: a fact
# written once by hand, in a file nothing verifies, about a value that lives
# somewhere else and moved.
#
# 1..3. README.md's "Version pins" table was wrong on EVERY row — pi `0.84.4`
# vs ARG PI_VERSION=0.85.1, pi-atelier `v0.10.0` vs v0.10.1, mempalace
# `3.8.0` vs 3.9.0. That table is the worst possible place for this: it
# exists precisely to be the reviewable record of what is deliberately
# frozen, so when it lies, the review it enables is worthless.
# 4. README.md carried a "Planned for an upcoming minor release" section
# listing typst PDF export, which had ALREADY SHIPPED, tagged with a
# self-contradicting "(shipped in Unreleased/base)" marker. The
# CHANGELOG had already documented three earlier instances of exactly
# this stale-"Unreleased"-pointer class (see its v1.8.7 notes).
# 5. DOCKER_HUB.md claimed "Node.js v22" while this release ships Node 24.
# This one is the reason the gate exists at all: DOCKER_HUB.md is
# PUBLISHED. `update-description` in docker-publish.yml POSTs it to Hub
# as full_description on every tag, so unlike README.md — which no
# workflow or gate reads — a stale claim here is what users see.
#
# WHY A GATE AND NOT "REMEMBER TO CHECK". DOCKER_HUB.md had gone eight releases
# (v1.8.6 → v1.9.0) without a touch. Nothing generates it and nothing verifies
# it; the only mechanism keeping it true was whoever remembered. That is the
# same failure mode check-skill-floor.sh was written for, and the same fix:
# convert "someone remembers" into "CI refuses".
#
# TWO CLASSES OF CHECK, DELIBERATELY. Checks 1-7 compare a doc string to a
# value that EXISTS IN THIS REPO, so they can never be wrong about the world and
# need no network, no token, and no built image. Checks 8-9 compare against what
# is PUBLISHED (Docker Hub's measured sizes; the ref labels baked into the last
# released image), because those claims have no in-repo anchor at all and had
# rotted for exactly that reason. They need the network and therefore SKIP,
# loudly and counted, when it is absent -- a skip is neither OK nor a failure,
# because printing an unverified claim as OK is the habit this file exists to
# break, while failing on a third party's uptime would make every release
# hostage to it. Claims that need a RUNNING CONTAINER (the "N mempalace_* tools"
# count, uncompressed on-disk sizes) are still not gated here; assert them in
# scripts/smoke-test.sh where a real image is available.
#
# DELIBERATELY NOT GATED: Dockerfile.base's `# BASE_REBUILD_DATE:` comment, which
# is also stale (2026-07-13, three base rebuilds ago). base_tag is a hash of
# Dockerfile.base's CONTENT plus rootfs/, comments included, so a gate that
# demanded that comment be current would force a ~60 min base rebuild on any
# release that touched no base files at all. Fix it when you are already
# rebuilding the base — then it is free. This is a real cost asymmetry, not
# laziness.
#
# EXIT CODES (same contract as lint-shell.sh and check-skill-floor.sh):
# 0 every checked claim matches
# 1 at least one claim has drifted
# 2 cannot run (a file or ARG this gate reads is missing/unparseable)
# A gate that cannot run must not pass, so a missing input is 2, never 0.
set -euo pipefail
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
cd "$REPO_ROOT"
README="README.md"
HUB="DOCKER_HUB.md"
DF_VARIANT="Dockerfile.variant"
DF_BASE="Dockerfile.base"
# Docker Hub rejects a full_description longer than this. docker-publish.yml has
# no size check of its own; it only notices via a non-200 from the API, i.e.
# after paying the whole build. Catching it here makes it a 2-second failure.
HUB_MAX_CHARS=25000
WARN_ONLY=0
FAILURES=0
SKIPS=0
# Tolerance for the published size claims (check 8), as a percentage OF THE
# MEASURED SIZE. The denominator matters: against the claim instead, the same
# drift reads as a different number, and an early draft of this gate took 20%
# from the claim-relative figure and would therefore have MISSED its own
# motivating case. Both bounds are measured, not guessed:
# - the rot that motivated this check: claimed 1.1 GB vs measured 1.37 GB
# = 19.7% off, so the threshold must sit BELOW that or the gate is theatre.
# - the largest legitimate skew, i.e. a claim describing the currently-published
# release while the next tag changes the size: v1.9.1's 1.37 GB against
# v1.9.2's measured 1.23 GB = 11.4% off, so the threshold must sit ABOVE that
# or every size-changing release trips it.
# 15% sits in that 11.4%-19.7% window. Widen it only with a measured reason, and
# re-derive both bounds if you do.
SIZE_TOLERANCE_PCT="${SIZE_TOLERANCE_PCT:-15}"
usage() {
cat <<'EOF'
Usage: check-doc-drift.sh [--warn-only] [-h|--help]
Compares hand-written claims in README.md and DOCKER_HUB.md against the build
files they describe (Dockerfile.base, Dockerfile.variant).
--warn-only Report drift but exit 0 (advisory use, e.g. a local pre-push hook).
Environment:
SKIP_SIZE_CHECK=1 skip check 8 (published size claims vs Docker Hub)
SKIP_REF_CHECK=1 skip check 9 (refs moved since the last release are named)
SIZE_TOLERANCE_PCT check 8 tolerance, default 15 (see comment for its bounds)
Exit: 0 = in sync, 1 = drift, 2 = cannot run.
EOF
}
while [ $# -gt 0 ]; do
case "$1" in
--warn-only) WARN_ONLY=1; shift ;;
-h|--help) usage; exit 0 ;;
*) echo "::error::unknown argument: $1" >&2; usage >&2; exit 2 ;;
esac
done
for f in "$README" "$HUB" "$DF_VARIANT" "$DF_BASE"; do
if [ ! -f "$f" ]; then
echo "::error::$f not found (cwd $PWD). Cannot evaluate doc drift, so this is exit 2, not a pass."
exit 2
fi
done
# Read `ARG NAME=value` from a Dockerfile. Exit 2 when absent: if the ARG this
# gate is built around has been renamed, the gate is measuring nothing and must
# say so rather than silently comparing against an empty string.
read_arg() {
local file="$1" name="$2" value
value="$(sed -n "s/^ARG ${name}=\\(.*\\)\$/\\1/p" "$file" | head -1)"
if [ -z "$value" ]; then
echo "::error::ARG ${name} not found in ${file}. It was probably renamed;" >&2
echo "::error::update check-doc-drift.sh to match, because this gate is now blind." >&2
exit 2
fi
printf '%s' "$value"
}
# One row of README's "Version pins" table: `| pi | `0.85.1` | ... |`
read_pin_row() {
sed -n "s/^| $1 | \`\\([^\`]*\`*\\)\` |.*/\\1/p" "$README" | head -1
}
fail() {
FAILURES=$((FAILURES + 1))
echo "::error::$1"
}
ok() { printf ' OK %s\n' "$1"; }
# A check that could not be EVALUATED, as distinct from one that passed.
# Deliberately neither ok() nor fail(): printing it as OK would launder an
# unmeasured claim into a passing one (the exact habit this file exists to
# break), while failing on a third party's uptime would make every release
# hostage to Docker Hub's API. Loud, counted, and surfaced in the summary.
skip() { SKIPS=$((SKIPS + 1)); printf ' SKIP %s\n' "$1"; }
echo "Checking hand-maintained doc claims against the build files they describe."
echo
# ---------------------------------------------------------------------------
# 1-3. README's version-pin table vs the ARGs it names by name.
# ---------------------------------------------------------------------------
check_pin() {
local label="$1" documented="$2" actual="$3" where="$4"
if [ -z "$documented" ]; then
fail "README.md: no '| $label |' row found in the version-pin table. Either the
table was restructured (update this gate) or the row was dropped (restore it)."
return
fi
if [ "$documented" != "$actual" ]; then
fail "README.md version-pin table is stale for $label: says '$documented',
$where says '$actual'. Fix the table — it is the reviewable record of what
this repo deliberately freezes, so a wrong row defeats its only purpose."
return
fi
ok "README pin $label = $actual"
}
PI_ACTUAL="$(read_arg "$DF_VARIANT" PI_VERSION)"
ATELIER_ACTUAL="$(read_arg "$DF_VARIANT" PI_ATELIER_REF)"
MEMPALACE_ACTUAL="$(read_arg "$DF_BASE" MEMPALACE_VERSION)"
check_pin pi "$(read_pin_row pi)" "$PI_ACTUAL" "ARG PI_VERSION in $DF_VARIANT"
check_pin pi-atelier "$(read_pin_row pi-atelier)" "$ATELIER_ACTUAL" "ARG PI_ATELIER_REF in $DF_VARIANT"
check_pin mempalace "$(read_pin_row mempalace)" "$MEMPALACE_ACTUAL" "ARG MEMPALACE_VERSION in $DF_BASE"
# ---------------------------------------------------------------------------
# 4. DOCKER_HUB.md's Node claim vs ARG NODE_VERSION. This is the published page,
# so it is the one whose staleness reaches users.
# ---------------------------------------------------------------------------
NODE_ACTUAL="$(read_arg "$DF_BASE" NODE_VERSION)"
NODE_DOCUMENTED="$(sed -n 's/.*\*\*Node\.js\*\* v\([0-9][0-9]*\).*/\1/p' "$HUB" | head -1)"
if [ -z "$NODE_DOCUMENTED" ]; then
fail "$HUB: could not find a '**Node.js** vNN' claim. If the wording changed,
update this gate; do not leave the published page unverified."
elif [ "$NODE_DOCUMENTED" != "$NODE_ACTUAL" ]; then
fail "$HUB claims Node v$NODE_DOCUMENTED but ARG NODE_VERSION=$NODE_ACTUAL.
This file is PUBLISHED to Docker Hub by update-description on every tag,
and it is read from the TAG — so fix it before tagging, not after."
else
ok "$HUB Node claim = v$NODE_ACTUAL"
fi
# ---------------------------------------------------------------------------
# 5. Placeholders CI will not substitute. docker-publish.yml substitutes exactly
# {{PI_VERSION}} and then greps for leftovers of that ONE token, so any other
# {{...}} sails through the guard and is published literally.
# ---------------------------------------------------------------------------
UNKNOWN_PLACEHOLDERS="$(grep -o '{{[A-Za-z0-9_]*}}' "$HUB" | sort -u | grep -v '^{{PI_VERSION}}$' || true)"
if [ -n "$UNKNOWN_PLACEHOLDERS" ]; then
fail "$HUB contains placeholders CI does not substitute, which would be
published verbatim: $(echo "$UNKNOWN_PLACEHOLDERS" | tr '\n' ' ')
docker-publish.yml only fills {{PI_VERSION}}; add substitution there first."
else
ok "$HUB has no placeholders beyond {{PI_VERSION}}"
fi
# Match only the UPPER_SNAKE placeholder convention CI uses. A bare '{{' search
# is WRONG here, and the first version of this check proved it by failing on
# README.md:900 — `docker inspect --format '{{json .Config.Labels}}'`, a Go
# template in a legitimate example, not a placeholder. The gate was wrong, not
# the doc. Keep this anchored to [A-Z] so Go/Jinja/Handlebars examples pass.
README_PLACEHOLDERS="$(grep -o '{{[A-Z][A-Z0-9_]*}}' "$README" | sort -u || true)"
if [ -n "$README_PLACEHOLDERS" ]; then
fail "$README contains placeholder(s) nothing substitutes, so they would render
literally for every reader: $(echo "$README_PLACEHOLDERS" | tr '\n' ' ')
Only DOCKER_HUB.md gets substitution, and only for {{PI_VERSION}}."
else
ok "$README has no unsubstituted placeholders"
fi
# ---------------------------------------------------------------------------
# 6. Hub full_description length.
# ---------------------------------------------------------------------------
HUB_CHARS="$(wc -c < "$HUB" | tr -d ' ')"
if [ "$HUB_CHARS" -gt "$HUB_MAX_CHARS" ]; then
fail "$HUB is $HUB_CHARS chars, over Docker Hub's $HUB_MAX_CHARS-char
full_description limit. update-description would fail with a non-200 AFTER
the full build. Trim it — this file is the essentials-only page, and
README.md is the long form on purpose."
else
ok "$HUB is $HUB_CHARS chars (limit $HUB_MAX_CHARS)"
fi
# ---------------------------------------------------------------------------
# 7. Stale "Unreleased" pointers. "Unreleased" is a CHANGELOG-only concept; in
# a user-facing doc it is always a pointer that outlived what it pointed at.
# This class has now bitten five times, hence a gate rather than vigilance.
# ---------------------------------------------------------------------------
STALE_MARKERS="$(grep -n 'Unreleased' "$README" "$HUB" || true)"
if [ -n "$STALE_MARKERS" ]; then
fail "'Unreleased' appears in a user-facing doc, which is always a stale
pointer once the thing ships (it has happened five times here):
${STALE_MARKERS//$'\n'/$'\n' }
State the fact directly, or move it to CHANGELOG.md where 'Unreleased' means something."
else
ok "no stale 'Unreleased' pointers in $README or $HUB"
fi
# ---------------------------------------------------------------------------
# 8. Published size claims vs Docker Hub's MEASURED full_size.
#
# Why this exists: every other claim in these docs is checked against a file
# in this repo, so it cannot rot without someone editing the thing it
# describes. The size claims had no such anchor -- nothing in the repo states
# the image size -- so they quietly went 24% wrong across eight releases
# (DOCKER_HUB.md said ~1.1 GB; :latest measured 1.37 GB on 2026-09-14).
# DOCKER_HUB.md is POSTed to Docker Hub by update-description, so that number
# is the first thing a stranger reads about this image.
#
# Hub's `full_size` tracks the FIRST manifest entry (amd64 here), NOT the sum
# across architectures -- measured: v1.9.2 full_size=1.228 GB, amd64=1.228,
# arm64=1.211, sum=2.439. That matches the table's per-arch "Size
# (compressed)" column, which is why full_size is the right field.
#
# NOT COVERED, deliberately: README.md's ~3.2 GB figures are UNCOMPRESSED
# on-disk sizes, and the registry API exposes compressed sizes only (layer
# sizes in a manifest are compressed; the config blob carries no uncompressed
# totals). Measuring them needs a real pull, so they are out of scope here --
# do not read a green check 8 as covering them.
# ---------------------------------------------------------------------------
# Shared by checks 8 and 9: which Hub repo, and its tag list (one request).
# Derive the repo from the doc's own rows rather than hardcoding it, so a
# rename cannot leave these checks silently probing a repo nobody publishes to.
# shellcheck disable=SC2016 # single quotes are deliberate: this is a sed
# script, and its \( \) groups and \1 backreference must reach sed unexpanded.
HUB_REPO_PATH="$(sed -n 's/^| `\([^:`]*\):[^`]*`.*/\1/p' "$HUB" | head -1)"
HUB_TAGS_JSON=""
HAVE_NET_TOOLS=0
if command -v curl >/dev/null 2>&1 && command -v python3 >/dev/null 2>&1; then
HAVE_NET_TOOLS=1
if [ -n "$HUB_REPO_PATH" ] && \
{ [ "${SKIP_SIZE_CHECK:-0}" != "1" ] || [ "${SKIP_REF_CHECK:-0}" != "1" ]; }; then
HUB_TAGS_JSON="$(curl -sS -m 20 \
"https://hub.docker.com/v2/repositories/${HUB_REPO_PATH}/tags/?page_size=100" \
2>/dev/null || true)"
fi
fi
if [ "${SKIP_SIZE_CHECK:-0}" = "1" ]; then
skip "size claims -- SKIP_SIZE_CHECK=1 was set"
elif [ "$HAVE_NET_TOOLS" = 0 ]; then
skip "size claims -- need both curl and python3 to measure them"
else
if [ -z "$HUB_REPO_PATH" ]; then
skip "size claims -- found no \`repo:tag\` image rows in $HUB to check"
else
if [ -z "$HUB_TAGS_JSON" ]; then
skip "size claims -- Docker Hub API unreachable (offline?); NOT verified"
else
SIZE_RC=0
# NO `|| true` on the python invocation: an early draft had one, and it
# swallowed the exit code so a printed DRIFT line still exited 0 -- a gate
# that reports the defect and passes anyway. The outer `|| SIZE_RC=$?` is
# what keeps `set -e` happy while preserving the code.
SIZE_OUT="$(HUB_MD="$HUB" HUB_JSON="$HUB_TAGS_JSON" TOL="$SIZE_TOLERANCE_PCT" \
python3 <<'PYEOF'
import json, os, re, sys
try:
data = json.loads(os.environ["HUB_JSON"])
except (ValueError, KeyError) as exc:
print(" SKIP size claims -- Hub API returned unparseable JSON (%s)" % exc)
sys.exit(3)
# full_size == first manifest entry (amd64), which is the per-arch number the
# table's "Size (compressed)" column claims. Verified against .images[] sizes.
sizes = {
r["name"]: r["full_size"] / 1e9
for r in data.get("results", [])
if isinstance(r.get("full_size"), int) and r.get("name")
}
if not sizes:
print(" SKIP size claims -- Hub API returned no usable tags")
sys.exit(3)
tol = float(os.environ["TOL"])
row = re.compile(r"^\|\s*`([^`:]+):([^`]+)`\s*\|[^|]*\|\s*~?([0-9]+(?:\.[0-9]+)?)\s*GB\s*\|")
checked = drift = 0
with open(os.environ["HUB_MD"], encoding="utf-8") as fh:
for line in fh:
m = row.match(line)
if not m:
continue # rows saying "same", and every non-image row
_repo, tag, claimed = m.group(1), m.group(2), float(m.group(3))
if "X.Y.Z" in tag:
continue # placeholder row; the concrete tag is checked instead
# base-<hash> is content-addressed and immutable, so its size is
# base-latest's by construction -- probe the alias that always exists.
probe = "base-latest" if tag.startswith("base-") else tag
actual = sizes.get(probe)
if actual is None:
print(" SKIP size %s -- tag '%s' not present on Hub" % (tag, probe))
continue
checked += 1
off = abs(claimed - actual) / actual * 100
if off <= tol:
print(" OK size %s claims ~%.2f GB, Hub measures %.2f GB (%.0f%% off)"
% (tag, claimed, actual, off))
else:
drift += 1
print(" DRIFT size %s claims ~%.2f GB but Hub measures %.2f GB"
" (%.0f%% off, tolerance %.0f%%)" % (tag, claimed, actual, off, tol))
if checked == 0:
print(" SKIP size claims -- no checkable rows resolved to a published tag")
sys.exit(3)
sys.exit(1 if drift else 0)
PYEOF
)" || SIZE_RC=$?
printf '%s\n' "$SIZE_OUT"
case "$SIZE_RC" in
0) : ;;
3) SKIPS=$((SKIPS + 1)) ;;
*)
fail "a published size claim in $HUB has drifted from what Docker Hub
actually serves (see DRIFT above). This page is POSTed to Docker Hub by
update-description, so it is the first size a stranger sees. Re-measure and
update the table:
curl -sS 'https://hub.docker.com/v2/repositories/${HUB_REPO_PATH}/tags/?page_size=100' |
jq -r '.results[] | \"\\(.name) \\(.full_size/1e9)\"'"
;;
esac
fi
fi
fi
# ---------------------------------------------------------------------------
# 9. Everything the NEXT build would bake differently from the LAST PUBLISHED
# release must be named in the CHANGELOG text above that release's heading.
#
# Why this exists, measured 2026-09-19: pi-extensions 25c1265 (a new `task`
# tool and a hook that blocks certain `fork` calls -- a change to how every
# agent in the container delegates work) and mempalace-toolkit 817b3a8 (the
# feed's mine deadline had never reached the transport) both reached this
# image through floating `*_REF=main` ARGs. Neither produced a diff in this
# repo, so nothing here asked for a CHANGELOG entry, and neither had one
# until a reader asked. This is the same shape as check 8: a fact with no
# in-repo anchor rots. The hand practice that existed for it -- the
# "Dependency audit" table in each release's notes ("Baked in vN | Upstream
# now") -- is precisely a "someone remembers" mechanism, and it had lapsed.
#
# How it measures, with no docker/crane/token: the last published `vX.Y.Z`
# is the highest such tag in Hub's tag list (shared with check 8); its
# amd64 config blob is read through the anonymous registry API (token ->
# manifest index -> per-arch manifest -> config) and carries one
# `se.jordbo.pi-devbox.<name>-ref` label per component, each holding the
# SHA that build-args actually baked (resolve-versions in docker-publish.yml
# turns every ref into a SHA before `docker build`). "What the next build
# would bake" is resolved the way that job does it: a 40-hex ARG is itself,
# a tag or branch is `git ls-remote`d (peeled `^{}` first -- an annotated
# tag's un-dereferenced SHA is the tag object, a false alarm this repo has
# already fallen for once), pi-studio is the highest semver tag, and
# `PI_VERSION` / `MEMPALACE_VERSION` are compared as literals against the
# `pi-version` / `mempalace-version` labels (the latter set in Dockerfile.base
# and inherited; absent on releases before it shipped, which reports SKIP).
#
# The rule: baked == would-bake is OK with no mention required. If they
# differ, the text ABOVE the last published version's `## ` heading -- i.e.
# `## Unreleased` plus any not-yet-published `## vX.Y.Z` section, which is
# what the release commit turns Unreleased into -- must contain the
# would-bake value's 7-char SHA prefix (or, for pi-studio, the tag name; for
# pi, the version string). Naming the SHA, not just the repo, is the point:
# it is what the audit table always recorded, and it makes the failure
# message's compare URL a copy-paste away from knowing what moved.
#
# Every upstream commit therefore re-reds this gate until the CHANGELOG
# names the new head. That is the intended cost: the thing that gets baked
# is the thing that gets named, and a typo-fix upstream costs one edited
# SHA here. Read from the TAG like everything else in these docs -- the
# release commit renames Unreleased, so the pending text still covers it.
#
# SKIPs, each counted: SKIP_REF_CHECK=1; no curl/python3; Hub unreachable;
# the release's labels unreadable; one component's upstream unreachable
# (that component only). A published tag whose heading is MISSING from the
# CHANGELOG is a failure, not a skip: that is drift in its own right.
# ---------------------------------------------------------------------------
if [ "${SKIP_REF_CHECK:-0}" = "1" ]; then
skip "ref moves -- SKIP_REF_CHECK=1 was set"
elif [ "$HAVE_NET_TOOLS" = 0 ]; then
skip "ref moves -- need both curl and python3 to read the published labels"
elif ! command -v git >/dev/null 2>&1; then
skip "ref moves -- need git (ls-remote) to resolve what the next build would bake"
elif [ -z "$HUB_REPO_PATH" ]; then
skip "ref moves -- found no \`repo:tag\` image rows in $HUB to locate the published image"
elif [ -z "$HUB_TAGS_JSON" ]; then
skip "ref moves -- Docker Hub API unreachable (offline?); NOT verified"
else
# One plain top-level assignment per ARG, on purpose: read_arg exits 2 on a
# missing ARG, and under `set -e` that only propagates from a bare
# `VAR="$(...)"`. Nested inside a heredoc's $(...) the exit would be swallowed
# by `cat`, and a renamed ARG would leave this check comparing a label against
# an empty string and reporting the component "unchanged".
TOOLKIT_REPO="$(read_arg "$DF_VARIANT" PI_TOOLKIT_REPO)"; TOOLKIT_REF="$(read_arg "$DF_VARIANT" PI_TOOLKIT_REF)"
EXTENSIONS_REPO="$(read_arg "$DF_VARIANT" PI_EXTENSIONS_REPO)"; EXTENSIONS_REF="$(read_arg "$DF_VARIANT" PI_EXTENSIONS_REF)"
FORK_REPO="$(read_arg "$DF_VARIANT" PI_FORK_REPO)"; FORK_REF="$(read_arg "$DF_VARIANT" PI_FORK_REF)"
OBSMEM_REPO="$(read_arg "$DF_VARIANT" PI_OBSMEM_REPO)"; OBSMEM_REF="$(read_arg "$DF_VARIANT" PI_OBSMEM_REF)"
ATELIER_REPO="$(read_arg "$DF_VARIANT" PI_ATELIER_REPO)"
MPTK_REPO="$(read_arg "$DF_BASE" MEMPALACE_TOOLKIT_REPO)"; MPTK_REF="$(read_arg "$DF_BASE" MEMPALACE_TOOLKIT_REF)"
STUDIO_REPO="$(read_arg "$DF_VARIANT" PI_STUDIO_REPO)"
SKILLSET_SNAPSHOT="$(read_arg "$DF_VARIANT" SKILLSET_SNAPSHOT_REF)"
# name|kind|repo|ref -- one line per label the variant image carries.
# kinds: ref = branch/tag/SHA resolved like resolve-versions does;
# studio = highest semver tag of the repo (label lives on <tag>-studio);
# literal = the ARG value IS the baked value (a SHA pin, a version).
REF_COMPONENTS="pi-toolkit|ref|$TOOLKIT_REPO|$TOOLKIT_REF
pi-extensions|ref|$EXTENSIONS_REPO|$EXTENSIONS_REF
pi-fork|ref|$FORK_REPO|$FORK_REF
pi-obsmem|ref|$OBSMEM_REPO|$OBSMEM_REF
pi-atelier|ref|$ATELIER_REPO|$ATELIER_ACTUAL
mempalace-toolkit|ref|$MPTK_REPO|$MPTK_REF
pi-studio|studio|$STUDIO_REPO|
skillset-snapshot|literal||$SKILLSET_SNAPSHOT
pi-version|literal||$PI_ACTUAL
mempalace-version|literal||$MEMPALACE_ACTUAL"
REF_RC=0
# Same discipline as check 8: no `|| true` on the python, or a printed DRIFT
# exits 0. Per-component SKIP lines are counted afterwards by grep, so a run
# that evaluated eight components and could not reach the ninth reports one
# skip, not a green tick over the ninth.
REF_OUT="$(HUB_REPO="$HUB_REPO_PATH" HUB_JSON="$HUB_TAGS_JSON" CHANGELOG="CHANGELOG.md" \
COMPONENTS="$REF_COMPONENTS" python3 <<'PYEOF'
import json, os, re, subprocess, sys, urllib.request, urllib.parse
SHA40 = re.compile(r"^[0-9a-f]{40}$")
SEMVER = re.compile(r"^v?[0-9]+\.[0-9]+\.[0-9]+$")
LABEL = "se.jordbo.pi-devbox."
def ver_key(tag):
return tuple(int(x) for x in tag.lstrip("v").split("."))
def http_json(url, headers=None, timeout=30):
req = urllib.request.Request(url, headers=headers or {})
with urllib.request.urlopen(req, timeout=timeout) as resp:
return json.loads(resp.read().decode("utf-8"))
def labels_of(repo, tag):
"""Config labels of <repo>:<tag>'s amd64 image via the anonymous registry API."""
tok = http_json(
"https://auth.docker.io/token?service=registry.docker.io&scope="
+ urllib.parse.quote(f"repository:{repo}:pull", safe=":")
)["token"]
hdr = {
"Authorization": f"Bearer {tok}",
"Accept": ", ".join([
"application/vnd.oci.image.index.v1+json",
"application/vnd.docker.distribution.manifest.list.v2+json",
"application/vnd.oci.image.manifest.v1+json",
"application/vnd.docker.distribution.manifest.v2+json",
]),
}
base = f"https://registry-1.docker.io/v2/{repo}"
man = http_json(f"{base}/manifests/{tag}", hdr)
if "manifests" in man: # multi-arch index: pick linux/amd64, as check 8 does
cands = [m for m in man["manifests"]
if m.get("platform", {}).get("architecture") == "amd64"
and m.get("platform", {}).get("os") == "linux"]
if not cands:
raise RuntimeError("no linux/amd64 entry in the manifest index")
man = http_json(f"{base}/manifests/{cands[0]['digest']}", hdr)
cfg = http_json(f"{base}/blobs/{man['config']['digest']}", hdr)
return cfg.get("config", {}).get("Labels") or {}
def ls_remote(repo, *patterns):
# GIT_TERMINAL_PROMPT=0: a repo flipped private must fail fast as a SKIP,
# not sit waiting for a username on a CI runner until the job times out.
env = dict(os.environ, GIT_TERMINAL_PROMPT="0")
out = subprocess.run(["git", "ls-remote", repo, *patterns], env=env,
capture_output=True, text=True, timeout=60, check=True).stdout
return {line.split("\t")[1]: line.split("\t")[0] for line in out.splitlines() if "\t" in line}
def resolve_ref(repo, ref):
"""What docker-publish.yml's resolve-versions would pass as the build-arg."""
if SHA40.match(ref):
return ref, ref
refs = ls_remote(repo, f"refs/heads/{ref}", f"refs/tags/{ref}", f"refs/tags/{ref}^{{}}")
for key in (f"refs/tags/{ref}^{{}}", f"refs/heads/{ref}", f"refs/tags/{ref}"):
if key in refs:
return refs[key], ref
raise RuntimeError(f"'{ref}' is neither a branch nor a tag of {repo}")
def resolve_studio(repo):
refs = ls_remote(repo, "refs/tags/*")
tags = {k[len("refs/tags/"):]: v for k, v in refs.items()}
names = sorted((t for t in tags if SEMVER.match(t)), key=ver_key)
if not names:
raise RuntimeError(f"no semver tag at {repo}")
tag = names[-1]
return tags.get(tag + "^{}", tags[tag]), tag
def compare_url(repo, a, b):
root = repo[:-4] if repo.endswith(".git") else repo
return f"{root}/compare/{a}...{b}"
try:
hub = json.loads(os.environ["HUB_JSON"])
except (ValueError, KeyError) as exc:
print(" SKIP ref moves -- Hub API returned unparseable JSON (%s)" % exc)
sys.exit(3)
released = sorted((r["name"] for r in hub.get("results", [])
if isinstance(r.get("name"), str) and re.fullmatch(r"v[0-9]+\.[0-9]+\.[0-9]+", r["name"])),
key=ver_key)
if not released:
print(" SKIP ref moves -- Hub lists no published vX.Y.Z tag to compare against")
sys.exit(3)
last = released[-1]
repo = os.environ["HUB_REPO"]
# The text every not-yet-published change lives in: everything above the last
# published version's heading. Its absence is drift, not a skip.
text = open(os.environ["CHANGELOG"], encoding="utf-8").read()
# (\s|$) rather than \b: a word boundary would accept "## v1.9.2-rc1" or
# "## v1.9.2-typo" as v1.9.2's heading. Caught by the sabotage test, not review.
m = re.search(r"^## v?%s(\s|$)" % re.escape(last.lstrip("v")), text, re.M)
if not m:
print(" DRIFT ref moves -- %s is the last PUBLISHED tag on Hub but %s has no '## %s' heading"
% (last, os.environ["CHANGELOG"], last))
sys.exit(1)
pending = text[:m.start()].lower()
try:
labels = labels_of(repo, last)
except Exception as exc: # network, auth, shape -- all "could not measure"
print(" SKIP ref moves -- could not read %s:%s's labels from the registry (%s); NOT verified"
% (repo, last, exc))
sys.exit(3)
studio_labels = None
checked = drift = 0
problems = []
for line in os.environ["COMPONENTS"].splitlines():
if not line.strip():
continue
name, kind, url, ref = line.split("|", 3)
# <name>-ref labels hold SHAs; names that already end in -version are the
# label (pi-version, mempalace-version) -- a version string, compared literally.
key = LABEL + name if name.endswith("-version") else LABEL + name + "-ref"
try:
if kind == "studio":
if studio_labels is None:
studio_labels = labels_of(repo, last + "-studio")
baked = studio_labels.get(key)
else:
baked = labels.get(key)
except Exception as exc:
print(" SKIP %-18s -- could not read %s:%s-studio's labels (%s)" % (name, repo, last, exc))
continue
if not baked:
print(" SKIP %-18s -- %s carries no %s label" % (name, last, key))
continue
try:
if kind == "ref":
now, shown = resolve_ref(url, ref)
elif kind == "studio":
now, shown = resolve_studio(url)
else:
now, shown = ref, ref
except Exception as exc:
print(" SKIP %-18s -- could not resolve what the next build would bake (%s)" % (name, exc))
continue
checked += 1
is_sha = bool(SHA40.match(now))
short = (lambda s: s[:7] if SHA40.match(s) else s)
if baked == now:
print(" OK %-18s unchanged since %s (%s)" % (name, last, short(now)))
continue
names = [now[:7].lower()] if is_sha else [now.lower()]
if kind == "studio":
names.append(shown.lower())
if any(n in pending for n in names):
print(" OK %-18s %s -> %s since %s, named above the %s heading"
% (name, short(baked), short(now), last, last))
continue
drift += 1
hint = compare_url(url, baked, now) if (url and is_sha and SHA40.match(baked)) else ""
problems.append(" %-18s %s -> %s%s" % (name, short(baked), short(now), (" " + hint) if hint else ""))
print(" DRIFT %-18s %s -> %s since %s, NOT named above the %s heading"
% (name, short(baked), short(now), last, last))
if problems:
print(" Name each new value (7-char SHA prefix, or the tag/version) in CHANGELOG.md above '## %s':" % last)
print("\n".join(problems))
if checked == 0 and drift == 0:
print(" SKIP ref moves -- no component could be evaluated")
sys.exit(3)
sys.exit(1 if drift else 0)
PYEOF
)" || REF_RC=$?
printf '%s\n' "$REF_OUT"
REF_SKIPS="$(printf '%s\n' "$REF_OUT" | grep -c '^ SKIP ' || true)"
case "$REF_RC" in
0) SKIPS=$((SKIPS + REF_SKIPS)) ;;
3) SKIPS=$((SKIPS + 1)) ;;
*)
SKIPS=$((SKIPS + REF_SKIPS))
fail "a component the next build would bake differently from the last published
release is not named in CHANGELOG.md (see DRIFT above). These reach the image
through floating refs, so nothing else in this repo records that they moved;
the CHANGELOG entry is the only place a reader of the next tag can learn it.
Name the new SHA (7 chars is enough) where you describe the change -- the
compare URL above shows what moved."
;;
esac
fi
echo
if [ "$FAILURES" -eq 0 ]; then
if [ "$SKIPS" -gt 0 ]; then
echo "OK: every checked doc claim matches the build files" \
"($SKIPS check(s) SKIPPED and therefore NOT verified -- see SKIP above)."
else
echo "OK: every checked doc claim matches the build files."
fi
exit 0
fi
echo "::error::$FAILURES doc claim(s) have drifted from the build files."
echo
echo "Docs are read from the TAG, not from main: docker-publish.yml checks out"
echo "github.ref, so a fix pushed after tagging does not reach the release or the"
echo "Hub page. Update the docs BEFORE you tag."
if [ "$WARN_ONLY" -eq 1 ]; then
echo "(--warn-only: exiting 0 anyway)"
exit 0
fi
exit 1
+7 -1
View File
@@ -32,10 +32,16 @@
#
# SEVERITY CHOICE
# -S error is 0 findings across this repo when clean, so it is free to add.
# -S warning is NOT free here (19x SC2088 tilde-in-quotes in
# -S warning is NOT free here (20x SC2088 tilde-in-quotes in
# recreate-sanity-check.sh, plus assorted SC2016 — both intentional), and a
# noisy gate trains people to ignore it. Error-only, matching the
# SHELLCHECK_OPTS philosophy in lint.yml.
# Reproduce the count before editing it (the `$ ` prefix is load-bearing: a
# comment whose first word is "shellcheck" is parsed as a DIRECTIVE, and a
# malformed one is SC1072/SC1073 at severity error — this gate caught exactly
# that when the line was first written without it):
# $ shellcheck -S warning -f gcc scripts/*.sh rootfs/usr/local/bin/* \
# entrypoint*.sh hooks/* | grep -c SC2088
#
# Usage: bash scripts/lint-shell.sh [root] (default root: repo top level)
set -uo pipefail
+105 -2
View File
@@ -11,7 +11,8 @@
# pi-observational-memory / (studio variant) pi-studio package
# registrations in settings.json packages[]
# - Shell defaults re-seeded from /etc/skel-devbox
# - /tmp/sshcm exists with mode 700 (ssh ControlMaster dir)
# - ssh ControlMaster works: /tmp/sshcm exists 700 AND the ControlPath that
# ssh actually resolves (ssh -G) is a writable directory
# - /opt toolkits intact
# - Known expected-absences don't regress
#
@@ -425,13 +426,115 @@ if [ -d /opt/pi-atelier ] && command -v jq >/dev/null 2>&1; then
fi
echo
echo "-- ssh ControlMaster dir --"
echo "-- ssh ControlMaster: socket dir + EFFECTIVE ControlPath --"
# TWO LAYERS, and the second is the one that has actually broken in the field.
#
# LAYER 1 (original check): /tmp/sshcm, the directory entrypoint-user.sh creates
# for the base image's system drop-in
# (/etc/ssh/ssh_config.d/00-devbox-controlmaster.conf).
#
# LAYER 2 (added 2026-09-15): the directory a config NAMES — which is not the
# same question, and asserting layer 1 is structurally blind to it. On
# emb-7kj4vr4g a durable ~/.pi/ssh/config pointed ControlPath at /tmp/ssh-cm
# (with a hyphen), a directory nothing in the image creates. EVERY ssh died
# unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory
# rc=255 with the remote command never running — while this script printed a
# green tick for layer 1, truthfully, about the wrong object.
#
# The same rc=255 has a second, independent cause already documented in prose in
# Dockerfile.base ("SSH client defaults" CAVEAT) and never verified anywhere: a
# per-host `ControlPath ~/.ssh/cm/%r@%h:%p` inherited from a bind-mounted
# READ-ONLY ~/.ssh. Measured to be the identical failure class:
# unix_listener: cannot bind to path ~/.ssh/cm/...: Read-only file system
# So do not guess which config wins — ask ssh. `ssh -G` applies real config
# precedence (first-obtained-value-wins, system drop-in, Include, -F override)
# and prints the fully expanded ControlPath. Require its parent to exist and be
# writable. Cost measured at 0.116 s for 48 hosts; -G never opens a connection.
if [ -d /tmp/sshcm ] && [ "$(stat -c %a /tmp/sshcm 2>/dev/null)" = "700" ]; then
pass "/tmp/sshcm exists with mode 700"
else
fail "/tmp/sshcm missing or not mode 700"
fi
# Probe one route. $1 = label, $2 = config to force with -F ("" = ssh's own
# default precedence), $3 = severity when a ControlPath dir is unusable.
#
# SEVERITY SPLIT IS DELIBERATE. The default route legitimately resolves into the
# read-only ~/.ssh on any host whose own config pins ControlPath there, and the
# supported workaround (`ssh -F ~/.ssh-local/config`) already exists — so that
# is a warn, not a fail. Failing it would paint this script red on every run of
# every device, and a check that fires benignly every time is one you learn to
# ignore. The sidecar route is the PRESCRIBED one, so there it is a hard fail.
_ssh_cm_probe() {
local label="$1" cfg="${2:-}" sev="${3:-fail}"
local h out cm cp dir n=0 shown
local bad=()
while IFS= read -r h; do
[ -n "$h" ] || continue
if [ -n "$cfg" ]; then
out=$(ssh -F "$cfg" -G "$h" 2>/dev/null) || continue
else
out=$(ssh -G "$h" 2>/dev/null) || continue
fi
cm=$(printf '%s\n' "$out" | awk '/^controlmaster /{print $2; exit}')
case "$cm" in '' | no | none | false) continue ;; esac
cp=$(printf '%s\n' "$out" | awk '/^controlpath /{print $2; exit}')
case "$cp" in '' | none) continue ;; esac
n=$((n + 1))
dir=$(dirname "$cp")
if [ ! -d "$dir" ] || [ ! -w "$dir" ]; then
bad+=("$h")
fi
done <<< "$SSH_CM_HOSTS"
# Cap the host list. A 50-name line is the noisy gate this repo already warns
# about in lint-shell.sh: unreadable output is ignored output. Six names plus
# a count is enough to identify the class and act.
if [ "${#bad[@]}" -gt 0 ]; then
shown="${bad[*]:0:6}"
if [ "${#bad[@]}" -gt 6 ]; then
shown="$shown (+$(( ${#bad[@]} - 6 )) more)"
fi
fi
if [ "$n" -eq 0 ]; then
warn "$label: no host resolves to ControlMaster on — effective ControlPath not exercised"
elif [ "${#bad[@]}" -eq 0 ]; then
pass "$label: $n ControlMaster host(s), every ControlPath dir exists and is writable"
elif [ "$sev" = "warn" ]; then
warn "$label: ${#bad[@]}/$n host(s) resolve ControlPath to a missing or unwritable dir [$shown] — expected when ~/.ssh/config pins ControlPath inside the read-only ~/.ssh; use 'ssh -F ~/.ssh-local/config' (see Dockerfile.base CAVEAT)"
else
fail "$label: ${#bad[@]}/$n host(s) resolve ControlPath to a missing or unwritable dir [$shown] — ssh dies rc=255 'unix_listener: cannot bind to path' and the remote command never runs"
fi
}
if command -v ssh >/dev/null 2>&1; then
# Concrete Host aliases only: patterns (*, ?) and negations (!) are not
# connectable targets, so `ssh -G` on them proves nothing.
_cm_cfgs=()
if [ -r "$HOME/.ssh/config" ]; then _cm_cfgs+=("$HOME/.ssh/config"); fi
if [ -r "$HOME/.ssh-local/config" ]; then _cm_cfgs+=("$HOME/.ssh-local/config"); fi
if [ "${#_cm_cfgs[@]}" -gt 0 ]; then
SSH_CM_HOSTS=$(awk 'tolower($1)=="host"{for(i=2;i<=NF;i++) if ($i !~ /[*?!]/) print $i}' \
"${_cm_cfgs[@]}" 2>/dev/null | sort -u)
else
SSH_CM_HOSTS=""
fi
if [ -z "$SSH_CM_HOSTS" ]; then
warn "no concrete Host aliases in ~/.ssh/config or ~/.ssh-local/config — effective ControlPath not verified"
else
_ssh_cm_probe "default ssh precedence" "" warn
if [ -r "$HOME/.ssh-local/config" ]; then
_ssh_cm_probe "ssh -F ~/.ssh-local/config" "$HOME/.ssh-local/config" fail
else
warn "~/.ssh-local/config absent — setup-lan-access.sh did not run; the prescribed multiplex route is unverified"
fi
fi
else
warn "ssh not on PATH — effective ControlPath not verified"
fi
echo
echo "-- Shell defaults re-seeded from /etc/skel-devbox --"
if [ -f "$HOME/.bash_aliases" ]; then
+134 -4
View File
@@ -28,6 +28,9 @@
# (human, --json, --quiet)
# - (studio variant only, auto-detected) pi-studio cloned + prebuilt
# client bundle present + registered via `pi install`
# - no foreign npm-11 platform packages (@esbuild, clipboard) beyond the host
# - no build-time npm cache (/root/.npm) shipped in the image
# - esbuild compiles + clipboard native loads at every install site
# - image size within threshold
set -euo pipefail
@@ -315,8 +318,24 @@ run "pi-toolkit clone" "test -d /opt/pi-toolkit && git -C /opt/pi-toolkit rev
run "pi-extensions clone" "test -d /opt/pi-extensions && git -C /opt/pi-extensions rev-parse --short HEAD"
run "pi-fork clone + node_modules" \
"test -f /opt/pi-fork/package.json && test -d /opt/pi-fork/node_modules"
run "pi-observational-memory clone + node_modules" \
"test -f /opt/pi-observational-memory/package.json && test -d /opt/pi-observational-memory/node_modules"
# om is checked differently from pi-fork ON PURPOSE. It declares ZERO runtime
# dependencies: 8 devDependencies (omitted by --omit=dev) and 4 peerDependencies,
# which pi itself provides. npm 10 still materialised a node_modules for it, but
# that directory held exactly ONE file (.package-lock.json, 4 KB) and no nested
# package.json at all — 20 empty scope dirs. npm 11 stopped creating it, so the
# old `test -d node_modules` assertion went red on v1.9.0 while nothing about om
# had changed or broken. It was asserting an npm artefact, not a property of the
# shipped software. What actually has to hold is that the entry point pi loads
# exists, so assert THAT, straight out of the manifest pi reads
# (package.json -> pi.extensions), rather than a hardcoded path that could drift.
run "pi-observational-memory clone + declared pi entry point" \
"test -f /opt/pi-observational-memory/package.json && \
node -e 'const p=require(\"/opt/pi-observational-memory/package.json\"),f=require(\"fs\"),h=require(\"path\"); \
const l=(p.pi&&p.pi.extensions)||[]; \
if(!l.length){console.error(\"package.json declares no pi.extensions\");process.exit(1)} \
for(const e of l){const t=h.resolve(\"/opt/pi-observational-memory\",e); \
if(!f.existsSync(t)){console.error(\"declared entry missing: \"+t);process.exit(1)}} \
console.log(\"entries ok: \"+l.join(\",\"))'"
# ...and that the clone carries the AUTH FIX, not merely that it exists. om's
# pre-flight hasUsableAuth() check silently disabled `recall` for ~8 weeks once
# pi moved to request-time SigV4 signing and stopped exposing a static Bedrock
@@ -578,6 +597,19 @@ if [ -n "$LBL" ] && [ "$LBL" != "<no value>" ]; then
else
printf " ❌ OCI label se.jordbo.pi-devbox.pi-extensions-ref missing or empty\n"; FAIL=$((FAIL+1))
fi
# mempalace-version is set in Dockerfile.base and INHERITED by the variant, so
# it states the pin of the base this image actually built on. It must equal the
# installed binary: the one way they diverge is a base built with
# INSTALL_MEMPALACE=false (label says 3.x, nothing installed) or an install that
# resolved to something other than the pin — both invisible to a label-only
# check. Same ground-truth rule as the manifest assertion above.
MP_LBL=$(docker inspect --format '{{ index .Config.Labels "se.jordbo.pi-devbox.mempalace-version" }}' "$IMAGE" 2>/dev/null || true)
MP_BIN=$(docker run --rm --entrypoint= "$IMAGE" sh -c 'mempalace --version 2>/dev/null | head -n1 | tr -d "\r"' 2>/dev/null || true); MP_BIN=${MP_BIN##* }
if [ -n "$MP_LBL" ] && [ "$MP_LBL" != "<no value>" ] && [ "$MP_LBL" = "$MP_BIN" ]; then
printf " ✅ OCI label se.jordbo.pi-devbox.mempalace-version=%s equals the installed core\n" "$MP_LBL"; PASS=$((PASS+1))
else
printf " ❌ OCI label se.jordbo.pi-devbox.mempalace-version=[%s] vs installed mempalace=[%s]\n" "$MP_LBL" "$MP_BIN"; FAIL=$((FAIL+1))
fi
# ── Runtime deployment (needs entrypoint to run) ──────────────────────
echo ""
@@ -678,7 +710,22 @@ exec_test "mempalace skill linked (fallback)" 'test -L $HOME/.agents/skills
# (forgotten bump AND re-vendored stale snapshot). Upstream content behind this
# refresh: the bare project-name wing convention and the <harness>@<device>
# added_by rule.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Diaries self-heal; plain drawers do not" "$f" && ! grep -q "Agent diaries live in" "$f" && echo ok'
#
# Unreleased: RE-PINNED again on refresh e9e09d9 -> e9e45f7. The retired pair was
# still green against the new snapshot (the diaries section was untouched), so it
# was blind to this refresh for the same reason the v1.8.13 pair was blind to
# that one. The replacement pair is unusually strong because BOTH witnesses come
# out of the same upstream commit: skillset e9e45f7 ADDED the "Withdrawing an ask
# you sent" bullet and DELETED the sentence "there is nothing anyone can do about
# it from the other end" that the new bullet contradicts. Directions were
# MEASURED against both files, not read off the diff: "Withdrawing an ask you
# sent" is new=1/old=0, "nothing anyone can do about it from the other end" is
# new=0/old=1. A canary whose negative witness was removed by the very commit it
# pins fails loudly on the OLD bytes instead of merely failing to notice them,
# which is the property every previous pair here lacked. Upstream content:
# requester-side ask withdrawal became DEPLOYED behaviour once v1.9.1 baked
# mempalace-toolkit e68ee20 (>= e2b060a) through the floating MEMPALACE_TOOLKIT_REF.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Withdrawing an ask you sent" "$f" && ! grep -q "nothing anyone can do about it from the other end" "$f" && echo ok'
# Link TARGETS, not just link existence: with no skillset mounted (as here) the
# baked tree must be what resolves, for all four vendored skills.
exec_test "vendored skills resolve to the baked tree (no skillset mounted)" \
@@ -697,9 +744,17 @@ exec_test "pi-devbox-version reports skill sources (all baked, no skillset here)
'out=$(pi-devbox-version)
echo "$out" | grep -q "skills:" || { echo "no skills section" >&2; exit 1; }
for s in mempalace pi-extensions pi-devbox-environment credential-incident-response; do
echo "$out" | grep -qE "^ $s +baked$" \
echo "$out" | grep -qE "^ $s +baked( \([^)]*\))?$" \
|| { echo "$s not reported as baked" >&2; exit 1; }
done; echo ok'
# The optional " (...)" above is what pi-extensions now appends to say WHICH copy
# shipped — "baked (package copy)", or a loud FALLBACK/MIXED annotation. Without
# allowing it, adding that annotation turned this assertion red on v1.9.0 even
# though the state it reported was the correct one. The suffix is deliberately
# matched loosely rather than pinned to "(package copy)", because WHICH copy
# shipped is already asserted authoritatively above, against the manifest field
# and its measured tree hash, and duplicating that here in a regex would just
# create a second place to update whenever the wording changes.
# The boot banner must NOT carry the section: entrypoint-user.sh prints the
# version FIRST, before the baked links exist and long before the skillset
# deploy + reconcile run last, so anything it said about skill sources would be
@@ -855,6 +910,55 @@ exec_test "pi-fork extensions floor is [] (forks cannot write to the palace)" \
exec_test "/tmp/sshcm dir mode 700 (ssh ControlMaster)" \
'test -d /tmp/sshcm && [ "$(stat -c %a /tmp/sshcm)" = "700" ] && echo ok'
# ── Build-time leftovers (npm 11 bloat sentinels) ─────────────────────
# Both of these are worth a PASS/FAIL assertion rather than a size-gate
# diagnostic, because the size gate has ~225 MB of deliberate margin: v1.9.1
# shipped +131 MB of pure build residue and stayed green. These name the
# residue directly, so a regression is legible instead of merely "bigger".
echo ""
echo "── Build-time leftovers ──"
# npm 11 installs EVERY optional platform package of a native dependency, not
# just the host's (it ignores os/cpu, --os/--cpu and npmrc os=/cpu=). Two
# families are affected and pruned in Dockerfile.variant: @esbuild/<platform>
# and @mariozechner/clipboard-<triple>. Keep-set is the host arch only, plus
# clipboard's gnu AND musl (its napi loader picks between them at runtime).
# Runs as root because the image declares no USER; that is also what lets the
# cache assertion below read /root.
run "no foreign platform packages (npm 11 sentinel)" \
'arch=$(node -p process.arch); bad=$(find /usr/lib/node_modules /opt -type d \( -regex ".*/@esbuild/[^/]+" -o -regex ".*/@mariozechner/clipboard-[^/]+" \) ! -name "linux-$arch" ! -name "clipboard-linux-$arch-gnu" ! -name "clipboard-linux-$arch-musl" -prune -print 2>/dev/null); if [ -n "$bad" ]; then echo "foreign platform dirs shipped:" >&2; echo "$bad" >&2; du -sm $bad 2>/dev/null | sort -rn | head -5 >&2; exit 1; fi; echo ok'
# The build's own npm download cache is not free: it lands in the layer that
# created it. v1.9.1 shipped 145 MB of /root/.npm/_cacache (35 MB in v1.8.14)
# — the largest single item in its +131 MB residual, and invisible to the
# size-gate diagnostics because those only looked under node_modules and /opt.
# Nothing at runtime reads it: root's cache, while the container runs as
# `developer`. NOTE the assertion must run as root or a permission error on
# mode-700 /root would make `test ! -d` pass for the wrong reason.
run "no build-time npm cache shipped (/root/.npm)" \
'test "$(id -u)" = "0" || { echo "assertion needs root to read /root" >&2; exit 1; }; if [ -e /root/.npm ]; then echo "/root/.npm shipped: $(du -sm /root/.npm | cut -f1) MB" >&2; exit 1; fi; echo ok'
# The prune's risk is not "too big" but "removed something needed", and only a
# FUNCTIONAL check covers that. These load the natives from every install site
# found in the image, so they also scale to the studio variant's third site.
#
# NOTE THE PATH-QUALIFIED require(). The obvious form, `node -e
# 'require("esbuild")...'`, resolves by walking up from the CURRENT DIRECTORY —
# so it fails with MODULE_NOT_FOUND from /workspace on a perfectly good image,
# because esbuild lives nested inside the pi trees and global installs are not
# on node's require path (NODE_PATH is unset). That exact command was left in a
# runbook as "if this fails, revert the release", and it duly failed for the
# wrong reason on the first machine that ran it. A check must fail only for the
# thing it is checking.
run "esbuild works at every install site (prune removed weight, not function)" \
'sites=$(find /usr/lib/node_modules /opt -type d -path "*/node_modules/esbuild" -prune 2>/dev/null); if [ -z "$sites" ]; then echo "no esbuild install found at all" >&2; exit 1; fi; for d in $sites; do node -e "require(\"$d\").transformSync(\"const x:number=1\",{loader:\"ts\"})" || { echo "esbuild broken at $d" >&2; exit 1; }; done; echo ok'
# Clipboard is the family pruned second, and its napi-rs loader picks its native
# binding at require() time — so a successful load IS the proof that the kept
# platform package is the one this image needs.
run "clipboard native loads at every install site" \
'sites=$(find /usr/lib/node_modules /opt -type d -path "*/node_modules/@mariozechner/clipboard" -prune 2>/dev/null); if [ -z "$sites" ]; then echo "no @mariozechner/clipboard install found at all" >&2; exit 1; fi; for d in $sites; do node -e "var c=require(\"$d\"); if (typeof c.setText !== \"function\") { throw new Error(\"native binding missing\"); }" || { echo "clipboard native broken at $d" >&2; exit 1; }; done; echo ok'
# ── Image size ────────────────────────────────────────────────────────
echo ""
echo "── Image size ──"
@@ -882,6 +986,32 @@ elif [ "$SIZE_MB" -le "$SIZE_THRESHOLD_MB" ]; then
printf " ✅ size: %d MB (threshold %d MB)\n" "$SIZE_MB" "$SIZE_THRESHOLD_MB"; PASS=$((PASS+1))
else
printf " ❌ size: %d MB exceeds threshold %d MB\n" "$SIZE_MB" "$SIZE_THRESHOLD_MB"; FAIL=$((FAIL+1))
# A bare "too big" verdict cost a full CI-log dig plus a local npm bisect to
# attribute the v1.9.0 overshoot (+431 MB, which turned out to be npm 11
# installing 26 @esbuild platform binaries per pi-coding-agent copy). The
# container already knows where its bytes are, so make it say so: the biggest
# layers, and the biggest directories under the paths that historically grow.
# Same principle as the run() helper above — a red assertion should carry its
# own diagnostic rather than send the next reader spelunking.
echo " ── largest layers (docker history) ──"
docker history --format '{{.Size}}\t{{.CreatedBy}}' "$IMAGE" 2>/dev/null \
| grep -vE '^0B' | head -12 | sed 's/^/ /' | cut -c1-160
echo " ── largest directories in the image ──"
docker run --rm --entrypoint sh "$IMAGE" -c \
'du -sm /usr/lib/node_modules/* /opt/* /usr/local/share/ms-playwright 2>/dev/null | sort -rn | head -12' \
2>/dev/null | sed 's/^/ /' || echo " (could not inspect directories)"
echo " ── build caches that should not be in the image ──"
# v1.9.1's residual was 145 MB of npm cache under /root, and the du list
# above cannot see it: it enumerates node_modules and /opt only. A gate whose
# diagnostic looks only where the bytes were LAST time sends the next reader
# spelunking again, so name the cache paths explicitly.
docker run --rm --entrypoint sh "$IMAGE" -c \
'du -sm /root/.npm /root/.cache /tmp/node-compile-cache /home/developer/.npm 2>/dev/null | sort -rn' \
2>/dev/null | sed 's/^/ /' || true
echo " ── foreign platform dirs (npm 11 regression sentinel) ──"
docker run --rm --entrypoint sh "$IMAGE" -c \
'find /usr/lib/node_modules /opt -type d \( -regex ".*/@esbuild/[^/]+" -o -regex ".*/@mariozechner/clipboard-[^/]+" \) -printf "%f\n" 2>/dev/null | sort | uniq -c | sort -rn | head' \
2>/dev/null | sed 's/^/ /' || true
fi
# ── Summary ───────────────────────────────────────────────────────────