Compare commits

..

13 Commits

Author SHA1 Message Date
joakimp 16fddebd43 docs(v1.9.4): the client/server skew is narrower than the release commit said
Lint / hadolint (push) Successful in 8s
Lint / skill-floor (push) Successful in 11s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 21s
Publish Docker Image / resolve-versions (push) Successful in 15s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-base (push) Successful in 1h4m42s
Publish Docker Image / smoke (push) Successful in 6m4s
Publish Docker Image / smoke-studio (push) Successful in 9m18s
Publish Docker Image / build-variant (push) Successful in 19m20s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / build-variant-studio (push) Successful in 24m15s
e3b38cd named "the feeder against the 3.9.0 hub" as the one path to exercise
before tagging, on the reasoning that 3.10.0's CLI writes now follow the daemon
write-routing policy. That was a reading of the mempalace changelog, not a
measurement of this image's feeder. Measured now:

  mempalace-pi-session in remote mode (MEMPALACE_REMOTE_URL set, which is how
  every fleet devbox runs) stages transcripts locally in python, rsyncs them to
  the hub host, and calls the hub's own `mempalace_mine` MCP tool over HTTP
  (run_remote_mine). The local `mempalace` CLI is required only in local mode
  (line 475) and invoked only on the local branch (line 1029). With a PATH shim
  logging every `mempalace` invocation, a 47-session `--dry-run` from this
  container logged ZERO calls; the shim's positive control logged one. The pi
  extension likewise speaks HTTP to the hub and spawns no local mempalace-mcp.

So the 3.10.0 CLI never runs against the 3.9.0 hub from this image. In remote
mode the client pin touches first-run `mempalace init` and the on-disk layout,
nothing else -- which is exactly what the ENV MEMPALACE_CONFIG_DIR change is
for. Both the Dockerfile.base comment and the CHANGELOG paragraph now say so,
and the "Still open" item becomes the thing that is actually unmeasured: a
first-boot acceptance of the built image from an empty ~/.mempalace volume.

Gates re-run on this tree: check-doc-drift.sh rc=0, lint-shell.sh rc=0,
hadolint 2.15.1 rc=0. lint.yml for e3b38cd: run 693, completed, success.
2026-09-22 14:51:22 +02:00
joakimp e3b38cdb0b release(v1.9.4): mempalace 3.10.0 with a pinned palace root, pi-atelier v0.10.3, pi held at 0.85.1
Lint / skill-floor (push) Successful in 8s
Lint / hadolint (push) Successful in 10s
Lint / doc-drift (push) Successful in 13s
Lint / actionlint (push) Successful in 22s
Three pins move or are deliberately held; the CHANGELOG's Unreleased section
becomes the v1.9.4 entry in this same commit, because since ee6cb9e the tag
build runs check-doc-drift.sh check 9 and every would-bake value must be named
above the last published heading.

mempalace 3.9.0 -> 3.10.0 (Dockerfile.base). Deferred at v1.9.3 for two
"Upgrade notes" items; both re-measured against the 3.10.0 wheel:

  - New installs resolve ~/.config/mempalace. The order is $MEMPALACE_CONFIG_DIR,
    then ~/.mempalace IF it holds config.json / people_map.json /
    palace/chroma.sqlite3, then XDG. An EMPTY ~/.mempalace fails that test, and
    an empty ~/.mempalace is what a freshly mounted devbox-palace volume (or
    entrypoint.sh's mkdir on a volume-less container) looks like at first boot.
    Measured with a fresh $HOME and `uvx mempalace==3.10.0 init`: config.json
    landed in ~/.config/mempalace, ~/.mempalace stayed empty, and
    entrypoint-user.sh's `[ ! -d ~/.mempalace/palace ]` would fire every start.
    Same run with MEMPALACE_CONFIG_DIR set: everything in ~/.mempalace.
    => ENV MEMPALACE_CONFIG_DIR=/home/${USER_NAME}/.mempalace, placed in the
    non-root-user section where USER_NAME is in scope, and entrypoint-user.sh
    reads PALACE_DIR from the same variable (old path as fallback). Existing
    volumes were safe either way; this is for first boots. MEMPALACE_PALACE_PATH
    is still honoured (config.py:927), so smoke-test.sh's stage test holds.
  - event_list defaults newest-first without a cursor. Server-side: the
    extension talks to synlig over MEMPALACE_REMOTE_URL, so it lands when the
    hub upgrades. mempalace-toolkit 2167a1b (floating main) makes every
    cursor-less call say order:"desc" so the mailbox reads the same window on
    either server version.

Corrected a stale sequencing comment while there: synlig serves the palace as
a `uv tool` under systemd (mempalace-serve.service; measured 2026-09-22:
mempalace 3.9.0, python 3.12.13, chromadb 1.5.9, no palace container), not via
docker-compose.mempalace.yml as the comment had said for two releases.

pi-atelier v0.10.1 -> v0.10.3 (Dockerfile.variant). Fixes only (both tags
released 2026-09-22); package.json at v0.10.3 still has zero runtime deps and no
build script, peerDeps still >=0.84.0. Annotated tags: the CHANGELOG names the
peeled SHA ed3837b, which is what resolve-versions bakes and check 9 compares.

pi HELD at 0.85.1 while 0.86.1 / 0.87.0 exist. 0.87.0 removed the
shouldStopAfterTurn agent option; pi-observational-memory 3.1.4 still uses it
and AgentContext.systemPrompt in observer/reflector/dropper (measured: 1 hit
per worker in src/agents on both the baked 3.1.3 tree and upstream master;
finishTurn 0). peerDeps are `*`, so it would break at runtime, not install.
Upstream #82, fix PR #83 open and unmerged. Bump pi and pi-obsmem together
once #83 ships. The other extensions are clean against the 0.86/0.87 lists.

Verified locally the way CI does: check-doc-drift.sh rc=0 with check 9
reporting all four moved components as named above the v1.9.3 heading;
lint-shell.sh rc=0; hadolint 2.15.1 (CI's pin, with .hadolint.yaml) rc=0 on
both Dockerfiles; PyPI 3.10.0 present, not yanked. Two claims caught by
re-measurement before commit and corrected in the text: a pi issue number
that does not exist on the 0.87.0 changelog line, and "8 shouldStopAfterTurn
hits" that was 3 (one per worker) in src/.

Still open before the tag: run the feeder from the built image against the
3.9.0 hub (dry-run first); recreate acceptance must show ~/.mempalace/palace
still passing and no ~/.config/mempalace appearing. After acceptance: upgrade
synlig's tool to 3.10.0.
2026-09-22 14:26:10 +02:00
joakimp ee6cb9e62a ci(release): a tag must not publish a component no CHANGELOG entry names
Lint / skill-floor (push) Successful in 9s
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 18s
Lint / doc-drift (push) Successful in 16s
lint-gate already existed to enforce "do not RELEASE a tree whose lint failed",
because lint.yml does not run on tag pushes. scripts/check-doc-drift.sh had the
same gap and it was never extended to cover it: check 9 ran only in lint.yml, so
no tag build has ever evaluated it. Add it to lint-gate, which resolve-versions
already needs, so it fails in ~8 s ahead of the 46-minute base build.

For this check the gap is strictly worse than it is for shellcheck. Shellcheck
judges the tree, so green on main is still green at the tag -- the bytes did not
move. Check 9 judges the tree against upstream NOW, and the floating refs it
watches move with no commit here at all, so a green reading on main carries no
information about tag time. v1.9.3 is the worked example: pi-observational-memory
moved cba0334 -> e7d77dc the day AFTER the tag, nothing went red, and it
surfaced only because someone ran the gate by hand.

Measured, not assumed:
- Same file lint.yml calls (one reference in each workflow), not a second copy.
- Works on a CI-shaped checkout: cloned --depth 1 --no-tags (0 tags, 1 commit),
  rc=0. It needs no local tags because `last` comes from the Hub tags API, not
  `git tag`, so the plain actions/checkout@v4 above is sufficient.
- Has teeth: deleting the Unreleased section from that clone gives rc=1 and
  names the component; restoring it gives rc=0.
- ~8 s (7.8-8.3 s measured), vs ~1 s for lint-shell.sh.

Residual, accepted: the gate resolves the refs seconds before resolve-versions
resolves them again, so an upstream push inside that window still slips past.
Check 9 on the next release names it then.
2026-09-21 23:23:29 +02:00
joakimp 0d324f1855 changelog: name pi-obsmem 3.1.4 (e7d77dc) as the next rebuild's implicit adoption
Lint / skill-floor (push) Successful in 13s
Lint / hadolint (push) Successful in 15s
Lint / doc-drift (push) Successful in 14s
Lint / actionlint (push) Successful in 23s
check-doc-drift.sh was rc=1: pi-observational-memory moved cba0334 (3.1.3) ->
e7d77dc (3.1.4) upstream on 2026-09-20, the day after the v1.9.3 tag, and
nothing named it above the v1.9.3 heading.

No code change is needed or wanted here. PI_OBSMEM_REF defaults to master in
Dockerfile.variant and CI's resolve-versions turns that into a SHA at build
time, so the next rebuild adopts this whether or not anyone acts -- which is
precisely why it has to be written down. There is no pin in this repo to bump,
so the CHANGELOG is the only record a reader of the next tag would have.

v1.9.3 itself has no defect: its manifest, image labels and baked
/opt/pi-observational-memory clone all agree on cba0334. The drift is
post-tag, so a green drift reading taken on 2026-09-19 was correct when taken.

Measured, not assumed: the range cba0334...e7d77dc is 4 commits whose only
substance is one line in src/agents/worker-stream.ts (preserve registry
receiver when resolving stream, PR #78 / issue #77) plus a 19-line regression
test; the rest is the version bump and release merge. peerDependencies are
unchanged (all four @earendil-works/* peers still *) and there is no engines
block, so no pi or node floor to clear.

Gate now rc=0: "OK pi-obsmem cba0334 -> e7d77dc since v1.9.3, named above the
v1.9.3 heading".
2026-09-21 23:15:11 +02:00
Joakim Persson 6c13f43ac8 release: date the v1.9.3 section for the tag
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 16s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 41m49s
Publish Docker Image / smoke (push) Successful in 8m45s
Publish Docker Image / smoke-studio (push) Successful in 9m8s
Publish Docker Image / build-variant (push) Successful in 17m42s
Publish Docker Image / update-description (push) Successful in 11s
Publish Docker Image / promote-base-latest (push) Successful in 26s
Publish Docker Image / build-variant-studio (push) Successful in 21m28s
CI checks out the tag, so the CHANGELOG heading has to name the version
before the tag exists (same convention as v1.9.1/v1.9.2). check-doc-drift
check 9 was sabotage-tested for exactly this rename and still passes:
everything above the published v1.9.2 heading is the pending text.

Contents of v1.9.3 (all measured, nothing inherited):
  - f25efa0  recreate-sanity-check resolves the ControlPath ssh actually
             uses (ssh -G); the ~/.pi/ssh/config hyphen typo deleted
  - b8d818e  Hub size claims corrected; check 8 gates them against full_size
  - c7d369f  promote-base-latest's conditional re-tag measured (run 669)
  - 50153e6  skill floor -> pi-extensions 25c1265 (task tool + fork-gate)
  - ea89605  changelog for the two floating-ref changes: pi-extensions
             25c1265 and mempalace-toolkit 817b3a8 (mine deadline)
  - cb6d9e5  check 9: what the next build bakes differently must be named
             in the CHANGELOG; first run found pi-obsmem 7b397f4->cba0334
  - b9057fd  LABEL se.jordbo.pi-devbox.mempalace-version (inherited)
  - 960aada  dependency audit 2026-09-19; mempalace 3.10.0 deferred

Floating refs this tag will bake, named per check 9: pi-toolkit 9c87ee8,
pi-extensions 25c1265, mempalace-toolkit 817b3a8, pi-observational-memory
cba0334 (3.1.3). Base REBUILDS (Dockerfile.base + rootfs/ changed), so the
BASE_REBUILD_DATE marker is refreshed now, when it is free -- it had been
stale since 2026-07-13 through two rebuilds.

Deliberately NOT in this tag: mempalace 3.10.0 (see the audit section for
the three measured preconditions).
2026-09-19 17:45:50 +02:00
Joakim Persson 960aada769 changelog: dependency audit 2026-09-19 — mempalace 3.10.0 measured and deliberately deferred
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 17s
Lint / skill-floor (push) Successful in 8s
Lint / doc-drift (push) Successful in 13s
Every row by direct command: npm/PyPI/GitHub release APIs, git ls-remote,
the v1.9.2 config labels via the registry API, and binaries in a running
v1.9.2 container for the floating tools. Only pin with a newer upstream is
mempalace 3.9.0 -> 3.10.0, and it is NOT a small bump: new installs move to
~/.config/mempalace (entrypoint-user.sh tests ~/.mempalace/palace and
entrypoint.sh persists ~/.mempalace, so a fresh container would initialise
outside the volume), and mempalace_event_list flips to newest-first by
default (the toolkit's deriveClosed passes no order, deriveOwed on 2 of 3
calls — server-default-dependent). Deferred to its own release with the
three measured preconditions listed.
2026-09-19 17:38:02 +02:00
Joakim Persson b9057fdc8c Dockerfile.base: LABEL se.jordbo.pi-devbox.mempalace-version, inherited by both variants
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 10s
Lint / doc-drift (push) Successful in 16s
Lint / actionlint (push) Successful in 21s
Closes the blind spot check 9 (cb6d9e5) named: no label recorded the palace
pin, so a MEMPALACE_VERSION bump — the one component whose skew against the
shared central palace is fleet-wide — could ship without a CHANGELOG line.

In Dockerfile.base, not Dockerfile.variant, deliberately: the value sits next
to the ARG that defines it (a copy in the variant is one more pin able to
drift); labels are inherited by every image built FROM the base, so no
build-arg to plumb through the variant's four call sites; and inheritance
means the label states the pin of the base the image ACTUALLY built on,
which is the question when base-decide cache-hits an older base. Both
mechanisms measured on the published v1.9.2 config blob rather than assumed:
maintainer + image.source (set only in Dockerfile.base) are present on the
variant image, and pi-version=0.85.1 is an ARG expanded inside a LABEL.

Intent, like every se.jordbo.pi-devbox.* label; the manifest's
mempalace_version (read from the installed binary) stays the ground truth,
and smoke-test.sh now asserts label == installed core — the one way they
diverge is a base built with INSTALL_MEMPALACE=false or an off-pin install,
both invisible to a label-only check.

check-doc-drift check 9 gains the component (literal, against
ARG MEMPALACE_VERSION in Dockerfile.base); the label-key rule generalises to
"names ending in -version are the label itself". Until a release carries the
label it reports a counted SKIP, not OK — measured: "v1.9.2 carries no
se.jordbo.pi-devbox.mempalace-version label", summary says 1 SKIPPED.

Costs nothing extra: this Unreleased already forces a base rebuild
(50153e6 rootfs/ skill floor). check-base-hash unchanged (no new *_REF).
2026-09-19 17:27:34 +02:00
Joakim Persson cb6d9e5dd0 check-doc-drift: check 9 — what the next build bakes differently must be named in the CHANGELOG
Lint / hadolint (push) Successful in 8s
Lint / skill-floor (push) Successful in 10s
Lint / doc-drift (push) Successful in 14s
Lint / actionlint (push) Successful in 19s
The two most recent Unreleased entries (pi-extensions 25c1265: task tool +
fork-gate; mempalace-toolkit 817b3a8: the mine deadline that never reached
the transport) reached this image through floating *_REF=main ARGs. Neither
produced a diff here, so nothing asked for a CHANGELOG line, and neither had
one until a reader asked. Same class as check 8 — a claim with no in-repo
anchor rots — and the hand practice that covered it (the per-release
"Dependency audit" table, Baked-in vs Upstream-now) is a someone-remembers
mechanism that had lapsed.

Method, no docker/crane/token: last published vX.Y.Z from Hub's tag list
(one request, now shared with check 8); its amd64 config blob via the
anonymous registry API carries one se.jordbo.pi-devbox.<name>-ref label per
component with the SHA the build-args baked. "Would bake" is resolved the way
resolve-versions does it: 40-hex ARG is itself, branch/tag via git ls-remote
with the peeled ^{} form preferred, pi-studio = highest semver tag (label on
<tag>-studio), PI_VERSION as a literal against the pi-version label. Nine
components, ~7.5 s. Unchanged: OK. Moved: the new value's 7-char prefix (tag
for studio, version for pi) must appear above the last published version's
"## " heading — Unreleased plus any not-yet-published section, which is what
the release commit turns Unreleased into, so the tag build passes on the same
text. A published tag with no heading is a failure, not a skip.

First run found a move nobody had recorded: pi-observational-memory 7b397f4
-> cba0334 (6 commits, 3.1.1 -> 3.1.3; memory workers now routed by the
model's exact provider instead of the first registered provider with a
matching api — upstream #70). Named in the CHANGELOG with the honest scope:
no expected effect on the shipped bedrock-only configuration, unmeasured.

Sabotage-tested, expectation written before each run:
  obsmem SHA removed from CHANGELOG ............ rc=1  DRIFT
  Unreleased renamed to "## v1.9.3 — …" ........ rc=0  (release-commit shape)
  "## v1.9.2" heading mangled to v1.9.2-typo .... rc=1  — this one FAILED first:
      \b accepted "v1.9.2-typo" as v1.9.2's heading; now (\s|$)
  bare "## v1.9.2" (no date) ................... rc=0
  ARG PI_OBSMEM_REPO renamed ................... rc=2  (blind gate must not pass;
      read_arg calls are bare top-level assignments on purpose — nested in a
      heredoc's $(...) the exit 2 is swallowed by cat)
  offline (proxy to a closed port) ............. rc=0, 2 SKIPs counted
  SKIP_REF_CHECK=1 ............................. rc=0, 1 SKIP counted
  restored ..................................... rc=0, files byte-identical
GIT_TERMINAL_PROMPT=0 on ls-remote: a repo flipped private fails in 0.3 s as
a SKIP instead of waiting for a username until the job times out.

Header, lint.yml job comment and AGENTS.md all still said this gate "needs
no network" — false since check 8 (2026-09-14). Now: two classes, hermetic
1-7 and published-state 8-9 that SKIP loudly and counted when offline.
Considered and not added: compose <-> .env.example variable cross-check (the
four mismatches are commented-out lines, the mempalace-server compose file's
own variables, and entrypoint-consumed ones — it would fire on nothing
wrong); "documented tag exists on Hub" (check 8 already SKIPs by name; a
hard fail would misreport the tag-to-publish window).
2026-09-19 17:15:12 +02:00
Joakim Persson ea896054df changelog: task tool + fork-gate (pi-extensions 25c1265), and the mine deadline that never reached the transport (mempalace-toolkit 817b3a8)
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 18s
Lint / doc-drift (push) Successful in 6s
Both land in the image through floating refs (PI_EXTENSIONS_REF=main,
MEMPALACE_TOOLKIT_REF=main), so neither shows up as a diff in this repo — the
changelog is the only place a reader of a pi-devbox tag learns that the
container's delegation and feed behaviour changed. Recorded under Unreleased
with the measurements: 5/5 real fork briefs redirected, live pi -p proof for
gate and tool, 2-of-6 → 6-of-6 on the toolkit's new timeout test.

Dockerfile.base: the "Stall protection" comment listed MEMPALACE_MCP_TIMEOUT_MS
as THE tool-call deadline; it now says the feed's mine carries its own.
Comment-only, but Dockerfile.base is hashed into base_tag; the rebuild was
already forced by 50153e6 (rootfs/ skill floor).
2026-09-19 17:00:59 +02:00
Joakim Persson 50153e65b7 skill floor: refresh vendored pi-extensions skill to pi-extensions@25c1265 (task tool + fork-gate)
Lint / hadolint (push) Successful in 11s
Lint / doc-drift (push) Successful in 10s
Lint / skill-floor (push) Successful in 10s
Lint / actionlint (push) Successful in 21s
check-skill-floor.sh: OK, tree_sha256 9b85a633… matches the package at main (25c1265).
Forces one base rebuild (base_tag hashes rootfs/), which is also what bakes
extensions/task.ts and fork-gate.ts into /opt/pi-extensions via the floating
PI_EXTENSIONS_REF=main; install.sh symlinks every extensions/*.ts on start.
2026-09-19 16:42:06 +02:00
Joakim Persson f25efa074d fix(sanity): verify the ControlPath ssh RESOLVES, not the dir the image creates
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 13s
Lint / doc-drift (push) Successful in 12s
Lint / actionlint (push) Successful in 24s
recreate-sanity-check.sh asserted `/tmp/sshcm exists with mode 700` and printed
a green tick while every ssh in the container died rc=255. It was right about
what it checked: the breakage was in the directory a CONFIG named, not the one
the image creates, and the old check could not see the disagreement between them.

Found on the v1.9.2 first boot on emb-7kj4vr4g. A durable ~/.pi/ssh/config,
hand-written into the ~/.pi named volume by the previous session so it would
survive the recreate, declared `ControlPath /tmp/ssh-cm/%C` — with a hyphen.
Nothing here creates that path; the canonical spelling is /tmp/sshcm, identical
in Dockerfile.base, entrypoint-user.sh, recreate-sanity-check.sh and
smoke-test.sh. ControlMaster auto with an unusable ControlPath does not degrade
to an unmultiplexed connection — it fails hard:

    unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory

rc=255, remote command never runs.

Now: resolve the ControlPath ssh itself would use via `ssh -G` and require its
parent to exist and be writable. -G applies real config precedence (first-value-
wins, the system drop-in, Include, an -F override), so it answers "which rule
captured this host" instead of re-implementing the guess, and it never opens a
connection — 0.116s for 48 hosts.

This also puts a check under a caveat documented in prose in Dockerfile.base
("SSH client defaults") and verified nowhere: a per-host ControlPath under a
read-only bind-mounted ~/.ssh gives the identical failure, "Read-only file
system". On the machine this was built on that is 16 of 50 hosts — freeipa-1..6,
gitea.egl.lan, runner-1..3, tor-ms22 — never reported by anything before.

Severity split is deliberate. Default ssh precedence legitimately lands in the
read-only ~/.ssh for any host whose own config pins it there, and the supported
workaround (ssh -F ~/.ssh-local/config, generated every start by
setup-lan-access.sh) already exists, so that is a warn. Failing it would paint
the script red on every run of every device, and a check that fires benignly
every time is one you learn to ignore — the same reasoning that keeps
lint-shell.sh at -S error. The sidecar route is prescribed, so there it is a
hard fail. Host lists cap at six names plus a count: unreadable output is
ignored output.

Teeth proven both directions, each sabotage confirmed by diff BEFORE the result
was believed: ControlPath -> nonexistent dir => rc=1; -> existing-but-read-only
dir => rc=1; restored => rc=0 with the sidecar byte-identical.

No config was added to fix the original problem — the fix was to DELETE
~/.pi/ssh/config, a third hand-maintained copy of what setup-lan-access.sh
already generates from version control on every container start, with fewer
features and one typo. The host-owned ~/.ssh/config is left alone on purpose:
those ~/.ssh/cm paths are correct on the host, where ~/.ssh is writable.

Also lint-shell.sh: the SC2088 count in the severity-choice rationale said 19
and is now 20, with the reproduce command recorded. That command needs a `$ `
prefix — a comment whose first word is "shellcheck" is parsed as a directive,
and the malformed one tripped SC1072/SC1073 at severity error. The gate caught
it on the very commit that introduced it.
2026-09-15 22:59:43 +02:00
joakimp c7d369f28d docs(changelog): measure promote-base-latest's conditional re-tag
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 20s
Closes the open question from v1.9.2's release verification. crane copy DOES
bump Hub's last_updated when it runs (run 669: digests differed, copy ran
17:10:04->17:10:08, base-latest.last_updated moved to 17:10). The
identical-digest path stays unproven by construction -- the job prints
'base-latest already current; nothing to promote.' and never copies -- so a
freshness assertion on the alias is conditional on base-decide's need_build,
not a blanket upgrade. Operational rule lives in the ci-release-watcher skill
(skillset c8034af); the pipeline property is recorded here.
2026-09-14 20:49:23 +02:00
joakimp b8d818ed99 fix(docs): correct the Hub size claims, and gate them so they cannot rot again
Lint / hadolint (push) Successful in 9s
Lint / doc-drift (push) Successful in 13s
Lint / skill-floor (push) Successful in 15s
Lint / actionlint (push) Successful in 18s
DOCKER_HUB.md claimed ~1.1 GB for :latest while Docker Hub served 1.37 GB --
20% wrong, drifting quietly across eight releases, and POSTed to Docker Hub by
update-description every time. It is the first number a stranger reads about
this image.

Root cause is structural, not carelessness: every OTHER claim check-doc-drift.sh
guards is anchored to a build file (a pin, an ARG, a placeholder), so it cannot
rot without someone editing the thing it describes. Nothing in this repo states
the image size, so nothing measured it.

Corrected against measured full_size after v1.9.2 published (amd64/arm64):
  :latest         ~1.1  -> ~1.23 GB   (1.228 / 1.211)
  :latest-studio  ~1.15 -> ~1.25 GB   (1.255 / 1.238)
  :base-latest    ~1.0  -> ~1.17 GB   (1.167 / 1.151)
full_size tracks the FIRST manifest entry (amd64), not the sum across arches --
measured: full_size=1.228, amd64=1.228, arm64=1.211, sum=2.439 -- which is what
the table's per-arch "Size (compressed)" column claims.

New check 8 gates those claims against Hub. Fails on drift; SKIPS LOUDLY via a
new skip() helper (counted, named in the summary) when curl/python3 are absent,
the API is unreachable, or SKIP_SIZE_CHECK=1. A skip is deliberately neither OK
nor a failure: printing an unverified claim as OK is the habit this file exists
to break, and failing on Docker Hub's uptime would hold releases hostage to a
third party. Not wired into hooks/pre-push, so pushes stay offline.

Two bugs caught by writing the expected exit code down before running the check:

  1. The tolerance would have missed its own motivating case. The percentage is
     computed against the MEASURED size, but the 20% first chosen came from the
     claim-relative figure; the real drift was |1.1-1.37|/1.37 = 19.7% and would
     have passed. Now 15%, inside a window with both bounds measured: above the
     largest legitimate skew (11.4%, a claim describing the published release
     while the next tag changes the size) and below the rot it must catch.

  2. A `|| true` on the python invocation made the gate FAIL OPEN -- it printed
     DRIFT and exited 0. Removed; the outer `|| SIZE_RC=$?` satisfies set -e
     without swallowing the code. Verified two-sided: 19% drift exits 1 at the
     default tolerance and 0 at SIZE_TOLERANCE_PCT=25, so the threshold does the
     work rather than the ordering.

NOT covered, and the script says so where a reader will see it: README.md's
~3.2 GB figures are UNCOMPRESSED and the registry exposes compressed sizes only,
so measuring them needs a real pull. A green check 8 says nothing about them.

Verified: check-doc-drift.sh rc=0 on the corrected tree, rc=1 on injected drift,
rc=0 with --warn-only, SKIP+rc=0 on the offline path (proxy to a dead port);
shellcheck clean at default severity (the sed single-quote SC2016 carries a
disable with its reason); lint-shell.sh 16 files clean.
2026-09-14 20:33:40 +02:00
14 changed files with 1427 additions and 60 deletions
+26
View File
@@ -174,6 +174,29 @@ jobs:
# #
# ~40 s, ahead of everything expensive, and it runs scripts/lint-shell.sh -- # ~40 s, ahead of everything expensive, and it runs scripts/lint-shell.sh --
# the same file lint.yml calls, not a second copy that drifts. # the same file lint.yml calls, not a second copy that drifts.
#
# scripts/check-doc-drift.sh is here for the same reason and closes the same
# gap -- and for it the gap is strictly worse. Shellcheck judges the TREE:
# green on main is still green at the tag, because the bytes did not move.
# Check 9 judges the tree against UPSTREAM NOW, and the floating refs it
# watches (PI_OBSMEM_REF=master and friends, which resolve-versions below turns
# into SHAs) move with no commit in this repo at all -- so a green reading on
# main carries no information about tag time, and that window is exactly where
# releases live. Worked example: pi-observational-memory moved cba0334 ->
# e7d77dc the day AFTER v1.9.3 was tagged. Nothing went red; it surfaced only
# because someone ran the gate by hand. Without this step a tag can publish a
# component that no CHANGELOG entry names, and the floating ref means no other
# file in the repo would record it either.
#
# Adds ~8 s. No token and no built image: it takes the last published vX.Y.Z
# from the Hub tags API, that release's baked labels from the anonymous
# registry API, and `git ls-remote`s each upstream -- so a plain checkout is
# enough, with no tags or history to fetch. Offline it SKIPs loudly and
# counted rather than passing, so an outage degrades it to a visible skip
# instead of a false green. Residual, accepted: it resolves the refs seconds
# before resolve-versions resolves them again, so an upstream push landing
# inside that window still slips through -- and check 9 on the NEXT release
# would then name it.
lint-gate: lint-gate:
runs-on: ubuntu-latest runs-on: ubuntu-latest
container: container:
@@ -189,6 +212,9 @@ jobs:
- name: "Shellcheck + syntax-check repository scripts (severity: error)" - name: "Shellcheck + syntax-check repository scripts (severity: error)"
run: bash scripts/lint-shell.sh run: bash scripts/lint-shell.sh
- name: Components the next build would bake differently are named in the CHANGELOG
run: bash scripts/check-doc-drift.sh
resolve-versions: resolve-versions:
# Gated: a defective tree must not reach a 46-minute base build. # Gated: a defective tree must not reach a 46-minute base build.
needs: [lint-gate] needs: [lint-gate]
+11 -4
View File
@@ -197,10 +197,17 @@ jobs:
# someone remembering. It is also read from the TAG, so a fix pushed to main # someone remembering. It is also read from the TAG, so a fix pushed to main
# after tagging never reaches the published page. # after tagging never reaches the published page.
# #
# Cheap and hermetic on purpose: every check compares a doc string against a # Two classes of check. 1-7 are hermetic: each compares a doc string against
# value that exists in this repo, so no network, no token, no built image, # a value that exists in this repo — no network, no token, no built image.
# and no sibling clone. Claims that genuinely need a running container (image # 8-9 compare against what is PUBLISHED, because those claims have no
# sizes, the "N mempalace_* tools" count) are deliberately left out — a gate # in-repo anchor and rotted for exactly that reason: 8 reads Docker Hub's
# measured sizes; 9 reads the ref labels baked into the last released image
# (anonymous registry API, no docker/crane) and `git ls-remote`s each
# floating upstream, then requires every component the next build would
# bake differently to be NAMED in the CHANGELOG above that release's
# heading. Both SKIP loudly and counted when offline — a skip is neither OK
# nor a failure. Claims that genuinely need a running container (the "N
# mempalace_* tools" count, uncompressed sizes) are still left out — a gate
# that cannot evaluate a claim honestly would have to guess, and a guessing # that cannot evaluate a claim honestly would have to guess, and a guessing
# gate is worse than none. Assert those in scripts/smoke-test.sh instead. # gate is worse than none. Assert those in scripts/smoke-test.sh instead.
# #
+13 -1
View File
@@ -103,7 +103,19 @@ re-brand of opencode-devbox's `pi-only` variant.
It compares README.md's version-pin table against the ARGs it names, and It compares README.md's version-pin table against the ARGs it names, and
DOCKER_HUB.md's Node claim against `ARG NODE_VERSION`, plus Hub's DOCKER_HUB.md's Node claim against `ARG NODE_VERSION`, plus Hub's
25 000-char limit, unsubstituted `{{PLACEHOLDERS}}`, and stale `Unreleased` 25 000-char limit, unsubstituted `{{PLACEHOLDERS}}`, and stale `Unreleased`
pointers in user-facing docs. pointers in user-facing docs. With network it also checks DOCKER_HUB.md's
size claims against Hub's measured sizes (check 8) and — check 9 — that
**every component the next build would bake differently from the last
published release is named in the CHANGELOG** above that release's heading:
it reads the `se.jordbo.pi-devbox.*-ref` labels off the published image and
`git ls-remote`s each floating `*_REF`. A red check 9 means an upstream
(pi-toolkit, pi-extensions, mempalace-toolkit, pi-fork,
pi-observational-memory, pi-studio) moved and no entry names the new SHA;
the failure prints the compare URL; a `PI_VERSION` or `MEMPALACE_VERSION`
bump is caught the same way via the `pi-version` / `mempalace-version`
labels. Name the 7-char SHA (or version) where you describe
the change — that is what the old "Dependency audit" tables recorded by
hand, now required.
**Why before and not after:** `docker-publish.yml` runs `actions/checkout@v4` **Why before and not after:** `docker-publish.yml` runs `actions/checkout@v4`
with no `ref:`, so every job reads `github.ref` — the **tag**. A doc fix with no `ref:`, so every job reads `github.ref` — the **tag**. A doc fix
+548
View File
@@ -11,6 +11,554 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
--- ---
## v1.9.4 — 2026-09-22
### Dependency audit (2026-09-22)
Every component checked against upstream by direct command, not assumed. "Baked"
is v1.9.3's published labels or, for the floating `*_VERSION=latest` tools, the
binaries in a running v1.9.3 container. The 13 GitHub-`latest` tools were
resolved through the same `/releases/latest` redirect the Dockerfile follows.
| Component | Baked in v1.9.3 | Upstream now | Action |
|---|---|---|---|
| **pi** | `0.85.1` (pinned) | **`0.87.0`** (0.86.0, 0.86.1, 0.87.0 since) | **held — blocked by pi-observational-memory, see below** |
| **pi-atelier** | `v0.10.1` (`734258b`) | **`v0.10.3`** (`ed3837b`, released 2026-09-22) | **bumped** |
| **mempalace** | `3.9.0` (pinned) | **`3.10.0`** (changelog dated 2026-09-15) | **bumped, with one adaptation** |
| **mempalace-toolkit** | `817b3a8` | **`2167a1b`** | floating `main`; the commit is this release's own (below) |
| **pi-observational-memory** | `cba0334` (3.1.3) | **`e7d77dc`** (3.1.4) | adopted implicitly via `master` (below) |
| pi-toolkit | `9c87ee8` | `9c87ee8` | none |
| pi-extensions | `25c1265` | `25c1265` | none |
| pi-fork | `e69725c` | `e69725c` | none |
| pi-studio (studio variant) | `e04fc7a` | `e04fc7a` highest semver tag | none |
| skillset (mempalace fallback snapshot) | `e9e45f7` | skillset at `1c5f960`; `check-doc-drift.sh` rc=0 (snapshot still byte-identical) | none |
| gitea-mcp, agent-browser, node | `1.7.0`, `0.38.1`, `v24.21.0` | identical | none |
| 13 floating `*_VERSION=latest` tools | — | all 13 identical to the running v1.9.3 binaries | none |
### mempalace 3.9.0 → 3.10.0, and why one `ENV` line comes with it
v1.9.3 deferred 3.10.0 for two "Upgrade notes" items. Both were re-measured
against the 3.10.0 wheel rather than the changelog; one needed an adaptation, the
other turned out to be the other side's problem.
**"New installs keep config and palace under `~/.config/mempalace`."** The
resolution order (`config.py`, `_default_config_dir`) is `$MEMPALACE_CONFIG_DIR`,
then `~/.mempalace` *if* it holds `config.json`, `people_map.json` or
`palace/chroma.sqlite3`, then XDG. An **empty** `~/.mempalace` fails that test —
and an empty `~/.mempalace` is exactly what a freshly mounted `devbox-palace`
volume looks like at first boot (and what `entrypoint.sh`'s `mkdir` leaves on a
volume-less container). Measured with a fresh `$HOME` and `uvx mempalace==3.10.0`:
`mempalace init` wrote `config.json` to `~/.config/mempalace`, `~/.mempalace`
stayed empty, and `entrypoint-user.sh`'s first-run test `[ ! -d ~/.mempalace/palace ]`
would have stayed true on every start. Same run with `MEMPALACE_CONFIG_DIR` set:
everything landed in `~/.mempalace`.
So `Dockerfile.base` now sets `ENV MEMPALACE_CONFIG_DIR=/home/${USER_NAME}/.mempalace`
(in the non-root-user section, where `USER_NAME` is in scope), and
`entrypoint-user.sh` reads `PALACE_DIR` from the same variable with the old path
as fallback — the first-run test and mempalace's own resolution can no longer
disagree. Existing volumes were safe either way (`config.json` is a legacy
marker; this container's volume holds one); the `ENV` is for first boots.
`palace_path` still defaults to `<config_dir>/palace` and `MEMPALACE_PALACE_PATH`
is still honoured (`config.py:927`), so `smoke-test.sh`'s stage-path test keeps
its meaning, and `recreate-sanity-check.sh`'s hard-coded `~/.mempalace/palace/chroma.sqlite3`
stays correct.
**"MCP event listing returns the newest events first when no cursor is given."**
This is **server-side**. The pi extension talks to the fleet hub over
`MEMPALACE_REMOTE_URL`, so `mempalace_event_list`'s default flips when *synlig*
upgrades, not when this image does. `mempalace-toolkit` `2167a1b` (this release)
makes every cursor-less `event_list` call in the extension say `order: "desc"`
explicitly — three of five did not — so the mailbox's `deriveOwed` and
`deriveClosed` read the same newest-N window against either server version. Its
selection was already order-independent (`isStrictlyAfter`); what the fix changes
is *which* events are in the window: on a 3.9.0 hub the cursor-less calls were
returning the **oldest** N, the latent truncation `deriveOwed` had already fixed
for two of its calls in toolkit `e2b060a` (2026-09-09).
Also in the upgrade notes, neither reaching this image: `mempalace rules` dropped
`--agent` (zero callers in pi-devbox, mempalace-toolkit, skillset, myconfigs);
`get_collection()` refuses unknown collection names (library callers only). MCP
tool-schema review — the regression class this pin exists for: no tool removed or
renamed; additive fields on search results (`filed_at` / `content_date`
provenance), `limit`/`offset` on `kg_timeline`, `last_modified` on drawers.
**Server state, measured over ssh on 2026-09-22:** synlig serves mempalace
**3.9.0** as a `uv tool` under `systemd` (`mempalace-serve.service`, python
3.12.13, chromadb 1.5.9) — *not* via `docker-compose.mempalace.yml`, which the
`Dockerfile.base` sequencing comment had claimed for two releases (corrected in
this commit). Client and server are level today; this image reintroduces skew
until synlig runs `uv tool upgrade mempalace` and restarts the unit. That is the
step that lights up the server-side changes above. The skew meanwhile is narrower
than the release commit's message claimed: the pi extension speaks HTTP to the hub
(no local `mempalace-mcp` is spawned), and the feeder in remote mode stages
locally in python, rsyncs, and calls the hub's own `mempalace_mine` tool.
Measured with a PATH shim in front of `mempalace`: a 47-session
`mempalace-pi-session --dry-run` from this container made **zero** local CLI calls
(the shim's positive control logged one). So 3.10.0's new CLI write-routing
policy never runs against the hub from this image; in remote mode the client pin
touches first-run `mempalace init` and the on-disk layout, nothing else.
### pi-atelier v0.10.1 → v0.10.3 (`734258b` → `ed3837b`)
Fixes only per the release notes for v0.10.2 and v0.10.3 (both 2026-09-22):
sidebar text/borders preserved beside inline images (#53), transcript images
hidden while capturing overlays are open, Workspace Pulse skips redundant
HEAD/diff when nothing tracked changed (#61), sidebar height from row counts
(#59), git/usage scans suspended while disabled, Display Revert + Undo ordering.
Checked before bumping: `package.json` at v0.10.3 still declares zero runtime
dependencies and no build script (the no-`npm install` reasoning in
`Dockerfile.variant` holds), `peerDependencies` still `>=0.84.0` (spans the
pinned pi 0.85.1), 13 commits all under `src/ tests/ docs/ scripts/` plus
metadata, no entry-point move. Both tags are annotated: the SHA above is the
peeled commit, which is what `resolve-versions` bakes and what check 9 compares.
### pi held at 0.85.1 (0.86.1 and 0.87.0 exist)
Not an oversight. pi 0.87.0 *"Removed the inherited `shouldStopAfterTurn` agent
option. Use `finishTurn` and return `{ action: "end" }` instead"*, and 0.86.0
moved provider stream inputs to `TranscriptContext` with the system prompt read
from `context.messages`. pi-observational-memory 3.1.4 still uses both the
removed option and `AgentContext.systemPrompt` in its observer, reflector and
dropper workers — measured in `src/agents/*/agent.ts` on both the baked 3.1.3
tree and upstream `master`: `shouldStopAfterTurn` once per worker, `finishTurn`
zero. Its `peerDependencies` are `*`, so nothing at install time would refuse;
under 0.87 the workers would lose their turn caps and their specialised prompts
at runtime. Upstream tracks it as pi-observational-memory
[#82](https://github.com/elpapi42/pi-observational-memory/issues/82) with fix PR
[#83](https://github.com/elpapi42/pi-observational-memory/pull/83) — opened
2026-09-21, mergeable, **not merged** at this writing. Bump pi and pi-obsmem
together once #83 ships in a release.
The other extensions were checked against the 0.86.0 and 0.87.0 breaking lists
and are clean: `ssh-controlmaster`'s `user_bash` handler already returns
`undefined | { operations }` (0.86.0's fail-closed contract), and its four
`registerTool` calls spread the built-in tools so they carry parameter schemas
(#9300). 0.86.1 as an intermediate was not tested — not worth the pty matrix for
a stop that #83 makes moot.
### Still open
- **First-boot acceptance on the built image** — the palace-path adaptation was
measured with the wheel, not the image. Start a throwaway container from the
published v1.9.4 with an *empty* `~/.mempalace` volume and confirm the
entrypoint's `mempalace init` lands `config.json` there, that
`~/.config/mempalace` does not appear, and that `MEMPALACE_CONFIG_DIR` is in
the container environment. On a real device the recreate checklist's
`recreate-sanity-check.sh` should keep passing `~/.mempalace/palace/chroma.sqlite3`.
- **synlig upgrade to 3.10.0** after v1.9.4 is accepted on one device:
`uv tool upgrade mempalace`, restart `mempalace-serve.service`, verify one
search and one `event_list`.
- **pi 0.87.x + pi-observational-memory ≥ 3.1.5** together, when #83 has shipped.
### pi-observational-memory 3.1.3 → 3.1.4 (`cba0334` → `e7d77dc`)
No change in this repo. `PI_OBSMEM_REF` defaults to `master`
(`Dockerfile.variant`) and CI resolves it to a SHA at build time
(`docker-publish.yml`, `resolve-versions`), so the next rebuild bakes this
whether or not anyone acts — which is exactly why it is written down here. The
floating ref means this CHANGELOG is the only place a reader of the next tag can
learn that the component moved.
v1.9.3 shipped `cba0334` (3.1.3) and was internally consistent: its manifest,
its image labels and the baked `/opt/pi-observational-memory` clone all agree on
that SHA. Upstream `master` then moved to `e7d77dc` (3.1.4) on **2026-09-20**,
the day *after* the v1.9.3 tag — so the drift is real but v1.9.3 has no defect,
and a green `check-doc-drift.sh` reading taken before that date was correct when
it was taken.
| | |
|---|---|
| Upstream range | [`cba0334...e7d77dc`](https://github.com/elpapi42/pi-observational-memory/compare/cba03347a60af8b8afbb35677ade4d6bb05d4c5a...e7d77dc9a8305acb8054124e47662b3c766c2321) — 4 commits |
| Substance | one line in `src/agents/worker-stream.ts` — *preserve registry receiver when resolving stream* (PR #78, upstream issue #77) — plus a 19-line regression test |
| Remainder | `chore(release): prepare 3.1.4` (version bump) and the release merge (PR #79) |
| `peerDependencies` | unchanged — all four `@earendil-works/*` peers still `*`, so no pi floor to clear |
| `engines` | absent in both — no node floor |
Because the ref floats, the SHA this repo actually bakes is whatever `master`
resolves to **at build time**. If upstream moves again before the next tag, the
value above is stale and `check-doc-drift.sh` will say so; re-run it immediately
before tagging rather than trusting this row.
---
## v1.9.3 — 2026-09-19
### Dependency audit (2026-09-19)
Every component checked against upstream by direct command, not assumed. "Baked"
is v1.9.2's published amd64 config labels (read through the registry API) or,
for the floating `*_VERSION=latest` tools, the binaries in a running v1.9.2
container.
| Component | Baked in v1.9.2 | Upstream now | Action |
|---|---|---|---|
| pi | `0.85.1` (pinned) | `0.85.1` is npm latest (2026-09-05) | none |
| pi-atelier | `v0.10.1` (pinned) | `v0.10.1` highest tag | none |
| **mempalace** | `3.9.0` (pinned) | **`3.10.0`** (2026-09-16) | **not adopted — see below** |
| skillset (mempalace fallback snapshot) | `e9e45f7` | `--check` OK: `skills/mempalace/SKILL.md` byte-identical at skillset `debc8f6` | none |
| **mempalace-toolkit** | `dab989b` | **`817b3a8`** | ships the mine-deadline fix (above) |
| **pi-toolkit** | `adfb553` | **`9c87ee8`** | ships the `task`-first AGENTS.md (above) |
| **pi-extensions** | `2610545` | **`25c1265`** | ships `task.ts` + `fork-gate.ts` (above) |
| pi-fork | `e69725c` | `e69725c` | none |
| **pi-observational-memory** | `7b397f4` (3.1.1) | **`cba0334`** (3.1.3) | adopted implicitly via `master` (above); peerDeps still `*`, no pi floor to clear |
| pi-studio (studio variant) | `e04fc7a` | `e04fc7a` highest semver tag | none |
| floating `*_VERSION=latest` tools | — | 14 of 16 already at latest; `uv` `0.12.13`→`0.12.17` (four patch releases, none with a Breaking section), `agent-browser` `0.37.1`→`0.38.1` (minor: `screenshot --if-changed`, `snapshot --delta`, persistent refs; `.1` is a recording-timing fix) | adopted implicitly by the rebuild; named here per this repo's floating-ref rule |
| node | major pin `24`, installed `v24.21.0` | `v24.21.0` newest 24.x | none |
**mempalace 3.10.0 is deliberately not in this release.** It is not a small
bump: a Rust exact-vector engine, four modules split into packages, and three
agent-facing contract changes, two of which touch this image directly —
- **New installs put config and palace under `~/.config/mempalace`.**
`entrypoint-user.sh` decides "first run" by `[ ! -d ~/.mempalace/palace ]`
and `entrypoint.sh` provisions `~/.mempalace` as the persisted path. On
3.10.0 a fresh container would initialise into `~/.config/mempalace` —
outside the volume — and the entrypoint's test would stay true on every
start. Whether `MEMPALACE_HOME`/config resolution honours the old path on a
*pre-existing* `~/.mempalace` is stated ("unchanged") but unmeasured here.
- **`mempalace_event_list` defaults to newest-first when no cursor or `order`
is given.** The toolkit's `deriveOwed` passes `order: "desc"` on two of its
three queries and `deriveClosed` on neither, so their windows are
server-default-dependent. That is a server-side skew (the hub is synlig's
stack, whose default image is `joakimp/pi-devbox:latest` — so this pin *is*
the server's next version), and the right fix is in the toolkit: pass
`order` explicitly on every `event_list` call so the derivation is the same
on 3.9 and 3.10 servers. Filed as follow-up; not a blocker for this tag.
- `mempalace rules` dropped `--agent`; `get_collection()` refuses unknown
names — no caller of either in this repo or the toolkit (grepped).
Adopting it needs its own release: the entrypoint path test, a measured
upgrade of an existing `~/.mempalace`, the toolkit `order` hardening, and the
synlig redeploy sequencing the pin comment in `Dockerfile.base` describes.
---
**New check 9 in `scripts/check-doc-drift.sh`: anything the next build would bake
differently from the last *published* release must be named in the CHANGELOG text
above that release's heading.** The two entries below this one are why. The
`task` tool and `fork-gate` (pi-extensions `25c1265`) and the mine-deadline fix
(mempalace-toolkit `817b3a8`) both reach this image through floating
`*_REF=main` ARGs, so neither produced a diff in this repo, nothing here asked
for a CHANGELOG line, and neither had one until a reader asked. Same shape as
check 8: a claim with no in-repo anchor rots. The hand practice that existed for
it — the "Dependency audit" table in each release's notes, *Baked in vN* against
*Upstream now* — is a "someone remembers" mechanism, and it had lapsed.
How it measures, with no `docker`, `crane` or token: the last published
`vX.Y.Z` is the highest such tag in Hub's list (one request, shared with
check 8); that tag's amd64 config blob is read through the anonymous registry
API (token → index → per-arch manifest → config) and carries one
`se.jordbo.pi-devbox.<name>-ref` label per component holding the SHA the
build-args actually baked. "What the next build would bake" is resolved the way
`resolve-versions` does it — a 40-hex ARG is itself, a branch or tag is
`git ls-remote`d with the peeled `^{}` form preferred (the un-dereferenced SHA of
an annotated tag is the tag object; this repo has raised that false alarm once
already), pi-studio is the highest semver tag read from `<tag>-studio`'s labels,
and `PI_VERSION` is compared as a literal against the `pi-version` label. Nine
components, 7.5 s.
The rule: unchanged needs no mention. Moved requires the new value's 7-char SHA
prefix (tag name for pi-studio, version string for pi) somewhere above the last
published version's `## ` heading — `## Unreleased` plus any not-yet-published
`## vX.Y.Z`, which is what the release commit turns Unreleased into, so the tag
build passes on the same text — sabotage-tested: renaming `## Unreleased` to
`## v1.9.3 — …` stays green; mangling the published `## v1.9.2` heading goes
red, and that test caught a `\b` that would have accepted `v1.9.2-typo` as the
heading (now `(\s|$)`). Naming the SHA rather than the repo is
deliberate: it is what the audit table always recorded, and it makes the failure
message's compare URL one click from knowing what moved. Every upstream commit
re-reds the gate until the CHANGELOG names the new head; that is the intended
cost — **the thing that gets baked is the thing that gets named.** A published
tag with no CHANGELOG heading is a failure, not a skip.
**First run found a move nobody had recorded.** `pi-observational-memory`
`7b397f4 → cba0334` (6 upstream commits, 2026-09-14..16, 3.1.1 → 3.1.3): the
memory workers' `streamSimple` lookup used to iterate every extension-registered
provider and take the first whose `api` matched the model's, so two providers
sharing an API type could route the observer to the wrong one (upstream #70);
it now asks for the model's exact provider and keeps the `api` match as a
consistency check. Reaches this image on the next variant build via
`PI_OBSMEM_REF=master`. No behaviour change expected on the shipped
configuration — every profile here uses the built-in `amazon-bedrock` provider
and no extension registers one — but that is an expectation, not a
measurement; the code path only differs when an extension has called
`registerProvider`.
Header and `lint.yml` comment corrected alongside: both still said this gate
"needs no network", which check 8 made false on 2026-09-14. There are now two
classes — hermetic checks 1–7, and published-state checks 8–9 that SKIP loudly
and counted when offline. Considered and not added, with reasons in the script:
a `docker-compose.yml` ↔ `.env.example` variable cross-check (the four
mismatches are commented-out lines, the mempalace-server compose file's own
documented variables, and two entrypoint-consumed variables — a gate there
would fire on nothing wrong), and a "documented tag exists on Hub" check
(check 8 already SKIPs a missing tag by name, and a hard fail would
misreport the window between tagging and publish).
**Blind spot closed while it was free: `se.jordbo.pi-devbox.mempalace-version`.**
No label recorded the palace pin, so check 9 could not see a `MEMPALACE_VERSION`
bump — checks 1–3 keep README's pin table consistent, but nothing required a
CHANGELOG line for the one component whose skew against the shared central
palace is fleet-wide. The label is set in `Dockerfile.base` next to the `ARG`
that defines it and **inherited** by both variants: no second copy of the pin to
drift, no build-arg to plumb through four variant call sites, and it states the
pin of the base the image *actually* built on — the question that matters when
`base-decide` cache-hits an older base. Intent, like every label here; the
manifest's `mempalace_version` stays the ground truth, and `smoke-test.sh` now
asserts label == installed binary (the one way they diverge is a base built with
`INSTALL_MEMPALACE=false`, or an install that resolved off-pin). Until a release
carries the label, check 9 reports that component as a counted SKIP, not OK;
costs nothing extra because this Unreleased already forces a base rebuild
(`rootfs/` skill floor).
---
**The rule "use `pi-task`, not `fork`, for a brief that carries a prohibition" was
written in the global `AGENTS.md` and in the pi-extensions skill, and it lost to
the `fork` tool's own description anyway.** Measured on tor-ms22, 2026-09-17: five
of five fork briefs in one session carried "do not"; one returned verbatim quotes
that did not exist in the source, and four had disjoint write boundaries that
fork cannot enforce — re-run as `pi-task`, all four passed their envelope. The
image now ships the rule where the decision is *made*, not where it is read about:
- **`fork-gate.ts`** (pi-extensions `25c1265`): a `tool_call` hook that blocks a
`fork` whose brief contains a prohibition (*do not / never / only …*), a write
boundary (*only touch / read-only / stay within …*) or a clause-initial
file-changing imperative (*Edit …, Commit …, Fix …*). The block reason the model
reads **is** the `task(...)` call to make instead. It matches wording, not
intent, and says so. `PI_FORK_GATE=off` logs instead of blocking.
- **`task.ts`** (same commit): `pi-task` registered as the `task` tool, with the
decision rule in its description and in `promptGuidelines` — which pi appends
to the **system prompt**, the one place compaction cannot remove it from.
Before spending a model run it rejects the two spec errors that make a boundary
violation certain (`write_allowed` not an *exact subset* of `roots` — pi-task
keys deltas by root string; a writable root nested inside a watched-only root,
whose porcelain would change every time) and serialises sibling tasks whose
roots overlap (parallel siblings saw each other's writes as violations,
2026-09-17). A FAIL verdict is a *result*; only the CLI refusing to run is an
error.
- **`pi-global-AGENTS.md`** (pi-toolkit `9c87ee8`): the delegation section now
opens "`task` first, `fork` second" — the one discriminator (what the child
sees), a three-question pre-flight before any `fork(...)`, a copy-paste minimal
call with the roots contract. This file is in the system prompt; the skill is
not, which is why the rule lives here.
- **Skill floor refreshed** to `25c1265` (`check-skill-floor.sh` OK, tree
`9b85a633…`); mirror in skillset `debc8f6`.
Why prose failed, as mechanisms: `fork` is a **tool** — its self-recommending
description ("exploration, implementation, testing, review…") is in the tool
list on every turn and survives compaction; the skill is gone after the first
compaction; `pi-task` was a CLI to be remembered and reached through `bash` with a
hand-written JSON spec. The asymmetry *widens* in exactly the long sessions where
fork is worst, and fork deletes its temp dir on exit, so its failures were found
only by re-verifying the narrative. Third time on this fleet that a rule held in
prose and violated in practice was fixed by moving it into a hook
(`check-secrets`, `check-egl-only`, now delegation).
Evidence, two-sided: `test/fork-gate.test.mjs` pins 15 must-block briefs
(including the real shapes) and 10 must-pass (including *"Write a summary…"*,
*"Report which files were modified…"*, *"Give me an update…"* — verbs a naive
list misfires on); the shipped classifier over the five real briefs from the
motivating session redirects 5/5. Live in `pi -p`: the fork was intercepted
before any child spawned (no `/tmp/pi-fork-*` directory), `task` returned PASS in
4 s / $0.014 with an evidence pointer and audit dir, a `usd=0.000001` budget came
back as a FAIL *result* (`isError=false`), and a nested-root spec was rejected
with no audit dir created.
Nothing to wire in this repo: `PI_EXTENSIONS_REF=main` floats and
`pi-extensions/install.sh` symlinks every `extensions/*.ts` on container start, so
the next build ships both. A container running today has neither —
`~/.pi/agent/extensions/` links into `/opt/pi-extensions`, which is the baked ref.
---
**`[mempalace ext] feed (tick) failed: mempalace remote request 'tools/call'
failed: timed out after 60000ms` is the same event as the `mine timed out after
30000ms` message the 2026-09 toolkit fix addressed, one deadline further down —
and that fix was incomplete.** mempalace-toolkit `817b3a8` (2026-09-18; v1.9.2
baked `dab989b`; the floating `MEMPALACE_TOOLKIT_REF=main` picks it up on the next
build). `MEMPALACE_FEED_MINE_TIMEOUT_MS` had been raised to 300 000 but was only
*raced* against `client.callTool("mempalace_mine")`; `callTool()` had no way to
carry a deadline, so every mine went out under the transport's generic
per-request timeout — `MEMPALACE_MCP_TIMEOUT_MS`, 60 000, the value the "Stall
protection" comment in `Dockerfile.base` documents — which fired first on every
honest 60 s+ mine on the shared single-writer palace. The 300 s was unreachable.
Over HTTP nothing is lost (the mine continues server-side and is idempotent);
over stdio it was worse than noise — that transport **kills the child** on
timeout, so there the mine really was aborted at 60 s.
Fix: `callTool(name, args, { timeoutMs })` on both transports, the feed passes
its own deadline down, plain calls keep 60 s (a *query* that slow is wedged; the
race stays as the liveness guard for a transport with its timeout disabled).
`scripts/test-mcp-call-timeout.sh` cuts `RemoteMcpClient` out of the shipped
file, drives it against a local JSON-RPC server that delays `tools/call`, and
asserts three things — a plain call rejects at the generic deadline, the override
outlives it, the override is itself a deadline: 2 of 6 fail on `dab989b`, 6 of 6
pass on `817b3a8`. `test-owed-withdrawal.sh` (17) and `check-mcp-client-sync.sh`
stay green; sync token `v1` untouched, since nothing in the protocol changed.
Reading for the fleet: on a fixed build that message means a mine exceeded
*five* minutes — look at palace size or a competing writer, not at the timeout.
---
**`scripts/recreate-sanity-check.sh` asserted that `/tmp/sshcm` exists while every
`ssh` in the container was dying `rc=255`, and it was right to — it was checking
the directory the *image* creates, and the breakage was in the directory a
*config* named.** Found on the v1.9.2 first boot on `emb-7kj4vr4g`. A durable
`~/.pi/ssh/config` (hand-written into the `~/.pi` named volume by the previous
session, so it would survive the recreate) declared `ControlPath
/tmp/ssh-cm/%C` — with a hyphen. Nothing in this repo creates that path; the
canonical directory is `/tmp/sshcm`, spelled the same way in four places
(`Dockerfile.base`, `entrypoint-user.sh`, `recreate-sanity-check.sh`,
`smoke-test.sh`). Result:
```
unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory
```
`rc=255`, and the remote command never ran at all — `ControlMaster auto` with an
unusable `ControlPath` fails hard rather than falling back to an unmultiplexed
connection. The old check passed truthfully, about the wrong object. **Two
independent facts about the same subsystem can both be true while the subsystem
is dead; a check that asserts only one of them cannot see their disagreement.**
**New in `recreate-sanity-check.sh`: resolve the ControlPath `ssh` itself would
use, via `ssh -G`, and require its parent to exist and be writable.** `-G`
applies real config precedence — first-obtained-value-wins, the system drop-in,
`Include`, an `-F` override — so it answers "which rule captured this host"
instead of re-implementing the guess. It never opens a connection: measured
0.116 s for 48 hosts.
This also puts a check under a caveat that had been documented in prose in
`Dockerfile.base` ("SSH client defaults") and verified nowhere: a per-host
`ControlPath ~/.ssh/cm/%r@%h:%p` inherited from a **read-only** bind-mounted
`~/.ssh` produces the identical failure, `cannot bind … Read-only file system`.
On the machine where this was built that is **16 of 50 hosts** — `freeipa-1..6`,
`gitea.egl.lan`, `runner-1..3`, `tor-ms22` and more — none of which had ever been
reported by anything.
**The two routes get different severities, deliberately.** Default `ssh`
precedence legitimately lands in the read-only `~/.ssh` on any host whose own
config pins it there, and the supported workaround (`ssh -F
~/.ssh-local/config`, generated every container start by `setup-lan-access.sh`)
already exists — so that is a `warn`. Making it a failure would paint the script
red on every run of every device, and **a check that fires benignly every time is
one you learn to ignore**, which is the same reasoning that keeps `lint-shell.sh`
at `-S error`. The sidecar route is the prescribed one, so there an unusable
directory is a hard `fail`. Host lists are capped at six names plus a count for
the same reason: unreadable output is ignored output.
Teeth proven in both directions, with each sabotage confirmed by `diff` *before*
the result was believed — a vacuous sabotage that silently fails to apply
reports "gate passed" and is worse than no test:
| Sabotage | Class | Result |
|---|---|---|
| `ControlPath` → nonexistent dir | the original hyphen bug | `✗` `rc=1` |
| `ControlPath` → existing but read-only dir | the `Dockerfile.base` caveat | `✗` `rc=1` |
| restored | — | `✓` `rc=0`, sidecar byte-identical |
No config was added to this repo to fix the original problem, because the fix was
to **delete** the offending file: `~/.pi/ssh/config` was a third hand-maintained
copy of what `setup-lan-access.sh` already generates from version control on
every start, with fewer features (no `known_hosts` sidecar, no
`StrictHostKeyChecking accept-new`) and one typo. The host-owned `~/.ssh/config`
is also left alone on purpose — those `~/.ssh/cm` paths are correct *on the host*,
where `~/.ssh` is writable, and the container-side override is the right layer.
---
**The size numbers on the Docker Hub page were the only claim in these docs with
nothing in the repo to check them against, and they had gone 20% wrong across
eight releases.** Every other claim `scripts/check-doc-drift.sh` guards is
anchored to a build file — a pin, an `ARG`, a placeholder — so it cannot rot
without someone editing the thing it describes. Nothing in this repo states the
image size, so `DOCKER_HUB.md`'s `~1.1 GB` simply drifted while the image grew,
and `update-description` POSTed it to Docker Hub every release. It is the first
number a stranger reads about this image.
Corrected against Docker Hub's measured `full_size`, 2026-09-14 after v1.9.2
published (amd64 / arm64, compressed):
| Row | Claimed | Measured | Now says |
|---|---|---|---|
| `:latest` | ~1.1 GB | 1.228 / 1.211 | ~1.23 GB |
| `:latest-studio` | ~1.15 GB | 1.255 / 1.238 | ~1.25 GB |
| `:base-latest`, `:base-<hash>` | ~1.0 GB | 1.167 / 1.151 | ~1.17 GB |
`full_size` is the right field because it tracks the **first manifest entry**
(amd64), not the sum across architectures — measured on v1.9.2:
`full_size=1.228`, `amd64=1.228`, `arm64=1.211`, `sum=2.439`. That matches the
table's per-arch "Size (compressed)" column.
**New check 8 in `scripts/check-doc-drift.sh`: size claims vs Hub's measured
`full_size`,** so this class cannot rot silently again. It fails on drift beyond
tolerance, and **skips loudly** — a new `skip()` helper, counted and named in the
summary — when `curl`/`python3` are missing, the API is unreachable, or
`SKIP_SIZE_CHECK=1`. Skips are deliberately neither `OK` nor a failure: printing
an unverified claim as OK is the habit this file exists to break, while failing
on Docker Hub's uptime would make every release hostage to a third party. Not in
`hooks/pre-push` (that runs `lint-shell.sh` only), so pushes do not hit the
network.
Two bugs were caught while building it, both by writing the expected exit code
down *before* running the check:
- **The tolerance would have missed its own motivating case.** The percentage is
computed against the *measured* size, but the 20% first chosen came from the
claim-relative figure. The real drift was `|1.1 − 1.37| / 1.37 = 19.7%` — it
would have passed. Now 15%, sitting inside a window whose bounds are both
measured: above the largest legitimate skew (a claim describing the published
release while the next tag changes the size — v1.9.1's 1.37 against v1.9.2's
1.23 = 11.4%) and below the rot it exists to catch (19.7%).
- **A `|| true` on the python invocation made the gate fail open.** It printed
`DRIFT` and exited 0 — a gate that reports the defect and passes anyway.
Removed; the outer `|| SIZE_RC=$?` is what satisfies `set -e` without
swallowing the code. Verified two-sided afterwards: a 19% drift exits 1 at the
default tolerance and 0 at `SIZE_TOLERANCE_PCT=25`, so the threshold is doing
the work rather than the ordering.
**Explicitly NOT covered:** `README.md`'s `~3.2 GB` figures are *uncompressed*
on-disk sizes, and the registry exposes compressed sizes only (manifest layer
sizes are compressed; the config blob carries no uncompressed totals). Measuring
them needs a real pull, so they remain unverified — a green check 8 says nothing
about them, and the script says so where a reader will see it.
**`promote-base-latest`'s conditional re-tag is now measured, which closes an
open question from the v1.9.2 release verification.** The job compares digests
and re-tags `base-latest` only if stale, and that short-circuit determines
whether a release watcher may assert *freshness* on the alias or merely
*existence*.
Measured on run 669: the digests differed (`want sha256:8c452575…`,
`have sha256:f34ad201…`), the job logged
`Promoting base-latest -> …:base-b5f2d03baae2`, `crane copy` ran
`17:10:04 → 17:10:08`, and Docker Hub's `base-latest.last_updated` moved to
`17:10`. So **a `crane copy` that runs does bump Hub's timestamp** — a manifest
re-tag is a tag write.
The identical-digest case remains *unproven by construction*: when `base-latest`
already resolves to the new `base-<hash>` the job prints
`base-latest already current; nothing to promote.` and never copies, so the
timestamp legitimately stays put. A freshness assertion would then report a
correct release as stale. The rule is therefore conditional on
`base-decide`'s `need_build` — strengthen it to fresh-required only when
`Dockerfile.base`, `rootfs/` or a folded `*_REF` has moved, and leave it as an
existence check for the variant-only and docs-only releases that cache-hit the
base. Operational guidance lives with the tooling that consumes it
(`ci-release-watcher` skill, skillset `c8034af`); recorded here because it is a
property of *this* pipeline.
Worth restating alongside it, because v1.9.2 proved it the useful way: tag
freshness is not the authoritative evidence that a base rebuild baked what you
expected. The `se.jordbo.pi-devbox.*-ref` labels are — readable straight from the
registry with no `docker` or `crane` (token → manifest index → per-arch manifest
→ config blob), and worth validating against the *previous* tag first, since a
reader that cannot show the old SHA cannot prove the new one.
---
## v1.9.2 — 2026-09-14 ## v1.9.2 — 2026-09-14
**The v1.9.1 residual is attributed and fixed: it was mostly npm's own download **The v1.9.1 residual is attributed and fixed: it was mostly npm's own download
+4 -4
View File
@@ -8,12 +8,12 @@ A self-contained Docker container for the [pi coding-agent](https://github.com/e
| Tag | Architectures | Size (compressed) | What you get | | Tag | Architectures | Size (compressed) | What you get |
|---|---|---|---| |---|---|---|---|
| `joakimp/pi-devbox:latest` | amd64, arm64 | ~1.1 GB | Self-contained: base + pi `{{PI_VERSION}}` + companions | | `joakimp/pi-devbox:latest` | amd64, arm64 | ~1.23 GB | Self-contained: base + pi `{{PI_VERSION}}` + companions |
| `joakimp/pi-devbox:vX.Y.Z` | amd64, arm64 | same | Pinned semver release | | `joakimp/pi-devbox:vX.Y.Z` | amd64, arm64 | same | Pinned semver release |
| `joakimp/pi-devbox:latest-studio` | amd64, arm64 | ~1.15 GB | `latest` + [pi-studio](https://github.com/omaclaren/pi-studio): browser prompt editor, KaTeX/Mermaid preview, tmux-backed literate REPLs | | `joakimp/pi-devbox:latest-studio` | amd64, arm64 | ~1.25 GB | `latest` + [pi-studio](https://github.com/omaclaren/pi-studio): browser prompt editor, KaTeX/Mermaid preview, tmux-backed literate REPLs |
| `joakimp/pi-devbox:vX.Y.Z-studio` | amd64, arm64 | same | Pinned semver studio release | | `joakimp/pi-devbox:vX.Y.Z-studio` | amd64, arm64 | same | Pinned semver studio release |
| `joakimp/pi-devbox:base-latest` | amd64, arm64 | ~1.0 GB | Base layer alias (internal building block; pull `:latest` instead) | | `joakimp/pi-devbox:base-latest` | amd64, arm64 | ~1.17 GB | Base layer alias (internal building block; pull `:latest` instead) |
| `joakimp/pi-devbox:base-<hash>` | amd64, arm64 | ~1.0 GB | Content-addressed base; immutable. Stable parent for variant rebuilds. | | `joakimp/pi-devbox:base-<hash>` | amd64, arm64 | ~1.17 GB | Content-addressed base; immutable. Stable parent for variant rebuilds. |
> **pi-studio (`-studio` tags):** launch with `/studio --no-browser --port 8765` inside a pi session. The server binds `127.0.0.1` **inside the container**, so reach it via host networking or a loopback bridge (and `ssh -L` for a remote host; mosh needs a parallel `ssh -L`). Full recipe: [README → Using pi-studio](https://gitea.jordbo.se/joakimp/pi-devbox#using-pi-studio--studio-variant). > **pi-studio (`-studio` tags):** launch with `/studio --no-browser --port 8765` inside a pi session. The server binds `127.0.0.1` **inside the container**, so reach it via host networking or a loopback bridge (and `ssh -L` for a remote host; mosh needs a parallel `ssh -L`). Full recipe: [README → Using pi-studio](https://gitea.jordbo.se/joakimp/pi-devbox#using-pi-studio--studio-variant).
+88 -11
View File
@@ -14,7 +14,7 @@
# content-addressed over this file, so any byte change invalidates the # content-addressed over this file, so any byte change invalidates the
# cache. Recommended cadence: once per release for security updates. # cache. Recommended cadence: once per release for security updates.
# #
# BASE_REBUILD_DATE: 2026-07-13 (Unreleased — agent-browser CLI + Playwright Chromium for headless browser automation; prior: typst PDF engine + xz-utils + pandoc typst-template default-font patch) # BASE_REBUILD_DATE: 2026-09-22 (v1.9.4 — mempalace 3.9.0 -> 3.10.0 with ENV MEMPALACE_CONFIG_DIR pinning the layout, mempalace-toolkit 2167a1b explicit event_list order; previous marker 2026-09-19 / v1.9.3)
# #
# ── Lineage note ───────────────────────────────────────────────────── # ── Lineage note ─────────────────────────────────────────────────────
# Adapted from opencode-devbox/Dockerfile.base (commit before v1.16.2). # Adapted from opencode-devbox/Dockerfile.base (commit before v1.16.2).
@@ -452,7 +452,10 @@ RUN ARCH=$(case "${TARGETARCH}" in amd64) echo "x86_64" ;; arm64) echo "aarch64"
# uninterruptibly. A stall-kill is no longer a permanent latch either: the # uninterruptibly. A stall-kill is no longer a permanent latch either: the
# next tool call respawns the server with capped exponential backoff (the # next tool call respawns the server with capped exponential backoff (the
# budget resets on any successful response). Tunables: # budget resets on any successful response). Tunables:
# MEMPALACE_MCP_TIMEOUT_MS (default 60000), MEMPALACE_MCP_INIT_TIMEOUT_MS # MEMPALACE_MCP_TIMEOUT_MS (default 60000; the feed's `mempalace_mine` carries
# its own longer MEMPALACE_FEED_MINE_TIMEOUT_MS, default 300000, since toolkit
# 817b3a8 — before that the 60 s deadline cut every honest mine off),
# MEMPALACE_MCP_INIT_TIMEOUT_MS
# (default 300000 — generous so a genuine first cold-open isn't killed), # (default 300000 — generous so a genuine first cold-open isn't killed),
# MEMPALACE_MCP_MAX_RESPAWNS (default 2; 0 disables self-heal), # MEMPALACE_MCP_MAX_RESPAWNS (default 2; 0 disables self-heal),
# MEMPALACE_MCP_RESPAWN_BACKOFF_MS (default 1000); timeouts of 0 disable. # MEMPALACE_MCP_RESPAWN_BACKOFF_MS (default 1000); timeouts of 0 disable.
@@ -529,14 +532,22 @@ ARG INSTALL_MEMPALACE=true
# the part that should stay manual. # the part that should stay manual.
# #
# Deployment sequencing note for whoever ships this bump: synlig (the shared # Deployment sequencing note for whoever ships this bump: synlig (the shared
# central palace host) serves mempalace 3.8.0 SERVER-SIDE via # central palace host) serves mempalace SERVER-SIDE as a `uv tool` install run
# docker-compose.mempalace.yml, which reuses this same devbox image. (Measured # by the systemd unit `mempalace-serve.service` (`python -m mempalace.mcp_server
# 2026-09-06 over ssh: synlig's UV_TOOL_DIR mempalace entry last changed # --transport http`), NOT via docker-compose.mempalace.yml — that compose file
# 2026-08-25 15:33 — this comment previously said 3.7.1, which was stale.) # exists in this repo but is not what runs there. (Measured 2026-09-22 over
# Bumping this ARG changes only the CLIENT version baked into pi-devbox # ssh: `uv tool list` -> mempalace v3.9.0, python 3.12.13, chromadb 1.5.9;
# images: it introduces client/server skew until synlig's compose stack is # `docker ps` matched no palace container. This comment previously said the
# separately rebuilt/redeployed with the new pin. Not something to code around # compose stack served 3.8.0, which was stale on both counts.) Bumping this ARG
# here — just sequence the redeploy. # changes only the CLIENT version baked into pi-devbox images: it introduces
# client/server skew until synlig's tool is upgraded (`uv tool upgrade
# mempalace` + restart the unit). Not something to code around here — just
# sequence the upgrade. And note which side OWNS what: MCP tool semantics
# (event_list ordering, kg_timeline pagination, search result fields) come
# from the SERVER the extension talks to over MEMPALACE_REMOTE_URL, so they
# change when synlig upgrades; only the local CLI (`mempalace init` at first
# run, the mempalace-pi-session feeder) and the on-disk layout under
# ~/.mempalace change when THIS pin does.
# #
# v1.8.13: 3.8.0 -> 3.9.0. Audited: no Breaking/Removed changelog headings. # v1.8.13: 3.8.0 -> 3.9.0. Audited: no Breaking/Removed changelog headings.
# Adopted mainly for #2281 (`mempalace_mine` accepts a single conversation # Adopted mainly for #2281 (`mempalace_mine` accepts a single conversation
@@ -548,7 +559,61 @@ ARG INSTALL_MEMPALACE=true
# (release awareness, `task create`/`task launch` MCP tools) are SERVER-side, # (release awareness, `task create`/`task launch` MCP tools) are SERVER-side,
# so they stay dark until synlig is redeployed — a client bump alone cannot # so they stay dark until synlig is redeployed — a client bump alone cannot
# light them up. # light them up.
ARG MEMPALACE_VERSION=3.9.0 #
# v1.9.4: 3.9.0 -> 3.10.0 (PyPI 2026-09-15). Deferred at v1.9.3 for two
# "Upgrade notes" items; both re-measured against the 3.10.0 wheel, one needed
# an adaptation:
# - "New installs keep config and palace under ~/.config/mempalace". The
# resolution order is $MEMPALACE_CONFIG_DIR, then ~/.mempalace IF it holds
# config.json / people_map.json / palace/chroma.sqlite3, then XDG. An
# EMPTY ~/.mempalace does not count — and an empty ~/.mempalace is exactly
# what a freshly mounted devbox-palace volume (or entrypoint.sh's mkdir on
# a volume-less container) looks like at first boot. Measured with a fresh
# $HOME: `mempalace init` wrote to ~/.config/mempalace, outside the
# persisted path, and entrypoint-user.sh's first-run test
# `[ ! -d ~/.mempalace/palace ]` would stay true on every start. With
# MEMPALACE_CONFIG_DIR set, everything landed in ~/.mempalace. Hence the
# ENV MEMPALACE_CONFIG_DIR below (in the non-root-user section, where
# ${USER_NAME} is in scope): first in the resolution order, so the
# heuristic never runs and the image's layout contract no longer depends
# on it. Existing volumes were safe either way (config.json is a legacy
# marker); the ENV is for first boots. palace_path still defaults to
# <config_dir>/palace and MEMPALACE_PALACE_PATH is still honoured (config.py
# :927), so scripts/smoke-test.sh's stage-path test keeps its meaning.
# - "MCP event listing returns the newest events first when no cursor is
# given". SERVER-side (see above), so it lands when synlig upgrades, not
# here. mempalace-toolkit 2167a1b made every cursor-less event_list call
# in the pi extension say `order: "desc"` explicitly, so the mailbox reads
# the same window against either server version.
# Also in the notes, neither reaching this image: `mempalace rules` dropped
# `--agent` (no caller in pi-devbox, mempalace-toolkit, skillset or myconfigs);
# `get_collection()` refuses unknown collection names (library callers only).
# MCP tool-schema review, as always: no tool removed or renamed; additive
# fields on search results (filed_at / content_date provenance), `limit` /
# `offset` on kg_timeline, `last_modified` on drawers. Skew while synlig stays
# on 3.9.0 is narrower than it looks: the pi extension speaks HTTP to the hub
# (no local mempalace-mcp is spawned), and the feeder in remote mode stages
# locally in python, rsyncs, and calls the hub's own `mempalace_mine` tool.
# Measured 2026-09-22 with a PATH shim in front of `mempalace`: a 47-session
# `mempalace-pi-session --dry-run` made ZERO local CLI calls (the shim's
# positive control logged one). So 3.10.0's new CLI write-routing policy never
# runs against the hub from this image; the client pin touches first-run
# `mempalace init` and the on-disk layout, nothing else in remote mode.
ARG MEMPALACE_VERSION=3.10.0
# Recorded as a label HERE, not in Dockerfile.variant, for three reasons: the
# value lives next to the ARG that defines it (a second copy in the variant
# would be one more pin able to drift, which is the class check-doc-drift.sh
# exists to catch); labels are inherited by every image built FROM this one, so
# both variants carry it with no build-arg to plumb through four call sites;
# and inheritance means the label states the pin of the base the variant
# ACTUALLY built on — which is the question when base-decide cache-hits an
# older base. Like every se.jordbo.pi-devbox.* label this records INTENT; the
# ground truth is /etc/pi-devbox/build-manifest.json's mempalace_version, read
# from the installed binary, and scripts/smoke-test.sh asserts the two agree.
# check-doc-drift.sh check 9 reads this off the last published image so that a
# pin bump must be named in the CHANGELOG — until this label ships, that
# component reports SKIP (label absent on the published release), not OK.
LABEL se.jordbo.pi-devbox.mempalace-version="${MEMPALACE_VERSION}"
ENV UV_TOOL_DIR=/opt/uv-tools ENV UV_TOOL_DIR=/opt/uv-tools
ENV UV_TOOL_BIN_DIR=/usr/local/bin ENV UV_TOOL_BIN_DIR=/usr/local/bin
RUN if [ "${INSTALL_MEMPALACE}" = "true" ]; then \ RUN if [ "${INSTALL_MEMPALACE}" = "true" ]; then \
@@ -815,6 +880,18 @@ print('chromadb embedding model warmed: all-MiniLM-L6-v2')" && \
ENV NPM_CONFIG_PREFIX=/home/${USER_NAME}/.pi/npm-global ENV NPM_CONFIG_PREFIX=/home/${USER_NAME}/.pi/npm-global
ENV PATH="/home/${USER_NAME}/.pi/npm-global/bin:${PATH}" ENV PATH="/home/${USER_NAME}/.pi/npm-global/bin:${PATH}"
# ── MemPalace config/palace root: pin it, do not let a heuristic pick it ──
# mempalace >= 3.10.0 resolves its config dir as $MEMPALACE_CONFIG_DIR, then
# ~/.mempalace ONLY if it already holds a config/palace, else ~/.config/mempalace
# (XDG). An empty ~/.mempalace — a fresh devbox-palace volume, or entrypoint.sh's
# mkdir on a volume-less container — fails that test, so a first boot would put
# the palace outside the persisted path and re-run first-run init forever. This
# ENV is first in the order, so the layout is what entrypoint.sh (mkdir),
# entrypoint-user.sh (first-run test), the feeder's <palace-root>/pi-stage and
# scripts/recreate-sanity-check.sh all already assume. Rationale and the
# measurement live with ARG MEMPALACE_VERSION above; keep the two in step.
ENV MEMPALACE_CONFIG_DIR=/home/${USER_NAME}/.mempalace
# ── Shell defaults (bash history, aliases, readline) ───────────────── # ── Shell defaults (bash history, aliases, readline) ─────────────────
RUN mkdir -p /etc/skel-devbox RUN mkdir -p /etc/skel-devbox
COPY rootfs/home/developer/.bash_aliases /etc/skel-devbox/.bash_aliases COPY rootfs/home/developer/.bash_aliases /etc/skel-devbox/.bash_aliases
+34 -2
View File
@@ -113,6 +113,27 @@ ARG USER_NAME=developer
# signature is ~5s of sustained CPU. Two-sided check: the atelier sidebar # signature is ~5s of sustained CPU. Two-sided check: the atelier sidebar
# painted ACTIVITY+WORKSPACE identically to the 0.84.4 control, so the test # painted ACTIVITY+WORKSPACE identically to the 0.84.4 control, so the test
# could distinguish "loaded" from "silently absent". # could distinguish "loaded" from "silently absent".
#
# v1.9.4: HELD at 0.85.1 while 0.86.1 and 0.87.0 exist upstream. 0.87.0's
# changelog: "Removed the inherited `shouldStopAfterTurn` agent option. Use
# `finishTurn` and return `{ action: "end" }` instead" (no issue number on
# that line); 0.86.0 moved provider stream inputs to `TranscriptContext`, with
# system prompts read from `context.messages`. pi-observational-memory 3.1.4
# (the `master` ref baked below) still uses both the old option and
# `AgentContext.systemPrompt` in its observer, reflector and dropper workers —
# measured 2026-09-22 in src/agents/*/agent.ts, both the baked 3.1.3 tree and
# upstream master 3.1.4: `shouldStopAfterTurn` 1 per worker (3), `finishTurn`
# 0, `systemPrompt` 1 per worker. Its peerDependencies are `*`,
# so nothing at install time would refuse; it would break at runtime (turn
# caps ignored, workers losing their specialised prompts). Upstream tracks it
# as pi-observational-memory #82 with fix PR #83 (opened 2026-09-21, mergeable,
# not merged at this writing). Bump pi and pi-obsmem TOGETHER once #83 has
# shipped in a release. The other extensions were checked against the 0.86.0
# and 0.87.0 breaking lists and are clean: ssh-controlmaster's `user_bash`
# handler already returns `undefined | { operations }` (0.86.0 fail-closed
# contract) and its registerTool calls spread the built-in tools so they
# carry parameter schemas (#9300). 0.86.1 as an intermediate is untested and
# not worth the pty matrix for a stop that #83 will make moot.
ARG PI_VERSION=0.85.1 ARG PI_VERSION=0.85.1
ARG PI_TOOLKIT_REF=main ARG PI_TOOLKIT_REF=main
ARG PI_EXTENSIONS_REF=main ARG PI_EXTENSIONS_REF=main
@@ -170,9 +191,20 @@ ARG PI_ATELIER_REPO=https://github.com/michaelmjhhhh/pi-atelier.git
# old and new pin. Included because it was already exercised: the pty matrix # old and new pin. Included because it was already exercised: the pty matrix
# for PI_VERSION above ran atelier v0.10.1 against pi 0.85.1 and painted the # for PI_VERSION above ran atelier v0.10.1 against pi 0.85.1 and painted the
# sidebar identically to v0.10.0. # sidebar identically to v0.10.0.
ARG PI_ATELIER_REF=v0.10.1 #
# v1.9.4: v0.10.1 -> v0.10.3 (both v0.10.2 and v0.10.3 released 2026-09-22).
# Fixes only per the release notes: sidebar text/borders preserved beside
# inline images (#53), transcript images hidden while capturing overlays are
# open, Workspace Pulse skips redundant HEAD/diff when nothing tracked changed
# (#61), sidebar height from row counts (#59), git/usage scans suspended while
# disabled, Display Revert + Undo ordering. Checked before bumping: package.json
# at v0.10.3 still declares zero runtime dependencies and no build script (so
# the no-`npm install` reasoning above holds) and peerDependencies are still
# pi >=0.84.0, so it still spans the pinned 0.85.1. 13 commits v0.10.1..v0.10.3,
# all under src/ tests/ docs/ scripts/ plus metadata; no entry-point move.
ARG PI_ATELIER_REF=v0.10.3
# Human-readable tag PI_ATELIER_REF was resolved from; recorded as a label. # Human-readable tag PI_ATELIER_REF was resolved from; recorded as a label.
ARG PI_ATELIER_VERSION=v0.10.1 ARG PI_ATELIER_VERSION=v0.10.3
RUN set -e && \ RUN set -e && \
# git_fetch_ref: clone-equivalent helper that accepts EITHER a branch name # git_fetch_ref: clone-equivalent helper that accepts EITHER a branch name
+13 -5
View File
@@ -572,7 +572,13 @@ ChromaDB ONNX embedding model so first-time semantic search is
instant. instant.
The palace data lives at `~/.mempalace/palace` on the host The palace data lives at `~/.mempalace/palace` on the host
(bind-mounted into the container). This means: (bind-mounted into the container). The image pins that layout with
`ENV MEMPALACE_CONFIG_DIR=/home/developer/.mempalace` (`Dockerfile.base`):
mempalace ≥ 3.10.0 would otherwise treat an *empty* `~/.mempalace` — a freshly
mounted volume at first boot — as "no install here" and put a new palace under
`~/.config/mempalace`, outside anything the compose files persist. With the
variable set, first in mempalace's resolution order, the location is a contract
rather than a heuristic. This means:
- A pi running on the host and a pi running inside this container see - A pi running on the host and a pi running inside this container see
the same palace. the same palace.
@@ -901,8 +907,10 @@ docker inspect --format '{{json .Config.Labels}}' joakimp/pi-devbox:latest | jq
``` ```
`org.opencontainers.image.{version,revision,created}` plus `org.opencontainers.image.{version,revision,created}` plus
`se.jordbo.pi-devbox.*-ref` record the intended pi version and companion `se.jordbo.pi-devbox.*-ref` and `se.jordbo.pi-devbox.*-version` record the
refs. The on-disk `/etc/pi-devbox/build-manifest.json` records **ground intended pi and mempalace versions and companion refs (`mempalace-version` is
set in `Dockerfile.base` and inherited, so it names the pin of the base the
image actually built on). The on-disk `/etc/pi-devbox/build-manifest.json` records **ground
truth** — the actual checked-out commit of each `/opt` clone, the live truth** — the actual checked-out commit of each `/opt` clone, the live
`pi --version`, and (from v1.8.6) the live `mempalace --version` of the `pi --version`, and (from v1.8.6) the live `mempalace --version` of the
installed palace core — so a tag is reconstructable after CI logs rotate: installed palace core — so a tag is reconstructable after CI logs rotate:
@@ -1138,8 +1146,8 @@ resolved to `latest` at build time:
| Component | Pin | Where | | Component | Pin | Where |
|---|---|---| |---|---|---|
| pi | `0.85.1` | `ARG PI_VERSION` — `Dockerfile.variant` | | pi | `0.85.1` | `ARG PI_VERSION` — `Dockerfile.variant` |
| pi-atelier | `v0.10.1` | `ARG PI_ATELIER_REF` — `Dockerfile.variant` | | pi-atelier | `v0.10.3` | `ARG PI_ATELIER_REF` — `Dockerfile.variant` |
| mempalace | `3.9.0` | `ARG MEMPALACE_VERSION` — `Dockerfile.base` | | mempalace | `3.10.0` | `ARG MEMPALACE_VERSION` — `Dockerfile.base` |
The objective is **not** to freeze versions. Bumping is routine — usually one The objective is **not** to freeze versions. Bumping is routine — usually one
line plus a changelog note. The objective is that adopting a new upstream line plus a changelog note. The objective is that adopting a new upstream
+8 -1
View File
@@ -100,7 +100,14 @@ fi
# existing data. `--yes` auto-accepts detected entities so the init is # existing data. `--yes` auto-accepts detected entities so the init is
# non-interactive. # non-interactive.
if command -v mempalace &>/dev/null && [ -d /workspace ]; then if command -v mempalace &>/dev/null && [ -d /workspace ]; then
PALACE_DIR="${HOME}/.mempalace" # Read the root from the same variable mempalace itself reads (set as an
# image ENV in Dockerfile.base since mempalace 3.10.0 started resolving
# ~/.config/mempalace for an EMPTY ~/.mempalace). The fallback keeps the
# historical location for anyone running this script with the ENV unset;
# the point of naming the variable here is that this test and mempalace's
# own resolution can no longer disagree about where the palace lives — a
# disagreement that would make this branch fire on every start.
PALACE_DIR="${MEMPALACE_CONFIG_DIR:-${HOME}/.mempalace}"
if [ ! -d "$PALACE_DIR/palace" ]; then if [ ! -d "$PALACE_DIR/palace" ]; then
echo "Initializing MemPalace for workspace (non-interactive)..." echo "Initializing MemPalace for workspace (non-interactive)..."
# </dev/null: mempalace init has an interactive "Mine this directory # </dev/null: mempalace init has an interactive "Mine this directory
@@ -1,17 +1,17 @@
--- ---
name: pi-extensions name: pi-extensions
description: >- description: >-
Use the pi extensions (pi-fork, pi-observational-memory, ssh-controlmaster) effectively in the pi coding agent harness. Load this skill only when running inside pi (detection - `fork` and `recall` are present in your tool list, or `pi --ssh` was used to start the session). pi-fork dispatches focused subtasks to forked agents at fast/balanced/deep effort tiers; pi-observational-memory compacts long sessions into recallable observations + reflections; ssh-controlmaster rewires pi's read/write/edit/bash tools to execute on a remote host over a multiplexed SSH connection. Also covers the context ladder L0-L4 and when to reach for the separate `pi-task` CLI instead of `fork` - isolated child, immutable spec, machine-checked envelope, write-boundary diff. This skill covers tier selection, task design, boundary discipline, when to use recall, and remote-pi mechanics. Use the pi extensions (pi-fork, pi-observational-memory, ssh-controlmaster) effectively in the pi coding agent harness. Load this skill only when running inside pi (detection - `fork` and `recall` are present in your tool list, or `pi --ssh` was used to start the session). pi-fork dispatches focused subtasks to forked agents at fast/balanced/deep effort tiers; pi-observational-memory compacts long sessions into recallable observations + reflections; ssh-controlmaster rewires pi's read/write/edit/bash tools to execute on a remote host over a multiplexed SSH connection. Also covers the context ladder L0-L4 and the `task` tool (pi-task: isolated child, immutable spec, machine-checked envelope, write-boundary diff) that is the DEFAULT for delegated work, with `fork` reserved for read-only exploration and parallel opinions, plus the `fork-gate` hook that enforces the split. This skill covers rung and tier selection, task design, boundary discipline, when to use recall, and remote-pi mechanics.
--- ---
# Pi Extensions: pi-fork, pi-observational-memory, ssh-controlmaster # Pi Extensions: pi-fork + task/fork-gate, pi-observational-memory, ssh-controlmaster
## When to Load This Skill ## When to Load This Skill
Load only when **both** of these are true: Load only when **both** of these are true:
1. You are running inside the **pi coding agent harness** (not Claude Code, not opencode, not any other harness). 1. You are running inside the **pi coding agent harness** (not Claude Code, not opencode, not any other harness).
2. The `fork` and/or `recall` tools appear in your available tool list, **or** the session was started with `pi --ssh ...`. 2. The `fork`, `task` and/or `recall` tools appear in your available tool list, **or** the session was started with `pi --ssh ...`.
If you do not see those tools, this skill does not apply — skip it. Other harnesses do not have these extensions and the patterns below will not work there. If you do not see those tools, this skill does not apply — skip it. Other harnesses do not have these extensions and the patterns below will not work there.
@@ -71,7 +71,78 @@ ssh-controlmaster is orthogonal but composes cleanly: when pi is operating remot
--- ---
## Part 1: pi-fork ## Part 1: delegating work — `task` and `fork`
### Decide the rung BEFORE the brief (read this first)
Two tools run a child agent. They differ in one thing, and it decides the
quality of what comes back: **what the child sees.**
| tool | child sees | rung | gives you | use for |
|---|---|---|---|---|
| **`task`** (pi-extensions `task.ts`, wraps `pi-task`) | **only your spec** — goal, named files, curated facts | L0–L2 | immutable spec, PASS/FAIL envelope, per-root boundary diff, audit dir, budgets | **any delegated work that changes files or must obey a rule** — the default |
| **`fork`** (pi-fork) | **your entire branch**, brief appended last | L4 | prose report, effort tiers, parallel dispatch from one message | read-only exploration that needs this conversation; N independent opinions |
**Pre-flight before any `fork(...)` — one *yes* makes it a `task`:**
1. Does the brief say *do not / only / never / must not*?
2. Will the child write, edit, commit or push anything?
3. Do I want a PASS/FAIL I can check, rather than prose?
Why the text alone did not work (and why this is now enforced): the rule above
lived in this skill and in the global AGENTS.md for months and was still
violated by agents that had just read it — five fork briefs in one session on
2026-09-17, all carrying "do not", one of which returned confident verbatim
quotes that did not exist. `fork` is a *tool*: its self-recommending
description ("implementation, testing, review…") is in the tool list every turn
and survives compaction; this skill is gone after the first compaction, and
`pi-task` was a CLI to be remembered. Two structural fixes shipped 2026-09-19:
- **`task` is a tool** (`extensions/task.ts`), so both rungs sit in the tool
list with the decision rule in their descriptions, and `promptGuidelines`
puts the rule in the system prompt where compaction cannot remove it. It
rejects, before spending a model run, the two spec errors that make a
violation certain (see "roots" below) and serialises overlapping tasks.
- **`fork-gate`** (`extensions/fork-gate.ts`) is a `tool_call` hook that BLOCKS
a fork whose brief contains a prohibition, a write boundary or a clause-initial
file-changing imperative, and returns the `task(...)` call to make instead.
It matches wording, not intent: a genuinely read-only exploration brief that
trips it is rephrased, and one that cannot be rephrased without its
prohibition needed `task` all along. `PI_FORK_GATE=off` logs instead of
blocking; `/ext` disables it entirely.
Minimal call:
```
task(id="slug", goal="…verbatim; the child has NO other context…",
deliverable="…exact shape wanted…", effort="fast|balanced|deep",
read_only=false,
roots=["/abs/repo/docs", "/abs/repo/src"], # WATCHED, each diffed alone
write_allowed=["/abs/repo/docs"], # exact subset of roots
facts=["verified fact"], files=["/abs/path/to/read"])
```
**Roots — the two errors the tool refuses up front.** Every root is diffed on
its own and a delta is allowed only if that *exact root string* is in
`write_allowed`. So (a) `write_allowed` must be a subset of `roots`, not a
subdirectory of one, and (b) a writable root must not lie inside a watched-only
root — the parent's porcelain would change and register a violation every time.
List the writable part as its own root and leave the enclosing repo out. This is
the shape the 2026-09-17 migration tasks used (sibling roots, `write_allowed`
naming four of them) and it passed cleanly.
**Overlap.** Sibling `task` calls whose roots overlap would see each other's
writes as violations; the tool runs them one after another automatically. Do
not rely on that for ordering *semantics* — if B needs A's output, call B after
A returns.
**What isolation does not fix.** L0 removes the *narrative* failures (parent
voice, invented continuity, ignored prohibitions). It does not remove
confabulation: an under-specified spec still gets a confident deliverable. The
report prints the evidence pointers under a "SPOT-CHECK THESE" heading for a
reason.
Everything below about tiers, brief design and boundary discipline applies to
**both** tools — a `task` spec is a brief too.
### Effort tier mapping ### Effort tier mapping
@@ -85,15 +156,15 @@ Configured in `~/.pi/agent/settings.json` under `pi-fork.effortProfiles`. The co
**Rule of thumb:** start at `balanced` unless you have a specific reason to go up or down. Going too cheap on a deep task wastes a fork; going too expensive on a mechanical task is just slow. **Rule of thumb:** start at `balanced` unless you have a specific reason to go up or down. Going too cheap on a deep task wastes a fork; going too expensive on a mechanical task is just slow.
### When to fork vs. do it yourself ### When to delegate vs. do it yourself
Fork when **any** of: Having chosen the rung above, delegate (either tool) when **any** of:
- The task requires reading many files whose contents you don't need to keep in your main context afterwards (the fork returns a dense summary; raw file contents stay in the fork's context and are discarded). - The task requires reading many files whose contents you don't need to keep in your main context afterwards (the fork returns a dense summary; raw file contents stay in the fork's context and are discarded).
- You want to run multiple analyses in **parallel** (especially: comparing N options, where independent reasoning is itself a signal — see "parallel forks" below). - You want to run multiple analyses in **parallel** (especially: comparing N options, where independent reasoning is itself a signal — see "parallel forks" below).
- The task is well-scoped enough to specify completely up front and well-bounded enough that returning a dense report is more useful than continuing the dialogue. - The task is well-scoped enough to specify completely up front and well-bounded enough that returning a dense report is more useful than continuing the dialogue.
- You are about to do something that would burn a lot of tokens on tool calls (long file reads, many bash invocations) whose output you will mostly discard. - You are about to do something that would burn a lot of tokens on tool calls (long file reads, many bash invocations) whose output you will mostly discard.
Don't fork when: Don't delegate when:
- The work fits in your current context budget without crowding out what comes next. - The work fits in your current context budget without crowding out what comes next.
- The task is exploratory and you'll need to iterate based on what you find (forking turns iteration into round-trips with full task-spec rewrites). - The task is exploratory and you'll need to iterate based on what you find (forking turns iteration into round-trips with full task-spec rewrites).
- You need to make decisions during the work that depend on context only the main thread has. - You need to make decisions during the work that depend on context only the main thread has.
@@ -175,11 +246,12 @@ sits at one extreme of it. Five rungs:
| **L3** | a **truncated tail** of the parent branch | *nothing implements this* — would need a new spec key plus `--session <trimmed snapshot>` | **no** | | **L3** | a **truncated tail** of the parent branch | *nothing implements this* — would need a new spec key plus `--session <trimmed snapshot>` | **no** |
| **L4** | the **entire** parent branch | `fork(task=…)` — `getHeader()+getBranch()`, no offset or limit anywhere in the call chain | yes | | **L4** | the **entire** parent branch | `fork(task=…)` — `getHeader()+getBranch()`, no offset or limit anywhere in the call chain | yes |
**`pi-task` is a CLI, not an extension — it will never appear in your tool list.** **`pi-task` is a CLI (`/opt/pi-toolkit/bin/pi-task`, `schema` prints the spec
Invoke it with `bash`: `/opt/pi-toolkit/bin/pi-task run <spec.json>` (source at fields); the `task` tool from pi-extensions wraps it** so it appears in your tool
`/workspace/pi-toolkit/bin/pi-task`, `schema` subcommand prints the spec fields). list next to `fork`. If the tool is absent, invoke the CLI with `bash`:
It reads an immutable JSON spec, and "inherit the session" is not expressible in `/opt/pi-toolkit/bin/pi-task run <spec.json>`. Either way it reads an immutable
that schema — the isolation is structural, not a request. JSON spec, and "inherit the session" is not expressible in that schema — the
isolation is structural, not a request.
**Choose the lowest rung that can do the job:** **Choose the lowest rung that can do the job:**
@@ -292,7 +364,17 @@ When entries conflict, **the most recent observation reflects the latest known s
## Quick Reference ## Quick Reference
``` ```
task(id, goal, deliverable, effort, read_only, roots, write_allowed, facts, files, commands, wall_s, usd)
- L0-L2: isolated child sees ONLY the spec — DEFAULT for work that writes or has rules
- roots[] = WATCHED (each diffed alone); write_allowed[] = exact subset of roots,
never nested inside a watched-only root (the tool rejects both errors up front)
- envelope must parse or the run FAILED; spot-check evidence pointers
- overlapping-root tasks are serialised; audit: ~/.pi/agent/pi-task/<stamp>-<id>/result.json
- CLI fallback: bash /opt/pi-toolkit/bin/pi-task run <spec.json> (schema | selftest | run --dry-run)
fork(task=..., effort=fast|balanced|deep) # L4: child inherits your WHOLE branch fork(task=..., effort=fast|balanced|deep) # L4: child inherits your WHOLE branch
- ONLY for read-only exploration needing this conversation, or N parallel opinions
- fork-gate BLOCKS briefs with do-not/only/never, write boundaries, or "edit/commit/fix …"
- state decision authority explicitly - state decision authority explicitly
- pass verified context up front - pass verified context up front
- specify deliverable shape - specify deliverable shape
@@ -302,13 +384,6 @@ fork(task=..., effort=fast|balanced|deep) # L4: child inherits your WHOLE b
- write-capable? demand "What I did NOT do", then verify from git/fs, not the report - write-capable? demand "What I did NOT do", then verify from git/fs, not the report
- prohibition in the brief => not a `fast` task - prohibition in the brief => not a `fast` task
bash: /opt/pi-toolkit/bin/pi-task run <spec> # L0-L2: isolated child, NOT a tool
- schema | selftest | run [--dry-run]
- context.facts (pasted) / .files (names only) / .commands
- roots[] = WATCHED, write_allowed[] = CHANGEABLE subset
- envelope must parse or the run FAILED
- audit + cost: ~/.pi/agent/pi-task/<stamp>-<id>/result.json
recall(id=<12-char-hex>) recall(id=<12-char-hex>)
- only when stakes justify the cost - only when stakes justify the cost
- id must already be visible in your context - id must already be visible in your context
@@ -317,8 +392,9 @@ recall(id=<12-char-hex>)
``` ```
~/.pi/agent/settings.json ~/.pi/agent/settings.json
pi-fork.effortProfiles — model + thinking-depth per tier pi-fork.effortProfiles — model + thinking-depth per tier (used by BOTH fork and task)
pi-fork.defaultEffort — usually "balanced" pi-fork.defaultEffort — usually "balanced"
env PI_FORK_GATE=off — fork-gate logs instead of blocking (default: block)
observational-memory.* — token thresholds, model, agentMaxTurns observational-memory.* — token thresholds, model, agentMaxTurns
observational-memory.debugLog: true — opt-in NDJSON telemetry at observational-memory.debugLog: true — opt-in NDJSON telemetry at
~/.pi/agent/observational-memory/debug/<session>.ndjson (off by default) ~/.pi/agent/observational-memory/debug/<session>.ndjson (off by default)
+461 -9
View File
@@ -29,14 +29,18 @@
# same failure mode check-skill-floor.sh was written for, and the same fix: # same failure mode check-skill-floor.sh was written for, and the same fix:
# convert "someone remembers" into "CI refuses". # convert "someone remembers" into "CI refuses".
# #
# WHY THESE FIVE CHECKS AND NOT MORE. Every check here compares a doc string to # TWO CLASSES OF CHECK, DELIBERATELY. Checks 1-7 compare a doc string to a
# a value that EXISTS IN THIS REPO, so it can never be wrong about the world and # value that EXISTS IN THIS REPO, so they can never be wrong about the world and
# needs no network, no token, and no built image. Claims that require a running # need no network, no token, and no built image. Checks 8-9 compare against what
# container to verify (image sizes, the "N mempalace_* tools" count) are # is PUBLISHED (Docker Hub's measured sizes; the ref labels baked into the last
# deliberately NOT gated: a check that cannot be evaluated honestly at lint time # released image), because those claims have no in-repo anchor at all and had
# would either be skipped or guessed, and a guessing gate is worse than none. # rotted for exactly that reason. They need the network and therefore SKIP,
# If you want those, assert them in scripts/smoke-test.sh where a real image is # loudly and counted, when it is absent -- a skip is neither OK nor a failure,
# available. # because printing an unverified claim as OK is the habit this file exists to
# break, while failing on a third party's uptime would make every release
# hostage to it. Claims that need a RUNNING CONTAINER (the "N mempalace_* tools"
# count, uncompressed on-disk sizes) are still not gated here; assert them in
# scripts/smoke-test.sh where a real image is available.
# #
# DELIBERATELY NOT GATED: Dockerfile.base's `# BASE_REBUILD_DATE:` comment, which # DELIBERATELY NOT GATED: Dockerfile.base's `# BASE_REBUILD_DATE:` comment, which
# is also stale (2026-07-13, three base rebuilds ago). base_tag is a hash of # is also stale (2026-07-13, three base rebuilds ago). base_tag is a hash of
@@ -69,6 +73,22 @@ HUB_MAX_CHARS=25000
WARN_ONLY=0 WARN_ONLY=0
FAILURES=0 FAILURES=0
SKIPS=0
# Tolerance for the published size claims (check 8), as a percentage OF THE
# MEASURED SIZE. The denominator matters: against the claim instead, the same
# drift reads as a different number, and an early draft of this gate took 20%
# from the claim-relative figure and would therefore have MISSED its own
# motivating case. Both bounds are measured, not guessed:
# - the rot that motivated this check: claimed 1.1 GB vs measured 1.37 GB
# = 19.7% off, so the threshold must sit BELOW that or the gate is theatre.
# - the largest legitimate skew, i.e. a claim describing the currently-published
# release while the next tag changes the size: v1.9.1's 1.37 GB against
# v1.9.2's measured 1.23 GB = 11.4% off, so the threshold must sit ABOVE that
# or every size-changing release trips it.
# 15% sits in that 11.4%-19.7% window. Widen it only with a measured reason, and
# re-derive both bounds if you do.
SIZE_TOLERANCE_PCT="${SIZE_TOLERANCE_PCT:-15}"
usage() { usage() {
cat <<'EOF' cat <<'EOF'
@@ -79,6 +99,11 @@ files they describe (Dockerfile.base, Dockerfile.variant).
--warn-only Report drift but exit 0 (advisory use, e.g. a local pre-push hook). --warn-only Report drift but exit 0 (advisory use, e.g. a local pre-push hook).
Environment:
SKIP_SIZE_CHECK=1 skip check 8 (published size claims vs Docker Hub)
SKIP_REF_CHECK=1 skip check 9 (refs moved since the last release are named)
SIZE_TOLERANCE_PCT check 8 tolerance, default 15 (see comment for its bounds)
Exit: 0 = in sync, 1 = drift, 2 = cannot run. Exit: 0 = in sync, 1 = drift, 2 = cannot run.
EOF EOF
} }
@@ -124,6 +149,13 @@ fail() {
ok() { printf ' OK %s\n' "$1"; } ok() { printf ' OK %s\n' "$1"; }
# A check that could not be EVALUATED, as distinct from one that passed.
# Deliberately neither ok() nor fail(): printing it as OK would launder an
# unmeasured claim into a passing one (the exact habit this file exists to
# break), while failing on a third party's uptime would make every release
# hostage to Docker Hub's API. Loud, counted, and surfaced in the summary.
skip() { SKIPS=$((SKIPS + 1)); printf ' SKIP %s\n' "$1"; }
echo "Checking hand-maintained doc claims against the build files they describe." echo "Checking hand-maintained doc claims against the build files they describe."
echo echo
@@ -227,9 +259,429 @@ else
ok "no stale 'Unreleased' pointers in $README or $HUB" ok "no stale 'Unreleased' pointers in $README or $HUB"
fi fi
# ---------------------------------------------------------------------------
# 8. Published size claims vs Docker Hub's MEASURED full_size.
#
# Why this exists: every other claim in these docs is checked against a file
# in this repo, so it cannot rot without someone editing the thing it
# describes. The size claims had no such anchor -- nothing in the repo states
# the image size -- so they quietly went 24% wrong across eight releases
# (DOCKER_HUB.md said ~1.1 GB; :latest measured 1.37 GB on 2026-09-14).
# DOCKER_HUB.md is POSTed to Docker Hub by update-description, so that number
# is the first thing a stranger reads about this image.
#
# Hub's `full_size` tracks the FIRST manifest entry (amd64 here), NOT the sum
# across architectures -- measured: v1.9.2 full_size=1.228 GB, amd64=1.228,
# arm64=1.211, sum=2.439. That matches the table's per-arch "Size
# (compressed)" column, which is why full_size is the right field.
#
# NOT COVERED, deliberately: README.md's ~3.2 GB figures are UNCOMPRESSED
# on-disk sizes, and the registry API exposes compressed sizes only (layer
# sizes in a manifest are compressed; the config blob carries no uncompressed
# totals). Measuring them needs a real pull, so they are out of scope here --
# do not read a green check 8 as covering them.
# ---------------------------------------------------------------------------
# Shared by checks 8 and 9: which Hub repo, and its tag list (one request).
# Derive the repo from the doc's own rows rather than hardcoding it, so a
# rename cannot leave these checks silently probing a repo nobody publishes to.
# shellcheck disable=SC2016 # single quotes are deliberate: this is a sed
# script, and its \( \) groups and \1 backreference must reach sed unexpanded.
HUB_REPO_PATH="$(sed -n 's/^| `\([^:`]*\):[^`]*`.*/\1/p' "$HUB" | head -1)"
HUB_TAGS_JSON=""
HAVE_NET_TOOLS=0
if command -v curl >/dev/null 2>&1 && command -v python3 >/dev/null 2>&1; then
HAVE_NET_TOOLS=1
if [ -n "$HUB_REPO_PATH" ] && \
{ [ "${SKIP_SIZE_CHECK:-0}" != "1" ] || [ "${SKIP_REF_CHECK:-0}" != "1" ]; }; then
HUB_TAGS_JSON="$(curl -sS -m 20 \
"https://hub.docker.com/v2/repositories/${HUB_REPO_PATH}/tags/?page_size=100" \
2>/dev/null || true)"
fi
fi
if [ "${SKIP_SIZE_CHECK:-0}" = "1" ]; then
skip "size claims -- SKIP_SIZE_CHECK=1 was set"
elif [ "$HAVE_NET_TOOLS" = 0 ]; then
skip "size claims -- need both curl and python3 to measure them"
else
if [ -z "$HUB_REPO_PATH" ]; then
skip "size claims -- found no \`repo:tag\` image rows in $HUB to check"
else
if [ -z "$HUB_TAGS_JSON" ]; then
skip "size claims -- Docker Hub API unreachable (offline?); NOT verified"
else
SIZE_RC=0
# NO `|| true` on the python invocation: an early draft had one, and it
# swallowed the exit code so a printed DRIFT line still exited 0 -- a gate
# that reports the defect and passes anyway. The outer `|| SIZE_RC=$?` is
# what keeps `set -e` happy while preserving the code.
SIZE_OUT="$(HUB_MD="$HUB" HUB_JSON="$HUB_TAGS_JSON" TOL="$SIZE_TOLERANCE_PCT" \
python3 <<'PYEOF'
import json, os, re, sys
try:
data = json.loads(os.environ["HUB_JSON"])
except (ValueError, KeyError) as exc:
print(" SKIP size claims -- Hub API returned unparseable JSON (%s)" % exc)
sys.exit(3)
# full_size == first manifest entry (amd64), which is the per-arch number the
# table's "Size (compressed)" column claims. Verified against .images[] sizes.
sizes = {
r["name"]: r["full_size"] / 1e9
for r in data.get("results", [])
if isinstance(r.get("full_size"), int) and r.get("name")
}
if not sizes:
print(" SKIP size claims -- Hub API returned no usable tags")
sys.exit(3)
tol = float(os.environ["TOL"])
row = re.compile(r"^\|\s*`([^`:]+):([^`]+)`\s*\|[^|]*\|\s*~?([0-9]+(?:\.[0-9]+)?)\s*GB\s*\|")
checked = drift = 0
with open(os.environ["HUB_MD"], encoding="utf-8") as fh:
for line in fh:
m = row.match(line)
if not m:
continue # rows saying "same", and every non-image row
_repo, tag, claimed = m.group(1), m.group(2), float(m.group(3))
if "X.Y.Z" in tag:
continue # placeholder row; the concrete tag is checked instead
# base-<hash> is content-addressed and immutable, so its size is
# base-latest's by construction -- probe the alias that always exists.
probe = "base-latest" if tag.startswith("base-") else tag
actual = sizes.get(probe)
if actual is None:
print(" SKIP size %s -- tag '%s' not present on Hub" % (tag, probe))
continue
checked += 1
off = abs(claimed - actual) / actual * 100
if off <= tol:
print(" OK size %s claims ~%.2f GB, Hub measures %.2f GB (%.0f%% off)"
% (tag, claimed, actual, off))
else:
drift += 1
print(" DRIFT size %s claims ~%.2f GB but Hub measures %.2f GB"
" (%.0f%% off, tolerance %.0f%%)" % (tag, claimed, actual, off, tol))
if checked == 0:
print(" SKIP size claims -- no checkable rows resolved to a published tag")
sys.exit(3)
sys.exit(1 if drift else 0)
PYEOF
)" || SIZE_RC=$?
printf '%s\n' "$SIZE_OUT"
case "$SIZE_RC" in
0) : ;;
3) SKIPS=$((SKIPS + 1)) ;;
*)
fail "a published size claim in $HUB has drifted from what Docker Hub
actually serves (see DRIFT above). This page is POSTed to Docker Hub by
update-description, so it is the first size a stranger sees. Re-measure and
update the table:
curl -sS 'https://hub.docker.com/v2/repositories/${HUB_REPO_PATH}/tags/?page_size=100' |
jq -r '.results[] | \"\\(.name) \\(.full_size/1e9)\"'"
;;
esac
fi
fi
fi
# ---------------------------------------------------------------------------
# 9. Everything the NEXT build would bake differently from the LAST PUBLISHED
# release must be named in the CHANGELOG text above that release's heading.
#
# Why this exists, measured 2026-09-19: pi-extensions 25c1265 (a new `task`
# tool and a hook that blocks certain `fork` calls -- a change to how every
# agent in the container delegates work) and mempalace-toolkit 817b3a8 (the
# feed's mine deadline had never reached the transport) both reached this
# image through floating `*_REF=main` ARGs. Neither produced a diff in this
# repo, so nothing here asked for a CHANGELOG entry, and neither had one
# until a reader asked. This is the same shape as check 8: a fact with no
# in-repo anchor rots. The hand practice that existed for it -- the
# "Dependency audit" table in each release's notes ("Baked in vN | Upstream
# now") -- is precisely a "someone remembers" mechanism, and it had lapsed.
#
# How it measures, with no docker/crane/token: the last published `vX.Y.Z`
# is the highest such tag in Hub's tag list (shared with check 8); its
# amd64 config blob is read through the anonymous registry API (token ->
# manifest index -> per-arch manifest -> config) and carries one
# `se.jordbo.pi-devbox.<name>-ref` label per component, each holding the
# SHA that build-args actually baked (resolve-versions in docker-publish.yml
# turns every ref into a SHA before `docker build`). "What the next build
# would bake" is resolved the way that job does it: a 40-hex ARG is itself,
# a tag or branch is `git ls-remote`d (peeled `^{}` first -- an annotated
# tag's un-dereferenced SHA is the tag object, a false alarm this repo has
# already fallen for once), pi-studio is the highest semver tag, and
# `PI_VERSION` / `MEMPALACE_VERSION` are compared as literals against the
# `pi-version` / `mempalace-version` labels (the latter set in Dockerfile.base
# and inherited; absent on releases before it shipped, which reports SKIP).
#
# The rule: baked == would-bake is OK with no mention required. If they
# differ, the text ABOVE the last published version's `## ` heading -- i.e.
# `## Unreleased` plus any not-yet-published `## vX.Y.Z` section, which is
# what the release commit turns Unreleased into -- must contain the
# would-bake value's 7-char SHA prefix (or, for pi-studio, the tag name; for
# pi, the version string). Naming the SHA, not just the repo, is the point:
# it is what the audit table always recorded, and it makes the failure
# message's compare URL a copy-paste away from knowing what moved.
#
# Every upstream commit therefore re-reds this gate until the CHANGELOG
# names the new head. That is the intended cost: the thing that gets baked
# is the thing that gets named, and a typo-fix upstream costs one edited
# SHA here. Read from the TAG like everything else in these docs -- the
# release commit renames Unreleased, so the pending text still covers it.
#
# SKIPs, each counted: SKIP_REF_CHECK=1; no curl/python3; Hub unreachable;
# the release's labels unreadable; one component's upstream unreachable
# (that component only). A published tag whose heading is MISSING from the
# CHANGELOG is a failure, not a skip: that is drift in its own right.
# ---------------------------------------------------------------------------
if [ "${SKIP_REF_CHECK:-0}" = "1" ]; then
skip "ref moves -- SKIP_REF_CHECK=1 was set"
elif [ "$HAVE_NET_TOOLS" = 0 ]; then
skip "ref moves -- need both curl and python3 to read the published labels"
elif ! command -v git >/dev/null 2>&1; then
skip "ref moves -- need git (ls-remote) to resolve what the next build would bake"
elif [ -z "$HUB_REPO_PATH" ]; then
skip "ref moves -- found no \`repo:tag\` image rows in $HUB to locate the published image"
elif [ -z "$HUB_TAGS_JSON" ]; then
skip "ref moves -- Docker Hub API unreachable (offline?); NOT verified"
else
# One plain top-level assignment per ARG, on purpose: read_arg exits 2 on a
# missing ARG, and under `set -e` that only propagates from a bare
# `VAR="$(...)"`. Nested inside a heredoc's $(...) the exit would be swallowed
# by `cat`, and a renamed ARG would leave this check comparing a label against
# an empty string and reporting the component "unchanged".
TOOLKIT_REPO="$(read_arg "$DF_VARIANT" PI_TOOLKIT_REPO)"; TOOLKIT_REF="$(read_arg "$DF_VARIANT" PI_TOOLKIT_REF)"
EXTENSIONS_REPO="$(read_arg "$DF_VARIANT" PI_EXTENSIONS_REPO)"; EXTENSIONS_REF="$(read_arg "$DF_VARIANT" PI_EXTENSIONS_REF)"
FORK_REPO="$(read_arg "$DF_VARIANT" PI_FORK_REPO)"; FORK_REF="$(read_arg "$DF_VARIANT" PI_FORK_REF)"
OBSMEM_REPO="$(read_arg "$DF_VARIANT" PI_OBSMEM_REPO)"; OBSMEM_REF="$(read_arg "$DF_VARIANT" PI_OBSMEM_REF)"
ATELIER_REPO="$(read_arg "$DF_VARIANT" PI_ATELIER_REPO)"
MPTK_REPO="$(read_arg "$DF_BASE" MEMPALACE_TOOLKIT_REPO)"; MPTK_REF="$(read_arg "$DF_BASE" MEMPALACE_TOOLKIT_REF)"
STUDIO_REPO="$(read_arg "$DF_VARIANT" PI_STUDIO_REPO)"
SKILLSET_SNAPSHOT="$(read_arg "$DF_VARIANT" SKILLSET_SNAPSHOT_REF)"
# name|kind|repo|ref -- one line per label the variant image carries.
# kinds: ref = branch/tag/SHA resolved like resolve-versions does;
# studio = highest semver tag of the repo (label lives on <tag>-studio);
# literal = the ARG value IS the baked value (a SHA pin, a version).
REF_COMPONENTS="pi-toolkit|ref|$TOOLKIT_REPO|$TOOLKIT_REF
pi-extensions|ref|$EXTENSIONS_REPO|$EXTENSIONS_REF
pi-fork|ref|$FORK_REPO|$FORK_REF
pi-obsmem|ref|$OBSMEM_REPO|$OBSMEM_REF
pi-atelier|ref|$ATELIER_REPO|$ATELIER_ACTUAL
mempalace-toolkit|ref|$MPTK_REPO|$MPTK_REF
pi-studio|studio|$STUDIO_REPO|
skillset-snapshot|literal||$SKILLSET_SNAPSHOT
pi-version|literal||$PI_ACTUAL
mempalace-version|literal||$MEMPALACE_ACTUAL"
REF_RC=0
# Same discipline as check 8: no `|| true` on the python, or a printed DRIFT
# exits 0. Per-component SKIP lines are counted afterwards by grep, so a run
# that evaluated eight components and could not reach the ninth reports one
# skip, not a green tick over the ninth.
REF_OUT="$(HUB_REPO="$HUB_REPO_PATH" HUB_JSON="$HUB_TAGS_JSON" CHANGELOG="CHANGELOG.md" \
COMPONENTS="$REF_COMPONENTS" python3 <<'PYEOF'
import json, os, re, subprocess, sys, urllib.request, urllib.parse
SHA40 = re.compile(r"^[0-9a-f]{40}$")
SEMVER = re.compile(r"^v?[0-9]+\.[0-9]+\.[0-9]+$")
LABEL = "se.jordbo.pi-devbox."
def ver_key(tag):
return tuple(int(x) for x in tag.lstrip("v").split("."))
def http_json(url, headers=None, timeout=30):
req = urllib.request.Request(url, headers=headers or {})
with urllib.request.urlopen(req, timeout=timeout) as resp:
return json.loads(resp.read().decode("utf-8"))
def labels_of(repo, tag):
"""Config labels of <repo>:<tag>'s amd64 image via the anonymous registry API."""
tok = http_json(
"https://auth.docker.io/token?service=registry.docker.io&scope="
+ urllib.parse.quote(f"repository:{repo}:pull", safe=":")
)["token"]
hdr = {
"Authorization": f"Bearer {tok}",
"Accept": ", ".join([
"application/vnd.oci.image.index.v1+json",
"application/vnd.docker.distribution.manifest.list.v2+json",
"application/vnd.oci.image.manifest.v1+json",
"application/vnd.docker.distribution.manifest.v2+json",
]),
}
base = f"https://registry-1.docker.io/v2/{repo}"
man = http_json(f"{base}/manifests/{tag}", hdr)
if "manifests" in man: # multi-arch index: pick linux/amd64, as check 8 does
cands = [m for m in man["manifests"]
if m.get("platform", {}).get("architecture") == "amd64"
and m.get("platform", {}).get("os") == "linux"]
if not cands:
raise RuntimeError("no linux/amd64 entry in the manifest index")
man = http_json(f"{base}/manifests/{cands[0]['digest']}", hdr)
cfg = http_json(f"{base}/blobs/{man['config']['digest']}", hdr)
return cfg.get("config", {}).get("Labels") or {}
def ls_remote(repo, *patterns):
# GIT_TERMINAL_PROMPT=0: a repo flipped private must fail fast as a SKIP,
# not sit waiting for a username on a CI runner until the job times out.
env = dict(os.environ, GIT_TERMINAL_PROMPT="0")
out = subprocess.run(["git", "ls-remote", repo, *patterns], env=env,
capture_output=True, text=True, timeout=60, check=True).stdout
return {line.split("\t")[1]: line.split("\t")[0] for line in out.splitlines() if "\t" in line}
def resolve_ref(repo, ref):
"""What docker-publish.yml's resolve-versions would pass as the build-arg."""
if SHA40.match(ref):
return ref, ref
refs = ls_remote(repo, f"refs/heads/{ref}", f"refs/tags/{ref}", f"refs/tags/{ref}^{{}}")
for key in (f"refs/tags/{ref}^{{}}", f"refs/heads/{ref}", f"refs/tags/{ref}"):
if key in refs:
return refs[key], ref
raise RuntimeError(f"'{ref}' is neither a branch nor a tag of {repo}")
def resolve_studio(repo):
refs = ls_remote(repo, "refs/tags/*")
tags = {k[len("refs/tags/"):]: v for k, v in refs.items()}
names = sorted((t for t in tags if SEMVER.match(t)), key=ver_key)
if not names:
raise RuntimeError(f"no semver tag at {repo}")
tag = names[-1]
return tags.get(tag + "^{}", tags[tag]), tag
def compare_url(repo, a, b):
root = repo[:-4] if repo.endswith(".git") else repo
return f"{root}/compare/{a}...{b}"
try:
hub = json.loads(os.environ["HUB_JSON"])
except (ValueError, KeyError) as exc:
print(" SKIP ref moves -- Hub API returned unparseable JSON (%s)" % exc)
sys.exit(3)
released = sorted((r["name"] for r in hub.get("results", [])
if isinstance(r.get("name"), str) and re.fullmatch(r"v[0-9]+\.[0-9]+\.[0-9]+", r["name"])),
key=ver_key)
if not released:
print(" SKIP ref moves -- Hub lists no published vX.Y.Z tag to compare against")
sys.exit(3)
last = released[-1]
repo = os.environ["HUB_REPO"]
# The text every not-yet-published change lives in: everything above the last
# published version's heading. Its absence is drift, not a skip.
text = open(os.environ["CHANGELOG"], encoding="utf-8").read()
# (\s|$) rather than \b: a word boundary would accept "## v1.9.2-rc1" or
# "## v1.9.2-typo" as v1.9.2's heading. Caught by the sabotage test, not review.
m = re.search(r"^## v?%s(\s|$)" % re.escape(last.lstrip("v")), text, re.M)
if not m:
print(" DRIFT ref moves -- %s is the last PUBLISHED tag on Hub but %s has no '## %s' heading"
% (last, os.environ["CHANGELOG"], last))
sys.exit(1)
pending = text[:m.start()].lower()
try:
labels = labels_of(repo, last)
except Exception as exc: # network, auth, shape -- all "could not measure"
print(" SKIP ref moves -- could not read %s:%s's labels from the registry (%s); NOT verified"
% (repo, last, exc))
sys.exit(3)
studio_labels = None
checked = drift = 0
problems = []
for line in os.environ["COMPONENTS"].splitlines():
if not line.strip():
continue
name, kind, url, ref = line.split("|", 3)
# <name>-ref labels hold SHAs; names that already end in -version are the
# label (pi-version, mempalace-version) -- a version string, compared literally.
key = LABEL + name if name.endswith("-version") else LABEL + name + "-ref"
try:
if kind == "studio":
if studio_labels is None:
studio_labels = labels_of(repo, last + "-studio")
baked = studio_labels.get(key)
else:
baked = labels.get(key)
except Exception as exc:
print(" SKIP %-18s -- could not read %s:%s-studio's labels (%s)" % (name, repo, last, exc))
continue
if not baked:
print(" SKIP %-18s -- %s carries no %s label" % (name, last, key))
continue
try:
if kind == "ref":
now, shown = resolve_ref(url, ref)
elif kind == "studio":
now, shown = resolve_studio(url)
else:
now, shown = ref, ref
except Exception as exc:
print(" SKIP %-18s -- could not resolve what the next build would bake (%s)" % (name, exc))
continue
checked += 1
is_sha = bool(SHA40.match(now))
short = (lambda s: s[:7] if SHA40.match(s) else s)
if baked == now:
print(" OK %-18s unchanged since %s (%s)" % (name, last, short(now)))
continue
names = [now[:7].lower()] if is_sha else [now.lower()]
if kind == "studio":
names.append(shown.lower())
if any(n in pending for n in names):
print(" OK %-18s %s -> %s since %s, named above the %s heading"
% (name, short(baked), short(now), last, last))
continue
drift += 1
hint = compare_url(url, baked, now) if (url and is_sha and SHA40.match(baked)) else ""
problems.append(" %-18s %s -> %s%s" % (name, short(baked), short(now), (" " + hint) if hint else ""))
print(" DRIFT %-18s %s -> %s since %s, NOT named above the %s heading"
% (name, short(baked), short(now), last, last))
if problems:
print(" Name each new value (7-char SHA prefix, or the tag/version) in CHANGELOG.md above '## %s':" % last)
print("\n".join(problems))
if checked == 0 and drift == 0:
print(" SKIP ref moves -- no component could be evaluated")
sys.exit(3)
sys.exit(1 if drift else 0)
PYEOF
)" || REF_RC=$?
printf '%s\n' "$REF_OUT"
REF_SKIPS="$(printf '%s\n' "$REF_OUT" | grep -c '^ SKIP ' || true)"
case "$REF_RC" in
0) SKIPS=$((SKIPS + REF_SKIPS)) ;;
3) SKIPS=$((SKIPS + 1)) ;;
*)
SKIPS=$((SKIPS + REF_SKIPS))
fail "a component the next build would bake differently from the last published
release is not named in CHANGELOG.md (see DRIFT above). These reach the image
through floating refs, so nothing else in this repo records that they moved;
the CHANGELOG entry is the only place a reader of the next tag can learn it.
Name the new SHA (7 chars is enough) where you describe the change -- the
compare URL above shows what moved."
;;
esac
fi
echo echo
if [ "$FAILURES" -eq 0 ]; then if [ "$FAILURES" -eq 0 ]; then
echo "OK: every checked doc claim matches the build files." if [ "$SKIPS" -gt 0 ]; then
echo "OK: every checked doc claim matches the build files" \
"($SKIPS check(s) SKIPPED and therefore NOT verified -- see SKIP above)."
else
echo "OK: every checked doc claim matches the build files."
fi
exit 0 exit 0
fi fi
+7 -1
View File
@@ -32,10 +32,16 @@
# #
# SEVERITY CHOICE # SEVERITY CHOICE
# -S error is 0 findings across this repo when clean, so it is free to add. # -S error is 0 findings across this repo when clean, so it is free to add.
# -S warning is NOT free here (19x SC2088 tilde-in-quotes in # -S warning is NOT free here (20x SC2088 tilde-in-quotes in
# recreate-sanity-check.sh, plus assorted SC2016 — both intentional), and a # recreate-sanity-check.sh, plus assorted SC2016 — both intentional), and a
# noisy gate trains people to ignore it. Error-only, matching the # noisy gate trains people to ignore it. Error-only, matching the
# SHELLCHECK_OPTS philosophy in lint.yml. # SHELLCHECK_OPTS philosophy in lint.yml.
# Reproduce the count before editing it (the `$ ` prefix is load-bearing: a
# comment whose first word is "shellcheck" is parsed as a DIRECTIVE, and a
# malformed one is SC1072/SC1073 at severity error — this gate caught exactly
# that when the line was first written without it):
# $ shellcheck -S warning -f gcc scripts/*.sh rootfs/usr/local/bin/* \
# entrypoint*.sh hooks/* | grep -c SC2088
# #
# Usage: bash scripts/lint-shell.sh [root] (default root: repo top level) # Usage: bash scripts/lint-shell.sh [root] (default root: repo top level)
set -uo pipefail set -uo pipefail
+105 -2
View File
@@ -11,7 +11,8 @@
# pi-observational-memory / (studio variant) pi-studio package # pi-observational-memory / (studio variant) pi-studio package
# registrations in settings.json packages[] # registrations in settings.json packages[]
# - Shell defaults re-seeded from /etc/skel-devbox # - Shell defaults re-seeded from /etc/skel-devbox
# - /tmp/sshcm exists with mode 700 (ssh ControlMaster dir) # - ssh ControlMaster works: /tmp/sshcm exists 700 AND the ControlPath that
# ssh actually resolves (ssh -G) is a writable directory
# - /opt toolkits intact # - /opt toolkits intact
# - Known expected-absences don't regress # - Known expected-absences don't regress
# #
@@ -425,13 +426,115 @@ if [ -d /opt/pi-atelier ] && command -v jq >/dev/null 2>&1; then
fi fi
echo echo
echo "-- ssh ControlMaster dir --" echo "-- ssh ControlMaster: socket dir + EFFECTIVE ControlPath --"
# TWO LAYERS, and the second is the one that has actually broken in the field.
#
# LAYER 1 (original check): /tmp/sshcm, the directory entrypoint-user.sh creates
# for the base image's system drop-in
# (/etc/ssh/ssh_config.d/00-devbox-controlmaster.conf).
#
# LAYER 2 (added 2026-09-15): the directory a config NAMES — which is not the
# same question, and asserting layer 1 is structurally blind to it. On
# emb-7kj4vr4g a durable ~/.pi/ssh/config pointed ControlPath at /tmp/ssh-cm
# (with a hyphen), a directory nothing in the image creates. EVERY ssh died
# unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory
# rc=255 with the remote command never running — while this script printed a
# green tick for layer 1, truthfully, about the wrong object.
#
# The same rc=255 has a second, independent cause already documented in prose in
# Dockerfile.base ("SSH client defaults" CAVEAT) and never verified anywhere: a
# per-host `ControlPath ~/.ssh/cm/%r@%h:%p` inherited from a bind-mounted
# READ-ONLY ~/.ssh. Measured to be the identical failure class:
# unix_listener: cannot bind to path ~/.ssh/cm/...: Read-only file system
# So do not guess which config wins — ask ssh. `ssh -G` applies real config
# precedence (first-obtained-value-wins, system drop-in, Include, -F override)
# and prints the fully expanded ControlPath. Require its parent to exist and be
# writable. Cost measured at 0.116 s for 48 hosts; -G never opens a connection.
if [ -d /tmp/sshcm ] && [ "$(stat -c %a /tmp/sshcm 2>/dev/null)" = "700" ]; then if [ -d /tmp/sshcm ] && [ "$(stat -c %a /tmp/sshcm 2>/dev/null)" = "700" ]; then
pass "/tmp/sshcm exists with mode 700" pass "/tmp/sshcm exists with mode 700"
else else
fail "/tmp/sshcm missing or not mode 700" fail "/tmp/sshcm missing or not mode 700"
fi fi
# Probe one route. $1 = label, $2 = config to force with -F ("" = ssh's own
# default precedence), $3 = severity when a ControlPath dir is unusable.
#
# SEVERITY SPLIT IS DELIBERATE. The default route legitimately resolves into the
# read-only ~/.ssh on any host whose own config pins ControlPath there, and the
# supported workaround (`ssh -F ~/.ssh-local/config`) already exists — so that
# is a warn, not a fail. Failing it would paint this script red on every run of
# every device, and a check that fires benignly every time is one you learn to
# ignore. The sidecar route is the PRESCRIBED one, so there it is a hard fail.
_ssh_cm_probe() {
local label="$1" cfg="${2:-}" sev="${3:-fail}"
local h out cm cp dir n=0 shown
local bad=()
while IFS= read -r h; do
[ -n "$h" ] || continue
if [ -n "$cfg" ]; then
out=$(ssh -F "$cfg" -G "$h" 2>/dev/null) || continue
else
out=$(ssh -G "$h" 2>/dev/null) || continue
fi
cm=$(printf '%s\n' "$out" | awk '/^controlmaster /{print $2; exit}')
case "$cm" in '' | no | none | false) continue ;; esac
cp=$(printf '%s\n' "$out" | awk '/^controlpath /{print $2; exit}')
case "$cp" in '' | none) continue ;; esac
n=$((n + 1))
dir=$(dirname "$cp")
if [ ! -d "$dir" ] || [ ! -w "$dir" ]; then
bad+=("$h")
fi
done <<< "$SSH_CM_HOSTS"
# Cap the host list. A 50-name line is the noisy gate this repo already warns
# about in lint-shell.sh: unreadable output is ignored output. Six names plus
# a count is enough to identify the class and act.
if [ "${#bad[@]}" -gt 0 ]; then
shown="${bad[*]:0:6}"
if [ "${#bad[@]}" -gt 6 ]; then
shown="$shown (+$(( ${#bad[@]} - 6 )) more)"
fi
fi
if [ "$n" -eq 0 ]; then
warn "$label: no host resolves to ControlMaster on — effective ControlPath not exercised"
elif [ "${#bad[@]}" -eq 0 ]; then
pass "$label: $n ControlMaster host(s), every ControlPath dir exists and is writable"
elif [ "$sev" = "warn" ]; then
warn "$label: ${#bad[@]}/$n host(s) resolve ControlPath to a missing or unwritable dir [$shown] — expected when ~/.ssh/config pins ControlPath inside the read-only ~/.ssh; use 'ssh -F ~/.ssh-local/config' (see Dockerfile.base CAVEAT)"
else
fail "$label: ${#bad[@]}/$n host(s) resolve ControlPath to a missing or unwritable dir [$shown] — ssh dies rc=255 'unix_listener: cannot bind to path' and the remote command never runs"
fi
}
if command -v ssh >/dev/null 2>&1; then
# Concrete Host aliases only: patterns (*, ?) and negations (!) are not
# connectable targets, so `ssh -G` on them proves nothing.
_cm_cfgs=()
if [ -r "$HOME/.ssh/config" ]; then _cm_cfgs+=("$HOME/.ssh/config"); fi
if [ -r "$HOME/.ssh-local/config" ]; then _cm_cfgs+=("$HOME/.ssh-local/config"); fi
if [ "${#_cm_cfgs[@]}" -gt 0 ]; then
SSH_CM_HOSTS=$(awk 'tolower($1)=="host"{for(i=2;i<=NF;i++) if ($i !~ /[*?!]/) print $i}' \
"${_cm_cfgs[@]}" 2>/dev/null | sort -u)
else
SSH_CM_HOSTS=""
fi
if [ -z "$SSH_CM_HOSTS" ]; then
warn "no concrete Host aliases in ~/.ssh/config or ~/.ssh-local/config — effective ControlPath not verified"
else
_ssh_cm_probe "default ssh precedence" "" warn
if [ -r "$HOME/.ssh-local/config" ]; then
_ssh_cm_probe "ssh -F ~/.ssh-local/config" "$HOME/.ssh-local/config" fail
else
warn "~/.ssh-local/config absent — setup-lan-access.sh did not run; the prescribed multiplex route is unverified"
fi
fi
else
warn "ssh not on PATH — effective ControlPath not verified"
fi
echo echo
echo "-- Shell defaults re-seeded from /etc/skel-devbox --" echo "-- Shell defaults re-seeded from /etc/skel-devbox --"
if [ -f "$HOME/.bash_aliases" ]; then if [ -f "$HOME/.bash_aliases" ]; then
+13
View File
@@ -597,6 +597,19 @@ if [ -n "$LBL" ] && [ "$LBL" != "<no value>" ]; then
else else
printf " ❌ OCI label se.jordbo.pi-devbox.pi-extensions-ref missing or empty\n"; FAIL=$((FAIL+1)) printf " ❌ OCI label se.jordbo.pi-devbox.pi-extensions-ref missing or empty\n"; FAIL=$((FAIL+1))
fi fi
# mempalace-version is set in Dockerfile.base and INHERITED by the variant, so
# it states the pin of the base this image actually built on. It must equal the
# installed binary: the one way they diverge is a base built with
# INSTALL_MEMPALACE=false (label says 3.x, nothing installed) or an install that
# resolved to something other than the pin — both invisible to a label-only
# check. Same ground-truth rule as the manifest assertion above.
MP_LBL=$(docker inspect --format '{{ index .Config.Labels "se.jordbo.pi-devbox.mempalace-version" }}' "$IMAGE" 2>/dev/null || true)
MP_BIN=$(docker run --rm --entrypoint= "$IMAGE" sh -c 'mempalace --version 2>/dev/null | head -n1 | tr -d "\r"' 2>/dev/null || true); MP_BIN=${MP_BIN##* }
if [ -n "$MP_LBL" ] && [ "$MP_LBL" != "<no value>" ] && [ "$MP_LBL" = "$MP_BIN" ]; then
printf " ✅ OCI label se.jordbo.pi-devbox.mempalace-version=%s equals the installed core\n" "$MP_LBL"; PASS=$((PASS+1))
else
printf " ❌ OCI label se.jordbo.pi-devbox.mempalace-version=[%s] vs installed mempalace=[%s]\n" "$MP_LBL" "$MP_BIN"; FAIL=$((FAIL+1))
fi
# ── Runtime deployment (needs entrypoint to run) ────────────────────── # ── Runtime deployment (needs entrypoint to run) ──────────────────────
echo "" echo ""