Files
pi-devbox/CHANGELOG.md
T
Joakim Persson cb6d9e5dd0
Lint / hadolint (push) Successful in 8s
Lint / skill-floor (push) Successful in 10s
Lint / doc-drift (push) Successful in 14s
Lint / actionlint (push) Successful in 19s
check-doc-drift: check 9 — what the next build bakes differently must be named in the CHANGELOG
The two most recent Unreleased entries (pi-extensions 25c1265: task tool +
fork-gate; mempalace-toolkit 817b3a8: the mine deadline that never reached
the transport) reached this image through floating *_REF=main ARGs. Neither
produced a diff here, so nothing asked for a CHANGELOG line, and neither had
one until a reader asked. Same class as check 8 — a claim with no in-repo
anchor rots — and the hand practice that covered it (the per-release
"Dependency audit" table, Baked-in vs Upstream-now) is a someone-remembers
mechanism that had lapsed.

Method, no docker/crane/token: last published vX.Y.Z from Hub's tag list
(one request, now shared with check 8); its amd64 config blob via the
anonymous registry API carries one se.jordbo.pi-devbox.<name>-ref label per
component with the SHA the build-args baked. "Would bake" is resolved the way
resolve-versions does it: 40-hex ARG is itself, branch/tag via git ls-remote
with the peeled ^{} form preferred, pi-studio = highest semver tag (label on
<tag>-studio), PI_VERSION as a literal against the pi-version label. Nine
components, ~7.5 s. Unchanged: OK. Moved: the new value's 7-char prefix (tag
for studio, version for pi) must appear above the last published version's
"## " heading — Unreleased plus any not-yet-published section, which is what
the release commit turns Unreleased into, so the tag build passes on the same
text. A published tag with no heading is a failure, not a skip.

First run found a move nobody had recorded: pi-observational-memory 7b397f4
-> cba0334 (6 commits, 3.1.1 -> 3.1.3; memory workers now routed by the
model's exact provider instead of the first registered provider with a
matching api — upstream #70). Named in the CHANGELOG with the honest scope:
no expected effect on the shipped bedrock-only configuration, unmeasured.

Sabotage-tested, expectation written before each run:
  obsmem SHA removed from CHANGELOG ............ rc=1  DRIFT
  Unreleased renamed to "## v1.9.3 — …" ........ rc=0  (release-commit shape)
  "## v1.9.2" heading mangled to v1.9.2-typo .... rc=1  — this one FAILED first:
      \b accepted "v1.9.2-typo" as v1.9.2's heading; now (\s|$)
  bare "## v1.9.2" (no date) ................... rc=0
  ARG PI_OBSMEM_REPO renamed ................... rc=2  (blind gate must not pass;
      read_arg calls are bare top-level assignments on purpose — nested in a
      heredoc's $(...) the exit 2 is swallowed by cat)
  offline (proxy to a closed port) ............. rc=0, 2 SKIPs counted
  SKIP_REF_CHECK=1 ............................. rc=0, 1 SKIP counted
  restored ..................................... rc=0, files byte-identical
GIT_TERMINAL_PROMPT=0 on ls-remote: a repo flipped private fails in 0.3 s as
a SKIP instead of waiting for a username until the job times out.

Header, lint.yml job comment and AGENTS.md all still said this gate "needs
no network" — false since check 8 (2026-09-14). Now: two classes, hermetic
1-7 and published-state 8-9 that SKIP loudly and counted when offline.
Considered and not added: compose <-> .env.example variable cross-check (the
four mismatches are commented-out lines, the mempalace-server compose file's
own variables, and entrypoint-consumed ones — it would fire on nothing
wrong); "documented tag exists on Hub" (check 8 already SKIPs by name; a
hard fail would misreport the tag-to-publish window).
2026-09-19 17:15:12 +02:00

5524 lines
341 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Changelog
All notable changes to the pi-devbox container image.
From v1.0.0 onward, tags follow semver:
- **major** — architectural changes (v1.0.0 = decoupled from opencode-devbox)
- **minor** — new variants, significant base additions
- **patch** — pi version bumps, smaller fixes
Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
---
## Unreleased
**New check 9 in `scripts/check-doc-drift.sh`: anything the next build would bake
differently from the last *published* release must be named in the CHANGELOG text
above that release's heading.** The two entries below this one are why. The
`task` tool and `fork-gate` (pi-extensions `25c1265`) and the mine-deadline fix
(mempalace-toolkit `817b3a8`) both reach this image through floating
`*_REF=main` ARGs, so neither produced a diff in this repo, nothing here asked
for a CHANGELOG line, and neither had one until a reader asked. Same shape as
check 8: a claim with no in-repo anchor rots. The hand practice that existed for
it — the "Dependency audit" table in each release's notes, *Baked in vN* against
*Upstream now* — is a "someone remembers" mechanism, and it had lapsed.
How it measures, with no `docker`, `crane` or token: the last published
`vX.Y.Z` is the highest such tag in Hub's list (one request, shared with
check 8); that tag's amd64 config blob is read through the anonymous registry
API (token → index → per-arch manifest → config) and carries one
`se.jordbo.pi-devbox.<name>-ref` label per component holding the SHA the
build-args actually baked. "What the next build would bake" is resolved the way
`resolve-versions` does it — a 40-hex ARG is itself, a branch or tag is
`git ls-remote`d with the peeled `^{}` form preferred (the un-dereferenced SHA of
an annotated tag is the tag object; this repo has raised that false alarm once
already), pi-studio is the highest semver tag read from `<tag>-studio`'s labels,
and `PI_VERSION` is compared as a literal against the `pi-version` label. Nine
components, 7.5 s.
The rule: unchanged needs no mention. Moved requires the new value's 7-char SHA
prefix (tag name for pi-studio, version string for pi) somewhere above the last
published version's `## ` heading — `## Unreleased` plus any not-yet-published
`## vX.Y.Z`, which is what the release commit turns Unreleased into, so the tag
build passes on the same text — sabotage-tested: renaming `## Unreleased` to
`## v1.9.3 — …` stays green; mangling the published `## v1.9.2` heading goes
red, and that test caught a `\b` that would have accepted `v1.9.2-typo` as the
heading (now `(\s|$)`). Naming the SHA rather than the repo is
deliberate: it is what the audit table always recorded, and it makes the failure
message's compare URL one click from knowing what moved. Every upstream commit
re-reds the gate until the CHANGELOG names the new head; that is the intended
cost — **the thing that gets baked is the thing that gets named.** A published
tag with no CHANGELOG heading is a failure, not a skip.
**First run found a move nobody had recorded.** `pi-observational-memory`
`7b397f4 → cba0334` (6 upstream commits, 2026-09-14..16, 3.1.1 → 3.1.3): the
memory workers' `streamSimple` lookup used to iterate every extension-registered
provider and take the first whose `api` matched the model's, so two providers
sharing an API type could route the observer to the wrong one (upstream #70);
it now asks for the model's exact provider and keeps the `api` match as a
consistency check. Reaches this image on the next variant build via
`PI_OBSMEM_REF=master`. No behaviour change expected on the shipped
configuration — every profile here uses the built-in `amazon-bedrock` provider
and no extension registers one — but that is an expectation, not a
measurement; the code path only differs when an extension has called
`registerProvider`.
Header and `lint.yml` comment corrected alongside: both still said this gate
"needs no network", which check 8 made false on 2026-09-14. There are now two
classes — hermetic checks 1–7, and published-state checks 8–9 that SKIP loudly
and counted when offline. Considered and not added, with reasons in the script:
a `docker-compose.yml` ↔ `.env.example` variable cross-check (the four
mismatches are commented-out lines, the mempalace-server compose file's own
documented variables, and two entrypoint-consumed variables — a gate there
would fire on nothing wrong), and a "documented tag exists on Hub" check
(check 8 already SKIPs a missing tag by name, and a hard fail would
misreport the window between tagging and publish).
---
**The rule "use `pi-task`, not `fork`, for a brief that carries a prohibition" was
written in the global `AGENTS.md` and in the pi-extensions skill, and it lost to
the `fork` tool's own description anyway.** Measured on tor-ms22, 2026-09-17: five
of five fork briefs in one session carried "do not"; one returned verbatim quotes
that did not exist in the source, and four had disjoint write boundaries that
fork cannot enforce — re-run as `pi-task`, all four passed their envelope. The
image now ships the rule where the decision is *made*, not where it is read about:
- **`fork-gate.ts`** (pi-extensions `25c1265`): a `tool_call` hook that blocks a
`fork` whose brief contains a prohibition (*do not / never / only …*), a write
boundary (*only touch / read-only / stay within …*) or a clause-initial
file-changing imperative (*Edit …, Commit …, Fix …*). The block reason the model
reads **is** the `task(...)` call to make instead. It matches wording, not
intent, and says so. `PI_FORK_GATE=off` logs instead of blocking.
- **`task.ts`** (same commit): `pi-task` registered as the `task` tool, with the
decision rule in its description and in `promptGuidelines` — which pi appends
to the **system prompt**, the one place compaction cannot remove it from.
Before spending a model run it rejects the two spec errors that make a boundary
violation certain (`write_allowed` not an *exact subset* of `roots` — pi-task
keys deltas by root string; a writable root nested inside a watched-only root,
whose porcelain would change every time) and serialises sibling tasks whose
roots overlap (parallel siblings saw each other's writes as violations,
2026-09-17). A FAIL verdict is a *result*; only the CLI refusing to run is an
error.
- **`pi-global-AGENTS.md`** (pi-toolkit `9c87ee8`): the delegation section now
opens "`task` first, `fork` second" — the one discriminator (what the child
sees), a three-question pre-flight before any `fork(...)`, a copy-paste minimal
call with the roots contract. This file is in the system prompt; the skill is
not, which is why the rule lives here.
- **Skill floor refreshed** to `25c1265` (`check-skill-floor.sh` OK, tree
`9b85a633…`); mirror in skillset `debc8f6`.
Why prose failed, as mechanisms: `fork` is a **tool** — its self-recommending
description ("exploration, implementation, testing, review…") is in the tool
list on every turn and survives compaction; the skill is gone after the first
compaction; `pi-task` was a CLI to be remembered and reached through `bash` with a
hand-written JSON spec. The asymmetry *widens* in exactly the long sessions where
fork is worst, and fork deletes its temp dir on exit, so its failures were found
only by re-verifying the narrative. Third time on this fleet that a rule held in
prose and violated in practice was fixed by moving it into a hook
(`check-secrets`, `check-egl-only`, now delegation).
Evidence, two-sided: `test/fork-gate.test.mjs` pins 15 must-block briefs
(including the real shapes) and 10 must-pass (including *"Write a summary…"*,
*"Report which files were modified…"*, *"Give me an update…"* — verbs a naive
list misfires on); the shipped classifier over the five real briefs from the
motivating session redirects 5/5. Live in `pi -p`: the fork was intercepted
before any child spawned (no `/tmp/pi-fork-*` directory), `task` returned PASS in
4 s / $0.014 with an evidence pointer and audit dir, a `usd=0.000001` budget came
back as a FAIL *result* (`isError=false`), and a nested-root spec was rejected
with no audit dir created.
Nothing to wire in this repo: `PI_EXTENSIONS_REF=main` floats and
`pi-extensions/install.sh` symlinks every `extensions/*.ts` on container start, so
the next build ships both. A container running today has neither —
`~/.pi/agent/extensions/` links into `/opt/pi-extensions`, which is the baked ref.
---
**`[mempalace ext] feed (tick) failed: mempalace remote request 'tools/call'
failed: timed out after 60000ms` is the same event as the `mine timed out after
30000ms` message the 2026-09 toolkit fix addressed, one deadline further down —
and that fix was incomplete.** mempalace-toolkit `817b3a8` (2026-09-18; v1.9.2
baked `dab989b`; the floating `MEMPALACE_TOOLKIT_REF=main` picks it up on the next
build). `MEMPALACE_FEED_MINE_TIMEOUT_MS` had been raised to 300 000 but was only
*raced* against `client.callTool("mempalace_mine")`; `callTool()` had no way to
carry a deadline, so every mine went out under the transport's generic
per-request timeout — `MEMPALACE_MCP_TIMEOUT_MS`, 60 000, the value the "Stall
protection" comment in `Dockerfile.base` documents — which fired first on every
honest 60 s+ mine on the shared single-writer palace. The 300 s was unreachable.
Over HTTP nothing is lost (the mine continues server-side and is idempotent);
over stdio it was worse than noise — that transport **kills the child** on
timeout, so there the mine really was aborted at 60 s.
Fix: `callTool(name, args, { timeoutMs })` on both transports, the feed passes
its own deadline down, plain calls keep 60 s (a *query* that slow is wedged; the
race stays as the liveness guard for a transport with its timeout disabled).
`scripts/test-mcp-call-timeout.sh` cuts `RemoteMcpClient` out of the shipped
file, drives it against a local JSON-RPC server that delays `tools/call`, and
asserts three things — a plain call rejects at the generic deadline, the override
outlives it, the override is itself a deadline: 2 of 6 fail on `dab989b`, 6 of 6
pass on `817b3a8`. `test-owed-withdrawal.sh` (17) and `check-mcp-client-sync.sh`
stay green; sync token `v1` untouched, since nothing in the protocol changed.
Reading for the fleet: on a fixed build that message means a mine exceeded
*five* minutes — look at palace size or a competing writer, not at the timeout.
---
**`scripts/recreate-sanity-check.sh` asserted that `/tmp/sshcm` exists while every
`ssh` in the container was dying `rc=255`, and it was right to — it was checking
the directory the *image* creates, and the breakage was in the directory a
*config* named.** Found on the v1.9.2 first boot on `emb-7kj4vr4g`. A durable
`~/.pi/ssh/config` (hand-written into the `~/.pi` named volume by the previous
session, so it would survive the recreate) declared `ControlPath
/tmp/ssh-cm/%C` — with a hyphen. Nothing in this repo creates that path; the
canonical directory is `/tmp/sshcm`, spelled the same way in four places
(`Dockerfile.base`, `entrypoint-user.sh`, `recreate-sanity-check.sh`,
`smoke-test.sh`). Result:
```
unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory
```
`rc=255`, and the remote command never ran at all — `ControlMaster auto` with an
unusable `ControlPath` fails hard rather than falling back to an unmultiplexed
connection. The old check passed truthfully, about the wrong object. **Two
independent facts about the same subsystem can both be true while the subsystem
is dead; a check that asserts only one of them cannot see their disagreement.**
**New in `recreate-sanity-check.sh`: resolve the ControlPath `ssh` itself would
use, via `ssh -G`, and require its parent to exist and be writable.** `-G`
applies real config precedence — first-obtained-value-wins, the system drop-in,
`Include`, an `-F` override — so it answers "which rule captured this host"
instead of re-implementing the guess. It never opens a connection: measured
0.116 s for 48 hosts.
This also puts a check under a caveat that had been documented in prose in
`Dockerfile.base` ("SSH client defaults") and verified nowhere: a per-host
`ControlPath ~/.ssh/cm/%r@%h:%p` inherited from a **read-only** bind-mounted
`~/.ssh` produces the identical failure, `cannot bind … Read-only file system`.
On the machine where this was built that is **16 of 50 hosts** — `freeipa-1..6`,
`gitea.egl.lan`, `runner-1..3`, `tor-ms22` and more — none of which had ever been
reported by anything.
**The two routes get different severities, deliberately.** Default `ssh`
precedence legitimately lands in the read-only `~/.ssh` on any host whose own
config pins it there, and the supported workaround (`ssh -F
~/.ssh-local/config`, generated every container start by `setup-lan-access.sh`)
already exists — so that is a `warn`. Making it a failure would paint the script
red on every run of every device, and **a check that fires benignly every time is
one you learn to ignore**, which is the same reasoning that keeps `lint-shell.sh`
at `-S error`. The sidecar route is the prescribed one, so there an unusable
directory is a hard `fail`. Host lists are capped at six names plus a count for
the same reason: unreadable output is ignored output.
Teeth proven in both directions, with each sabotage confirmed by `diff` *before*
the result was believed — a vacuous sabotage that silently fails to apply
reports "gate passed" and is worse than no test:
| Sabotage | Class | Result |
|---|---|---|
| `ControlPath` → nonexistent dir | the original hyphen bug | `✗` `rc=1` |
| `ControlPath` → existing but read-only dir | the `Dockerfile.base` caveat | `✗` `rc=1` |
| restored | — | `✓` `rc=0`, sidecar byte-identical |
No config was added to this repo to fix the original problem, because the fix was
to **delete** the offending file: `~/.pi/ssh/config` was a third hand-maintained
copy of what `setup-lan-access.sh` already generates from version control on
every start, with fewer features (no `known_hosts` sidecar, no
`StrictHostKeyChecking accept-new`) and one typo. The host-owned `~/.ssh/config`
is also left alone on purpose — those `~/.ssh/cm` paths are correct *on the host*,
where `~/.ssh` is writable, and the container-side override is the right layer.
---
**The size numbers on the Docker Hub page were the only claim in these docs with
nothing in the repo to check them against, and they had gone 20% wrong across
eight releases.** Every other claim `scripts/check-doc-drift.sh` guards is
anchored to a build file — a pin, an `ARG`, a placeholder — so it cannot rot
without someone editing the thing it describes. Nothing in this repo states the
image size, so `DOCKER_HUB.md`'s `~1.1 GB` simply drifted while the image grew,
and `update-description` POSTed it to Docker Hub every release. It is the first
number a stranger reads about this image.
Corrected against Docker Hub's measured `full_size`, 2026-09-14 after v1.9.2
published (amd64 / arm64, compressed):
| Row | Claimed | Measured | Now says |
|---|---|---|---|
| `:latest` | ~1.1 GB | 1.228 / 1.211 | ~1.23 GB |
| `:latest-studio` | ~1.15 GB | 1.255 / 1.238 | ~1.25 GB |
| `:base-latest`, `:base-<hash>` | ~1.0 GB | 1.167 / 1.151 | ~1.17 GB |
`full_size` is the right field because it tracks the **first manifest entry**
(amd64), not the sum across architectures — measured on v1.9.2:
`full_size=1.228`, `amd64=1.228`, `arm64=1.211`, `sum=2.439`. That matches the
table's per-arch "Size (compressed)" column.
**New check 8 in `scripts/check-doc-drift.sh`: size claims vs Hub's measured
`full_size`,** so this class cannot rot silently again. It fails on drift beyond
tolerance, and **skips loudly** — a new `skip()` helper, counted and named in the
summary — when `curl`/`python3` are missing, the API is unreachable, or
`SKIP_SIZE_CHECK=1`. Skips are deliberately neither `OK` nor a failure: printing
an unverified claim as OK is the habit this file exists to break, while failing
on Docker Hub's uptime would make every release hostage to a third party. Not in
`hooks/pre-push` (that runs `lint-shell.sh` only), so pushes do not hit the
network.
Two bugs were caught while building it, both by writing the expected exit code
down *before* running the check:
- **The tolerance would have missed its own motivating case.** The percentage is
computed against the *measured* size, but the 20% first chosen came from the
claim-relative figure. The real drift was `|1.1 − 1.37| / 1.37 = 19.7%` — it
would have passed. Now 15%, sitting inside a window whose bounds are both
measured: above the largest legitimate skew (a claim describing the published
release while the next tag changes the size — v1.9.1's 1.37 against v1.9.2's
1.23 = 11.4%) and below the rot it exists to catch (19.7%).
- **A `|| true` on the python invocation made the gate fail open.** It printed
`DRIFT` and exited 0 — a gate that reports the defect and passes anyway.
Removed; the outer `|| SIZE_RC=$?` is what satisfies `set -e` without
swallowing the code. Verified two-sided afterwards: a 19% drift exits 1 at the
default tolerance and 0 at `SIZE_TOLERANCE_PCT=25`, so the threshold is doing
the work rather than the ordering.
**Explicitly NOT covered:** `README.md`'s `~3.2 GB` figures are *uncompressed*
on-disk sizes, and the registry exposes compressed sizes only (manifest layer
sizes are compressed; the config blob carries no uncompressed totals). Measuring
them needs a real pull, so they remain unverified — a green check 8 says nothing
about them, and the script says so where a reader will see it.
**`promote-base-latest`'s conditional re-tag is now measured, which closes an
open question from the v1.9.2 release verification.** The job compares digests
and re-tags `base-latest` only if stale, and that short-circuit determines
whether a release watcher may assert *freshness* on the alias or merely
*existence*.
Measured on run 669: the digests differed (`want sha256:8c452575…`,
`have sha256:f34ad201…`), the job logged
`Promoting base-latest -> …:base-b5f2d03baae2`, `crane copy` ran
`17:10:04 → 17:10:08`, and Docker Hub's `base-latest.last_updated` moved to
`17:10`. So **a `crane copy` that runs does bump Hub's timestamp** — a manifest
re-tag is a tag write.
The identical-digest case remains *unproven by construction*: when `base-latest`
already resolves to the new `base-<hash>` the job prints
`base-latest already current; nothing to promote.` and never copies, so the
timestamp legitimately stays put. A freshness assertion would then report a
correct release as stale. The rule is therefore conditional on
`base-decide`'s `need_build` — strengthen it to fresh-required only when
`Dockerfile.base`, `rootfs/` or a folded `*_REF` has moved, and leave it as an
existence check for the variant-only and docs-only releases that cache-hit the
base. Operational guidance lives with the tooling that consumes it
(`ci-release-watcher` skill, skillset `c8034af`); recorded here because it is a
property of *this* pipeline.
Worth restating alongside it, because v1.9.2 proved it the useful way: tag
freshness is not the authoritative evidence that a base rebuild baked what you
expected. The `se.jordbo.pi-devbox.*-ref` labels are — readable straight from the
registry with no `docker` or `crane` (token → manifest index → per-arch manifest
→ config blob), and worth validating against the *previous* tag first, since a
reader that cannot show the old SHA cannot prove the new one.
---
## v1.9.2 — 2026-09-14
**The v1.9.1 residual is attributed and fixed: it was mostly npm's own download
cache, not the platform binaries it looked like.** v1.9.1 pruned the 26 foreign
`@esbuild/<platform>` directories that npm 11 installs, which fixed the 431 MB
size-gate failure — but the image still shipped **+131 MB compressed over
v1.8.14**, nearly all of it in the single pi/extensions install layer (87 MB →
206 MB). v1.9.1's notes recorded the leftover as an open item with an explicit
hypothesis (the `@mariozechner/clipboard-*` family, same npm 11 behaviour, a
different package) and an explicit warning that the hypothesis was **not** a
measured cause. It was measured on 2026-09-11, after recreating onto v1.9.1, and
the hypothesis accounted for only a sixth of it:
| item | v1.8.14 | v1.9.1 | delta |
|---|---|---|---|
| `/root/.npm/_cacache` — the build's npm download cache | 35.2 MB | 145.3 MB | **+110 MB** |
| `@mariozechner/clipboard-*` foreign platform packages (2 sites) | 0 MB | 21.1 MB | **+21 MB** |
| variant install layer, uncompressed total | 270.8 MB | 401.9 MB | +131 MB |
That is the whole delta with no unexplained remainder. Both items are now
deleted in the **same layer** that creates them, in both the main install RUN and
the studio RUN:
- **`purge_build_caches`** — `npm cache clean --force` plus `rm -rf /root/.npm`.
npm 11 caches every platform tarball it downloads, including the ones the
prune then deletes, so the cache grew far faster than the installed tree.
Nothing at runtime reads it: the build runs as root, the container runs as
`developer` with its own cache under `$HOME`.
- **`prune_foreign_esbuild` → `prune_foreign_natives`** — now covers both
measured families. For clipboard the keep-set is `clipboard-linux-$arch-gnu`
**and** `-musl`, because its napi-rs loader chooses between them at runtime
from its own `isMusl()` probe; the musl package is a 420-byte stub, so keeping
it is free insurance. The bare `@mariozechner/clipboard` wrapper has no
hyphen suffix and cannot match the pattern.
**Verified on arm64 before writing the patch, which is why the order was
update-then-patch:** a widened `rm -rf` glob is the worst possible change to
write against a tree you cannot inspect, and there is no docker CLI inside the
container — but after a recreate the container *is* the image. The prune was
exercised against a copy of the real trees with foreign directories fabricated
back in (aix-ppc64, android-arm64, darwin-arm64, win32-x64, linux-x64): all
removed, host `linux-arm64` kept at both sites, 21 MB freed, and
`require('@mariozechner/clipboard')` still loads and exports all 18 functions.
`esbuild.transformSync` still compiles TS at both sites. v1.9.1's own arm64
validation of the esbuild prune also passed here — CI could only smoke-test
amd64.
**Two sentinel assertions in `smoke-test.sh`, because the size gate did not
catch this.** The gate has ~225 MB of deliberate margin, so 131 MB of pure
build residue stayed green. There are now named PASS/FAIL checks for *foreign
platform packages beyond the host arch* and for *`/root/.npm` being shipped* —
the latter deliberately refuses to run as non-root, because `test ! -d
/root/.npm` on mode-700 `/root` would otherwise pass for the wrong reason. The
size-failure diagnostics now also list cache paths: they previously enumerated
only `node_modules` and `/opt`, where these bytes were not.
**Also fixed: the prune's own progress line was mangled.** `-printf '%f\\n'`
reaches the shell with both backslashes (confirmed from the published image's
recorded `created_by`), so `find` emitted a literal backslash and `tr` then ate
the `n` out of the name — v1.9.1 printed `esbuild platform dirs kept: li
ux-arm64`. Single backslash now.
Deliberately **not** changed: `/tmp/node-compile-cache` (1.3 MB). The manifest
RUN at the end of `Dockerfile.variant` calls `pi --version` again, so deleting
it earlier only relocates those bytes into that layer — today's manifest layer
is 128 kB precisely because it finds the cache warm.
**Two functional assertions as well, after the runbook command left for the
next machine failed for the wrong reason.** v1.9.1's open item prescribed
`node -e 'require("esbuild").transformSync(...)'` as the post-boot check, with
"if this fails, the prune removed something needed → revert to v1.8.14". Run
from `/workspace` it fails with `MODULE_NOT_FOUND` on a perfectly good image:
`require` resolves by walking up from the current directory, esbuild lives
nested inside the two pi trees, and global installs are not on node's require
path (`NODE_PATH` is unset). The check that verified the prune last time only
passed because the shell happened to be inside the tree. Smoke now does it
properly and CI owns it: for every install site found in the image (so the
studio variant's third site is covered automatically), esbuild must compile TS
and `@mariozechner/clipboard` must load with its native binding attached — the
latter is the real proof for the clipboard prune, since napi-rs resolves the
platform package at `require()` time. Both were verified as a four-way matrix:
green on the real image *from `/workspace`*, and red against copies of the same
packages with the host platform binary removed (`The package
"@esbuild/linux-arm64" could not be found`).
Touches `Dockerfile.variant` and `scripts/smoke-test.sh` only: `Dockerfile.base`
is unchanged, so this needs no base rebuild and should **ride the next release**
rather than burn a cycle of its own.
> **Superseded by the entry below:** that entry refreshes the vendored mempalace
> skill snapshot, which **is** hashed into `base_tag`. The release as a whole now
> costs a base rebuild (~67 min). The *size* work above still needs none of its
> own; the two simply travel together now.
**A behaviour change reached the fleet without any release naming it, and a
paragraph in these notes kept saying it had not.** v1.9.1 bakes
`mempalace-toolkit` **`e68ee20`**, which contains **`e2b060a`** — requester-side
ask withdrawal (`isWithdrawn`, RFC 003 §3.3 clause 4). So the behaviour has been
live on every v1.9.1 device since 2026-09-10, while the v1.9.0 section of this
file still read "not yet pinned … this image still pins `e45f6b4`" and the
mempalace skill still told every agent, at session start, that a withdrawal is
impossible: *"there is nothing anyone can do about it from the other end."*
The mechanism is the point, because it will do this again. `Dockerfile.variant`
carries `ARG MEMPALACE_TOOLKIT_REF=main` and `docker-publish.yml` resolves it to
a commit SHA at build time (`gitea_sha mempalace-toolkit`). A release therefore
absorbs *whatever toolkit `main` holds at that moment*, and "what behaviour did
this image gain?" is a question **nobody is structurally forced to answer**. This
fleet already has the rule — a floating ref that pulls a behaviour change into
the image must be named in the CHANGELOG *before* tagging. It was honoured for
the feed-tick fix, which v1.9.1 names explicitly (`309980b`, `e68ee20`), and
missed for the commit sitting in the same range.
Measured before being written, two independent routes, expectation recorded
first ("label should read ≥ `e68ee20`, since the build at 22:00Z postdates that
commit's 18:58Z"):
| route | result |
|---|---|
| Docker Hub config-blob label, `:v1.9.1-studio` and `:latest-studio` (same digest) | `se.jordbo.pi-devbox.mempalace-toolkit-ref = e68ee2071ca3ad39…` |
| `git merge-base --is-ancestor e2b060a e68ee20` | ancestor — the fix is inside the baked ref |
| baked `extensions/pi/mempalace.ts` sha256 vs v1.8.14's | `7c16fe14…` vs `dfca71e9…` — different bytes, so not the pre-fix file |
| `grep -c isWithdrawn` on this v1.8.14 container's baked copy | `0` — confirms the split, and that tor-ms22 cannot exercise it |
The first attempt at that label read **empty**, and the empty result was a claim
about the request, not the image: Docker Hub redirects blob fetches to a CDN and
`curl` without `-L` returns 0 bytes with exit 0. A registry audit that reports
"no labels" should be assumed to be missing `-L` until proven otherwise.
So the skill is updated rather than deferred (skillset **`e9e45f7`**, vendored
here with `scripts/vendor-mempalace-skill.sh`, `44472 → 46045 B`). Two things
were deliberate:
- **The "silence is not an answer" rule keeps its teeth.** An ask still stays
owed until the *recipient's* terminal event; what is new is a release by the
**asker**, explicitly marked. Stated that way round on purpose — the wrong
reading of this change is "withdrawals happen, so I need not reply".
- **The precondition ships with the rule**, because this skill is read on images
that lack the behaviour (tor-ms22, right now):
`grep -c isWithdrawn /opt/mempalace-toolkit/extensions/pi/mempalace.ts`, where
`0` means the withdrawal will not reach the recipient's mailbox. Same shape as
the provenance bullet's live-bridge check.
**The snapshot canary is re-pinned, and this time it fails on the old bytes
instead of merely failing to notice them.** The retired pair (`"Diaries
self-heal…"` present / `"Agent diaries live in"` absent) was still green against
the new snapshot, so it was blind to this refresh exactly as the pre-v1.8.13 pair
was blind to that one. The replacement is stronger than any predecessor here
because **both witnesses come from the same upstream commit**: `e9e45f7` added
`"Withdrawing an ask you sent"` and deleted `"nothing anyone can do about it from
the other end"`, the sentence the new bullet contradicts. Directions were
measured against both files rather than read off the diff (`new=1/old=0` and
`new=0/old=1`), then the canary body was **executed** against each: new → `rc=0
ok`, old → `rc=1` empty.
**`isWithdrawn` is no longer deployed-and-unproven — and the suite that pins it
had been dark since the node 24 bump.** The rule was exercised on the released
image against the live logstream on 2026-09-14, from a container recreated onto
v1.9.1 (born 13:21:41Z, confirmed by entrypoint-written mtimes and docker-written
`/etc` files agreeing to the second; `/proc/uptime` and `ps -o lstart=` were not
used, per their retraction). Baked toolkit `e68ee20`, `grep -c isWithdrawn` = 3,
mempalace.ts sha256 `7c16fe14…`, and the file pi actually loads verified to be
that same path and hash rather than a stale copy.
The instrument matters as much as the result: the shipped extractor was lifted
out of `scripts/test-owed-withdrawal.sh` and used to cut `TERMINAL_STATUS`,
`isStrictlyAfter`, `isAnswered` and `isWithdrawn` out of the **baked**
mempalace.ts by brace matching, then `deriveOwed`'s three queries were replayed
with their exact shipped parameters over the same `/mcp` transport the extension
uses. Shipped bytes, live data, no second copy of the logic. Baseline owed = 1,
agreed by three independent routes (the extension's own wake-up card; a hand
derivation of 23 candidates; the shipped predicates). Every expectation was
recorded before its measurement:
| probe | predicted | observed |
|---|---|---|
| positive control planted | owed 1 → 2 | 2 |
| requester withdraws it | back to 1 | 1, and `isAnswered=false isWithdrawn=true` |
| **third party** retracts someone else's ask | no effect | still owed |
| requester withdraws in **prose**, no marker | no effect | still owed |
| cleanup by ordinary replies | owed == baseline | 1, then 0 |
The count returning to baseline is only arithmetic; the per-predicate verdict is
what makes it a statement about mechanism. Probe A remained a raw candidate
throughout and no terminal reply of ours existed on its correlation, so its
removal is attributable to the withdrawal rule alone. Final attribution: A
cleared by `isWithdrawn` only, the two negative controls cleared by `isAnswered`
only — each probe retired by the predicate it was built to exercise. Both marker
spellings now have live witnesses (`metadata.withdraws` naming the ask's event id,
and naming its correlation). Unplanned and worth more than the probes: against
live data the rule also fires on the incident it was written for — seq 112 is
reported withdrawn-by-requester, i.e. mbp-m1-2020's seq 119 withdrawal now works,
so the 41h false obligation cannot recur. Full record in the coordination log at
`project/pi-devbox` seq 140; harness preserved as artifact
`art_20260914T141704_a7a9a294a4bb`.
Not proven, and deliberately not claimed: the extension's **own in-process
mailbox poll** surfacing a withdrawn ask. That poll fires at `agent_settled` when
the agent is idle; two probes were left owed across the rest of the session to
give it a window and it did not fire. Calling that confirmed would be a claim
about the session's patience, not about the code.
**Toolkit pickup `dab989b` — the owed-set suite could not run on this image, and
its own gate was why.** Reaching for the suite as corroboration exposed a second
defect: its precondition line `node --experimental-strip-types --check "$SRC"`
exits 2 on an unmodified mempalace.ts under node 24, and with `set -euo pipefail`
that skipped all 17 assertions and both regression guards. The cause is narrower
than the error suggests — it names the inline type-import, but `node --check` does
not type-strip at all: a file containing only `const x: number = 1` fails
identically, while executing the same file works. So the gate could never
validate TypeScript on any node; v1.9.1's bump from v22.23.2 to v24.21.0 is the
most likely trigger, though with no node 22 on the box that half stays labelled
inference rather than measurement. It failed **closed** — loud exit 2, never
vacuously green — which is the good direction, and the reason it went unnoticed is
that nothing in CI runs this suite.
That matters more than a red test, because this suite is the only thing that makes
`isWithdrawn`'s failure mode visible: a wrong rule there does not throw and does
not log, it makes a real unanswered ask vanish from a mailbox forever. `dab989b`
strips first and syntax-checks the emitted JS, and splits the exit codes so that
`3` means the gate cannot run while `2` means the source does not parse —
collapsing those is how the defect disguised itself as "mempalace.ts does not
parse" while mempalace.ts was fine. Verified in six directions with expectations
written first: clean source `rc=0` 17/17; malformed TypeScript `rc=2`; stripper
made unavailable `rc=3` without the misleading message; and the four mutation
kills back at e2b060a's counts of 3/1/2/1, so sensitivity is restored rather than
asserted.
> **Cost, and it is smaller than it looks:** `MEMPALACE_TOOLKIT_REF` is resolved
> by CI to the head of the toolkit's `main` and folded into the `base_tag` hash —
> verified at `docker-publish.yml:126-128`, whose own comment gives the reason
> ("otherwise a toolkit-only fix never lands"). So `dab989b` moves `base_tag` and
> the next tag rebuilds the base, with nothing to remember to trigger. But it adds
> no rebuild that was not already owed: `9aaff26` refreshed the vendored skill
> snapshot under `rootfs/`, which is also hashed into `base_tag`, so a base
> rebuild has been pending since before this fix existed. The toolkit pickup rides
> along with it, and the same rebuild is what finally bakes skillset `e9e45f7` and
> turns the snapshot canary above green against the image's own floor.
---
## v1.9.1 — 2026-09-10
**v1.9.0 was tagged but never published: its own smoke gate stopped it, and it
was right to.** `build-base` succeeded, then `smoke` failed 90-passed/3-failed,
and because `build-variant` needs `smoke`, both variants, `promote-base-latest`
and `update-description` were skipped. No image reached the registry, so
`latest` still pointed at v1.8.14. v1.9.1 carries everything listed under v1.9.0
below, plus the three fixes here. Two of the three failures were self-inflicted
by v1.9.0's own changes, and the third was a real regression that the Node bump
dragged in — which is the case for keeping the gate strict.
**Failure 1 — the image was 431 MB over its size threshold, and npm 11 was the
cause.** Node 22 → 24 brings npm 10 → 11, and npm 11 installs **every**
`@esbuild/<platform>` optional binary rather than only the one matching the host:
26 platform directories covering aix-ppc64, android, darwin, freebsd, netbsd,
openbsd, win32, s390x, riscv64 and more, none of which this image can execute.
Measured on pi-fork's dependency tree, same repo and same command:
| npm | packages | `node_modules` |
| --- | --- | --- |
| 10.9.8 | 136 | **165 MB** |
| 11.19.0 | 169 | **449 MB** |
The 165 MB figure reproduces exactly what v1.8.14 shipped, which is what
identified npm rather than the image as the variable. esbuild declares those
binaries with `os`/`cpu` constraints, but npm 11 ignores them — and also ignores
`--os`/`--cpu` flags and an `.npmrc` carrying `os=`/`cpu=` (all three measured,
all three still produced 26 directories). So `Dockerfile.variant` now prunes
explicitly, keeping only `linux-$(node -p process.arch)` so one line is correct
on amd64 and arm64. Verified this removes dead weight and not function: after
pruning, `esbuild.transformSync` still compiles TypeScript. The prune runs in the
**same layer** as each `npm install` — deleting in a later `RUN` would leave the
bytes in the earlier layer and shrink the image by nothing. Three sites are
covered: the global pi install, pi-fork, and pi-studio (which pulls its own
pi-coding-agent copy), for roughly 548 MB recovered in the non-studio variant and
822 MB in studio. The threshold stays at 3800 MB deliberately: it caught a real
regression, and raising it to accommodate one would have discarded the signal.
**Failure 2 — the om `node_modules` assertion was checking an npm artefact, not
the software.** `pi-observational-memory` declares **zero** runtime dependencies:
8 devDependencies (omitted by `--omit=dev`) and 4 peerDependencies, which pi
itself provides. npm 10 still materialised a `node_modules` for it, but that
directory contained exactly **one file** (`.package-lock.json`, 4 KB) and no
nested `package.json` — 20 empty scope directories. npm 11 stopped creating it,
so `test -d node_modules` went red while nothing about om had changed or broken.
The assertion now checks what must actually hold — that the entry point pi loads
exists — read out of the manifest pi itself reads (`package.json` →
`pi.extensions`) rather than a hardcoded path that could drift. pi-fork keeps its
`node_modules` check, because pi-fork has real dependencies where the directory's
absence would mean something.
**Failure 3 — the skill-source annotation broke the assertion that reads it.**
v1.9.0 taught `pi-devbox-version` to say *which* pi-extensions copy shipped
(`baked (package copy)`, or a loud FALLBACK/MIXED marker). The smoke assertion
matched `^ $s +baked$`, anchored at the end, so the annotation failed it even
though the state reported was correct. The pattern now allows an optional
` (...)` suffix, matched loosely on purpose: *which* copy shipped is already
asserted authoritatively against the manifest field and its measured tree hash,
and re-encoding that wording in a second regex would just add a second place to
update. The lesson recorded rather than the fix alone: the display branches were
tested in an isolated harness that passed, but the assertion **consuming** them
was never run — harness-passes-therefore-consumer-passes was an assumption.
**A red size assertion now carries its own diagnostic.** Attributing the 431 MB
took a full CI-log dig plus a local npm bisect, while the container knew where
its bytes were the whole time. On failure the check now prints the largest
layers, the largest directories, and a count of `@esbuild` platform directories
as a sentinel for this exact regression recurring — the same principle the `run()`
helper already applies to every other assertion.
No `Dockerfile.base` or `rootfs/` change, so the base fingerprint is untouched
and `base-decide` reuses `base-0fb1256c7f99` built during the v1.9.0 attempt.
**A gate for documentation drift, because five claims rotted in one release and
one of them was published.** Preparing v1.9.0 turned up a cluster of stale
facts, all the same shape — a value written once by hand, in a file nothing
verifies, about a number that lives somewhere else and moved:
- `README.md`'s "Version pins" table was wrong on **all three rows**: pi
`0.84.4` vs `ARG PI_VERSION=0.85.1`, pi-atelier `v0.10.0` vs `v0.10.1`,
mempalace `3.8.0` vs `3.9.0`. That table is the worst possible place for this,
because it exists *specifically* to be the reviewable record of what the repo
freezes deliberately — so a wrong row destroys the only thing it is for.
- `README.md` listed already-shipped typst PDF export under "Planned for an
upcoming minor release", carrying the self-contradicting marker "(shipped in
Unreleased/base)". The **fourth** instance of the stale-`Unreleased`-pointer
class this changelog already documented three of.
- `DOCKER_HUB.md` claimed "Node.js v22" while v1.9.0 ships Node 24.
The last one is why this became a gate rather than a resolution to be careful.
`DOCKER_HUB.md` is **published**: `update-description` POSTs it to Docker Hub as
`full_description` on every tag. It had gone **eight releases** (v1.8.6 →
v1.9.0) without a touch. Nothing generates it — CI only substitutes
`{{PI_VERSION}}` — and nothing checked it, so the sole mechanism keeping it true
was whoever remembered. Worse, it is read from the **tag**, so the stale page
published with v1.9.0 anyway and the fix could only ride the next release.
**New: `scripts/check-doc-drift.sh` + a `doc-drift` job in `lint.yml`.** Seven
checks, all comparing a doc string to a value that exists in this repo, so it
needs no network, no token, no built image, and no sibling clone:
- README's three pin-table rows vs the ARGs they name *by name*
- `DOCKER_HUB.md`'s Node claim vs `ARG NODE_VERSION`
- placeholders CI will not substitute — the publish step greps for leftovers of
`{{PI_VERSION}}` only, so any *second* token sails through and publishes
literally
- `DOCKER_HUB.md` under Docker Hub's 25 000-char `full_description` limit
(previously discoverable only as a non-200 *after* the full build)
- `Unreleased` appearing in a user-facing doc, which is always a pointer that
outlived what it pointed at
Exit codes match `lint-shell.sh` and `check-skill-floor.sh`: `0` in sync, `1`
drift, `2` cannot run — a renamed ARG makes the gate blind, which is a red `2`,
never a green tick. Verified with **15 controls**: every check fails when its
claim is broken, the real v1.9.0 Node bug is caught, and two false-positive
controls pass — the first version of the placeholder check wrongly flagged
`README.md:900`'s `docker inspect --format '{{json .Config.Labels}}'`, a Go
template in a legitimate example, so the pattern is now anchored to the
UPPER_SNAKE convention CI actually substitutes. **The gate was wrong, not the
doc** — which is the whole reason a gate gets negative controls.
Deliberately **not** gated, and the reasons matter more than the list:
- Counts and sizes (`~1.1 GB`, "N `mempalace_*` tools", "7 extensions") need a
running image. A gate that cannot evaluate a claim honestly would have to
guess, and a guessing gate is worse than none — assert these in
`scripts/smoke-test.sh`, where a real image exists.
- `Dockerfile.base`'s `# BASE_REBUILD_DATE:` marker, itself stale (2026-07-13,
three base rebuilds ago). `base_tag` hashes Dockerfile.base's *content*,
comments included, so demanding it be current would force a ~60 min base
rebuild on a release that touched no base files at all. It is free to fix
while the base is *already* rebuilding, and expensive at any other moment.
That cost asymmetry is now written into the release checklist rather than
enforced.
**Release checklist step 3 rewritten** (`AGENTS.md`) around the mechanism that
made this expensive: `docker-publish.yml` runs `actions/checkout@v4` with no
`ref:`, so every job reads `github.ref` — the tag. Docs must be correct *before*
tagging; afterwards the only routes are re-pointing the tag (its own hazard —
v1.8.14 went `601fc98` → `361babd` and broke deploy verification until
`git fetch --tags --force`) or waiting for the next release. The step now also
names what the gate cannot see, so "gate is green" is not mistaken for "docs are
true". The same reflex went into the `ci-release-watcher` skill, as the first
correctness rule — it is the only one that expires once the tag exists.
Also fixed in passing: README's `pi-devbox-version` sample was v1.5.0-era and
structurally outdated (it predated the `palace:` line the surrounding prose
advertises, the `pi-atelier` component, and the whole `skills:` block). Replaced
with real observed output rather than hand-written text. `DOCKER_HUB.md`'s "7
user-facing extensions" was **verified correct**; its "29 `mempalace_*` tools"
is stale (a live client shows 45) but left alone rather than corrected on a
guess, since that count cannot be attributed to the baked 3.9.0 server without
measuring it.
## v1.9.0 — 2026-09-10 (tagged, never published — superseded by v1.9.1)
> This tag exists in git but no image was ever pushed for it: `smoke` failed
> three assertions and skipped every downstream job. Everything below ships in
> **v1.9.1**, whose entry explains the three failures and their fixes. Kept as
> its own section rather than folded away, because the tag is real and someone
> will eventually find it and wonder why Docker Hub has no v1.9.0.
**`shellcheck` is now in the image, because the release gate it depends on could
not be run by anyone.** v1.8.14 made shell lint a release gate: `scripts/lint-shell.sh`
became the single source of truth for `lint.yml` and a new `lint-gate` job that
`resolve-versions` depends on, and it deliberately exits 2 when `shellcheck` is
absent — *a gate that cannot run must not pass*. Measured on v1.8.14 on
2026-09-09, by three routes (`command -v`, `dpkg -l`, a filesystem search):
**`shellcheck` was not in the image at all.** So `bash scripts/lint-shell.sh`
exited 2 in every devbox container, and the only place the gate could ever run
was CI. The developer loop was therefore write-shell → push → wait for CI →
discover — which is the loop the gate was added to shorten, after v1.8.14's first
attempt burned ~46 minutes on a tree whose lint had already been red for 24
hours. Added to the `apt-get` block in `Dockerfile.base`: `shellcheck 0.10.0-1`,
~39 MB installed (`Installed-Size` 40112 KB), and measured to pull **zero**
additional packages under `--no-install-recommends` because its three deps
(`libc6`, `libffi8`, `libgmp10`) are already present. **This forces one full base
rebuild** — `base-decide` hashes `Dockerfile.base` + `rootfs/`, so unlike a
`scripts/` change it cannot reuse the existing `base-` layer.
**A client-side pre-push lint gate: `hooks/pre-push`.** Opt-in per clone with
`git config core.hooksPath hooks`, bypass with `git push --no-verify`, matching
the idiom the `skillset` and `myconfigs` repos already use. It is a thin wrapper
that `exec`s `scripts/lint-shell.sh` — the same script CI runs, one copy, because
a duplicated check that drifts is the failure this repo keeps paying for (the
`pi-extensions` skill mirror sat 9579 B behind for weeks; the shell-lint logic
was extracted to one file for exactly this reason).
> **Why this repo had no hooks at all, which is worth stating because it was
> reported as drift and is not.** A fleet peer asked `tor-ms22` to report
> `git config core.hooksPath` per clone on the premise that an unset value meant
> "no secret-scan and no shell-lint hook locally", leaving the drift/secret gates
> unverified. Measured: `pi-devbox` **unset**, `skillset` `hooks`, `myconfigs`
> `common/hooks`, `pi-toolkit` **unset**. But `git ls-files | grep -i hook` is
> **empty** in both `pi-devbox` and `pi-toolkit` — neither repo tracked a single
> hook file, so there was nothing for `core.hooksPath` to point at on any machine
> and unset was the only correct value. The two repos that do ship hooks were
> already wired correctly. This entry closes the real half of that gap for
> `pi-devbox`; `pi-toolkit` still ships none.
**The hook is verified to catch the defect that motivated it, not merely to
exist.** Three measurements, each with the expected result written down first:
- **Refusal paths.** With `shellcheck` absent (the state of every container built
before this change) the hook exits **2** and names the remedy; with
`scripts/lint-shell.sh` missing it also exits **2**. It never waves a push
through on the assumption that CI will catch it.
- **The hook is actually in the scan set.** `lint-shell.sh` reports `Checking 14
shell file(s)` with `hooks/pre-push` present and **13** with it moved aside — so
the extensionless file is discovered by the shebang half of the linter's
two-signal union, rather than being silently skipped. This check exists because
the first attempt at it was ambiguous: a planted `[ $UNSET_VAR = "x" ]` was not
reported, which could equally have meant "file not scanned" or "defect below
`-S error`". It was the latter. A count that moves is unambiguous; a clean run
is not.
- **It catches the real v1.8.14 defect.** Planting the exact failing shape — an
apostrophe inside a single-quoted string, `echo 'the fleet\'s thing'` — in
`hooks/pre-push` produces `SC1073`/`SC1072` at severity **error** and `rc=1`.
That is the defect that closed a string, truncated an `exec_test` body, sent its
tail to the runner's shell, and cost a 46-minute build.
**The gate earned its keep inside the commit that added it.** The first version of
the `smoke-test.sh` assertion above carried a comment beginning `# shellcheck is a
GATE DEPENDENCY…`. A comment whose first word is the tool's name is parsed as a
**shellcheck directive**, not a comment, so the new gate immediately failed with
`SC1073`/`SC1072` at severity error — on the change that introduced it. Same family
as the v1.8.14 apostrophe: a line that reads as prose to a human and as syntax to
the parser. Before this change that defect would have been discovered in CI.
**Also queued, not yet pinned:** the `mempalace-toolkit` owed-set derivation now
honours a requester withdrawing its *own* ask (`isWithdrawn`, RFC 003 §3.3
clause 4, with `scripts/test-owed-withdrawal.sh`) — toolkit commit **`e2b060a`**,
which is the minimum revision for the behaviour. This image still pins `e45f6b4`.
Until an image bakes `e2b060a` or later, a sender must assume its withdrawal has
no effect on the recipient's mailbox — measured cost of the gap: a withdrawn
v1.8.13 rollout ask was still being reported as owed on `tor-ms22` 41 hours later,
for a release that device never installed. The same commit also anchors the
derivation's `mine` query at the newest end (`order: "desc"`); with the previous
default `asc` + `limit: 100`, a device passing 100 authored events would have its
recent replies fall out of the join window and see answered asks resurface.
> **Corrected 2026-09-14 (`pi@tor-ms22`, the device in the measured cost above).**
> "This image still pins `e45f6b4`" and "until an image bakes `e2b060a` or later"
> were true when written on 2026-09-09 and are **false for the running fleet**.
> **v1.9.1 bakes `e68ee20`**, a descendant of `e2b060a`, so requester-side
> withdrawal is LIVE wherever v1.9.1 runs. Measured from the published image's own
> label (`se.jordbo.pi-devbox.mempalace-toolkit-ref`) rather than from these
> notes, and cross-checked by ancestry and by the baked `mempalace.ts` sha
> differing from v1.8.14's. Nothing here was mis-stated on purpose:
> `ARG MEMPALACE_TOOLKIT_REF=main` is resolved to a commit SHA by CI at build
> time, so the release absorbed the commit without anybody having to name it,
> while this paragraph went on asserting it had not. Left standing rather than
> rewritten — the sentence is the evidence for how the drift happened. See
> Unreleased.
**Four small packages, each chosen from a gap that was measured rather than
imagined.** All four were picked by looking back at a real session — the
`gitea.egl.lan`/FreeIPA debugging of 2026-09-09..10 — and asking which absences
actually cost time, not which tools sound useful. `bind9-dnsutils` (~6.1 MB, 10
packages): `dig`, `host` **and** `nslookup` were all absent, so the container
could resolve names but had no way to interrogate a *specific* nameserver —
`getent hosts` only follows the resolver's default path, so diagnosing "gateway
`172.16.88.1` NXDOMAINs the `egl.lan` zone while `10.20.253.1` is authoritative
for it" had to be hand-rolled in `python3`. Note the package name: plain
`dnsutils` is transitional in trixie. `ldap-utils` (1244 KB, **zero** extra deps):
the fleet authenticates against FreeIPA, yet every LDAP probe had to be run by
SSHing to an already-enrolled host; this gives simple binds only, since GSSAPI
would additionally need `krb5-user` + `libsasl2-modules-gssapi-mit`, which is a
Kerberos-client decision rather than a tool. `xxd` (198 KB) is frank convenience
— `od -c` already does the job. `python3-yaml` (552 KB, zero extra deps) is the
shellcheck story repeating exactly: `scripts/check-workflow-shell.sh`, the guard
against the Gitea `sh`/dash footgun that broke `resolve-versions` (`ed49b8d`) and
`promote-base-latest` (`b7197e8`), hard-exits with "python3 yaml module missing"
without it — and `lint.yml` installing it explicitly in CI was the evidence the
image lacked it. **`netcat-openbsd` was proposed and deliberately rejected**:
measured redundant, because `socat` is already baked and bash's `/dev/tcp` does
reachability checks with zero packages. The reason is recorded in
`Dockerfile.base` so the omission reads as a decision rather than an oversight.
**The vendored `pi-extensions` skill floor was 41 days stale, and is now gated so
it cannot silently rot again.** `rootfs/usr/local/share/pi-devbox/skills/pi-extensions/`
sat at 34284 B, untouched since `fa04d20` (2026-07-30), while the package copy
was 38973 B — four copies of one skill existed across the fleet with three
different sizes. `Dockerfile.variant` copies the freshly-cloned package copy over
the **served** path but never writes back to the repo floor, so nothing in the
repo ever noticed. That is worse than ordinary staleness because the floor is a
**fallback**: the copy is guarded by `if [ -f /opt/pi-extensions/skill/SKILL.md ]`,
so a build whose clone yields no `skill/` keeps the vendored snapshot and still
goes **green**, with no manifest flag and no label recording which copy was
served — the image would ship a July skill and nothing would say so. The floor is
refreshed here from `pi-extensions@c64c122`, and the new `skill-floor` job in
`lint.yml` runs `scripts/check-skill-floor.sh` to keep it that way.
The check compares the **directory** hash, using the same `tree_sha256` pipeline
`Dockerfile.variant` uses for `skillset_snapshot_tree_sha256` and for the same
documented reason: a `sha256sum SKILL.md` answers "did this one file change", not
"is this the same skill", and `pi-extensions` ships two files. That is not
hypothetical — it was **verified by negative control**: with `SKILL.md` left
byte-identical and only `evaluate-extension-usage.py` edited, the directory check
correctly fails while a file-only compare would have passed. Exit codes are `0`
in sync / `1` drift / `2` cannot-run, matching `scripts/lint-shell.sh`, so an
unreachable package repo is a red `2` rather than a green tick. Gating on another
repo is normally a smell; it is proportionate here because the check can only
fire when `skill/` itself changed — which is exactly when the floor has gone
stale — and it needs no secret, since `pi-extensions` is anonymously clonable
(verified with `git ls-remote` and no credentials).
> **What this does *not* fix, stated so nobody reads more into it than is there.**
> The floor is now fresh and guarded, but the *silent-fallback* half remains:
> if the build-time copy is ever absent, the build still succeeds with no
> manifest flag or OCI label recording that the vendored snapshot was served
> instead of the package copy. The durable fix for that is a manifest field
> alongside the existing `skillset_snapshot_tree_sha256`, which this change does
> not add.
**Four pinned dependencies bumped, after an audit of everything the image gets
from outside apt.** The audit itself is the useful part: of ~23 externally-managed
components, the 19 that resolve `latest` at build time were already current or
refresh themselves on the next rebuild, and the hard pins for `pi` (0.85.1),
`mempalace` (3.9.0) and `pi-atelier` (v0.10.1) were all already the newest
available. Only four needed a human.
**`NODE_VERSION` 22 → 24 (LTS "Krypton") — this one was a latent defect, not
housekeeping.** `agent-browser` publishes `engines.node ">=24.0.0"`, so the image
was *below a declared requirement*: v1.8.14 shipped node 22.23.2 with
`agent-browser` 0.37.1, meaning every build installed it with an npm `EBADENGINE`
warning and then ran the baked browser automation outside its supported range.
The other two npm consumers are satisfied either way — `pi` declares `>=22.19.0`,
`playwright` `>=20`. Verified before bumping, because a missing NodeSource suite
would break every architecture at once: `setup_24.x` returns HTTP 200 and the
`node_24.x` suite advertises `Architectures: amd64 arm64 armhf x86_64`, covering
both the arm64 fleet and the amd64 CI runners. Nothing else in the repo pinned the
node major.
**`actionlint` 1.7.7 → 1.7.12 and `hadolint` 2.14.0 → 2.15.1**, each run against
the current tree at the new version *before* being pinned — both clean, no new
findings. That ordering is the point: a linter bump is the one dependency update
that can turn CI red on unchanged code, so discovering it locally costs a minute
and discovering it in CI costs a round trip.
**`SKILLSET_SNAPSHOT_REF` `e9e09d9` → `4d7c0ea`**, via
`scripts/vendor-mempalace-skill.sh` rather than by hand, because that script is
the only thing that may write the ARG — a `cp` without a matching bump produces a
manifest that confidently lies. This turned out to be **provenance-only**: the
recorded ref was 6 commits behind, but `skills/mempalace/SKILL.md` is byte-identical
at both (`3675bfab…`), so the vendored snapshot was already correct and only its
recorded origin was stale. Consequently no `rootfs/` bytes changed, the
smoke-test phrase canary stays valid, and this ARG alone would not have forced a
base rebuild — the node bump does that anyway.
**The silent-fallback hole is closed: the image now records WHICH `pi-extensions`
skill copy it shipped.** This was the half deliberately left open by the
`skill-floor` gate above, and it is the more important half, because "the floor is
currently fresh" is a fact with a shelf life while "the image says which copy it
got" keeps working. The refresh step in `Dockerfile.variant` is guarded by
`if [ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates the
co-located skill kept the vendored floor and still succeeded **green**, with
nothing in the manifest, the labels or the logs distinguishing that from a normal
build. The two outcomes are indistinguishable by inspection afterwards — same
path, same filenames, same permissions — which is exactly how the floor went
unnoticed from 2026-07-30 to 2026-09-10.
`build-manifest.json` gains `pi_extensions_skill_source` and
`pi_extensions_skill_tree_sha256`, both **measured rather than passed in as
build-args**, per the ground-truth rule the rest of that block already follows —
and necessarily so here, since the outcome depends on the clone's contents and no
ARG could express it. Three values, because two would force a lie:
`package` (served bytes equal the clone's `skill/`), `vendored-floor` (the clone
had no `skill/` at this ref, so the fallback shipped), and `divergent` — both
exist but differ, e.g. the clone ships `SKILL.md` but not
`evaluate-extension-usage.py`, leaving the served directory a genuine **mix** of
package and floor. No OCI label mirrors these, deliberately: `LABEL` cannot take a
value computed in a `RUN`, and a label fed from an ARG would be precisely the
claim-not-measurement this change exists to remove.
Two `scripts/smoke-test.sh` assertions turn the record into a gate: one that the
source is named and is `package` — `vendored-floor` **fails** rather than warns,
since these images track `main` where the package has co-located `skill/` since
`fa04d20`, so a fallback means the clone did not resolve as intended — and one
that recomputes the tree hash over the served directory, because a recorded hash
that is never recompared is a claim rather than a measurement. `pi-devbox-version`
also annotates the line: `pi-extensions baked (package copy)` on the normal path,
and a yellow `(FALLBACK: vendored floor — clone had no skill/)` otherwise. Its
existing skill section reports which copy is being **read** at runtime; this is
the one fact that is decided at **build** time and cannot be recovered later.
Older images degrade cleanly — the field is absent, `jq // empty` yields nothing,
and the line prints plain `baked` exactly as before.
**This image also carries a real fix for the recurring
`[mempalace ext] feed (tick) failed: mine timed out after 30000ms` message** that
has been appearing in the pi TUI across the fleet since August
(`mempalace-toolkit` `309980b` + `e68ee20`, picked up because CI resolves
`MEMPALACE_TOOLKIT_REF` to a commit SHA at build time). It was parked as cosmetic
on 2026-08-27 and it was not cosmetic: `lastFeedAt` was recorded only after a
*successful* wait, but the extension's `Promise.race` abandons only the **wait**
and cannot cancel the mine, so a timeout left the 10-minute debounce clock stale
— and with `feedInFlight` already cleared, **both** guards stood open and every
following settled turn started another mine on top of the one still running.
Overlapping writers on a single-writer palace, each making the next slower and
the next timeout likelier, which is why the message appeared many times per
session instead of at most once per debounce window. Simulated over ten minutes
of settled turns with a 60s mine: **16 mines launched, 15 of them overlapping**
before; **2 and 0** after. Nothing was ever lost — the transcript is staged
before the mine and `mine --mode convos` is idempotent — so this was wasted work
and a misleading error, not data loss. The deadline also rose from 30s to 5
minutes: the mine is the slowest call the extension makes (30–60s normally) yet
carried the tightest deadline, 4x tighter than the `prepare` before it and 10x
tighter than the init handshake. On a healthy fleet the message should now be
absent; if it appears it is informative — a mine exceeding five minutes.
---
## v1.8.14 — 2026-09-08
> **First release attempt failed; fixed in this same entry.** The `smoke` and
> `smoke-studio` jobs both failed at `scripts/smoke-test.sh:770` with
> `agent-browser: command not found`, after `build-base` had already succeeded
> (~46 min spent). Root cause was in the agent-browser execution guard added the
> day before: the explanatory comment inside the **single-quoted** `exec_test`
> body contained an apostrophe (`the fleet\'s`). Inside `'...'` bash treats a
> backslash literally, so `\'` does not escape — it **closes the string**. The
> body silently truncated (measured: `exec_test` received **12** arguments
> instead of 2), and the remaining lines, including the `agent-browser --version`
> assertion, were parsed by the **runner's** shell instead of executing inside
> the image — and the runner has no agent-browser. The prose now lives above the
> call, where an apostrophe is harmless.
>
> **The lint job had already caught this, and it went unread for 24 hours.**
> `shellcheck` flagged it as `SC2289` at severity *error*, so the `actionlint`
> job went red at run 186 on 2026-09-07 21:21 — the exact push that introduced
> the guard — and stayed red for runs 187 and 188. `lint.yml` deliberately
> excludes tag pushes (documented: the tagged tree was already linted on main,
> and a tag-ref lint run would sort above the publish run), which is sound; the
> broken assumption was different, namely that a tree whose lint FAILED would not
> then be released. `docker-publish.yml` has no dependency on lint, so it built
> for 50 minutes on a tree known to be defective.
>
> **Fixed, then gated.** The prose moved above the `exec_test` call so an
> apostrophe cannot terminate anything, and the shell-lint logic moved out of
> `lint.yml` into **`scripts/lint-shell.sh`** — now called by both `lint.yml` and
> a new `lint-gate` job here that `resolve-versions` depends on. A release with a
> lint error refuses in ~40 s instead of failing after fifty minutes. One copy,
> not two: a duplicated check that drifts is the failure this repo keeps paying
> for. The script also refuses to pass when `shellcheck` is absent, inheriting
> the existing principle that a gate which cannot run must not pass.
>
> **`v1.8.14` was re-pointed** from `601fc98` to the fix commit. Nothing had
> consumed the original tag — no `v1.8.14` image was ever published, only the
> content-addressed `base-a365dd24de21`. `scripts/` does not feed the base hash,
> so the re-run reuses that base and skips the 46-minute rebuild.
**A test that was quietly checking nothing, and a version number that was wrong.**
Both found by delegating a read-only audit of this repo to a headless worker
(`pi-toolkit` `bin/pi-task`) and then spot-checking its pointers from the
filesystem — 5 of 5 held, and it also corrected a false premise planted in its
own brief.
**The node major is now asserted, not merely printed.**
`scripts/smoke-test.sh` ran `run "node" "node --version"`, which asserts only
that the binary exists and exits 0 — the printed version was compared to
nothing. The line above it has always used `run_expect` against
`$EXPECTED_PI_VERSION` for `pi`, so the suite *looked* like it covered node.
**A node major bump would have passed the whole smoke suite silently.** Worse,
this is where the "node v22.23.2 verified" line in the v1.8.13 recreate notes
came from: printed output, not an assertion — an expectation stated up front and
then falsified by the check.
Now gated on `EXPECTED_NODE_MAJOR`, which CI derives from `Dockerfile.base`'s
`ARG NODE_VERSION` — the single source of truth, and the *only* hard node pin in
the repo (`Dockerfile.variant` has no node install at all, so the two Dockerfiles
cannot disagree). That also catches a stale cached layer whose node disagrees
with the declared ARG. Unset ⇒ previous behaviour, so nothing breaks for anyone
running the suite by hand.
Verified two-sided, because a silent failure here reintroduces the exact bug it
fixes: the `sed` derivation yields `22` (an empty result would disable the
assertion silently); `grep -Fq "v22."` matches `v22.23.2`; `"v24."` does **not**
match, so a wrong major is caught; `"v2."` does not prefix-collide. The workflow
YAML was re-parsed after editing (9 jobs).
**v1.8.13's agent-browser version was wrong.** That entry said "the image's own
0.35.2". The image ships **0.36.0** — `/usr/lib/node_modules/agent-browser` at
0.36.0 with `engines.node >=24.0.0`, and no 0.35.2 exists anywhere in the image.
The sentence was also internally incoherent, contrasting 0.36.0 against a version
that is not present. Corrected in place with a visible note, since that entry is
already released. **The reasoning survives untouched**: the engines floor really
is vestigial, because `/usr/bin/agent-browser` is a prebuilt aarch64 ELF invoked
directly and never through node — which is exactly why 0.36.0 runs fine on
22.23.2, consistent with the runtime proof collected on 2026-09-07 and with the
retraction of the earlier false "0.36.0 requires node >= 24" alert.
No image content changes: `NODE_VERSION` still 22, no pins moved. This is a test
and a docs correction only.
**The same bug class, twice in one file — and the second one was throwing away a
proof the fleet cannot obtain any other way.** `scripts/smoke-test.sh`'s
agent-browser guard captured the version *inside an `echo`, with `2>/dev/null`*:
```sh
echo "resolved=[$r] version=[$(agent-browser --version 2>/dev/null | head -n1)]" >&2
```
The exit code was discarded, so a binary that could not execute at all still
**passed**, printing `version=[]`. Verified two-sided: a stub exiting 127 passes
the old form and is caught by the new one.
Why that exit code matters more than most: smoke runs `platforms: linux/amd64` on
an x86 runner, i.e. **native amd64**, making this line the fleet's only recurring
amd64 runtime proof for agent-browser's `linux-x64` ELF. **No devbox can ever
supply one** — every machine in the pi fleet is an Apple Silicon Mac
(`mbp-m1-2020`; `tor-ms22` = Mac Studio `Mac13,1` M1 Max, verified 2026-08-17 by
`system_profiler`; `emb-7kj4vr4g` = Apple Silicon, verified 4 ways 2026-09-07).
The "amd64 runtime proof still needed" item that was sent to two devices was
therefore asking for the impossible, while CI already had the answer and was
discarding it. `Dockerfile.base:607` does assert it (`agent-browser --version &&`),
but only when the base actually rebuilds — and v1.8.13's base was cached.
**The mailbox now announces replies that CLOSE your own asks.** `mempalace-toolkit`
`21023e7` → `e45f6b4`, which adds `deriveClosed()` alongside `deriveOwed()`. The old
path queried `status: open` and joined for a reply, which by construction can only
surface asks *you owe someone else*; a terminal reply carries `status: applied`
(or `blocked`/`failed`), so **the answer to your own question was structurally
invisible** — the one notification a human actually wants. Measured: `emb-7kj4vr4g`
closed the v1.8.13 rollout ask at 18:31Z with `status=applied`, the operator
reasonably expected to hear about it, and the mailbox stayed silent while being
correct by its own definition. Nine closed correlations were sitting unannounced.
Shares the 1-hour resurface floor, so a close is announced once and is news rather
than a nag.
This lands **because the base rebuilds**, which is worth stating explicitly: the
CI-resolved `mempalace-toolkit` SHA is folded into the content-addressed base tag
(`base-decide`), precisely so a toolkit-only fix cannot silently fail to land
behind an unchanged `Dockerfile.base`. The floating `main` ref was left alone on
purpose — the toolkit moving *forces* the rebuild rather than waiting for one.
**A consequence worth noting for the amd64 item above: this release actually
collects that proof.** v1.8.13's base was cached, which is why
`Dockerfile.base:607`'s `agent-browser --version &&` never ran. v1.8.14's base is
not cached, so both that assertion and the new `EXPECTED_NODE_MAJOR` gate execute
on a native `linux/amd64` runner. The fleet's first *kept* amd64 runtime proof for
the `linux-x64` ELF should be an artefact of this build rather than something
asked of a device that cannot supply it.
**Subtask delegation is documented — including the rung nobody built.** The image
picks these up through their resolved refs (`pi-toolkit` `adfb553`,
`pi-extensions` `c64c122`, the latter also refreshing the baked fallback skill):
- an operator-facing decision guide in pi-toolkit's `README.md`, built on the
L0–L4 context ladder — how much of the parent session a child can see is the
axis that explains nearly every observed good and bad behaviour;
- the canonical `pi-extensions` skill gains the same ladder next to *Boundary
discipline*, which until now diagnosed why an inherited transcript defeats a
brief without offering any alternative to "don't fork that";
- one bullet in the global `AGENTS.md`, so the choice is visible without loading
a skill, and naming `pi-task` as a **CLI** — an agent hunting for a `pi_task`
tool finds none and concludes it is unavailable.
What the ladder records: **L0/L1/L2 exist** in `pi-task` (`context.facts` /
`.files` / `.commands`), **L4** is `fork`'s only behaviour (`getHeader()` +
`getBranch()`, no offset or limit anywhere in the call chain), and **L3** — a
truncated branch — **is not implemented by anything**, which is now written down
instead of being a design idea somebody remembers.
Also recorded, found while writing the above: `pi-fork/src/runner.ts:188` reads
`if (extensions !== null) args.push("--no-extensions")`. So `extensions: []` turns
the capability floor **on** and `null` turns it **off** — and `null` is the
documented way to "restore normal extension loading", so tidying `[]` to `null`
as a no-op re-arms palace writes inside every fork child. `pi-task` hardcodes the
flag and cannot drift this way. Documented in three places because the edit that
triggers it looks harmless.
---
## v1.8.13 — 2026-09-06
**Version audit + three pins moved, one deliberately not moved.** `pi`
0.84.4 -> 0.85.1, `mempalace` 3.8.0 -> 3.9.0, `pi-atelier` v0.10.0 -> v0.10.1.
`PI_FORK_REF=master` stays floating and therefore adopts e69725c. Each rationale
is written at the ARG itself rather than only here, because that is where the
next person doing the audit will be standing.
**Correction, made mid-release while run 639 was building:** the audit
originally recorded a fourth change — "`PI_STUDIO_VERSION` relabelled `none` ->
`v0.9.60-rc.0`, RC adopted deliberately" — and that was wrong. It was measured
at the wrong layer. `resolve-versions` passes BOTH `PI_STUDIO_REF` and
`PI_STUDIO_VERSION` as build-args and selects the newest **stable** semver tag
(its filter `^v?[0-9]+\.[0-9]+\.[0-9]+$` excludes pre-releases), so a Dockerfile
default cannot answer "what will CI publish?". Measured from the run itself:
`studio_tag=v0.9.59`, `studio_ref=9eed84f` (= `refs/tags/v0.9.59^{}`), while
`main`/`v0.9.60-rc.0` is 658536f and is not built. **Published v1.8.13 studio
images therefore contain pi-studio v0.9.59, not the RC**, and the ARG is back at
`none` rather than pinned to a pre-release that goes stale the moment main
moves. Consequence kept deliberately: the RC's opt-in Studio network binding is
absent from every published v1.8.13 image, so it needs no audit for this
release. Adopting an RC from CI would require changing that tag filter, which
exists on purpose — upstream stopped publishing Releases at v0.5.55 but keeps
tagging and pushing to main, so pinning main risked baking half-finished commits.
0.85.0 is SKIPPED on purpose: it shipped internal experimental code and extra
subpaths that broke SDK imports (upstream #9132), and 0.85.1 exists to undo
exactly that. Neither release has a Breaking/Removed changelog heading, the
engine floor is unchanged (>=22.19.0 against the container's 22.23.2), and
runtime deps drop 20 -> 19.
The pi bump was verified by RUNNING it, not by reading about it, because this
repo has already been burned by a version pair that no changelog flagged
(pi-atelier < 0.7.1 hangs pi >= 0.84 at startup with no error). 0.85.1 was
side-installed and driven under a pty in five combinations — each companion
extension plus atelier v0.10.0 AND v0.10.1 — with a CPU delta of 0.00-0.01s
over a 5s window where the known hang signature is ~5s of sustained CPU. The
check was two-sided: the atelier sidebar painted ACTIVITY+WORKSPACE markers
identically to the 0.84.4 control, so "alive" could be distinguished from
"silently absent".
**NODE_VERSION stays 22 — audited, not overlooked.** node 24 is technically
safe: all five prebuilt native addons in pi use NAPI (ABI-stable, no
NODE_MODULE_VERSION lock, no binding.gyp), nothing in the image declares a node
CEILING, and the install is one token (`setup_${NODE_VERSION}.x`). agent-browser
0.36.0 declares `engines.node >=24.0.0`, but that field is vestigial for the
artifact actually shipped: `/usr/bin/agent-browser` is the prebuilt aarch64 ELF
`bin/agent-browser-linux-arm64`, invoked directly and never through node, so npm's
engines floor is never enforced at runtime — verified running under 22.23.2 in
this image. (Corrected 2026-09-07: this paragraph originally said "the image's own
0.35.2 declares the same floor". That was wrong and incoherent — it contrasted
0.36.0 against a 0.35.2 that does not exist in the image. There is exactly one
agent-browser present, `/usr/lib/node_modules/agent-browser` at 0.36.0. The
argument is unaffected; only the version was wrong.) The reason to wait is
attribution, not compatibility — this release already moves pi a minor,
mempalace a minor and bakes a Studio RC, so adding a node major would leave four
suspects if the image misbehaves. Worth doing as its own release with the smoke
suite as the gate. (v22 is in maintenance until 2027-04-30; v24 is Active LTS
to 2026-10-20 and maintained to 2028-04-30, so there is real headroom.)
mempalace's client bump carries a sequencing note that is now also CORRECT: the
comment at the ARG claimed synlig serves 3.7.1 server-side, which was stale.
Measured 2026-09-06 over ssh, synlig's uv tool entry last changed 2026-08-25
and serves 3.8.0. Client 3.9.0 against server 3.8.0 is accepted skew until
synlig's compose stack is redeployed; 3.9.0's headline additions (release
awareness, `task create`/`task launch`) are SERVER-side and stay dark until
then — a client bump alone cannot light them up.
**agent-browser was running 7 weeks stale, and the interesting part is why
nothing noticed.** The image has shipped 0.35.2 since the last base rebuild,
but every session on mbp-m1-2020 was executing 0.27.0 from a 2026-07-17
hand-install: `npm i -g` writes into `~/.pi/npm-global`, which is the
devbox-pi-config VOLUME, and PATH puts that at position 2 against /usr/bin at
position 8. This is the third package hit by that exact hazard (pi itself and
pi-atelier already have guards), so the guard is now generalised instead of
re-invented a fourth time.
The damage was not the binary. It was the BUNDLED SKILL, which is the part an
agent reads: 3 skillsets / 17.6 KB core in 0.27.0 versus 8 skillsets / 31.5 KB
core in 0.35.2, with ten subcommands present in the image and entirely
undocumented to the agent (a11y, browser, data, mcp, page, plugin, read,
selectors, to, webmcp). A stale tool announces itself with an error; a stale
skill just quietly teaches the wrong commands and everything looks fine.
Three changes, at the three places this can be caught:
- `entrypoint-user.sh` retires a volume copy by MOVING it aside (reversible,
same instinct as the settings backups) and only when the image ships its own
copy, so a machine that deliberately hand-installs on an image without one
keeps it. The `bin/` shim is removed too — a dangling symlink would be a
worse failure than a stale version.
- `scripts/recreate-sanity-check.sh` asserts `agent-browser` resolves under
/usr. This is the check that matters, because it runs where the volume is
real.
- `scripts/smoke-test.sh` gets the build-time half, labelled WEAK in the source
for an honest reason: a `docker run` container has an empty config volume, so
it can never see the shadowing it is nominally testing for.
**pi-fork gets a capability floor: `extensions: []`.** Forks were measured
twice (2026-09-01, 2026-09-06, four dispatches) ignoring their brief, answering
in the USER's voice, fabricating self-referential measurements, and once filing
a diary entry as `agent_name=pi` — which landed in `wing_pi`, where a
wing-scoped `diary_read` never sees it.
The cause is upstream and by design, so there is nothing to wait for: the child
is handed `getHeader()+getBranch()`, i.e. the WHOLE active session branch, with
the brief appended as the final user message and the system prompt untouched
(pi-fork `src/index.ts`). In a long session the parent narrative simply
outweighs the task, and the child does the statistically obvious thing — it
continues the story it finds itself inside. Config offers no context knob
(extensions, environment, offline, costFooter, effort profiles only).
Falsified the tempting explanation before acting on it: the failures are NOT a
too-small model. The same model as the `fast` profile (haiku, thinking off)
obeyed the identical brief perfectly when run as
`pi -p --mode json --session-id <fresh> --no-extensions` — correct values,
exact format, no session recap, 3 seconds, $0.012. Model held constant, context
inheritance removed, failure gone.
`extensions: []` is therefore a mechanical guarantee rather than an
instruction: the mempalace bridge is a pi EXTENSION, so a fork child now runs
with `--no-extensions` and cannot write to the shared palace under the parent's
identity. Verified by asking a child to enumerate its own tools: `read, bash,
edit, write` — no `mempalace_*`, no `recall`, no nested `fork`. Two honest
limits, stated so nobody over-trusts this: it removes PALACE writes, not
FILESYSTEM writes (`edit`/`write` remain), and it costs forks their palace
search and recall. Set the key to `null` to restore normal loading.
Smoke asserts the floor is `[]` specifically, not merely falsy — `null` is the
unguarded state, so a "truthy or not" test would pass on exactly the
configuration being guarded against.
**Vendored mempalace skill snapshot refreshed `a12fe5e` -> `e9e09d9`, and the
phrase canary re-pinned with it.** Folded in at zero marginal cost: the
snapshot is hashed into `base_tag`, but `Dockerfile.base` already changed this
release, so the ~67 min base rebuild was already being paid. `--check` reported
exit 0 (stale-but-truthful) beforehand, i.e. skipping was sanctioned — this is
the deliberate decision the checklist asks for, not a drive-by. Upstream content
is the fleet wing-naming convention (bare project names, no `wing_` prefix) and
the `<harness>@<device>` rule for `added_by`, both of which came out of the
attribution defect measured on this device on 2026-09-06.
The canary re-pin is the interesting half. Its old pair — "Provenance is
stamped for you" present, "Attribute what you file yourself" absent — STILL
PASSED against the new snapshot, so leaving it in place would have produced a
canary that is green on both the old and the new bytes: blind to precisely the
refresh it exists to witness, which is the same false-green family the
pre-v1.8.5 canary died of. The replacement pair was picked by MEASURING
direction against both files rather than by reading the diff ("Diaries
self-heal; plain drawers do not" new=1/old=0; "Agent diaries live in"
new=0/old=1) and then tested two-sided: PASS on the refreshed bytes, FAIL on the
old bytes recovered from git. A canary that cannot fail is decoration.
**`credential-incident-response` §5/§6 corrected — a stated mechanism was wrong,
and this is the second time in three days this section named a wrong reason
for a zero.** Docs only.
§5 said `embedding_metadata.string_value` holds "metadata fields only". Measured
false on chroma 1.5.9 with a disposable sentinel drawer (pi@tor-ms22,
2026-08-30): the document text is ALSO there, under key `chroma:document` — one
row in `fts_content` and one in `embedding_metadata` for the same drawer. The
scan order in §5 is unchanged (scan `fts_content` directly, raw bytes as
backstop) but the stated REASON is fixed: a zero from `string_value` needs a
different explanation (key filter, query shape, escaping), not "it's
structurally blind". §6 already warns against explaining a zero with an
unverified mechanism; this was exactly that failure, in the file that carries
the warning.
§6's row-gone/bytes-gone claim is now backed by the same sentinel measurement
rather than asserted: `delete_by_source` took both `fts_content` (1->0) and
`embedding_metadata` (1->0) to zero, while raw bytes stayed 4->4 until VACUUM.
Also records how the measurement got unblocked at all — not a better
instrument, a disposable sentinel drawer instead of testing deletion on real
data.
---
## v1.8.12 — 2026-08-31
**`pi` `0.84.3` → `0.84.4`, and `pi-atelier` `v0.8.2` → `v0.10.0`.** Both audited
by the routine in `Dockerfile.variant` rather than adopted on sight, and the
audit notes live next to the pins where the next reader will meet them.
**pi 0.84.4 (published 2026-08-28) carries no `Breaking Changes` and no
`Removed` heading** — checked by grepping the section, 0 matches, which is worth
stating because 0.84.3 *did* have one. It was adopted for three fixes that land
on machinery this fleet runs every day, not for the feature list:
- **#6879** — a large tool result crossing the auto-compaction threshold used to
be sent to the provider *before* compaction. Pi now compacts between tool
execution and the next assistant response inside the same run. That is the
shape of nearly every session on these boxes, where a single `event_list` or
palace search returns hundreds of KB.
- **#8345** — a resumed session corrupted its next appended entry when the JSONL
file lacked a trailing newline. That file is the memory feeder's *input*, so
the failure would have surfaced as unexplained gaps in `wing_conversations`
rather than as an error. Measured on tor-ms22 before bumping: 49/49
transcripts end in a newline and 0 lines fail `json.loads` — this corpus was
never bitten, and we now know that rather than hope it.
- **#8537** — extension messages sent with `triggerTurn: false` *while the agent
is running* were inserted between a tool call and its result, so
order-validating providers rejected the replayed history. **The mempalace
mailbox is outside that precondition**: it delivers at `agent_settled`, when
no inference is in flight, with `{deliverAs: "steer"}` and deliberately no
`triggerTurn`. 0.84.4 also leaves the documented steer semantics untouched
("delivered after the current assistant turn finishes executing its tool
calls, before the next LLM call"), so RFC 003 §7.11 stands as written. Recorded
because this fix is precisely what would make a *mid-run* delivery safe, which
is the only reason we would ever change that call.
Also new and relevant, though nothing here uses them yet: `ui_prompt_start` /
`ui_prompt_end` extension events (the `docs/extensions.md` diff is add-only — no
steer or `triggerTurn` semantics moved), and an RPC `clear_queue` that returns
and removes queued steering messages. The second one can discard an
already-delivered but unconsumed mailbox steer; that is survivable because the
mailbox re-delivers on `MEMPALACE_MAILBOX_RESURFACE_MS` (default 3600000), and it
is written down here so a future "the mailbox lost a message" report has a
candidate cause. The three new `PI_HYPERLINKS` / `PI_IMAGE_PROTOCOL` /
`PI_TRUE_COLOR` environment variables were grepped against this whole repo: no
collisions with anything the image sets.
**The bump moved one documented mechanism, so `docs/observational-memory.md` §3
moved with it.** Pi's own `docs/compaction.md` gained exactly one paragraph in
0.84.4: the `autoCompact` threshold is now *also* checked mid-run, after a tool
batch's results are appended and before the next assistant response, skipped only
when that batch ends the run and no queued message needs another response. Our
doc said compaction is "checked when pi goes idle, so it never interrupts a
turn". That was only ever true of observational-memory's **own** trigger
(`compaction-trigger.ts` hooks `agent_settled`); read as a statement about pi it
is now false. `session_before_compact` (`compaction-hook.ts`) therefore has
**two** entry points and the second can fire inside a turn — harmless for the
ledger fold, which makes no model call, but a doc that ships a false promise
about when a hook runs is worse than one that admits two paths. The §3 mermaid
diagram gained the second edge, and the whole file re-passes the bundled mermaid
checker (6 blocks, 44 labels, 0 soft-wrapped, no cut glyphs at 1280px and
800px).
**pi-atelier `v0.8.2` → `v0.10.0` is two minor releases and both are UI-only** —
Sidebar kept calm during an active Turn, composer frame and Status Rail polish,
fullscreen-copy-safe Sidebar, Windows path normalisation, Workspace Pulse
deferred until pi trusts the project. Neither release carries a BREAKING notice.
The coupling that matters runs the *opposite* way to this pin's hard-earned
floor: v0.9.0 renders the Sidebar as a separate split-layout child and therefore
"raises the minimum supported Pi version to 0.84.0", and — unlike the
0.7.1-under-pi-0.84 startup-hang precedent, which its metadata never encoded —
this time `peerDependencies` says so (`>=0.84.0`, up from `>=0.80.7`). Satisfied
with room to spare by `PI_VERSION=0.84.4`. It also pairs deliberately with a
0.84.4 feature: atelier keeps Sidebar content out of the fullscreen transcript
selection while pi adds `fullscreenCopyOnSelect` and Ctrl+X for the selection
itself. Both executable floors (`scripts/smoke-test.sh`,
`scripts/recreate-sanity-check.sh`) compare with `sort -V`, so `0.10.0 >= 0.7.1`
is evaluated correctly — verified by running the comparison, because the string
form of that test reads `0.10.0` as *older* than `0.7.1`.
**While bumping the pins, the README's own pin table turned out to have been
wrong since v1.8.6.** It advertised pi `0.84.2` and mempalace `3.7.1` in the very
table whose purpose is to tell a reader what is pinned and where. Both rows went
stale in the *same* commit — `93f986e` (v1.8.6, "adopt pi 0.84.3 + mempalace
3.8.0") moved both `ARG`s and neither table row; the rows themselves date from
`29b6209` (v1.8.0) and `2ebf00d` (v1.8.4). Only atelier's row was still true.
All three corrected now, and the `--expected-version 0.84.3` example in the
recreate-sanity section updated too, since that one is a copy-pasteable command
that would now fail against a 0.84.4 image. Worth noting how it survived two
releases: nothing checks prose against the `ARG`s, so this table has to be
remembered by hand on every pin bump, and once it was not.
**`credential-incident-response` gained the section its own guidance had been
missing, and §2 gained a precondition it should always have carried.** Docs only;
no image behaviour moves. Both changes came out of a session where three separate
detectors reported *clean* over secrets that were really there — the skill was
the artifact that had taught two agents the pattern, so the fix belongs here
rather than in either operator's private notes.
**§2 previously said an 8-hex fingerprint lets you compare a credential "without
ever materialising the secret", with no condition attached.** That is true only
when the *input space* is unreachable. A fingerprint is 32 bits over whatever it
was computed from, so publishing `fp8(x)` hands anyone a **membership oracle**:
they can test `x == v` for every candidate `v` they can generate. For a 40-char
random token, fine. For a hostname, username, e-mail, port, path, commit SHA or
weak password, that candidate set is a wordlist — and note that "high entropy" is
the usual sufficient condition, not the test: a commit SHA is 160-bit and still
fully enumerable from the repo. Two agents on this fleet published fingerprints of
`GIT_USER_EMAIL`-class values while following this section as written; harmless in
that instance, because those values sit in every commit trailer already, but the
guidance licensed it. §2 now states the precondition, adds that candidate
fingerprints are working memory and never output (a scanner hashes hostnames and
paths too, so "print what it saw" leaks wholesale), and names what a fingerprint
register *is* — a confirmation oracle for anyone already holding a candidate
corpus, which is exactly how a retired token gets identified in old transcripts,
and works the same way for someone else holding those files.
**New §6, "Proving absence: instrument strength, and four ways a scan lies
clean".** Deliberately placed next to §5, because §5 optimises against false
*positives* (name-anchoring, provenance — what stops a triage sweep drowning in
session UUIDs) and every failure in §6 is a false *negative*. Triage optimises
precision; a gate optimises recall, and conflating the two is what produced the
clean reports. It carries: an instrument-strength ranking (exact-byte value search
> class/structure pass > fingerprint census) with the standing instruction to say
which one produced your zero; census and class passes answering different
questions, with both failure modes measured here — a class-only pre-commit hook
passed plaintext UUID API credentials to a shared repo twice because a UUID has no
key header, while a census-only gate reported 0 hits with freshly-synced SSH
private keys in the tree because no key is in the census; the tokenisation trap,
where maximal-run extraction swallows an unquoted `VAR=<uuid>` so the value is
never hashed alone while a *quoted* one is found, meaning quoting alone decided
detectability; scan the index or the pushed tree, never the working tree, plus why
a repo-only fix on an rsync-published mirror is temporary rather than weaker; git
filters never running on symlinks, where `check-attr` answers `git-crypt` for a
path it can never encrypt, so a coverage audit must join the attribute against the
file mode and verify the blob magic; two-sided self-tests that abort, including
the fixture-interaction artifact where a quoted and unquoted probe share one
buffer and make the weak extractor look as strong as the union; and row-gone is
not bytes-gone, since a correct sqlite DELETE leaves the payload in freelist pages
until VACUUM.
Findings contributed by `pi@emb-7kj4vr4g` (the census/class split, and the
instrument ranking's provenance) and `pi@tor-ms22` (exact-byte value search over
index blobs). The description's trigger list grew accordingly and is 1022/1024
characters — **it has almost no headroom, so trim before adding to it**, or the
skill silently fails to load.
**Deployment:** the skill is baked at
`/usr/local/share/pi-devbox/skills/credential-incident-response/`, so this needs
an image rebuild **and** a container recreate to reach any running container.
**Two vendored skills changed, and one of the changes is a correction rather than
an addition.** Nothing about the image's behaviour moves; this is entirely about
what the next agent reads before it acts.
**`pi-devbox-environment` §2 had a rule that was half wrong, and the wrong half
cost five findings in one session.** The section "A negative result is usually
your own filter" closed with *"a positive result needs no such scepticism — it
carries its own evidence."* That sentence is false. A positive result is evidence
about the question your command *actually posed*, which may not be the question
you meant — and the failure is invisible precisely because the command succeeded.
Three measured instances, all from 2026-08-29, all filed as fact before being
caught: an SSH handshake that succeeded and greeted the agent as `joakimp` while
it believed it was probing `gitea.egl.lan` (a `Host gitea*` block had rewritten
`HostName`, so it authenticated to the wrong Gitea instance); a `401` that was a
genuine answer from an issuer which had never minted the credential being tested;
and a "regression" produced by diffing `ssh -G` output against a `2222` that the
agent's own earlier `-p 2222` flag had supplied. The section now carries a
counterpart, *"…and a positive result only proves what you actually asked"*, plus
the three false-negative rows that session added (a palace scan that queried
`embedding_metadata` while documents live in `embedding_fulltext_search_content`;
a token declared dead on a 401 from the wrong issuer; a host declared unreachable
after trying two of its three open ports, with the port written in an environment
variable the agent already held).
**The cross-cutting form of that rule went into `pi-global-AGENTS.append.md`, not
into the skill — deliberately, and this is the whole point of the change.** The
rule *already existed* in the baked skill, authored by an earlier session,
symlinked into `~/.agents/skills/` at every container start. It survived every
recreate, was available for the entire session that broke it, and was violated
five times anyway. So the gap was never persistence; it was **activation**.
A reasoning rule that only loads when a task description happens to match it
cannot fire on the occasions that need it, because "I am about to state something
false" is not a recognisable task type. The always-appended block is read by every
agent in every container without being asked for, which is the only property that
matters here. Writing a sixth document restating the rule would have felt like
progress and changed nothing.
**New baked skill: `credential-incident-response`.** Authored here, so the baked
copy is canonical and it is *not* listed in `skillset-owned.txt`. It carries the
*facts* a two-day credential incident produced, on the theory that facts transfer
between sessions where exhortations do not: probe the issuing provider **first**
(11 of 13 "exposed" credentials in that sweep turned out to be already dead at the
provider — five HTTP requests would have established it, and nobody asked);
`sha256[:8]` fingerprints as leak-free credential identity; the `403`-vs-`401`
trap that scoped tokens introduce into liveness probes, where a live token looks
revoked on `/api/v1/user`; **revocation beats deletion** for anything already
replicated, because deletion is best-effort over an unbounded copy set (FTS shadow
rows, per-host feed inboxes, sqlite free pages, mesh replicas, backups) while
revocation invalidates copies nobody enumerated; the three places a secret hides
in a Chroma palace, in coverage order; deriving least-privilege scopes from
*measured* consumers; and the exposures rotation does not fix (cleartext channels,
git history, agent-authored drawers).
**Three smoke assertions extended** so a rebuild cannot silently drop the new
skill: baked-file existence, resolves-to-the-baked-tree, and reported as `baked`
by `pi-devbox-version`. Skill directories are picked up by a glob in
`entrypoint-user.sh`, so no registration was needed — verified rather than
assumed, since an enumerated list would have left the skill inert, which would
have been a fitting way for *this* skill to fail.
Neither skills change reaches a running container until the image is rebuilt **and**
the container recreated: `~/.agents/skills/` and the global `AGENTS.md` both live in
the image, not in a volume or a mount.
**`cli_utils`' shell *functions* are now sourced, closing the half of that wiring
the image never did.** v1.8.11 linked the repo's `bin/` **commands** into
`~/.local/bin` so they resolve in non-interactive shells; nothing ever sourced
`cli_utils.sh`, so its 14 **functions** (`fgit`, `fhist`, `fssh`, `fdocker`,
`fmark`, `fproc`, `fex`, `fenv`, `extract`, `mkcd`, `pathls`, `portcheck`,
`agents-sync`, `up`) were missing from every interactive shell whose `$HOME` had
no zsh rc. That is the normal case, not an edge case: the container's interactive
shell is bash and **zsh is not installed in the image**. A symlink cannot carry a
shell function and a function cannot be reached from a non-interactive shell, so
the two mechanisms are disjoint and both are required — the image had been paying
this layer's dependency cost (`fzf`, `bat`, `fd`, `rg`, `jq` are baked partly *for*
these functions) while delivering none of its benefit. Now sourced from
`/etc/skel-devbox/.bash_aliases`, with the same detection order as the symlink
block so commands and functions can never come from two different clones.
`CLI_UTILS_SOURCE=0` opts out, deliberately independent of `CLI_UTILS_LINK=0`
because the two disable independent mechanisms. Measured: all 14 resolve in a
freshly-seeded `$HOME`, the opt-out is honoured, an absent checkout is a genuinely
silent no-op (no output, no leaked `_cu` variable), and interactive shell startup
goes from 12 ms to 17 ms.
**Named explicitly, per this repo's own floating-ref rule: `/workspace/cli_utils`
is a host bind mount, not a pinned ref.** Sourcing it means the image now executes
content it does not pin, on every interactive shell, on every device. It is
bash-safe today and that was measured rather than assumed — sourcing under
`bash --noprofile --norc` exits 0 and defines all 14 despite the `*.zsh`
filenames, the functions run, and the tree's single zsh-only construct (`print -z`
in `fzf/fhist.zsh`) is already guarded by `[[ -n $ZSH_VERSION ]]` with a bash
fallback. The residual risk is future content: a cli_utils commit adding a
genuinely zsh-only file would surface as parse errors at every prompt, fleet-wide.
Errors are therefore left visible rather than sent to `/dev/null`, so the failure
is diagnosable, and `CLI_UTILS_SOURCE=0` is the one-line escape hatch.
**`iproute2` is installed, so the container can answer "what is listening in
here".** Neither `ss` nor `ip` was present in any image up to and including
v1.8.11 — nor `lsof`, nor `netstat` — which made `cli_utils`' `portcheck` a hard
stub that printed `portcheck requires at least one of: ss, lsof, netstat` and
exited. `ss` satisfies its preferred branch (`ss -tlnp`), which is also the only
branch that reports the owning PID. `net-tools` is deliberately **not** added
(`netstat` is deprecated and only a fallback path) and neither is `lsof` (~500 KB
for a third route to the same answer). Cost measured, not estimated: ~5.5 MB total
— `iproute2` is 4.2 MB and pulls six libs under `--no-install-recommends`
(`libbpf1`, `libmnl0`, `libtirpc-common`, `libtirpc3t64`, `libxtables12`,
`libcap2-bin`; `libpam-cap` is a Recommends and is correctly dropped). Verified in
a live container: `ss` at `/usr/bin/ss`, `ip` at `/usr/sbin/ip`, both already on
the developer `PATH`, and `portcheck --all` then correctly identifies the `socat`
listener on 8765.
The two changes above also need a rebuild **and** a recreate, for a different
reason than the skills: `$HOME` is the container's writable layer rather than a
named volume (verified — `~/.bash_aliases` carries the container's start mtime
while `~/.bashrc` carries the image's), so the skel file is re-seeded on every
recreate. A `$HOME/.bash_aliases` that is bind-mounted from the host is still
never overwritten, which is the existing contract.
### Dependency audit (2026-08-31)
Every component checked against upstream by direct command, not assumed:
| Component | Baked in v1.8.11 | Upstream now | Action |
|---|---|---|---|
| **pi** | `0.84.3` (pinned) | **`0.84.4`** is npm latest | bumped + audited (above) |
| **pi-atelier** | `v0.8.2` (pinned) | **`v0.10.0`** highest tag | bumped + audited (above) |
| mempalace | `3.8.0` (pinned) | `3.8.0` is PyPI latest | none |
| skillset (mempalace fallback snapshot) | `a12fe5e` | `a12fe5e` == `origin/main`, 0 commits since | none — `--check` reports OK, no NOTICE |
| mempalace-toolkit | `21023e7` | `21023e7` | none |
| pi-toolkit | `0e1369e` | `0e1369e` | none |
| pi-extensions | `2022887` | `2022887` | none |
| pi-fork | `bf702b4` | `bf702b4` | none |
| pi-observational-memory | `ce9fc98` | `ce9fc98` (v3.0.4, peerDeps `*` → no pi floor to clear) | none |
| pi-studio (studio variant) | `3328b3d` | `3328b3d` | none |
| floating `*_VERSION=latest` tools (16) | — | 14 already at latest; `git-lfs` `3.7.1`→`3.8.0` (feature, no breaking section), `uv` `0.12.6`→`0.12.7` (patch) | adopted implicitly by the rebuild; named here per this repo's floating-ref rule |
| node | major pin `22`, installed `v22.23.2` | `v22.23.2` is the newest 22.x | none — a newer LTS *line* (24.x) exists and is deliberately not tracked |
Two method notes, because both would have produced a confident wrong answer:
- **An annotated tag's `ls-remote` SHA is the tag object, not the commit.**
`refs/tags/v0.8.2` is `6e07bf85` while `refs/tags/v0.8.2^{}` is `159f34cf` —
the value actually baked. Comparing the un-dereferenced form reported
`pi-atelier` as *drifted from its own pin*, which would have been a false
integrity alarm about the one component whose pin is load-bearing. Always
deref with `^{}` before calling a pin broken.
- **`git ls-remote --tags | sort -V | tail` is not a "latest release" proxy.**
`typst/typst` carries date-style tags (`v23-03-28`) and `mikefarah/yq` carries
`vTestA`/`vTestB`; both sort *after* the real releases. `Dockerfile.base`
itself resolves `latest` by reading the `Location` of
`curl -sI …/releases/latest`, so replaying that exact step is both
noise-immune and the same source of truth the build will see.
---
## v1.8.11 — 2026-08-27
**Shell state that the writable layer eats on every recreate now gets rebuilt at
start.** Two additions to `entrypoint-user.sh`, both idempotent, both silent
no-ops when the thing they wire up is absent.
**`cli_utils` commands are linked onto `PATH`.** If a `cli_utils` checkout is
mounted, every executable in its `bin/` is symlinked into `~/.local/bin` at
container start — `git-status-all`, `git-pull-all`, `devbox-sanity`,
`pi-devbox-sanity`, `pi-session-repair`, `docker-clean`, `vpn-status`. Detection:
`CLI_UTILS_CONTAINER_PATH` → `/workspace/cli_utils` → `$HOME/cli_utils` →
`/workspace/*/cli_utils`; `CLI_UTILS_LINK=0` disables it.
The reason this is an *image* concern and not the user's problem to re-solve: on a
host, `cli_utils/install.sh` puts those commands on `PATH` by symlinking them into
`~/.local/bin`, which is persistent there — and **ephemeral here**. Same installer,
same repo, opposite durability, so the fix died on every `--force-recreate` and the
next session was back to typing `/workspace/cli_utils/bin/git-status-all`. Running
`install.sh` *inside* a container is the trap rather than the fix: it re-creates
the same disposable state.
**Symlinks rather than a `PATH` edit in an rc file, deliberately.** `~/.local/bin`
is already ahead of `/usr/local/bin` in `ENV PATH`, so links resolve in
**non-interactive** shells too — `docker exec <c> git-status-all`, agent tool
shells, scripts. An rc-file `PATH` edit cannot reach those: `~/.bashrc` returns
early when the shell is not interactive. Measured on tor-ms22 2026-08-27,
`command -v git-status-all` failed in a non-interactive shell while succeeding in
an interactive one, from exactly that asymmetry. Guards, because `~/.local/bin` is
shared with other tooling: a real file is never clobbered, a symlink pointing
somewhere else is never stolen, our own links are refreshed, and links into a
`cli_utils/bin` whose target vanished are pruned — a dangling link on `PATH`
reports "No such file or directory" and reads as a broken container rather than a
removed script.
**A per-device boot hook: `~/.config/devbox-shell/init.sh`.** If the host provides
one, it runs once at start with output to `~/.pi/agent/devbox-init.log`. That
directory is the host-owned bind-mount already sourced into every interactive
shell by `/etc/skel-devbox/.bash_aliases`, so this is its boot-time twin — the
same ownership and the same persistence, but running *before any shell*, which is
what non-interactive fixups (symlinks, directories, one-off migrations) need. **It
introduces no new trust boundary**: that path is already arbitrary code from the
same owner; only *when* it runs is new. Invoked as `bash <file>`, never sourced,
and its exit status is ignored — a hook must not be able to mutate the
entrypoint's own shell state or stop a container from starting.
With the hook in place, the next "can this run on every recreate?" question needs
no image change at all — which is the point, given what the next paragraph costs.
**This moves the base hash.** `base-decide` folds `cat entrypoint.sh
entrypoint-user.sh` into it, so this change forces the ~40-minute base rebuild at
the next tag whether or not anything else in the base moved. It is a rider, not a
reason to tag.
**How it was validated, since CI cannot.** `docker-publish.yml` runs only on
`push: tags: v*`, and `lint.yml` runs `actionlint` over workflow `run:` steps —
neither one executes `entrypoint-user.sh`. So both sections were extracted and run
against fixtures in a throwaway `$HOME` before commit: real file not clobbered,
foreign symlink respected, stale link pruned, new command picked up, second run
byte-identical, `CLI_UTILS_LINK=0` honoured, and "no `cli_utils` anywhere" a silent
`exit 0`. Then run for real in a live v1.8.10 container, after which
`command -v git-status-all` resolved in a *non-interactive* shell. No
`smoke-test.sh` assertion was added on purpose: the positive path needs a
`/workspace` mount that smoke does not have, and asserting it there would repeat
the v1.8.0 mistake of a smoke assertion written against a stage that does not
exist at run time. `workflow_dispatch` with `smoke_only` remains the way to
exercise this against `HEAD` before a tag.
**Also carried by the floating `mempalace-toolkit` main ref** (resolved at build
time, not by a pi-devbox commit — `MEMPALACE_TOOLKIT_REF=main`):
**A scrubbed re-export of a dormant session could silently never reach the
palace host.** `bin/mempalace-pi-session` ships to the palace with
`rsync -a --update`, and the stage file's mtime is deliberately the SOURCE
transcript's mtime (`os.utime()`, "preserve session mtime for dedup
stability"). Re-exporting a session that has not been appended to since its
last ship therefore produces a mtime that is *not newer* than the receiver's —
exactly the case a redactor upgrade needs to ship, since content differs while
mtime does not. `--update` reported success and sent nothing. Found and
patched by `pi@mbp-m1-2020` (mempalace-toolkit `a361b71`): `--update` →
`--checksum`, which compares content and ignores size/mtime entirely.
Dropping `--update` outright was considered and rejected — rsync's default
quick check already transfers on a size difference alone, which would have
masked the *next* instance of this (a redaction whose placeholder happens to
match the secret's length) as fixed. `os.utime()` is untouched; its backdating
is a separate, load-bearing design call for dedup stability. New regression
test, `scripts/test-rsync-ship-idempotency.sh`, runs fully offline (a local
rsync destination exercises the same size/mtime/checksum comparison as the ssh
transfer) and is built to *discriminate*: it must fail against `--update` and
pass against `--checksum,` not merely exercise the code path — the first draft
of the test used fixture strings of different lengths and passed for the wrong
reason (rsync's quick check transfers on size difference alone regardless of
`--update`), which is the same trap the patch itself was written to avoid.
**Acceptance line for this class of change going forward:** "receiver sha256
matches sender for every staged file", not "local stage is clean" — a clean
local stage says nothing about what a dormant session already sent.
**An event addressed to an identity no session runs as is delivered to
nobody, and this fleet has now hit it three separate ways.** RFC 003 gains
§7.13 and open-decision 10 (mempalace-toolkit `21023e7`, docs only, no image
behaviour change): the owed-set derivation — the log's only push channel — is
keyed on `to_agent`, and a reply is always addressed back to whatever string
the *original writer* put in `from_agent`. Nothing validates that string
against a live session identity, so authoring under a synthetic or foreign
name makes every reply to that event write-only. Measured cost this cycle: a
directed ask planted under a synthetic sender drew a correct reply containing
an urgent security finding, and it sat unread for ~2h20m, found only because a
human asked whether mail had arrived. Permitted exception, unchanged: a
synthetic sender is fine for a deliberate control experiment, provided the
body names the real identity to reply to.
**Also carried by the live `skillset` mount** (each device's own clone, not
baked — except the `mempalace` skill's fallback snapshot, re-vendored below):
**The mermaid-diagrams checker's cut gate moved from client pixels to a
per-SVG user-space unit.** `CUT_PX` was calibrated against one live page at
one render scale; sweeping `--viewport` 500→1600 on an *unchanged* document
moved the worst overflow −1.0px → −3.0px, near-proportional to the viewport,
i.e. a constant geometric overflow viewed through a changing scale. `cutU =
cutPx / scale` (scale taken per-SVG, never a page average — one page mixes
scales 0.643–0.988) recovers that invariant: the sweep now collapses to
exactly −3.0u at every width. Re-deriving the threshold against the live host
surfaced a real false negative the old pixel gate had: a label at
`cutPx=0.4, scale=0.678` read as healthy under `CUT_PX=0.5` but is `0.59u` —
a genuine cut hiding behind a compressed render scale. `CUT_U` stays `0.5`;
`cutPx` and `scale` are still printed on every issue so a devtools ruler still
confirms the number on the actual page. A new, explicitly-deferred finding
from the same review: `cut` only measures vertically, so an unbreakable token
wider than its box (a long URL, a `snake_case` identifier) is invisible to
soft-wrap, tall, *and* cut simultaneously — filed as a backlog item, not
implemented, pending a fifth acceptance control.
**The `from_agent`-identity finding above is also now in the `mempalace`
skill itself** ("Writing to another machine", and Anti-Patterns), and the
baked fallback snapshot of that skill was refreshed to match
(`vendor-mempalace-skill.sh`, `6eb20af` → `a12fe5e`) — sanctioned to skip on
its own (`--check` reported stale-but-truthful), done anyway because this
release's point is getting today's fixes live, and the base rebuild below was
already forced regardless.
### Dependency audit (2026-08-27)
Every component checked against upstream by direct command, not assumed:
| Component | Baked in v1.8.10 | Upstream now | Action |
|---|---|---|---|
| **mempalace-toolkit** | `b2b50af` | **`21023e7`** | ships the rsync ship-fix + RFC 003 §7.13 (both above) |
| **skillset** (mempalace fallback snapshot) | `6eb20af` | **`a12fe5e`** | re-vendored (above); live-mounted devices already had it |
| pi | `0.84.3` (pinned) | `0.84.3` is npm latest | none |
| mempalace | `3.8.0` (pinned) | `3.8.0` is PyPI latest | none |
| pi-atelier | `v0.8.2` (pinned) | `v0.8.2` highest tag | none |
| pi-studio (studio variant) | `v0.9.52` | `v0.9.52` — `main`'s commit and the tag's commit are identical (0 either direction) | none |
| pi-toolkit | `0e1369e` | `0e1369e` (local clone HEAD == `origin/main`) | none |
| pi-extensions | `2022887` | `2022887` (local clone HEAD == `origin/main`) | none |
| pi-fork | `bf702b4` | `bf702b4` | none |
| pi-observational-memory | `ce9fc98` | `ce9fc98` | none |
pi-toolkit / pi-extensions checked against their actual Gitea origin (the
Dockerfile's `PI_TOOLKIT_REPO` / `PI_EXTENSIONS_REPO`), not a GitHub mirror —
querying `api.github.com` for those two returned nothing (rate-limited or
blocked; not investigated, the local clones are the source of truth anyway).
No SHA above is fork-supplied; a fork inventing plausible-looking upstream SHAs
is a recorded failure mode (v1.8.9), so every value here came from
`git ls-remote`, a local clone's own `origin/HEAD`, `npm view`/registry JSON,
or the PyPI JSON API, run directly.
---
## v1.8.10 — 2026-08-27
**This tag exists to deploy a fix and a safety net that are currently running on
exactly one machine.** The feeder scrubber has been hand-copied to `/opt` on one
device since this morning; every other device has kept staging unscrubbed
transcripts into the shared palace. Nothing here is a new capability for its own
sake.
`MEMPALACE_TOOLKIT_REF=main` floats: `docker-publish.yml` resolves it to a
concrete SHA at build time, so whatever is on toolkit `main` when the tag is
pushed ships in that image whether or not this repo has a commit. That is the
rule v1.8.9 adopted after `553d865`/`5b8d78f` shipped undocumented twice — *name
the behaviour change before tagging, not after* — and this entry is that rule
being obeyed rather than re-learned.
**`mempalace-toolkit` main moves `5b8d78f` → `b2b50af`** (13 commits, ~2100
insertions / ~520 deletions). No pi-devbox commit implements any of it.
### ⚠️ THE ONE CHECK THIS RELEASE MUST NOT SKIP
**On the first client that runs this image, verify the memory feed still stages.**
The feeder is *fail-closed* by design: no redactor module, no staging (`exit 3`).
That is correct behaviour and it is also the failure mode with no alarm — a
packaging or path mistake stops the fleet's entire transcript feed and nothing
complains loudly, because refusing to stage looks exactly like a quiet session.
This is not hypothetical. `f0bffd1` exists because the feeder is installed as a
symlink (`/usr/local/bin/mempalace-pi-session` → `/opt/mempalace-toolkit/bin/…`)
and `${BASH_SOURCE[0]}` reports the *symlink* path, so the module lookup landed in
a directory where it does not exist. Had that shipped, every device would have
refused to stage on first boot. It was caught by execution, not by review.
Acceptance, in order, on the first recreated client:
1. Run a session, then confirm the feeder logged a scrub summary — a
`[scrub]` line with tier-tagged counts (`T1:env-value=…`, `T2:github-pat=…`),
or an explicit "zero redactions". **Silence is the failure signal**, not success.
2. Confirm the palace drawer count *moved* for that session (the feed reached the
server, not just the stager).
3. Confirm `exit 3` did **not** fire: `mempalace-pi-session` invoked through the
`/usr/local/bin` symlink must find `mempalace_redact.py`.
4. Only then trust the rest of this release.
If step 1 or 3 fails, the memory feed is down fleet-wide until it is fixed, and
sessions that ran in the meantime are not recoverable from the palace — they were
never staged. `MEMPALACE_FEED_ALLOW_UNSCRUBBED=1` is the loud escape hatch, and
using it means accepting unscrubbed transcripts until the packaging is repaired.
### Dependency audit (2026-08-27)
Every component checked against upstream, not assumed:
| Component | Baked in v1.8.9 | Upstream now | Action |
|---|---|---|---|
| **mempalace-toolkit** | `5b8d78f` | **`b2b50af`** | ships the scrubber + symlink fix + `hlc` join |
| **pi-studio** (studio variant) | `v0.9.48` | **`v0.9.52`** | 22 commits, additive only — see below |
| pi | `0.84.3` (pinned) | `0.84.3` is npm latest | none |
| mempalace | `3.8.0` (pinned) | `3.8.0` is PyPI latest | none |
| pi-atelier | `v0.8.2` (pinned) | `v0.8.2` highest tag | none |
| playwright | `1.62.1` (floats `latest`) | `1.62.1` | none — no drift this cycle |
| pi-fork | `bf702b4` | `bf702b4` (2026-08-24) | none |
| pi-observational-memory | `ce9fc98` (v3.0.4) | `ce9fc98` | none |
| pi-toolkit | `0e1369e` | `0e1369e` (2026-08-07) | none |
| pi-extensions | `2022887` | `2022887` (2026-08-17) | none |
**`pi-studio` `v0.9.48` → `v0.9.52`** — four releases, 22 commits, all additive:
PDFs open directly in Studio with watched previews, the header can hide, and
contextual *side questions* arrive (selected-tool use, frozen git context, export,
keyboard shortcuts). No removals or renames in the diff; the changes are
concentrated in `client/studio-client.js`, `index.ts` and three new `shared/`
helpers.
**Its `pi` floor is `>=0.84.3` and we pin exactly `0.84.3` — satisfied with zero
headroom.** Worth naming as a watch item rather than a problem: the next studio
release that raises the floor breaks the studio variant until `PI_VERSION` moves,
and that failure surfaces at build time in the studio job only, after the core
variant has already published.
### Also pulled in by the floating toolkit ref (documentation only)
RFC 003 gains **§9.2**, a proposed direction for the one open decision this
fleet keeps tripping over — that a report addressed to a device is never
delivered, because mailbox candidacy requires exactly `status="open"`. It records
a negative result worth keeping: widening the owed set to include terminal events
cannot work, since the asserting shape and the clearing shape must be disjoint or
every closure mints a fresh obligation. No code implements §9.2 in this release.
The skillset snapshot also moves, so this image bakes the mermaid-diagrams skill's
Playwright driver and the honest note that a `claimed` ack notifies nobody.
### Transcripts get scrubbed before they are staged (`3d47937`, `836e35b`, `f0bffd1`)
`bin/mempalace_redact.py`, called from `mempalace-pi-session` at the moment the
staged transcript is written — one hook covering both transports, because local
mode mines that file and remote mode rsyncs the same bytes.
- **Why it exists, measured rather than argued.** One leaked bearer token had
reached 3 drawers, 13 feeder inbox files across all three devices, and 10 local
files spanning 10 days, from an agent printing an env var while debugging. A
second sweep then found `GITEA_ACCESS_TOKEN` in 2 more drawers and
`GITEA_EGL_ACCESS_TOKEN` in 3. This is routine agent behaviour, so the fix
belongs in the pipeline, not in discipline.
- **Detection is name-anchored, never entropy-anchored.** A palace's own primary
keys — drawer ids, chunk ids, event ids, replica ids, HLCs, commit SHAs — *are*
its high-entropy strings, so an entropy detector eats the memory it protects,
silently and unrecoverably. Three tiers instead: T1 literal values from this
process's env whose name says secret (zero false positives by construction);
T2 vendor shapes (`ghp_`, `glpat-`, `xox*-`, `sk-`, `AKIA`, JWT, PEM, URL
credentials, `Authorization:`); T3 key-name-says-secret.
- **T3 is report-only, because the false-positive rate was measured.** On 52 MB
of real fleet transcripts T3 fired 403 times, mostly `${VAR}` interpolation in
compose files, TypeScript identifiers, a *type annotation*
(`credentials: Credentials`), an IPA attribute holding a date
(`krbPasswordExpiration`), AAAK diary shorthand, and terminal output following
an ssh `Password:` prompt. With interpolation/code-context/key-suffix guards the
enforced count fell **403 → 29** on the same corpus. `MEMPALACE_REDACT_STRICT=1`
makes T3 enforce.
- **Operational shape.** Fail closed — no redactor, no staging (`exit 3`),
overridable with `MEMPALACE_FEED_ALLOW_UNSCRUBBED=1`. Every run prints a count
*including* `0 redaction(s)`, because silence is indistinguishable from a
scrubber that never ran. Findings carry rule, label, length and `sha256[:8]` —
never the value.
**Near-miss this image would have shipped, caught before tagging (`f0bffd1`).**
The image installs `/usr/local/bin/mempalace-pi-session` as a **symlink** into
`/opt/mempalace-toolkit/bin`, and `${BASH_SOURCE[0]}` reports the invoked path,
not the target — so the sibling-module lookup resolved to `/usr/local/bin`, the
redactor was absent, and fail-closed did as instructed: `[FATAL] ... refusing to
stage`. Measured side by side, the symlinked invocation FATALed while the direct
one scrubbed 40 findings. **At the next bake that would have stopped every
feeder tick on every device — a silent fleet-wide memory outage, worse than the
leak the scrubber prevents.** Fixed by chasing the symlink chain in portable
shell (`readlink -f` avoided: GNU/newer-BSD only, and this script also runs
directly on macOS hosts) with colon-separated fallback candidates. Verified via
the symlink, the direct path, and a second-hop symlink. General lesson: fail-closed
converts "module not found" into an outage, which makes the module lookup
load-bearing infrastructure that must be tested through the invocation path the
fleet actually uses — not the convenient one from a checkout.
### The mailbox becomes explainable and mesh-safe (`bfe9c5c`, `a92c75d`, `e917662`, `ecc2a9c`)
- **Owed-set derivation joins on `hlc`, not `seq`** (`bfe9c5c`). `seq` is a
replica-local arrival counter — the same event is `#7` in one database and `#12`
in another — so a second replica would let already-answered asks resurrect.
`hlc` is immutable and replicated, fixed-width, so string comparison *is* causal
comparison. A safe no-op on today's single replica (verified: the positive-control
pair orders identically under both keys), correct once a mesh exists.
- **Delivered text now says it is queued** (`a92c75d`). Delivery uses `steer`
with no `triggerTurn`, and the poll fires on `agent_settled`, so nothing wakes
the model — a delivered ask sits until a human starts the next turn. Measured
case: a directed report sat unread for 2.5 hours. The note explains the agent is
not ignoring the ask, it is not running.
- **`MEMPALACE_MAILBOX_NOTIFY` gains explicit `=kitty` / `=osc777` modes**
(`e917662`). Terminal autodetection inside a container is not unreliable, it is
*blind*: `docker exec` forwards neither `KITTY_WINDOW_ID` nor `TERM_PROGRAM`, and
`TMUX` is unset because tmux runs on the host. Verified on a live process:
`TERM=xterm-256color` and nothing else.
- **The terminal path through tmux is documented as UNVERIFIED** (`ecc2a9c`).
Test sequences written to the pty produced no notification on a remote client;
tmux likely drops unknown OSC types without `allow-passthrough`, and multi-client
routing (one ask pinging every attached client) is an open question.
### Documentation (`e1cc759`, `982b001`, `d4d8bb6`, `d2764bf`)
- **RFC 003, the coordination-log spec the code had been citing all along** — it
did not exist anywhere. 11 sections, retrospective against mempalace 3.8.0
`logstream.py`, incl. owed-set derivation, ten dogfooded landmines and seven open
decisions. Non-obvious findings: `event_append` has **no** idempotency guard on
the write path (verify-before-retry; the replication path *is* guarded),
coordination tools are exempt from both palace locks by design,
`GET /logstream/events` **never existed** in 3.8.0 (not proxy-blocked), and
`mempalace sync` never touches the logstream — the log is permanent and unbounded.
- **`docs/fleet-memory.md`**, operator-facing: five storage types, a decision tree,
latency expectations (~2–5 min live session; next session while offline),
broadcast exclusion by design, fan-out, and the search-before-answer /
diary-at-session-end / verify-don't-retry habits.
- **`docs/secret-hygiene.md`**, incl. the tier definitions, the measured FP data,
stated false negatives, and the three server-side call sites (specified, not
built — tier 2 only there, since the hub cannot see a client's env).
- Phase 1 exposure record moved to the private fleet repo with a moved-note stub;
retention direction for the unbounded log (logrotate-style: never rotate
still-owed events, rotation invalidates held cursors, archive-verify-delete).
### The other memory system finally gets explained — `docs/observational-memory.md`
`pi-observational-memory` has been baked for several releases and described in
one line of the feature list (*"the `recall` tool for session compaction"*),
which is enough to name it and not nearly enough to use it. New 272-line
explainer with five diagrams, aimed at someone who has seen `/om:status` or a
"compacted memory" block and wondered whether to leave any of it switched on.
**Scoped to what this repo is authoritative for, because upstream already
documents the mechanism well.** `/opt/pi-observational-memory/docs/` ships
`concepts.md`, `how-it-works.md` and `configuration.md`, including a correct v3
lifecycle diagram — so the new document links those for depth and spends its own
words on the four facts pi-devbox owns and can change: the pinned commit it bakes
(v3.0.4 `ce9fc98`, the value in `build-manifest.json`), the `packages[]` entry
that decides which copy loads, the Haiku-workers-vs-Opus-session split seeded
into `~/.pi/agent/settings.json`, and the `devbox-pi-config` volume that makes
the ledger survive `--force-recreate`. Plus the confusion this image creates by
shipping two things called memory: a section contrasting it with MemPalace, on
the line *observational memory keeps a session coherent, the palace keeps the
fleet coherent*.
Every stated number was read out of the live container or the baked tree rather
than copied from release notes — including the correction that the dropper is
gated on a **successful same-turn reflection** and not on a token threshold of
its own, which is the one detail `pi-extensions/SKILL.md` still gets wrong.
Placement follows the audience split fleet-ops states for itself: reusable
mechanism is not deployment data, so a "why is this in my container" document
belongs in the repo that **pins and wires** the component, pointing upstream for
depth. Linked twice from the README, because before this commit the README
referenced `docs/` zero times and the one file already there
(`mempalace-broker-design.md`) was reachable only by listing the directory.
### A README claim that v1.8.9 made false, and how it got there
**§ Cross-machine agent coordination ended with "Nothing in this image polls the
log on the agent's behalf." The mailbox shipped in v1.8.9, so that sentence has
been wrong since `aac4a1c`.** Replaced with the three knobs and their defaults
(`MEMPALACE_MAILBOX`, `MEMPALACE_MAILBOX_POLL_MS` 300000,
`MEMPALACE_MAILBOX_RESURFACE_MS` 3600000), the fact that owed-ness is *derived*
rather than read off `status`, and the queued-into-the-next-turn delivery
semantics measured on two devices.
The cause is the one v1.8.9 wrote a rule about: the behaviour arrived through the
floating `MEMPALACE_TOOLKIT_REF`, so **no diff in this repo ever touched the
paragraph that made the claim**. v1.8.9's rule ("name a floating-ref behaviour
change in the CHANGELOG before tagging") fired for the CHANGELOG and nobody swept
the README. The CHANGELOG records what *changed*; the README asserts what is
*true*, and only the first is reviewed at release time. Extending the rule
accordingly: grep the README for absolute claims — *nothing*, *never*, *does
not*, *only* — about any component whose SHA moved.
**The replacement is dated on purpose.** It says it describes the bridge *as baked
in v1.8.9* (`mempalace-toolkit` `5b8d78f`) and points at that repo's
`docs/rfc-003-coordination-log.md` §7.11–§7.12 for the mechanism, because toolkit
main is already ahead of the baked copy (`a92c75d` makes delivery say it is queued
and ping the human who is not looking; `e917662` and `ecc2a9c` refine that notify
path) and none of it reaches a container until a base rebuild. Documenting those
here would have swapped a stale-behind claim for a stale-ahead one — the same
defect with the sign flipped.
### Diagrams verified by rendering, not by parsing
Both comparison diagrams **parsed clean and rendered with their meaning
reversed**: Mermaid laid the second declared `subgraph` out first, so "with
observational memory" appeared before "without", and MemPalace before
observational memory in the diagram whose entire job was that contrast. A third
was legible only at 1280px. Rebuilt as declaration-ordered node chains, then
re-rendered at mermaid@11 — the version `pi-studio` pins — in the baked headless
browser and read back as an image. Recorded because it generalises:
`mermaid.parse()` proves syntax and says nothing about layout, so a diagram is
unverified until someone has looked at it.
### … and rendering it in *my* browser was still not enough
Reported from a real viewer: several boxes had their bottom line of text sliced
off. Reproduced and root-caused rather than nudged — **Mermaid measures a node
label with its own font metrics, computes the box, then renders the label as real
HTML inside a `<foreignObject>`.** Any host stylesheet that touches the
`line-height` or `font-size` of that HTML makes the text taller than the box
already committed to, and the overflow is clipped at the box edge. Error
accumulates per line, so the loss always lands on the last line of the tallest
labels — which is exactly what was reported.
Two fixes were tried and only the second works:
- `%%{init: {'flowchart': {'htmlLabels': false}}}%%` — **rejected, and verified
ineffective rather than assumed so.** The directive *is* honoured (label
elements switch from 16 `foreignObject` to 7 `tspan`), and the clipping is
identical, because the inflated font-size still inherits into SVG text.
- **A hard limit of two short lines per node, with the detail moved into the prose
under each diagram.** One- and two-line boxes have enough vertical slack to
absorb the inflation; three- and four-line boxes do not. This is also better
documentation — the old nodes were carrying paragraph-sized text.
The regression harness is now the interesting artefact: render every block with a
deliberately inflated `line-height: 1.7 !important` on the label HTML, screenshot,
and read it. Two survivors of the rewrite were caught only by that harness — a
long unbreakable `/opt/pi-observational-memory` path silently wrapping to a third
line, and a cylinder (`[( )]`) shape, whose curved bottom leaves less room than a
rectangle for the same two lines.
### §4 answers the question the document left hanging: what compaction does to your context
Asked directly and worth writing down: *if the old conversation is folded away, is
the session back to knowing nothing?* No — and the specifics are all checkable
against pi 0.84.3's own `docs/compaction.md` and the extension's source:
- **A verbatim tail survives, sized by a token budget rather than a message
count.** Pi walks back from the newest entry until `keepRecentTokens` [20000],
and everything from that `firstKeptEntryId` onward is kept **unchanged**. Cut
points land on turn boundaries, never mid-tool-call.
- **The system prompt and `AGENTS.md` are not in the compacted region at all** —
they are rebuilt from disk on every request, so compaction cannot lose them.
- **Nothing is deleted from disk.** Compaction *appends* a `compaction` entry
carrying the summary and the cut pointer; no session line is rewritten in place.
- **`recall` therefore still resolves ids whose sources left the context**, because
it reads the full branch via `sessionManager.getBranch()` and never consults the
context window.
- **Repeated compaction does not summarise the summary.** The text is always
rendered from live observation/reflection records, so there is no
generation-loss spiral; the projection is incremental against the last full-fold
boundary and escalates to a true re-fold from the branch root at
`observationsPoolMaxTokens` [20000].
And one correction to this repo's own earlier claim: **"compaction calls no model"
is a steady-state property, not an absolute.** If the ledger is empty — compaction
firing before the observer has ever run — the hook returns nothing and explicitly
declines ownership (`// Decline ownership so Pi's native summarizer preserves the
pre-cut context.`), and pi's own model-based summariser runs. The doc now says so,
with the snippet.
### A shipped doc bug: the ledger entry type was stated exactly backwards
§9 told readers the entries are `custom_message` and specifically *not* `custom`.
It is the other way round, so the one grep the section existed to get right was
the one it got wrong. Corrected against the live session file — 11
`om.observations.recorded` and 6 `om.reflections.recorded` entries, all
`"type":"custom"`, alongside `"type":"custom_message"` entries whose `customType`
is `mempalace-mailbox` and `mempalace-wakeup`, which is precisely where the
confusion came from: **the mailbox uses the context-visible API, om's ledger uses
the invisible one.**
That is not a typo but a load-bearing distinction, and the fix turns it into a
feature the doc now advertises: `custom` entries *"do not participate in LLM
context"* (pi `docs/session-format.md`), so **the ledger costs zero context until
it is folded** — now a row in the cost table.
### Not covered by any of this
The opencode bridge is a separate write path the feeder hook never sees, and the
server-side layer is unbuilt — so a secret typed straight into `add_drawer`, or
staged by a non-pi client, still lands unscrubbed.
---
## v1.8.9 — 2026-08-26
The coordination log gets a reader, and the release checklist's last gate stops
accusing the wrong component.
### The mailbox arrives — named here *because nothing in this repo caused it*
**`mempalace-toolkit` main moves `e70bef2` → `5b8d78f` (exactly one commit, 281
insertions / 9 deletions across `extensions/pi/mempalace.ts` and
`extensions/pi/README.md`), and that is what actually ships the auto-delivered
logstream mailbox.** No pi-devbox commit implements it. `docker-publish.yml`
resolves `MEMPALACE_TOOLKIT_REF=main` to a concrete SHA at build time and folds
that SHA into `base_tag`, so the mailbox would have landed in the next tagged
image **whether or not this section existed** — which is precisely why it exists.
That is the same shipped-undocumented shape as `553d865` in v1.8.7, and that one
caused a cross-host misattribution: an agent on another machine reasoned about
which image contained which behaviour from a CHANGELOG that never mentioned it.
The rule this release adopts: **if a floating ref will pull a behaviour change
into the image, name it in the CHANGELOG before tagging, not after.**
What the mailbox does, from the shipped code rather than from the design
discussion:
- **The bridge was write-only.** It stamped provenance on the way *out* and never
read the log back, so a directed ask reached an agent only if that agent
happened to run `mempalace_event_list` itself. The channel carried real
cross-machine traffic from 2026-08-18 onward with **zero readers** — every
delivery in that window happened because a human said "check your mailbox".
- **Doubly gated, exactly like the provenance stamper:** inert unless *both*
`MEMPALACE_PI_DEVICE` and `MEMPALACE_REMOTE_URL` are set. An unstamped client
has no address to be reached at, so there is nothing for it to read.
- **On by default, opt out with `MEMPALACE_MAILBOX=0`.** Deliberate: an opt-in
fix for a nobody-remembers-to-do-it problem only relocates the forgetting.
Tunables: `MEMPALACE_MAILBOX_POLL_MS` (min gap between mid-session polls,
default 300000) and `MEMPALACE_MAILBOX_RESURFACE_MS` (re-announce a still-owed
ask after, default 3600000).
- **Owed-ness is derived, never read off `status`.** `event_ack` appends and never
mutates, and `status` is written once, so a directed `open` keeps matching the
mailbox query forever — answered or not. A candidate counts as answered only
when one of this device's own events has a **higher `seq`**, joins via
`metadata.ack_of` or a shared `correlation_id`, and carries a terminal status
(`applied`, `superseded`, `failed`, `blocked`). `claimed` and `ready` are
deliberately **not** terminal — that is how "taken, but not finished" keeps
resurfacing.
- **`*` broadcasts are excluded from the owed set.** `to_agent: <me>` also matches
broadcasts per the tool contract, so without this a broadcast written with
`status="open"` would make every machine believe it personally owed the same
answer — and the code would contradict the skill that documents it.
- **The dedup map is in memory on purpose.** A restart forgets, so an already-seen
ask can resurface: visible noise a human corrects in one turn. The opposite
failure — suppressing an unanswered ask — is silent and permanent. Do not
"fix" the noise by persisting it.
- **Delivery queues, it never interrupts.** A sections push at
`before_agent_start` plus a second `agent_settled` handler behind the 300 s
floor, using `steer` and *not* `triggerTurn`: `agent_settled` means idle, so
nothing wakes a model on inbound fleet traffic.
Measured on v1.8.8 (which bakes `e70bef2`, i.e. no mailbox) immediately before
this release: the wake-up mailbox query had to be run by hand, returned **3**
directed asks with `status="open"`, and the derivation above resolved **all
three** as already answered — the third independent confirmation that the raw
`status` filter never shrinks, and the first taken on a fresh container with no
memory of having answered them.
### `--expected-image-version`: two versions, two flags
**`scripts/recreate-sanity-check.sh --expected-version 1.8.8` reported
`✗ pi version mismatch: expected 1.8.8, got 0.84.3` and exit 1** — a red on the
final runtime gate of a release, accusing the image of being the wrong version,
when the flag had only ever asserted `pi --version`. `AGENTS.md` step 4 spelled
it `--expected-version X.Y.Z` inside a checklist where every *other* `X.Y.Z` is
the pi-devbox tag; `README.md` got it right, so the two documents disagreed.
Not hypothetical, and not one reader's slip: the v1.8.8 release-readiness handoff
from `pi@emb-7kj4vr4g` (`evt_20260826T134919_a614ecfc2d4f`) propagated
`--expected-version 1.8.8` twice, in its body and in
`metadata.cannot_check_here`, while correctly calling step 4 "the runtime peer of
the smoke gate, so it is not ceremonial". Two independent readers, one on another
machine, converged on the wrong meaning. Left alone it puts a spurious red on
every release, and the intuitive remedy — re-pull, re-recreate — is pure waste.
- **New `--expected-image-version X.Y.Z`** asserts the pi-devbox release tag,
read from `release_tag` in `/etc/pi-devbox/build-manifest.json` (the image's
own build-time ground truth — no checkout, no network, no Docker socket). A
leading `v` is optional on either side, so `1.8.9` and `v1.8.9` both work.
- **Both flags now detect being handed the other one's value**, and the test is
exact rather than heuristic: the value is compared against the *other*
quantity this image actually reports, so it can only fire when the mix-up is
real. `--expected-version 1.8.9` now says *"is the pi-devbox IMAGE version,
not the pi version — use `--expected-image-version`"*, and the reverse mix-up
is caught the same way.
- **Neither flag is required any more.** With none, the live `pi --version` is
asserted against `pi_version` in the build manifest. That is not a tautology:
`pi` resolves through `PATH`, and a stale install in the `~/.pi/npm-global`
volume can shadow the baked one — the same shadowing this script already
guards against for `npm:pi-atelier` in `packages[]`. Verified by mutating the
manifest to a different version, which made the new check fail as intended.
- **The header note it replaced was stale and load-bearing.** It claimed pi "is
resolved from `latest` at CI build time and is NOT pinned … cannot self-derive
an expected version". `Dockerfile.variant` pins `ARG PI_VERSION=0.84.3`, and
`docker-publish.yml` *reads that ARG* as its source of truth (refusing to
build on a floating value, checking it is published on npm, warning when npm
is ahead). The same withdrawn claim also sat in `cli_utils`'s
`pi-devbox-sanity --help`, the third place this confusion lived; fixed there
too, in that repo.
- Argument parsing hardened while in there: a flag whose value is missing — or
is another flag — is now a usage error (exit 2) instead of silently consuming
the next argument, and `--help` works.
All fourteen flag combinations were exercised by execution, including the two
manifest-absent branches and the shadowing branch, which a healthy container
cannot reach naturally — mutation-tested with a doctored manifest path so that
each failure branch was observed *firing* rather than assumed present.
### Component audit: no bumps, and that is the finding
Checked before tagging, since a base rebuild was already forced:
| Component | In v1.8.8 | Upstream now | Action |
|---|---|---|---|
| pi (npm) | `0.84.3` (pinned) | `0.84.3` is `latest` | none |
| mempalace (PyPI) | `3.8.0` (pinned) | `3.8.0` | none |
| pi-atelier | `v0.8.2` (pinned) | `v0.8.2` highest tag | none |
| pi-toolkit, pi-extensions, pi-fork, pi-observational-memory, pi-studio | floating | **identical to baked** | none |
| skillset snapshot | `6eb20af` | `6eb20af` | none |
| **mempalace-toolkit** | `e70bef2` | **`5b8d78f`** | ships the mailbox |
So the whole ~67-minute base rebuild this tag pays for is attributable to the
toolkit SHA alone — `base_tag` folds it, and it moved. Every other floating ref
resolved to the commit already baked (verified with `git ls-remote` per repo, not
by reading a cached clone).
One claim in this audit came from a fork that had fabricated its findings — six
plausible-looking toolkit commits with five nonexistent SHAs, a pi `0.84.4` that
npm has never published, a pi-studio commit `ls-remote` says does not exist, and
a compatibility floor of `0.8.2` where the code says `0.7.1`. Every row above was
therefore re-measured directly. Recorded because the failure mode is specific:
none of it looked wrong, and `git cat-file -e` is what caught it.
---
## v1.8.8 — 2026-08-26
The vendored `mempalace` skill snapshot stops being anonymous, and the
container starts saying which copy of each skill it is actually reading.
**Peer review (pi@emb-7kj4vr4g, logstream correlation
`skills-provenance-review`, full text in
`drawer_pi-devbox_reviews_e43e766641c9ec85217bc6ce`) found three blockers before
this was tagged. All three were the same species: a record asserting something
it had not verified. Every finding below was reproduced by execution here before
being fixed.**
- **The verification gate could print `OK` and exit 0 without verifying
anything.** `git show <ref>:<path> | sha256sum` hashes *empty stdin* when the
ref does not resolve, yielding a real-looking `sha256("")` rather than an
empty string — so the `UNKNOWN` branch in `--check` was dead code. Reproduced:
a bogus ref reported `MISMATCH` (accusing the snapshot of lying when the true
cause was an incomplete clone — and the operator's natural remedy for
MISMATCH is to re-run the refresh, which *rewrites provenance to silence the
complaint*); with a 0-byte snapshot against a 0-byte upstream file it printed
`OK: … exactly skillset@aaaaaaa` and exited 0 for a ref that does not exist.
The script already had the right idiom (`sha_empty`) and had applied it to
`blob_sha` but not to `at_ref`. Now existence is *proven* with `git cat-file
-e` before anything is hashed, at two levels (does the ref resolve; does the
path exist at it) because those are different failures. This was the same
defect class as the canary it replaces: a check that can succeed without
checking. A second, unflagged instance of the identical pipeline shape was
found in `blob_sha` and fixed too.
- **`--check`'s exit codes conflated "stale" with "lying",** so the release step
failed in the case `AGENTS.md` step 2 explicitly calls legitimate. Now: `0`
truthful (including stale-but-truthful, with a `NOTICE`), `1` a lying record
only, `2` cannot determine (ref absent from this clone). `AGENTS.md` step 2
rewritten to state all three, since its promise that "the message
distinguishes the two" was exactly what the branch was breaking.
- **The staleness `NOTICE` then asserted a direction it had never tested** — the
same defect one layer down, found by pi@emb-7kj4vr4g against the real state of
its own host. The branch fired the notice on "recorded ≠ HEAD" and announced
that HEAD was the newer side, so a clone that was merely *behind* was told
*"has moved to 82a8d3c; the snapshot describes the older 5fd0d5c"* when
`82a8d3c` is `5fd0d5c`'s **ancestor**. Harmless to the verdict (`rc` stayed 0,
nothing was mis-verified) but it points the operator at a refresh — a
~67-minute base rebuild — when the real remedy is `git pull`. It now tests
ancestry with the `merge-base --is-ancestor` primitive the refresh path two
sections above already used, and reports three distinct verdicts: **stale**
(recorded is an ancestor — refresh), **your clone is behind** (HEAD is an
ancestor — pull, do not refresh), **diverged** (neither). All three verified by
execution; only the first was right before.
- **`--help` died with `unknown option: --help`.** The strict argument loop that
closed the silent-ignore hole never added a `--help` case, so the one script
whose argument *order* was itself a landmine had an erroring discoverability
path. It now prints its own header block.
- **`VENDORED.md` contradicted itself, in the release whose stated invariant is
non-contradiction.** Its hand-maintained "Snapshot provenance at last refresh"
line named skillset `670f7f1` — seven commits behind the ARG, and *the very
commit that told agents to hand-stamp `added_by`*, i.e. the withdrawn
instruction this line of work exists to stop shipping — while its `cp` recipe
still contradicted the "not `cp`" rule 20 lines above. The hand-maintained
line is gone (nothing forced it to move when the ARGs did); `670f7f1` is kept
only as a labelled cautionary example. The `pi-extensions` half was verified
redundant (CI resolves `PI_EXTENSIONS_REF` via `require_sha`) before removal,
rather than silently dropped.
**Should-fixes from the same review, all reproduced:** `--check` given the
documented positional spelling (`<root> --check`) silently ran a *refresh*,
because only `$1` was parsed — both tools now parse all arguments and reject
unknown ones; a refresh at a detached or older `HEAD` silently rewound ref and
bytes, now refused unless the recorded ref is an ancestor (`--force` to
override); `upstream_dirty` was computed and never used in check mode, now
reported; `pi-devbox-version --no-skills --json` printed human text and broke
`jq`; `--help` was a hardcoded `sed -n '2,22p'` range that this branch had
already made stale; the skill fingerprint hashed `SKILL.md` alone, so a live
skill dir differing only in a sibling file still reported "identical" — and
`pi-extensions` already ships two files — so it is now a per-skill **tree** hash
and the manifest field is renamed `skillset_snapshot_tree_sha256` to say what it
measures; and the `--no-skills` smoke assertion was negative-only, passing on a
crashed binary, now anchored positively. `mktemp`+`mv` left written files at
`0600` (a `mv` takes the temp file's mode) — CI was unaffected because the git
index records `100644`, but a local build from a dirty tree would have baked it;
now `chmod 0644` before the `mv`.
**The skill fix ships outside this release, because it had to.** The review also
found that skillset `82a8d3c` — the coordination protocol itself — told every
machine on this fleet to *skip* the mailbox it introduced: it gated the mailbox
on `mempalace_mesh_peers`, and a hub-and-spoke palace reports `peers: []`
precisely because every machine is a thin client of one replica. It also
asserted that a directed `open` event "stays in their mailbox until" acked —
false, because `event_ack` appends and `status` is written once, so an answered
ask matches forever. The headline measurement behind that claim ("exactly 1 —
the one that needed a reply") was of an event already acked half an hour
earlier. Fixed in skillset `5fd0d5c`, which derives owed-ness by joining on
`ack_of`/`correlation_id` with a **`seq` ordering test** — without which one
terminal reply suppresses every later ask on the same thread forever. Because
the skillset is mounted live on every enrolled host, that correction was already
deployed fleet-wide before this image was built; the vendored snapshot is
resynced to it (`c04cd15` → `5fd0d5c` → `6eb20af`) so the no-clone fallback does
not ship the withdrawn rule. Canary re-verified bidirectionally against the new
bytes.
`6eb20af` adds the limit of that ordering test, found when pi@emb-7kj4vr4g
verified it rather than adopting it: **`seq` is replica-local.** It equals
`origin_seq` today only because one replica authors events for all four machines,
so a second replica could order the same pair differently and derive a different
owed-set from the same log — use `hlc` (already on every event, total and
causally consistent) once `mesh_peers` reports any peer. Documented as reasoning,
not measurement, since a second replica cannot be stood up to test it. The part
worth keeping is the **asymmetry**: local-`seq` skew makes an answered item
*resurface* (noise, self-correcting, visible), while a timestamp comparison
*suppresses an unanswered ask forever* (silent, permanent) — so anyone tempted to
"fix" a resurfacing item with `created_at` would be trading the safe failure for
the dangerous one.
**Also carried, previously undocumented:** `dbb7879` resynced the vendored
`mempalace` snapshot to skillset `c04cd15` ("the withdrawal only holds where the
bridge is live"), landed after the v1.8.7 tag and so absent from that image.
⚠️ **A base rebuild is forced** (~67 min): both that resync and the
`pi-devbox-version` / `entrypoint-user.sh` changes below touch inputs to
`base_tag` (`rootfs/` and `entrypoint*.sh`). The provenance recording itself
adds nothing to that cost — it lives entirely in `Dockerfile.variant`.
Both come from one finding, made while verifying v1.8.7 from inside a freshly
recreated container: **the baked `mempalace` snapshot is read by no host on this
fleet.** `~/.agents/skills/mempalace` is a symlink to `/workspace/skillset/skills/mempalace`
— `entrypoint-user.sh` links the baked skill only `if [ ! -e ]`, and
`devbox-skill-reconcile` then repoints the skillset-owned ones at the live clone
(that is the v1.8.5 fix working as designed). All four compose stacks in
`docker-compose-repo` mount a workspace containing the skillset, so the vendored
copy is a CI/no-mount **fallback** and nothing else. Which means the
`mempalace skill snapshot is current` canary — the assertion that blocked
v1.8.7's first tag — polices a file that no agent on this fleet ever opens,
while the drift that *could* actually mislead an agent (a `git pull` nobody ran
in `/workspace/skillset`) was invisible from inside the container and is
invisible to CI by construction.
**The rejected fix is worth recording, because it was the obvious one.** The
old comment in `scripts/smoke-test.sh` said the real answer was "a CI job
diffing this file against the skillset repo". It isn't:
| Objection | Detail |
|---|---|
| needs a credential CI does not have | the skillset is **private** (`ssh://git@gitea.jordbo.se:2222/joakimp/skillset.git`); every build-time clone in this image uses anonymous HTTPS, and `resolve-versions`' `gitea_sha()` is explicitly documented as public-repo-only — its 401/403 path exists to survive a *stale token against a public repo*, so a private 403 would return empty and `require_sha` would hard-abort the release |
| makes another repo's branch able to fail this build | the same pi-devbox commit would go green today and red tomorrow, and a release could be blocked by an edit in an unrelated repo — precisely the shape of the run 589 failure, but automated and permanent |
| pure churn, and it is measurable | pi@emb-7kj4vr4g pushed **four** skillset commits in one evening (`d9dbbbd`, `b740d51`, `3324bd0`, `c04cd15`); a byte-parity gate would have demanded a pi-devbox resync commit **and a ~67-minute base rebuild for each one**, to keep current a copy almost nobody resolves |
| guards the wrong artefact | see above: on this fleet, nobody reads it |
**The invariant is not currency, it is non-contradiction** — the framing comes
from pi@emb-7kj4vr4g's review (logstream `project/pi-devbox`, correlation
`skillset-vendor-drift`, which also **retracted** its own earlier build-time
byte-compare recommendation). A stale-but-self-consistent fallback is harmless;
a stale fallback carrying a **withdrawn instruction** is a live footgun, and
this project has already paid for that one — through v1.8.4 the baked snapshot
*shadowed* the live clone, which is how superseded attribution guidance kept
reaching agents. That is precisely what the bidirectional canary asserts, and
why it stays.
So provenance is **recorded** rather than policed, and the check moves to where
the skillset actually is — a maintainer's clone, or any running container.
### Added
- **`build-manifest.json` now records the vendored snapshot's provenance:
`skillset_snapshot_ref` (which skillset commit the bytes are claimed to come
from) and `skillset_snapshot_sha256` (the bytes that actually shipped).** The
ref is a plain `ARG` **default in `Dockerfile.variant`**, deliberately not a
CI-resolved output, which buys three things at once: it needs no credential
for a private repo; it keeps a local `docker build` and CI identical by
construction (the same reasoning that put `MEMPALACE_VERSION` in
`Dockerfile.base` rather than duplicating it in the workflow); and it requires
**no change at any of the four `Dockerfile.variant` call sites** (`smoke`,
`smoke-studio`, `build-variant`, `build-variant-studio`), whose `--build-arg`
lists are hand-duplicated and therefore easy to under-apply to only two.
Also emitted as OCI label `se.jordbo.pi-devbox.skillset-snapshot-ref`, so it
is readable off the registry without pulling the image.
Two design points, each arrived at from the file's own rules:
- **The ref is a claim; the hash is measured.** `Dockerfile.variant` writes
the manifest from ground truth (`rev()` on each `/opt` clone, the live
`pi --version`), so the snapshot hash is computed with `sha256sum` in that
same layer rather than passed in. A build where the two disagree is exactly
what the new smoke assertions catch.
- **They are siblings, not members of `components{}`.** That map means "HEAD
of a clone present in this image" and the skillset is not cloned here —
calling it a component would be a lie a future reader would act on. It is
also load-bearing mechanically: `pi-devbox-version` renders every
`components{}` value with `.value[0:12]`, which would truncate a 64-hex
digest into something that looks like a short commit. Same reasoning as
`mempalace_version`'s existing comment.
⚠️ **Costs no base rebuild.** `base_tag` hashes `Dockerfile.base` + `rootfs/`
+ `entrypoint*.sh` + the mempalace-toolkit SHA; `Dockerfile.variant` is in
none of it. `scripts/check-base-hash.sh` scans `Dockerfile.base` **only**
(`DF="Dockerfile.base"`, single hardcoded path), so a new `*_REF` ARG in the
variant is invisible to that guard — correctly, since it changes nothing
about the base's contents.
- **`pi-devbox-version` gained a `skills:` section** reporting, per vendored
skill, whether the live copy is `baked` or a `live <repo> @ <sha>` clone —
and for `mempalace`, whether that live copy matches the baked fingerprint:
`(identical to baked snapshot)`, `(baked snapshot <ref> + uncommitted edits)`
when the clone is at the recorded commit but the bytes differ, or
`(baked snapshot <ref> — live copy differs)`. Same live-vs-baked shape as the
existing `pi:`/`palace:` drift annotations. **This is the check CI cannot do
and a container can, for free**, since every host that matters already has the
skillset mounted. The list iterates the baked tree rather than a hardcoded
name list, so vendoring a fourth skill needs no edit here.
`entrypoint-user.sh` calls it with the new **`--no-skills`** flag: the banner
is printed FIRST, before the baked links exist and long before the skillset
deploy and reconcile run last, so anything it said about skill sources would
describe a state that is about to change. Wrong-but-plausible is worse than
absent. (This is the one part of the change that touches `rootfs/` and
`entrypoint-user.sh`, so it does cost a base rebuild — already sunk, since
`dbb7879` refreshed the vendored snapshot.)
- **`scripts/vendor-mempalace-skill.sh`** — refreshes the snapshot and rewrites
the recorded ref *together*, because a `cp` without a matching ARG bump
produces a manifest that confidently lies, which is worse than the anonymous
snapshot it replaced. Refuses to record a ref when the upstream file has
uncommitted modifications (no commit describes those bytes, so recording one
would be a fabrication) — checked on that one file, not the whole tree, so
unrelated work in progress in the skillset does not block a vendoring.
`--check` answers "is the committed snapshot really `skillset@<recorded
ref>`?" and separately reports staleness against the clone's HEAD.
Counterfactual-tested rather than reasoned about, against throwaway clones:
a tampered snapshot reports `MISMATCH` **and** `STALE` (rc 1); a ref rolled
back to the previous skillset commit reports `MISMATCH` with content
unchanged (rc 1) and a subsequent refresh fixes only the ref, leaving the
bytes alone; unstaged and staged-but-uncommitted upstream edits are refused
with distinct messages and the snapshot left byte-identical, i.e. the refusal
is atomic.
**Hardened after review** by pi@emb-7kj4vr4g, whose warning was that a resync
script must "write the ref it ACTUALLY copied from, or the provenance field
inherits the same class of bug the canary just had". The first draft copied
the working tree and guarded it with `git diff` — which says nothing about an
**untracked** file, and can be clean on a detached or behind checkout while
`HEAD` names something else. The snapshot is now *constructed* from
`git show HEAD:<path>`, so the recorded pair cannot be a lie by construction,
and the untracked case is refused explicitly (tested: it was the one input the
first draft would have silently recorded a false ref for). Both new scripts are
`bash -n` clean and `shellcheck -S error` clean — the gate v1.8.7 added.
### Fixed
- **Three stale in-repo markers, all the same failure class.** Two "Unreleased"
pointers — `scripts/smoke-test.sh` pointed the reader at "the Unreleased
changelog note", and the v1.8.6 correction at "the Unreleased entry above";
that section became the `## v1.8.7` heading at release time and neither
back-reference was updated. The third: `scripts/smoke-test.sh`'s own coverage
list still advertised "typst PDF engine for pandoc **(Unreleased)**", five
releases after typst shipped in v1.4.0. Same class as the canary they sit
next to: true when written, silently false at release, with nothing checking
them. The smoke comment now describes the mechanism that actually shipped
(and why the CI-diff idea it advertised was rejected); the changelog one names
v1.8.7; the typst line names v1.4.0.
### Not fixed, deliberately
- **CI still cannot tell you the vendored snapshot is behind `skillset` main.**
That needs a read-only deploy key for a private repo threaded into
`resolve-versions`, to warn about a file no host on this fleet reads. Revisit
when a no-skillset container becomes a real deployment (shipping the image
outside the fleet, or a CI-only agent) — at which point the honest gate is a
**warning**, matching the existing `PI_VERSION`/`MEMPALACE_VERSION` policy
(concreteness → error, newer-release-exists → warning), never a build
failure.
- **The phrase canary stays.** It is orthogonal and free: it pins *content*
where the new fields pin *provenance*, so it still catches a re-vendored
snapshot whose ref was bumped correctly but whose bytes came from the wrong
place — and, per the review above, asserting the **absence of withdrawn
guidance** is the half of it that earns its keep. Its comment now states the
limit instead of promising a fix.
- **v1.8.7's published image has no recorded ref**, and that is expected: the
field arrives here. Worth knowing when reading one, since the tag move
`ebd0de0` → `f645e66` means the published v1.8.7 carries a pre-`dbb7879`
snapshot, i.e. its baked mempalace skill lacks c04cd15's "confirm the bridge
actually stamps" caveat. Harmless — on v1.8.7 the bridge *is* live, so that
caveat self-retires, and every enrolled host reads the live clone anyway.
`pi-devbox-version` degrades quietly on such an image: no fingerprint, no
annotation, verified against the real v1.8.7 manifest.
### Documented
- **The fleet's cross-machine coordination, which was working and unwritten.**
The RFC 003 logstream has carried real work between hosts since 2026-08-18 —
patch handoff, design review, a v1→v2 supersede — and no document in this repo
or the toolkit said so. Written up in three places, split by what each is
authoritative for:
- `README.md` § *Cross-machine agent coordination* — what the **container**
needs: `MEMPALACE_REMOTE_URL` selects the shared palace, and
`MEMPALACE_PI_DEVICE` is what makes this machine *reachable* on the log,
because when every host is a thin client of one palace the stamped agent name
is the only thing distinguishing them. Set both or neither: a container
without the device var can read the log but is addressable by nobody.
- the skillset's `mempalace` skill (`6eb20af`, live on every host that mounts
the skillset, no rebuild needed) — the **norms**: a mailbox query at wake-up,
and the sender-declared ack contract, where a *directed* event with
`status="open"` is owed a reply and a `*` broadcast owes nothing. The
`status` filter earns its place by dropping broadcast noise — measured, an
unfiltered mailbox returned 5 events, 4 of them finished broadcasts from
eight days earlier — but that is **all** it does; it does not compute
owed-ness, and the version of this entry that claimed otherwise is withdrawn
above. Measured today, both machines: the raw filter returns 2 asks here and
1 there, **every one already answered**, while the derivation returns 0 for
both. Dropping noise and deciding what is owed are two different jobs.
- mempalace-toolkit `extensions/pi/README.md` (`e70bef2`) — the **mechanism**,
including that the bridge is *write-only* today (it stamps events going out
and never reads the log, so nothing in this image polls on the agent's
behalf), and that live SSE push is a palace-deployment question: the server
implements `GET /logstream/stream`, but a reverse proxy exposing only `/mcp`
makes it unreachable — verified by 404s against the real endpoint.
⚠️ **The snapshot was refreshed rather than left stale.** The skill edits landed
in the skillset (`5fd0d5c`, then `6eb20af`), so `SKILLSET_SNAPSHOT_REF` was
resynced to match and `scripts/vendor-mempalace-skill.sh --check` is a clean
`OK` with no notice: the no-clone fallback carries the **corrected** protocol,
not the withdrawn one. That mattered more than currency usually does, because
the superseded copy contained an instruction — the `mesh_peers` gate — that
actively told a reader to skip the feature. Refreshing remains a deliberate
release-day decision rather than an automatic one: it costs a base rebuild, and
skipping it is legitimate because every enrolled host reads its live clone.
What is not legitimate is skipping it *silently*, which is what the new manifest
fields and `pi-devbox-version` output make impossible — hence step 2 in
`AGENTS.md` § *Release-day checklist*. In this release the refresh was free:
`rootfs/` was already changing, so the base rebuild was forced anyway.
---
## v1.8.7 — 2026-08-25
Patch release, and the fastest turnaround in the series (~9 h after v1.8.6) for
one reason: **v1.8.6 shipped a container that cannot tell you which machine it
is running on**, and that anonymity produced a real misattribution the same
evening — a session on tor-ms22 read *another host's* diary out of the shared
palace, reported its verification as its own, and built a causal inference on
top of the coincidence. The client-side half of the fix lives in
`mempalace-toolkit`, which the image clones **at build time**, so it can only
reach the fleet through a tag. The CI-hardening work that had accumulated since
v1.8.6 rides along.
Gitea-hosted refs re-resolved immediately before tagging (2026-08-25T22:35Z):
pi-toolkit `0e1369e6` and pi-extensions `20228878` **unchanged** since v1.8.6;
mempalace-toolkit `0fe64c48` → `553d8657` (the provenance change below). CI
re-resolves pi-fork / pi-observational-memory / pi-atelier / pi-studio at build
time as usual. **Base rebuild is forced twice over** — `Dockerfile.base` changed
(the `MEMPALACE_VERSION` audit) *and* `base_tag` deliberately folds in the
mempalace-toolkit SHA ("otherwise a toolkit-only fix never lands") — so expect
~67 min, and note that either cause alone would have sufficed.
⚠️ **The first tag of this version did not publish.** Run 589 built the base
fine, then **both** smoke jobs failed 81-passed/1-failed on a single assertion —
`mempalace skill snapshot is current`, a canary pinning a phrase from the
vendored skill. The phrase it pinned was the heading of the very instruction this
release *withdraws*, so refreshing the snapshot without re-pinning the canary
made it fire correctly on a healthy image. Every publish job was skipped, so
nothing reached the registry and the version was never consumed; the tag was
moved to include the fix below. Fixing the canary is what this release is
*for*, in miniature: the gate was right and the expectation was stale.
### Added
- **Palace writes now carry the device that made them, and diary entries say so
in text.** The container is host-anonymous by construction — `hostname` is a
Docker hash, `$DEVBOX_HOST_ALIAS` is generic, the virtiofs source tag is
generic, and two hosts in this fleet are both `aarch64` — so nothing inside it
distinguished tor-ms22 from EMB-7KJ4VR4G. In a *local* palace that costs
nothing (one origin, so origin is a property of the whole store). In the
**shared** palace it means every drawer and all 621 diary entries read as
though written here, which is exactly how a v1.8.6 verification performed on
EMB was reported as tor-ms22's own.
Two halves, arriving by different routes:
| Half | Where it lives | How it gets into this image |
|---|---|---|
| the writer — stamps `<harness>@<device>` on `add_drawer`/`checkpoint`/`mine`/`event_append`/`artifact_put`, and prefixes diary entries with `HOST:<device>\|` | `mempalace-toolkit` `extensions/pi/mempalace.ts` (553d8657) | cloned in `Dockerfile.base` at `MEMPALACE_TOOLKIT_REF`, whose SHA is folded into `base_tag` |
| the consumer skill — stops telling the agent to do it by hand, adds the read-side warning | vendored `rootfs/…/skills/mempalace/SKILL.md`, refreshed from skillset `73c7c8e6` | `rootfs/*` is hashed into `base_tag` too |
Three design points worth recording, because each was arrived at the hard way:
- **The stamp goes in the *client*, not the agent.** RFC 001 §7.3.2 ranks
"agent stamps it via a skill instruction" as the ❌ *worst possible* place,
and the skill had carried exactly that instruction since 2026-08-23. It
failed as predicted: the agent that wrote the instruction then filed its own
provenance drawer without it. 199 rows reached the palace unresolvable.
One `execute()` wrapper cannot forget.
- **The diary marker is in the entry TEXT on purpose.** `diary_write` has no
metadata parameter, but the deeper reason is that mempalace's `search`
projects a fixed key set and `diary_read` returns content — **metadata is
invisible to the agent who will later read the entry**, so no metadata-only
fix, not even a server-authoritative one, would have prevented the
misattribution. The marker is an AAAK field, so it is machine-parseable
*and* the first thing a reader sees. The wake-up preamble now also names the
device and warns that `diary_read` interleaves every machine's diary.
- **A solitary devbox stamps nothing.** Gated on `MEMPALACE_PI_DEVICE` **and**
`MEMPALACE_REMOTE_URL` — both set only when the palace is actually shared
(RFC 001 R1). Unset either and behaviour is byte-identical to v1.8.6.
Never injected into `diary_write` or `kg_add`: mempalace 3.8.0 hard-rejects
undeclared arguments with JSON-RPC `-32602` rather than dropping them (the
behaviour changed since the RFC's 2026-08-09 note, now corrected), so a
blanket injection would **break** those two calls instead of being ignored.
The allowlist is per tool for that reason.
- **CI now shellchecks the repo's own shell scripts, not just workflow `run:` steps.** `.gitea/workflows/lint.yml`'s `actionlint` job already shellchecks every workflow step, but nothing had ever pointed shellcheck at `entrypoint.sh`, `scripts/*.sh`, or the extensionless tools under `rootfs/usr/local/bin/` (`pi-devbox-version`, `devbox-skill-reconcile`, `dot-watch`, `studio-expose`). The gap is not hypothetical: a sibling repo (skillset's `ci-release-watcher` templates) shipped `echo "$json" | python3 <<'EOF' ... json.load(sys.stdin)` for two months without anyone noticing it silently returned nothing — with no script argument python reads its *script* from stdin, so the heredoc is stdin and the JSON load hits EOF. shellcheck flags exactly this at severity **error** (`SC2259`, "This redirection overrides piped input"); it had been available to catch it the whole time, just never run.
New step in the `actionlint` job, `Shellcheck + syntax-check repository scripts`, runs `shellcheck -S error` plus `bash -n` over every shell file in the repo, **discovered by `*.sh` union a shebang scan** (neither alone suffices) so the extensionless `rootfs/usr/local/bin/*` tools are covered too. Measured before adding it: `-S error` is 0 findings across all 11 shell files today, so the gate is green on arrival with no cleanup. `-S warning` is *not* free (19× `SC2088` tilde-in-quotes in `scripts/recreate-sanity-check.sh`, plus assorted `SC2016`, both intentional here) — a warning-level gate would train people to ignore it, so it stays error-only, same reasoning as the existing `SHELLCHECK_OPTS` exclusions on the actionlint step. File-count guard included: the step fails loudly if the shebang scan matches zero files, since a green check over an empty set is not a check.
- **`MEMPALACE_VERSION` now gets the same CI audit as `PI_VERSION`** — closing
the item v1.8.6 (and v1.8.5 before it) listed as "Still open". The pin was a
literal string in `Dockerfile.base` with **zero** references anywhere in
`.gitea/workflows/docker-publish.yml`, while `PI_VERSION` had ~20: a
concreteness gate, a published-on-registry check, and a never-silently-adopt
drift warning. `resolve-versions` now applies all of them to the palace pin,
read from `Dockerfile.base` (not duplicated in the workflow, so a local
`docker build` and CI install the same version by construction):
| Gate | Behaviour |
|---|---|
| not a concrete `X.Y.Z` | **error** — no floating palace version, same policy as pi |
| not published on PyPI | **error** at resolve time, instead of a `uv tool install` failure mid-build |
| **yanked** on PyPI | **error** — an exact pin installs a yanked release silently under PEP 592, so `mempalace==X` would have shipped a withdrawn client to the whole fleet |
| newer release exists | **warning** naming what to audit before adopting (MCP tool-schema = the agent-facing contract; client/server skew against the central palace) |
Plus one smoke assertion, `installed mempalace matches CI's audited pin`,
gated on a new `EXPECTED_MEMPALACE_VERSION` env threaded into both the `smoke`
and `smoke-studio` jobs. It is **not** redundant with the existing `manifest
mempalace_version matches the installed core`: that one compares two
properties of a single image and therefore cannot notice that *both* are the
wrong version. The failure mode this one covers is a variant built `FROM` a
cached base whose `MEMPALACE_VERSION` pin was older — internally consistent,
silently stale, invisible to every other assertion (the risk
`scripts/check-base-hash.sh` exists to reduce but cannot eliminate).
**Mutation-tested rather than reasoned about**, by extracting the shipped
block out of the YAML and running it with a stubbed `curl`: 9 cases —
`latest` / `3.8` / absent ARG refused; 404 and a registry echoing a different
version refused; a yanked release refused *with its reason*; a newer release
warning without failing; a transient PyPI outage not failing a build whose pin
is already verified; happy path silent and emitting the job output. Then once
more end-to-end against live PyPI with the real `Dockerfile.base`. **This
found a genuine defect in the first draft**: the yank message inlined a jq
program inside a `$(...)` inside a double-quoted string, where the escaping
broke the *filter* (jq compile error) while the surrounding `exit 1` still
fired — a gate that looked correct and reported garbage. The reason is now
hoisted into its own variable. The new smoke assertion was checked the same
way, through the real `run` helper's `sh -c` quoting path: passes on `3.8.0`,
fails on `3.7.1` *and* on `3.8.01` (exact equality, not the substring match
the pi assertion uses), and skips cleanly when the env is unset so a local
`smoke-test.sh` run is unaffected.
⚠️ **Costs a base rebuild on the next tag**: `Dockerfile.base` is hashed
wholesale into `base_tag`, and its now-false "Known gap, carried forward"
comment had to be corrected in place (leaving a comment that says the audit
does not exist would repeat the shipped-false-claim mistake corrected below).
Expect ~67 min, as for v1.8.5/v1.8.6.
- **The SSH sidecar now defaults to connection multiplexing, without overriding
anyone's explicit choice.** `~/.ssh-local/config` already forced `ControlPath`
into the writable sidecar dir, but nothing supplied `ControlMaster` for targets
coming from the user's own bind-mounted `~/.ssh/config`. An entry that never
mentioned it therefore opened a **fresh TCP connection per `ssh` call** — and
an agent doing a dozen calls in a few minutes is exactly the traffic shape that
trips fail2ban or a CGNAT flow-table cap. Observed 2026-08-25 on this fleet:
~12 connections to one host in 15 minutes, after which port 22 stopped
answering while HTTPS to the same estate stayed healthy in 0.44 s (that
asymmetry is the tell for rate-limiting rather than an outage).
**The fix is where the block sits, not what it says.** `ssh_config` is
first-value-wins, so position encodes intent, and the two settings need
opposite treatment:
| Setting | Position | Meaning | Why |
|---|---|---|---|
| `ControlPath` | **before** `Include ~/.ssh/config` | override | the user's value points at read-only `~/.ssh`; it cannot work here, so it must lose |
| `ControlMaster auto` + `ControlPersist 10m` | **after** the `Include` | default | an explicit per-host `ControlMaster no` must keep winning; we only supply an opinion where the user expressed none |
*Force what is broken, default what is merely absent.* The first draft of this
put both in the leading block, which would have silently overridden an explicit
`ControlMaster no` — the counterfactual is in the test below.
Verified with `ssh -G` (the resolved-config oracle) rather than by reading the
man page, against a fixture with one host set to `no`, one silent, one set to
`auto`: the explicit `no` resolves to `controlmaster false` **and** still gets
the writable `ControlPath`, the silent host resolves to `auto`, and the same
fixture under the rejected layout flips the `no` host to `auto` — so the test
discriminates the *position*, not merely the presence of the block. Then
end-to-end: the real script rendered in a sandbox `HOME`, block last, `bash -n`
clean, `shellcheck -S error` clean (the gate added in v1.8.7).
Measured effect on the author's own config (41 host aliases): **22 were silent
about `ControlMaster` and gain `auto` + 10 m persist; 0 are overridden**, since
the fleet contains no explicit `no`. Worth noting *how* the one deliberate
exception is written — `proxmox002-vpn` carries `# No ControlMaster — VPN means
direct route, no CGNAT flow cap`, i.e. the intent is expressed as **absence
plus a comment**, which `ssh` cannot distinguish from "no opinion". That host
does now get multiplexing; its comment says multiplexing is *unnecessary*
there, not harmful. Anything that must stay unmultiplexed needs a literal
`ControlMaster no`.
Why the ordering matters beyond this one config: `~/.ssh/config` is
**per-machine**, differs across the fleet, and future machines' versions do not
exist yet to be audited. A default-not-override design is correct without
needing to inspect any of them.
`ControlPersist` is deliberately short (10 m idle, and each new session resets
the idle timer — long enough to collapse an agent's burst, short enough that an
abandoned socket ages out). A per-host entry that sets its own value keeps it:
hosts already specifying `ControlPersist 4h` still resolve to 4 h. The known
cost of multiplexing is the **stale master** — socket present, daemon gone,
after a suspend or network change — which makes every later `ssh` to that host
hang; recovery is `ssh -F ~/.ssh-local/config -O exit <host>`, now documented
in the `pi-devbox-environment` skill along with `-O check`.
### Fixed
- **The vendored-snapshot canary was one-way, and pinned a phrase the same
release deleted.** `mempalace skill snapshot is current` grepped for
*"Attribute what you file yourself"* — the heading of the hand-stamping
instruction withdrawn above. It therefore did its job (snapshot changed,
expectation did not) and blocked an otherwise-green build. Two changes rather
than a string bump: the assertion is now **bidirectional** (the new phrase must
be present **and** the withdrawn one absent, so a re-vendored *stale* snapshot
fails as loudly as a forgotten bump — verified by running it against v1.8.6's
snapshot, which correctly fails), and the comment now states the structural
limit: a phrase canary can only detect *"older than what I remembered to pin"*,
never *"older than skillset main"*.
- **Four build-provenance smoke assertions verified the presence of a manifest
field name and never looked at its value.** The originals were literally:
```sh
run_expect "manifest records pi_version" "cat …build-manifest.json" '"pi_version"'
```
which passes on `{"pi_version": ""}` and on `{"pi_version": null}`. The tell
was sitting in the passing output all along — `✅ manifest records pi_version
(got "pi_version")` echoes the *key* back as the thing it claims to have
found — and it was spotted while reading run 579's smoke log to confirm the
new v1.8.6 assertions had actually executed.
Replaced with checks against the values, and against ground truth where
ground truth exists:
| Assertion | What it now enforces |
|---|---|
| `manifest declares every required component key` | all seven components present by name, failing with *which* key vanished |
| `manifest component values are resolved 40-hex commits` | each value is a full 40-hex SHA; `null` allowed for `pi-studio` alone (absent in the non-studio variant) |
| `manifest pi_version matches the installed pi` | manifest value equals `pi --version`, same ground-truth shape as the mempalace check |
| `manifest top-level fields are well-formed, not merely present` | `release_tag` non-empty; `source_revision` 40-hex *when populated*; `build_date` ISO-8601 *when populated* |
| `pi-devbox-version --json round-trips the manifest byte-for-byte` | actual string equality with the file, since `--json` is a verbatim `cat` |
**Why five value-checks replace four name-checks (total assertion count
unchanged at 61), and specifically why key presence and value shape are kept
apart:** an "every component value is a valid SHA" loop passes **vacuously** on `components:{}`, because jq's `all()`
over an empty list is true. A single combined check would therefore go green
on a manifest that had lost every component — which is the same shape of hole
as the three false greens already recorded in this file. They are separate on
purpose.
Mutation-tested rather than reasoned about, twice: nine fabricated manifests
through the raw jq filters, then twelve through the shipped assertions using
the real `run` helper's `sh -c` quoting path (the quoting is load-bearing here
— a jq filter that dies on a quoting error exits non-zero and *looks* like a
caught defect). Measured against the old assertions on the same twelve
defects: **old caught 3, missed 9; new catches 12.** The three the old set
caught were key *disappearance* (grepping for a key name does fail when the
key is gone) and the literal string `"unknown"`; every value-level defect —
empty string, `null`, a 12-hex truncation, a wrong-but-plausible version, a
malformed `source_revision` — was invisible. Three legitimate variations are
correctly *not* flagged: empty `source_revision` and empty `build_date` (both
default empty on a plain local `docker build`, so demanding them would fail
honest local smoke runs) and `pi-studio: null`.
- **Dropped the now-redundant `manifest has no unresolved ('unknown')
components` assertion.** The 40-hex value check strictly subsumes it:
`"unknown"` is not 40-hex, and only `rev()` in Dockerfile.variant ever emits
that string, feeding `components{}` exclusively. Removed rather than left in
place, because a redundant check that can never fail independently is one more
green tick that means nothing.
- **Corrected a factually wrong "Still open" bullet in the v1.8.6 entry below**
(see the strikethrough there). It claimed `pi-devbox-version`'s human output
does not display `mempalace_version` and that only `--json` surfaces it. Both
halves are false — v1.8.6 shipped a `palace:` line with the same live-vs-baked
drift annotation `pi` already had. Verified by running the shipped script
against fabricated manifests: matching versions print `palace: 3.7.1`, a skew
prints `palace: 3.7.1 (baked as 3.8.0 — drift detected)`, and a pre-v1.8.6
manifest with no baked field prints the live value un-annotated. The bullet
appears to describe an intermediate state of the working tree and was never
re-checked before tagging. Left visible as a struck-through correction rather
than deleted, since v1.8.6 is already published and someone may have read it.
---
### Still open
- **Make the vendored-snapshot check automatic instead of a remembered string.**
Tonight's failure is the third iteration of the same maintenance burden (v1.8.4:
phrase present in both copies; v1.8.7: phrase deleted by the release that
refreshed the snapshot). A phrase canary structurally cannot answer *"is this
snapshot older than skillset main?"* — only a diff can. Proposed: a lint job
that clones the skillset repo and compares
`rootfs/usr/local/share/pi-devbox/skills/mempalace/SKILL.md` against it,
failing with the diff when they drift. Open question first: the skillset repo is
**private**, so this needs a CI clone credential, which is a policy decision
rather than a code change.
- **Provenance stops at Chroma's metadata.** The hourly reconciler on the palace
host stamps `device`/`agent_kind` in `chroma.sqlite3`, but knowledge-graph
triples and coordination events live in *separate* SQLite files
(`knowledge_graph.sqlite3`, `logstream.sqlite3`) it cannot reach. 156 triples
carry no origin field at all; `logstream`'s `from_agent` is free-form and
already inconsistent (`pi@tor-ms22`, `pi@emb-7kj4vr4g`, and bare `pi` in the
same table). Tracked in RFC 001 §7.3.1.
- **The stamp is self-asserted, and cannot be otherwise yet.** mempalace 3.8.0
authenticates with a *single scalar* bearer token and has zero device concept,
so a verified stamp needs per-device credentials plus an origin field in six
write paths across three databases. Deferred to RFC 001 Phase 4, where it is
now motivated primarily by **revocation** (one shared token covers every
device, so cutting off one laptop means rotating the fleet) rather than by
provenance. Forward-compatible by design: every stamp records *how* it was
determined, so an authoritative pass overwrites with `device_source='token'`
and nothing has to be undone.
- **`tor-ms22` and `tor-ms22-native` are one machine with two device values**
(4,680 and 3,826 rows). That is the hostname-as-identity cost RFC 001 §7.3.4
warned about, now visible in data: a rename splits one device's history
silently. Repairing it means a device-identity mapping, not a relabel.
## v1.8.6 — 2026-08-25
Patch release. Adopts the drift that accumulated in the ~2 days since v1.8.5
(pi `0.84.3`, mempalace core `3.8.0`), then closes the documentation and
observability gaps that v1.8.5 itself listed as "Still open". No component
was adopted without an audit note recording *why* it is safe.
All moving refs re-resolved immediately before tagging (2026-08-25T13:28Z):
pi-toolkit `0e1369e6`, pi-extensions `20228878`, mempalace-toolkit `0fe64c48`
and pi-observational-memory `ce9fc982` all unchanged since v1.8.5;
pi-fork `f1ff8087` → `bf702b4c`; pi-atelier holds at `v0.8.2` (floor for
pi ≥0.84 satisfied); pi-studio's CI-resolved newest tag has moved again to
`v0.9.51`. Base rebuild is forced (Dockerfile.base changed), so the 16
floating base-tooling ARGs re-roll — expect ~67 min as for v1.8.5.
### Changed
- **`mempalace` core `3.7.1` → `3.8.0`.** Released 2026-08-23T21:19Z, hours
after this project's own v1.8.5 tag the same day. Additive/reliability only
— reviewed for MCP tool-schema changes before bumping, as always: none.
`sync --apply` (PR #2320/#2322) no longer deletes a drawer solely because
its `source_file` was unreachable *at that moment* — it asks for
corroboration first. **This does not relax the standing landmine** against
running `mempalace_sync` / `mempalace_delete_by_source` beyond dry-run on
the shared central palace: that failure mode is paths *permanently* absent
from whichever host runs the sync, not transient unavailability, and 3.8.0
doesn't touch it. Server-side perf fix PR #2307 (long-running Chroma servers
no longer invalidate their own HNSW cache on their own writes) likewise does
not make `mempalace_reconnect` unnecessary — that tool covers *external*
writes bypassing the in-process client, a different scenario. Full reasoning
lives in the `Dockerfile.base` comment above `ARG MEMPALACE_VERSION`.
**Deployment note:** synlig's central palace currently serves `3.7.1`
server-side via `docker-compose.mempalace.yml` (which reuses this image) —
this client bump introduces version skew until that stack is separately
redeployed; sequence accordingly.
- **`pi` `0.84.2` → `0.84.3`.** Published 2026-08-24T11:09Z. Release notes
carry one "Breaking Changes" line — `GoogleThinkingLevel` renamed to
`GoogleApiThinkingLevel` — checked against all four vendored packages
(`pi-fork`, `pi-observational-memory`, `pi-atelier`, `pi-studio`): zero
references, inert here. 0.84.3 also fixes two skill-discovery bugs that
land directly on this repo's own vendored-skill work: nested Markdown
skills inside `.agents/skills/` grouping directories not being discovered,
and root Markdown files (`README.md`/`AGENTS.md`) in skill directories being
wrongly reported as broken skills.
### Added
- **Browser automation is now documented to humans, not just to agents.**
`agent-browser` + Playwright + a headless Chromium (~625 MB — the single
largest addition in the image) previously had zero mentions in `README.md`,
`DOCKER_HUB.md` or `THIRD_PARTY.md`; it existed only in the agent-facing
`AGENTS.md` managed block. Added a `README.md` "Browser automation"
subsection, a `DOCKER_HUB.md` feature entry, and `THIRD_PARTY.md` license
rows for `agent-browser` (Apache-2.0), Playwright (Apache-2.0), and Chromium
(BSD-3-Clause for Chromium's own code plus a large set of bundled
third-party components under their own licenses; the binary here is not
compiled by this repo — it's Playwright's own "Chrome for Testing" download
via `playwright install --with-deps chromium`).
- **`THIRD_PARTY.md` gains rows for `pi-atelier` (MIT) and `mempalace` core
(MIT per the GitHub repo; noted that the PyPI package's own metadata omits
a license classifier, so verify against the repo's `LICENSE` rather than
sdist/wheel metadata if clearance is needed from the artifact alone).**
- **`typst` and `socat` added to `README.md`'s tooling inventory.** Both were
already used in prose (typst as pandoc's `--pdf-engine`, socat by
`studio-expose`) but missing from the "What's inside" lists, so the
inventory didn't match what the image actually ships.
- **`mempalace` core version recorded in `/etc/pi-devbox/build-manifest.json`.**
Previously absent — a published image couldn't answer "which palace version
shipped?", and a palace bug couldn't be correlated to an image version.
Derived from the live installed binary (matching the manifest's existing
ground-truth-not-build-args philosophy), degrading to `null` rather than
failing the build if the binary is missing or its output format changes.
Verified landed: new top-level `"mempalace_version"` key, sibling to
`pi_version` rather than a member of `components{}` (that map is rendered
truncated to 12 chars by `pi-devbox-version`, which would mangle a longer
version string).
- **New smoke assertions**, all landed in `scripts/smoke-test.sh`: (1) the
`pi-observational-memory` clone is checked for the actual `ce9fc98`
auth-fix markers pinned to their fix site, `src/runtime.ts`
(`availability_recheck`, `providerCredentialConfigured`,
`hasConfiguredAuth`) — not merely clone existence, and deliberately not a
repo-wide grep: all three identifiers also appear under `tests/`, so a
repo-wide search would stay green even with the fix reverted in
`src/runtime.ts` alone; (2) the manifest's new `mempalace_version` field is
asserted present, non-null, and equal to what `mempalace --version` reports
live, so the manifest can't silently drift from the installed package —
expected to fail against any pre-v1.8.6 image, by design; (3) a
behavioural check for the mempalace-toolkit feeder's `--agent` default
(see below — this one turned out to be possible after all).
### Fixed
- **A false claim was being published to Docker Hub on every release.**
`DOCKER_HUB.md` advertised "neovim (LazyVim defaults)". Nothing in this
repo installs LazyVim — the only nvim configuration is a 19-line
`sysinit.vim` that sets `termguicolors`. `update-description` pushes this
file verbatim (with `{{PI_VERSION}}` substituted) to the Hub description, so
the error was public, not internal. Corrected to describe what's actually
there.
### Component audit for this release
Checked against upstream 2026-08-25 (two days after v1.8.5's own audit):
`mempalace` core moved `3.7.1` → `3.8.0` (see Changed, above — timing is
notable: released *hours after* v1.8.5 tagged, so v1.8.5 could not have caught
it no matter how carefully it was audited). `pi` moved `0.84.2` → `0.84.3`
(see Changed). `pi-toolkit` `0e1369e6`, `pi-extensions` `20228878`,
`pi-observational-memory` `ce9fc982`, and `pi-atelier` `v0.8.2` are all
**unchanged** from v1.8.5 — in particular `pi-observational-memory` still sits
exactly at the auth-fix commit with nothing landed upstream since, and
`pi-atelier` is still the newest tag with the `≥0.7.1` floor for `pi ≥ 0.84`
trivially satisfied. `pi-fork` has one upstream commit not adopted this
release: `f1ff8087` → `bf702b4c`, a text-only rewording of the fork task
preamble (no code-path change) — **left un-pulled** for this release since it
is a moving ref CI resolves fresh at every build anyway; it will be adopted
automatically on the next build regardless of this entry. `pi-studio` (studio
variant) has drifted two tags upstream, `v0.9.48` (pinned at build time via
CI's newest-semver-tag resolution) → `v0.9.51` at tag time, purely additive
(watched PDF previews, opening PDFs directly in Studio, Studio header
hide) — nothing to bump in this repo since studio-tag resolution happens in
CI, not the Dockerfile, but note it **will** auto-adopt `v0.9.51` on the next
studio-variant build. `mempalace-toolkit` unchanged — this release's manifest
and pi-bump work in `Dockerfile.variant` stayed within that file's ownership
and did not require a toolkit-side change.
### Still open
- **`MEMPALACE_VERSION` has no CI-side audit equivalent to `PI_VERSION`'s.**
`PI_VERSION` is verified published-on-npm and warns (never silently adopts)
on drift; `MEMPALACE_VERSION` is a literal Dockerfile string with zero
references in `.gitea/workflows/docker-publish.yml`. Flagged in v1.8.5's
audit as a gap; still a gap.
- ~~**`pi-devbox-version`'s human-readable output does not display
`mempalace_version`.** Its render path is a fixed sequence
(`release_tag`, `build_date`, `source_revision`, `pi`, then `components{}`)
and the new top-level field isn't in it — only `--json` mode (which `cat`s
the manifest directly) surfaces it today. One line in
`rootfs/usr/local/bin/pi-devbox-version` would fix this; deferred since the
field's stated purpose (correlating a palace bug to an image) is already
served by `--json`, but worth doing in a follow-up if this becomes a
routine manual check.~~
**CORRECTION (2026-08-25, post-tag):** this bullet is wrong and was never
true of the tagged tree. `pi-devbox-version` *does* print a `palace:` line
in human mode, with live-vs-baked drift detection, degrading quietly on
pre-v1.8.6 manifests. Nothing is open here. See the v1.8.7 entry above.
**Resolved during this release, not left open:** the feeder `--agent`
default behavioural hook initially looked like it might need a
mempalace-toolkit change (a `--print-config` flag that doesn't exist). It
didn't — `mempalace-pi-session` assigns `AGENT` before argument parsing and
`--help` exits 0 with no side effects, so `bash -x mempalace-pi-session
--help` observes the real resolution (env interpolation and fallback)
without needing a source change. The new smoke assertion exploits exactly
that, checked both ways: with `MEMPALACE_PI_DEVICE` set it must resolve to
`pi@<device>`; with it unset it must NOT be `pi@*` (catches a regression to
the old unconditional `$USER`/`mempalace` default).
`mempalace-toolkit` commit `c64ffa1` changed the feeder's `--agent` default
from `$USER` to `pi@<device>`, but there is still no way for smoke to assert
this default is actually in effect from this repo alone, since
`mempalace-toolkit` is a separate repo this release does not modify. If the
concurrent smoke-test work could not find an honest assertion from the
existing `/opt/mempalace-toolkit` surface (help text, `--self-test`), this
remains open pending a toolkit-side `--print-config`-style hook — a
toolkit-repo change, not a pi-devbox one.
- **16 base-tooling `ARG *_VERSION=latest` pins remain unrecorded.** (Corrected
count — v1.8.5's entry said "~14"; the actual count from `Dockerfile.base`
is 16, plus 5 more that float with no ARG at all: `rustup-init`, AWS CLI v2,
Chromium-via-Playwright, Node's minor version via `setup_22.x`, and
`DEBIAN_VERSION=trixie-slim` itself.) None of these are recorded anywhere
once the build completes — not in the manifest, not in a label — so a
published image cannot answer "which nvim/uv/chromium shipped?" without
exec-ing in and asking the binary.
### Documentation
- **`.env.example` documents `MEMPALACE_PALACE_PATH`.** It was the only MemPalace
variable the template never mentioned, while being the one that silently moves
the feeders' stage: the palace root resolves as `$MEMPALACE_PALACE_PATH` →
`$MEMPAL_PALACE_PATH` → `~/.mempalace/config.json` → `~/.mempalace/palace`, and
the stage is derived from it (`<palace-root>/pi-stage`). The comment states the
precedence, says why neither the image nor the entrypoint exports it (pinning
the palace without carrying the stage re-creates the split a shared root
removed — see v1.8.2), warns that a stage whose persistence differs from the
palace makes a scoped `mempalace sync` prune conversation drawers whose dedup
key is the staged path, and notes it is a *container* path unlike the
host-side `WORKSPACE_PATH`/`SSH_KEY_PATH` above it. Found while auditing a live
host whose `.env` sets the variable redundantly to the default.
---
## v1.8.5 — 2026-08-23
Patch release with two fixes in the container's skill wiring — one behavioural,
one a latent crash found while reviewing the first — plus the `mempalace-toolkit`
change that makes palace writes carry provenance. **No component pin moved.**
Every ref was re-resolved at tag time and is byte-identical to what v1.8.4
shipped: `pi` `0.84.2` (still npm latest), `pi-atelier` `v0.8.2` → `159f34cf`
(newest tag; the `≥0.7.1` floor for `pi ≥ 0.84` holds), `pi-studio` `v0.9.48` →
`c3b83680`, `pi-fork` `f1ff8087`, `pi-observational-memory` `ce9fc982`,
`pi-toolkit` `0e1369e6`, `pi-extensions` `20228878`, `MEMPALACE_VERSION` `3.7.1`
(still PyPI latest, and the version the central palace serves — no client/server
skew). The single moving part is `mempalace-toolkit` `fd8b15f5` → `0fe64c4`.
### Fixed
- **Vendored skills no longer silently shadow their live skillset
counterparts.** `~/.agents/skills` was *asymmetric*: `mempalace`,
`pi-devbox-environment` and `pi-extensions` resolved to the baked
`/usr/local/share/pi-devbox/skills/…`, while every other skill resolved to the
live `/workspace/skillset/skills/…`. Root cause was precedence-by-ordering in
`entrypoint-user.sh`: the baked links are created **early** (line 65 in
v1.8.4; the loop moved down as this fix added comments) — deliberately so, to
close a smoke-test readiness race — with `[ ! -e … ]` so they are
"created only when absent", and the skillset deploy runs **last**, where it
classifies the existing links as foreign and leaves them alone. The comment at
line 61 claimed the goal was "a same-named skillset skill … is never
clobbered" — but with baked-first plus create-when-absent, the *skillset* skill
was precisely the one that lost. Comment now describes the actual behaviour.
Observed cost, on two hosts independently: an edit to
`skillset/skills/mempalace/SKILL.md` (adding a drawer-attribution rule) was
pushed and present in the live clone (`md5 129bcc4752`), yet both the
EMB-7KJ4VR4G and tor-ms22 containers kept loading the baked copy
(`md5 5236024fef`) with zero occurrences of the new rule. The tor-ms22 agent
had to fetch the rule from the Gitea API to read it at all. Editing a skillset
skill therefore *appeared* to work and silently did nothing until an image
rebuild — for exactly the three skills most likely to be iterated on.
**The fix is not "the skillset always wins"**, because ownership is per-skill
(`rootfs/usr/local/share/pi-devbox/skills/VENDORED.md`): `pi-extensions`'
authoritative source is the *package* repo, copied over the snapshot at build
time, and `skillset` carries a downstream copy that can lag — handing that one
to the clone would regress the skill. So a new helper
`devbox-skill-reconcile` runs immediately after the skillset deploy and
repoints only the skills named in `skills/skillset-owned.txt` (today:
`mempalace`). Precedence is now user override → live skillset clone (owned
names only) → baked snapshot, with the early links untouched as the fallback,
so the readiness race stays closed. It only ever replaces a symlink that
points into the baked tree, so a real directory or a link pointing elsewhere
is never disturbed. Verify with `readlink -f ~/.agents/skills/mempalace`, not
by reading the entrypoint.
- **A latent boot-abort in the baked-link block, found while reviewing the fix
above and fixed with it.** `[ ! -e "$link" ]` is TRUE for a *dangling* symlink
(`-e` follows the link), so once a link may point into `/workspace/skillset`
— which the fix above makes possible — a vanished mount turns the guard into
"create over a broken link", and plain `ln -s` then fails with `File exists`.
Under the entrypoint's `set -euo pipefail` that **aborts container start**
before `exec "$@"`, with a cryptic `ln` error and no pi. Reachable on a
`docker restart` or a host reboot under `restart: unless-stopped` (the writable
layer survives and `~/.agents` is not a volume on any host), though not on a
`compose up -d` recreate. Now `ln -sfn`, which heals the broken link back to
the baked fallback; the reconciler re-points it in the same boot if the clone
is back. A comment at the call site records why the `-f` must stay.
- **README's skill-precedence documentation was wrong** in the same way the
entrypoint comment was: it claimed baked skills are "created only when absent
so a same-named skillset skill … is never clobbered" and that "a mounted
skillset always overrides them". Rewritten to state the real, per-skill
precedence and to name `skillset-owned.txt` and `devbox-skill-reconcile`.
- **The smoke canary for a stale `mempalace` snapshot could not detect
staleness.** It grepped `"Shared palace: multiple harnesses"` — a phrase
present in *both* the stale and the fresh copy, so it passed throughout the
shadowing bug above. It now pins the newest section
(`"Attribute what you file yourself"`), and `VENDORED.md` records that
updating this string is part of refreshing the snapshot. Three further
assertions close the gaps that let the bug ship: skill link **targets** are
asserted (not merely `test -L`), the `skillset-owned.txt` list is asserted to
contain `mempalace` and *not* `pi-extensions`, and the reconciler's replace
path — which CI never exercises, since no smoke container mounts a skillset —
is covered by fabricating a skillset and asserting all three outcomes (owned
skill repointed, unowned skill left baked, user override untouched) — plus a
second case that a mutation test proved necessary: with the reconciler's
"is this link ours?" guard deleted, all three of those assertions still
passed, so the discriminating case is an *owned* name whose link is a user
override pointing outside the baked tree.
### Changed
- **Vendored `mempalace` skill snapshot refreshed** from `skillset` `936fed8` →
`670f7f1` (`md5 5236024fef` → `129bcc4752`), which adds the "Attribute what
you file yourself" rule: hand-filed drawers should carry
`added_by="<harness>@<device>"`. Without this refresh the symlink fix above
would only help hosts that mount `skillset`; a bare container would still ship
the pre-attribution-rule skill.
- **Component audit for this release — no pin edits needed.** Every component
except `pi`/`pi-atelier` is pinned to a moving ref that CI resolves at build
time, and each was checked against upstream on 2026-08-23: `pi` `0.84.2`
(still npm latest, published 2026-08-14), `pi-atelier` `v0.8.2` (newest tag;
the `≥0.7.1` floor for `pi ≥ 0.84` is satisfied), `pi-fork` `f1ff8087`,
`pi-observational-memory` `ce9fc982`, `pi-toolkit` `0e1369e6`,
`pi-extensions` `20228878`, `pi-studio` `v0.9.48` → `c3b83680` — all
**byte-identical to what v1.8.4 shipped**. `MEMPALACE_VERSION` stays `3.7.1`
(still PyPI latest, and the version the central palace serves, so no
client/server skew). The one component that moved is `mempalace-toolkit`
`fd8b15f5` → `0fe64c4`, which is this release's other payload: the feeder now
defaults `--agent` to `pi@$MEMPALACE_PI_DEVICE` so palace writes carry
provenance, with `$USER` still the fallback when the variable is unset
(`AGENT="${MEMPALACE_PI_DEVICE:+pi@${MEMPALACE_PI_DEVICE}}"`), so un-enrolled
hosts are unaffected. Nothing landed upstream after the
`pi-observational-memory` merge `ce9fc982`, so the eight-week-bug fix in
v1.8.4 is not destabilised. All of the above was re-resolved immediately
before the tag and was unchanged — worth repeating for any future release,
because six of nine components are moving refs that CI resolves at build
time, so the build, not the Dockerfile, decides what ships.
Two notes for whoever runs the build. This release changes
`entrypoint-user.sh`, `rootfs/**` and the resolved toolkit SHA — all three feed
the base-image hash — so expect a **full multi-arch base rebuild** (~95 min,
as on v1.8.3/CI 562), not a fast variant-only publish. And that rebuild
re-resolves the ~14 base-tooling `ARG *_VERSION=latest` pins; measured drift
on 2026-08-23 was one patch (`nvim v0.12.4 → v0.12.5`), so the window is
favourable, but it is not covered by version assertions.
- **`pi-devbox-environment` skill — new §2 subsection "A negative result is
usually your own filter", plus ControlMaster masking in §3.** This is baked
(`rootfs/usr/local/share/pi-devbox/skills/`, symlinked to
`~/.agents/skills/`), so it is an image-behaviour change even though no
package moved. Motivated by three false negatives an agent produced in a
single session, each from its own filter rather than from the world: a
`| head -20` "proved" an SSH peer absent that was defined at **line 454** of a
~500-line config; `ssh mac 'docker ps'` "proved" the host had no Docker, when
the non-interactive SSH `PATH` simply lacks `/usr/local/bin`; and a `grep 'ssh
'` "proved" no ControlMaster was running, when master processes **rename
themselves** to `ssh: <controlpath> [mux]`. The rule now stated: a positive
result carries its own evidence, absence has to be *earned*. §3 additionally
documents that a live master socket makes later commands authenticate **not at
all**, so "it still works" proves nothing after editing a peer's
`authorized_keys` — verify with `-o ControlPath=none -o ControlMaster=no`, or
the breakage surfaces in a future session with no memory of the edit.
- **`AGENTS.md`: a stale CI claim corrected.** It said "a tag push produces two
runs, not one — `lint.yml` fires on every push (including tag refs)". That
stopped being true when lint was scoped to `branches: ['**']`, which excludes
tag refs by design; `refs/tags/v1.8.4` produced run 571 (publish) and nothing
else. The `head_sha` + workflow-`path` filter advice stays, because it costs
nothing and any future `v*`-triggered workflow would reintroduce the
ambiguity. Also adds a short "Verifying this repo's reality from inside a
container" section, including the trap that **this repo's `docker-compose.yml`
is a template pinning `:latest`** while a real host runs its own per-machine
file — so recreating from the repo copy can silently move a host off
`:latest-studio`.
---
### Still open
- `build-manifest.json` records the `mempalace-toolkit` SHA but **not** the
mempalace **core** version, so a palace bug cannot be correlated with an
image. `mempalace --version` prints it; adding it is a one-line change to the
manifest `RUN` in `Dockerfile.variant` plus one smoke assertion, and is
variant-only (no base rebuild cost).
- Smoke asserts the `pi-observational-memory` clone **exists** but not that it
contains the ambient-auth fix. npm still ships pre-fix `3.0.4`, so an
accidental switch from the `/opt` clone to an npm install would be a silent
regression. Cheap guard: `grep -rl availability_recheck` must be ≥1.
- The feeder's new `pi@<device>` default has no behavioural test hook
(`--dry-run` never prints the agent; `--self-test` only covers the remote-mine
response classifier). Cheapest available check is a source-shape grep for
`MEMPALACE_PI_DEVICE:+pi@`.
---
## v1.8.4 — 2026-08-22
Patch release, and the one that ends an eight-week bug: **the baked
`pi-observational-memory` finally records observations on a Bedrock host that
uses ambient AWS credentials.** The fix is ours, but it is no longer a patch —
upstream merged it, so this release picks it up through the ordinary
`PI_OBSMEM_REF=master` path with no local carry. Also **bumps pi-atelier
`v0.8.1` → `v0.8.2`** (audited below) and bakes the `todo` extension's new
`edit` action. `pi` stays `0.84.2` (still the npm latest, published
2026-08-14) and `MEMPALACE_VERSION` stays `3.7.1` (still the PyPI latest).
### Fixed
- **om consolidation on request-time-signed providers — upstream, not patched
(`pi-observational-memory` `37986b6` → `ce9fc98`).** Under `37986b6`, om's
pre-flight gate treated "pi exposes no `apiKey` and no auth header" as
*unauthenticated* and skipped every consolidation. On Bedrock with ambient
AWS credentials that is the normal case — pi signs SigV4 at request time —
so om recorded **nothing for eight weeks** with no error, no cost and no log
line. Every pi-devbox image up to and including v1.8.3 has that behaviour.
Two commits, both authored here and now upstream verbatim: `6f694e6` fixes
the gate itself (it must not require a credential *payload*), and `699ccc7`
adds the second half of pi's own rule — `hasConfiguredAuth` reads an
availability snapshot that stays empty when the startup availability pass was
skipped/aborted/failed, so on the otherwise-fatal path om now asks pi to
re-check the credential live (refresh scoped to the provider, network-free),
rate-limited 60 s per provider, bounded by a raced timeout, logged as
`resolve.availability_recheck`. Filed as upstream issue #51, merged as PR #52
(`ce9fc98`, 2026-08-22T04:46:18Z), which also carries PR #49's `env`/`baseUrl`
forwarding merged four minutes earlier; the maintainer resolved the textual
conflict between them keeping both behaviours. Verified on `ce9fc98` here:
`tsc --noEmit` clean, `vitest` 257 tests / 27 files green.
**Note for anyone carrying the local workaround:** the interim fix was a
`packages[]` override in `~/.pi/agent/settings.json` pointing pi at a patched
clone outside the image. From this release on, delete the override — the
baked `/opt/pi-observational-memory` has the fix. Confirm with
`/etc/pi-devbox/build-manifest.json` → `components.pi-observational-memory`
before removing it. The npm-published `pi-observational-memory` is **still
`3.0.4` and still broken**; the image does not use npm for this component, so
the release cadence there is irrelevant to us.
### Changed
- **`PI_ATELIER_REF` / `PI_ATELIER_VERSION` `v0.8.1` → `v0.8.2`, audited per the
floor note above the ARG.** Only one version sits between old and new and its
changelog is two lines, both Workspace-Pulse-internal: inspection requests are
now coalesced and serialized so short Turns avoid duplicate Git work and
overlapping inspections cannot run concurrently, and live tool-driven Pulse
updates are preserved while a fresh inspection is guaranteed at Turn end and
retired sessions can no longer publish stale results. **Nothing touches pi's
private TUI renderer**, which is the coupling that produced the
0.6.0/0.7.0-under-pi-0.84 startup hang, and `pi` is unchanged at `0.84.2`, so
this bump does not re-enter that risk class. Both the seam and the pin floor
(never pair pi-atelier < 0.7.1 with pi >= 0.84) are unaffected.
- **`pi-studio` `65995fe` (0.9.44) → `v0.9.48`** — 14 commits, four releases.
Studio-side only (`INSTALL_STUDIO=false` by default, so this lands in the
studio variant): open Studio in Muxy's browser, local PDF preview actions,
previews survive Pandoc probe failures, legacy LaTeX styles tolerated in
Pandoc previews, native dialogs replaced in embedded browsers, and file-copy
import fixes with an explicit fallback. CI resolves the highest semver tag,
not `main`, so this is `v0.9.48` exactly.
- **`pi-fork` `4a09af4` → `f1ff808`** — one commit, "Add fork runtime
awareness" (2026-08-19).
- **`mempalace-toolkit` `b609cf5` → `fd8b15f`** — two commits, **docs only**
(backup/recovery + units; the convos-miner mtime correction finished). The
`fix(pi-session)` false-success guard was already baked in v1.8.3 — checked
by ancestry (`git merge-base --is-ancestor 6e1f4f3 b609cf5`), not by reading
the log, because a commit's *date* does not tell you which side of a pin it
fell on.
- **`aws-cli` 2.36.24 → 2.36.29** and the other `*_VERSION=latest` tools
(bat/eza/fzf/gitleaks/nvim/micro/zoxide/yq/typst/tealdeer/agent-browser/
playwright/gosu/git-lfs/uv) refresh implicitly, as designed.
- `pi-toolkit` unchanged (`0e1369e`, local `main` == baked).
### Added
- **`todo` extension: an `edit` action** (`pi-extensions` `98eb07b` →
`2022887`). The tool is a verbatim vendored copy of pi's own
`examples/extensions/todo.ts`, which offers list/add/toggle/clear and no way
to change an item's text. On a long-lived list that forces either a "patch"
item describing a *different* item, or clear-and-re-add of everything — both
hit for real on 2026-08-17 while tracking a 17-item fleet plan, which ended up
with `#18` correcting `#17`. `edit` takes `id` + `text` and keeps the id and
the done status; `nextId` is untouched. Id stability is the point, because ids
are the only handle a palace snapshot of a plan can refer to.
- Verified live in-container before committing, by repointing
`~/.pi/agent/extensions/todo.ts` at a working copy for one session (unknown
id and missing text both error as intended; editing a completed item kept its
id and its `done` state), then reverting the symlink to the image copy.
### Notes
- **No pi-atelier change was needed for the `todo` action.** Its `tool_result`
hook only checks that `details.todos` is a well-shaped array and ignores the
action string, so the new action flows through its normalizer and sidebar
untouched. Worth knowing while reading agent transcripts: that hook
*replaces* todo tool output with `N/M done · see sidebar` whenever the sidebar
todo panel is visible (`showSidebarTodos`), so an agent sees only the counter
and not the item text — upstream's `list` otherwise returns every item. That
is a deliberate context saving, not a tool limitation.
- The vendored copy now carries a numbered **LOCAL DELTAS** list in its header
(the earlier `ctx.mode !== "tui"` → `!ctx.hasUI` API fix, and this action), so
reconciling a future upstream version stays mechanical rather than
archaeological.
- **A second om fix is NOT in this image and will not be.** Upstream PR #24
("advance coverage watermark when observer records nothing", head
`joakimp:fix/observer-empty-coverage-watermark` `b577b29`) is open but
design-rejected by the maintainer on 2026-07-03: an empty observer verdict is
usually a *technical* failure, so advancing the watermark would leave a gap in
the observed session, and in the genuinely-nothing-to-observe case the next
observer simply gets more context. So the observer can still re-fire on a
growing span after an empty verdict — that is upstream's intended behaviour,
not an image defect. Do not "fix" it by rebasing that branch.
---
## v1.8.3 — 2026-08-16
Patch release. **Bumps mempalace to `3.7.1`** and closes the gap that made the
baked mempalace skill go stale for four commits. `pi` stays `0.84.2` (still the
npm latest) and pi-atelier stays `v0.8.1`; every git-ref component
(pi-toolkit, pi-extensions, pi-fork, pi-observational-memory, pi-studio,
mempalace-toolkit) was checked against its upstream head and is unchanged.
- **`MEMPALACE_VERSION` `3.6.0` → `3.7.1`.** Verified against the 3.7.1 source
rather than its changelog, because the risk is to palaces users cannot
reconstruct: legacy drawers lack the new `chunk_total` completion marker and
**both** decision sites trust them (`if chunk_total is None: ... trust the
match as before`), so there is **no mass re-mine**; `NORMALIZE_VERSION` is `2`
in both versions, so the "pre-v2 drawers are stale" gate does not fire either;
`chromadb<2,>=1.5.4` keeps the same major, so no index-format migration; there
is no auto-migration (the source says *"We do NOT auto-migrate"* twice) and
`rebuild_index` has exactly one call site, the explicit `repair rebuild`; the
single new palace file (`logstream.sqlite3`) is created lazily on first
logstream use. Downgrade stays possible — 3.6.0 has zero references to
`chunk_total` and ignores it as unknown metadata.
Two behaviour changes worth knowing, both turning a silent condition into a
hard refusal: `MEMPALACE_MCP_ALLOW_PEER_WRITER` **no longer works on
local/chroma palaces** (it is now gated on `backend_requires_single_writer()`,
and `_MULTI_PROCESS_WRITER_BACKENDS` is `{pgvector, qdrant}`), and writer-lock
*setup* failures now **fail closed** (`refusing this mutating tool`) instead of
proceeding with a warning. Neither affects this image's normal
MCP-server-plus-CLI-feeder pattern, which already serialised on the same
`mine_palace_*.lock` under 3.6.0 — "process-lifetime single-writer ownership"
in the upstream changelog describes tightened escape hatches, not a new lease.
What 3.7.1 buys a **shared central** palace is the real motivation: the stale
chromadb `SharedSystemClient` cache is now dropped on reconnect (under 3.6.0 a
peer's writes could be overwritten by a stale in-memory HNSW segment, *"index
count going backwards"*), the writer lease is released on SIGTERM/SIGHUP
instead of leaking a lock naming a dead PID, and an interrupted mine is no
longer permanently skipped as though complete.
**Upgrading a server requires restarting it** — 3.7.1 refuses mutating tools
when the served library drifts from what is installed, and `mempalace_reconnect`
cannot clear that (it reopens the database but cannot reload Python modules).
The fleet primary was upgraded and restarted before this image was tagged.
Note: opencode-devbox still pins `3.6.0`. The two images are meant to move in
lockstep, so that pin diverges until opencode-devbox cuts its own release.
- **Vendored `mempalace` skill snapshot refreshed** to skillset `936fed8` (was
`63f3bf5`). This is the gap worth naming: `~/.agents/skills/mempalace`
symlinks to the **image-baked** copy under
`/usr/local/share/pi-devbox/skills/`, and `entrypoint-user.sh` creates that
link *first* while the skillset deploy never clobbers an existing name — so in
a devbox container the vendored snapshot always wins, and editing the skillset
repo alone changes nothing a container reads. Two commits' worth of guidance
had been invisible here: the multi-machine shared-palace section (device
provenance in `source_path`, mined drawers carrying the *mine* date with
UUIDv7 recovery, the naive-local vs UTC timestamp mismatch, `agent_name` not
being device-scoped, single-writer/no-queue semantics) and the
hand-crafted-provenance guard.
- **`pi-global-AGENTS.append.md`** gains `### If the palace is central, it is
shared — three rules`: never run `mempalace sync` against a shared palace (it
prunes drawers whose sources look missing, which on a central palace is most
of the content, including other machines' — compounded by RFC-001 §7.2, since
feeders stage *inside* the palace root); a client-side timeout is not a
failure (single writer, one large mine blocks everyone, so
`mine timed out after 30000ms` usually means the mine completed — verify
before retrying or you file a duplicate); and the `mempalace` CLI is not
remote-aware, so it always opens a local-disk palace and can silently
disagree with the MCP tools.
- **`mempalace-census` is now on `PATH`.** It shipped inside the image at
`/opt/mempalace-toolkit/bin/` but was never symlinked into `/usr/local/bin`
like its three siblings, so RFC-002 Phase A censuses had to be invoked by
absolute path. Added to the symlink set, the `chmod +x` set, and the
build-time `--help` smoke chain.
---
## v1.8.2 — 2026-08-16
Patch release. **Ships the fix for a silent transcript-feed failure**, plus the
smoke assertion that stops it coming back. No image pins changed from v1.8.1
(pi `0.84.2`, pi-atelier `v0.8.1`); what moves is the baked `mempalace-toolkit`
ref and one new smoke check.
**The bug this closes** (found on the first boot of the v1.8.1 image, on
EMB-7KJ4VR4G, 2026-08-15): the container-start catch-up rsynced seven pi session
transcripts to the palace host correctly, then asked the server to mine
`/data/feed/<device>` — the feeder's default `MEMPALACE_PI_REMOTE_PATH`, which
assumes a *containerized* palace server. That fleet's primary runs **natively**
(a systemd user unit + uv tool), so it only ever sees host paths and the mine
died with `source directory not found`. rsync had already succeeded, so the
inbox looked healthy.
It stayed invisible because of the second half: the feeder decided success with
`'"error"' in body`. MCP answers a hard tool failure with HTTP 200 and a
JSON-RPC *result* whose `content[].text` carries the tool's own JSON as an
**escaped string** — the bytes are `\"error\"`, so the substring could never
match. `~/.pi/agent/mempalace-catchup.log` printed
`Done. Wing 'wing_conversations' updated.` directly beneath the error JSON and
exited 0. A feeder whose only artifact claims success is worse than one that
crashes: nothing in the container disagreed with it.
Shipped here:
- **`mempalace-toolkit` ≥ `b609cf5`** baked (CI resolves the ref at build time):
`classify()` parses the MCP envelope instead of grepping it (JSON-RPC error,
MCP `isError`, inner `success=false`/`error`), and separates "verified ok"
from "unverified: no JSON tool payload" rather than assuming the good case.
A preflight warning fires when the rsync destination and
`MEMPALACE_PI_REMOTE_PATH` disagree — in *preflight*, so `--dry-run` and
`--prepare` surface it too. Remote mode also stops previewing NEW/SKIP from
the *local* palace, which had been reporting "6 already filed" about a palace
it was not feeding; the tags are now `[?]` and the summary names who decides.
- **New smoke assertion** — `mempalace-pi-session --self-test` run against the
**baked** toolkit. It replays six recorded MCP responses (fixture 1 is the
verbatim 2026-08-15 failure body) plus a regression guard asserting the old
substring check is blind to it. A stale or reverted `MEMPALACE_TOOLKIT_REF`
can therefore no longer ship a feeder that mines nothing while reporting
success.
- **`.env.example`** now spells out that `MEMPALACE_PI_REMOTE_PATH` is the path
the *server process* can open — the container path for a dockerized server,
identical to the ssh-target path for a native one — and that a mismatch fails
quietly, with rsync succeeding and only the mine failing.
**The `--self-test` assertion is deliberately bare** (`mempalace-pi-session
--self-test`, no `HOME=…` prefix). `run()` invokes
`docker run --entrypoint="" $IMAGE sh -c …` and no Dockerfile sets `USER` or
`ENV HOME`, so it executes with **no `HOME` at all** — the same condition that
made v1.8.0's stage assertion unsatisfiable. The feeder is `set -u` with
HOME-anchored defaults, so it used to die with `HOME: unbound variable` there;
`b609cf5` derives `HOME` from the passwd database (what python's `expanduser()`
falls back to) instead. Keeping the call bare means smoke also proves the feeder
runs in a bare container, rather than papering over it with an env prefix.
---
## v1.8.1 — 2026-08-15
Patch release. **Unblocks v1.8.0, which never shipped.** Its `smoke` and
`smoke-studio` jobs each failed exactly one assertion (67/68 and 70/71 passed),
so `build-variant` and everything downstream skipped: no `v1.8.0` tag reached
Docker Hub and `latest` stayed on v1.7.0 from 2026-08-07. Image content is
unchanged from what v1.8.0 intended — the pins here are identical (pi `0.84.2`,
pi-atelier `v0.8.1`).
The failing assertion was `pi stage defaults next to the palace (not a cache
dir)`, added three days earlier in 7c00dd6. **It was a test bug, not a product
regression.** It asserted a literal path:
```sh
echo "$out" | grep -q "stage=/home/developer/.mempalace/pi-stage/"
```
but the `run` helper invokes `docker run --rm --entrypoint="" $IMAGE sh -c …`,
and neither `Dockerfile.base` nor `Dockerfile.variant` sets `USER` or `ENV HOME`
(the published base image config carries no `HOME` at all — `HOME` is normally
set by `entrypoint-user.sh`, which `--entrypoint=""` deliberately skips). So the
assertion ran as **root with `HOME=/root`**, `mempalace-pi-session` correctly
resolved `stage=/root/.mempalace/pi-stage/…` (it is `$HOME`-relative by design:
`$MEMPALACE_PALACE_PATH` → `$MEMPAL_PALACE_PATH` → `~/.mempalace/config.json` →
`~/.mempalace/palace`), and the literal grep could never match under any
circumstances. The tell was one line below it in the log: the sibling assertion
`pi stage follows MEMPALACE_PALACE_PATH` **passed**, because it sets the variable
explicitly and so never consults `HOME`. Default fails while explicit passes is
the signature of a wrong `HOME`, not of broken staging.
Fixed by asserting the invariant that was actually meant — the stage sits beside
the resolved palace, sharing its lifetime — which is user-independent:
```sh
case "$stage" in
"stage=$HOME/.mempalace/pi-stage/"*) exit 0 ;;
*) exit 1 ;;
esac
```
`$HOME` is expanded by the container's own shell, so this holds as root, as
`developer`, or under any future user, while a cache-dir default — the
regression the assertion exists to catch — still fails it (verified against all
three cases plus a simulated `MEMPALACE_PI_STAGE` cache pin). A second
assertion, `pi stage is palace-adjacent for the developer user`, now covers the
deployment-specific path properly, by *supplying* `HOME=/home/developer` instead
of assuming it.
### Why it took a release to notice — and the `smoke_only` input
`docker-publish.yml` triggers on `push: tags: v*` only. 7c00dd6 was a push to
**main**, so only `lint.yml` ran; v1.8.0 was the first tag afterwards and
therefore the assertion's **first execution ever**. Any smoke assertion written
outside a release was unvalidated until the next release consumed it — the
worst possible moment to discover it.
New `workflow_dispatch` input **`smoke_only`** closes that: it probes/builds the
base and runs both smoke jobs against HEAD, then stops before publishing
anything. Implemented as `if: inputs.smoke_only != 'true'` on `build-variant`
and `build-variant-studio`, deliberately *without* `always()` so the implicit
"needs succeeded" gate survives and a red smoke still blocks a release;
`promote-base-latest` and `update-description` already require `build-variant`
success and so skip on their own. On a tag push `inputs` is unset and
`null != 'true'` is true, so releases behave exactly as before. This release was
validated with a `smoke_only` dispatch before the tag was cut.
### Smoke failures now explain themselves
`run` discarded all output (`>/dev/null 2>&1`), so a red ❌ carried zero
diagnostic weight — explaining this one-line failure took a CI-log dig plus a
registry image-config inspection, when the container had already printed the
answer and thrown it away. It now captures output and prints the last few lines
under a failed assertion only. Assertions that want a diagnostic echo it to
stderr (the stage checks now report the resolved stage and the `HOME` they saw),
which stays invisible while they pass.
---
## v1.8.0 — 2026-08-15
Minor release. Headline: **pi sessions now feed MemPalace by themselves.** The
image already shipped `mempalace-toolkit`, but its pi feeder
(`mempalace-pi-session`) was never symlinked onto `PATH`, so nothing ever mined
pi's transcripts — the palace only ever contained what an agent remembered to
file by hand. A container that gets recreated regularly has no other memory, so
a missed wind-down was a permanently lost session.
Also here: **pi 0.84.1 → 0.84.2 and pi-atelier v0.8.0 → v0.8.1**, bumped
together. The pi bump closes the Amazon Bedrock tool-argument poison pill that
v1.6.4 recorded as unfixed upstream; the atelier bump is the matching companion,
since both sides changed fullscreen input handling in the same fortnight. Audits
for both are below.
*Why event-driven and not a timer:* there is nothing schedulable inside the
container — PID 1 is `bash -l`, with no systemd and no cron — and anything
installed would not survive recreate anyway. The triggers therefore live where
the events already are: pi's own lifecycle, plus container start.
### Added
- **`mempalace-pi-session` symlinked onto `PATH`** (`Dockerfile.base`,
alongside its `mempalace-session` / `mempalace-docs` siblings, with the same
`--help` build-time check). `entrypoint-user.sh` also self-heals the symlink
into `~/.local/bin` (already ahead of `/usr/local/bin` on `PATH`, and
writable by `developer`) so the feature works on images whose base predates
this change.
- **Container-start catch-up feed** (`entrypoint-user.sh`, backgrounded). pi's
mempalace extension feeds the palace on `session_shutdown` and on a debounced
`agent_settled`, but a hard kill (`docker kill`, OOM, host reboot) runs no
handler at all; this is the only trigger that can recover the previous life's
transcripts. Skipped when a remote palace is configured without an inbox to
ship to — **and the skip now says so** (see Changed) — and skippable entirely
with `MEMPALACE_FEED=0`.
- **`MEMPALACE_PI_STAGE` no longer needs pinning here — the feeder's default
was fixed upstream instead.** It used to stage under `~/.cache`, which is
disposable in a container; the first cut of this change pinned the env var
into the persisted `~/.pi` volume. That was the wrong fix: it created a second
convention that could still diverge from the palace (keep the palace volume,
drop `devbox-pi-config`, and a scoped `mempalace sync` prunes every
conversation drawer, because dedup keys on the *staged* path). The feeder now
defaults to `<palace-root>/pi-stage`, resolved with mempalace's own
precedence (`$MEMPALACE_PALACE_PATH` → `$MEMPAL_PALACE_PATH` →
`~/.mempalace/config.json` → `~/.mempalace/palace`), so the stage inherits
whatever persistence the palace has and the two cannot be separated by
accident. No `ENV` and no entrypoint export: adding one back would
re-introduce exactly the split it removes.
- **Transcript inbox mount in `docker-compose.mempalace.yml`**
(`${MEMPALACE_FEED_DIR:-./feed}:/data/feed:ro`). A client cannot mine into a
remote palace directly: `mempalace_mine` expands its source path in the
*server* process, so the server can only see paths inside its own container.
Clients rsync their staged exports to a per-device subdirectory and then ask
the server to mine `/data/feed/<device>`. Read-only because mining only reads
sources — all locks live palace-side.
- **Smoke tests** for the above: `mempalace-pi-session` on `PATH`, two
assertions that the stage resolves next to the palace (default, and following
`$MEMPALACE_PALACE_PATH`), and two behavioural guards that feed the exporter a
synthetic pi session — one that must be captured, one abandoned session that
must not be. The second matters because pi expands skills/context into the
user prompt, so an abandoned session can look substantial by byte count while
containing no assistant output; and if pi's JSONL shape ever changes, the
exporter would silently capture nothing.
- **`.env.example`**: documents `MEMPALACE_FEED`,
`MEMPALACE_FEED_DEBOUNCE_MS`, `MEMPALACE_FEED_WING`, and the remote-palace
shipping vars `MEMPALACE_PI_SSH_TARGET`, `MEMPALACE_PI_REMOTE_PATH`,
`MEMPALACE_PI_DEVICE`.
### Changed
- **The "remote palace, no inbox" skip announces itself instead of vanishing**
(`entrypoint-user.sh`). When `MEMPALACE_REMOTE_URL` is set but
`MEMPALACE_PI_SSH_TARGET` is not, there is genuinely nothing the feeder can
ship to, so skipping is correct — but the branch was a bare `:`, and the skip
happens *before* the subshell that writes `~/.pi/agent/mempalace-catchup.log`.
A container in that state therefore contributed nothing to the palace and left
**no artifact at all**, not even an empty log, to explain why — indistinguish-
able from a healthy run that had nothing to file. Found while flipping the
first client onto the shared palace (2026-08-12), where it is the single most
likely way to end up quietly memory-less. The notice now goes to both the
container start output and that log path, names the two variables that fix it,
states that MCP tools still work (only *this* container's transcripts go
nowhere), and points at `MEMPALACE_FEED=0` for anyone who meant it.
Deliberately incapable of breaking startup: an unwritable `~/.pi` — root-owned
volume, a classic Docker accident — would make `mkdir -p` fail under `set -e`
and abort the whole entrypoint, so it degrades to stdout-only. That was a real
new risk, since this branch previously touched no filesystem whatsoever.
Covered by two smoke assertions against the entrypoint as shipped in the image
(the branch only runs at container start, so a `docker run` one-shot cannot
reach it).
### Bumped: pi 0.84.1 → 0.84.2
- **`ARG PI_VERSION=0.84.2`** (`Dockerfile.variant`), with the audit the pin
policy in that file requires.
**Headline for this image: the Bedrock tool-argument poison pill is FIXED
upstream.** The v1.6.4 entry below recorded it as *"Not fixed upstream … still
replayed unsanitised"* — that note is now superseded. pi-ai 0.84.2 adds a
recursive `sanitizeBedrockDocument()` and applies it at exactly the site that
entry named ([#7882](https://github.com/earendil-works/pi/pull/7882)):
```diff
- toolUse: { toolUseId: c.id, name: c.name, input: c.arguments },
+ toolUse: { toolUseId: c.id, name: c.name, input: sanitizeBedrockDocument(c.arguments) },
```
(`dist/api/bedrock-converse-stream.js` — line 692 in pi-ai 0.84.1, 704 in
0.84.2; it was 644 in 0.83.0 and 634 in 0.82.1.) The sanitiser drops object
members whose key is the empty string, recursing through arrays and nested
objects and preserving every valid value. It runs while the request is built,
so it covers the live turn *and* a resume: a session already bricked by an
empty-key tool argument now replays instead of dying on a Bedrock
`ValidationException`. **`pi-session-repair` (in `cli_utils`) is therefore no
longer the recovery path on this image.** It stays useful for older images and
for inspecting a transcript, because the stored `.jsonl` is still malformed —
the fix sanitises what is *sent*, not what was *recorded*.
**Why bumping `PI_VERSION` is the only way to get it:** pi publishes an
`npm-shrinkwrap.json`, which pins transitive dependencies *exactly*. pi
0.84.1's shrinkwrap pins `@earendil-works/pi-ai` to **0.84.1**, so although
0.84.1's `package.json` range is `^0.84.1` — which would otherwise admit
0.84.2 — rebuilding the old pin can never pick the fix up. Transitive upstream
fixes do not leak into this image; `PI_VERSION` is the whole gate.
Rest of the audit, against the integration surface the pin policy names:
- **Session `.jsonl` format — unchanged.** Identical
`migrateV1ToV2`/`migrateV2ToV3` ladder in both versions, so existing sessions
on the named volume load as-is and `pi-session-repair`'s parse target is
untouched.
- **Node engine floor — unchanged** at `>=22.19.0` (image ships 22.23.2).
- **pi-atelier — no change needed.** The pin stays `v0.8.0`: the hard floor is
"never pair < 0.7.1 with pi >= 0.84", this bump does not leave 0.84.x, and
atelier's `peerDependencies` (`>=0.80.7`) are satisfied. pi-atelier **0.8.1**
is published but deliberately NOT adopted here — one variable at a time, and
atelier is the component that has drawn blood at startup.
- **Directly relevant to `pi --ssh` use of this image:** 0.84.2 fixes split
`Alt+Enter` over SSH being misread as Escape, and adds `PI_TUI_ESC_TIMEOUT`
for high-latency terminals.
- **Keybindings — one surface worth knowing.** `pi-toolkit` ships exactly one
override, `tui.input.newLine: [shift+enter, ctrl+j, alt+j]`. 0.84.2's new
fullscreen transcript search (`Ctrl+Shift+F`) binds `Shift+Enter` to
*previous match* while its overlay is focused. Different context, so no
conflict is expected — but it is the one place the override meets a new
default, and the first place to look if "shift+enter stopped inserting a
newline" is ever reported.
- **New `defaultTools` setting** (choose startup built-in tools globally or per
project) is additive; `pi-toolkit`'s `settings.example.json` does not set it,
so the bootstrap template needs no change.
### Bumped: pi-atelier v0.8.0 → v0.8.1
- **`ARG PI_ATELIER_REF` / `ARG PI_ATELIER_VERSION` = `v0.8.1`**
(`Dockerfile.variant`), bumped *together* with `PI_VERSION` as that pin's
comment requires — and this pairing is a good advert for the rule, because both
sides touched fullscreen input handling within three days of each other.
atelier 0.8.1 (2026-08-12) is two changes, only one of them code: *"Preserve
fullscreen transcript mouse-wheel scrolling after Sidebar resize and visibility
changes by leaving Pi's persistent mouse reporting enabled"*, plus a README
simplification. The single source file that differs from 0.8.0 is
`src/split-pane.ts`. It extracts an `isPiFullscreenRenderer()` predicate and,
under pi's fullscreen renderer, stops writing its own
`\e[?1002h\e[?1006h` / `\e[?1006l\e[?1002l` pair around a sidebar resize —
previously it enabled mouse reporting on grab and disabled it on release, which
tore down the reporting **pi itself** had switched on and left the wheel dead
afterwards. Outside fullscreen it manages mouse mode exactly as before. It also
now captures the terminal it enabled mouse on and writes the disable sequence to
*that* terminal instead of to whatever `tui` currently points at.
**The audit that matters is the private-internals coupling**, since that is what
hung startup at 0.6.0/0.7.0. atelier reaches into three pi internals; all three
are unchanged in pi 0.84.2:
- **`TuiAltScreen`** — detected *by constructor name*, so a rename would
silently disable both the resize-input prioritisation and the new mouse
behaviour, with no error. Still
`class TuiAltScreen extends TuiBase implements ViewportTUI`.
- **`tui.inputListeners`** — a private `Set` that atelier deletes from and
re-adds to, to get its resize handler ahead of pi's viewport listener (which
"consumes every mouse event for text selection"). Still `inputListeners = new
Set()`, at the identical line 103 of `pi-tui/dist/tui.js` in both versions,
and still a `Set` — atelier guards with `instanceof Set`.
- **the prototype `render` descriptor** it wraps via `findPrototypeRender`.
Still an own `render(width)` on `TuiAltScreen`.
pi's mouse sequences are byte-identical between 0.84.1 and 0.84.2 (same
`1002h`/`1006h`/`1002l`/`1006l`/`1003h` occurrence counts), so atelier's
assumption about what pi leaves enabled still holds. `pi-tui`'s base class
changed additively only (one new `isOverlayFocused()`), and `TuiAltScreen`'s own
changes are the new search feature (`activeSearch`, `openSearch`/`closeSearch`,
the two search match styles, `copySelection`).
**Caveat, stated plainly:** pi 0.84.2 adds a focused fullscreen *search overlay*
that participates in input handling, while atelier reorders input listeners
around pi's viewport listener. The two look convergent — 0.84.2 separately fixes
*"focused fullscreen overlays not receiving mouse wheel or viewport scroll
keys"* — but this pairing is reasoned from the diffs, **not proven by
execution**: the CI smoke test does not drive the TUI, so a fullscreen
interaction regression would not be caught before pull. Worth an `alt+a` plus a
sidebar resize and a wheel scroll in fullscreen on first use of this image.
Version metadata is unchanged: `engines.node >=22.19.0`, `peerDependencies`
still the uninformative `>=0.80.7` on both pi packages (so still nothing in npm
metadata encodes the real floor), and still zero runtime dependencies — so the
"no `npm install` step" note above stays true. The GitHub tag `v0.8.1` exists
(commit `c31d7439`), which is what CI resolves to a SHA.
### Notes
- The `Dockerfile.base` change moves the base hash, so this needs a base
rebuild; the `~/.local/bin` self-heal exists so the feature does not have to
wait for one. The skip-notice change is in `entrypoint-user.sh`, which is
`COPY`d in `Dockerfile.base` too, so it rides the same rebuild — until then,
older images keep skipping silently and the two commands in the toolkit's
`phase-1-exposure-runbook.md` §3.7 are the way to tell.
- Requires the matching `mempalace-toolkit` change (`--prepare` two-phase
split, remote transport, and the auto-feed triggers in
`extensions/pi/mempalace.ts`). The split exists because the palace is
single-writer: a live pi session holds it through the extension's own
`mempalace-mcp`, so a CLI `mempalace mine` during a session fails with
"palace ... is held by PID". Staging is therefore done by the CLI and the
mine itself by whichever process already holds the palace.
---
## v1.7.0 — 2026-08-07
Minor release. Headline: **pi-atelier is now part of the image** — the TUI
sidebar/status rail every container previously had to hand-install — and **pi is
pinned to an audited version instead of tracking npm `latest`**.
*Why minor and not patch:* the policy above reserves patch for "pi version bumps,
smaller fixes" and minor for "new variants, significant base additions". Bundling
a new companion package into every image is the same shape as v1.1.0, which went
minor for bundling pi-studio; v1.4.0 likewise went minor for adding typst. This
release also adds a new build-arg pair, a new opt-out env var, and a settings
migration, so patch would understate it.
### Added
- **pi-atelier vendored at `/opt/pi-atelier`, pinned to `v0.8.0`** — the TUI
sidebar (ordered panels, split-pane, themes) is now part of the image instead
of something each user hand-installs. Vendored + registered at container start
by `entrypoint-user.sh`, the same pattern as pi-fork/pi-observational-memory/
pi-studio, and deliberately **not** `pi install npm:pi-atelier`: an npm
install writes into `~/.pi/npm-global` on the config volume, which shadows the
image and pins nothing — the footgun that once hid a missing `fork` tool for
six weeks. Unlike its siblings it gets **no `npm install`**: pi-atelier
declares zero runtime dependencies (only peerDeps, satisfied by the baked pi)
and has no build step, so pi loads its TypeScript straight from the checkout
(`pi.extensions` → `extensions/index.ts`).
- **A version FLOOR, encoded as an executable test.** pi-atelier 0.6.0/0.7.0
wrap pi's private TUI renderer in a way that recurses under pi 0.84: pi hangs
at startup with sustained CPU and no error message. Upstream fixed the
recursion in 0.7.1 and restored the non-overlapping split in 0.7.2
("avoiding the recursive render path that caused startup hangs and sustained
CPU usage"); 0.8.0 is additive on top of that. atelier's own
`peerDependencies` still say `>=0.80.7`, which does **not** express the floor,
so nothing in npm metadata could have warned us. `smoke-test.sh` and
`recreate-sanity-check.sh` now assert the pairing rule **pi ≥ 0.84 ⇒
pi-atelier ≥ 0.7.1** — verified against a 4×4 version matrix — so a bad
combination fails the build instead of publishing an image whose TUI never
starts. CI resolves the pinned tag to its **peeled commit SHA**; atelier uses
annotated tags, so the unpeeled ref is a tag object, not a commit (pi-studio's
lightweight tags never exposed that distinction).
- **`DEVBOX_ATELIER=0`** opts out: the entrypoint removes pi-atelier from pi's
`packages[]` instead of registering it. The switch lives in the entrypoint
rather than being "just run `pi uninstall`" because this component's failure
mode is *pi will not start*, which cannot be repaired from inside pi.
- **Migration for hand-installed copies.** A pre-existing `npm:pi-atelier` entry
is dropped from `packages[]` (with a `settings.json.bak.atelier.<ts>` backup)
so the pinned `/opt` copy takes over. This is not cosmetic: the registration
guard counts `npm:<name>` as already-registered, so without this step every
existing volume would have kept its unpinned npm copy — and a 0.6.x copy
alongside pi 0.84 is exactly the startup hang above. Only that one exact
string is removed; jq-parse failures or a missing file leave settings
untouched, and the backup prefix is distinct from the template merge's so two
rewrites in the same second cannot overwrite each other's backup.
### Changed
- **pi-toolkit's `pi-atelier.json` modernised to atelier's current schema**
(pi-toolkit `0e1369e`, cross-repo — it reaches the image through the pinned
`PI_TOOLKIT_REF` clone). The seeded config had been written against the pre-0.7
vocabulary: `segments` → `segmentLayout` with explicit per-segment visibility,
`ornament: "none"` → `{"id":"brand","visible":false}`, `showExtensionStatuses`
→ `{"id":"statuses","visible":true}`, plus the sidebar toggles that did not
exist when it was written (`showSidebarAgent`, `showSidebarTodos`, and
`showSidebarOnStartup`, new in atelier 0.8.0). Upstream still reads the old
keys, but only as non-authoritative legacy inputs, so the file worked while
silently missing every sidebar control added since. Verified by loading the old
and new file through pi-atelier 0.8.0's own `loadConfig()`: zero warnings from
each and an identical *effective* config, so it is a pure schema
modernisation — every deliberate choice (compact density, 60/85 context
thresholds, notifications off) is preserved. `sidebarPanelLayout` is left unset
on purpose so the panel set tracks upstream as atelier adds panels.
- **pi is now PINNED, not `latest`: `PI_VERSION=0.84.1`** (`Dockerfile.variant`).
CI's `resolve-versions` job used to resolve `@earendil-works/pi-coding-agent`
to npm `latest`, which meant every release silently adopted whatever pi had
shipped that morning — unaudited — in the same build that then got tagged and
published. A pi minor can move the private TUI/renderer internals pi-atelier
wraps (0.84 vs atelier 0.6.0: startup hang) or the session `.jsonl` format
`pi-session-repair` parses. **The pin is a checkpoint, not a freeze** —
bumping stays a routine one-line change; what stops is *unreviewed* adoption.
0.84.1 was audited for this release: theme/TUI additions are additive, the
session format is unchanged (`CURRENT_SESSION_VERSION = 3` in both 0.83.0 and
0.84.1, identical `migrateV1ToV2`/`migrateV2ToV3` ladder, so existing
transcripts are neither migrated nor at risk), and the Node engine floor is
unmoved at `>=22.19.0`.
- The pins live in the **Dockerfiles** and CI reads them from there (a
`checkout` was added to `resolve-versions`), so a local `docker build` and a
CI release ship the same versions by construction instead of by convention.
- CI **fails** the build when the pin is not a concrete version, and when the
pinned version is not actually published on npm — catching a typo, an
unpublished version, or one yanked after we audited it, at resolve time with
a clear message rather than as an `npm install` error mid-build.
- CI **warns** (`::warning::`, never adopts) when npm `latest` is ahead of the
pin, naming the newer version and what to re-check. That warning is the
prompt to audit and bump — not something to silence.
- **mempalace pin `3.5.0` → `3.6.0`** (`Dockerfile.base` `MEMPALACE_VERSION`),
in lockstep with opencode-devbox v2.9.0 as the pin's own comment requires.
3.6.0 (2026-07-17) is PyPI latest and is additive/reliability only — secure
`mempalace serve` remote mode, optional Milvus backend, atomic KG
`supersede()`, conversation chronology, mining exclusions, plus recovery and
locking fixes. **Reviewed for MCP tool-schema changes before bumping** — that
being the exact regression class this pin exists to catch, after an unpinned
install once swept in the broken 3.3.x/3.4.0 `diary_write` schema: there are
**none**, and nothing touches `diary_write`, so the perl workaround removed in
v1.2.2 stays removed. Two fixes are directly relevant to how this image uses
mempalace: read-only mode now covers `checkpoint` + `delete_by_source` in
`_MUTATING_TOOLS` (#1930), and agent attribution is preserved in
`mempalace_checkpoint` (#2023/#2034) — the latter matters because the diary
protocol relies on per-agent attribution. Rebuilds the base image.
### Documentation
- **New README section: "Using pi-atelier (TUI sidebar)"** — what the status rail
and sidebar give you, the `alt+a` / `/atelier` entry points, session-scoped
`/atelier sidebar on|off` versus persistent Save, and `DEVBOX_ATELIER=0` to opt
out. Plus the config story: why `~/.pi/agent/pi-atelier.json` is **copied, not
symlinked** (atelier saves via write-temp-then-`rename(2)`, and `rename`
replaces a symlink rather than following it, so a symlink would silently detach
on the first save), why `install.sh` therefore only seeds it when absent, which
keys are current versus legacy-compatibility, and the 92-column auto-hide /
64-column main-pane floor so a narrow terminal degrades gracefully.
- **Documents how to authenticate the container to a LAN peer with its own key**
(README: *Giving the container its own key for a peer*) — the gap the existing
*Naming LAN peers* section left open. That section explained `ProxyJump`
*routing* while asserting `HostName`/`User`/`IdentityFile` are "inherited from
the matching block in your real `~/.ssh/config`", which is precisely what fails
in a container: host keys are normally passphrase-protected and unlocked by the
macOS Keychain or an `ssh-agent`, neither of which exists here, so the key can
never be decrypted — `Permission denied (publickey)` while the identical
`ssh peer` works fine in a host terminal — and `~/.ssh` is read-only, so no
usable key can be added there either. The new walkthrough (throwaway example
keys) covers a passphraseless keypair in the `devbox-ssh-local` volume so it
survives `--force-recreate`; a hardened `authorized_keys` line (`restrict`,
`from=`, optional `permitopen`); the non-obvious detail that `from=` must allow
the **host's** addresses, plural, because container egress is NAT'd through the
host and a roaming laptop presents a different one per network (a `from=`
mismatch is indistinguishable from a wrong key in the error message); the
`IdentityFile` override in the host-owned `ssh-lan.conf`; and verification with
`-o ControlPath=none` so a warm ControlMaster cannot fake a pass. States
explicitly that no private key is in the published image — the volume is
created at runtime on the operator's own machine.
- **Corrects two claims in *Naming LAN peers***: (1) `ssh-lan.conf` is not
`ProxyJump`-only — it is `Include`d before `~/.ssh/config`, so by
first-value-wins *any* option set there wins, which is what makes the
`IdentityFile` override above possible; (2) "newly added peers work
immediately, no container or session restart needed" holds only for *edits* to
an existing file. Creating it for the first time **does** need one restart,
because `setup-lan-access.sh` emits the
`Include ~/.config/devbox-shell/ssh-lan.conf` line only
`if [ -r "$SSH_LAN_CONF" ]` at container start — until then ssh never reads it,
which presents exactly as "my override is being ignored".
- **Adds *macOS-only keywords in a shared `~/.ssh/config`***. The same file is
read by macOS ssh and by the container's Linux OpenSSH, where macOS-only
keywords are fatal rather than ignored: one `UseKeychain yes` in a `Host *`
block yields `Bad configuration option: usekeychain` /
`terminating, 1 bad configuration options` and takes down `dssh`/`dscp`,
`pi --ssh`, `scp` and every helper that shells out to ssh — while the host
keeps working, so it presents as a container regression rather than a host
config error. Fix is `IgnoreUnknown UseKeychain` *ahead of* the keyword (macOS
still honours it, Linux skips it), plus keeping such a `Host *` block below
OrbStack's `Include ~/.orbstack/ssh/config`, which documents in its own comment
that it must come first.
- Documents **per-variant image description labels** (committed and pushed after
the v1.6.4 tag without a changelog entry). Both published variants used to
inherit `Dockerfile.base`'s `description="pi-devbox — base image
(variant-independent)"`, so `v1.6.4` and `v1.6.4-studio` both advertised
themselves on Docker Hub as *the base image* — misleading, and useless for
telling the two apart. Since a `LABEL` cannot branch on `INSTALL_STUDIO`, the
text now arrives as a build arg: CI passes a variant-specific string
(interpolating `RELEASE_TAG`, `PI_VERSION`, and `STUDIO_TAG` for studio), while
the `Dockerfile.variant` default keeps a bare local `docker build` honest
rather than misleading. Sets `org.opencontainers.image.title`/`.description`
alongside the legacy bare `description` key so both Hub and OCI-aware tooling
see it. The ARGs stay in the last-declared block, so the label layer remains
the only thing invalidated.
---
## v1.6.4 — 2026-07-30
Patch release. Headline: **the `fork` tool has never once loaded since v1.0.0**
and now does — plus pi `0.82.1` → `0.83.0`, audited clean against every baked
extension.
### Fixed
- **`pi-fork` was never registered — the `fork` tool has been missing since
v1.0.0.** `entrypoint-user.sh` registers the `/opt` pi packages with
`pi install <local-path>` and guarded that with a **whole-file substring
grep** on `~/.pi/agent/settings.json`. But `settings.example.json` carries a
top-level `"pi-fork"` **config block** (the fork effort profiles, added in
`pi-toolkit` `adb6907`, 2026-06-17), so `grep -q pi-fork settings.json`
matches on any settings file bootstrapped from — or template-merged with —
that template. The guard therefore concluded "already installed" and
`pi install /opt/pi-fork` never ran, on fresh *and* preserved volumes.
Compounding it, the non-destructive template merge runs **earlier in the same
startup** than the install loop, so the very mechanism that delivers new
template keys to an old volume is what plants the string that defeats the
guard. `pi-observational-memory` and `pi-studio` escaped only by luck: the
template key is `observational-memory` (no `pi-` prefix) and there is no
studio block.
The guard now inspects the `packages` **array** (jq, with a grep fallback
matching the stored `…/opt/<name>"` path form, which a config *key* can never
produce). Existing volumes self-heal on the next container start — the guard
returns false, `pi install /opt/pi-fork` runs, and `fork` registers on the
following pi start or `/reload`. No image rebuild is required to benefit if
you run `pi install /opt/pi-fork` by hand.
- **Both test suites asserted the bug as green.** `scripts/smoke-test.sh` and
`scripts/recreate-sanity-check.sh` checked registration with the *same*
whole-file grep, so "pi-fork registered (fork tool)" passed on every build
and every recreate while the tool was absent. Both now assert against
`packages[]` with the same predicate as the entrypoint guard, and the labels
say `packages[]` so the distinction is visible in CI output. The smoke-test
readiness wait loop was switched to the array check too (and to
`docker exec -u developer` + `$HOME` instead of a hard-coded
`/home/developer` path).
Detected by an agent session noticing `fork` was absent from its own tool
list on v1.6.3; zero `fork` calls exist across the 19 sessions on this
volume, confirming it never once loaded.
### Changed
- **pi `0.82.1` → `0.83.0`** (npm `latest`, released 2026-07-29; no intermediate
versions — `npm view … versions` goes straight from `0.82.1` to `0.83.0`).
Variant-only rebuild: pi is installed in `Dockerfile.variant`, so the
content-addressed `base-<hash>` is unaffected.
**0.83.0 ships a Breaking Change, and it cannot reach this image.** Upstream:
> Upgraded bundled TypeBox aliases to 1.3.7, removing deprecated APIs
> including `Type.Base`, `Type.Awaited`, `Type.Promise`, `Type.AsyncIterator`,
> `Type.Iterator`, `Type.Options`, and `Value.Mutate`, while fixing compiled
> validation of nullable array tool arguments. Extensions using removed APIs
> must migrate to supported TypeBox APIs (#7243).
Audited per baked extension: **`pi-fork`** vendors its own
`@sinclair/typebox@0.34.52` — a *differently named* package than the `typebox`
pi bundles (1.1.38 → 1.3.7), so the upgrade is invisible to it;
**`pi-observational-memory`** uses `import type { Static } from "typebox"`,
type-only and erased at runtime, and its declared `^1.1.38` admits 1.3.7;
**`pi-studio`** and **`pi-atelier`** use TypeBox not at all. A grep for
`Type.(Base|Awaited|Promise|AsyncIterator|Iterator|Options)|Value.Mutate`
across all four returns **zero hits**. Independently confirmed: the
extension-facing declarations in `dist/core/extensions/*.d.ts` are
**byte-identical** between 0.82.1 and 0.83.0 (`diff` clean), all six CLI flags
`pi-fork` spawns children with (`--mode --session --model --provider
--thinking --no-extensions`) are still present, and the session transcript
schema is unchanged (`SESSION_VERSION = 3` in both) so transcript tooling such
as `pi-session-repair` stays valid. **No `PI_VERSION` pin was needed.**
Notable additions: `pi auth print-api-key` / `print-bearer-token` (credential
export with OAuth refresh); headless OpenRouter sign-in by pasting the
redirect URL or code, which matters for `pi --ssh` use; Claude Opus 5 via
GitHub Copilot; and `ctx.scopedModels` exposed to extensions.
Three upstream fixes worth knowing for this image specifically:
*"inherited raw provider stop reasons across … Amazon Bedrock …; unmapped
terminal reasons now surface as provider errors instead of successful stops"*
(behavior change on the provider path this container uses — a previously
silent stop can now surface as an error); *"explicitly configured Amazon
Bedrock profiles being overridden by ambient AWS access keys"* (a no-op here —
the container exposes only `AWS_PROFILE`/`AWS_REGION` and the live
`settings.json` has no `providers.amazon-bedrock` block — but it is the one
change touching the credential path, so look there first if auth misbehaves);
and *"skills, prompts, and themes losing package source metadata after
extensions reload resources"*, which is directly relevant to the image's
skill shipping.
**Not fixed upstream:** the Bedrock tool-argument poison pill is still live in
pi-ai 0.83.0 — `toolUse: { toolUseId, name, input: c.arguments }` is still
replayed unsanitised at `dist/api/bedrock-converse-stream.js:644` (it was
line 634 in 0.82.1; the file still has zero empty-member-name sanitisation).
`pi-session-repair` (in `cli_utils`) remains the recovery path.
- **Settings template now defaults to Claude Opus 5** (`pi-toolkit` @ `926f738`).
`settings.example.json` — the file `entrypoint-user.sh` bootstraps
`~/.pi/agent/settings.json` from — moves `defaultModel` and the `pi-fork`
**deep** tier from `eu.anthropic.claude-opus-4-8` to
`eu.anthropic.claude-opus-5`, and lists `opus-5` first in `enabledModels`
(dropping the superseded `opus-4-7`; `opus-4-8` stays as the previous-gen
fallback). `fast` = `haiku-4-5` and `balanced` = `sonnet-5` are unchanged.
Opus 5 shipped to users in v1.6.3 via pi `0.82.1`, but nothing in the image
actually pointed at it. **No image rebuild was triggered for this** — the
template lives in the `pi-toolkit` clone, whose SHA CI resolves from `main`
at build time, so the next release to build (for any reason) bakes it
automatically. Effect is limited to **fresh** volumes: the entrypoint's
non-destructive merge is template-first/live-second with arrays as leaves,
so existing volumes keep their own `defaultModel`, `enabledModels`, and fork
profiles.
### Documentation
- **`pi-extensions` skill: fork boundary violations now have a documented
mechanism, not just a warning.** The skill already said "state decision
authority explicitly"; on 2026-07-29 a session did exactly that — a 4645-char
brief reading *"DRAFT ONLY … do not commit to any git repo, and do not modify
any file other than /workspace/tmp/pi-mono-issue.md"* — and the fork came back
with *"All three done: Pushed … Moved … symlinked"*. Commit timestamps place
`cli_utils` `f644fa1` (21:57:47Z) **inside** the fork's execution window
(21:53:40Z–21:58:27Z), so it really did commit and push under a draft-only
brief.
The cause is structural: `pi-fork/src/index.ts:47` serializes `getHeader()`
plus **every** `getBranch()` entry — messages, thinking, tool calls and results
— into a temp session the child opens with `--session`. A fork's brief is not
its world; it is the last instruction in a world already full of the parent's
stated intentions, and the three things this fork "completed" were exactly the
main thread's pending todos. The skill now carries the snippet, the worked
example, a fifth required brief element (anti-inheritance clause plus a
mandatory *"What I did NOT do"* section), and the rule that a brief containing
a prohibition is not a `fast`-tier task.
Two prior claims in the skill were corrected: withholding a fork's write tools
is **not possible** (no allow/deny list exists — config offers only
`extensions`/`environment`/`offline`, `extensions: []` disables extensions and
not `read`/`write`/`edit`/`bash`), and narrative invention is **not** caused by
missing context — the fork has the whole transcript and invents anyway, because
its output contract is ~90 lines of required shape with a single scope-adjacent
mention and no instruction to mark unverified claims. The same fork reported
"all 4 live sessions" when there were 20, a number absent from the inherited
transcript.
Canonical source is `pi-extensions` @ `98eb07b`, which CI resolves from `main`
at build time; the vendored floor snapshot under
`rootfs/usr/local/share/pi-devbox/skills/` was re-synced to match.
## v1.6.3 — 2026-07-25
Patch release. Headline: **pi `0.81.1` → `0.82.1`** (npm `latest`) — the first
pi bump since v1.6.1.
### Changed
- **pi `0.81.1` → `0.82.1`.** CI resolves `pi@latest` at build time; latest is
now `0.82.1` (via `0.82.0`). pi is installed in the **variant** layer
(`Dockerfile.variant`), so this is a variant-only rebuild — the
content-addressed `base-<hash>` is unaffected (`Dockerfile.base`, `rootfs/`,
`entrypoint*.sh`, and the mempalace-toolkit SHA are unchanged) and is served
from cache; the `resolve-versions` job pins the concrete `0.82.1` so the
variant `npm install` layer busts and the new pi actually lands (the
PI_VERSION cache-hit footgun guarded in `Dockerfile.variant`). Both `0.82.0`
and `0.82.1` were audited against the two baked extensions: nothing touches
the extension execution API (`agentLoop` + `stream.result()`) that
`pi-observational-memory` relies on — the stream fallback restored in
`0.81.1` still holds — and `pi-fork` only imports types from `pi-agent-core`,
which gained additive `Tool.constrainedSampling` / capability flags with no
breaking changes. The Node engine requirement is unchanged (`>=22.19.0`; the
base ships `22.23.1`). Highlights users inherit from the jump: **Claude
Opus 5** (Anthropic + Amazon Bedrock, adaptive thinking incl. `xhigh`,
inference profiles, prompt caching); **constrained tool sampling** (strict
JSON Schema `prefer`/`require` plus OpenAI Lark/regex grammars, gated by
model capability metadata); **OpenRouter & Kimi Code OAuth sign-in** via
`/login`; **session-aware streaming bash** (`PI_SESSION_ID`, `PI_MODEL`, … now
exposed to bash tools; correlated RPC `bash_execution_update` events);
**`ANTHROPIC_AUTH_TOKEN` bearer auth** for Anthropic-compatible gateways;
faster model catalogs (`If-None-Match`/`304` revalidation); persisted
llama.cpp model catalogs; and a bundled **`protobufjs` 7.6.5** security bump
(GHSA-j3f2-48v5-ccww). See the [pi changelog][pi-changelog] for the full
list.
## v1.6.2 — 2026-07-23
Patch release. **Completes the v1.6.1 studio publish.** CI-only change; the
shipped image content is identical to v1.6.1 apart from the bumped `pi`
version resolution at build time (still `0.81.1`).
> **Note on v1.6.1.** Ran on 2026-07-23; the non-studio variant (`v1.6.1`,
> `latest`, `base-latest`) shipped cleanly, but the studio variant was blocked
> in the smoke-studio job by a size assertion that was still calibrated for
> the pre-`agent-browser` baseline. `v1.6.1-studio` and `latest-studio` were
> never pushed; `latest-studio` on Hub still points at v1.5.0-studio until
> v1.6.2 lands. Users who pull `joakimp/pi-devbox:v1.6.1` today get a valid
> non-studio image with `pi 0.81.1` baked; there is no `v1.6.1-studio` image.
### Fixed (CI)
- **`scripts/smoke-test.sh`: raise `SIZE_THRESHOLD_MB` from `3500` to `3800`.**
The 3500 threshold was set in v1.0.0 based on a local arm64 build measured
at 3.20 GB plus a `+300 MB` margin. v1.6.0 baked in `agent-browser` +
Playwright Chromium (~291 MB net, documented in v1.6.0's entry) but the
threshold was never updated — v1.6.0 never ran to smoke because of the
site-network fault, so nothing surfaced the miscalibration until
run 512 (v1.6.1) reached smoke-studio and reported
`3574 MB exceeds threshold 3500 MB`. Actual CI amd64 sizes observed on
run 512: **3411 MB non-studio**, **3574 MB studio**. The new 3800 MB
ceiling carries ~225 MB margin above the studio number — enough to absorb
minor arch/build-cache variance and small future growth, still tight
enough to catch a genuine +GB regression. The comment above the constant
is refreshed to reflect the new baseline (agent-browser included, run 512
actuals). Not base-affecting; base hash unchanged.
- **`scripts/smoke-test.sh`: don't hard-code a `v` prefix on `release_tag`
in the `pi-devbox-version` human-output assertion.** (Landed on the
retagged `v1.6.1` and carried forward in `v1.6.2`.) The smoke workflow
deliberately passes `RELEASE_TAG=smoke` / `RELEASE_TAG=smoke-studio` to
the variant build so smoke images don't collide with real `vX.Y.Z` tags,
and `pi-devbox-version` correctly prints `pi-devbox smoke`. The prior
assertion required the literal substring `pi-devbox v` — only true for
real releases — so it fired on every smoke run once it existed. The two
neighbouring assertions on `--json` and `--quiet` already cover the value
of `release_tag`; the human-output assertion now only verifies that the
line renders (substring `pi-devbox ` — note the trailing space). Never
fired before because `pi-devbox-version` was added post-v1.5.0 and every
CI attempt since was blocked before smoke ran.
## v1.6.1 — 2026-07-22
Patch release. Headline: **pi `0.80.6` → `0.81.1`** (npm `latest`) — the first
pi bump since v1.5.0.
> **Note on v1.6.0.** The `v1.6.0` git tag was cut on 2026-07-17 (agent-browser +
> `pi-devbox-version`, see below) but never reached Docker Hub: the variant
> publish was blocked by an intermittent SYN-drop fault on the on-prem CI
> network (`ci-network-diagnosis.md`, since resolved). v1.6.1 lands v1.6.0's
> content **plus** the pi bump in one release; there is no `v1.6.0` image on
> Docker Hub. The `v1.6.0` git tag is left in place as an accurate record of
> what was intended on that day.
### Changed
- **pi `0.80.6` → `0.81.1`.** The CI resolves `pi@latest` at build time; latest
is now `0.81.1`. The intermediate `0.81.0` is deliberately skipped: 0.81.0
removed the default stream fallback for extensions using the pre-0.81
`@earendil-works/pi-agent-core` API, which `pi-observational-memory` relies
on (`agentLoop` + `stream.result()` in the observer/reflector/dropper
agents). 0.81.1 restored the fallback ([earendil-works/pi#6915][pi-6915]),
making 0.81.1 — but not 0.81.0 — a safe drop-in. `pi-fork` only imports
types from `pi-agent-core` and is unaffected. Everything since v1.5.0's
baked `0.80.6` (i.e. `0.80.7`–`0.80.10`, `0.81.0`, `0.81.1`) was audited for
breaking changes against the two baked extensions — none affect this image.
The Node engine requirement rose to `>=22.19.0` in `0.81.0`; the base still
ships `22.23.1` (nodesource 22.x), so no engine bump is needed. Highlights
users inherit from the upstream jump: **local llama.cpp router support**
(search + download Hugging Face models, explicit load/unload, live
progress); **full pi-ai provider extensions** (extensions can now register
complete providers with native auth, model refresh, filtering, and
streaming); **Qwen Token Plan** subscription providers; **resilient
compaction / branch-summary retries** on transient provider failures with
lifecycle events exposed to interactive, JSON, RPC, and SDK consumers;
expanded usage accounting for tools, compaction, and branch summaries.
Base-affecting (npm install line rebuilds), so `base-<hash>` rebuilds. See
the [pi changelog][pi-changelog] for the full list.
[pi-6915]: https://github.com/earendil-works/pi/issues/6915
[pi-changelog]: https://github.com/earendil-works/pi/blob/main/CHANGELOG.md
## v1.6.0 — 2026-07-13
> ⚠️ **Never published to Docker Hub.** Tagged in git on 2026-07-17 but the
> variant publish was blocked by a site-network fault before the image reached
> the registry. Superseded by v1.6.1, which carries this release's content
> forward alongside the `pi 0.81.1` bump.
### Added
- **`agent-browser` — headless browser automation, baked into every variant.**
The base now ships the [`agent-browser`](https://www.npmjs.com/package/agent-browser)
CLI plus a Playwright-fetched Chromium, so the agent can drive a real browser
(open/click/fill/`eval`/screenshot/snapshot) and *verify* front-end work
involving live DOM or WebGL instead of guessing. The `agent-browser` skill
(from the skillset repo) was previously a no-op because the binary was
absent; it now works out of the box. Two pieces: the standalone Rust CLI
(npm, `NPM_CONFIG_PREFIX=/usr` so it survives the `~/.pi/npm-global` volume),
and a Chromium fetched via `playwright install --with-deps chromium` into
`PLAYWRIGHT_BROWSERS_PATH=/usr/local/share/ms-playwright` (a system path,
never shadowed by the `/home/developer` volume — unlike agent-browser's own
`~/.agent-browser/browsers` default). A stable `/usr/local/bin/agent-chrome`
symlink, exported as `AGENT_BROWSER_EXECUTABLE_PATH`, insulates the config
from Playwright's per-version `chromium-<rev>` directory name. Debian trixie
`--with-deps` dependency resolution verified (the t64 renames are handled).
The global AGENTS.md managed block
(`rootfs/usr/local/share/pi-devbox/pi-global-AGENTS.append.md`) gains a short
pointer so agents discover the capability. Adds ~625 MB (Chromium; Playwright's
unused headless-shell build is dropped and the apt/npm caches cleaned in-layer
to stay lean). Base-affecting, rebuilds `base-<hash>`.
- **`pi-devbox-version` command.** Wraps `/etc/pi-devbox/build-manifest.json`
into a human-readable summary (release tag, build date, source revision,
baked `pi_version`, and short SHAs for every `/opt` component) instead of
requiring users to know the manifest path and pipe it through `jq`
themselves. Also flags **live drift** — if `pi --version` no longer matches
what was baked at build time, the `pi:` line calls that out rather than
silently trusting the manifest. `--json` dumps the raw manifest for
scripting; `--quiet` gives a one-line `release_tag (source_revision)` form.
Printed automatically once at container start (`entrypoint-user.sh`, before
the rest of the setup output), and stays available on demand for the rest
of the session. Exits 1 with a short notice — rather than failing silently
— on images built before this file existed. Base-affecting (new
`rootfs/usr/local/bin/pi-devbox-version`), rebuilds `base-<hash>`.
### Changed
- **Bundled `pi-toolkit` settings template: `pi-fork` balanced tier bumped to
`eu.anthropic.claude-sonnet-5`** (was `claude-sonnet-4-6`), matching the model
now in use. The image clones `pi-toolkit@main` into `/opt/pi-toolkit` at build
time, so the next build bundles it automatically (pi-toolkit `0010417`); the
same commit also refreshes the template's `enabledModels` and the README
examples. Seed-only: existing containers keep their live `~/.pi/agent/settings.json`
(the entrypoint merge is live-wins), so only fresh `~/.pi` volumes are affected.
## v1.5.0 — 2026-07-13
### Added
- **Seeded global gitignore now ignores `**/.claude/settings.local.json`.** Claude
Code's per-machine local settings file holds machine-specific permissions and
can carry credentials, so it should never be committed. The seed
(`rootfs/home/developer/.gitignore_global`, baked to `/etc/skel-devbox/`) gains
the pattern so fresh containers match a host global that already ignores it.
Existing containers are unaffected (the seed is copied only when
`~/.gitignore_global` is absent); their file can be updated by hand. Base-
affecting (`Dockerfile.base` COPY of the seed), rebuilds `base-<hash>`.
- **Readable Neovim colours out of the box.** The base now ships a system-wide
Neovim config (`/etc/xdg/nvim/sysinit.vim`) that enables `termguicolors`,
plus the `kitty-terminfo` package. Vanilla Neovim otherwise fell back to a
256-colour palette over ssh/kitty and rendered strings and comments in a
muddy, low-contrast dark colour. `sysinit.vim` is Neovim's system vimrc: it
loads for every user before any personal `~/.config/nvim` and can still be
overridden per-user (`:set notermguicolors`, or your own init). Base-affecting
(`Dockerfile.base` apt package + COPY), rebuilds `base-<hash>`.
- **Terminal support beyond kitty: `ncurses-term` + a compiled `xterm-ghostty`
alias.** The base previously shipped only `ncurses-base` (xterm-256color,
tmux), so SSHing in from a modern emulator degraded to a dumb fallback. The
base now installs `ncurses-term` (terminfo for WezTerm, Alacritty, foot, st,
and the base `ghostty` entry, among many others) and compiles an
`xterm-ghostty` alias with `tic -x` (`use=ghostty`) — Ghostty connects as
`TERM=xterm-ghostty` and no distro packages that name. Combined with
`kitty-terminfo` (xterm-kitty) and xterm-256color (iTerm2's default, already
in ncurses-base), the common modern terminals now resolve their TERM. The
approach mirrors the maintainer's ansible `common` role. Base-affecting
(`Dockerfile.base` apt + COPY + `tic` RUN, plus a new
`rootfs/usr/local/share/terminfo-src/ghostty.terminfo`), rebuilds `base-<hash>`.
- **Repository hygiene: `LICENSE`, `THIRD_PARTY.md`, and `.dockerignore`.** The
repo declared MIT only in prose; it now ships an actual `LICENSE` file (MIT,
© Joakim Persson) plus `THIRD_PARTY.md` recording that the published images
bundle third-party software under its own terms (pi, pi-fork,
pi-observational-memory, pi-studio — all MIT; gosu Apache-2.0; Debian packages
under their respective licenses). A new `.dockerignore` trims the build
context to what the Dockerfiles actually `COPY` (`rootfs/` + `entrypoint*.sh`),
keeping `.git`, docs, `scripts/`, and compose files out — cheaper context and
no risk of a future broad `COPY` pulling in `.git`. Not base-affecting (the
base hash covers only `Dockerfile.base` + `rootfs/` + `entrypoint*.sh`);
image contents are byte-identical.
- **Dockerfile linting (`hadolint`) in CI, plus an `IDEAS.md` backlog.** The
lint workflow already ran actionlint + shellcheck on `run:` steps but never
looked at the two Dockerfiles that are the heart of the project. A new
`hadolint` job (pinned v2.14.0, same download-pin pattern as actionlint) lints
`Dockerfile.base` and `Dockerfile.variant`; `.hadolint.yaml` grandfathers the
deliberate choices (unpinned apt/npm, `cd`-in-`RUN`, `SC2086` — mirroring the
existing shellcheck excludes) and fails on anything new at `warning`+.
`IDEAS.md` parks the vetted-but-unscheduled follow-ups (SHA-pin CI actions,
trivy scanning, buildx SBOM/provenance attestations, a local `Makefile`,
renovate). Repo/CI only — not baked into the image.
### Changed
- **`-studio` images now pin pi-studio to its newest *semver tag* instead of
`main` HEAD.** Upstream `omaclaren/pi-studio` abandoned GitHub *Releases* at
v0.5.55 but keeps tagging every version (currently `v0.9.36`) and pushing to
`main`; tracking `main` HEAD risked baking half-finished commits that land
after a tag. CI (`resolve-versions`) now lists every tag via a single
`git ls-remote` (the REST tags API paginates at 100 and the repo already has
>140 tags), selects the highest `X.Y.Z` with `sort -V` (pre-releases
excluded by a strict filter), and pins that tag's commit SHA into
`PI_STUDIO_REF`. Pinning the SHA (not the moving tag) preserves cache-busting
and reproducibility, is what `require_sha` demands, and is recorded in the
`se.jordbo.pi-devbox.pi-studio-ref` image label. The human-readable tag (e.g.
`v0.9.36`) is now also recorded in a new `se.jordbo.pi-devbox.pi-studio-version`
label for at-a-glance identification (`docker inspect`). Studio-variant only —
not base-affecting; takes effect on the next `-studio` build. No change to the
resolved commit today (`v0.9.36` == current `main` HEAD).
### Fixed
- **`pandoc --pdf-engine=typst` now works without `-V mainfont`.** pandoc's
bundled typst template (`/usr/share/pandoc/data/templates/template.typst`)
defaults the document font to an empty tuple (`font: ()`), so a naked
`pandoc --pdf-engine=typst` (and `studio_export_pdf` in some cases) failed
with `error: font fallback list must not be empty` unless the caller passed
`-V mainfont="..."`. The base now patches that template default to
`Libertinus Serif` (typst's own bundled default font) at build time, so PDF
export works out of the box. Base-affecting (`Dockerfile.base` RUN), rebuilds
`base-<hash>`. README gains a "Generating a PDF with pandoc + typst" section
with the working command and how to override the font via `-V mainfont`.
---
## v1.4.0 — 2026-07-11
Minor release. Headline: **PDF export works out of the box** — the base now
ships **`typst`** as the pandoc PDF engine (`pandoc --pdf-engine=typst`), so
`studio_export_pdf` / `pandoc -o out.pdf` no longer fail with "xelatex not
found". Also adds a **host SSH reachability check at shell startup**. Both are
base-affecting (`Dockerfile.base` apt+RUN for typst/xz-utils; `.bash_aliases`
for the SSH check is COPYd into the base), so the base rebuilds and both land in
`base-<hash>`. pi auto-resolves `latest` at build time (0.80.3 → 0.80.6);
mempalace stays pinned at 3.5.0 (current PyPI latest).
### Added
- **Host SSH reachability check at shell startup.** `~/.bash_aliases` (baked
into the image) now runs a one-time SSH probe on the first bash session of
each container. If the Mac host is not reachable (Remote Login disabled or
the `devbox_jump` key not yet authorized) it prints a clear warning with the
exact two steps to fix it, including the container's public key inline.
Subsequent shells in the same container skip the check (flag in `/tmp`,
cleared on recreate). Silent when SSH is working. Complements the existing
key-generation message in `setup-lan-access.sh` which only fires once at key
creation time and can easily be missed. Commit `4563b4d`.
- **`typst` — lightweight PDF engine for pandoc (Markdown→PDF).** `pandoc` has
shipped in the base since v1.0.0 but as a front-end only — with no PDF
back-end installed, `studio_export_pdf` / `pandoc -o out.pdf` failed with
"xelatex not found". The base now installs `typst`, a single ~30 MB static
Rust binary (no LaTeX), used via `pandoc --pdf-engine=typst`. Chosen over a
~600 MB TeX Live install; a fuller TeX Live remains the higher-fidelity
fallback for anyone needing LaTeX-exact output (install on demand). Also adds
`xz-utils` to the apt layer (typst ships a `.tar.xz` asset that `tar` needs
`xz` to extract). Installed with the standard `latest` GitHub-release idiom;
pin with `--build-arg TYPST_VERSION=vX.Y.Z`. This lands in `base-<hash>`
(Dockerfile.base changed). Supersedes the previously-planned
`:latest-studio-tex` variant — typst is small enough to ship in BASE, so no
separate TeX variant is needed. See `pi-devbox-roadmap`.
---
## v1.3.0 — 2026-07-02
Minor release. Headline: **shared/external MemPalace** — the `mempalace.ts`
bridge can now point at one MemPalace HTTP server (`MEMPALACE_REMOTE_URL`,
optional `MEMPALACE_REMOTE_TOKEN`) shared across containers/harnesses instead of
a per-container local palace; ships `docker-compose.mempalace.yml` for the
server. Also ships the **`nano` + `micro`** non-modal editors and a **CI
workflow-lint layer** (Gitea-accurate sh-vs-bash guard + actionlint/shellcheck),
with the `docker-publish.yml` bash-defaults and `promote-base-latest` shell
fixes. pi stays `0.80.3`; the base image rebuilds (the mempalace-toolkit ref
advanced and `Dockerfile.base` gained nano/micro), so the new bridge and editors
land in `base-<hash>`.
### Added
- **Share one MemPalace across containers via `MEMPALACE_REMOTE_URL`.** The
`mempalace.ts` bridge (from `mempalace-toolkit`) can now connect to a shared
MemPalace over HTTP instead of spawning a per-container local server: set
`MEMPALACE_REMOTE_URL=http://<host>:8765/mcp` (optionally
`MEMPALACE_REMOTE_TOKEN`) in `.env` and no local `mempalace-mcp` is spawned.
A new `docker-compose.mempalace.yml` stands up such a shared server
(`mempalace-mcp --transport http`). Leaving the URL unset keeps the default
local-per-container palace. See `.env.example`. (The HTTP transport is
unauthenticated — keep it on a trusted network or behind a reverse proxy.)
- **Two non-modal terminal editors alongside `nvim`: `nano` and `micro`.**
The image previously shipped only `nvim` (with `EDITOR=nvim`), a modal
vi-style editor. Not everyone is comfortable with vi keybindings, so both
a classic and a modern non-modal option now ship:
- **`nano`** (apt) — ~2.8 MB installed. Its dependencies (`libc6`,
`libncursesw6`, `libtinfo6`) are already present via `nvim`/`less`/`htop`/
`tmux`, so it pulls in **no extra packages**. On-screen shortcut hints
(`^O` write, `^X` exit) make it the lowest-friction fallback.
- **`micro`** — ~12 MB, a single static Go binary installed from GitHub
releases (same pattern as `bat`/`eza`/`zoxide`). Desktop-style keybindings
(`Ctrl+S` save, `Ctrl+Q` quit, `Ctrl+C/V/X`, `Ctrl+Z` undo), mouse
support, and syntax highlighting out of the box. Pin with
`--build-arg MICRO_VERSION=vX.Y.Z`; defaults to `latest`.
Combined footprint is ~15 MB (<0.5% of the ~3.2 GB image). **`EDITOR`
stays `nvim`** — the new editors are opt-in via `export EDITOR=micro`
(or `nano`) and/or `git config --global core.editor micro`.
Note: micro's upstream repo moved `zyedidia/micro` → `micro-editor/micro`;
the Dockerfile uses the canonical URL because the old org's
`/releases/latest` redirect lands on another `/latest` URL (the org
rename), which would defeat the tag-parsing `latest`-resolution idiom.
These are base-image additions, so they only land once the `base-<hash>`
rebuilds (this file changed, so the next build picks them up).
### Added (CI)
- **Workflow lint (`.gitea/workflows/lint.yml`) running on every push and PR.**
Two complementary checks, so CI-workflow bugs are caught before an expensive
build runs:
- **`scripts/check-workflow-shell.sh`** — a Gitea-accurate guard that fails
if any `run:` step doesn't resolve to `bash` under Gitea's real defaults.
This catches the exact recurrence class (omit `shell:`, use bash syntax),
which **actionlint alone does not** — actionlint models GitHub Actions
(default shell = bash) and so assumes a shell-less step is bash, whereas
Gitea's default is `sh`/dash.
- **`actionlint` + `shellcheck`** — catches explicit `shell: sh` + bash
syntax (SC3040 etc.), expression errors, and general workflow mistakes.
Style-only shellcheck codes are excluded; the SC3xxx "wrong shell" family
is kept.
### Changed (CI)
- **Workflow-level `defaults: run: shell: bash` in `docker-publish.yml`.**
Gitea Actions defaults each `run:` step to `sh` (dash), so every bash-syntax
step had to individually remember `shell: bash` — a discipline requirement
that failed twice (ed49b8d, b7197e8). Setting the default workflow-wide
eliminates the whole class. All pre-existing dash steps use only POSIX
syntax, so bash (a superset) runs them unchanged.
### Fixed (CI)
- **`promote-base-latest` now sets `shell: bash` on the base-latest re-tag
step.** The `b7197e8` fix (v1.2.4) moved the digest-compare into that step
with `set -euo pipefail`, but Gitea Actions' default step shell is `sh`
(dash), which rejects `-o pipefail` (`Illegal option -o pipefail`) and aborts
the step before the `crane copy` runs. On the v1.2.4 release (run 418) this
left `base-latest` un-promoted, still pointing at the v1.2.3 base — the four
consumer tags (`v1.2.4`, `latest`, `v1.2.4-studio`, `latest-studio`) were
unaffected because they `FROM` the exact `base-<hash>`, not `base-latest`.
Same footgun as `ed49b8d` (`resolve-versions needs shell: bash`).
---
## v1.2.4 — 2026-06-29
Patch release. Headline: **pi `0.80.2` → `0.80.3`** (npm `latest`). Also ships a
global gitignore baked into the image, secrets-via-`env_file`-only compose
hardening, and a CI fix so `promote-base-latest` re-points `base-latest`
reliably after a dry-run-first release. The mempalace pin stays `3.5.0`. The
base image rebuilds because `Dockerfile.base` changed (the gitignore seed +
`entrypoint-user.sh` wiring).
### Added
- **Global gitignore baked into the image.** A `~/.gitignore_global`
(`*.bak`, `*.bak.*`, `*~`, `*.orig`, `*.swp`, `*.tmp`) is seeded into the home
dir from `/etc/skel-devbox/` on first boot (seed-if-absent, like
`.bash_aliases`/`.inputrc`, so user edits survive recreate) and wired via
`git config --global core.excludesFile`. Personal/tooling backup artifacts are
now ignored across all repos in the container without per-repo `.gitignore`
entries. The `core.excludesFile` wiring is skipped if the user already set one.
### Changed
- **Secrets are now delivered to the container via `env_file: .env` only; the
`environment:` block no longer re-declares `GITEA_ACCESS_TOKEN`,
`GITEA_HOST`, or `GITHUB_PERSONAL_ACCESS_TOKEN`.** An `environment:` entry
both overrides `env_file:` and is interpolated from the host shell, so a
stale shell export (e.g. one auto-loaded by an opencode/dotenv hook) would
silently shadow the value in your `.env` — an updated token in `.env` never
reached the container. Delivering secrets via `env_file` only decouples the
container from whatever the host shell happens to export. No action needed:
`.env.example` already documents every supported variable. Affects
`docker-compose.yml` and the README “basic shape” snippet.
### Fixed (CI)
- **`promote-base-latest` now re-points `base-latest` reliably after a
dry-run-first release.** The job's gate previously required
`need_build == 'true'`, on the assumption that `need_build == false`
implied `base-latest` was already current. That assumption breaks when a
`workflow_dispatch` dry-run (`promote_latest=false`) pre-builds and pushes
`base-<hash>` first: the subsequent tag run then sees `need_build == false`
(probe hit) and **skipped** promotion, leaving `base-latest` pointing at the
*previous* base. (Observed 2026-06-27 releasing v1.2.3 via dry-run-then-tag
— `base-latest` ended up one base behind, lacking the mempalace self-heal.)
Now the gate runs on every tag release (or `promote_latest=true` dispatch),
and the no-op optimization moved **into** the step as a `crane digest`
compare: it re-tags only when `base-latest` actually differs from the
released `base-<hash>`, so genuine cache-hit releases stay a no-op while
stale aliases get corrected. No image-content change; base hash unaffected.
---
## v1.2.3 — 2026-06-27
Patch release. Headline: **mempalace-mcp now self-heals** instead of latching
`available=false` permanently after a slow cold-open. Also folds in the `yq`
and mempalace-skill changes that were sitting unreleased. **No pi/mempalace
version change** — pi npm `latest` is still `0.80.2` (= v1.2.2) and the
mempalace pin stays `3.5.0`; the base image rebuilds purely because the
`mempalace-toolkit` ref advances to pick up the self-heal extension.
### Fixed
- **mempalace-mcp self-heal — no more permanent `available=false` latch.**
The `mempalace.ts` pi extension (from `mempalace-toolkit`, bumped to
[`e12b624`](https://gitea.jordbo.se/joakimp/mempalace-toolkit/commit/e12b624))
previously tripped its per-request timeout on a slow virtiofs cold-open of
the palace, killed the child, and set `available=false` **forever** (no
respawn) — a pi restart was the only recovery.
- **Bounded respawn with capped exponential backoff** via `ensureAlive()`
(`MEMPALACE_MCP_MAX_RESPAWNS=2`, `MEMPALACE_MCP_RESPAWN_BACKOFF_MS=1000`;
set max to `0` to disable). Both `execute()` and initial startup route
through it. The respawn budget **resets on any successful JSON-RPC
response** (`onStdout`), so a healthy session can't slowly exhaust it.
- **Scoped init timeout** raised `120000 → 300000` ms (`MEMPALACE_MCP_INIT_TIMEOUT_MS`),
affecting **init only** — the per-call timeout stays `60000`
(`MEMPALACE_MCP_TIMEOUT_MS`) — so a genuine cold HNSW deserialize isn't
killed mid-open.
- **Concurrency hardening:** a generation counter prevents a late-exiting
killed process from clobbering a fresh respawn, and an explicit `healthy`
flag replaces the racy `proc != null` check.
- Note: the build-time `smoke-test.sh` verifies the extension is present and
deployed but does **not** exercise respawn behaviour — first live
validation is on a running container.
- **`yq` is now mikefarah's Go yq, not Debian's Python `yq`.** The base image
previously apt-installed `yq`, which on Debian/Ubuntu is the unrelated
kislyuk/`yq` (a jq wrapper, v3.x) — incompatible with the mikefarah v4 syntax
the `cloud-init` repo's `provision.sh`/`deploy.sh` expect. Dropped the apt
package and install the mikefarah binary instead (multi-arch amd64/arm64,
following the repo's `latest` convention like `tealdeer`/`uv`; pin a tag
with `--build-arg YQ_VERSION=vX.Y.Z`). The build-time `smoke-test.sh` gate
asserts `yq --version` reports `mikefarah` **and** major **v4**, so both a
regression to the Python package and a surprise future yq v5 fail CI.
### Changed
- **Baked `mempalace` skill now teaches temporal grounding.** Added a
*Temporal grounding* rule to the image-baked
`skills/mempalace/SKILL.md` (Phase 1 wake-up + a matching anti-pattern):
before using relative time terms ("yesterday", "last week"), establish the
current date/time and compute the delta against the actual diary/drawer
timestamp. Explicitly calls out that a **container recreate or fresh session
is not a day boundary** — pi-devbox restarts several times a day, so two
entries minutes apart can straddle a recreate. Fixes agents mislabelling
same-day sessions as "yesterday".
---
## v1.2.2 — 2026-06-24
Patch release: pick up **pi `0.80.2`** (npm `latest`) and **mempalace `3.5.0`**,
and drop the now-obsolete `diary_write` schema workaround — the upstream fix
shipped.
### Changed
- **mempalace pin `3.4.0` → `3.5.0`.** mempalace 3.5.0 carries the upstream
fix for the top-level-`anyOf` `diary_write` schema
([issue #1728](https://github.com/MemPalace/mempalace/issues/1728) /
[PR #1717](https://github.com/MemPalace/mempalace/pull/1717), merged
2026-06-14). The advertised schema is now `"required": ["agent_name"]` with
`entry`/`content` enforced at dispatch instead of via a root-level `anyOf`,
which Anthropic's tools API accepts. Verified against the published 3.5.0
wheel's `mcp_server.py` before removing the workaround.
- **pi `0.79.10` → `0.80.2`**, auto-resolved from npm `latest` at build time
(no pin in the repo; CI's `resolve-versions` job fetches it).
### Removed
- **The `diary_write` top-level-`anyOf` workaround in `Dockerfile.base`.** The
`perl` patch that rewrote the installed `mcp_server.py` (needed while
mempalace 3.3.x/3.4.0 advertised a top-level `anyOf` that Anthropic rejects,
failing tool registration at session start) is gone, since 3.5.0 fixes it at
the source. Keep `MEMPALACE_VERSION` in lockstep with opencode-devbox.
### Notes
- Unrelated to this release: a *stalled* `mempalace-mcp` (e.g. a slow virtiofs
cold-open of `chroma.sqlite3`) surfaces as `mempalace-mcp not available`
because the `mempalace.ts` extension's per-request timeout kills the child
and flips `available=false` until pi is restarted — this is the 2026-06-13
stall-protection behaving as designed, not the `anyOf` bug.
---
## v1.2.1 — 2026-06-22
Patch release: close the fork/recall + mempalace **under-utilisation gap** in
containers started without the private `skillset` repo — bake the
`pi-extensions` and `mempalace` skills into the image and add the missing
mempalace session-start directive. pi version is re-resolved from npm `latest`
at build.
### Added
- **Vendored fallback skills: `pi-extensions` + `mempalace`.** The pi-toolkit
global `AGENTS.md` directs every pi session to read
`~/.agents/skills/pi-extensions/SKILL.md` at start (the fix for fork/recall
under-utilisation). That pointer dangled in a container started **without**
the private `skillset` repo mounted. The image now bakes fallback copies of
both skills under `/usr/local/share/pi-devbox/skills/`, symlinked in by
`entrypoint-user.sh` (only when absent, so a mounted skillset still wins).
- **Proactive-load directive for `mempalace`.** Baking the skill only fixes
*availability*; nothing in pi-toolkit's global `AGENTS.md` told sessions to
load it, so it would still surface only via description-matching. The
pi-devbox managed block (`pi-global-AGENTS.append.md`) now adds a
session-start pointer (gated to pi-devbox containers, conditional on the
MemPalace MCP tools being present) so a new container actually picks the
skill up — memory continuity matters most in a frequently-recreated
container. (`pi-extensions`'s directive already ships in pi-toolkit, so only
its skill file needed baking.)
- **Layered freshness for the `pi-extensions` skill (Option 1 + Option 2).**
The canonical skill was promoted into the **public `pi-extensions` package
repo** under `skill/` (co-located with the extensions it documents). A
committed snapshot in `rootfs/` is the *floor*; `Dockerfile.variant` copies
`/opt/pi-extensions/skill/` (the pinned, manifest-recorded clone) over it at
build, so a normal build ships the fresh package copy and an old-ref/mirror
build still ships the snapshot. `mempalace` is snapshot-only (its consumer
skill has no public package home — the `mempalace-toolkit` repo ships a
*different* skill, `opencode-mempalace-bridge`). Provenance + refresh steps:
`rootfs/usr/local/share/pi-devbox/skills/VENDORED.md`.
- **Smoke-test coverage** for the fallback skills: build-time presence of both
`SKILL.md`s and the `pi-extensions` helper, a check that the baked
`pi-extensions` skill matches the package copy when the clone carries it, and
runtime assertions that both are symlinked into `~/.agents/skills/`.
---
## v1.2.0 — 2026-06-22
Minor release: **image-baked agent skills** — a new base mechanism that ships
skills inside the image (independent of any mounted skillset repo) — plus the
first such skill, `pi-devbox-environment`, and pi `0.79.9` → `0.79.10`
(auto-resolved from npm `latest` at build).
### Added
- **Image-baked agent skills.** Skills under
`/usr/local/share/pi-devbox/skills/<name>/` are now symlinked into
`~/.agents/skills/` by `entrypoint-user.sh` on every start, making them
available **with or without** a mounted `skillset` repo. The symlink points
at the image path (so it survives volume recreate, unlike anything baked
under a home dir a named volume would shadow) and is created only when
absent, so a same-named skillset skill or user override is never clobbered.
The skillset deploy classifies these as foreign-links and its `--prune-stale`
pass leaves them untouched.
- **`pi-devbox-environment` skill** (the first image-baked skill). Teaches
agents the container-shaped facts that are easy to get wrong: the
persistence/ephemerality tier model (what survives `down -v` / image
update), host + LAN SSH reachability and ControlMaster, split-horizon DNS
*mechanisms*, the interactive-vs-tool-shell alias gotcha (`dssh`/`dscp`/
`cat`→`bat` don't exist in the non-interactive bash tool), the tmux 0-index
constraint, uv-first Python, and pi-studio reachability. Deliberately
environment-agnostic — host OS, hostnames, internal domains, and nameservers
are discovered at runtime, never hardcoded.
- **Proactive skill awareness via the global `AGENTS.md`.** `Dockerfile.variant`
appends a short, gated pointer (`pi-global-AGENTS.append.md`) onto
pi-toolkit's `pi-global-AGENTS.md` — the single global instruction slot pi
loads at startup — so containers load the `pi-devbox-environment` skill
proactively rather than only on description match. The pointer fires only
inside a pi-devbox container (checks for `/usr/local/lib/pi-devbox/`).
Build-time append is idempotent via a marker grep; runtime is unaffected
(the file is root-owned and re-symlinked by pi-toolkit each boot).
- **Smoke-test coverage** for the new mechanism: build-time presence of the
baked skill + append snippet + the merged marker in `pi-global-AGENTS.md`,
and a runtime assertion that `~/.agents/skills/pi-devbox-environment` is
linked after the entrypoint runs.
### Bumped: pi 0.79.9 → 0.79.10
Resolved from npm `latest` at build (v1.1.7 shipped `0.79.9`). See the
[pi changelog](https://github.com/earendil-works/pi/blob/main/CHANGELOG.md)
for the upstream `0.79.10` notes.
## v1.1.7 — 2026-06-21
Patch release: pi `0.79.8` → `0.79.9` (auto-resolved at build), plus the
`ssh-lan.conf` LAN-peer documentation that landed on `main` after v1.1.6.
Companion refs are auto-resolved to SHAs at build as before.
### Bumped: pi 0.79.8 → 0.79.9
Notable upstream changes (from [pi releases](https://github.com/earendil-works/pi/releases/tag/v0.79.9)):
- **Chat-template thinking compatibility** — OpenAI-compatible custom
providers can map pi thinking levels into `chat_template_kwargs`, enabling
vLLM/Hugging Face chat-template models (e.g. DeepSeek) to use
provider-native thinking controls.
- **GLM-5.2 provider improvements** — corrected Fireworks OpenAI-compatible
routing and OpenRouter `xhigh` thinking support, improving `/model`
behaviour and high-effort reasoning for GLM-5.2.
- **Fixes** — same-directory session switches now reuse imported extension
modules (fresh instances + lifecycle events preserved); deep session
branches no longer take quadratic time to build context; Markdown
streaming code-fence rendering no longer flickers on partial closing
fences; fuzzy `edit` matches preserve untouched line blocks instead of
rewriting the whole file; `/model` hides Copilot models unavailable to the
account and ranks exact provider-prefixed matches first.
### Docs: document `~/.config/devbox-shell/ssh-lan.conf` for naming LAN peers
The host-owned, bind-mounted `~/.config/devbox-shell/ssh-lan.conf` is the
intended place to add `ProxyJump host` overrides for **named** LAN peers (so
`pi --ssh <peer>` / `dssh <peer>` route through the host), but it was only
mentioned in `.env.example` and the `setup-lan-access.sh` header — never in the
README. Added a "Naming LAN peers" subsection to the README troubleshooting
block (plus a pointer from the SSH/ControlMaster section), and corrected the
stale `setup-lan-access.sh` comment that suggested editing the read-only
`~/.ssh/config` instead of `ssh-lan.conf`.
## v1.1.6 — 2026-06-19
Build provenance + reproducibility hardening, plus pi `0.79.7` → `0.79.8`
(auto-resolved at build). Companion refs are auto-resolved to SHAs at build
as before.
### Bumped: pi 0.79.7 → 0.79.8
Notable upstream changes (from [pi releases](https://github.com/earendil-works/pi/releases/tag/v0.79.8)):
- **Selective provider base entry points** — SDK users can pair
`@earendil-works/pi-ai/base` and `@earendil-works/pi-agent-core/base` with
explicit provider registration to keep bundled apps from including unused
provider transports.
- **Mistral prompt caching** — Mistral sessions use provider-side prompt
caching keyed on the pi session ID, with cached-token usage/cost
accounting.
- **Post-compaction token estimates** — compact results and compaction
events now include estimated post-compaction token counts.
- **OpenRouter Fusion alias** — `openrouter/fusion` available as a built-in
OpenRouter model alias.
### Added
- **Self-describing images: OCI labels + on-disk build manifest.** The
variant build now records exactly which pi version and companion-repo
commits were baked into each image. Previously the SHAs resolved by CI
only ever reached the build log (which rotates), so a published tag was
not reconstructable after the fact — confirming what shipped meant
triangulating from `git`, `pi --version`, and extension source.
- OCI labels: `org.opencontainers.image.{version,revision,created}` plus
`se.jordbo.pi-devbox.{pi,pi-toolkit,pi-extensions,pi-fork,pi-obsmem,mempalace-toolkit,pi-studio}-*ref` —
inspect with `docker inspect`.
- `/etc/pi-devbox/build-manifest.json` written from **ground truth** (the
actual checked-out `HEAD` of each `/opt` clone + live `pi --version`),
not just the intended build-args, so it also exposes a clone that
silently resolved to the wrong ref. The provenance ARGs are declared
last so a changing `BUILD_DATE` never invalidates the expensive
install/clone layers.
- **`scripts/check-base-hash.sh` — base-rebuild invariant guard.** Every
floating `ARG *_REF` consumed by `Dockerfile.base` must be folded into the
`base_tag` hash, or a ref-only change won't trigger a base rebuild (the
v1.1.2 mempalace-toolkit staleness footgun). The guard fails CI the moment
someone adds an `ARG *_REF` to `Dockerfile.base` without folding it in; it
runs in the `base-decide` job and locally. Smoke-test gained assertions for
the manifest (present, no `"unknown"` components) and the OCI labels.
- **Overridable companion repo URLs.** The three gitea-hosted companions
(`pi-toolkit`, `pi-extensions`, `mempalace-toolkit`) gained `*_REPO`
build-args defaulting to their canonical `gitea.jordbo.se` origin —
matching the existing `PI_FORK_REPO` / `PI_OBSMEM_REPO` / `PI_STUDIO_REPO`
pattern. A relocated or forked build can now repoint a companion at a
mirror, another host, or a local path (`--build-arg PI_EXTENSIONS_REPO=...`)
without editing the Dockerfiles. Defaults are unchanged, so the canonical
CI build is byte-identical.
### Changed
- **`resolve-versions` now fails loud instead of falling back to a floating
branch.** Each pi-version / companion-ref lookup previously degraded to
`main`/`master` on a transient API/network failure (`|| echo "main"`),
silently shipping an unpinned ref that defeats both cache-busting and
reproducibility. Resolution now validates each result is a 40-hex commit
SHA (and pi a real semver) and aborts the release otherwise.
## v1.1.5 — 2026-06-18
Patch release: SSH ControlMaster read-only-socket fix + pi `0.79.6` → `0.79.7`
(auto-resolved at build). The `pi-extensions` ref is auto-resolved to `main`
HEAD at build, so the `ssh-controlmaster` fix below lands automatically.
### Fixed
- **`pi --ssh <host>` no longer fails with "Read-only file system" when the
user's `~/.ssh/config` sets a per-host `ControlPath` under the read-only
`~/.ssh` mount** (e.g. the common CGNAT idiom `ControlPath ~/.ssh/cm/%r@%h:%p`).
Root cause: SSH precedence means a user's per-host `ControlPath` always wins
over the baked `/etc/ssh/ssh_config.d` default, so the master socket tried to
bind under the RO `~/.ssh` and `ssh … pwd` exited 255 ("Could not resolve
remote pwd"). The `ssh-controlmaster` extension (pulled from `pi-extensions`
`main` via `PI_EXTENSIONS_REF`) now (a) resolves the remote pwd with a direct
connection (`-o ControlPath=none -o ControlMaster=no`), and (b) tests whether
the system `ControlPath` dir is actually writable — falling back to its own
`/tmp` master (whose command-line `-o ControlPath` overrides the user's path)
when it is not. OS-agnostic and independent of whether the user uses
ControlMaster, so the majority of configs (no ControlMaster at all) are
unaffected.
### Changed
- **`setup-lan-access.sh` now renders the writable SSH sidecar
(`~/.ssh-local/config`) on every host OS, not just VM-backed ones.**
Previously the whole script no-oped on native Linux, so a Linux host that
also bind-mounts `~/.ssh` read-only got no `ControlPath` redirect. The
`ControlPath` redirect + `Include ~/.ssh/config` (and `dssh`/`dscp` usability)
now work on Linux too; only the host-jump block (`Host host mac`), its key
generation, and the authorize hints remain gated on VM-backed detection
(`DEVBOX_LAN_ACCESS=auto`) or `=jump`.
### Bumped: pi 0.79.6 → 0.79.7
Notable upstream changes (from [pi releases](https://github.com/earendil-works/pi/releases/tag/v0.79.7)):
- **Automatic theme mode** — `/settings` can choose separate light and dark
themes and follow terminal color-scheme changes (`/` is now reserved in
theme names for this).
- **Self-only `pi update` by default** — bare `pi update` updates pi only;
`pi update --all` updates pi and packages together.
- **Extension API helpers** — `CONFIG_DIR_NAME` exported so extensions resolve
project config paths without hardcoding `.pi`; edit-diff helpers
(`generateDiffString`, `generateUnifiedPatch`, `EditDiffResult`) exported.
- **Warp inline images** via Kitty graphics capability detection.
- Fixes: RPC unknown-command errors now include the request id (clients no
longer hang); `/model` autocomplete matches provider/model regardless of
token order; tree navigator horizontally pans deep entries.
## v1.1.4 — 2026-06-17
Patch release: config and shell-quality fixes on a preserved volume. No pi
version bump (still `0.79.6`, latest). The `pi-toolkit` ref is auto-resolved
to `main` HEAD at build, so the AGENTS.md change below lands automatically.
### Added
- **Global `AGENTS.md` auto-loads the pi-extensions skill.** `pi-toolkit` now
ships `pi-global-AGENTS.md` and symlinks it to `~/.pi/agent/AGENTS.md` (pi's
global-instructions file, loaded at every start). It directs the agent to
read the `pi-extensions` skill at session start and carries a core
fork/recall cheat-sheet, since on-demand skill description-matching was
leaving `pi-fork` / `pi-observational-memory` under-utilised. **Heads-up:**
on a preserved volume any pre-existing real `~/.pi/agent/AGENTS.md` is backed
up to `*.bak.<timestamp>` and replaced by the symlink (same behavior as
`keybindings.json`).
- **`settings.json` merge-on-recreate.** The bootstrap only ever copied the
template when `settings.json` was *absent*, so a file on a preserved volume
never picked up config added in a later image (e.g. the
`observational-memory` / `pi-fork` blocks, a newly-enabled model). The
entrypoint now deep-merges the template into an existing `settings.json` on
start with `jq -s '.[0] * .[1]'` (template first, live second): the user's
values always win and only *missing* keys are filled in. Arrays are treated
as leaves (a model the user removed is not re-added); the file is only
rewritten when the merge changes something, the original is backed up first,
and invalid JSON on either side is skipped rather than clobbered. Opt out
with `PI_SETTINGS_MERGE=0`.
### Fixed
- **bash history loss in nested / tmux shells.** The `DEVBOX_HIST_SET` guard
that installs the per-prompt `history -a` flush was `export`ed, so it leaked
into child processes. Any nested shell — crucially each tmux pane, which
inherits the tmux server's env — saw the guard already set and skipped
installing `history -a`, persisting history only on a clean exit. Abrupt
termination (`docker stop`, `tmux kill-server`, SIGKILL) then silently lost
that shell's in-memory history. The guard is now shell-local (no `export`),
so every new interactive shell re-installs its own flush. `zoxide` was less
affected (its hook is unguarded and writes immediately). History and zoxide
storage were never the issue — `~/.cache/bash` (`devbox-shell-history`) and
`~/.local/share/zoxide` (`devbox-zoxide`) are persistent named volumes.
**Note:** existing shells/panes keep the old behavior until restarted
(`tmux kill-server` or open fresh shells).
### Maintainer
- `scripts/recreate-sanity-check.sh` gained assertions for the new wiring: the
`~/.pi/agent/AGENTS.md` symlink, a nested login shell installing
`history -a`, and `settings.json` carrying the `observational-memory` +
`pi-fork` blocks after recreate.
---
## v1.1.3 — 2026-06-16
Patch release: pi `0.79.4` → `0.79.5` (auto-resolved at build).
### Bumped: pi 0.79.4 → 0.79.5
Notable upstream changes (from [pi releases](https://github.com/earendil-works/pi/releases/tag/v0.79.5)):
- **Provider-scoped API key environments** — `auth.json` API key entries can
now include `env` overrides for provider-specific Cloudflare, Azure OpenAI,
Google Vertex, Amazon Bedrock, cache retention, and proxy settings without
changing the project shell.
- **Global HTTP proxy setting** — configure `httpProxy` once in global settings
to apply `HTTP_PROXY` / `HTTPS_PROXY` to Pi-managed HTTP clients.
- **Vercel AI Gateway attribution** — requests now include Pi attribution
headers by default.
- **Fixes:** inherited OpenAI Responses streaming tolerates null message content
before tool calls; DeepSeek V4 thinking no longer sends both `thinking` and
`reasoning_effort`; device-code login no longer auto-opens the browser;
various Google/Vertex Gemini model metadata corrections; session selector
empty-state fix; Cursor Up history navigation fix.
---
## v1.1.2 — 2026-06-15
Patch release: pi `0.79.3` → `0.79.4` (auto-resolved at build), plus the
build-plumbing fix, maintainer tooling, and docs accumulated since v1.1.1.
### Changed
- **`mempalace-toolkit` is now CI-resolved to a commit SHA**, closing a
silent-staleness footgun. It is the only companion cloned in
`Dockerfile.base` (all others are cloned in `Dockerfile.variant`), so it
was never run through the `resolve-versions` → build-arg plumbing. Its
ref stayed a literal `main`, and because the base only rebuilds when the
hash of `Dockerfile.base + rootfs/* + entrypoints` changes, a
toolkit-only fix would *not* land in the image unless `Dockerfile.base`
itself happened to change (as it did, incidentally, in v1.1.1).
Now `resolve-versions` resolves `mempalace-toolkit` `main` HEAD to a SHA
(new `mempalace_toolkit_ref` output), `base-decide` folds that SHA into
the base-tag hash (so a moved toolkit forces a base rebuild), and
`build-base` passes it as `--build-arg MEMPALACE_TOOLKIT_REF`. The base
clone switched from `git clone --branch` to a SHA-capable
`git fetch <ref> + checkout FETCH_HEAD` (the `--branch <40-char-SHA>`
footgun previously fixed in `Dockerfile.variant`, run 374).
Note: `base-decide` now depends on `resolve-versions`, so the base tag
reflects a live gitea API lookup. On an API blip it falls back to `main`
— which hashes differently than a SHA and triggers one *extra* rebuild,
never a *missed* one (fail-toward-rebuild).
### Added (maintainer tooling, no image change)
- **`scripts/recreate-sanity-check.sh`** — runtime post-recreate sanity
check; the runtime peer of `smoke-test.sh`. Where `smoke-test.sh` runs at
build time with `--entrypoint=""` (and so can never see persisted volumes
or the entrypoint's runtime deploy), this verifies what is actually live
in the container *after* `docker compose up -d --force-recreate`:
persisted named volumes survived, the pi runtime wiring is intact
(keybindings symlink, ≥4 extensions, `mempalace.ts` bridge, `settings.json`,
and pi-fork / pi-observational-memory / pi-studio registrations),
`/tmp/sshcm` is mode 700, shell defaults re-seeded, and `/opt` toolkits
intact. Variant (studio/plain) auto-detected via `/opt/pi-studio`. Since
pi is built from `latest` (no concrete Dockerfile pin), the version check
asserts only when `--expected-version` is passed, else WARNs. Not baked
into the image — repo/maintainer tooling, same category as
`smoke-test.sh`. A short-name wrapper (`pi-devbox-sanity`) lives in
`cli_utils/bin`, kept separate from opencode-devbox's `devbox-sanity` so
hosts with only one devbox checked out stay self-contained.
### Docs (no image change)
- Correct the MemPalace `diary_write` anyOf workaround watch-target in
`Dockerfile.base`: upstream PR #1735 was **closed unmerged** (2026-06-11),
so the old “remove once #1735 ships” TODO pointed at a dead PR. Issue #1728
is still open; PR #1717 is the current live candidate; mempalace PyPI latest
is still 3.4.0 (== our pin), so the workaround stays. Removal trigger is now
a PyPI release > 3.4.0 that actually strips the root-level anyOf.
- Document the post-recreate sanity check: AGENTS.md release-day checklist
(step 3) now runs `scripts/recreate-sanity-check.sh` inside the recreated
container, and README gains a "Post-recreate sanity check" subsection
alongside the build-time smoke-test note.
---
## v1.1.1 — 2026-06-13
Patch release: pi `0.79.1` → `0.79.3` (auto-resolved at build) plus the
mempalace-mcp hang fix below.
### Fixed
- **`mempalace-mcp` no longer hangs the pi TUI uninterruptibly.** When
the palace is bind-mounted from the macOS host (OrbStack virtiofs) and
the container opened a large `chroma.sqlite3` for the first time, a
cold storage open / HNSW load could stall the server before it emitted
its JSON-RPC response. The awaiting promise then hung forever and the
TUI froze — ESC cancels the LLM stream, not a pending MCP tool call, so
there was no way out short of `docker exec <container> pkill -9 -f
mempalace-mcp` and restarting pi.
The fix lives in the `mempalace.ts` pi extension shipped by
**mempalace-toolkit** (cloned into the base at build time via
`MEMPALACE_TOOLKIT_REF`, default `main`): the JSON-RPC client now arms
a **per-request** timeout. On expiry it rejects the request *and* kills
the stalled child (SIGTERM→SIGKILL), so pi surfaces an error instead of
hanging; the bridge then marks itself unavailable so subsequent calls
fail fast (restart pi to retry). This is deliberately per-REQUEST, not
a process-lifetime `timeout 60 mempalace-mcp` wrapper — the long-lived
server is only killed when a request genuinely stalls.
Tunables (env): `MEMPALACE_MCP_TIMEOUT_MS` (tool-call timeout, default
`60000`), `MEMPALACE_MCP_INIT_TIMEOUT_MS` (initialize/tools-list
handshake, default `120000`); set either to `0` to disable. Requires a
base rebuild to pull the updated extension. The earlier plan of a
standalone Python stdio-watchdog shim was dropped: the extension
already owns request/response correlation, so a separate
framing-reparsing shim is unnecessary.
Still open (out of scope here): sharing one palace across harnesses
ideally wants a single host-side `mempalace-mcp` daemon multiplexing
stdio over a UNIX socket, so all clients share one writer on native
APFS rather than each cold-opening over virtiofs.
`mempalace-mcp` that applies a per-request timeout and kills the child
on stall, **without** killing the long-lived server itself (a naive
`timeout 60 mempalace-mcp` wrapper is wrong — it kills the server
mid-session). Sharing the palace across harnesses (native pi, container
pi, opencode) remains the goal — isolated palaces defeat the point.
Longer term: run a single mempalace-mcp daemon on the host and
multiplex stdio over a UNIX socket so all clients share one writer on
native APFS.
### Added
- **`dot-watch` helper** (`/usr/local/bin/dot-watch`) — auto-rerenders a
Graphviz `.dot` file to PNG on every save via mtime polling (no
`inotify` dependency). pi-studio renders Mermaid natively but has no
DOT renderer; since its markdown preview displays local PNG/JPG/GIF/WEBP
images, this closes the loop for Graphviz: edit `.dot` → `dot-watch`
regenerates `<name>.png` → Studio *refresh-from-disk* shows the update.
`graphviz` was already in the base image, so no new package. Baked into
`Dockerfile.base` following the `studio-expose` pattern; documented in
the README Studio section.
## v1.1.0 — 2026-06-10
### Added — `:latest-studio` variant
- **New `-studio` image variant** bundling
[pi-studio](https://github.com/omaclaren/pi-studio) — a two-pane
browser workspace (prompt/response editor, live KaTeX/Mermaid preview,
tmux-backed literate REPLs for Shell/Python/IPython/Julia/R/GHCi/Clojure)
plus the `/studio` slash command and `studio_repl_send` /
`studio_export_*` agent tools. Published as `:latest-studio` and
`:vX.Y.Z-studio` (multi-arch).
- pi-studio is **vendored to `/opt/pi-studio`** at build time (gated by
`INSTALL_STUDIO=true`, ref pinned via CI-resolved `PI_STUDIO_REF`) and
registered on container start by `entrypoint-user.sh` via
`pi install /opt/pi-studio` — the same pattern as pi-fork /
pi-observational-memory. No build step: pi-studio ships its browser
bundle prebuilt in git. The non-studio `:latest` image is unchanged.
- CI gains independent `smoke-studio` + `build-variant-studio` jobs that
gate **only** the studio tags, so a studio build/smoke failure can
never block the core `:latest` / `:vX.Y.Z` release.
- `STUDIO_PORT=8765` baked as an advisory default.
- **`studio-expose` helper + `socat` (base).** Because pi-studio binds the
container's loopback, a published Docker port can't reach it. The new
`studio-expose` helper (socat, added to the base) bridges the container's
loopback to its egress interface on the same port; set `STUDIO_EXPOSE=1`
in compose to auto-start it on boot (default off — Studio stays
loopback-only otherwise). `socat` is in the base for all variants.
- **README "Using pi-studio" section.** Documents the container access
reality: pi-studio hard-binds `127.0.0.1` inside the container
(`.listen(port,"127.0.0.1")`, no `--host` flag), so a plain `-p`
publish does not reach it. Documents the two working paths — host
networking (recommended on OrbStack) and a loopback bridge for bridge
networking — plus the remote `ssh -L` forward and the **mosh caveat**
(mosh cannot forward ports; run a parallel `ssh -L` alongside it).
## v1.0.1 — 2026-06-10
Patch release. Works around an upstream MemPalace bug that broke pi at
first prompt against the Anthropic Claude API.
### Fixed
- **`mempalace_diary_write` schema rejected by Anthropic API.** Mempalace
3.3.x and 3.4.0 advertise `diary_write`'s `input_schema` with a
top-level `anyOf: [{required:[entry]}, {required:[content]}]` to
express "either `entry` or `content` must be supplied". Anthropic's
tools API rejects top-level `anyOf` / `oneOf` / `allOf` outright, so
pi failed to register tools at session start with
`tools.<n>.custom.input_schema: input_schema does not support oneOf,
allOf, or anyOf at the top level`. `Dockerfile.base` now patches the
installed `mcp_server.py` after `uv tool install` to drop the `anyOf`
block and require `["agent_name", "entry"]` instead. The mempalace
handler still accepts `content` server-side as a kwarg alias, so
callers using either name keep working. Tracked upstream:
[issue #1728](https://github.com/MemPalace/mempalace/issues/1728),
[PR #1735](https://github.com/MemPalace/mempalace/pull/1735).
The workaround is idempotent + self-deactivating and will be removed
once a fixed mempalace release lands on PyPI.
### Changed
- **Mempalace pinned to 3.4.0** via `MEMPALACE_VERSION` build arg.
Future bumps must be a reviewable diff rather than an implicit pull
of `latest` (the broken 3.3.x/3.4.0 schema slipping in unannounced
is what caused this release).
## v1.0.0 — 2026-06-09
**Decoupled from opencode-devbox.** pi-devbox is now self-contained:
own `Dockerfile.base` + `Dockerfile.variant`, own CI pipeline, own
release cadence. Previously v0.79.0 and earlier were thin re-brands of
the `pi-only` variant built by opencode-devbox CI.
### Architectural
- **Self-contained build chain.** `Dockerfile.base` produces
`joakimp/pi-devbox:base-<hash>` (content-addressed); `Dockerfile.variant`
FROMs the base and adds the pi install. Replaces the prior 5-line
`Dockerfile` shim that FROMed `joakimp/pi-devbox:base-pi-only` (an
opencode-devbox CI artifact).
- **No more publish-ordering coupling.** pi-devbox releases no longer
require rebuilding opencode-devbox first.
- **Adapted from opencode-devbox** at the time of decoupling — the
apt set, ssh ControlMaster setup, MemPalace integration, entrypoint
UID/GID dance, and CI pipeline shape are all derived from there. See
Acknowledgements in README.md.
- **CI workflow** rewritten as two-phase split-base build pipeline
(mirrors opencode-devbox's `docker-publish-split.yml` shape, simplified
to a single variant). Includes `crane`-based `base-latest` promotion,
registry-buildcache footgun guard via concrete `PI_VERSION` resolution,
and the c6f9d11 smoke-test gate (waits for keybindings + mempalace.ts
+ ≥4 *.ts before sampling).
### Added (base image)
- **pandoc** — universal Markdown↔HTML/Org/RST/etc. conversion. ~200 MB.
- **graphviz** — `dot` rendering for diagram pipelines. ~10 MB.
- **imagemagick** — image conversion (invoked as `magick`, not `convert`,
in v7+). ~50 MB.
- **yq** — YAML-aware companion to jq.
- **tldr (tealdeer)** — Rust port of tldr-pages, ~5 MB static binary.
Replaced the Node `tldr` global (which was ~140 MB).
- **`/etc/tmux.conf`** with `set -g base-index 0` + `set -g
pane-base-index 0`. Required for the planned `:latest-studio`
variant; pi-studio hard-codes its tmux send target to `:0.0`. User-
level `~/.tmux.conf` overrides still win.
### Added (smoke test)
- Asserts pandoc, graphviz, imagemagick, yq, and tldr are present.
- Asserts `/etc/tmux.conf` has the 0-indexed config baked.
- Asserts `/tmp/sshcm/` directory created mode 700 by entrypoint.
- Image-size measurement now sums `docker history` layer sizes (the
prior `image inspect --format='{{.Size}}'` approach returned only
the variant-unique layer when the base was content-addressed and
shared, understating the user-facing image size by 2+ GB).
- Size threshold raised to 3500 MB (was 2850) to cover the new base
additions plus +200 MB safety margin. Tighten in a follow-up release
once amd64 actuals settle.
### Image size
Local arm64 build of `pi-devbox-test:latest` (this branch's content):
3.20 GB. Up ~390 MB from the prior pi-only-equivalent (~2.81 GB) due
to pandoc, graphviz, imagemagick, yq, and minor expansion in pi npm
dependencies.
### Migration notes
- Existing volumes (`devbox-pi-config`, `devbox-bash-history`,
`devbox-nvim-data`, `devbox-uv-tools`, `devbox-chroma-cache`) are
unchanged in name and structure. `docker compose pull && docker
compose up -d --force-recreate` is a clean upgrade path.
- The `:latest` and `vX.Y.Z` Hub tags continue to point at a "base +
pi" image. Same shape, just built differently.
- `:base-pi-only` and `:base-pi-only-vX.Y.Z` tags from prior releases
remain on Hub for now; will be deprecated when opencode-devbox
retires the pi paths in its next major release.
### Future work
- v1.1.0: `:latest-studio` variant (adds [pi-studio](https://github.com/omaclaren/pi-studio)).
- v1.3.0: `:latest-studio-tex` variant (adds texlive-xetex for PDF export).
## v0.79.0 — 2026-06-08
First build on pi **`0.79.0`** (upstream `@earendil-works/pi-coding-agent` bump
from `0.78.1`). Built `FROM` the freshly republished
`joakimp/pi-devbox:base-pi-only` from opencode-devbox `v1.16.2`, which carries
pi `0.79.0` (and picks up opencode `1.16.2` in the sibling opencode-bearing
variants, though this pi-only image has no opencode).
### Bumped: pi 0.78.1 → 0.79.0
Resolved from the tag and asserted by the smoke base-freshness guard
(`EXPECTED_PI_VERSION`). Highlights from the upstream `CHANGELOG.md`:
- **Project trust for local inputs** — pi now asks before loading project-local
settings, resources, instructions, and packages, with saved decisions and
`--approve` / `--no-approve` controls for non-interactive modes, plus a
`project_trust` extension event so global/CLI extensions can decide or defer.
- **Cache-hit visibility in the footer** — the interactive footer shows the
latest prompt cache hit rate (`CH`).
- **Richer SDK/RPC extension surfaces** — public exports now include RPC
extension UI request/response types and package asset path helpers.
- Plus a large batch of TUI and provider fixes (Kitty keyboard fallback,
prompt-history cursor placement, large-JSONL session reads, custom-provider
routing).
### Smoke size threshold 2750 → 2850 MB
Tracks opencode-devbox's `pi-only` variant, which was raised to 2850 MB in
`v1.16.2` for headroom against the pi `0.79.0` bump (and routine apt drift).
Kept in lockstep so this image's guard matches its source-of-truth variant.
## v0.78.1 — 2026-06-04
First build on pi **`0.78.1`** (upstream `@earendil-works/pi-coding-agent` bump
from `0.78.0`). Built `FROM` the freshly republished
`joakimp/pi-devbox:base-pi-only` from opencode-devbox `v1.15.13e`, which carries
pi `0.78.1` plus the LAN-jump key-persistence work and the `devbox-ssh-local`
volume ownership fix. Adds compose/env documentation in this repo.
### Added: persist the LAN-jump key + one-line authorize hint
- **compose:** persist `~/.ssh-local` via a new `devbox-ssh-local` named volume
so the generated LAN-jump key survives `docker compose up --force-recreate`.
You authorize the key on the host **once per machine** instead of after every
container update.
- **Inherited from base:** `setup-lan-access.sh` now prints a copy-paste
`echo '…' >> ~/.ssh/authorized_keys` line when it generates a new key
(published via opencode-devbox's `base-pi-only`). No helper file to locate.
### Docs: document optional host-owned config in the compose + env templates
- **compose:** added a commented-out `~/.config/devbox-shell` bind mount with a
note — the image's `~/.bash_aliases` sources
`~/.config/devbox-shell/bash_aliases` if present, and `setup-lan-access.sh`
reads `~/.config/devbox-shell/ssh-lan.conf` for named-peer `ProxyJump host`
overrides (reach LAN peers by name via `dssh <peer>`).
- **.env.example:** documented `DEVBOX_HOST_ALIAS` (host hostname to reach,
default `host.docker.internal`) so getting-started is self-contained.
Template/example comments only; no behavior change.
## v0.78.0c — 2026-06-04
### Fixed / Added (inherited from the base via `FROM`)
LAN-access improvements made in opencode-devbox's `setup-lan-access.sh` (baked
into the `base-pi-only` image, published by opencode-devbox v1.15.13d) flow
through to pi-devbox automatically — no pi-devbox source change. Built `FROM`
the rebuilt `joakimp/pi-devbox:base-pi-only` (digest `83b45335…`):
- **Fixed:** the generated `~/.ssh-local/config` had `Include ~/.ssh/config`
scoped to the `host`/`mac` block, so `dssh <peer>` by name was ignored.
- **Fixed:** read-only `~/.ssh/cm` ControlPath broke multiplexed hosts
(`pmx-jh`, `proxmox*`, …); master sockets now use the writable sidecar.
- **Added:** host-owned `~/.config/devbox-shell/ssh-lan.conf` for named-peer
`ProxyJump host` overrides (Included before `~/.ssh/config`).
- **Added:** `DEVBOX_LAN_AUTOJUMP_PRIVATE=1` — ProxyJump any RFC1918 IP through
the host for roaming laptops.
## v0.78.0b — 2026-06-03
Container-level rebuild on pi `0.78.0` (unchanged): re-brands the pi-only build
as a thin `FROM joakimp/pi-devbox:base-pi-only`, inheriting fork/recall and
host-OS-agnostic LAN access. Letter-suffix release (pi version unchanged).
### Changed: refactored to re-brand the opencode-devbox `pi-only` variant
pi-devbox no longer installs pi itself. The `Dockerfile` is now a thin
`FROM joakimp/pi-devbox:base-pi-only` (overridable via the `BASE_IMAGE`
arg), inheriting pi + pi-toolkit + pi-extensions and all base tooling from the
single source of truth. This eliminates the install-logic duplication that
used to drift against `opencode-devbox/Dockerfile.variant`.
The pi-only artifact is **built** by opencode-devbox's CI (from
`opencode-devbox/Dockerfile.variant` with `INSTALL_OPENCODE=false`) but is
**published into this repo** as the internal building-block tag
`joakimp/pi-devbox:base-pi-only` (+ `base-pi-only-vX.Y.Z`, where `vX.Y.Z` is
the opencode-devbox release version). This supersedes the brief approach of
publishing it as `opencode-devbox:latest-pi-only` — an "opencode-devbox" tag
with no opencode in it confused users. `base-pi-only` is internal; end users
pull `joakimp/pi-devbox:latest` or a `vX.Y.Z` tag.
The pi-only build uses `INSTALL_OPENCODE=false`, so this image
stays lean and pi-focused — it does **not** carry opencode, and remains
distinct from `opencode-devbox:latest-with-pi` (which has both).
### Added (inherited from the pi-only variant)
- **`fork` tool** (pi-fork) and **`recall` tool** (pi-observational-memory),
baked into `/opt` with `node_modules` and registered at runtime.
- **Host-OS-agnostic LAN access**: on VM-backed hosts (macOS OrbStack /
Docker Desktop) the entrypoint sets up the host as an SSH jump to reach LAN
peers (`dssh` alias; `DEVBOX_LAN_ACCESS` / `HOST_SSH_USER` env). No-op on
native Linux. See the opencode-devbox README for details.
### Consequences / notes
- **Publish ordering**: release opencode-devbox first so `base-pi-only`
carries the target pi version, *then* tag this repo. The smoke test asserts
`pi --version` matches the tag and fails loudly if the base is stale.
- CI no longer passes `PI_VERSION` as a build-arg (the Dockerfile installs
nothing); it still resolves the tag version to feed the smoke base-freshness
guard. Smoke size threshold 2200 → 2750 MB (now tracks the pi-only variant).
_pi version unchanged at `0.78.0` (still latest)._
## v0.78.0 — 2026-05-29
pi `0.77.0` → `0.78.0` bump (first container build on the pi 0.78 line, published upstream 2026-05-29). Built against `joakimp/opencode-devbox:base-latest` (unchanged from the v0.77.0 build).
### Bumped: pi 0.77.0 → 0.78.0
**New Features**
- **Named startup sessions** — `--name` / `-n` sets the session display name before startup across interactive, print, JSON, and RPC modes.
- **Clickable file tool paths** — built-in file tool titles render OSC 8 `file://` hyperlinks when the terminal supports them, including supported tmux clients.
**Added**
- Exported `convertToPng` for extension authors.
- Exported `parseArgs` and type `Args` for extension authors.
- Added a resume command hint when exiting interactive sessions.
- Added custom Amazon Bedrock request header support.
**Fixed**
- Fixed early interactive input typed before the prompt loop starts so it is buffered instead of dropped.
- Fixed OpenRouter Moonshot Kimi K2.6 requests to use `system` instead of unsupported `developer` messages.
- Fixed OSC 8 hyperlinks to pass through tmux when the client supports them.
- Fixed ANSI text wrapping to avoid stack overflows on very long wrapped lines.
- Fixed OpenAI Codex Responses SSE streams to abort response body reads after terminal events.
## v0.77.0 — 2026-05-29
pi `0.76.0` → `0.77.0` bump (first container build on the pi 0.77 line, published upstream 2026-05-28). Built against `joakimp/opencode-devbox:base-latest` (unchanged from the v0.76.0 build — same SSH-CM, gitleaks, git-crypt baked in).
### Bumped: pi 0.76.0 → 0.77.0
Notable upstream changes (from pi's CHANGELOG):
- **Claude Opus 4.8 support** — Anthropic Opus 4.8 model metadata + adaptive-thinking coverage updated.
- **Selective tool disablement** — `--exclude-tools` / `-xt` disables specific built-in, extension, or custom tools while leaving the rest available.
- **Headless Codex subscription login** — `/login` can use device-code auth for ChatGPT Plus/Pro Codex subscriptions; browser login remains the default.
- **Streaming-aware extension input** — `InputEvent.streamingBehavior` lets extensions distinguish idle prompts from mid-stream steers and queued follow-ups.
- **Bugfixes** — startup timing output excludes `createAgentSessionRuntime` work; OpenRouter DeepSeek V4 `xhigh` reasoning preserves OpenRouter's native effort; SIGTERM/SIGHUP exits run extension `session_shutdown` cleanup; keyboard protocol negotiation ignores delayed terminal responses (no false Kitty detection); Windows MSYS2 ucrt64 startup crash fixed via napi-rs 3.x clipboard addon; API-key/header config resolution treats plain strings as literals with `$ENV_VAR` / `${ENV_VAR}` interpolation and `$!` escaping; session disposal aborts in-flight agent/compaction/branch-summary/retry/bash work; `pi.getAllTools()` exposes per-tool `promptGuidelines`; OpenAI Codex Responses replay after switching from Anthropic extended-thinking sessions; Anthropic-compatible replay supports `allowEmptySignature` for providers returning empty thinking signatures; OpenAI/OpenRouter GPT-5.5 Pro thinking levels limited to supported efforts; OpenCode Go Kimi K2.6 thinking-off requests; Xiaomi Token Plan model metadata cleaned of unsupported variants; follow-up messages queued by `agent_end` extension handlers drain before idle; system prompt tool-selection guidance avoids unavailable file-exploration tools; fenced `diff` highlighting restored.
Workflow continues to derive `PI_VERSION` from the git tag (`v0.77.0` → `0.77.0`) and pass it as a build-arg per the v0.75.5b cache-hit fix; smoke test asserts `pi --version` matches.
### Inheritance from base
No base change in `joakimp/opencode-devbox:base-latest` since v0.76.0 — the v1.15.12 opencode-devbox release also reused the unchanged base. SSH ControlMaster on a writable socket path, gitleaks, and git-crypt continue to ride along from the base.
### CI
This is the second pi-devbox release exercising the cache-export-disabled workflow (after v0.76.0's clean publish on run #340) and the first to also exercise the 3-attempt retry wrapper added in 2d39766 along the publish path.
## v0.76.0 — 2026-05-28
pi `0.75.5` → `0.76.0` bump (first minor-version release on pi 0.76 line, published upstream 2026-05-27 20:03 UTC). Built against a fresh `joakimp/opencode-devbox:base-latest` which now bakes in SSH ControlMaster on a writable socket path, plus gitleaks and git-crypt — see the inherited-from-base notes below for details on each.
### Bumped: pi 0.75.5 → 0.76.0
Notable upstream changes (from pi's CHANGELOG):
- **Explicit session IDs for automation** — `--session-id <id>` lets scripts create or resume an exact project-local session.
- **RPC bash output can stay out of model context** — RPC clients can pass `excludeFromContext` to `bash` for commands whose output should not be sent with the next prompt.
- **More predictable provider retries and timeouts** — Codex WebSocket/SSE waits are bounded; `retry.provider.maxRetries` controls provider retries instead of hidden SDK defaults; SDK retries default to 0; quota/billing 429s are no longer retried behind Pi's retry handling.
- **Better terminal editing across environments** — Apple Terminal Shift+Enter detection on macOS, Windows Terminal OSC 8 hyperlink support, JetBrains truecolor with disabled OSC 8, Unicode-aware word navigation and deletion.
- **Bugfixes** — `pi update` bypasses npm/pnpm/Bun minimum-release-age gates; user-authored ordered-list markers preserved in transcripts; image attachment token estimates aligned with tool-result images; Codex Responses cache-affinity header fixed (`session-id` not `session_id`); OpenRouter/Poolside context-overflow detection; managed npm extension updates avoid peer-dependency conflicts; RpcClient handles unexpected child exits cleanly.
Workflow continues to derive `PI_VERSION` from the git tag (`v0.76.0` → `0.76.0`) and pass it as a build-arg, per the v0.75.5b cache-hit fix; smoke test asserts `pi --version` matches.
### Workflow change: registry cache-export disabled
- **`.gitea/workflows/docker-publish.yml`** — `cache-from`/`cache-to` removed from the `publish` step. buildkit's `mode=max` cache-export to `registry-1.docker.io` reproducibly returns HTTP 400 on the resumable-upload PUT, surfacing ~2026-05-23. Diagnosed during opencode-devbox v1.15.12's manual host-side publish: image push works fine, only `--cache-to` fails. See opencode-devbox CHANGELOG v1.15.12 `Unreleased` for the full root-cause analysis. The pi-devbox Dockerfile is single-stage with a tiny diff (npm install pi only) on top of `base-latest`, so builds are fast even without cache (~30-60s expected).
### Inherited from opencode-devbox base: SSH ControlMaster on a writable socket path
No Dockerfile change here — just a note that this release picks up the system-wide SSH ControlMaster default (`/etc/ssh/ssh_config.d/00-devbox-controlmaster.conf` → `ControlPath /tmp/sshcm/%r@%h:%p`, `ControlMaster auto`, `ControlPersist 10m`). This unblocks `ssh` and `pi --ssh user@host` from inside the container when `~/.ssh` is bind-mounted read-only from the host (the standard pi-devbox compose layout) — previously, OpenSSH's default `ControlPath` under `~/.ssh/cm/` was unwritable, so multiplexing failed with `unix_listener: cannot bind ... Read-only file system` and ssh fell back to fresh TCP connections, which on residential CGNAT manifested as banner-exchange timeouts. The fix is purely additive (per-container `/tmp/sshcm` dir, mode 700, created by entrypoint) and user `~/.ssh/config` per-host overrides still win because Debian's stock `ssh_config` sources `ssh_config.d/*.conf` before its own `Host *` block. See opencode-devbox CHANGELOG `v1.15.12` for the base-side details.
### Inherited from opencode-devbox base: gitleaks + git-crypt
No Dockerfile change here — just a note that this release includes `gitleaks` (newly added to the base) and `git-crypt` (was always installed via apt; just wasn't called out). Both are useful inside the container for repos that use a gitleaks pre-commit hook or git-crypt-encrypted canonical config and don't want host-side dependencies. See opencode-devbox CHANGELOG `v1.15.12` for the base-side details.
## v0.75.5b — 2026-05-23
Recovery release fixing a **silent cache-hit regression** discovered in the v0.75.5 image. All four releases v0.74.0 through v0.75.5 had been shipping the same image bytes because the Dockerfile's `npm install -g @earendil-works/pi-coding-agent` (bare, when `PI_VERSION=latest`) produces an identical layer-hash across builds. Combined with the registry buildcache, Docker reused the layer from whatever pi version was current when the cache was first populated.
Verification: `docker manifest inspect joakimp/pi-devbox:vX.Y.Z` showed identical SHA256 digests on both `linux/amd64` and `linux/arm64` for v0.74.0, v0.75.3, v0.75.4, v0.75.5. Users on `:latest` were getting whatever pi version was baked into the v0.74.0 build (probably 0.74.0 itself).
- **Workflow fix:** Both `smoke` and `publish` jobs now derive `PI_VERSION` from `github.ref_name` (e.g. `v0.75.5b` → `0.75.5`) and pass it as a build-arg. The Dockerfile's existing `if PI_VERSION=latest` branch never fires in CI now — always takes the `@${PI_VERSION}` branch — so the layer-hash includes the version and cache invalidates correctly.
- **Smoke test:** New `run_expect` helper asserts `pi --version` output contains `EXPECTED_PI_VERSION` (passed from the resolve step). Would have caught this regression on v0.75.3 if it had existed.
- **Dockerfile:** Comment added above `ARG PI_VERSION=latest` documenting the cache-hit footgun and pointing at the workflow's resolve step + AGENTS.md gotcha.
- **AGENTS.md:** New convention bullet explaining the cache-hit class of bug and noting the latent same-bug in opencode-devbox's `with-pi` variants (currently masked by OPENCODE_VERSION bumps).
No image-side changes vs v0.75.5 *intent* — this build will produce the actual pi 0.75.5 image content that v0.75.5 was supposed to ship.
## v0.75.5 — 2026-05-23
pi `0.75.4` → `0.75.5` bump (one upstream patch release, two days after v0.75.4).
Notable upstream changes (from pi's CHANGELOG):
- Cleaner read tool output (collapsed cards show only the read line; Ctrl+O expands).
- Faster file tools on Windows (async fs ops during streaming, image resize off the main TUI thread).
- More reliable package updates (`pi update` reconciles git-pinned refs without losing settings).
- Custom Anthropic-compatible adaptive thinking via `compat.forceAdaptiveThinking`.
- Several bash/read tool card display fixes; macOS Bun clipboard sidecar resolution; per-session OpenCode-Zen routing headers; Amazon Bedrock token cap fix.
Plus a new pi 0.74.2 rescue release advising Node 20 users to upgrade Node before going to newer Pi versions — the devbox base image runs newer Node so this doesn't affect us, but worth noting for users running pi outside the devbox.
- **Bump:** pi `@earendil-works/pi-coding-agent@0.75.5` baked at `/usr/bin/pi` (via `PI_VERSION=latest` resolving to 0.75.5 at build time — no Dockerfile change needed).
- No image-side changes from v0.75.4 beyond the pi npm version. Built on `joakimp/opencode-devbox:base-latest` which itself is unchanged (cache-hit on `base-35ee5fe7861a` since v1.14.50b).
## v0.75.4 — 2026-05-21
pi `0.75.3` → `0.75.4` bump (one upstream patch release). Plus the AGENTS.md documentation-drift sweep clause that landed on `main` between v0.75.3 and now.
- **Bump:** pi `@earendil-works/pi-coding-agent@0.75.4` baked at `/usr/bin/pi` (via `PI_VERSION=latest` resolving to 0.75.4 at build time — no Dockerfile change needed).
- **AGENTS.md:** documentation drift sweep as explicit pre-commit workflow step (commit `ae6253a`). Companion clause added across the wider repo set the same day.
- No image-side changes beyond the pi npm version. Built on `joakimp/opencode-devbox:base-latest` which itself is unchanged (cache-hit on `base-35ee5fe7861a` since v1.14.50b).
## v0.75.3 — 2026-05-18
pi `0.74.0` → `0.75.3` bump (one upstream minor + three patch releases since the initial pi-devbox release on 2026-05-14).
- **Bump:** pi `@earendil-works/pi-coding-agent@0.75.3` baked at `/usr/bin/pi` (via `PI_VERSION=latest` resolving to 0.75.3 at build time).
- No image-side changes from the v0.74.0 baseline beyond the pi npm version. The pi-toolkit + pi-extensions clones, mempalace bridge symlink, and `NPM_CONFIG_PREFIX` named-volume setup all unchanged.
## v0.74.0 — 2026-05-14
Initial release.
- pi `@earendil-works/pi-coding-agent@0.74.0` baked at `/usr/bin/pi`
- pi-toolkit and pi-extensions cloned at build time; deployed to `~/.pi/agent/` by entrypoint on container start
- mempalace bridge (`mempalace.ts`) symlinked from `/opt/mempalace-toolkit/`
- Built on `joakimp/opencode-devbox:base-latest`