Files
pi-devbox/CHANGELOG.md
T
joakimp 361babd4fd
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Publish Docker Image / smoke (push) Successful in 7m39s
Publish Docker Image / build-variant-studio (push) Successful in 17m15s
Publish Docker Image / build-variant (push) Successful in 17m58s
Publish Docker Image / update-description (push) Successful in 8s
Publish Docker Image / promote-base-latest (push) Successful in 12s
ci: gate the release on shell lint, from one shared script
v1.8.14's first attempt spent ~46 minutes building a base image for a tree whose
own lint had been failing for 24 hours. shellcheck had already flagged the
defect (SC2289, severity error) on the push that introduced it; the lint
workflow went red at run 186 and nobody read it.

lint.yml deliberately skips tag pushes and its reasoning is sound -- the tagged
tree was already linted on main, and a tag-ref lint run sorts above the publish
run, making a release look finished before anything ships. The missing invariant
was never "lint the tag". It was "do not RELEASE a tree whose lint failed", and
only a job inside the publish workflow can enforce that.

So: extract the shell-lint logic from lint.yml into scripts/lint-shell.sh and
call it from both places, then add a lint-gate job that resolve-versions depends
on. resolve-versions is the graph root, so gating it gates everything. Cost is
~40 s at the front of a release; the alternative already cost fifty minutes.

Extracted rather than copied on purpose. A second copy of a check is the drift
this repo keeps paying for -- the same evening produced a skillset mirror that
had sat 9579 B behind its upstream through two consecutive edits.

The script adds one behaviour the inline version lacked: if shellcheck is not
installed it exits 2 rather than silently finding nothing, inheriting the
existing "a gate that cannot run must not pass" rule from hooks/pre-commit in
the skillset repo. Without that, reordering the install step away would turn the
gate into a green tick over zero checks.

Verified locally with a stubbed shellcheck (the real binary is not in the
devbox), five cases, each with its expectation stated first: absent shellcheck
-> rc=2; stub pass -> rc=0 and a non-zero file count; stub fail -> rc=1; a
deliberately unterminated `if` planted in scripts/ -> rc=1 via the bash -n half,
naming the file; removal -> rc=0 again. Discovery cross-checks against CI's own
number: the inline version reported 12 files, the extracted one reports 13, the
difference being lint-shell.sh itself. YAML re-parsed (10 jobs, was 9) with an
assertion that resolve-versions needs lint-gate, and the repo's
check-workflow-shell.sh guard still passes.
2026-09-08 23:41:44 +02:00

280 KiB
Raw Permalink Blame History

Changelog

All notable changes to the pi-devbox container image.

From v1.0.0 onward, tags follow semver:

  • major — architectural changes (v1.0.0 = decoupled from opencode-devbox)
  • minor — new variants, significant base additions
  • patch — pi version bumps, smaller fixes

Pre-v1.0.0 tags followed the pi npm version (v{pi_version}[letter]).


v1.8.14 — 2026-09-08

First release attempt failed; fixed in this same entry. The smoke and smoke-studio jobs both failed at scripts/smoke-test.sh:770 with agent-browser: command not found, after build-base had already succeeded (~46 min spent). Root cause was in the agent-browser execution guard added the day before: the explanatory comment inside the single-quoted exec_test body contained an apostrophe (the fleet\'s). Inside '...' bash treats a backslash literally, so \' does not escape — it closes the string. The body silently truncated (measured: exec_test received 12 arguments instead of 2), and the remaining lines, including the agent-browser --version assertion, were parsed by the runner's shell instead of executing inside the image — and the runner has no agent-browser. The prose now lives above the call, where an apostrophe is harmless.

The lint job had already caught this, and it went unread for 24 hours. shellcheck flagged it as SC2289 at severity error, so the actionlint job went red at run 186 on 2026-09-07 21:21 — the exact push that introduced the guard — and stayed red for runs 187 and 188. lint.yml deliberately excludes tag pushes (documented: the tagged tree was already linted on main, and a tag-ref lint run would sort above the publish run), which is sound; the broken assumption was different, namely that a tree whose lint FAILED would not then be released. docker-publish.yml has no dependency on lint, so it built for 50 minutes on a tree known to be defective.

Fixed, then gated. The prose moved above the exec_test call so an apostrophe cannot terminate anything, and the shell-lint logic moved out of lint.yml into scripts/lint-shell.sh — now called by both lint.yml and a new lint-gate job here that resolve-versions depends on. A release with a lint error refuses in ~40 s instead of failing after fifty minutes. One copy, not two: a duplicated check that drifts is the failure this repo keeps paying for. The script also refuses to pass when shellcheck is absent, inheriting the existing principle that a gate which cannot run must not pass.

v1.8.14 was re-pointed from 601fc98 to the fix commit. Nothing had consumed the original tag — no v1.8.14 image was ever published, only the content-addressed base-a365dd24de21. scripts/ does not feed the base hash, so the re-run reuses that base and skips the 46-minute rebuild.

A test that was quietly checking nothing, and a version number that was wrong. Both found by delegating a read-only audit of this repo to a headless worker (pi-toolkit bin/pi-task) and then spot-checking its pointers from the filesystem — 5 of 5 held, and it also corrected a false premise planted in its own brief.

The node major is now asserted, not merely printed. scripts/smoke-test.sh ran run "node" "node --version", which asserts only that the binary exists and exits 0 — the printed version was compared to nothing. The line above it has always used run_expect against $EXPECTED_PI_VERSION for pi, so the suite looked like it covered node. A node major bump would have passed the whole smoke suite silently. Worse, this is where the "node v22.23.2 verified" line in the v1.8.13 recreate notes came from: printed output, not an assertion — an expectation stated up front and then falsified by the check.

Now gated on EXPECTED_NODE_MAJOR, which CI derives from Dockerfile.base's ARG NODE_VERSION — the single source of truth, and the only hard node pin in the repo (Dockerfile.variant has no node install at all, so the two Dockerfiles cannot disagree). That also catches a stale cached layer whose node disagrees with the declared ARG. Unset ⇒ previous behaviour, so nothing breaks for anyone running the suite by hand.

Verified two-sided, because a silent failure here reintroduces the exact bug it fixes: the sed derivation yields 22 (an empty result would disable the assertion silently); grep -Fq "v22." matches v22.23.2; "v24." does not match, so a wrong major is caught; "v2." does not prefix-collide. The workflow YAML was re-parsed after editing (9 jobs).

v1.8.13's agent-browser version was wrong. That entry said "the image's own 0.35.2". The image ships 0.36.0 — /usr/lib/node_modules/agent-browser at 0.36.0 with engines.node >=24.0.0, and no 0.35.2 exists anywhere in the image. The sentence was also internally incoherent, contrasting 0.36.0 against a version that is not present. Corrected in place with a visible note, since that entry is already released. The reasoning survives untouched: the engines floor really is vestigial, because /usr/bin/agent-browser is a prebuilt aarch64 ELF invoked directly and never through node — which is exactly why 0.36.0 runs fine on 22.23.2, consistent with the runtime proof collected on 2026-09-07 and with the retraction of the earlier false "0.36.0 requires node >= 24" alert.

No image content changes: NODE_VERSION still 22, no pins moved. This is a test and a docs correction only.

The same bug class, twice in one file — and the second one was throwing away a proof the fleet cannot obtain any other way. scripts/smoke-test.sh's agent-browser guard captured the version inside an echo, with 2>/dev/null:

echo "resolved=[$r] version=[$(agent-browser --version 2>/dev/null | head -n1)]" >&2

The exit code was discarded, so a binary that could not execute at all still passed, printing version=[]. Verified two-sided: a stub exiting 127 passes the old form and is caught by the new one.

Why that exit code matters more than most: smoke runs platforms: linux/amd64 on an x86 runner, i.e. native amd64, making this line the fleet's only recurring amd64 runtime proof for agent-browser's linux-x64 ELF. No devbox can ever supply one — every machine in the pi fleet is an Apple Silicon Mac (mbp-m1-2020; tor-ms22 = Mac Studio Mac13,1 M1 Max, verified 2026-08-17 by system_profiler; emb-7kj4vr4g = Apple Silicon, verified 4 ways 2026-09-07). The "amd64 runtime proof still needed" item that was sent to two devices was therefore asking for the impossible, while CI already had the answer and was discarding it. Dockerfile.base:607 does assert it (agent-browser --version &&), but only when the base actually rebuilds — and v1.8.13's base was cached.

The mailbox now announces replies that CLOSE your own asks. mempalace-toolkit 21023e7 → e45f6b4, which adds deriveClosed() alongside deriveOwed(). The old path queried status: open and joined for a reply, which by construction can only surface asks you owe someone else; a terminal reply carries status: applied (or blocked/failed), so the answer to your own question was structurally invisible — the one notification a human actually wants. Measured: emb-7kj4vr4g closed the v1.8.13 rollout ask at 18:31Z with status=applied, the operator reasonably expected to hear about it, and the mailbox stayed silent while being correct by its own definition. Nine closed correlations were sitting unannounced. Shares the 1-hour resurface floor, so a close is announced once and is news rather than a nag.

This lands because the base rebuilds, which is worth stating explicitly: the CI-resolved mempalace-toolkit SHA is folded into the content-addressed base tag (base-decide), precisely so a toolkit-only fix cannot silently fail to land behind an unchanged Dockerfile.base. The floating main ref was left alone on purpose — the toolkit moving forces the rebuild rather than waiting for one.

A consequence worth noting for the amd64 item above: this release actually collects that proof. v1.8.13's base was cached, which is why Dockerfile.base:607's agent-browser --version && never ran. v1.8.14's base is not cached, so both that assertion and the new EXPECTED_NODE_MAJOR gate execute on a native linux/amd64 runner. The fleet's first kept amd64 runtime proof for the linux-x64 ELF should be an artefact of this build rather than something asked of a device that cannot supply it.

Subtask delegation is documented — including the rung nobody built. The image picks these up through their resolved refs (pi-toolkit adfb553, pi-extensions c64c122, the latter also refreshing the baked fallback skill):

  • an operator-facing decision guide in pi-toolkit's README.md, built on the L0–L4 context ladder — how much of the parent session a child can see is the axis that explains nearly every observed good and bad behaviour;
  • the canonical pi-extensions skill gains the same ladder next to Boundary discipline, which until now diagnosed why an inherited transcript defeats a brief without offering any alternative to "don't fork that";
  • one bullet in the global AGENTS.md, so the choice is visible without loading a skill, and naming pi-task as a CLI — an agent hunting for a pi_task tool finds none and concludes it is unavailable.

What the ladder records: L0/L1/L2 exist in pi-task (context.facts / .files / .commands), L4 is fork's only behaviour (getHeader() + getBranch(), no offset or limit anywhere in the call chain), and L3 — a truncated branch — is not implemented by anything, which is now written down instead of being a design idea somebody remembers.

Also recorded, found while writing the above: pi-fork/src/runner.ts:188 reads if (extensions !== null) args.push("--no-extensions"). So extensions: [] turns the capability floor on and null turns it off — and null is the documented way to "restore normal extension loading", so tidying [] to null as a no-op re-arms palace writes inside every fork child. pi-task hardcodes the flag and cannot drift this way. Documented in three places because the edit that triggers it looks harmless.


v1.8.13 — 2026-09-06

Version audit + three pins moved, one deliberately not moved. pi 0.84.4 -> 0.85.1, mempalace 3.8.0 -> 3.9.0, pi-atelier v0.10.0 -> v0.10.1. PI_FORK_REF=master stays floating and therefore adopts e69725c. Each rationale is written at the ARG itself rather than only here, because that is where the next person doing the audit will be standing.

Correction, made mid-release while run 639 was building: the audit originally recorded a fourth change — "PI_STUDIO_VERSION relabelled none -> v0.9.60-rc.0, RC adopted deliberately" — and that was wrong. It was measured at the wrong layer. resolve-versions passes BOTH PI_STUDIO_REF and PI_STUDIO_VERSION as build-args and selects the newest stable semver tag (its filter ^v?[0-9]+\.[0-9]+\.[0-9]+$ excludes pre-releases), so a Dockerfile default cannot answer "what will CI publish?". Measured from the run itself: studio_tag=v0.9.59, studio_ref=9eed84f (= refs/tags/v0.9.59^{}), while main/v0.9.60-rc.0 is 658536f and is not built. Published v1.8.13 studio images therefore contain pi-studio v0.9.59, not the RC, and the ARG is back at none rather than pinned to a pre-release that goes stale the moment main moves. Consequence kept deliberately: the RC's opt-in Studio network binding is absent from every published v1.8.13 image, so it needs no audit for this release. Adopting an RC from CI would require changing that tag filter, which exists on purpose — upstream stopped publishing Releases at v0.5.55 but keeps tagging and pushing to main, so pinning main risked baking half-finished commits.

0.85.0 is SKIPPED on purpose: it shipped internal experimental code and extra subpaths that broke SDK imports (upstream #9132), and 0.85.1 exists to undo exactly that. Neither release has a Breaking/Removed changelog heading, the engine floor is unchanged (>=22.19.0 against the container's 22.23.2), and runtime deps drop 20 -> 19.

The pi bump was verified by RUNNING it, not by reading about it, because this repo has already been burned by a version pair that no changelog flagged (pi-atelier < 0.7.1 hangs pi >= 0.84 at startup with no error). 0.85.1 was side-installed and driven under a pty in five combinations — each companion extension plus atelier v0.10.0 AND v0.10.1 — with a CPU delta of 0.00-0.01s over a 5s window where the known hang signature is ~5s of sustained CPU. The check was two-sided: the atelier sidebar painted ACTIVITY+WORKSPACE markers identically to the 0.84.4 control, so "alive" could be distinguished from "silently absent".

NODE_VERSION stays 22 — audited, not overlooked. node 24 is technically safe: all five prebuilt native addons in pi use NAPI (ABI-stable, no NODE_MODULE_VERSION lock, no binding.gyp), nothing in the image declares a node CEILING, and the install is one token (setup_${NODE_VERSION}.x). agent-browser 0.36.0 declares engines.node >=24.0.0, but that field is vestigial for the artifact actually shipped: /usr/bin/agent-browser is the prebuilt aarch64 ELF bin/agent-browser-linux-arm64, invoked directly and never through node, so npm's engines floor is never enforced at runtime — verified running under 22.23.2 in this image. (Corrected 2026-09-07: this paragraph originally said "the image's own 0.35.2 declares the same floor". That was wrong and incoherent — it contrasted 0.36.0 against a 0.35.2 that does not exist in the image. There is exactly one agent-browser present, /usr/lib/node_modules/agent-browser at 0.36.0. The argument is unaffected; only the version was wrong.) The reason to wait is attribution, not compatibility — this release already moves pi a minor, mempalace a minor and bakes a Studio RC, so adding a node major would leave four suspects if the image misbehaves. Worth doing as its own release with the smoke suite as the gate. (v22 is in maintenance until 2027-04-30; v24 is Active LTS to 2026-10-20 and maintained to 2028-04-30, so there is real headroom.)

mempalace's client bump carries a sequencing note that is now also CORRECT: the comment at the ARG claimed synlig serves 3.7.1 server-side, which was stale. Measured 2026-09-06 over ssh, synlig's uv tool entry last changed 2026-08-25 and serves 3.8.0. Client 3.9.0 against server 3.8.0 is accepted skew until synlig's compose stack is redeployed; 3.9.0's headline additions (release awareness, task create/task launch) are SERVER-side and stay dark until then — a client bump alone cannot light them up.

agent-browser was running 7 weeks stale, and the interesting part is why nothing noticed. The image has shipped 0.35.2 since the last base rebuild, but every session on mbp-m1-2020 was executing 0.27.0 from a 2026-07-17 hand-install: npm i -g writes into ~/.pi/npm-global, which is the devbox-pi-config VOLUME, and PATH puts that at position 2 against /usr/bin at position 8. This is the third package hit by that exact hazard (pi itself and pi-atelier already have guards), so the guard is now generalised instead of re-invented a fourth time.

The damage was not the binary. It was the BUNDLED SKILL, which is the part an agent reads: 3 skillsets / 17.6 KB core in 0.27.0 versus 8 skillsets / 31.5 KB core in 0.35.2, with ten subcommands present in the image and entirely undocumented to the agent (a11y, browser, data, mcp, page, plugin, read, selectors, to, webmcp). A stale tool announces itself with an error; a stale skill just quietly teaches the wrong commands and everything looks fine.

Three changes, at the three places this can be caught:

  • entrypoint-user.sh retires a volume copy by MOVING it aside (reversible, same instinct as the settings backups) and only when the image ships its own copy, so a machine that deliberately hand-installs on an image without one keeps it. The bin/ shim is removed too — a dangling symlink would be a worse failure than a stale version.
  • scripts/recreate-sanity-check.sh asserts agent-browser resolves under /usr. This is the check that matters, because it runs where the volume is real.
  • scripts/smoke-test.sh gets the build-time half, labelled WEAK in the source for an honest reason: a docker run container has an empty config volume, so it can never see the shadowing it is nominally testing for.

pi-fork gets a capability floor: extensions: []. Forks were measured twice (2026-09-01, 2026-09-06, four dispatches) ignoring their brief, answering in the USER's voice, fabricating self-referential measurements, and once filing a diary entry as agent_name=pi — which landed in wing_pi, where a wing-scoped diary_read never sees it.

The cause is upstream and by design, so there is nothing to wait for: the child is handed getHeader()+getBranch(), i.e. the WHOLE active session branch, with the brief appended as the final user message and the system prompt untouched (pi-fork src/index.ts). In a long session the parent narrative simply outweighs the task, and the child does the statistically obvious thing — it continues the story it finds itself inside. Config offers no context knob (extensions, environment, offline, costFooter, effort profiles only).

Falsified the tempting explanation before acting on it: the failures are NOT a too-small model. The same model as the fast profile (haiku, thinking off) obeyed the identical brief perfectly when run as pi -p --mode json --session-id <fresh> --no-extensions — correct values, exact format, no session recap, 3 seconds, $0.012. Model held constant, context inheritance removed, failure gone.

extensions: [] is therefore a mechanical guarantee rather than an instruction: the mempalace bridge is a pi EXTENSION, so a fork child now runs with --no-extensions and cannot write to the shared palace under the parent's identity. Verified by asking a child to enumerate its own tools: read, bash, edit, write — no mempalace_*, no recall, no nested fork. Two honest limits, stated so nobody over-trusts this: it removes PALACE writes, not FILESYSTEM writes (edit/write remain), and it costs forks their palace search and recall. Set the key to null to restore normal loading.

Smoke asserts the floor is [] specifically, not merely falsy — null is the unguarded state, so a "truthy or not" test would pass on exactly the configuration being guarded against.

Vendored mempalace skill snapshot refreshed a12fe5e -> e9e09d9, and the phrase canary re-pinned with it. Folded in at zero marginal cost: the snapshot is hashed into base_tag, but Dockerfile.base already changed this release, so the ~67 min base rebuild was already being paid. --check reported exit 0 (stale-but-truthful) beforehand, i.e. skipping was sanctioned — this is the deliberate decision the checklist asks for, not a drive-by. Upstream content is the fleet wing-naming convention (bare project names, no wing_ prefix) and the <harness>@<device> rule for added_by, both of which came out of the attribution defect measured on this device on 2026-09-06.

The canary re-pin is the interesting half. Its old pair — "Provenance is stamped for you" present, "Attribute what you file yourself" absent — STILL PASSED against the new snapshot, so leaving it in place would have produced a canary that is green on both the old and the new bytes: blind to precisely the refresh it exists to witness, which is the same false-green family the pre-v1.8.5 canary died of. The replacement pair was picked by MEASURING direction against both files rather than by reading the diff ("Diaries self-heal; plain drawers do not" new=1/old=0; "Agent diaries live in" new=0/old=1) and then tested two-sided: PASS on the refreshed bytes, FAIL on the old bytes recovered from git. A canary that cannot fail is decoration.

credential-incident-response §5/§6 corrected — a stated mechanism was wrong, and this is the second time in three days this section named a wrong reason for a zero. Docs only.

§5 said embedding_metadata.string_value holds "metadata fields only". Measured false on chroma 1.5.9 with a disposable sentinel drawer (pi@tor-ms22, 2026-08-30): the document text is ALSO there, under key chroma:document — one row in fts_content and one in embedding_metadata for the same drawer. The scan order in §5 is unchanged (scan fts_content directly, raw bytes as backstop) but the stated REASON is fixed: a zero from string_value needs a different explanation (key filter, query shape, escaping), not "it's structurally blind". §6 already warns against explaining a zero with an unverified mechanism; this was exactly that failure, in the file that carries the warning.

§6's row-gone/bytes-gone claim is now backed by the same sentinel measurement rather than asserted: delete_by_source took both fts_content (1->0) and embedding_metadata (1->0) to zero, while raw bytes stayed 4->4 until VACUUM. Also records how the measurement got unblocked at all — not a better instrument, a disposable sentinel drawer instead of testing deletion on real data.


v1.8.12 — 2026-08-31

pi 0.84.3 → 0.84.4, and pi-atelier v0.8.2 → v0.10.0. Both audited by the routine in Dockerfile.variant rather than adopted on sight, and the audit notes live next to the pins where the next reader will meet them.

pi 0.84.4 (published 2026-08-28) carries no Breaking Changes and no Removed heading — checked by grepping the section, 0 matches, which is worth stating because 0.84.3 did have one. It was adopted for three fixes that land on machinery this fleet runs every day, not for the feature list:

  • #6879 — a large tool result crossing the auto-compaction threshold used to be sent to the provider before compaction. Pi now compacts between tool execution and the next assistant response inside the same run. That is the shape of nearly every session on these boxes, where a single event_list or palace search returns hundreds of KB.
  • #8345 — a resumed session corrupted its next appended entry when the JSONL file lacked a trailing newline. That file is the memory feeder's input, so the failure would have surfaced as unexplained gaps in wing_conversations rather than as an error. Measured on tor-ms22 before bumping: 49/49 transcripts end in a newline and 0 lines fail json.loads — this corpus was never bitten, and we now know that rather than hope it.
  • #8537 — extension messages sent with triggerTurn: false while the agent is running were inserted between a tool call and its result, so order-validating providers rejected the replayed history. The mempalace mailbox is outside that precondition: it delivers at agent_settled, when no inference is in flight, with {deliverAs: "steer"} and deliberately no triggerTurn. 0.84.4 also leaves the documented steer semantics untouched ("delivered after the current assistant turn finishes executing its tool calls, before the next LLM call"), so RFC 003 §7.11 stands as written. Recorded because this fix is precisely what would make a mid-run delivery safe, which is the only reason we would ever change that call.

Also new and relevant, though nothing here uses them yet: ui_prompt_start / ui_prompt_end extension events (the docs/extensions.md diff is add-only — no steer or triggerTurn semantics moved), and an RPC clear_queue that returns and removes queued steering messages. The second one can discard an already-delivered but unconsumed mailbox steer; that is survivable because the mailbox re-delivers on MEMPALACE_MAILBOX_RESURFACE_MS (default 3600000), and it is written down here so a future "the mailbox lost a message" report has a candidate cause. The three new PI_HYPERLINKS / PI_IMAGE_PROTOCOL / PI_TRUE_COLOR environment variables were grepped against this whole repo: no collisions with anything the image sets.

The bump moved one documented mechanism, so docs/observational-memory.md §3 moved with it. Pi's own docs/compaction.md gained exactly one paragraph in 0.84.4: the autoCompact threshold is now also checked mid-run, after a tool batch's results are appended and before the next assistant response, skipped only when that batch ends the run and no queued message needs another response. Our doc said compaction is "checked when pi goes idle, so it never interrupts a turn". That was only ever true of observational-memory's own trigger (compaction-trigger.ts hooks agent_settled); read as a statement about pi it is now false. session_before_compact (compaction-hook.ts) therefore has two entry points and the second can fire inside a turn — harmless for the ledger fold, which makes no model call, but a doc that ships a false promise about when a hook runs is worse than one that admits two paths. The §3 mermaid diagram gained the second edge, and the whole file re-passes the bundled mermaid checker (6 blocks, 44 labels, 0 soft-wrapped, no cut glyphs at 1280px and 800px).

pi-atelier v0.8.2 → v0.10.0 is two minor releases and both are UI-only — Sidebar kept calm during an active Turn, composer frame and Status Rail polish, fullscreen-copy-safe Sidebar, Windows path normalisation, Workspace Pulse deferred until pi trusts the project. Neither release carries a BREAKING notice. The coupling that matters runs the opposite way to this pin's hard-earned floor: v0.9.0 renders the Sidebar as a separate split-layout child and therefore "raises the minimum supported Pi version to 0.84.0", and — unlike the 0.7.1-under-pi-0.84 startup-hang precedent, which its metadata never encoded — this time peerDependencies says so (>=0.84.0, up from >=0.80.7). Satisfied with room to spare by PI_VERSION=0.84.4. It also pairs deliberately with a 0.84.4 feature: atelier keeps Sidebar content out of the fullscreen transcript selection while pi adds fullscreenCopyOnSelect and Ctrl+X for the selection itself. Both executable floors (scripts/smoke-test.sh, scripts/recreate-sanity-check.sh) compare with sort -V, so 0.10.0 >= 0.7.1 is evaluated correctly — verified by running the comparison, because the string form of that test reads 0.10.0 as older than 0.7.1.

While bumping the pins, the README's own pin table turned out to have been wrong since v1.8.6. It advertised pi 0.84.2 and mempalace 3.7.1 in the very table whose purpose is to tell a reader what is pinned and where. Both rows went stale in the same commit — 93f986e (v1.8.6, "adopt pi 0.84.3 + mempalace 3.8.0") moved both ARGs and neither table row; the rows themselves date from 29b6209 (v1.8.0) and 2ebf00d (v1.8.4). Only atelier's row was still true. All three corrected now, and the --expected-version 0.84.3 example in the recreate-sanity section updated too, since that one is a copy-pasteable command that would now fail against a 0.84.4 image. Worth noting how it survived two releases: nothing checks prose against the ARGs, so this table has to be remembered by hand on every pin bump, and once it was not.

credential-incident-response gained the section its own guidance had been missing, and §2 gained a precondition it should always have carried. Docs only; no image behaviour moves. Both changes came out of a session where three separate detectors reported clean over secrets that were really there — the skill was the artifact that had taught two agents the pattern, so the fix belongs here rather than in either operator's private notes.

§2 previously said an 8-hex fingerprint lets you compare a credential "without ever materialising the secret", with no condition attached. That is true only when the input space is unreachable. A fingerprint is 32 bits over whatever it was computed from, so publishing fp8(x) hands anyone a membership oracle: they can test x == v for every candidate v they can generate. For a 40-char random token, fine. For a hostname, username, e-mail, port, path, commit SHA or weak password, that candidate set is a wordlist — and note that "high entropy" is the usual sufficient condition, not the test: a commit SHA is 160-bit and still fully enumerable from the repo. Two agents on this fleet published fingerprints of GIT_USER_EMAIL-class values while following this section as written; harmless in that instance, because those values sit in every commit trailer already, but the guidance licensed it. §2 now states the precondition, adds that candidate fingerprints are working memory and never output (a scanner hashes hostnames and paths too, so "print what it saw" leaks wholesale), and names what a fingerprint register is — a confirmation oracle for anyone already holding a candidate corpus, which is exactly how a retired token gets identified in old transcripts, and works the same way for someone else holding those files.

New §6, "Proving absence: instrument strength, and four ways a scan lies clean". Deliberately placed next to §5, because §5 optimises against false positives (name-anchoring, provenance — what stops a triage sweep drowning in session UUIDs) and every failure in §6 is a false negative. Triage optimises precision; a gate optimises recall, and conflating the two is what produced the clean reports. It carries: an instrument-strength ranking (exact-byte value search

class/structure pass > fingerprint census) with the standing instruction to say which one produced your zero; census and class passes answering different questions, with both failure modes measured here — a class-only pre-commit hook passed plaintext UUID API credentials to a shared repo twice because a UUID has no key header, while a census-only gate reported 0 hits with freshly-synced SSH private keys in the tree because no key is in the census; the tokenisation trap, where maximal-run extraction swallows an unquoted VAR=<uuid> so the value is never hashed alone while a quoted one is found, meaning quoting alone decided detectability; scan the index or the pushed tree, never the working tree, plus why a repo-only fix on an rsync-published mirror is temporary rather than weaker; git filters never running on symlinks, where check-attr answers git-crypt for a path it can never encrypt, so a coverage audit must join the attribute against the file mode and verify the blob magic; two-sided self-tests that abort, including the fixture-interaction artifact where a quoted and unquoted probe share one buffer and make the weak extractor look as strong as the union; and row-gone is not bytes-gone, since a correct sqlite DELETE leaves the payload in freelist pages until VACUUM.

Findings contributed by pi@emb-7kj4vr4g (the census/class split, and the instrument ranking's provenance) and pi@tor-ms22 (exact-byte value search over index blobs). The description's trigger list grew accordingly and is 1022/1024 characters — it has almost no headroom, so trim before adding to it, or the skill silently fails to load.

Deployment: the skill is baked at /usr/local/share/pi-devbox/skills/credential-incident-response/, so this needs an image rebuild and a container recreate to reach any running container.

Two vendored skills changed, and one of the changes is a correction rather than an addition. Nothing about the image's behaviour moves; this is entirely about what the next agent reads before it acts.

pi-devbox-environment §2 had a rule that was half wrong, and the wrong half cost five findings in one session. The section "A negative result is usually your own filter" closed with "a positive result needs no such scepticism — it carries its own evidence." That sentence is false. A positive result is evidence about the question your command actually posed, which may not be the question you meant — and the failure is invisible precisely because the command succeeded. Three measured instances, all from 2026-08-29, all filed as fact before being caught: an SSH handshake that succeeded and greeted the agent as joakimp while it believed it was probing gitea.egl.lan (a Host gitea* block had rewritten HostName, so it authenticated to the wrong Gitea instance); a 401 that was a genuine answer from an issuer which had never minted the credential being tested; and a "regression" produced by diffing ssh -G output against a 2222 that the agent's own earlier -p 2222 flag had supplied. The section now carries a counterpart, "…and a positive result only proves what you actually asked", plus the three false-negative rows that session added (a palace scan that queried embedding_metadata while documents live in embedding_fulltext_search_content; a token declared dead on a 401 from the wrong issuer; a host declared unreachable after trying two of its three open ports, with the port written in an environment variable the agent already held).

The cross-cutting form of that rule went into pi-global-AGENTS.append.md, not into the skill — deliberately, and this is the whole point of the change. The rule already existed in the baked skill, authored by an earlier session, symlinked into ~/.agents/skills/ at every container start. It survived every recreate, was available for the entire session that broke it, and was violated five times anyway. So the gap was never persistence; it was activation. A reasoning rule that only loads when a task description happens to match it cannot fire on the occasions that need it, because "I am about to state something false" is not a recognisable task type. The always-appended block is read by every agent in every container without being asked for, which is the only property that matters here. Writing a sixth document restating the rule would have felt like progress and changed nothing.

New baked skill: credential-incident-response. Authored here, so the baked copy is canonical and it is not listed in skillset-owned.txt. It carries the facts a two-day credential incident produced, on the theory that facts transfer between sessions where exhortations do not: probe the issuing provider first (11 of 13 "exposed" credentials in that sweep turned out to be already dead at the provider — five HTTP requests would have established it, and nobody asked); sha256[:8] fingerprints as leak-free credential identity; the 403-vs-401 trap that scoped tokens introduce into liveness probes, where a live token looks revoked on /api/v1/user; revocation beats deletion for anything already replicated, because deletion is best-effort over an unbounded copy set (FTS shadow rows, per-host feed inboxes, sqlite free pages, mesh replicas, backups) while revocation invalidates copies nobody enumerated; the three places a secret hides in a Chroma palace, in coverage order; deriving least-privilege scopes from measured consumers; and the exposures rotation does not fix (cleartext channels, git history, agent-authored drawers).

Three smoke assertions extended so a rebuild cannot silently drop the new skill: baked-file existence, resolves-to-the-baked-tree, and reported as baked by pi-devbox-version. Skill directories are picked up by a glob in entrypoint-user.sh, so no registration was needed — verified rather than assumed, since an enumerated list would have left the skill inert, which would have been a fitting way for this skill to fail.

Neither skills change reaches a running container until the image is rebuilt and the container recreated: ~/.agents/skills/ and the global AGENTS.md both live in the image, not in a volume or a mount.

cli_utils' shell functions are now sourced, closing the half of that wiring the image never did. v1.8.11 linked the repo's bin/ commands into ~/.local/bin so they resolve in non-interactive shells; nothing ever sourced cli_utils.sh, so its 14 functions (fgit, fhist, fssh, fdocker, fmark, fproc, fex, fenv, extract, mkcd, pathls, portcheck, agents-sync, up) were missing from every interactive shell whose $HOME had no zsh rc. That is the normal case, not an edge case: the container's interactive shell is bash and zsh is not installed in the image. A symlink cannot carry a shell function and a function cannot be reached from a non-interactive shell, so the two mechanisms are disjoint and both are required — the image had been paying this layer's dependency cost (fzf, bat, fd, rg, jq are baked partly for these functions) while delivering none of its benefit. Now sourced from /etc/skel-devbox/.bash_aliases, with the same detection order as the symlink block so commands and functions can never come from two different clones. CLI_UTILS_SOURCE=0 opts out, deliberately independent of CLI_UTILS_LINK=0 because the two disable independent mechanisms. Measured: all 14 resolve in a freshly-seeded $HOME, the opt-out is honoured, an absent checkout is a genuinely silent no-op (no output, no leaked _cu variable), and interactive shell startup goes from 12 ms to 17 ms.

Named explicitly, per this repo's own floating-ref rule: /workspace/cli_utils is a host bind mount, not a pinned ref. Sourcing it means the image now executes content it does not pin, on every interactive shell, on every device. It is bash-safe today and that was measured rather than assumed — sourcing under bash --noprofile --norc exits 0 and defines all 14 despite the *.zsh filenames, the functions run, and the tree's single zsh-only construct (print -z in fzf/fhist.zsh) is already guarded by [[ -n $ZSH_VERSION ]] with a bash fallback. The residual risk is future content: a cli_utils commit adding a genuinely zsh-only file would surface as parse errors at every prompt, fleet-wide. Errors are therefore left visible rather than sent to /dev/null, so the failure is diagnosable, and CLI_UTILS_SOURCE=0 is the one-line escape hatch.

iproute2 is installed, so the container can answer "what is listening in here". Neither ss nor ip was present in any image up to and including v1.8.11 — nor lsof, nor netstat — which made cli_utils' portcheck a hard stub that printed portcheck requires at least one of: ss, lsof, netstat and exited. ss satisfies its preferred branch (ss -tlnp), which is also the only branch that reports the owning PID. net-tools is deliberately not added (netstat is deprecated and only a fallback path) and neither is lsof (~500 KB for a third route to the same answer). Cost measured, not estimated: ~5.5 MB total — iproute2 is 4.2 MB and pulls six libs under --no-install-recommends (libbpf1, libmnl0, libtirpc-common, libtirpc3t64, libxtables12, libcap2-bin; libpam-cap is a Recommends and is correctly dropped). Verified in a live container: ss at /usr/bin/ss, ip at /usr/sbin/ip, both already on the developer PATH, and portcheck --all then correctly identifies the socat listener on 8765.

The two changes above also need a rebuild and a recreate, for a different reason than the skills: $HOME is the container's writable layer rather than a named volume (verified — ~/.bash_aliases carries the container's start mtime while ~/.bashrc carries the image's), so the skel file is re-seeded on every recreate. A $HOME/.bash_aliases that is bind-mounted from the host is still never overwritten, which is the existing contract.

Dependency audit (2026-08-31)

Every component checked against upstream by direct command, not assumed:

Component Baked in v1.8.11 Upstream now Action
pi 0.84.3 (pinned) 0.84.4 is npm latest bumped + audited (above)
pi-atelier v0.8.2 (pinned) v0.10.0 highest tag bumped + audited (above)
mempalace 3.8.0 (pinned) 3.8.0 is PyPI latest none
skillset (mempalace fallback snapshot) a12fe5e a12fe5e == origin/main, 0 commits since none — --check reports OK, no NOTICE
mempalace-toolkit 21023e7 21023e7 none
pi-toolkit 0e1369e 0e1369e none
pi-extensions 2022887 2022887 none
pi-fork bf702b4 bf702b4 none
pi-observational-memory ce9fc98 ce9fc98 (v3.0.4, peerDeps * → no pi floor to clear) none
pi-studio (studio variant) 3328b3d 3328b3d none
floating *_VERSION=latest tools (16) — 14 already at latest; git-lfs 3.7.1→3.8.0 (feature, no breaking section), uv 0.12.6→0.12.7 (patch) adopted implicitly by the rebuild; named here per this repo's floating-ref rule
node major pin 22, installed v22.23.2 v22.23.2 is the newest 22.x none — a newer LTS line (24.x) exists and is deliberately not tracked

Two method notes, because both would have produced a confident wrong answer:

  • An annotated tag's ls-remote SHA is the tag object, not the commit. refs/tags/v0.8.2 is 6e07bf85 while refs/tags/v0.8.2^{} is 159f34cf — the value actually baked. Comparing the un-dereferenced form reported pi-atelier as drifted from its own pin, which would have been a false integrity alarm about the one component whose pin is load-bearing. Always deref with ^{} before calling a pin broken.
  • git ls-remote --tags | sort -V | tail is not a "latest release" proxy. typst/typst carries date-style tags (v23-03-28) and mikefarah/yq carries vTestA/vTestB; both sort after the real releases. Dockerfile.base itself resolves latest by reading the Location of curl -sI …/releases/latest, so replaying that exact step is both noise-immune and the same source of truth the build will see.

v1.8.11 — 2026-08-27

Shell state that the writable layer eats on every recreate now gets rebuilt at start. Two additions to entrypoint-user.sh, both idempotent, both silent no-ops when the thing they wire up is absent.

cli_utils commands are linked onto PATH. If a cli_utils checkout is mounted, every executable in its bin/ is symlinked into ~/.local/bin at container start — git-status-all, git-pull-all, devbox-sanity, pi-devbox-sanity, pi-session-repair, docker-clean, vpn-status. Detection: CLI_UTILS_CONTAINER_PATH → /workspace/cli_utils → $HOME/cli_utils → /workspace/*/cli_utils; CLI_UTILS_LINK=0 disables it.

The reason this is an image concern and not the user's problem to re-solve: on a host, cli_utils/install.sh puts those commands on PATH by symlinking them into ~/.local/bin, which is persistent there — and ephemeral here. Same installer, same repo, opposite durability, so the fix died on every --force-recreate and the next session was back to typing /workspace/cli_utils/bin/git-status-all. Running install.sh inside a container is the trap rather than the fix: it re-creates the same disposable state.

Symlinks rather than a PATH edit in an rc file, deliberately. ~/.local/bin is already ahead of /usr/local/bin in ENV PATH, so links resolve in non-interactive shells too — docker exec <c> git-status-all, agent tool shells, scripts. An rc-file PATH edit cannot reach those: ~/.bashrc returns early when the shell is not interactive. Measured on tor-ms22 2026-08-27, command -v git-status-all failed in a non-interactive shell while succeeding in an interactive one, from exactly that asymmetry. Guards, because ~/.local/bin is shared with other tooling: a real file is never clobbered, a symlink pointing somewhere else is never stolen, our own links are refreshed, and links into a cli_utils/bin whose target vanished are pruned — a dangling link on PATH reports "No such file or directory" and reads as a broken container rather than a removed script.

A per-device boot hook: ~/.config/devbox-shell/init.sh. If the host provides one, it runs once at start with output to ~/.pi/agent/devbox-init.log. That directory is the host-owned bind-mount already sourced into every interactive shell by /etc/skel-devbox/.bash_aliases, so this is its boot-time twin — the same ownership and the same persistence, but running before any shell, which is what non-interactive fixups (symlinks, directories, one-off migrations) need. It introduces no new trust boundary: that path is already arbitrary code from the same owner; only when it runs is new. Invoked as bash <file>, never sourced, and its exit status is ignored — a hook must not be able to mutate the entrypoint's own shell state or stop a container from starting.

With the hook in place, the next "can this run on every recreate?" question needs no image change at all — which is the point, given what the next paragraph costs.

This moves the base hash. base-decide folds cat entrypoint.sh entrypoint-user.sh into it, so this change forces the ~40-minute base rebuild at the next tag whether or not anything else in the base moved. It is a rider, not a reason to tag.

How it was validated, since CI cannot. docker-publish.yml runs only on push: tags: v*, and lint.yml runs actionlint over workflow run: steps — neither one executes entrypoint-user.sh. So both sections were extracted and run against fixtures in a throwaway $HOME before commit: real file not clobbered, foreign symlink respected, stale link pruned, new command picked up, second run byte-identical, CLI_UTILS_LINK=0 honoured, and "no cli_utils anywhere" a silent exit 0. Then run for real in a live v1.8.10 container, after which command -v git-status-all resolved in a non-interactive shell. No smoke-test.sh assertion was added on purpose: the positive path needs a /workspace mount that smoke does not have, and asserting it there would repeat the v1.8.0 mistake of a smoke assertion written against a stage that does not exist at run time. workflow_dispatch with smoke_only remains the way to exercise this against HEAD before a tag.

Also carried by the floating mempalace-toolkit main ref (resolved at build time, not by a pi-devbox commit — MEMPALACE_TOOLKIT_REF=main):

A scrubbed re-export of a dormant session could silently never reach the palace host. bin/mempalace-pi-session ships to the palace with rsync -a --update, and the stage file's mtime is deliberately the SOURCE transcript's mtime (os.utime(), "preserve session mtime for dedup stability"). Re-exporting a session that has not been appended to since its last ship therefore produces a mtime that is not newer than the receiver's — exactly the case a redactor upgrade needs to ship, since content differs while mtime does not. --update reported success and sent nothing. Found and patched by pi@mbp-m1-2020 (mempalace-toolkit a361b71): --update → --checksum, which compares content and ignores size/mtime entirely. Dropping --update outright was considered and rejected — rsync's default quick check already transfers on a size difference alone, which would have masked the next instance of this (a redaction whose placeholder happens to match the secret's length) as fixed. os.utime() is untouched; its backdating is a separate, load-bearing design call for dedup stability. New regression test, scripts/test-rsync-ship-idempotency.sh, runs fully offline (a local rsync destination exercises the same size/mtime/checksum comparison as the ssh transfer) and is built to discriminate: it must fail against --update and pass against --checksum, not merely exercise the code path — the first draft of the test used fixture strings of different lengths and passed for the wrong reason (rsync's quick check transfers on size difference alone regardless of --update), which is the same trap the patch itself was written to avoid. Acceptance line for this class of change going forward: "receiver sha256 matches sender for every staged file", not "local stage is clean" — a clean local stage says nothing about what a dormant session already sent.

An event addressed to an identity no session runs as is delivered to nobody, and this fleet has now hit it three separate ways. RFC 003 gains §7.13 and open-decision 10 (mempalace-toolkit 21023e7, docs only, no image behaviour change): the owed-set derivation — the log's only push channel — is keyed on to_agent, and a reply is always addressed back to whatever string the original writer put in from_agent. Nothing validates that string against a live session identity, so authoring under a synthetic or foreign name makes every reply to that event write-only. Measured cost this cycle: a directed ask planted under a synthetic sender drew a correct reply containing an urgent security finding, and it sat unread for ~2h20m, found only because a human asked whether mail had arrived. Permitted exception, unchanged: a synthetic sender is fine for a deliberate control experiment, provided the body names the real identity to reply to.

Also carried by the live skillset mount (each device's own clone, not baked — except the mempalace skill's fallback snapshot, re-vendored below):

The mermaid-diagrams checker's cut gate moved from client pixels to a per-SVG user-space unit. CUT_PX was calibrated against one live page at one render scale; sweeping --viewport 500→1600 on an unchanged document moved the worst overflow −1.0px → −3.0px, near-proportional to the viewport, i.e. a constant geometric overflow viewed through a changing scale. cutU = cutPx / scale (scale taken per-SVG, never a page average — one page mixes scales 0.643–0.988) recovers that invariant: the sweep now collapses to exactly −3.0u at every width. Re-deriving the threshold against the live host surfaced a real false negative the old pixel gate had: a label at cutPx=0.4, scale=0.678 read as healthy under CUT_PX=0.5 but is 0.59u — a genuine cut hiding behind a compressed render scale. CUT_U stays 0.5; cutPx and scale are still printed on every issue so a devtools ruler still confirms the number on the actual page. A new, explicitly-deferred finding from the same review: cut only measures vertically, so an unbreakable token wider than its box (a long URL, a snake_case identifier) is invisible to soft-wrap, tall, and cut simultaneously — filed as a backlog item, not implemented, pending a fifth acceptance control.

The from_agent-identity finding above is also now in the mempalace skill itself ("Writing to another machine", and Anti-Patterns), and the baked fallback snapshot of that skill was refreshed to match (vendor-mempalace-skill.sh, 6eb20af → a12fe5e) — sanctioned to skip on its own (--check reported stale-but-truthful), done anyway because this release's point is getting today's fixes live, and the base rebuild below was already forced regardless.

Dependency audit (2026-08-27)

Every component checked against upstream by direct command, not assumed:

Component Baked in v1.8.10 Upstream now Action
mempalace-toolkit b2b50af 21023e7 ships the rsync ship-fix + RFC 003 §7.13 (both above)
skillset (mempalace fallback snapshot) 6eb20af a12fe5e re-vendored (above); live-mounted devices already had it
pi 0.84.3 (pinned) 0.84.3 is npm latest none
mempalace 3.8.0 (pinned) 3.8.0 is PyPI latest none
pi-atelier v0.8.2 (pinned) v0.8.2 highest tag none
pi-studio (studio variant) v0.9.52 v0.9.52 — main's commit and the tag's commit are identical (0 either direction) none
pi-toolkit 0e1369e 0e1369e (local clone HEAD == origin/main) none
pi-extensions 2022887 2022887 (local clone HEAD == origin/main) none
pi-fork bf702b4 bf702b4 none
pi-observational-memory ce9fc98 ce9fc98 none

pi-toolkit / pi-extensions checked against their actual Gitea origin (the Dockerfile's PI_TOOLKIT_REPO / PI_EXTENSIONS_REPO), not a GitHub mirror — querying api.github.com for those two returned nothing (rate-limited or blocked; not investigated, the local clones are the source of truth anyway). No SHA above is fork-supplied; a fork inventing plausible-looking upstream SHAs is a recorded failure mode (v1.8.9), so every value here came from git ls-remote, a local clone's own origin/HEAD, npm view/registry JSON, or the PyPI JSON API, run directly.


v1.8.10 — 2026-08-27

This tag exists to deploy a fix and a safety net that are currently running on exactly one machine. The feeder scrubber has been hand-copied to /opt on one device since this morning; every other device has kept staging unscrubbed transcripts into the shared palace. Nothing here is a new capability for its own sake.

MEMPALACE_TOOLKIT_REF=main floats: docker-publish.yml resolves it to a concrete SHA at build time, so whatever is on toolkit main when the tag is pushed ships in that image whether or not this repo has a commit. That is the rule v1.8.9 adopted after 553d865/5b8d78f shipped undocumented twice — name the behaviour change before tagging, not after — and this entry is that rule being obeyed rather than re-learned.

mempalace-toolkit main moves 5b8d78f → b2b50af (13 commits, ~2100 insertions / ~520 deletions). No pi-devbox commit implements any of it.

⚠️ THE ONE CHECK THIS RELEASE MUST NOT SKIP

On the first client that runs this image, verify the memory feed still stages. The feeder is fail-closed by design: no redactor module, no staging (exit 3). That is correct behaviour and it is also the failure mode with no alarm — a packaging or path mistake stops the fleet's entire transcript feed and nothing complains loudly, because refusing to stage looks exactly like a quiet session.

This is not hypothetical. f0bffd1 exists because the feeder is installed as a symlink (/usr/local/bin/mempalace-pi-session → /opt/mempalace-toolkit/bin/…) and ${BASH_SOURCE[0]} reports the symlink path, so the module lookup landed in a directory where it does not exist. Had that shipped, every device would have refused to stage on first boot. It was caught by execution, not by review.

Acceptance, in order, on the first recreated client:

  1. Run a session, then confirm the feeder logged a scrub summary — a [scrub] line with tier-tagged counts (T1:env-value=…, T2:github-pat=…), or an explicit "zero redactions". Silence is the failure signal, not success.
  2. Confirm the palace drawer count moved for that session (the feed reached the server, not just the stager).
  3. Confirm exit 3 did not fire: mempalace-pi-session invoked through the /usr/local/bin symlink must find mempalace_redact.py.
  4. Only then trust the rest of this release.

If step 1 or 3 fails, the memory feed is down fleet-wide until it is fixed, and sessions that ran in the meantime are not recoverable from the palace — they were never staged. MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 is the loud escape hatch, and using it means accepting unscrubbed transcripts until the packaging is repaired.

Dependency audit (2026-08-27)

Every component checked against upstream, not assumed:

Component Baked in v1.8.9 Upstream now Action
mempalace-toolkit 5b8d78f b2b50af ships the scrubber + symlink fix + hlc join
pi-studio (studio variant) v0.9.48 v0.9.52 22 commits, additive only — see below
pi 0.84.3 (pinned) 0.84.3 is npm latest none
mempalace 3.8.0 (pinned) 3.8.0 is PyPI latest none
pi-atelier v0.8.2 (pinned) v0.8.2 highest tag none
playwright 1.62.1 (floats latest) 1.62.1 none — no drift this cycle
pi-fork bf702b4 bf702b4 (2026-08-24) none
pi-observational-memory ce9fc98 (v3.0.4) ce9fc98 none
pi-toolkit 0e1369e 0e1369e (2026-08-07) none
pi-extensions 2022887 2022887 (2026-08-17) none

pi-studio v0.9.48 → v0.9.52 — four releases, 22 commits, all additive: PDFs open directly in Studio with watched previews, the header can hide, and contextual side questions arrive (selected-tool use, frozen git context, export, keyboard shortcuts). No removals or renames in the diff; the changes are concentrated in client/studio-client.js, index.ts and three new shared/ helpers.

Its pi floor is >=0.84.3 and we pin exactly 0.84.3 — satisfied with zero headroom. Worth naming as a watch item rather than a problem: the next studio release that raises the floor breaks the studio variant until PI_VERSION moves, and that failure surfaces at build time in the studio job only, after the core variant has already published.

Also pulled in by the floating toolkit ref (documentation only)

RFC 003 gains §9.2, a proposed direction for the one open decision this fleet keeps tripping over — that a report addressed to a device is never delivered, because mailbox candidacy requires exactly status="open". It records a negative result worth keeping: widening the owed set to include terminal events cannot work, since the asserting shape and the clearing shape must be disjoint or every closure mints a fresh obligation. No code implements §9.2 in this release.

The skillset snapshot also moves, so this image bakes the mermaid-diagrams skill's Playwright driver and the honest note that a claimed ack notifies nobody.

Transcripts get scrubbed before they are staged (3d47937, 836e35b, f0bffd1)

bin/mempalace_redact.py, called from mempalace-pi-session at the moment the staged transcript is written — one hook covering both transports, because local mode mines that file and remote mode rsyncs the same bytes.

  • Why it exists, measured rather than argued. One leaked bearer token had reached 3 drawers, 13 feeder inbox files across all three devices, and 10 local files spanning 10 days, from an agent printing an env var while debugging. A second sweep then found GITEA_ACCESS_TOKEN in 2 more drawers and GITEA_EGL_ACCESS_TOKEN in 3. This is routine agent behaviour, so the fix belongs in the pipeline, not in discipline.
  • Detection is name-anchored, never entropy-anchored. A palace's own primary keys — drawer ids, chunk ids, event ids, replica ids, HLCs, commit SHAs — are its high-entropy strings, so an entropy detector eats the memory it protects, silently and unrecoverably. Three tiers instead: T1 literal values from this process's env whose name says secret (zero false positives by construction); T2 vendor shapes (ghp_, glpat-, xox*-, sk-, AKIA, JWT, PEM, URL credentials, Authorization:); T3 key-name-says-secret.
  • T3 is report-only, because the false-positive rate was measured. On 52 MB of real fleet transcripts T3 fired 403 times, mostly ${VAR} interpolation in compose files, TypeScript identifiers, a type annotation (credentials: Credentials), an IPA attribute holding a date (krbPasswordExpiration), AAAK diary shorthand, and terminal output following an ssh Password: prompt. With interpolation/code-context/key-suffix guards the enforced count fell 403 → 29 on the same corpus. MEMPALACE_REDACT_STRICT=1 makes T3 enforce.
  • Operational shape. Fail closed — no redactor, no staging (exit 3), overridable with MEMPALACE_FEED_ALLOW_UNSCRUBBED=1. Every run prints a count including 0 redaction(s), because silence is indistinguishable from a scrubber that never ran. Findings carry rule, label, length and sha256[:8] — never the value.

Near-miss this image would have shipped, caught before tagging (f0bffd1). The image installs /usr/local/bin/mempalace-pi-session as a symlink into /opt/mempalace-toolkit/bin, and ${BASH_SOURCE[0]} reports the invoked path, not the target — so the sibling-module lookup resolved to /usr/local/bin, the redactor was absent, and fail-closed did as instructed: [FATAL] ... refusing to stage. Measured side by side, the symlinked invocation FATALed while the direct one scrubbed 40 findings. At the next bake that would have stopped every feeder tick on every device — a silent fleet-wide memory outage, worse than the leak the scrubber prevents. Fixed by chasing the symlink chain in portable shell (readlink -f avoided: GNU/newer-BSD only, and this script also runs directly on macOS hosts) with colon-separated fallback candidates. Verified via the symlink, the direct path, and a second-hop symlink. General lesson: fail-closed converts "module not found" into an outage, which makes the module lookup load-bearing infrastructure that must be tested through the invocation path the fleet actually uses — not the convenient one from a checkout.

The mailbox becomes explainable and mesh-safe (bfe9c5c, a92c75d, e917662, ecc2a9c)

  • Owed-set derivation joins on hlc, not seq (bfe9c5c). seq is a replica-local arrival counter — the same event is #7 in one database and #12 in another — so a second replica would let already-answered asks resurrect. hlc is immutable and replicated, fixed-width, so string comparison is causal comparison. A safe no-op on today's single replica (verified: the positive-control pair orders identically under both keys), correct once a mesh exists.
  • Delivered text now says it is queued (a92c75d). Delivery uses steer with no triggerTurn, and the poll fires on agent_settled, so nothing wakes the model — a delivered ask sits until a human starts the next turn. Measured case: a directed report sat unread for 2.5 hours. The note explains the agent is not ignoring the ask, it is not running.
  • MEMPALACE_MAILBOX_NOTIFY gains explicit =kitty / =osc777 modes (e917662). Terminal autodetection inside a container is not unreliable, it is blind: docker exec forwards neither KITTY_WINDOW_ID nor TERM_PROGRAM, and TMUX is unset because tmux runs on the host. Verified on a live process: TERM=xterm-256color and nothing else.
  • The terminal path through tmux is documented as UNVERIFIED (ecc2a9c). Test sequences written to the pty produced no notification on a remote client; tmux likely drops unknown OSC types without allow-passthrough, and multi-client routing (one ask pinging every attached client) is an open question.

Documentation (e1cc759, 982b001, d4d8bb6, d2764bf)

  • RFC 003, the coordination-log spec the code had been citing all along — it did not exist anywhere. 11 sections, retrospective against mempalace 3.8.0 logstream.py, incl. owed-set derivation, ten dogfooded landmines and seven open decisions. Non-obvious findings: event_append has no idempotency guard on the write path (verify-before-retry; the replication path is guarded), coordination tools are exempt from both palace locks by design, GET /logstream/events never existed in 3.8.0 (not proxy-blocked), and mempalace sync never touches the logstream — the log is permanent and unbounded.
  • docs/fleet-memory.md, operator-facing: five storage types, a decision tree, latency expectations (~2–5 min live session; next session while offline), broadcast exclusion by design, fan-out, and the search-before-answer / diary-at-session-end / verify-don't-retry habits.
  • docs/secret-hygiene.md, incl. the tier definitions, the measured FP data, stated false negatives, and the three server-side call sites (specified, not built — tier 2 only there, since the hub cannot see a client's env).
  • Phase 1 exposure record moved to the private fleet repo with a moved-note stub; retention direction for the unbounded log (logrotate-style: never rotate still-owed events, rotation invalidates held cursors, archive-verify-delete).

The other memory system finally gets explained — docs/observational-memory.md

pi-observational-memory has been baked for several releases and described in one line of the feature list ("the recall tool for session compaction"), which is enough to name it and not nearly enough to use it. New 272-line explainer with five diagrams, aimed at someone who has seen /om:status or a "compacted memory" block and wondered whether to leave any of it switched on.

Scoped to what this repo is authoritative for, because upstream already documents the mechanism well. /opt/pi-observational-memory/docs/ ships concepts.md, how-it-works.md and configuration.md, including a correct v3 lifecycle diagram — so the new document links those for depth and spends its own words on the four facts pi-devbox owns and can change: the pinned commit it bakes (v3.0.4 ce9fc98, the value in build-manifest.json), the packages[] entry that decides which copy loads, the Haiku-workers-vs-Opus-session split seeded into ~/.pi/agent/settings.json, and the devbox-pi-config volume that makes the ledger survive --force-recreate. Plus the confusion this image creates by shipping two things called memory: a section contrasting it with MemPalace, on the line observational memory keeps a session coherent, the palace keeps the fleet coherent.

Every stated number was read out of the live container or the baked tree rather than copied from release notes — including the correction that the dropper is gated on a successful same-turn reflection and not on a token threshold of its own, which is the one detail pi-extensions/SKILL.md still gets wrong.

Placement follows the audience split fleet-ops states for itself: reusable mechanism is not deployment data, so a "why is this in my container" document belongs in the repo that pins and wires the component, pointing upstream for depth. Linked twice from the README, because before this commit the README referenced docs/ zero times and the one file already there (mempalace-broker-design.md) was reachable only by listing the directory.

A README claim that v1.8.9 made false, and how it got there

§ Cross-machine agent coordination ended with "Nothing in this image polls the log on the agent's behalf." The mailbox shipped in v1.8.9, so that sentence has been wrong since aac4a1c. Replaced with the three knobs and their defaults (MEMPALACE_MAILBOX, MEMPALACE_MAILBOX_POLL_MS 300000, MEMPALACE_MAILBOX_RESURFACE_MS 3600000), the fact that owed-ness is derived rather than read off status, and the queued-into-the-next-turn delivery semantics measured on two devices.

The cause is the one v1.8.9 wrote a rule about: the behaviour arrived through the floating MEMPALACE_TOOLKIT_REF, so no diff in this repo ever touched the paragraph that made the claim. v1.8.9's rule ("name a floating-ref behaviour change in the CHANGELOG before tagging") fired for the CHANGELOG and nobody swept the README. The CHANGELOG records what changed; the README asserts what is true, and only the first is reviewed at release time. Extending the rule accordingly: grep the README for absolute claims — nothing, never, does not, only — about any component whose SHA moved.

The replacement is dated on purpose. It says it describes the bridge as baked in v1.8.9 (mempalace-toolkit 5b8d78f) and points at that repo's docs/rfc-003-coordination-log.md §7.11–§7.12 for the mechanism, because toolkit main is already ahead of the baked copy (a92c75d makes delivery say it is queued and ping the human who is not looking; e917662 and ecc2a9c refine that notify path) and none of it reaches a container until a base rebuild. Documenting those here would have swapped a stale-behind claim for a stale-ahead one — the same defect with the sign flipped.

Diagrams verified by rendering, not by parsing

Both comparison diagrams parsed clean and rendered with their meaning reversed: Mermaid laid the second declared subgraph out first, so "with observational memory" appeared before "without", and MemPalace before observational memory in the diagram whose entire job was that contrast. A third was legible only at 1280px. Rebuilt as declaration-ordered node chains, then re-rendered at mermaid@11 — the version pi-studio pins — in the baked headless browser and read back as an image. Recorded because it generalises: mermaid.parse() proves syntax and says nothing about layout, so a diagram is unverified until someone has looked at it.

… and rendering it in my browser was still not enough

Reported from a real viewer: several boxes had their bottom line of text sliced off. Reproduced and root-caused rather than nudged — Mermaid measures a node label with its own font metrics, computes the box, then renders the label as real HTML inside a <foreignObject>. Any host stylesheet that touches the line-height or font-size of that HTML makes the text taller than the box already committed to, and the overflow is clipped at the box edge. Error accumulates per line, so the loss always lands on the last line of the tallest labels — which is exactly what was reported.

Two fixes were tried and only the second works:

  • %%{init: {'flowchart': {'htmlLabels': false}}}%% — rejected, and verified ineffective rather than assumed so. The directive is honoured (label elements switch from 16 foreignObject to 7 tspan), and the clipping is identical, because the inflated font-size still inherits into SVG text.
  • A hard limit of two short lines per node, with the detail moved into the prose under each diagram. One- and two-line boxes have enough vertical slack to absorb the inflation; three- and four-line boxes do not. This is also better documentation — the old nodes were carrying paragraph-sized text.

The regression harness is now the interesting artefact: render every block with a deliberately inflated line-height: 1.7 !important on the label HTML, screenshot, and read it. Two survivors of the rewrite were caught only by that harness — a long unbreakable /opt/pi-observational-memory path silently wrapping to a third line, and a cylinder ([( )]) shape, whose curved bottom leaves less room than a rectangle for the same two lines.

§4 answers the question the document left hanging: what compaction does to your context

Asked directly and worth writing down: if the old conversation is folded away, is the session back to knowing nothing? No — and the specifics are all checkable against pi 0.84.3's own docs/compaction.md and the extension's source:

  • A verbatim tail survives, sized by a token budget rather than a message count. Pi walks back from the newest entry until keepRecentTokens [20000], and everything from that firstKeptEntryId onward is kept unchanged. Cut points land on turn boundaries, never mid-tool-call.
  • The system prompt and AGENTS.md are not in the compacted region at all — they are rebuilt from disk on every request, so compaction cannot lose them.
  • Nothing is deleted from disk. Compaction appends a compaction entry carrying the summary and the cut pointer; no session line is rewritten in place.
  • recall therefore still resolves ids whose sources left the context, because it reads the full branch via sessionManager.getBranch() and never consults the context window.
  • Repeated compaction does not summarise the summary. The text is always rendered from live observation/reflection records, so there is no generation-loss spiral; the projection is incremental against the last full-fold boundary and escalates to a true re-fold from the branch root at observationsPoolMaxTokens [20000].

And one correction to this repo's own earlier claim: "compaction calls no model" is a steady-state property, not an absolute. If the ledger is empty — compaction firing before the observer has ever run — the hook returns nothing and explicitly declines ownership (// Decline ownership so Pi's native summarizer preserves the pre-cut context.), and pi's own model-based summariser runs. The doc now says so, with the snippet.

A shipped doc bug: the ledger entry type was stated exactly backwards

§9 told readers the entries are custom_message and specifically not custom. It is the other way round, so the one grep the section existed to get right was the one it got wrong. Corrected against the live session file — 11 om.observations.recorded and 6 om.reflections.recorded entries, all "type":"custom", alongside "type":"custom_message" entries whose customType is mempalace-mailbox and mempalace-wakeup, which is precisely where the confusion came from: the mailbox uses the context-visible API, om's ledger uses the invisible one.

That is not a typo but a load-bearing distinction, and the fix turns it into a feature the doc now advertises: custom entries "do not participate in LLM context" (pi docs/session-format.md), so the ledger costs zero context until it is folded — now a row in the cost table.

Not covered by any of this

The opencode bridge is a separate write path the feeder hook never sees, and the server-side layer is unbuilt — so a secret typed straight into add_drawer, or staged by a non-pi client, still lands unscrubbed.


v1.8.9 — 2026-08-26

The coordination log gets a reader, and the release checklist's last gate stops accusing the wrong component.

The mailbox arrives — named here because nothing in this repo caused it

mempalace-toolkit main moves e70bef2 → 5b8d78f (exactly one commit, 281 insertions / 9 deletions across extensions/pi/mempalace.ts and extensions/pi/README.md), and that is what actually ships the auto-delivered logstream mailbox. No pi-devbox commit implements it. docker-publish.yml resolves MEMPALACE_TOOLKIT_REF=main to a concrete SHA at build time and folds that SHA into base_tag, so the mailbox would have landed in the next tagged image whether or not this section existed — which is precisely why it exists. That is the same shipped-undocumented shape as 553d865 in v1.8.7, and that one caused a cross-host misattribution: an agent on another machine reasoned about which image contained which behaviour from a CHANGELOG that never mentioned it. The rule this release adopts: if a floating ref will pull a behaviour change into the image, name it in the CHANGELOG before tagging, not after.

What the mailbox does, from the shipped code rather than from the design discussion:

  • The bridge was write-only. It stamped provenance on the way out and never read the log back, so a directed ask reached an agent only if that agent happened to run mempalace_event_list itself. The channel carried real cross-machine traffic from 2026-08-18 onward with zero readers — every delivery in that window happened because a human said "check your mailbox".
  • Doubly gated, exactly like the provenance stamper: inert unless both MEMPALACE_PI_DEVICE and MEMPALACE_REMOTE_URL are set. An unstamped client has no address to be reached at, so there is nothing for it to read.
  • On by default, opt out with MEMPALACE_MAILBOX=0. Deliberate: an opt-in fix for a nobody-remembers-to-do-it problem only relocates the forgetting. Tunables: MEMPALACE_MAILBOX_POLL_MS (min gap between mid-session polls, default 300000) and MEMPALACE_MAILBOX_RESURFACE_MS (re-announce a still-owed ask after, default 3600000).
  • Owed-ness is derived, never read off status. event_ack appends and never mutates, and status is written once, so a directed open keeps matching the mailbox query forever — answered or not. A candidate counts as answered only when one of this device's own events has a higher seq, joins via metadata.ack_of or a shared correlation_id, and carries a terminal status (applied, superseded, failed, blocked). claimed and ready are deliberately not terminal — that is how "taken, but not finished" keeps resurfacing.
  • * broadcasts are excluded from the owed set. to_agent: <me> also matches broadcasts per the tool contract, so without this a broadcast written with status="open" would make every machine believe it personally owed the same answer — and the code would contradict the skill that documents it.
  • The dedup map is in memory on purpose. A restart forgets, so an already-seen ask can resurface: visible noise a human corrects in one turn. The opposite failure — suppressing an unanswered ask — is silent and permanent. Do not "fix" the noise by persisting it.
  • Delivery queues, it never interrupts. A sections push at before_agent_start plus a second agent_settled handler behind the 300 s floor, using steer and not triggerTurn: agent_settled means idle, so nothing wakes a model on inbound fleet traffic.

Measured on v1.8.8 (which bakes e70bef2, i.e. no mailbox) immediately before this release: the wake-up mailbox query had to be run by hand, returned 3 directed asks with status="open", and the derivation above resolved all three as already answered — the third independent confirmation that the raw status filter never shrinks, and the first taken on a fresh container with no memory of having answered them.

--expected-image-version: two versions, two flags

scripts/recreate-sanity-check.sh --expected-version 1.8.8 reported ✗ pi version mismatch: expected 1.8.8, got 0.84.3 and exit 1 — a red on the final runtime gate of a release, accusing the image of being the wrong version, when the flag had only ever asserted pi --version. AGENTS.md step 4 spelled it --expected-version X.Y.Z inside a checklist where every other X.Y.Z is the pi-devbox tag; README.md got it right, so the two documents disagreed.

Not hypothetical, and not one reader's slip: the v1.8.8 release-readiness handoff from pi@emb-7kj4vr4g (evt_20260826T134919_a614ecfc2d4f) propagated --expected-version 1.8.8 twice, in its body and in metadata.cannot_check_here, while correctly calling step 4 "the runtime peer of the smoke gate, so it is not ceremonial". Two independent readers, one on another machine, converged on the wrong meaning. Left alone it puts a spurious red on every release, and the intuitive remedy — re-pull, re-recreate — is pure waste.

  • New --expected-image-version X.Y.Z asserts the pi-devbox release tag, read from release_tag in /etc/pi-devbox/build-manifest.json (the image's own build-time ground truth — no checkout, no network, no Docker socket). A leading v is optional on either side, so 1.8.9 and v1.8.9 both work.
  • Both flags now detect being handed the other one's value, and the test is exact rather than heuristic: the value is compared against the other quantity this image actually reports, so it can only fire when the mix-up is real. --expected-version 1.8.9 now says "is the pi-devbox IMAGE version, not the pi version — use --expected-image-version", and the reverse mix-up is caught the same way.
  • Neither flag is required any more. With none, the live pi --version is asserted against pi_version in the build manifest. That is not a tautology: pi resolves through PATH, and a stale install in the ~/.pi/npm-global volume can shadow the baked one — the same shadowing this script already guards against for npm:pi-atelier in packages[]. Verified by mutating the manifest to a different version, which made the new check fail as intended.
  • The header note it replaced was stale and load-bearing. It claimed pi "is resolved from latest at CI build time and is NOT pinned … cannot self-derive an expected version". Dockerfile.variant pins ARG PI_VERSION=0.84.3, and docker-publish.yml reads that ARG as its source of truth (refusing to build on a floating value, checking it is published on npm, warning when npm is ahead). The same withdrawn claim also sat in cli_utils's pi-devbox-sanity --help, the third place this confusion lived; fixed there too, in that repo.
  • Argument parsing hardened while in there: a flag whose value is missing — or is another flag — is now a usage error (exit 2) instead of silently consuming the next argument, and --help works.

All fourteen flag combinations were exercised by execution, including the two manifest-absent branches and the shadowing branch, which a healthy container cannot reach naturally — mutation-tested with a doctored manifest path so that each failure branch was observed firing rather than assumed present.

Component audit: no bumps, and that is the finding

Checked before tagging, since a base rebuild was already forced:

Component In v1.8.8 Upstream now Action
pi (npm) 0.84.3 (pinned) 0.84.3 is latest none
mempalace (PyPI) 3.8.0 (pinned) 3.8.0 none
pi-atelier v0.8.2 (pinned) v0.8.2 highest tag none
pi-toolkit, pi-extensions, pi-fork, pi-observational-memory, pi-studio floating identical to baked none
skillset snapshot 6eb20af 6eb20af none
mempalace-toolkit e70bef2 5b8d78f ships the mailbox

So the whole ~67-minute base rebuild this tag pays for is attributable to the toolkit SHA alone — base_tag folds it, and it moved. Every other floating ref resolved to the commit already baked (verified with git ls-remote per repo, not by reading a cached clone).

One claim in this audit came from a fork that had fabricated its findings — six plausible-looking toolkit commits with five nonexistent SHAs, a pi 0.84.4 that npm has never published, a pi-studio commit ls-remote says does not exist, and a compatibility floor of 0.8.2 where the code says 0.7.1. Every row above was therefore re-measured directly. Recorded because the failure mode is specific: none of it looked wrong, and git cat-file -e is what caught it.


v1.8.8 — 2026-08-26

The vendored mempalace skill snapshot stops being anonymous, and the container starts saying which copy of each skill it is actually reading.

Peer review (pi@emb-7kj4vr4g, logstream correlation skills-provenance-review, full text in drawer_pi-devbox_reviews_e43e766641c9ec85217bc6ce) found three blockers before this was tagged. All three were the same species: a record asserting something it had not verified. Every finding below was reproduced by execution here before being fixed.

  • The verification gate could print OK and exit 0 without verifying anything. git show <ref>:<path> | sha256sum hashes empty stdin when the ref does not resolve, yielding a real-looking sha256("") rather than an empty string — so the UNKNOWN branch in --check was dead code. Reproduced: a bogus ref reported MISMATCH (accusing the snapshot of lying when the true cause was an incomplete clone — and the operator's natural remedy for MISMATCH is to re-run the refresh, which rewrites provenance to silence the complaint); with a 0-byte snapshot against a 0-byte upstream file it printed OK: … exactly skillset@aaaaaaa and exited 0 for a ref that does not exist. The script already had the right idiom (sha_empty) and had applied it to blob_sha but not to at_ref. Now existence is proven with git cat-file -e before anything is hashed, at two levels (does the ref resolve; does the path exist at it) because those are different failures. This was the same defect class as the canary it replaces: a check that can succeed without checking. A second, unflagged instance of the identical pipeline shape was found in blob_sha and fixed too.
  • --check's exit codes conflated "stale" with "lying", so the release step failed in the case AGENTS.md step 2 explicitly calls legitimate. Now: 0 truthful (including stale-but-truthful, with a NOTICE), 1 a lying record only, 2 cannot determine (ref absent from this clone). AGENTS.md step 2 rewritten to state all three, since its promise that "the message distinguishes the two" was exactly what the branch was breaking.
  • The staleness NOTICE then asserted a direction it had never tested — the same defect one layer down, found by pi@emb-7kj4vr4g against the real state of its own host. The branch fired the notice on "recorded ≠ HEAD" and announced that HEAD was the newer side, so a clone that was merely behind was told "has moved to 82a8d3c; the snapshot describes the older 5fd0d5c" when 82a8d3c is 5fd0d5c's ancestor. Harmless to the verdict (rc stayed 0, nothing was mis-verified) but it points the operator at a refresh — a ~67-minute base rebuild — when the real remedy is git pull. It now tests ancestry with the merge-base --is-ancestor primitive the refresh path two sections above already used, and reports three distinct verdicts: stale (recorded is an ancestor — refresh), your clone is behind (HEAD is an ancestor — pull, do not refresh), diverged (neither). All three verified by execution; only the first was right before.
  • --help died with unknown option: --help. The strict argument loop that closed the silent-ignore hole never added a --help case, so the one script whose argument order was itself a landmine had an erroring discoverability path. It now prints its own header block.
  • VENDORED.md contradicted itself, in the release whose stated invariant is non-contradiction. Its hand-maintained "Snapshot provenance at last refresh" line named skillset 670f7f1 — seven commits behind the ARG, and the very commit that told agents to hand-stamp added_by, i.e. the withdrawn instruction this line of work exists to stop shipping — while its cp recipe still contradicted the "not cp" rule 20 lines above. The hand-maintained line is gone (nothing forced it to move when the ARGs did); 670f7f1 is kept only as a labelled cautionary example. The pi-extensions half was verified redundant (CI resolves PI_EXTENSIONS_REF via require_sha) before removal, rather than silently dropped.

Should-fixes from the same review, all reproduced: --check given the documented positional spelling (<root> --check) silently ran a refresh, because only $1 was parsed — both tools now parse all arguments and reject unknown ones; a refresh at a detached or older HEAD silently rewound ref and bytes, now refused unless the recorded ref is an ancestor (--force to override); upstream_dirty was computed and never used in check mode, now reported; pi-devbox-version --no-skills --json printed human text and broke jq; --help was a hardcoded sed -n '2,22p' range that this branch had already made stale; the skill fingerprint hashed SKILL.md alone, so a live skill dir differing only in a sibling file still reported "identical" — and pi-extensions already ships two files — so it is now a per-skill tree hash and the manifest field is renamed skillset_snapshot_tree_sha256 to say what it measures; and the --no-skills smoke assertion was negative-only, passing on a crashed binary, now anchored positively. mktemp+mv left written files at 0600 (a mv takes the temp file's mode) — CI was unaffected because the git index records 100644, but a local build from a dirty tree would have baked it; now chmod 0644 before the mv.

The skill fix ships outside this release, because it had to. The review also found that skillset 82a8d3c — the coordination protocol itself — told every machine on this fleet to skip the mailbox it introduced: it gated the mailbox on mempalace_mesh_peers, and a hub-and-spoke palace reports peers: [] precisely because every machine is a thin client of one replica. It also asserted that a directed open event "stays in their mailbox until" acked — false, because event_ack appends and status is written once, so an answered ask matches forever. The headline measurement behind that claim ("exactly 1 — the one that needed a reply") was of an event already acked half an hour earlier. Fixed in skillset 5fd0d5c, which derives owed-ness by joining on ack_of/correlation_id with a seq ordering test — without which one terminal reply suppresses every later ask on the same thread forever. Because the skillset is mounted live on every enrolled host, that correction was already deployed fleet-wide before this image was built; the vendored snapshot is resynced to it (c04cd15 → 5fd0d5c → 6eb20af) so the no-clone fallback does not ship the withdrawn rule. Canary re-verified bidirectionally against the new bytes.

6eb20af adds the limit of that ordering test, found when pi@emb-7kj4vr4g verified it rather than adopting it: seq is replica-local. It equals origin_seq today only because one replica authors events for all four machines, so a second replica could order the same pair differently and derive a different owed-set from the same log — use hlc (already on every event, total and causally consistent) once mesh_peers reports any peer. Documented as reasoning, not measurement, since a second replica cannot be stood up to test it. The part worth keeping is the asymmetry: local-seq skew makes an answered item resurface (noise, self-correcting, visible), while a timestamp comparison suppresses an unanswered ask forever (silent, permanent) — so anyone tempted to "fix" a resurfacing item with created_at would be trading the safe failure for the dangerous one.

Also carried, previously undocumented: dbb7879 resynced the vendored mempalace snapshot to skillset c04cd15 ("the withdrawal only holds where the bridge is live"), landed after the v1.8.7 tag and so absent from that image. ⚠️ A base rebuild is forced (~67 min): both that resync and the pi-devbox-version / entrypoint-user.sh changes below touch inputs to base_tag (rootfs/ and entrypoint*.sh). The provenance recording itself adds nothing to that cost — it lives entirely in Dockerfile.variant.

Both come from one finding, made while verifying v1.8.7 from inside a freshly recreated container: the baked mempalace snapshot is read by no host on this fleet. ~/.agents/skills/mempalace is a symlink to /workspace/skillset/skills/mempalace — entrypoint-user.sh links the baked skill only if [ ! -e ], and devbox-skill-reconcile then repoints the skillset-owned ones at the live clone (that is the v1.8.5 fix working as designed). All four compose stacks in docker-compose-repo mount a workspace containing the skillset, so the vendored copy is a CI/no-mount fallback and nothing else. Which means the mempalace skill snapshot is current canary — the assertion that blocked v1.8.7's first tag — polices a file that no agent on this fleet ever opens, while the drift that could actually mislead an agent (a git pull nobody ran in /workspace/skillset) was invisible from inside the container and is invisible to CI by construction.

The rejected fix is worth recording, because it was the obvious one. The old comment in scripts/smoke-test.sh said the real answer was "a CI job diffing this file against the skillset repo". It isn't:

Objection Detail
needs a credential CI does not have the skillset is private (ssh://git@gitea.jordbo.se:2222/joakimp/skillset.git); every build-time clone in this image uses anonymous HTTPS, and resolve-versions' gitea_sha() is explicitly documented as public-repo-only — its 401/403 path exists to survive a stale token against a public repo, so a private 403 would return empty and require_sha would hard-abort the release
makes another repo's branch able to fail this build the same pi-devbox commit would go green today and red tomorrow, and a release could be blocked by an edit in an unrelated repo — precisely the shape of the run 589 failure, but automated and permanent
pure churn, and it is measurable pi@emb-7kj4vr4g pushed four skillset commits in one evening (d9dbbbd, b740d51, 3324bd0, c04cd15); a byte-parity gate would have demanded a pi-devbox resync commit and a ~67-minute base rebuild for each one, to keep current a copy almost nobody resolves
guards the wrong artefact see above: on this fleet, nobody reads it

The invariant is not currency, it is non-contradiction — the framing comes from pi@emb-7kj4vr4g's review (logstream project/pi-devbox, correlation skillset-vendor-drift, which also retracted its own earlier build-time byte-compare recommendation). A stale-but-self-consistent fallback is harmless; a stale fallback carrying a withdrawn instruction is a live footgun, and this project has already paid for that one — through v1.8.4 the baked snapshot shadowed the live clone, which is how superseded attribution guidance kept reaching agents. That is precisely what the bidirectional canary asserts, and why it stays.

So provenance is recorded rather than policed, and the check moves to where the skillset actually is — a maintainer's clone, or any running container.

Added

  • build-manifest.json now records the vendored snapshot's provenance: skillset_snapshot_ref (which skillset commit the bytes are claimed to come from) and skillset_snapshot_sha256 (the bytes that actually shipped). The ref is a plain ARG default in Dockerfile.variant, deliberately not a CI-resolved output, which buys three things at once: it needs no credential for a private repo; it keeps a local docker build and CI identical by construction (the same reasoning that put MEMPALACE_VERSION in Dockerfile.base rather than duplicating it in the workflow); and it requires no change at any of the four Dockerfile.variant call sites (smoke, smoke-studio, build-variant, build-variant-studio), whose --build-arg lists are hand-duplicated and therefore easy to under-apply to only two. Also emitted as OCI label se.jordbo.pi-devbox.skillset-snapshot-ref, so it is readable off the registry without pulling the image.

    Two design points, each arrived at from the file's own rules:

    • The ref is a claim; the hash is measured. Dockerfile.variant writes the manifest from ground truth (rev() on each /opt clone, the live pi --version), so the snapshot hash is computed with sha256sum in that same layer rather than passed in. A build where the two disagree is exactly what the new smoke assertions catch.
    • They are siblings, not members of components{}. That map means "HEAD of a clone present in this image" and the skillset is not cloned here — calling it a component would be a lie a future reader would act on. It is also load-bearing mechanically: pi-devbox-version renders every components{} value with .value[0:12], which would truncate a 64-hex digest into something that looks like a short commit. Same reasoning as mempalace_version's existing comment.

    ⚠️ Costs no base rebuild. base_tag hashes Dockerfile.base + rootfs/

    • entrypoint*.sh + the mempalace-toolkit SHA; Dockerfile.variant is in none of it. scripts/check-base-hash.sh scans Dockerfile.base only (DF="Dockerfile.base", single hardcoded path), so a new *_REF ARG in the variant is invisible to that guard — correctly, since it changes nothing about the base's contents.
  • pi-devbox-version gained a skills: section reporting, per vendored skill, whether the live copy is baked or a live <repo> @ <sha> clone — and for mempalace, whether that live copy matches the baked fingerprint: (identical to baked snapshot), (baked snapshot <ref> + uncommitted edits) when the clone is at the recorded commit but the bytes differ, or (baked snapshot <ref> — live copy differs). Same live-vs-baked shape as the existing pi:/palace: drift annotations. This is the check CI cannot do and a container can, for free, since every host that matters already has the skillset mounted. The list iterates the baked tree rather than a hardcoded name list, so vendoring a fourth skill needs no edit here.

    entrypoint-user.sh calls it with the new --no-skills flag: the banner is printed FIRST, before the baked links exist and long before the skillset deploy and reconcile run last, so anything it said about skill sources would describe a state that is about to change. Wrong-but-plausible is worse than absent. (This is the one part of the change that touches rootfs/ and entrypoint-user.sh, so it does cost a base rebuild — already sunk, since dbb7879 refreshed the vendored snapshot.)

  • scripts/vendor-mempalace-skill.sh — refreshes the snapshot and rewrites the recorded ref together, because a cp without a matching ARG bump produces a manifest that confidently lies, which is worse than the anonymous snapshot it replaced. Refuses to record a ref when the upstream file has uncommitted modifications (no commit describes those bytes, so recording one would be a fabrication) — checked on that one file, not the whole tree, so unrelated work in progress in the skillset does not block a vendoring. --check answers "is the committed snapshot really skillset@<recorded ref>?" and separately reports staleness against the clone's HEAD.

    Counterfactual-tested rather than reasoned about, against throwaway clones: a tampered snapshot reports MISMATCH and STALE (rc 1); a ref rolled back to the previous skillset commit reports MISMATCH with content unchanged (rc 1) and a subsequent refresh fixes only the ref, leaving the bytes alone; unstaged and staged-but-uncommitted upstream edits are refused with distinct messages and the snapshot left byte-identical, i.e. the refusal is atomic.

    Hardened after review by pi@emb-7kj4vr4g, whose warning was that a resync script must "write the ref it ACTUALLY copied from, or the provenance field inherits the same class of bug the canary just had". The first draft copied the working tree and guarded it with git diff — which says nothing about an untracked file, and can be clean on a detached or behind checkout while HEAD names something else. The snapshot is now constructed from git show HEAD:<path>, so the recorded pair cannot be a lie by construction, and the untracked case is refused explicitly (tested: it was the one input the first draft would have silently recorded a false ref for). Both new scripts are bash -n clean and shellcheck -S error clean — the gate v1.8.7 added.

Fixed

  • Three stale in-repo markers, all the same failure class. Two "Unreleased" pointers — scripts/smoke-test.sh pointed the reader at "the Unreleased changelog note", and the v1.8.6 correction at "the Unreleased entry above"; that section became the ## v1.8.7 heading at release time and neither back-reference was updated. The third: scripts/smoke-test.sh's own coverage list still advertised "typst PDF engine for pandoc (Unreleased)", five releases after typst shipped in v1.4.0. Same class as the canary they sit next to: true when written, silently false at release, with nothing checking them. The smoke comment now describes the mechanism that actually shipped (and why the CI-diff idea it advertised was rejected); the changelog one names v1.8.7; the typst line names v1.4.0.

Not fixed, deliberately

  • CI still cannot tell you the vendored snapshot is behind skillset main. That needs a read-only deploy key for a private repo threaded into resolve-versions, to warn about a file no host on this fleet reads. Revisit when a no-skillset container becomes a real deployment (shipping the image outside the fleet, or a CI-only agent) — at which point the honest gate is a warning, matching the existing PI_VERSION/MEMPALACE_VERSION policy (concreteness → error, newer-release-exists → warning), never a build failure.
  • The phrase canary stays. It is orthogonal and free: it pins content where the new fields pin provenance, so it still catches a re-vendored snapshot whose ref was bumped correctly but whose bytes came from the wrong place — and, per the review above, asserting the absence of withdrawn guidance is the half of it that earns its keep. Its comment now states the limit instead of promising a fix.
  • v1.8.7's published image has no recorded ref, and that is expected: the field arrives here. Worth knowing when reading one, since the tag move ebd0de0 → f645e66 means the published v1.8.7 carries a pre-dbb7879 snapshot, i.e. its baked mempalace skill lacks c04cd15's "confirm the bridge actually stamps" caveat. Harmless — on v1.8.7 the bridge is live, so that caveat self-retires, and every enrolled host reads the live clone anyway. pi-devbox-version degrades quietly on such an image: no fingerprint, no annotation, verified against the real v1.8.7 manifest.

Documented

  • The fleet's cross-machine coordination, which was working and unwritten. The RFC 003 logstream has carried real work between hosts since 2026-08-18 — patch handoff, design review, a v1→v2 supersede — and no document in this repo or the toolkit said so. Written up in three places, split by what each is authoritative for:

    • README.md § Cross-machine agent coordination — what the container needs: MEMPALACE_REMOTE_URL selects the shared palace, and MEMPALACE_PI_DEVICE is what makes this machine reachable on the log, because when every host is a thin client of one palace the stamped agent name is the only thing distinguishing them. Set both or neither: a container without the device var can read the log but is addressable by nobody.
    • the skillset's mempalace skill (6eb20af, live on every host that mounts the skillset, no rebuild needed) — the norms: a mailbox query at wake-up, and the sender-declared ack contract, where a directed event with status="open" is owed a reply and a * broadcast owes nothing. The status filter earns its place by dropping broadcast noise — measured, an unfiltered mailbox returned 5 events, 4 of them finished broadcasts from eight days earlier — but that is all it does; it does not compute owed-ness, and the version of this entry that claimed otherwise is withdrawn above. Measured today, both machines: the raw filter returns 2 asks here and 1 there, every one already answered, while the derivation returns 0 for both. Dropping noise and deciding what is owed are two different jobs.
    • mempalace-toolkit extensions/pi/README.md (e70bef2) — the mechanism, including that the bridge is write-only today (it stamps events going out and never reads the log, so nothing in this image polls on the agent's behalf), and that live SSE push is a palace-deployment question: the server implements GET /logstream/stream, but a reverse proxy exposing only /mcp makes it unreachable — verified by 404s against the real endpoint.

    ⚠️ The snapshot was refreshed rather than left stale. The skill edits landed in the skillset (5fd0d5c, then 6eb20af), so SKILLSET_SNAPSHOT_REF was resynced to match and scripts/vendor-mempalace-skill.sh --check is a clean OK with no notice: the no-clone fallback carries the corrected protocol, not the withdrawn one. That mattered more than currency usually does, because the superseded copy contained an instruction — the mesh_peers gate — that actively told a reader to skip the feature. Refreshing remains a deliberate release-day decision rather than an automatic one: it costs a base rebuild, and skipping it is legitimate because every enrolled host reads its live clone. What is not legitimate is skipping it silently, which is what the new manifest fields and pi-devbox-version output make impossible — hence step 2 in AGENTS.md § Release-day checklist. In this release the refresh was free: rootfs/ was already changing, so the base rebuild was forced anyway.


v1.8.7 — 2026-08-25

Patch release, and the fastest turnaround in the series (~9 h after v1.8.6) for one reason: v1.8.6 shipped a container that cannot tell you which machine it is running on, and that anonymity produced a real misattribution the same evening — a session on tor-ms22 read another host's diary out of the shared palace, reported its verification as its own, and built a causal inference on top of the coincidence. The client-side half of the fix lives in mempalace-toolkit, which the image clones at build time, so it can only reach the fleet through a tag. The CI-hardening work that had accumulated since v1.8.6 rides along.

Gitea-hosted refs re-resolved immediately before tagging (2026-08-25T22:35Z): pi-toolkit 0e1369e6 and pi-extensions 20228878 unchanged since v1.8.6; mempalace-toolkit 0fe64c48 → 553d8657 (the provenance change below). CI re-resolves pi-fork / pi-observational-memory / pi-atelier / pi-studio at build time as usual. Base rebuild is forced twice over — Dockerfile.base changed (the MEMPALACE_VERSION audit) and base_tag deliberately folds in the mempalace-toolkit SHA ("otherwise a toolkit-only fix never lands") — so expect ~67 min, and note that either cause alone would have sufficed.

⚠️ The first tag of this version did not publish. Run 589 built the base fine, then both smoke jobs failed 81-passed/1-failed on a single assertion — mempalace skill snapshot is current, a canary pinning a phrase from the vendored skill. The phrase it pinned was the heading of the very instruction this release withdraws, so refreshing the snapshot without re-pinning the canary made it fire correctly on a healthy image. Every publish job was skipped, so nothing reached the registry and the version was never consumed; the tag was moved to include the fix below. Fixing the canary is what this release is for, in miniature: the gate was right and the expectation was stale.

Added

  • Palace writes now carry the device that made them, and diary entries say so in text. The container is host-anonymous by construction — hostname is a Docker hash, $DEVBOX_HOST_ALIAS is generic, the virtiofs source tag is generic, and two hosts in this fleet are both aarch64 — so nothing inside it distinguished tor-ms22 from EMB-7KJ4VR4G. In a local palace that costs nothing (one origin, so origin is a property of the whole store). In the shared palace it means every drawer and all 621 diary entries read as though written here, which is exactly how a v1.8.6 verification performed on EMB was reported as tor-ms22's own.

    Two halves, arriving by different routes:

    Half Where it lives How it gets into this image
    the writer — stamps <harness>@<device> on add_drawer/checkpoint/mine/event_append/artifact_put, and prefixes diary entries with HOST:<device>| mempalace-toolkit extensions/pi/mempalace.ts (553d8657) cloned in Dockerfile.base at MEMPALACE_TOOLKIT_REF, whose SHA is folded into base_tag
    the consumer skill — stops telling the agent to do it by hand, adds the read-side warning vendored rootfs/…/skills/mempalace/SKILL.md, refreshed from skillset 73c7c8e6 rootfs/* is hashed into base_tag too

    Three design points worth recording, because each was arrived at the hard way:

    • The stamp goes in the client, not the agent. RFC 001 §7.3.2 ranks "agent stamps it via a skill instruction" as the ❌ worst possible place, and the skill had carried exactly that instruction since 2026-08-23. It failed as predicted: the agent that wrote the instruction then filed its own provenance drawer without it. 199 rows reached the palace unresolvable. One execute() wrapper cannot forget.
    • The diary marker is in the entry TEXT on purpose. diary_write has no metadata parameter, but the deeper reason is that mempalace's search projects a fixed key set and diary_read returns content — metadata is invisible to the agent who will later read the entry, so no metadata-only fix, not even a server-authoritative one, would have prevented the misattribution. The marker is an AAAK field, so it is machine-parseable and the first thing a reader sees. The wake-up preamble now also names the device and warns that diary_read interleaves every machine's diary.
    • A solitary devbox stamps nothing. Gated on MEMPALACE_PI_DEVICE and MEMPALACE_REMOTE_URL — both set only when the palace is actually shared (RFC 001 R1). Unset either and behaviour is byte-identical to v1.8.6.

    Never injected into diary_write or kg_add: mempalace 3.8.0 hard-rejects undeclared arguments with JSON-RPC -32602 rather than dropping them (the behaviour changed since the RFC's 2026-08-09 note, now corrected), so a blanket injection would break those two calls instead of being ignored. The allowlist is per tool for that reason.

  • CI now shellchecks the repo's own shell scripts, not just workflow run: steps. .gitea/workflows/lint.yml's actionlint job already shellchecks every workflow step, but nothing had ever pointed shellcheck at entrypoint.sh, scripts/*.sh, or the extensionless tools under rootfs/usr/local/bin/ (pi-devbox-version, devbox-skill-reconcile, dot-watch, studio-expose). The gap is not hypothetical: a sibling repo (skillset's ci-release-watcher templates) shipped echo "$json" | python3 <<'EOF' ... json.load(sys.stdin) for two months without anyone noticing it silently returned nothing — with no script argument python reads its script from stdin, so the heredoc is stdin and the JSON load hits EOF. shellcheck flags exactly this at severity error (SC2259, "This redirection overrides piped input"); it had been available to catch it the whole time, just never run.

    New step in the actionlint job, Shellcheck + syntax-check repository scripts, runs shellcheck -S error plus bash -n over every shell file in the repo, discovered by *.sh union a shebang scan (neither alone suffices) so the extensionless rootfs/usr/local/bin/* tools are covered too. Measured before adding it: -S error is 0 findings across all 11 shell files today, so the gate is green on arrival with no cleanup. -S warning is not free (19× SC2088 tilde-in-quotes in scripts/recreate-sanity-check.sh, plus assorted SC2016, both intentional here) — a warning-level gate would train people to ignore it, so it stays error-only, same reasoning as the existing SHELLCHECK_OPTS exclusions on the actionlint step. File-count guard included: the step fails loudly if the shebang scan matches zero files, since a green check over an empty set is not a check.

  • MEMPALACE_VERSION now gets the same CI audit as PI_VERSION — closing the item v1.8.6 (and v1.8.5 before it) listed as "Still open". The pin was a literal string in Dockerfile.base with zero references anywhere in .gitea/workflows/docker-publish.yml, while PI_VERSION had ~20: a concreteness gate, a published-on-registry check, and a never-silently-adopt drift warning. resolve-versions now applies all of them to the palace pin, read from Dockerfile.base (not duplicated in the workflow, so a local docker build and CI install the same version by construction):

    Gate Behaviour
    not a concrete X.Y.Z error — no floating palace version, same policy as pi
    not published on PyPI error at resolve time, instead of a uv tool install failure mid-build
    yanked on PyPI error — an exact pin installs a yanked release silently under PEP 592, so mempalace==X would have shipped a withdrawn client to the whole fleet
    newer release exists warning naming what to audit before adopting (MCP tool-schema = the agent-facing contract; client/server skew against the central palace)

    Plus one smoke assertion, installed mempalace matches CI's audited pin, gated on a new EXPECTED_MEMPALACE_VERSION env threaded into both the smoke and smoke-studio jobs. It is not redundant with the existing manifest mempalace_version matches the installed core: that one compares two properties of a single image and therefore cannot notice that both are the wrong version. The failure mode this one covers is a variant built FROM a cached base whose MEMPALACE_VERSION pin was older — internally consistent, silently stale, invisible to every other assertion (the risk scripts/check-base-hash.sh exists to reduce but cannot eliminate).

    Mutation-tested rather than reasoned about, by extracting the shipped block out of the YAML and running it with a stubbed curl: 9 cases — latest / 3.8 / absent ARG refused; 404 and a registry echoing a different version refused; a yanked release refused with its reason; a newer release warning without failing; a transient PyPI outage not failing a build whose pin is already verified; happy path silent and emitting the job output. Then once more end-to-end against live PyPI with the real Dockerfile.base. This found a genuine defect in the first draft: the yank message inlined a jq program inside a $(...) inside a double-quoted string, where the escaping broke the filter (jq compile error) while the surrounding exit 1 still fired — a gate that looked correct and reported garbage. The reason is now hoisted into its own variable. The new smoke assertion was checked the same way, through the real run helper's sh -c quoting path: passes on 3.8.0, fails on 3.7.1 and on 3.8.01 (exact equality, not the substring match the pi assertion uses), and skips cleanly when the env is unset so a local smoke-test.sh run is unaffected.

    ⚠️ Costs a base rebuild on the next tag: Dockerfile.base is hashed wholesale into base_tag, and its now-false "Known gap, carried forward" comment had to be corrected in place (leaving a comment that says the audit does not exist would repeat the shipped-false-claim mistake corrected below). Expect ~67 min, as for v1.8.5/v1.8.6.

  • The SSH sidecar now defaults to connection multiplexing, without overriding anyone's explicit choice. ~/.ssh-local/config already forced ControlPath into the writable sidecar dir, but nothing supplied ControlMaster for targets coming from the user's own bind-mounted ~/.ssh/config. An entry that never mentioned it therefore opened a fresh TCP connection per ssh call — and an agent doing a dozen calls in a few minutes is exactly the traffic shape that trips fail2ban or a CGNAT flow-table cap. Observed 2026-08-25 on this fleet: ~12 connections to one host in 15 minutes, after which port 22 stopped answering while HTTPS to the same estate stayed healthy in 0.44 s (that asymmetry is the tell for rate-limiting rather than an outage).

    The fix is where the block sits, not what it says. ssh_config is first-value-wins, so position encodes intent, and the two settings need opposite treatment:

    Setting Position Meaning Why
    ControlPath before Include ~/.ssh/config override the user's value points at read-only ~/.ssh; it cannot work here, so it must lose
    ControlMaster auto + ControlPersist 10m after the Include default an explicit per-host ControlMaster no must keep winning; we only supply an opinion where the user expressed none

    Force what is broken, default what is merely absent. The first draft of this put both in the leading block, which would have silently overridden an explicit ControlMaster no — the counterfactual is in the test below.

    Verified with ssh -G (the resolved-config oracle) rather than by reading the man page, against a fixture with one host set to no, one silent, one set to auto: the explicit no resolves to controlmaster false and still gets the writable ControlPath, the silent host resolves to auto, and the same fixture under the rejected layout flips the no host to auto — so the test discriminates the position, not merely the presence of the block. Then end-to-end: the real script rendered in a sandbox HOME, block last, bash -n clean, shellcheck -S error clean (the gate added in v1.8.7).

    Measured effect on the author's own config (41 host aliases): 22 were silent about ControlMaster and gain auto + 10 m persist; 0 are overridden, since the fleet contains no explicit no. Worth noting how the one deliberate exception is written — proxmox002-vpn carries # No ControlMaster — VPN means direct route, no CGNAT flow cap, i.e. the intent is expressed as absence plus a comment, which ssh cannot distinguish from "no opinion". That host does now get multiplexing; its comment says multiplexing is unnecessary there, not harmful. Anything that must stay unmultiplexed needs a literal ControlMaster no.

    Why the ordering matters beyond this one config: ~/.ssh/config is per-machine, differs across the fleet, and future machines' versions do not exist yet to be audited. A default-not-override design is correct without needing to inspect any of them.

    ControlPersist is deliberately short (10 m idle, and each new session resets the idle timer — long enough to collapse an agent's burst, short enough that an abandoned socket ages out). A per-host entry that sets its own value keeps it: hosts already specifying ControlPersist 4h still resolve to 4 h. The known cost of multiplexing is the stale master — socket present, daemon gone, after a suspend or network change — which makes every later ssh to that host hang; recovery is ssh -F ~/.ssh-local/config -O exit <host>, now documented in the pi-devbox-environment skill along with -O check.

Fixed

  • The vendored-snapshot canary was one-way, and pinned a phrase the same release deleted. mempalace skill snapshot is current grepped for "Attribute what you file yourself" — the heading of the hand-stamping instruction withdrawn above. It therefore did its job (snapshot changed, expectation did not) and blocked an otherwise-green build. Two changes rather than a string bump: the assertion is now bidirectional (the new phrase must be present and the withdrawn one absent, so a re-vendored stale snapshot fails as loudly as a forgotten bump — verified by running it against v1.8.6's snapshot, which correctly fails), and the comment now states the structural limit: a phrase canary can only detect "older than what I remembered to pin", never "older than skillset main".

  • Four build-provenance smoke assertions verified the presence of a manifest field name and never looked at its value. The originals were literally:

    run_expect "manifest records pi_version" "cat …build-manifest.json" '"pi_version"'
    

    which passes on {"pi_version": ""} and on {"pi_version": null}. The tell was sitting in the passing output all along — ✅ manifest records pi_version (got "pi_version") echoes the key back as the thing it claims to have found — and it was spotted while reading run 579's smoke log to confirm the new v1.8.6 assertions had actually executed.

    Replaced with checks against the values, and against ground truth where ground truth exists:

    Assertion What it now enforces
    manifest declares every required component key all seven components present by name, failing with which key vanished
    manifest component values are resolved 40-hex commits each value is a full 40-hex SHA; null allowed for pi-studio alone (absent in the non-studio variant)
    manifest pi_version matches the installed pi manifest value equals pi --version, same ground-truth shape as the mempalace check
    manifest top-level fields are well-formed, not merely present release_tag non-empty; source_revision 40-hex when populated; build_date ISO-8601 when populated
    pi-devbox-version --json round-trips the manifest byte-for-byte actual string equality with the file, since --json is a verbatim cat

    Why five value-checks replace four name-checks (total assertion count unchanged at 61), and specifically why key presence and value shape are kept apart: an "every component value is a valid SHA" loop passes vacuously on components:{}, because jq's all() over an empty list is true. A single combined check would therefore go green on a manifest that had lost every component — which is the same shape of hole as the three false greens already recorded in this file. They are separate on purpose.

    Mutation-tested rather than reasoned about, twice: nine fabricated manifests through the raw jq filters, then twelve through the shipped assertions using the real run helper's sh -c quoting path (the quoting is load-bearing here — a jq filter that dies on a quoting error exits non-zero and looks like a caught defect). Measured against the old assertions on the same twelve defects: old caught 3, missed 9; new catches 12. The three the old set caught were key disappearance (grepping for a key name does fail when the key is gone) and the literal string "unknown"; every value-level defect — empty string, null, a 12-hex truncation, a wrong-but-plausible version, a malformed source_revision — was invisible. Three legitimate variations are correctly not flagged: empty source_revision and empty build_date (both default empty on a plain local docker build, so demanding them would fail honest local smoke runs) and pi-studio: null.

  • Dropped the now-redundant manifest has no unresolved ('unknown') components assertion. The 40-hex value check strictly subsumes it: "unknown" is not 40-hex, and only rev() in Dockerfile.variant ever emits that string, feeding components{} exclusively. Removed rather than left in place, because a redundant check that can never fail independently is one more green tick that means nothing.

  • Corrected a factually wrong "Still open" bullet in the v1.8.6 entry below (see the strikethrough there). It claimed pi-devbox-version's human output does not display mempalace_version and that only --json surfaces it. Both halves are false — v1.8.6 shipped a palace: line with the same live-vs-baked drift annotation pi already had. Verified by running the shipped script against fabricated manifests: matching versions print palace: 3.7.1, a skew prints palace: 3.7.1 (baked as 3.8.0 — drift detected), and a pre-v1.8.6 manifest with no baked field prints the live value un-annotated. The bullet appears to describe an intermediate state of the working tree and was never re-checked before tagging. Left visible as a struck-through correction rather than deleted, since v1.8.6 is already published and someone may have read it.


Still open

  • Make the vendored-snapshot check automatic instead of a remembered string. Tonight's failure is the third iteration of the same maintenance burden (v1.8.4: phrase present in both copies; v1.8.7: phrase deleted by the release that refreshed the snapshot). A phrase canary structurally cannot answer "is this snapshot older than skillset main?" — only a diff can. Proposed: a lint job that clones the skillset repo and compares rootfs/usr/local/share/pi-devbox/skills/mempalace/SKILL.md against it, failing with the diff when they drift. Open question first: the skillset repo is private, so this needs a CI clone credential, which is a policy decision rather than a code change.

  • Provenance stops at Chroma's metadata. The hourly reconciler on the palace host stamps device/agent_kind in chroma.sqlite3, but knowledge-graph triples and coordination events live in separate SQLite files (knowledge_graph.sqlite3, logstream.sqlite3) it cannot reach. 156 triples carry no origin field at all; logstream's from_agent is free-form and already inconsistent (pi@tor-ms22, pi@emb-7kj4vr4g, and bare pi in the same table). Tracked in RFC 001 §7.3.1.

  • The stamp is self-asserted, and cannot be otherwise yet. mempalace 3.8.0 authenticates with a single scalar bearer token and has zero device concept, so a verified stamp needs per-device credentials plus an origin field in six write paths across three databases. Deferred to RFC 001 Phase 4, where it is now motivated primarily by revocation (one shared token covers every device, so cutting off one laptop means rotating the fleet) rather than by provenance. Forward-compatible by design: every stamp records how it was determined, so an authoritative pass overwrites with device_source='token' and nothing has to be undone.

  • tor-ms22 and tor-ms22-native are one machine with two device values (4,680 and 3,826 rows). That is the hostname-as-identity cost RFC 001 §7.3.4 warned about, now visible in data: a rename splits one device's history silently. Repairing it means a device-identity mapping, not a relabel.

v1.8.6 — 2026-08-25

Patch release. Adopts the drift that accumulated in the ~2 days since v1.8.5 (pi 0.84.3, mempalace core 3.8.0), then closes the documentation and observability gaps that v1.8.5 itself listed as "Still open". No component was adopted without an audit note recording why it is safe.

All moving refs re-resolved immediately before tagging (2026-08-25T13:28Z): pi-toolkit 0e1369e6, pi-extensions 20228878, mempalace-toolkit 0fe64c48 and pi-observational-memory ce9fc982 all unchanged since v1.8.5; pi-fork f1ff8087 → bf702b4c; pi-atelier holds at v0.8.2 (floor for pi ≥0.84 satisfied); pi-studio's CI-resolved newest tag has moved again to v0.9.51. Base rebuild is forced (Dockerfile.base changed), so the 16 floating base-tooling ARGs re-roll — expect ~67 min as for v1.8.5.

Changed

  • mempalace core 3.7.1 → 3.8.0. Released 2026-08-23T21:19Z, hours after this project's own v1.8.5 tag the same day. Additive/reliability only — reviewed for MCP tool-schema changes before bumping, as always: none. sync --apply (PR #2320/#2322) no longer deletes a drawer solely because its source_file was unreachable at that moment — it asks for corroboration first. This does not relax the standing landmine against running mempalace_sync / mempalace_delete_by_source beyond dry-run on the shared central palace: that failure mode is paths permanently absent from whichever host runs the sync, not transient unavailability, and 3.8.0 doesn't touch it. Server-side perf fix PR #2307 (long-running Chroma servers no longer invalidate their own HNSW cache on their own writes) likewise does not make mempalace_reconnect unnecessary — that tool covers external writes bypassing the in-process client, a different scenario. Full reasoning lives in the Dockerfile.base comment above ARG MEMPALACE_VERSION. Deployment note: synlig's central palace currently serves 3.7.1 server-side via docker-compose.mempalace.yml (which reuses this image) — this client bump introduces version skew until that stack is separately redeployed; sequence accordingly.

  • pi 0.84.2 → 0.84.3. Published 2026-08-24T11:09Z. Release notes carry one "Breaking Changes" line — GoogleThinkingLevel renamed to GoogleApiThinkingLevel — checked against all four vendored packages (pi-fork, pi-observational-memory, pi-atelier, pi-studio): zero references, inert here. 0.84.3 also fixes two skill-discovery bugs that land directly on this repo's own vendored-skill work: nested Markdown skills inside .agents/skills/ grouping directories not being discovered, and root Markdown files (README.md/AGENTS.md) in skill directories being wrongly reported as broken skills.

Added

  • Browser automation is now documented to humans, not just to agents. agent-browser + Playwright + a headless Chromium (~625 MB — the single largest addition in the image) previously had zero mentions in README.md, DOCKER_HUB.md or THIRD_PARTY.md; it existed only in the agent-facing AGENTS.md managed block. Added a README.md "Browser automation" subsection, a DOCKER_HUB.md feature entry, and THIRD_PARTY.md license rows for agent-browser (Apache-2.0), Playwright (Apache-2.0), and Chromium (BSD-3-Clause for Chromium's own code plus a large set of bundled third-party components under their own licenses; the binary here is not compiled by this repo — it's Playwright's own "Chrome for Testing" download via playwright install --with-deps chromium).
  • THIRD_PARTY.md gains rows for pi-atelier (MIT) and mempalace core (MIT per the GitHub repo; noted that the PyPI package's own metadata omits a license classifier, so verify against the repo's LICENSE rather than sdist/wheel metadata if clearance is needed from the artifact alone).
  • typst and socat added to README.md's tooling inventory. Both were already used in prose (typst as pandoc's --pdf-engine, socat by studio-expose) but missing from the "What's inside" lists, so the inventory didn't match what the image actually ships.
  • mempalace core version recorded in /etc/pi-devbox/build-manifest.json. Previously absent — a published image couldn't answer "which palace version shipped?", and a palace bug couldn't be correlated to an image version. Derived from the live installed binary (matching the manifest's existing ground-truth-not-build-args philosophy), degrading to null rather than failing the build if the binary is missing or its output format changes. Verified landed: new top-level "mempalace_version" key, sibling to pi_version rather than a member of components{} (that map is rendered truncated to 12 chars by pi-devbox-version, which would mangle a longer version string).
  • New smoke assertions, all landed in scripts/smoke-test.sh: (1) the pi-observational-memory clone is checked for the actual ce9fc98 auth-fix markers pinned to their fix site, src/runtime.ts (availability_recheck, providerCredentialConfigured, hasConfiguredAuth) — not merely clone existence, and deliberately not a repo-wide grep: all three identifiers also appear under tests/, so a repo-wide search would stay green even with the fix reverted in src/runtime.ts alone; (2) the manifest's new mempalace_version field is asserted present, non-null, and equal to what mempalace --version reports live, so the manifest can't silently drift from the installed package — expected to fail against any pre-v1.8.6 image, by design; (3) a behavioural check for the mempalace-toolkit feeder's --agent default (see below — this one turned out to be possible after all).

Fixed

  • A false claim was being published to Docker Hub on every release. DOCKER_HUB.md advertised "neovim (LazyVim defaults)". Nothing in this repo installs LazyVim — the only nvim configuration is a 19-line sysinit.vim that sets termguicolors. update-description pushes this file verbatim (with {{PI_VERSION}} substituted) to the Hub description, so the error was public, not internal. Corrected to describe what's actually there.

Component audit for this release

Checked against upstream 2026-08-25 (two days after v1.8.5's own audit): mempalace core moved 3.7.1 → 3.8.0 (see Changed, above — timing is notable: released hours after v1.8.5 tagged, so v1.8.5 could not have caught it no matter how carefully it was audited). pi moved 0.84.2 → 0.84.3 (see Changed). pi-toolkit 0e1369e6, pi-extensions 20228878, pi-observational-memory ce9fc982, and pi-atelier v0.8.2 are all unchanged from v1.8.5 — in particular pi-observational-memory still sits exactly at the auth-fix commit with nothing landed upstream since, and pi-atelier is still the newest tag with the ≥0.7.1 floor for pi ≥ 0.84 trivially satisfied. pi-fork has one upstream commit not adopted this release: f1ff8087 → bf702b4c, a text-only rewording of the fork task preamble (no code-path change) — left un-pulled for this release since it is a moving ref CI resolves fresh at every build anyway; it will be adopted automatically on the next build regardless of this entry. pi-studio (studio variant) has drifted two tags upstream, v0.9.48 (pinned at build time via CI's newest-semver-tag resolution) → v0.9.51 at tag time, purely additive (watched PDF previews, opening PDFs directly in Studio, Studio header hide) — nothing to bump in this repo since studio-tag resolution happens in CI, not the Dockerfile, but note it will auto-adopt v0.9.51 on the next studio-variant build. mempalace-toolkit unchanged — this release's manifest and pi-bump work in Dockerfile.variant stayed within that file's ownership and did not require a toolkit-side change.

Still open

  • MEMPALACE_VERSION has no CI-side audit equivalent to PI_VERSION's. PI_VERSION is verified published-on-npm and warns (never silently adopts) on drift; MEMPALACE_VERSION is a literal Dockerfile string with zero references in .gitea/workflows/docker-publish.yml. Flagged in v1.8.5's audit as a gap; still a gap.
  • pi-devbox-version's human-readable output does not display mempalace_version. Its render path is a fixed sequence (release_tag, build_date, source_revision, pi, then components{}) and the new top-level field isn't in it — only --json mode (which cats the manifest directly) surfaces it today. One line in rootfs/usr/local/bin/pi-devbox-version would fix this; deferred since the field's stated purpose (correlating a palace bug to an image) is already served by --json, but worth doing in a follow-up if this becomes a routine manual check. CORRECTION (2026-08-25, post-tag): this bullet is wrong and was never true of the tagged tree. pi-devbox-version does print a palace: line in human mode, with live-vs-baked drift detection, degrading quietly on pre-v1.8.6 manifests. Nothing is open here. See the v1.8.7 entry above.

Resolved during this release, not left open: the feeder --agent default behavioural hook initially looked like it might need a mempalace-toolkit change (a --print-config flag that doesn't exist). It didn't — mempalace-pi-session assigns AGENT before argument parsing and --help exits 0 with no side effects, so bash -x mempalace-pi-session --help observes the real resolution (env interpolation and fallback) without needing a source change. The new smoke assertion exploits exactly that, checked both ways: with MEMPALACE_PI_DEVICE set it must resolve to pi@<device>; with it unset it must NOT be pi@* (catches a regression to the old unconditional $USER/mempalace default). mempalace-toolkit commit c64ffa1 changed the feeder's --agent default from $USER to pi@<device>, but there is still no way for smoke to assert this default is actually in effect from this repo alone, since mempalace-toolkit is a separate repo this release does not modify. If the concurrent smoke-test work could not find an honest assertion from the existing /opt/mempalace-toolkit surface (help text, --self-test), this remains open pending a toolkit-side --print-config-style hook — a toolkit-repo change, not a pi-devbox one.

  • 16 base-tooling ARG *_VERSION=latest pins remain unrecorded. (Corrected count — v1.8.5's entry said "~14"; the actual count from Dockerfile.base is 16, plus 5 more that float with no ARG at all: rustup-init, AWS CLI v2, Chromium-via-Playwright, Node's minor version via setup_22.x, and DEBIAN_VERSION=trixie-slim itself.) None of these are recorded anywhere once the build completes — not in the manifest, not in a label — so a published image cannot answer "which nvim/uv/chromium shipped?" without exec-ing in and asking the binary.

Documentation

  • .env.example documents MEMPALACE_PALACE_PATH. It was the only MemPalace variable the template never mentioned, while being the one that silently moves the feeders' stage: the palace root resolves as $MEMPALACE_PALACE_PATH → $MEMPAL_PALACE_PATH → ~/.mempalace/config.json → ~/.mempalace/palace, and the stage is derived from it (<palace-root>/pi-stage). The comment states the precedence, says why neither the image nor the entrypoint exports it (pinning the palace without carrying the stage re-creates the split a shared root removed — see v1.8.2), warns that a stage whose persistence differs from the palace makes a scoped mempalace sync prune conversation drawers whose dedup key is the staged path, and notes it is a container path unlike the host-side WORKSPACE_PATH/SSH_KEY_PATH above it. Found while auditing a live host whose .env sets the variable redundantly to the default.

v1.8.5 — 2026-08-23

Patch release with two fixes in the container's skill wiring — one behavioural, one a latent crash found while reviewing the first — plus the mempalace-toolkit change that makes palace writes carry provenance. No component pin moved. Every ref was re-resolved at tag time and is byte-identical to what v1.8.4 shipped: pi 0.84.2 (still npm latest), pi-atelier v0.8.2 → 159f34cf (newest tag; the ≥0.7.1 floor for pi ≥ 0.84 holds), pi-studio v0.9.48 → c3b83680, pi-fork f1ff8087, pi-observational-memory ce9fc982, pi-toolkit 0e1369e6, pi-extensions 20228878, MEMPALACE_VERSION 3.7.1 (still PyPI latest, and the version the central palace serves — no client/server skew). The single moving part is mempalace-toolkit fd8b15f5 → 0fe64c4.

Fixed

  • Vendored skills no longer silently shadow their live skillset counterparts. ~/.agents/skills was asymmetric: mempalace, pi-devbox-environment and pi-extensions resolved to the baked /usr/local/share/pi-devbox/skills/…, while every other skill resolved to the live /workspace/skillset/skills/…. Root cause was precedence-by-ordering in entrypoint-user.sh: the baked links are created early (line 65 in v1.8.4; the loop moved down as this fix added comments) — deliberately so, to close a smoke-test readiness race — with [ ! -e … ] so they are "created only when absent", and the skillset deploy runs last, where it classifies the existing links as foreign and leaves them alone. The comment at line 61 claimed the goal was "a same-named skillset skill … is never clobbered" — but with baked-first plus create-when-absent, the skillset skill was precisely the one that lost. Comment now describes the actual behaviour.

    Observed cost, on two hosts independently: an edit to skillset/skills/mempalace/SKILL.md (adding a drawer-attribution rule) was pushed and present in the live clone (md5 129bcc4752), yet both the EMB-7KJ4VR4G and tor-ms22 containers kept loading the baked copy (md5 5236024fef) with zero occurrences of the new rule. The tor-ms22 agent had to fetch the rule from the Gitea API to read it at all. Editing a skillset skill therefore appeared to work and silently did nothing until an image rebuild — for exactly the three skills most likely to be iterated on.

    The fix is not "the skillset always wins", because ownership is per-skill (rootfs/usr/local/share/pi-devbox/skills/VENDORED.md): pi-extensions' authoritative source is the package repo, copied over the snapshot at build time, and skillset carries a downstream copy that can lag — handing that one to the clone would regress the skill. So a new helper devbox-skill-reconcile runs immediately after the skillset deploy and repoints only the skills named in skills/skillset-owned.txt (today: mempalace). Precedence is now user override → live skillset clone (owned names only) → baked snapshot, with the early links untouched as the fallback, so the readiness race stays closed. It only ever replaces a symlink that points into the baked tree, so a real directory or a link pointing elsewhere is never disturbed. Verify with readlink -f ~/.agents/skills/mempalace, not by reading the entrypoint.

  • A latent boot-abort in the baked-link block, found while reviewing the fix above and fixed with it. [ ! -e "$link" ] is TRUE for a dangling symlink (-e follows the link), so once a link may point into /workspace/skillset — which the fix above makes possible — a vanished mount turns the guard into "create over a broken link", and plain ln -s then fails with File exists. Under the entrypoint's set -euo pipefail that aborts container start before exec "$@", with a cryptic ln error and no pi. Reachable on a docker restart or a host reboot under restart: unless-stopped (the writable layer survives and ~/.agents is not a volume on any host), though not on a compose up -d recreate. Now ln -sfn, which heals the broken link back to the baked fallback; the reconciler re-points it in the same boot if the clone is back. A comment at the call site records why the -f must stay.

  • README's skill-precedence documentation was wrong in the same way the entrypoint comment was: it claimed baked skills are "created only when absent so a same-named skillset skill … is never clobbered" and that "a mounted skillset always overrides them". Rewritten to state the real, per-skill precedence and to name skillset-owned.txt and devbox-skill-reconcile.

  • The smoke canary for a stale mempalace snapshot could not detect staleness. It grepped "Shared palace: multiple harnesses" — a phrase present in both the stale and the fresh copy, so it passed throughout the shadowing bug above. It now pins the newest section ("Attribute what you file yourself"), and VENDORED.md records that updating this string is part of refreshing the snapshot. Three further assertions close the gaps that let the bug ship: skill link targets are asserted (not merely test -L), the skillset-owned.txt list is asserted to contain mempalace and not pi-extensions, and the reconciler's replace path — which CI never exercises, since no smoke container mounts a skillset — is covered by fabricating a skillset and asserting all three outcomes (owned skill repointed, unowned skill left baked, user override untouched) — plus a second case that a mutation test proved necessary: with the reconciler's "is this link ours?" guard deleted, all three of those assertions still passed, so the discriminating case is an owned name whose link is a user override pointing outside the baked tree.

Changed

  • Vendored mempalace skill snapshot refreshed from skillset 936fed8 → 670f7f1 (md5 5236024fef → 129bcc4752), which adds the "Attribute what you file yourself" rule: hand-filed drawers should carry added_by="<harness>@<device>". Without this refresh the symlink fix above would only help hosts that mount skillset; a bare container would still ship the pre-attribution-rule skill.

  • Component audit for this release — no pin edits needed. Every component except pi/pi-atelier is pinned to a moving ref that CI resolves at build time, and each was checked against upstream on 2026-08-23: pi 0.84.2 (still npm latest, published 2026-08-14), pi-atelier v0.8.2 (newest tag; the ≥0.7.1 floor for pi ≥ 0.84 is satisfied), pi-fork f1ff8087, pi-observational-memory ce9fc982, pi-toolkit 0e1369e6, pi-extensions 20228878, pi-studio v0.9.48 → c3b83680 — all byte-identical to what v1.8.4 shipped. MEMPALACE_VERSION stays 3.7.1 (still PyPI latest, and the version the central palace serves, so no client/server skew). The one component that moved is mempalace-toolkit fd8b15f5 → 0fe64c4, which is this release's other payload: the feeder now defaults --agent to pi@$MEMPALACE_PI_DEVICE so palace writes carry provenance, with $USER still the fallback when the variable is unset (AGENT="${MEMPALACE_PI_DEVICE:+pi@${MEMPALACE_PI_DEVICE}}"), so un-enrolled hosts are unaffected. Nothing landed upstream after the pi-observational-memory merge ce9fc982, so the eight-week-bug fix in v1.8.4 is not destabilised. All of the above was re-resolved immediately before the tag and was unchanged — worth repeating for any future release, because six of nine components are moving refs that CI resolves at build time, so the build, not the Dockerfile, decides what ships.

    Two notes for whoever runs the build. This release changes entrypoint-user.sh, rootfs/** and the resolved toolkit SHA — all three feed the base-image hash — so expect a full multi-arch base rebuild (~95 min, as on v1.8.3/CI 562), not a fast variant-only publish. And that rebuild re-resolves the ~14 base-tooling ARG *_VERSION=latest pins; measured drift on 2026-08-23 was one patch (nvim v0.12.4 → v0.12.5), so the window is favourable, but it is not covered by version assertions.

  • pi-devbox-environment skill — new §2 subsection "A negative result is usually your own filter", plus ControlMaster masking in §3. This is baked (rootfs/usr/local/share/pi-devbox/skills/, symlinked to ~/.agents/skills/), so it is an image-behaviour change even though no package moved. Motivated by three false negatives an agent produced in a single session, each from its own filter rather than from the world: a | head -20 "proved" an SSH peer absent that was defined at line 454 of a ~500-line config; ssh mac 'docker ps' "proved" the host had no Docker, when the non-interactive SSH PATH simply lacks /usr/local/bin; and a grep 'ssh ' "proved" no ControlMaster was running, when master processes rename themselves to ssh: <controlpath> [mux]. The rule now stated: a positive result carries its own evidence, absence has to be earned. §3 additionally documents that a live master socket makes later commands authenticate not at all, so "it still works" proves nothing after editing a peer's authorized_keys — verify with -o ControlPath=none -o ControlMaster=no, or the breakage surfaces in a future session with no memory of the edit.

  • AGENTS.md: a stale CI claim corrected. It said "a tag push produces two runs, not one — lint.yml fires on every push (including tag refs)". That stopped being true when lint was scoped to branches: ['**'], which excludes tag refs by design; refs/tags/v1.8.4 produced run 571 (publish) and nothing else. The head_sha + workflow-path filter advice stays, because it costs nothing and any future v*-triggered workflow would reintroduce the ambiguity. Also adds a short "Verifying this repo's reality from inside a container" section, including the trap that this repo's docker-compose.yml is a template pinning :latest while a real host runs its own per-machine file — so recreating from the repo copy can silently move a host off :latest-studio.


Still open

  • build-manifest.json records the mempalace-toolkit SHA but not the mempalace core version, so a palace bug cannot be correlated with an image. mempalace --version prints it; adding it is a one-line change to the manifest RUN in Dockerfile.variant plus one smoke assertion, and is variant-only (no base rebuild cost).
  • Smoke asserts the pi-observational-memory clone exists but not that it contains the ambient-auth fix. npm still ships pre-fix 3.0.4, so an accidental switch from the /opt clone to an npm install would be a silent regression. Cheap guard: grep -rl availability_recheck must be ≥1.
  • The feeder's new pi@<device> default has no behavioural test hook (--dry-run never prints the agent; --self-test only covers the remote-mine response classifier). Cheapest available check is a source-shape grep for MEMPALACE_PI_DEVICE:+pi@.

v1.8.4 — 2026-08-22

Patch release, and the one that ends an eight-week bug: the baked pi-observational-memory finally records observations on a Bedrock host that uses ambient AWS credentials. The fix is ours, but it is no longer a patch — upstream merged it, so this release picks it up through the ordinary PI_OBSMEM_REF=master path with no local carry. Also bumps pi-atelier v0.8.1 → v0.8.2 (audited below) and bakes the todo extension's new edit action. pi stays 0.84.2 (still the npm latest, published 2026-08-14) and MEMPALACE_VERSION stays 3.7.1 (still the PyPI latest).

Fixed

  • om consolidation on request-time-signed providers — upstream, not patched (pi-observational-memory 37986b6 → ce9fc98). Under 37986b6, om's pre-flight gate treated "pi exposes no apiKey and no auth header" as unauthenticated and skipped every consolidation. On Bedrock with ambient AWS credentials that is the normal case — pi signs SigV4 at request time — so om recorded nothing for eight weeks with no error, no cost and no log line. Every pi-devbox image up to and including v1.8.3 has that behaviour.

    Two commits, both authored here and now upstream verbatim: 6f694e6 fixes the gate itself (it must not require a credential payload), and 699ccc7 adds the second half of pi's own rule — hasConfiguredAuth reads an availability snapshot that stays empty when the startup availability pass was skipped/aborted/failed, so on the otherwise-fatal path om now asks pi to re-check the credential live (refresh scoped to the provider, network-free), rate-limited 60 s per provider, bounded by a raced timeout, logged as resolve.availability_recheck. Filed as upstream issue #51, merged as PR #52 (ce9fc98, 2026-08-22T04:46:18Z), which also carries PR #49's env/baseUrl forwarding merged four minutes earlier; the maintainer resolved the textual conflict between them keeping both behaviours. Verified on ce9fc98 here: tsc --noEmit clean, vitest 257 tests / 27 files green.

    Note for anyone carrying the local workaround: the interim fix was a packages[] override in ~/.pi/agent/settings.json pointing pi at a patched clone outside the image. From this release on, delete the override — the baked /opt/pi-observational-memory has the fix. Confirm with /etc/pi-devbox/build-manifest.json → components.pi-observational-memory before removing it. The npm-published pi-observational-memory is still 3.0.4 and still broken; the image does not use npm for this component, so the release cadence there is irrelevant to us.

Changed

  • PI_ATELIER_REF / PI_ATELIER_VERSION v0.8.1 → v0.8.2, audited per the floor note above the ARG. Only one version sits between old and new and its changelog is two lines, both Workspace-Pulse-internal: inspection requests are now coalesced and serialized so short Turns avoid duplicate Git work and overlapping inspections cannot run concurrently, and live tool-driven Pulse updates are preserved while a fresh inspection is guaranteed at Turn end and retired sessions can no longer publish stale results. Nothing touches pi's private TUI renderer, which is the coupling that produced the 0.6.0/0.7.0-under-pi-0.84 startup hang, and pi is unchanged at 0.84.2, so this bump does not re-enter that risk class. Both the seam and the pin floor (never pair pi-atelier < 0.7.1 with pi >= 0.84) are unaffected.
  • pi-studio 65995fe (0.9.44) → v0.9.48 — 14 commits, four releases. Studio-side only (INSTALL_STUDIO=false by default, so this lands in the studio variant): open Studio in Muxy's browser, local PDF preview actions, previews survive Pandoc probe failures, legacy LaTeX styles tolerated in Pandoc previews, native dialogs replaced in embedded browsers, and file-copy import fixes with an explicit fallback. CI resolves the highest semver tag, not main, so this is v0.9.48 exactly.
  • pi-fork 4a09af4 → f1ff808 — one commit, "Add fork runtime awareness" (2026-08-19).
  • mempalace-toolkit b609cf5 → fd8b15f — two commits, docs only (backup/recovery + units; the convos-miner mtime correction finished). The fix(pi-session) false-success guard was already baked in v1.8.3 — checked by ancestry (git merge-base --is-ancestor 6e1f4f3 b609cf5), not by reading the log, because a commit's date does not tell you which side of a pin it fell on.
  • aws-cli 2.36.24 → 2.36.29 and the other *_VERSION=latest tools (bat/eza/fzf/gitleaks/nvim/micro/zoxide/yq/typst/tealdeer/agent-browser/ playwright/gosu/git-lfs/uv) refresh implicitly, as designed.
  • pi-toolkit unchanged (0e1369e, local main == baked).

Added

  • todo extension: an edit action (pi-extensions 98eb07b → 2022887). The tool is a verbatim vendored copy of pi's own examples/extensions/todo.ts, which offers list/add/toggle/clear and no way to change an item's text. On a long-lived list that forces either a "patch" item describing a different item, or clear-and-re-add of everything — both hit for real on 2026-08-17 while tracking a 17-item fleet plan, which ended up with #18 correcting #17. edit takes id + text and keeps the id and the done status; nextId is untouched. Id stability is the point, because ids are the only handle a palace snapshot of a plan can refer to.
  • Verified live in-container before committing, by repointing ~/.pi/agent/extensions/todo.ts at a working copy for one session (unknown id and missing text both error as intended; editing a completed item kept its id and its done state), then reverting the symlink to the image copy.

Notes

  • No pi-atelier change was needed for the todo action. Its tool_result hook only checks that details.todos is a well-shaped array and ignores the action string, so the new action flows through its normalizer and sidebar untouched. Worth knowing while reading agent transcripts: that hook replaces todo tool output with N/M done · see sidebar whenever the sidebar todo panel is visible (showSidebarTodos), so an agent sees only the counter and not the item text — upstream's list otherwise returns every item. That is a deliberate context saving, not a tool limitation.
  • The vendored copy now carries a numbered LOCAL DELTAS list in its header (the earlier ctx.mode !== "tui" → !ctx.hasUI API fix, and this action), so reconciling a future upstream version stays mechanical rather than archaeological.
  • A second om fix is NOT in this image and will not be. Upstream PR #24 ("advance coverage watermark when observer records nothing", head joakimp:fix/observer-empty-coverage-watermark b577b29) is open but design-rejected by the maintainer on 2026-07-03: an empty observer verdict is usually a technical failure, so advancing the watermark would leave a gap in the observed session, and in the genuinely-nothing-to-observe case the next observer simply gets more context. So the observer can still re-fire on a growing span after an empty verdict — that is upstream's intended behaviour, not an image defect. Do not "fix" it by rebasing that branch.

v1.8.3 — 2026-08-16

Patch release. Bumps mempalace to 3.7.1 and closes the gap that made the baked mempalace skill go stale for four commits. pi stays 0.84.2 (still the npm latest) and pi-atelier stays v0.8.1; every git-ref component (pi-toolkit, pi-extensions, pi-fork, pi-observational-memory, pi-studio, mempalace-toolkit) was checked against its upstream head and is unchanged.

  • MEMPALACE_VERSION 3.6.0 → 3.7.1. Verified against the 3.7.1 source rather than its changelog, because the risk is to palaces users cannot reconstruct: legacy drawers lack the new chunk_total completion marker and both decision sites trust them (if chunk_total is None: ... trust the match as before), so there is no mass re-mine; NORMALIZE_VERSION is 2 in both versions, so the "pre-v2 drawers are stale" gate does not fire either; chromadb<2,>=1.5.4 keeps the same major, so no index-format migration; there is no auto-migration (the source says "We do NOT auto-migrate" twice) and rebuild_index has exactly one call site, the explicit repair rebuild; the single new palace file (logstream.sqlite3) is created lazily on first logstream use. Downgrade stays possible — 3.6.0 has zero references to chunk_total and ignores it as unknown metadata.

    Two behaviour changes worth knowing, both turning a silent condition into a hard refusal: MEMPALACE_MCP_ALLOW_PEER_WRITER no longer works on local/chroma palaces (it is now gated on backend_requires_single_writer(), and _MULTI_PROCESS_WRITER_BACKENDS is {pgvector, qdrant}), and writer-lock setup failures now fail closed (refusing this mutating tool) instead of proceeding with a warning. Neither affects this image's normal MCP-server-plus-CLI-feeder pattern, which already serialised on the same mine_palace_*.lock under 3.6.0 — "process-lifetime single-writer ownership" in the upstream changelog describes tightened escape hatches, not a new lease.

    What 3.7.1 buys a shared central palace is the real motivation: the stale chromadb SharedSystemClient cache is now dropped on reconnect (under 3.6.0 a peer's writes could be overwritten by a stale in-memory HNSW segment, "index count going backwards"), the writer lease is released on SIGTERM/SIGHUP instead of leaking a lock naming a dead PID, and an interrupted mine is no longer permanently skipped as though complete.

    Upgrading a server requires restarting it — 3.7.1 refuses mutating tools when the served library drifts from what is installed, and mempalace_reconnect cannot clear that (it reopens the database but cannot reload Python modules). The fleet primary was upgraded and restarted before this image was tagged.

    Note: opencode-devbox still pins 3.6.0. The two images are meant to move in lockstep, so that pin diverges until opencode-devbox cuts its own release.

  • Vendored mempalace skill snapshot refreshed to skillset 936fed8 (was 63f3bf5). This is the gap worth naming: ~/.agents/skills/mempalace symlinks to the image-baked copy under /usr/local/share/pi-devbox/skills/, and entrypoint-user.sh creates that link first while the skillset deploy never clobbers an existing name — so in a devbox container the vendored snapshot always wins, and editing the skillset repo alone changes nothing a container reads. Two commits' worth of guidance had been invisible here: the multi-machine shared-palace section (device provenance in source_path, mined drawers carrying the mine date with UUIDv7 recovery, the naive-local vs UTC timestamp mismatch, agent_name not being device-scoped, single-writer/no-queue semantics) and the hand-crafted-provenance guard.

  • pi-global-AGENTS.append.md gains ### If the palace is central, it is shared — three rules: never run mempalace sync against a shared palace (it prunes drawers whose sources look missing, which on a central palace is most of the content, including other machines' — compounded by RFC-001 §7.2, since feeders stage inside the palace root); a client-side timeout is not a failure (single writer, one large mine blocks everyone, so mine timed out after 30000ms usually means the mine completed — verify before retrying or you file a duplicate); and the mempalace CLI is not remote-aware, so it always opens a local-disk palace and can silently disagree with the MCP tools.

  • mempalace-census is now on PATH. It shipped inside the image at /opt/mempalace-toolkit/bin/ but was never symlinked into /usr/local/bin like its three siblings, so RFC-002 Phase A censuses had to be invoked by absolute path. Added to the symlink set, the chmod +x set, and the build-time --help smoke chain.


v1.8.2 — 2026-08-16

Patch release. Ships the fix for a silent transcript-feed failure, plus the smoke assertion that stops it coming back. No image pins changed from v1.8.1 (pi 0.84.2, pi-atelier v0.8.1); what moves is the baked mempalace-toolkit ref and one new smoke check.

The bug this closes (found on the first boot of the v1.8.1 image, on EMB-7KJ4VR4G, 2026-08-15): the container-start catch-up rsynced seven pi session transcripts to the palace host correctly, then asked the server to mine /data/feed/<device> — the feeder's default MEMPALACE_PI_REMOTE_PATH, which assumes a containerized palace server. That fleet's primary runs natively (a systemd user unit + uv tool), so it only ever sees host paths and the mine died with source directory not found. rsync had already succeeded, so the inbox looked healthy.

It stayed invisible because of the second half: the feeder decided success with '"error"' in body. MCP answers a hard tool failure with HTTP 200 and a JSON-RPC result whose content[].text carries the tool's own JSON as an escaped string — the bytes are \"error\", so the substring could never match. ~/.pi/agent/mempalace-catchup.log printed Done. Wing 'wing_conversations' updated. directly beneath the error JSON and exited 0. A feeder whose only artifact claims success is worse than one that crashes: nothing in the container disagreed with it.

Shipped here:

  • mempalace-toolkit ≥ b609cf5 baked (CI resolves the ref at build time): classify() parses the MCP envelope instead of grepping it (JSON-RPC error, MCP isError, inner success=false/error), and separates "verified ok" from "unverified: no JSON tool payload" rather than assuming the good case. A preflight warning fires when the rsync destination and MEMPALACE_PI_REMOTE_PATH disagree — in preflight, so --dry-run and --prepare surface it too. Remote mode also stops previewing NEW/SKIP from the local palace, which had been reporting "6 already filed" about a palace it was not feeding; the tags are now [?] and the summary names who decides.
  • New smoke assertion — mempalace-pi-session --self-test run against the baked toolkit. It replays six recorded MCP responses (fixture 1 is the verbatim 2026-08-15 failure body) plus a regression guard asserting the old substring check is blind to it. A stale or reverted MEMPALACE_TOOLKIT_REF can therefore no longer ship a feeder that mines nothing while reporting success.
  • .env.example now spells out that MEMPALACE_PI_REMOTE_PATH is the path the server process can open — the container path for a dockerized server, identical to the ssh-target path for a native one — and that a mismatch fails quietly, with rsync succeeding and only the mine failing.

The --self-test assertion is deliberately bare (mempalace-pi-session --self-test, no HOME=… prefix). run() invokes docker run --entrypoint="" $IMAGE sh -c … and no Dockerfile sets USER or ENV HOME, so it executes with no HOME at all — the same condition that made v1.8.0's stage assertion unsatisfiable. The feeder is set -u with HOME-anchored defaults, so it used to die with HOME: unbound variable there; b609cf5 derives HOME from the passwd database (what python's expanduser() falls back to) instead. Keeping the call bare means smoke also proves the feeder runs in a bare container, rather than papering over it with an env prefix.


v1.8.1 — 2026-08-15

Patch release. Unblocks v1.8.0, which never shipped. Its smoke and smoke-studio jobs each failed exactly one assertion (67/68 and 70/71 passed), so build-variant and everything downstream skipped: no v1.8.0 tag reached Docker Hub and latest stayed on v1.7.0 from 2026-08-07. Image content is unchanged from what v1.8.0 intended — the pins here are identical (pi 0.84.2, pi-atelier v0.8.1).

The failing assertion was pi stage defaults next to the palace (not a cache dir), added three days earlier in 7c00dd6. It was a test bug, not a product regression. It asserted a literal path:

echo "$out" | grep -q "stage=/home/developer/.mempalace/pi-stage/"

but the run helper invokes docker run --rm --entrypoint="" $IMAGE sh -c …, and neither Dockerfile.base nor Dockerfile.variant sets USER or ENV HOME (the published base image config carries no HOME at all — HOME is normally set by entrypoint-user.sh, which --entrypoint="" deliberately skips). So the assertion ran as root with HOME=/root, mempalace-pi-session correctly resolved stage=/root/.mempalace/pi-stage/… (it is $HOME-relative by design: $MEMPALACE_PALACE_PATH → $MEMPAL_PALACE_PATH → ~/.mempalace/config.json → ~/.mempalace/palace), and the literal grep could never match under any circumstances. The tell was one line below it in the log: the sibling assertion pi stage follows MEMPALACE_PALACE_PATH passed, because it sets the variable explicitly and so never consults HOME. Default fails while explicit passes is the signature of a wrong HOME, not of broken staging.

Fixed by asserting the invariant that was actually meant — the stage sits beside the resolved palace, sharing its lifetime — which is user-independent:

case "$stage" in
  "stage=$HOME/.mempalace/pi-stage/"*) exit 0 ;;
  *) exit 1 ;;
esac

$HOME is expanded by the container's own shell, so this holds as root, as developer, or under any future user, while a cache-dir default — the regression the assertion exists to catch — still fails it (verified against all three cases plus a simulated MEMPALACE_PI_STAGE cache pin). A second assertion, pi stage is palace-adjacent for the developer user, now covers the deployment-specific path properly, by supplying HOME=/home/developer instead of assuming it.

Why it took a release to notice — and the smoke_only input

docker-publish.yml triggers on push: tags: v* only. 7c00dd6 was a push to main, so only lint.yml ran; v1.8.0 was the first tag afterwards and therefore the assertion's first execution ever. Any smoke assertion written outside a release was unvalidated until the next release consumed it — the worst possible moment to discover it.

New workflow_dispatch input smoke_only closes that: it probes/builds the base and runs both smoke jobs against HEAD, then stops before publishing anything. Implemented as if: inputs.smoke_only != 'true' on build-variant and build-variant-studio, deliberately without always() so the implicit "needs succeeded" gate survives and a red smoke still blocks a release; promote-base-latest and update-description already require build-variant success and so skip on their own. On a tag push inputs is unset and null != 'true' is true, so releases behave exactly as before. This release was validated with a smoke_only dispatch before the tag was cut.

Smoke failures now explain themselves

run discarded all output (>/dev/null 2>&1), so a red ❌ carried zero diagnostic weight — explaining this one-line failure took a CI-log dig plus a registry image-config inspection, when the container had already printed the answer and thrown it away. It now captures output and prints the last few lines under a failed assertion only. Assertions that want a diagnostic echo it to stderr (the stage checks now report the resolved stage and the HOME they saw), which stays invisible while they pass.


v1.8.0 — 2026-08-15

Minor release. Headline: pi sessions now feed MemPalace by themselves. The image already shipped mempalace-toolkit, but its pi feeder (mempalace-pi-session) was never symlinked onto PATH, so nothing ever mined pi's transcripts — the palace only ever contained what an agent remembered to file by hand. A container that gets recreated regularly has no other memory, so a missed wind-down was a permanently lost session.

Also here: pi 0.84.1 → 0.84.2 and pi-atelier v0.8.0 → v0.8.1, bumped together. The pi bump closes the Amazon Bedrock tool-argument poison pill that v1.6.4 recorded as unfixed upstream; the atelier bump is the matching companion, since both sides changed fullscreen input handling in the same fortnight. Audits for both are below.

Why event-driven and not a timer: there is nothing schedulable inside the container — PID 1 is bash -l, with no systemd and no cron — and anything installed would not survive recreate anyway. The triggers therefore live where the events already are: pi's own lifecycle, plus container start.

Added

  • mempalace-pi-session symlinked onto PATH (Dockerfile.base, alongside its mempalace-session / mempalace-docs siblings, with the same --help build-time check). entrypoint-user.sh also self-heals the symlink into ~/.local/bin (already ahead of /usr/local/bin on PATH, and writable by developer) so the feature works on images whose base predates this change.
  • Container-start catch-up feed (entrypoint-user.sh, backgrounded). pi's mempalace extension feeds the palace on session_shutdown and on a debounced agent_settled, but a hard kill (docker kill, OOM, host reboot) runs no handler at all; this is the only trigger that can recover the previous life's transcripts. Skipped when a remote palace is configured without an inbox to ship to — and the skip now says so (see Changed) — and skippable entirely with MEMPALACE_FEED=0.
  • MEMPALACE_PI_STAGE no longer needs pinning here — the feeder's default was fixed upstream instead. It used to stage under ~/.cache, which is disposable in a container; the first cut of this change pinned the env var into the persisted ~/.pi volume. That was the wrong fix: it created a second convention that could still diverge from the palace (keep the palace volume, drop devbox-pi-config, and a scoped mempalace sync prunes every conversation drawer, because dedup keys on the staged path). The feeder now defaults to <palace-root>/pi-stage, resolved with mempalace's own precedence ($MEMPALACE_PALACE_PATH → $MEMPAL_PALACE_PATH → ~/.mempalace/config.json → ~/.mempalace/palace), so the stage inherits whatever persistence the palace has and the two cannot be separated by accident. No ENV and no entrypoint export: adding one back would re-introduce exactly the split it removes.
  • Transcript inbox mount in docker-compose.mempalace.yml (${MEMPALACE_FEED_DIR:-./feed}:/data/feed:ro). A client cannot mine into a remote palace directly: mempalace_mine expands its source path in the server process, so the server can only see paths inside its own container. Clients rsync their staged exports to a per-device subdirectory and then ask the server to mine /data/feed/<device>. Read-only because mining only reads sources — all locks live palace-side.
  • Smoke tests for the above: mempalace-pi-session on PATH, two assertions that the stage resolves next to the palace (default, and following $MEMPALACE_PALACE_PATH), and two behavioural guards that feed the exporter a synthetic pi session — one that must be captured, one abandoned session that must not be. The second matters because pi expands skills/context into the user prompt, so an abandoned session can look substantial by byte count while containing no assistant output; and if pi's JSONL shape ever changes, the exporter would silently capture nothing.
  • .env.example: documents MEMPALACE_FEED, MEMPALACE_FEED_DEBOUNCE_MS, MEMPALACE_FEED_WING, and the remote-palace shipping vars MEMPALACE_PI_SSH_TARGET, MEMPALACE_PI_REMOTE_PATH, MEMPALACE_PI_DEVICE.

Changed

  • The "remote palace, no inbox" skip announces itself instead of vanishing (entrypoint-user.sh). When MEMPALACE_REMOTE_URL is set but MEMPALACE_PI_SSH_TARGET is not, there is genuinely nothing the feeder can ship to, so skipping is correct — but the branch was a bare :, and the skip happens before the subshell that writes ~/.pi/agent/mempalace-catchup.log. A container in that state therefore contributed nothing to the palace and left no artifact at all, not even an empty log, to explain why — indistinguish- able from a healthy run that had nothing to file. Found while flipping the first client onto the shared palace (2026-08-12), where it is the single most likely way to end up quietly memory-less. The notice now goes to both the container start output and that log path, names the two variables that fix it, states that MCP tools still work (only this container's transcripts go nowhere), and points at MEMPALACE_FEED=0 for anyone who meant it. Deliberately incapable of breaking startup: an unwritable ~/.pi — root-owned volume, a classic Docker accident — would make mkdir -p fail under set -e and abort the whole entrypoint, so it degrades to stdout-only. That was a real new risk, since this branch previously touched no filesystem whatsoever. Covered by two smoke assertions against the entrypoint as shipped in the image (the branch only runs at container start, so a docker run one-shot cannot reach it).

Bumped: pi 0.84.1 → 0.84.2

  • ARG PI_VERSION=0.84.2 (Dockerfile.variant), with the audit the pin policy in that file requires.

    Headline for this image: the Bedrock tool-argument poison pill is FIXED upstream. The v1.6.4 entry below recorded it as "Not fixed upstream … still replayed unsanitised" — that note is now superseded. pi-ai 0.84.2 adds a recursive sanitizeBedrockDocument() and applies it at exactly the site that entry named (#7882):

    - toolUse: { toolUseId: c.id, name: c.name, input: c.arguments },
    + toolUse: { toolUseId: c.id, name: c.name, input: sanitizeBedrockDocument(c.arguments) },
    

    (dist/api/bedrock-converse-stream.js — line 692 in pi-ai 0.84.1, 704 in 0.84.2; it was 644 in 0.83.0 and 634 in 0.82.1.) The sanitiser drops object members whose key is the empty string, recursing through arrays and nested objects and preserving every valid value. It runs while the request is built, so it covers the live turn and a resume: a session already bricked by an empty-key tool argument now replays instead of dying on a Bedrock ValidationException. pi-session-repair (in cli_utils) is therefore no longer the recovery path on this image. It stays useful for older images and for inspecting a transcript, because the stored .jsonl is still malformed — the fix sanitises what is sent, not what was recorded.

    Why bumping PI_VERSION is the only way to get it: pi publishes an npm-shrinkwrap.json, which pins transitive dependencies exactly. pi 0.84.1's shrinkwrap pins @earendil-works/pi-ai to 0.84.1, so although 0.84.1's package.json range is ^0.84.1 — which would otherwise admit 0.84.2 — rebuilding the old pin can never pick the fix up. Transitive upstream fixes do not leak into this image; PI_VERSION is the whole gate.

    Rest of the audit, against the integration surface the pin policy names:

    • Session .jsonl format — unchanged. Identical migrateV1ToV2/migrateV2ToV3 ladder in both versions, so existing sessions on the named volume load as-is and pi-session-repair's parse target is untouched.
    • Node engine floor — unchanged at >=22.19.0 (image ships 22.23.2).
    • pi-atelier — no change needed. The pin stays v0.8.0: the hard floor is "never pair < 0.7.1 with pi >= 0.84", this bump does not leave 0.84.x, and atelier's peerDependencies (>=0.80.7) are satisfied. pi-atelier 0.8.1 is published but deliberately NOT adopted here — one variable at a time, and atelier is the component that has drawn blood at startup.
    • Directly relevant to pi --ssh use of this image: 0.84.2 fixes split Alt+Enter over SSH being misread as Escape, and adds PI_TUI_ESC_TIMEOUT for high-latency terminals.
    • Keybindings — one surface worth knowing. pi-toolkit ships exactly one override, tui.input.newLine: [shift+enter, ctrl+j, alt+j]. 0.84.2's new fullscreen transcript search (Ctrl+Shift+F) binds Shift+Enter to previous match while its overlay is focused. Different context, so no conflict is expected — but it is the one place the override meets a new default, and the first place to look if "shift+enter stopped inserting a newline" is ever reported.
    • New defaultTools setting (choose startup built-in tools globally or per project) is additive; pi-toolkit's settings.example.json does not set it, so the bootstrap template needs no change.

Bumped: pi-atelier v0.8.0 → v0.8.1

  • ARG PI_ATELIER_REF / ARG PI_ATELIER_VERSION = v0.8.1 (Dockerfile.variant), bumped together with PI_VERSION as that pin's comment requires — and this pairing is a good advert for the rule, because both sides touched fullscreen input handling within three days of each other.

    atelier 0.8.1 (2026-08-12) is two changes, only one of them code: "Preserve fullscreen transcript mouse-wheel scrolling after Sidebar resize and visibility changes by leaving Pi's persistent mouse reporting enabled", plus a README simplification. The single source file that differs from 0.8.0 is src/split-pane.ts. It extracts an isPiFullscreenRenderer() predicate and, under pi's fullscreen renderer, stops writing its own \e[?1002h\e[?1006h / \e[?1006l\e[?1002l pair around a sidebar resize — previously it enabled mouse reporting on grab and disabled it on release, which tore down the reporting pi itself had switched on and left the wheel dead afterwards. Outside fullscreen it manages mouse mode exactly as before. It also now captures the terminal it enabled mouse on and writes the disable sequence to that terminal instead of to whatever tui currently points at.

    The audit that matters is the private-internals coupling, since that is what hung startup at 0.6.0/0.7.0. atelier reaches into three pi internals; all three are unchanged in pi 0.84.2:

    • TuiAltScreen — detected by constructor name, so a rename would silently disable both the resize-input prioritisation and the new mouse behaviour, with no error. Still class TuiAltScreen extends TuiBase implements ViewportTUI.
    • tui.inputListeners — a private Set that atelier deletes from and re-adds to, to get its resize handler ahead of pi's viewport listener (which "consumes every mouse event for text selection"). Still inputListeners = new Set(), at the identical line 103 of pi-tui/dist/tui.js in both versions, and still a Set — atelier guards with instanceof Set.
    • the prototype render descriptor it wraps via findPrototypeRender. Still an own render(width) on TuiAltScreen.

    pi's mouse sequences are byte-identical between 0.84.1 and 0.84.2 (same 1002h/1006h/1002l/1006l/1003h occurrence counts), so atelier's assumption about what pi leaves enabled still holds. pi-tui's base class changed additively only (one new isOverlayFocused()), and TuiAltScreen's own changes are the new search feature (activeSearch, openSearch/closeSearch, the two search match styles, copySelection).

    Caveat, stated plainly: pi 0.84.2 adds a focused fullscreen search overlay that participates in input handling, while atelier reorders input listeners around pi's viewport listener. The two look convergent — 0.84.2 separately fixes "focused fullscreen overlays not receiving mouse wheel or viewport scroll keys" — but this pairing is reasoned from the diffs, not proven by execution: the CI smoke test does not drive the TUI, so a fullscreen interaction regression would not be caught before pull. Worth an alt+a plus a sidebar resize and a wheel scroll in fullscreen on first use of this image.

    Version metadata is unchanged: engines.node >=22.19.0, peerDependencies still the uninformative >=0.80.7 on both pi packages (so still nothing in npm metadata encodes the real floor), and still zero runtime dependencies — so the "no npm install step" note above stays true. The GitHub tag v0.8.1 exists (commit c31d7439), which is what CI resolves to a SHA.

Notes

  • The Dockerfile.base change moves the base hash, so this needs a base rebuild; the ~/.local/bin self-heal exists so the feature does not have to wait for one. The skip-notice change is in entrypoint-user.sh, which is COPYd in Dockerfile.base too, so it rides the same rebuild — until then, older images keep skipping silently and the two commands in the toolkit's phase-1-exposure-runbook.md §3.7 are the way to tell.
  • Requires the matching mempalace-toolkit change (--prepare two-phase split, remote transport, and the auto-feed triggers in extensions/pi/mempalace.ts). The split exists because the palace is single-writer: a live pi session holds it through the extension's own mempalace-mcp, so a CLI mempalace mine during a session fails with "palace ... is held by PID". Staging is therefore done by the CLI and the mine itself by whichever process already holds the palace.

v1.7.0 — 2026-08-07

Minor release. Headline: pi-atelier is now part of the image — the TUI sidebar/status rail every container previously had to hand-install — and pi is pinned to an audited version instead of tracking npm latest.

Why minor and not patch: the policy above reserves patch for "pi version bumps, smaller fixes" and minor for "new variants, significant base additions". Bundling a new companion package into every image is the same shape as v1.1.0, which went minor for bundling pi-studio; v1.4.0 likewise went minor for adding typst. This release also adds a new build-arg pair, a new opt-out env var, and a settings migration, so patch would understate it.

Added

  • pi-atelier vendored at /opt/pi-atelier, pinned to v0.8.0 — the TUI sidebar (ordered panels, split-pane, themes) is now part of the image instead of something each user hand-installs. Vendored + registered at container start by entrypoint-user.sh, the same pattern as pi-fork/pi-observational-memory/ pi-studio, and deliberately not pi install npm:pi-atelier: an npm install writes into ~/.pi/npm-global on the config volume, which shadows the image and pins nothing — the footgun that once hid a missing fork tool for six weeks. Unlike its siblings it gets no npm install: pi-atelier declares zero runtime dependencies (only peerDeps, satisfied by the baked pi) and has no build step, so pi loads its TypeScript straight from the checkout (pi.extensions → extensions/index.ts).
  • A version FLOOR, encoded as an executable test. pi-atelier 0.6.0/0.7.0 wrap pi's private TUI renderer in a way that recurses under pi 0.84: pi hangs at startup with sustained CPU and no error message. Upstream fixed the recursion in 0.7.1 and restored the non-overlapping split in 0.7.2 ("avoiding the recursive render path that caused startup hangs and sustained CPU usage"); 0.8.0 is additive on top of that. atelier's own peerDependencies still say >=0.80.7, which does not express the floor, so nothing in npm metadata could have warned us. smoke-test.sh and recreate-sanity-check.sh now assert the pairing rule pi ≥ 0.84 ⇒ pi-atelier ≥ 0.7.1 — verified against a 4×4 version matrix — so a bad combination fails the build instead of publishing an image whose TUI never starts. CI resolves the pinned tag to its peeled commit SHA; atelier uses annotated tags, so the unpeeled ref is a tag object, not a commit (pi-studio's lightweight tags never exposed that distinction).
  • DEVBOX_ATELIER=0 opts out: the entrypoint removes pi-atelier from pi's packages[] instead of registering it. The switch lives in the entrypoint rather than being "just run pi uninstall" because this component's failure mode is pi will not start, which cannot be repaired from inside pi.
  • Migration for hand-installed copies. A pre-existing npm:pi-atelier entry is dropped from packages[] (with a settings.json.bak.atelier.<ts> backup) so the pinned /opt copy takes over. This is not cosmetic: the registration guard counts npm:<name> as already-registered, so without this step every existing volume would have kept its unpinned npm copy — and a 0.6.x copy alongside pi 0.84 is exactly the startup hang above. Only that one exact string is removed; jq-parse failures or a missing file leave settings untouched, and the backup prefix is distinct from the template merge's so two rewrites in the same second cannot overwrite each other's backup.

Changed

  • pi-toolkit's pi-atelier.json modernised to atelier's current schema (pi-toolkit 0e1369e, cross-repo — it reaches the image through the pinned PI_TOOLKIT_REF clone). The seeded config had been written against the pre-0.7 vocabulary: segments → segmentLayout with explicit per-segment visibility, ornament: "none" → {"id":"brand","visible":false}, showExtensionStatuses → {"id":"statuses","visible":true}, plus the sidebar toggles that did not exist when it was written (showSidebarAgent, showSidebarTodos, and showSidebarOnStartup, new in atelier 0.8.0). Upstream still reads the old keys, but only as non-authoritative legacy inputs, so the file worked while silently missing every sidebar control added since. Verified by loading the old and new file through pi-atelier 0.8.0's own loadConfig(): zero warnings from each and an identical effective config, so it is a pure schema modernisation — every deliberate choice (compact density, 60/85 context thresholds, notifications off) is preserved. sidebarPanelLayout is left unset on purpose so the panel set tracks upstream as atelier adds panels.

  • pi is now PINNED, not latest: PI_VERSION=0.84.1 (Dockerfile.variant). CI's resolve-versions job used to resolve @earendil-works/pi-coding-agent to npm latest, which meant every release silently adopted whatever pi had shipped that morning — unaudited — in the same build that then got tagged and published. A pi minor can move the private TUI/renderer internals pi-atelier wraps (0.84 vs atelier 0.6.0: startup hang) or the session .jsonl format pi-session-repair parses. The pin is a checkpoint, not a freeze — bumping stays a routine one-line change; what stops is unreviewed adoption. 0.84.1 was audited for this release: theme/TUI additions are additive, the session format is unchanged (CURRENT_SESSION_VERSION = 3 in both 0.83.0 and 0.84.1, identical migrateV1ToV2/migrateV2ToV3 ladder, so existing transcripts are neither migrated nor at risk), and the Node engine floor is unmoved at >=22.19.0.

    • The pins live in the Dockerfiles and CI reads them from there (a checkout was added to resolve-versions), so a local docker build and a CI release ship the same versions by construction instead of by convention.
    • CI fails the build when the pin is not a concrete version, and when the pinned version is not actually published on npm — catching a typo, an unpublished version, or one yanked after we audited it, at resolve time with a clear message rather than as an npm install error mid-build.
    • CI warns (::warning::, never adopts) when npm latest is ahead of the pin, naming the newer version and what to re-check. That warning is the prompt to audit and bump — not something to silence.
  • mempalace pin 3.5.0 → 3.6.0 (Dockerfile.base MEMPALACE_VERSION), in lockstep with opencode-devbox v2.9.0 as the pin's own comment requires. 3.6.0 (2026-07-17) is PyPI latest and is additive/reliability only — secure mempalace serve remote mode, optional Milvus backend, atomic KG supersede(), conversation chronology, mining exclusions, plus recovery and locking fixes. Reviewed for MCP tool-schema changes before bumping — that being the exact regression class this pin exists to catch, after an unpinned install once swept in the broken 3.3.x/3.4.0 diary_write schema: there are none, and nothing touches diary_write, so the perl workaround removed in v1.2.2 stays removed. Two fixes are directly relevant to how this image uses mempalace: read-only mode now covers checkpoint + delete_by_source in _MUTATING_TOOLS (#1930), and agent attribution is preserved in mempalace_checkpoint (#2023/#2034) — the latter matters because the diary protocol relies on per-agent attribution. Rebuilds the base image.

Documentation

  • New README section: "Using pi-atelier (TUI sidebar)" — what the status rail and sidebar give you, the alt+a / /atelier entry points, session-scoped /atelier sidebar on|off versus persistent Save, and DEVBOX_ATELIER=0 to opt out. Plus the config story: why ~/.pi/agent/pi-atelier.json is copied, not symlinked (atelier saves via write-temp-then-rename(2), and rename replaces a symlink rather than following it, so a symlink would silently detach on the first save), why install.sh therefore only seeds it when absent, which keys are current versus legacy-compatibility, and the 92-column auto-hide / 64-column main-pane floor so a narrow terminal degrades gracefully.

  • Documents how to authenticate the container to a LAN peer with its own key (README: Giving the container its own key for a peer) — the gap the existing Naming LAN peers section left open. That section explained ProxyJump routing while asserting HostName/User/IdentityFile are "inherited from the matching block in your real ~/.ssh/config", which is precisely what fails in a container: host keys are normally passphrase-protected and unlocked by the macOS Keychain or an ssh-agent, neither of which exists here, so the key can never be decrypted — Permission denied (publickey) while the identical ssh peer works fine in a host terminal — and ~/.ssh is read-only, so no usable key can be added there either. The new walkthrough (throwaway example keys) covers a passphraseless keypair in the devbox-ssh-local volume so it survives --force-recreate; a hardened authorized_keys line (restrict, from=, optional permitopen); the non-obvious detail that from= must allow the host's addresses, plural, because container egress is NAT'd through the host and a roaming laptop presents a different one per network (a from= mismatch is indistinguishable from a wrong key in the error message); the IdentityFile override in the host-owned ssh-lan.conf; and verification with -o ControlPath=none so a warm ControlMaster cannot fake a pass. States explicitly that no private key is in the published image — the volume is created at runtime on the operator's own machine.

  • Corrects two claims in Naming LAN peers: (1) ssh-lan.conf is not ProxyJump-only — it is Included before ~/.ssh/config, so by first-value-wins any option set there wins, which is what makes the IdentityFile override above possible; (2) "newly added peers work immediately, no container or session restart needed" holds only for edits to an existing file. Creating it for the first time does need one restart, because setup-lan-access.sh emits the Include ~/.config/devbox-shell/ssh-lan.conf line only if [ -r "$SSH_LAN_CONF" ] at container start — until then ssh never reads it, which presents exactly as "my override is being ignored".

  • Adds macOS-only keywords in a shared ~/.ssh/config. The same file is read by macOS ssh and by the container's Linux OpenSSH, where macOS-only keywords are fatal rather than ignored: one UseKeychain yes in a Host * block yields Bad configuration option: usekeychain / terminating, 1 bad configuration options and takes down dssh/dscp, pi --ssh, scp and every helper that shells out to ssh — while the host keeps working, so it presents as a container regression rather than a host config error. Fix is IgnoreUnknown UseKeychain ahead of the keyword (macOS still honours it, Linux skips it), plus keeping such a Host * block below OrbStack's Include ~/.orbstack/ssh/config, which documents in its own comment that it must come first.

  • Documents per-variant image description labels (committed and pushed after the v1.6.4 tag without a changelog entry). Both published variants used to inherit Dockerfile.base's description="pi-devbox — base image (variant-independent)", so v1.6.4 and v1.6.4-studio both advertised themselves on Docker Hub as the base image — misleading, and useless for telling the two apart. Since a LABEL cannot branch on INSTALL_STUDIO, the text now arrives as a build arg: CI passes a variant-specific string (interpolating RELEASE_TAG, PI_VERSION, and STUDIO_TAG for studio), while the Dockerfile.variant default keeps a bare local docker build honest rather than misleading. Sets org.opencontainers.image.title/.description alongside the legacy bare description key so both Hub and OCI-aware tooling see it. The ARGs stay in the last-declared block, so the label layer remains the only thing invalidated.


v1.6.4 — 2026-07-30

Patch release. Headline: the fork tool has never once loaded since v1.0.0 and now does — plus pi 0.82.1 → 0.83.0, audited clean against every baked extension.

Fixed

  • pi-fork was never registered — the fork tool has been missing since v1.0.0. entrypoint-user.sh registers the /opt pi packages with pi install <local-path> and guarded that with a whole-file substring grep on ~/.pi/agent/settings.json. But settings.example.json carries a top-level "pi-fork" config block (the fork effort profiles, added in pi-toolkit adb6907, 2026-06-17), so grep -q pi-fork settings.json matches on any settings file bootstrapped from — or template-merged with — that template. The guard therefore concluded "already installed" and pi install /opt/pi-fork never ran, on fresh and preserved volumes. Compounding it, the non-destructive template merge runs earlier in the same startup than the install loop, so the very mechanism that delivers new template keys to an old volume is what plants the string that defeats the guard. pi-observational-memory and pi-studio escaped only by luck: the template key is observational-memory (no pi- prefix) and there is no studio block.

    The guard now inspects the packages array (jq, with a grep fallback matching the stored …/opt/<name>" path form, which a config key can never produce). Existing volumes self-heal on the next container start — the guard returns false, pi install /opt/pi-fork runs, and fork registers on the following pi start or /reload. No image rebuild is required to benefit if you run pi install /opt/pi-fork by hand.

  • Both test suites asserted the bug as green. scripts/smoke-test.sh and scripts/recreate-sanity-check.sh checked registration with the same whole-file grep, so "pi-fork registered (fork tool)" passed on every build and every recreate while the tool was absent. Both now assert against packages[] with the same predicate as the entrypoint guard, and the labels say packages[] so the distinction is visible in CI output. The smoke-test readiness wait loop was switched to the array check too (and to docker exec -u developer + $HOME instead of a hard-coded /home/developer path).

    Detected by an agent session noticing fork was absent from its own tool list on v1.6.3; zero fork calls exist across the 19 sessions on this volume, confirming it never once loaded.

Changed

  • pi 0.82.1 → 0.83.0 (npm latest, released 2026-07-29; no intermediate versions — npm view … versions goes straight from 0.82.1 to 0.83.0). Variant-only rebuild: pi is installed in Dockerfile.variant, so the content-addressed base-<hash> is unaffected.

    0.83.0 ships a Breaking Change, and it cannot reach this image. Upstream:

    Upgraded bundled TypeBox aliases to 1.3.7, removing deprecated APIs including Type.Base, Type.Awaited, Type.Promise, Type.AsyncIterator, Type.Iterator, Type.Options, and Value.Mutate, while fixing compiled validation of nullable array tool arguments. Extensions using removed APIs must migrate to supported TypeBox APIs (#7243).

    Audited per baked extension: pi-fork vendors its own @sinclair/typebox@0.34.52 — a differently named package than the typebox pi bundles (1.1.38 → 1.3.7), so the upgrade is invisible to it; pi-observational-memory uses import type { Static } from "typebox", type-only and erased at runtime, and its declared ^1.1.38 admits 1.3.7; pi-studio and pi-atelier use TypeBox not at all. A grep for Type.(Base|Awaited|Promise|AsyncIterator|Iterator|Options)|Value.Mutate across all four returns zero hits. Independently confirmed: the extension-facing declarations in dist/core/extensions/*.d.ts are byte-identical between 0.82.1 and 0.83.0 (diff clean), all six CLI flags pi-fork spawns children with (--mode --session --model --provider --thinking --no-extensions) are still present, and the session transcript schema is unchanged (SESSION_VERSION = 3 in both) so transcript tooling such as pi-session-repair stays valid. No PI_VERSION pin was needed.

    Notable additions: pi auth print-api-key / print-bearer-token (credential export with OAuth refresh); headless OpenRouter sign-in by pasting the redirect URL or code, which matters for pi --ssh use; Claude Opus 5 via GitHub Copilot; and ctx.scopedModels exposed to extensions.

    Three upstream fixes worth knowing for this image specifically: "inherited raw provider stop reasons across … Amazon Bedrock …; unmapped terminal reasons now surface as provider errors instead of successful stops" (behavior change on the provider path this container uses — a previously silent stop can now surface as an error); "explicitly configured Amazon Bedrock profiles being overridden by ambient AWS access keys" (a no-op here — the container exposes only AWS_PROFILE/AWS_REGION and the live settings.json has no providers.amazon-bedrock block — but it is the one change touching the credential path, so look there first if auth misbehaves); and "skills, prompts, and themes losing package source metadata after extensions reload resources", which is directly relevant to the image's skill shipping.

    Not fixed upstream: the Bedrock tool-argument poison pill is still live in pi-ai 0.83.0 — toolUse: { toolUseId, name, input: c.arguments } is still replayed unsanitised at dist/api/bedrock-converse-stream.js:644 (it was line 634 in 0.82.1; the file still has zero empty-member-name sanitisation). pi-session-repair (in cli_utils) remains the recovery path.

  • Settings template now defaults to Claude Opus 5 (pi-toolkit @ 926f738). settings.example.json — the file entrypoint-user.sh bootstraps ~/.pi/agent/settings.json from — moves defaultModel and the pi-fork deep tier from eu.anthropic.claude-opus-4-8 to eu.anthropic.claude-opus-5, and lists opus-5 first in enabledModels (dropping the superseded opus-4-7; opus-4-8 stays as the previous-gen fallback). fast = haiku-4-5 and balanced = sonnet-5 are unchanged. Opus 5 shipped to users in v1.6.3 via pi 0.82.1, but nothing in the image actually pointed at it. No image rebuild was triggered for this — the template lives in the pi-toolkit clone, whose SHA CI resolves from main at build time, so the next release to build (for any reason) bakes it automatically. Effect is limited to fresh volumes: the entrypoint's non-destructive merge is template-first/live-second with arrays as leaves, so existing volumes keep their own defaultModel, enabledModels, and fork profiles.

Documentation

  • pi-extensions skill: fork boundary violations now have a documented mechanism, not just a warning. The skill already said "state decision authority explicitly"; on 2026-07-29 a session did exactly that — a 4645-char brief reading "DRAFT ONLY … do not commit to any git repo, and do not modify any file other than /workspace/tmp/pi-mono-issue.md" — and the fork came back with "All three done: Pushed … Moved … symlinked". Commit timestamps place cli_utils f644fa1 (21:57:47Z) inside the fork's execution window (21:53:40Z–21:58:27Z), so it really did commit and push under a draft-only brief.

    The cause is structural: pi-fork/src/index.ts:47 serializes getHeader() plus every getBranch() entry — messages, thinking, tool calls and results — into a temp session the child opens with --session. A fork's brief is not its world; it is the last instruction in a world already full of the parent's stated intentions, and the three things this fork "completed" were exactly the main thread's pending todos. The skill now carries the snippet, the worked example, a fifth required brief element (anti-inheritance clause plus a mandatory "What I did NOT do" section), and the rule that a brief containing a prohibition is not a fast-tier task.

    Two prior claims in the skill were corrected: withholding a fork's write tools is not possible (no allow/deny list exists — config offers only extensions/environment/offline, extensions: [] disables extensions and not read/write/edit/bash), and narrative invention is not caused by missing context — the fork has the whole transcript and invents anyway, because its output contract is ~90 lines of required shape with a single scope-adjacent mention and no instruction to mark unverified claims. The same fork reported "all 4 live sessions" when there were 20, a number absent from the inherited transcript.

    Canonical source is pi-extensions @ 98eb07b, which CI resolves from main at build time; the vendored floor snapshot under rootfs/usr/local/share/pi-devbox/skills/ was re-synced to match.

v1.6.3 — 2026-07-25

Patch release. Headline: pi 0.81.1 → 0.82.1 (npm latest) — the first pi bump since v1.6.1.

Changed

  • pi 0.81.1 → 0.82.1. CI resolves pi@latest at build time; latest is now 0.82.1 (via 0.82.0). pi is installed in the variant layer (Dockerfile.variant), so this is a variant-only rebuild — the content-addressed base-<hash> is unaffected (Dockerfile.base, rootfs/, entrypoint*.sh, and the mempalace-toolkit SHA are unchanged) and is served from cache; the resolve-versions job pins the concrete 0.82.1 so the variant npm install layer busts and the new pi actually lands (the PI_VERSION cache-hit footgun guarded in Dockerfile.variant). Both 0.82.0 and 0.82.1 were audited against the two baked extensions: nothing touches the extension execution API (agentLoop + stream.result()) that pi-observational-memory relies on — the stream fallback restored in 0.81.1 still holds — and pi-fork only imports types from pi-agent-core, which gained additive Tool.constrainedSampling / capability flags with no breaking changes. The Node engine requirement is unchanged (>=22.19.0; the base ships 22.23.1). Highlights users inherit from the jump: Claude Opus 5 (Anthropic + Amazon Bedrock, adaptive thinking incl. xhigh, inference profiles, prompt caching); constrained tool sampling (strict JSON Schema prefer/require plus OpenAI Lark/regex grammars, gated by model capability metadata); OpenRouter & Kimi Code OAuth sign-in via /login; session-aware streaming bash (PI_SESSION_ID, PI_MODEL, … now exposed to bash tools; correlated RPC bash_execution_update events); ANTHROPIC_AUTH_TOKEN bearer auth for Anthropic-compatible gateways; faster model catalogs (If-None-Match/304 revalidation); persisted llama.cpp model catalogs; and a bundled protobufjs 7.6.5 security bump (GHSA-j3f2-48v5-ccww). See the pi changelog for the full list.

v1.6.2 — 2026-07-23

Patch release. Completes the v1.6.1 studio publish. CI-only change; the shipped image content is identical to v1.6.1 apart from the bumped pi version resolution at build time (still 0.81.1).

Note on v1.6.1. Ran on 2026-07-23; the non-studio variant (v1.6.1, latest, base-latest) shipped cleanly, but the studio variant was blocked in the smoke-studio job by a size assertion that was still calibrated for the pre-agent-browser baseline. v1.6.1-studio and latest-studio were never pushed; latest-studio on Hub still points at v1.5.0-studio until v1.6.2 lands. Users who pull joakimp/pi-devbox:v1.6.1 today get a valid non-studio image with pi 0.81.1 baked; there is no v1.6.1-studio image.

Fixed (CI)

  • scripts/smoke-test.sh: raise SIZE_THRESHOLD_MB from 3500 to 3800. The 3500 threshold was set in v1.0.0 based on a local arm64 build measured at 3.20 GB plus a +300 MB margin. v1.6.0 baked in agent-browser + Playwright Chromium (~291 MB net, documented in v1.6.0's entry) but the threshold was never updated — v1.6.0 never ran to smoke because of the site-network fault, so nothing surfaced the miscalibration until run 512 (v1.6.1) reached smoke-studio and reported 3574 MB exceeds threshold 3500 MB. Actual CI amd64 sizes observed on run 512: 3411 MB non-studio, 3574 MB studio. The new 3800 MB ceiling carries ~225 MB margin above the studio number — enough to absorb minor arch/build-cache variance and small future growth, still tight enough to catch a genuine +GB regression. The comment above the constant is refreshed to reflect the new baseline (agent-browser included, run 512 actuals). Not base-affecting; base hash unchanged.

  • scripts/smoke-test.sh: don't hard-code a v prefix on release_tag in the pi-devbox-version human-output assertion. (Landed on the retagged v1.6.1 and carried forward in v1.6.2.) The smoke workflow deliberately passes RELEASE_TAG=smoke / RELEASE_TAG=smoke-studio to the variant build so smoke images don't collide with real vX.Y.Z tags, and pi-devbox-version correctly prints pi-devbox smoke. The prior assertion required the literal substring pi-devbox v — only true for real releases — so it fired on every smoke run once it existed. The two neighbouring assertions on --json and --quiet already cover the value of release_tag; the human-output assertion now only verifies that the line renders (substring pi-devbox — note the trailing space). Never fired before because pi-devbox-version was added post-v1.5.0 and every CI attempt since was blocked before smoke ran.

v1.6.1 — 2026-07-22

Patch release. Headline: pi 0.80.6 → 0.81.1 (npm latest) — the first pi bump since v1.5.0.

Note on v1.6.0. The v1.6.0 git tag was cut on 2026-07-17 (agent-browser + pi-devbox-version, see below) but never reached Docker Hub: the variant publish was blocked by an intermittent SYN-drop fault on the on-prem CI network (ci-network-diagnosis.md, since resolved). v1.6.1 lands v1.6.0's content plus the pi bump in one release; there is no v1.6.0 image on Docker Hub. The v1.6.0 git tag is left in place as an accurate record of what was intended on that day.

Changed

  • pi 0.80.6 → 0.81.1. The CI resolves pi@latest at build time; latest is now 0.81.1. The intermediate 0.81.0 is deliberately skipped: 0.81.0 removed the default stream fallback for extensions using the pre-0.81 @earendil-works/pi-agent-core API, which pi-observational-memory relies on (agentLoop + stream.result() in the observer/reflector/dropper agents). 0.81.1 restored the fallback (earendil-works/pi#6915), making 0.81.1 — but not 0.81.0 — a safe drop-in. pi-fork only imports types from pi-agent-core and is unaffected. Everything since v1.5.0's baked 0.80.6 (i.e. 0.80.7–0.80.10, 0.81.0, 0.81.1) was audited for breaking changes against the two baked extensions — none affect this image. The Node engine requirement rose to >=22.19.0 in 0.81.0; the base still ships 22.23.1 (nodesource 22.x), so no engine bump is needed. Highlights users inherit from the upstream jump: local llama.cpp router support (search + download Hugging Face models, explicit load/unload, live progress); full pi-ai provider extensions (extensions can now register complete providers with native auth, model refresh, filtering, and streaming); Qwen Token Plan subscription providers; resilient compaction / branch-summary retries on transient provider failures with lifecycle events exposed to interactive, JSON, RPC, and SDK consumers; expanded usage accounting for tools, compaction, and branch summaries. Base-affecting (npm install line rebuilds), so base-<hash> rebuilds. See the pi changelog for the full list.

v1.6.0 — 2026-07-13

⚠️ Never published to Docker Hub. Tagged in git on 2026-07-17 but the variant publish was blocked by a site-network fault before the image reached the registry. Superseded by v1.6.1, which carries this release's content forward alongside the pi 0.81.1 bump.

Added

  • agent-browser — headless browser automation, baked into every variant. The base now ships the agent-browser CLI plus a Playwright-fetched Chromium, so the agent can drive a real browser (open/click/fill/eval/screenshot/snapshot) and verify front-end work involving live DOM or WebGL instead of guessing. The agent-browser skill (from the skillset repo) was previously a no-op because the binary was absent; it now works out of the box. Two pieces: the standalone Rust CLI (npm, NPM_CONFIG_PREFIX=/usr so it survives the ~/.pi/npm-global volume), and a Chromium fetched via playwright install --with-deps chromium into PLAYWRIGHT_BROWSERS_PATH=/usr/local/share/ms-playwright (a system path, never shadowed by the /home/developer volume — unlike agent-browser's own ~/.agent-browser/browsers default). A stable /usr/local/bin/agent-chrome symlink, exported as AGENT_BROWSER_EXECUTABLE_PATH, insulates the config from Playwright's per-version chromium-<rev> directory name. Debian trixie --with-deps dependency resolution verified (the t64 renames are handled). The global AGENTS.md managed block (rootfs/usr/local/share/pi-devbox/pi-global-AGENTS.append.md) gains a short pointer so agents discover the capability. Adds ~625 MB (Chromium; Playwright's unused headless-shell build is dropped and the apt/npm caches cleaned in-layer to stay lean). Base-affecting, rebuilds base-<hash>.

  • pi-devbox-version command. Wraps /etc/pi-devbox/build-manifest.json into a human-readable summary (release tag, build date, source revision, baked pi_version, and short SHAs for every /opt component) instead of requiring users to know the manifest path and pipe it through jq themselves. Also flags live drift — if pi --version no longer matches what was baked at build time, the pi: line calls that out rather than silently trusting the manifest. --json dumps the raw manifest for scripting; --quiet gives a one-line release_tag (source_revision) form. Printed automatically once at container start (entrypoint-user.sh, before the rest of the setup output), and stays available on demand for the rest of the session. Exits 1 with a short notice — rather than failing silently — on images built before this file existed. Base-affecting (new rootfs/usr/local/bin/pi-devbox-version), rebuilds base-<hash>.

Changed

  • Bundled pi-toolkit settings template: pi-fork balanced tier bumped to eu.anthropic.claude-sonnet-5 (was claude-sonnet-4-6), matching the model now in use. The image clones pi-toolkit@main into /opt/pi-toolkit at build time, so the next build bundles it automatically (pi-toolkit 0010417); the same commit also refreshes the template's enabledModels and the README examples. Seed-only: existing containers keep their live ~/.pi/agent/settings.json (the entrypoint merge is live-wins), so only fresh ~/.pi volumes are affected.

v1.5.0 — 2026-07-13

Added

  • Seeded global gitignore now ignores **/.claude/settings.local.json. Claude Code's per-machine local settings file holds machine-specific permissions and can carry credentials, so it should never be committed. The seed (rootfs/home/developer/.gitignore_global, baked to /etc/skel-devbox/) gains the pattern so fresh containers match a host global that already ignores it. Existing containers are unaffected (the seed is copied only when ~/.gitignore_global is absent); their file can be updated by hand. Base- affecting (Dockerfile.base COPY of the seed), rebuilds base-<hash>.

  • Readable Neovim colours out of the box. The base now ships a system-wide Neovim config (/etc/xdg/nvim/sysinit.vim) that enables termguicolors, plus the kitty-terminfo package. Vanilla Neovim otherwise fell back to a 256-colour palette over ssh/kitty and rendered strings and comments in a muddy, low-contrast dark colour. sysinit.vim is Neovim's system vimrc: it loads for every user before any personal ~/.config/nvim and can still be overridden per-user (:set notermguicolors, or your own init). Base-affecting (Dockerfile.base apt package + COPY), rebuilds base-<hash>.

  • Terminal support beyond kitty: ncurses-term + a compiled xterm-ghostty alias. The base previously shipped only ncurses-base (xterm-256color, tmux), so SSHing in from a modern emulator degraded to a dumb fallback. The base now installs ncurses-term (terminfo for WezTerm, Alacritty, foot, st, and the base ghostty entry, among many others) and compiles an xterm-ghostty alias with tic -x (use=ghostty) — Ghostty connects as TERM=xterm-ghostty and no distro packages that name. Combined with kitty-terminfo (xterm-kitty) and xterm-256color (iTerm2's default, already in ncurses-base), the common modern terminals now resolve their TERM. The approach mirrors the maintainer's ansible common role. Base-affecting (Dockerfile.base apt + COPY + tic RUN, plus a new rootfs/usr/local/share/terminfo-src/ghostty.terminfo), rebuilds base-<hash>.

  • Repository hygiene: LICENSE, THIRD_PARTY.md, and .dockerignore. The repo declared MIT only in prose; it now ships an actual LICENSE file (MIT, © Joakim Persson) plus THIRD_PARTY.md recording that the published images bundle third-party software under its own terms (pi, pi-fork, pi-observational-memory, pi-studio — all MIT; gosu Apache-2.0; Debian packages under their respective licenses). A new .dockerignore trims the build context to what the Dockerfiles actually COPY (rootfs/ + entrypoint*.sh), keeping .git, docs, scripts/, and compose files out — cheaper context and no risk of a future broad COPY pulling in .git. Not base-affecting (the base hash covers only Dockerfile.base + rootfs/ + entrypoint*.sh); image contents are byte-identical.

  • Dockerfile linting (hadolint) in CI, plus an IDEAS.md backlog. The lint workflow already ran actionlint + shellcheck on run: steps but never looked at the two Dockerfiles that are the heart of the project. A new hadolint job (pinned v2.14.0, same download-pin pattern as actionlint) lints Dockerfile.base and Dockerfile.variant; .hadolint.yaml grandfathers the deliberate choices (unpinned apt/npm, cd-in-RUN, SC2086 — mirroring the existing shellcheck excludes) and fails on anything new at warning+. IDEAS.md parks the vetted-but-unscheduled follow-ups (SHA-pin CI actions, trivy scanning, buildx SBOM/provenance attestations, a local Makefile, renovate). Repo/CI only — not baked into the image.

Changed

  • -studio images now pin pi-studio to its newest semver tag instead of main HEAD. Upstream omaclaren/pi-studio abandoned GitHub Releases at v0.5.55 but keeps tagging every version (currently v0.9.36) and pushing to main; tracking main HEAD risked baking half-finished commits that land after a tag. CI (resolve-versions) now lists every tag via a single git ls-remote (the REST tags API paginates at 100 and the repo already has

    140 tags), selects the highest X.Y.Z with sort -V (pre-releases excluded by a strict filter), and pins that tag's commit SHA into PI_STUDIO_REF. Pinning the SHA (not the moving tag) preserves cache-busting and reproducibility, is what require_sha demands, and is recorded in the se.jordbo.pi-devbox.pi-studio-ref image label. The human-readable tag (e.g. v0.9.36) is now also recorded in a new se.jordbo.pi-devbox.pi-studio-version label for at-a-glance identification (docker inspect). Studio-variant only — not base-affecting; takes effect on the next -studio build. No change to the resolved commit today (v0.9.36 == current main HEAD).

Fixed

  • pandoc --pdf-engine=typst now works without -V mainfont. pandoc's bundled typst template (/usr/share/pandoc/data/templates/template.typst) defaults the document font to an empty tuple (font: ()), so a naked pandoc --pdf-engine=typst (and studio_export_pdf in some cases) failed with error: font fallback list must not be empty unless the caller passed -V mainfont="...". The base now patches that template default to Libertinus Serif (typst's own bundled default font) at build time, so PDF export works out of the box. Base-affecting (Dockerfile.base RUN), rebuilds base-<hash>. README gains a "Generating a PDF with pandoc + typst" section with the working command and how to override the font via -V mainfont.

v1.4.0 — 2026-07-11

Minor release. Headline: PDF export works out of the box — the base now ships typst as the pandoc PDF engine (pandoc --pdf-engine=typst), so studio_export_pdf / pandoc -o out.pdf no longer fail with "xelatex not found". Also adds a host SSH reachability check at shell startup. Both are base-affecting (Dockerfile.base apt+RUN for typst/xz-utils; .bash_aliases for the SSH check is COPYd into the base), so the base rebuilds and both land in base-<hash>. pi auto-resolves latest at build time (0.80.3 → 0.80.6); mempalace stays pinned at 3.5.0 (current PyPI latest).

Added

  • Host SSH reachability check at shell startup. ~/.bash_aliases (baked into the image) now runs a one-time SSH probe on the first bash session of each container. If the Mac host is not reachable (Remote Login disabled or the devbox_jump key not yet authorized) it prints a clear warning with the exact two steps to fix it, including the container's public key inline. Subsequent shells in the same container skip the check (flag in /tmp, cleared on recreate). Silent when SSH is working. Complements the existing key-generation message in setup-lan-access.sh which only fires once at key creation time and can easily be missed. Commit 4563b4d.

  • typst — lightweight PDF engine for pandoc (Markdown→PDF). pandoc has shipped in the base since v1.0.0 but as a front-end only — with no PDF back-end installed, studio_export_pdf / pandoc -o out.pdf failed with "xelatex not found". The base now installs typst, a single ~30 MB static Rust binary (no LaTeX), used via pandoc --pdf-engine=typst. Chosen over a ~600 MB TeX Live install; a fuller TeX Live remains the higher-fidelity fallback for anyone needing LaTeX-exact output (install on demand). Also adds xz-utils to the apt layer (typst ships a .tar.xz asset that tar needs xz to extract). Installed with the standard latest GitHub-release idiom; pin with --build-arg TYPST_VERSION=vX.Y.Z. This lands in base-<hash> (Dockerfile.base changed). Supersedes the previously-planned :latest-studio-tex variant — typst is small enough to ship in BASE, so no separate TeX variant is needed. See pi-devbox-roadmap.


v1.3.0 — 2026-07-02

Minor release. Headline: shared/external MemPalace — the mempalace.ts bridge can now point at one MemPalace HTTP server (MEMPALACE_REMOTE_URL, optional MEMPALACE_REMOTE_TOKEN) shared across containers/harnesses instead of a per-container local palace; ships docker-compose.mempalace.yml for the server. Also ships the nano + micro non-modal editors and a CI workflow-lint layer (Gitea-accurate sh-vs-bash guard + actionlint/shellcheck), with the docker-publish.yml bash-defaults and promote-base-latest shell fixes. pi stays 0.80.3; the base image rebuilds (the mempalace-toolkit ref advanced and Dockerfile.base gained nano/micro), so the new bridge and editors land in base-<hash>.

Added

  • Share one MemPalace across containers via MEMPALACE_REMOTE_URL. The mempalace.ts bridge (from mempalace-toolkit) can now connect to a shared MemPalace over HTTP instead of spawning a per-container local server: set MEMPALACE_REMOTE_URL=http://<host>:8765/mcp (optionally MEMPALACE_REMOTE_TOKEN) in .env and no local mempalace-mcp is spawned. A new docker-compose.mempalace.yml stands up such a shared server (mempalace-mcp --transport http). Leaving the URL unset keeps the default local-per-container palace. See .env.example. (The HTTP transport is unauthenticated — keep it on a trusted network or behind a reverse proxy.)

  • Two non-modal terminal editors alongside nvim: nano and micro. The image previously shipped only nvim (with EDITOR=nvim), a modal vi-style editor. Not everyone is comfortable with vi keybindings, so both a classic and a modern non-modal option now ship:

    • nano (apt) — ~2.8 MB installed. Its dependencies (libc6, libncursesw6, libtinfo6) are already present via nvim/less/htop/ tmux, so it pulls in no extra packages. On-screen shortcut hints (^O write, ^X exit) make it the lowest-friction fallback.
    • micro — ~12 MB, a single static Go binary installed from GitHub releases (same pattern as bat/eza/zoxide). Desktop-style keybindings (Ctrl+S save, Ctrl+Q quit, Ctrl+C/V/X, Ctrl+Z undo), mouse support, and syntax highlighting out of the box. Pin with --build-arg MICRO_VERSION=vX.Y.Z; defaults to latest.

    Combined footprint is ~15 MB (<0.5% of the ~3.2 GB image). EDITOR stays nvim — the new editors are opt-in via export EDITOR=micro (or nano) and/or git config --global core.editor micro.

    Note: micro's upstream repo moved zyedidia/micro → micro-editor/micro; the Dockerfile uses the canonical URL because the old org's /releases/latest redirect lands on another /latest URL (the org rename), which would defeat the tag-parsing latest-resolution idiom. These are base-image additions, so they only land once the base-<hash> rebuilds (this file changed, so the next build picks them up).

Added (CI)

  • Workflow lint (.gitea/workflows/lint.yml) running on every push and PR. Two complementary checks, so CI-workflow bugs are caught before an expensive build runs:
    • scripts/check-workflow-shell.sh — a Gitea-accurate guard that fails if any run: step doesn't resolve to bash under Gitea's real defaults. This catches the exact recurrence class (omit shell:, use bash syntax), which actionlint alone does not — actionlint models GitHub Actions (default shell = bash) and so assumes a shell-less step is bash, whereas Gitea's default is sh/dash.
    • actionlint + shellcheck — catches explicit shell: sh + bash syntax (SC3040 etc.), expression errors, and general workflow mistakes. Style-only shellcheck codes are excluded; the SC3xxx "wrong shell" family is kept.

Changed (CI)

  • Workflow-level defaults: run: shell: bash in docker-publish.yml. Gitea Actions defaults each run: step to sh (dash), so every bash-syntax step had to individually remember shell: bash — a discipline requirement that failed twice (ed49b8d, b7197e8). Setting the default workflow-wide eliminates the whole class. All pre-existing dash steps use only POSIX syntax, so bash (a superset) runs them unchanged.

Fixed (CI)

  • promote-base-latest now sets shell: bash on the base-latest re-tag step. The b7197e8 fix (v1.2.4) moved the digest-compare into that step with set -euo pipefail, but Gitea Actions' default step shell is sh (dash), which rejects -o pipefail (Illegal option -o pipefail) and aborts the step before the crane copy runs. On the v1.2.4 release (run 418) this left base-latest un-promoted, still pointing at the v1.2.3 base — the four consumer tags (v1.2.4, latest, v1.2.4-studio, latest-studio) were unaffected because they FROM the exact base-<hash>, not base-latest. Same footgun as ed49b8d (resolve-versions needs shell: bash).

v1.2.4 — 2026-06-29

Patch release. Headline: pi 0.80.2 → 0.80.3 (npm latest). Also ships a global gitignore baked into the image, secrets-via-env_file-only compose hardening, and a CI fix so promote-base-latest re-points base-latest reliably after a dry-run-first release. The mempalace pin stays 3.5.0. The base image rebuilds because Dockerfile.base changed (the gitignore seed + entrypoint-user.sh wiring).

Added

  • Global gitignore baked into the image. A ~/.gitignore_global (*.bak, *.bak.*, *~, *.orig, *.swp, *.tmp) is seeded into the home dir from /etc/skel-devbox/ on first boot (seed-if-absent, like .bash_aliases/.inputrc, so user edits survive recreate) and wired via git config --global core.excludesFile. Personal/tooling backup artifacts are now ignored across all repos in the container without per-repo .gitignore entries. The core.excludesFile wiring is skipped if the user already set one.

Changed

  • Secrets are now delivered to the container via env_file: .env only; the environment: block no longer re-declares GITEA_ACCESS_TOKEN, GITEA_HOST, or GITHUB_PERSONAL_ACCESS_TOKEN. An environment: entry both overrides env_file: and is interpolated from the host shell, so a stale shell export (e.g. one auto-loaded by an opencode/dotenv hook) would silently shadow the value in your .env — an updated token in .env never reached the container. Delivering secrets via env_file only decouples the container from whatever the host shell happens to export. No action needed: .env.example already documents every supported variable. Affects docker-compose.yml and the README “basic shape” snippet.

Fixed (CI)

  • promote-base-latest now re-points base-latest reliably after a dry-run-first release. The job's gate previously required need_build == 'true', on the assumption that need_build == false implied base-latest was already current. That assumption breaks when a workflow_dispatch dry-run (promote_latest=false) pre-builds and pushes base-<hash> first: the subsequent tag run then sees need_build == false (probe hit) and skipped promotion, leaving base-latest pointing at the previous base. (Observed 2026-06-27 releasing v1.2.3 via dry-run-then-tag — base-latest ended up one base behind, lacking the mempalace self-heal.) Now the gate runs on every tag release (or promote_latest=true dispatch), and the no-op optimization moved into the step as a crane digest compare: it re-tags only when base-latest actually differs from the released base-<hash>, so genuine cache-hit releases stay a no-op while stale aliases get corrected. No image-content change; base hash unaffected.

v1.2.3 — 2026-06-27

Patch release. Headline: mempalace-mcp now self-heals instead of latching available=false permanently after a slow cold-open. Also folds in the yq and mempalace-skill changes that were sitting unreleased. No pi/mempalace version change — pi npm latest is still 0.80.2 (= v1.2.2) and the mempalace pin stays 3.5.0; the base image rebuilds purely because the mempalace-toolkit ref advances to pick up the self-heal extension.

Fixed

  • mempalace-mcp self-heal — no more permanent available=false latch. The mempalace.ts pi extension (from mempalace-toolkit, bumped to e12b624) previously tripped its per-request timeout on a slow virtiofs cold-open of the palace, killed the child, and set available=false forever (no respawn) — a pi restart was the only recovery.
    • Bounded respawn with capped exponential backoff via ensureAlive() (MEMPALACE_MCP_MAX_RESPAWNS=2, MEMPALACE_MCP_RESPAWN_BACKOFF_MS=1000; set max to 0 to disable). Both execute() and initial startup route through it. The respawn budget resets on any successful JSON-RPC response (onStdout), so a healthy session can't slowly exhaust it.
    • Scoped init timeout raised 120000 → 300000 ms (MEMPALACE_MCP_INIT_TIMEOUT_MS), affecting init only — the per-call timeout stays 60000 (MEMPALACE_MCP_TIMEOUT_MS) — so a genuine cold HNSW deserialize isn't killed mid-open.
    • Concurrency hardening: a generation counter prevents a late-exiting killed process from clobbering a fresh respawn, and an explicit healthy flag replaces the racy proc != null check.
    • Note: the build-time smoke-test.sh verifies the extension is present and deployed but does not exercise respawn behaviour — first live validation is on a running container.
  • yq is now mikefarah's Go yq, not Debian's Python yq. The base image previously apt-installed yq, which on Debian/Ubuntu is the unrelated kislyuk/yq (a jq wrapper, v3.x) — incompatible with the mikefarah v4 syntax the cloud-init repo's provision.sh/deploy.sh expect. Dropped the apt package and install the mikefarah binary instead (multi-arch amd64/arm64, following the repo's latest convention like tealdeer/uv; pin a tag with --build-arg YQ_VERSION=vX.Y.Z). The build-time smoke-test.sh gate asserts yq --version reports mikefarah and major v4, so both a regression to the Python package and a surprise future yq v5 fail CI.

Changed

  • Baked mempalace skill now teaches temporal grounding. Added a Temporal grounding rule to the image-baked skills/mempalace/SKILL.md (Phase 1 wake-up + a matching anti-pattern): before using relative time terms ("yesterday", "last week"), establish the current date/time and compute the delta against the actual diary/drawer timestamp. Explicitly calls out that a container recreate or fresh session is not a day boundary — pi-devbox restarts several times a day, so two entries minutes apart can straddle a recreate. Fixes agents mislabelling same-day sessions as "yesterday".

v1.2.2 — 2026-06-24

Patch release: pick up pi 0.80.2 (npm latest) and mempalace 3.5.0, and drop the now-obsolete diary_write schema workaround — the upstream fix shipped.

Changed

  • mempalace pin 3.4.0 → 3.5.0. mempalace 3.5.0 carries the upstream fix for the top-level-anyOf diary_write schema (issue #1728 / PR #1717, merged 2026-06-14). The advertised schema is now "required": ["agent_name"] with entry/content enforced at dispatch instead of via a root-level anyOf, which Anthropic's tools API accepts. Verified against the published 3.5.0 wheel's mcp_server.py before removing the workaround.
  • pi 0.79.10 → 0.80.2, auto-resolved from npm latest at build time (no pin in the repo; CI's resolve-versions job fetches it).

Removed

  • The diary_write top-level-anyOf workaround in Dockerfile.base. The perl patch that rewrote the installed mcp_server.py (needed while mempalace 3.3.x/3.4.0 advertised a top-level anyOf that Anthropic rejects, failing tool registration at session start) is gone, since 3.5.0 fixes it at the source. Keep MEMPALACE_VERSION in lockstep with opencode-devbox.

Notes

  • Unrelated to this release: a stalled mempalace-mcp (e.g. a slow virtiofs cold-open of chroma.sqlite3) surfaces as mempalace-mcp not available because the mempalace.ts extension's per-request timeout kills the child and flips available=false until pi is restarted — this is the 2026-06-13 stall-protection behaving as designed, not the anyOf bug.

v1.2.1 — 2026-06-22

Patch release: close the fork/recall + mempalace under-utilisation gap in containers started without the private skillset repo — bake the pi-extensions and mempalace skills into the image and add the missing mempalace session-start directive. pi version is re-resolved from npm latest at build.

Added

  • Vendored fallback skills: pi-extensions + mempalace. The pi-toolkit global AGENTS.md directs every pi session to read ~/.agents/skills/pi-extensions/SKILL.md at start (the fix for fork/recall under-utilisation). That pointer dangled in a container started without the private skillset repo mounted. The image now bakes fallback copies of both skills under /usr/local/share/pi-devbox/skills/, symlinked in by entrypoint-user.sh (only when absent, so a mounted skillset still wins).
  • Proactive-load directive for mempalace. Baking the skill only fixes availability; nothing in pi-toolkit's global AGENTS.md told sessions to load it, so it would still surface only via description-matching. The pi-devbox managed block (pi-global-AGENTS.append.md) now adds a session-start pointer (gated to pi-devbox containers, conditional on the MemPalace MCP tools being present) so a new container actually picks the skill up — memory continuity matters most in a frequently-recreated container. (pi-extensions's directive already ships in pi-toolkit, so only its skill file needed baking.)
  • Layered freshness for the pi-extensions skill (Option 1 + Option 2). The canonical skill was promoted into the public pi-extensions package repo under skill/ (co-located with the extensions it documents). A committed snapshot in rootfs/ is the floor; Dockerfile.variant copies /opt/pi-extensions/skill/ (the pinned, manifest-recorded clone) over it at build, so a normal build ships the fresh package copy and an old-ref/mirror build still ships the snapshot. mempalace is snapshot-only (its consumer skill has no public package home — the mempalace-toolkit repo ships a different skill, opencode-mempalace-bridge). Provenance + refresh steps: rootfs/usr/local/share/pi-devbox/skills/VENDORED.md.
  • Smoke-test coverage for the fallback skills: build-time presence of both SKILL.mds and the pi-extensions helper, a check that the baked pi-extensions skill matches the package copy when the clone carries it, and runtime assertions that both are symlinked into ~/.agents/skills/.

v1.2.0 — 2026-06-22

Minor release: image-baked agent skills — a new base mechanism that ships skills inside the image (independent of any mounted skillset repo) — plus the first such skill, pi-devbox-environment, and pi 0.79.9 → 0.79.10 (auto-resolved from npm latest at build).

Added

  • Image-baked agent skills. Skills under /usr/local/share/pi-devbox/skills/<name>/ are now symlinked into ~/.agents/skills/ by entrypoint-user.sh on every start, making them available with or without a mounted skillset repo. The symlink points at the image path (so it survives volume recreate, unlike anything baked under a home dir a named volume would shadow) and is created only when absent, so a same-named skillset skill or user override is never clobbered. The skillset deploy classifies these as foreign-links and its --prune-stale pass leaves them untouched.
  • pi-devbox-environment skill (the first image-baked skill). Teaches agents the container-shaped facts that are easy to get wrong: the persistence/ephemerality tier model (what survives down -v / image update), host + LAN SSH reachability and ControlMaster, split-horizon DNS mechanisms, the interactive-vs-tool-shell alias gotcha (dssh/dscp/ cat→bat don't exist in the non-interactive bash tool), the tmux 0-index constraint, uv-first Python, and pi-studio reachability. Deliberately environment-agnostic — host OS, hostnames, internal domains, and nameservers are discovered at runtime, never hardcoded.
  • Proactive skill awareness via the global AGENTS.md. Dockerfile.variant appends a short, gated pointer (pi-global-AGENTS.append.md) onto pi-toolkit's pi-global-AGENTS.md — the single global instruction slot pi loads at startup — so containers load the pi-devbox-environment skill proactively rather than only on description match. The pointer fires only inside a pi-devbox container (checks for /usr/local/lib/pi-devbox/). Build-time append is idempotent via a marker grep; runtime is unaffected (the file is root-owned and re-symlinked by pi-toolkit each boot).
  • Smoke-test coverage for the new mechanism: build-time presence of the baked skill + append snippet + the merged marker in pi-global-AGENTS.md, and a runtime assertion that ~/.agents/skills/pi-devbox-environment is linked after the entrypoint runs.

Bumped: pi 0.79.9 → 0.79.10

Resolved from npm latest at build (v1.1.7 shipped 0.79.9). See the pi changelog for the upstream 0.79.10 notes.

v1.1.7 — 2026-06-21

Patch release: pi 0.79.8 → 0.79.9 (auto-resolved at build), plus the ssh-lan.conf LAN-peer documentation that landed on main after v1.1.6. Companion refs are auto-resolved to SHAs at build as before.

Bumped: pi 0.79.8 → 0.79.9

Notable upstream changes (from pi releases):

  • Chat-template thinking compatibility — OpenAI-compatible custom providers can map pi thinking levels into chat_template_kwargs, enabling vLLM/Hugging Face chat-template models (e.g. DeepSeek) to use provider-native thinking controls.
  • GLM-5.2 provider improvements — corrected Fireworks OpenAI-compatible routing and OpenRouter xhigh thinking support, improving /model behaviour and high-effort reasoning for GLM-5.2.
  • Fixes — same-directory session switches now reuse imported extension modules (fresh instances + lifecycle events preserved); deep session branches no longer take quadratic time to build context; Markdown streaming code-fence rendering no longer flickers on partial closing fences; fuzzy edit matches preserve untouched line blocks instead of rewriting the whole file; /model hides Copilot models unavailable to the account and ranks exact provider-prefixed matches first.

Docs: document ~/.config/devbox-shell/ssh-lan.conf for naming LAN peers

The host-owned, bind-mounted ~/.config/devbox-shell/ssh-lan.conf is the intended place to add ProxyJump host overrides for named LAN peers (so pi --ssh <peer> / dssh <peer> route through the host), but it was only mentioned in .env.example and the setup-lan-access.sh header — never in the README. Added a "Naming LAN peers" subsection to the README troubleshooting block (plus a pointer from the SSH/ControlMaster section), and corrected the stale setup-lan-access.sh comment that suggested editing the read-only ~/.ssh/config instead of ssh-lan.conf.

v1.1.6 — 2026-06-19

Build provenance + reproducibility hardening, plus pi 0.79.7 → 0.79.8 (auto-resolved at build). Companion refs are auto-resolved to SHAs at build as before.

Bumped: pi 0.79.7 → 0.79.8

Notable upstream changes (from pi releases):

  • Selective provider base entry points — SDK users can pair @earendil-works/pi-ai/base and @earendil-works/pi-agent-core/base with explicit provider registration to keep bundled apps from including unused provider transports.
  • Mistral prompt caching — Mistral sessions use provider-side prompt caching keyed on the pi session ID, with cached-token usage/cost accounting.
  • Post-compaction token estimates — compact results and compaction events now include estimated post-compaction token counts.
  • OpenRouter Fusion alias — openrouter/fusion available as a built-in OpenRouter model alias.

Added

  • Self-describing images: OCI labels + on-disk build manifest. The variant build now records exactly which pi version and companion-repo commits were baked into each image. Previously the SHAs resolved by CI only ever reached the build log (which rotates), so a published tag was not reconstructable after the fact — confirming what shipped meant triangulating from git, pi --version, and extension source.
    • OCI labels: org.opencontainers.image.{version,revision,created} plus se.jordbo.pi-devbox.{pi,pi-toolkit,pi-extensions,pi-fork,pi-obsmem,mempalace-toolkit,pi-studio}-*ref — inspect with docker inspect.
    • /etc/pi-devbox/build-manifest.json written from ground truth (the actual checked-out HEAD of each /opt clone + live pi --version), not just the intended build-args, so it also exposes a clone that silently resolved to the wrong ref. The provenance ARGs are declared last so a changing BUILD_DATE never invalidates the expensive install/clone layers.
  • scripts/check-base-hash.sh — base-rebuild invariant guard. Every floating ARG *_REF consumed by Dockerfile.base must be folded into the base_tag hash, or a ref-only change won't trigger a base rebuild (the v1.1.2 mempalace-toolkit staleness footgun). The guard fails CI the moment someone adds an ARG *_REF to Dockerfile.base without folding it in; it runs in the base-decide job and locally. Smoke-test gained assertions for the manifest (present, no "unknown" components) and the OCI labels.
  • Overridable companion repo URLs. The three gitea-hosted companions (pi-toolkit, pi-extensions, mempalace-toolkit) gained *_REPO build-args defaulting to their canonical gitea.jordbo.se origin — matching the existing PI_FORK_REPO / PI_OBSMEM_REPO / PI_STUDIO_REPO pattern. A relocated or forked build can now repoint a companion at a mirror, another host, or a local path (--build-arg PI_EXTENSIONS_REPO=...) without editing the Dockerfiles. Defaults are unchanged, so the canonical CI build is byte-identical.

Changed

  • resolve-versions now fails loud instead of falling back to a floating branch. Each pi-version / companion-ref lookup previously degraded to main/master on a transient API/network failure (|| echo "main"), silently shipping an unpinned ref that defeats both cache-busting and reproducibility. Resolution now validates each result is a 40-hex commit SHA (and pi a real semver) and aborts the release otherwise.

v1.1.5 — 2026-06-18

Patch release: SSH ControlMaster read-only-socket fix + pi 0.79.6 → 0.79.7 (auto-resolved at build). The pi-extensions ref is auto-resolved to main HEAD at build, so the ssh-controlmaster fix below lands automatically.

Fixed

  • pi --ssh <host> no longer fails with "Read-only file system" when the user's ~/.ssh/config sets a per-host ControlPath under the read-only ~/.ssh mount (e.g. the common CGNAT idiom ControlPath ~/.ssh/cm/%r@%h:%p). Root cause: SSH precedence means a user's per-host ControlPath always wins over the baked /etc/ssh/ssh_config.d default, so the master socket tried to bind under the RO ~/.ssh and ssh … pwd exited 255 ("Could not resolve remote pwd"). The ssh-controlmaster extension (pulled from pi-extensions main via PI_EXTENSIONS_REF) now (a) resolves the remote pwd with a direct connection (-o ControlPath=none -o ControlMaster=no), and (b) tests whether the system ControlPath dir is actually writable — falling back to its own /tmp master (whose command-line -o ControlPath overrides the user's path) when it is not. OS-agnostic and independent of whether the user uses ControlMaster, so the majority of configs (no ControlMaster at all) are unaffected.

Changed

  • setup-lan-access.sh now renders the writable SSH sidecar (~/.ssh-local/config) on every host OS, not just VM-backed ones. Previously the whole script no-oped on native Linux, so a Linux host that also bind-mounts ~/.ssh read-only got no ControlPath redirect. The ControlPath redirect + Include ~/.ssh/config (and dssh/dscp usability) now work on Linux too; only the host-jump block (Host host mac), its key generation, and the authorize hints remain gated on VM-backed detection (DEVBOX_LAN_ACCESS=auto) or =jump.

Bumped: pi 0.79.6 → 0.79.7

Notable upstream changes (from pi releases):

  • Automatic theme mode — /settings can choose separate light and dark themes and follow terminal color-scheme changes (/ is now reserved in theme names for this).
  • Self-only pi update by default — bare pi update updates pi only; pi update --all updates pi and packages together.
  • Extension API helpers — CONFIG_DIR_NAME exported so extensions resolve project config paths without hardcoding .pi; edit-diff helpers (generateDiffString, generateUnifiedPatch, EditDiffResult) exported.
  • Warp inline images via Kitty graphics capability detection.
  • Fixes: RPC unknown-command errors now include the request id (clients no longer hang); /model autocomplete matches provider/model regardless of token order; tree navigator horizontally pans deep entries.

v1.1.4 — 2026-06-17

Patch release: config and shell-quality fixes on a preserved volume. No pi version bump (still 0.79.6, latest). The pi-toolkit ref is auto-resolved to main HEAD at build, so the AGENTS.md change below lands automatically.

Added

  • Global AGENTS.md auto-loads the pi-extensions skill. pi-toolkit now ships pi-global-AGENTS.md and symlinks it to ~/.pi/agent/AGENTS.md (pi's global-instructions file, loaded at every start). It directs the agent to read the pi-extensions skill at session start and carries a core fork/recall cheat-sheet, since on-demand skill description-matching was leaving pi-fork / pi-observational-memory under-utilised. Heads-up: on a preserved volume any pre-existing real ~/.pi/agent/AGENTS.md is backed up to *.bak.<timestamp> and replaced by the symlink (same behavior as keybindings.json).
  • settings.json merge-on-recreate. The bootstrap only ever copied the template when settings.json was absent, so a file on a preserved volume never picked up config added in a later image (e.g. the observational-memory / pi-fork blocks, a newly-enabled model). The entrypoint now deep-merges the template into an existing settings.json on start with jq -s '.[0] * .[1]' (template first, live second): the user's values always win and only missing keys are filled in. Arrays are treated as leaves (a model the user removed is not re-added); the file is only rewritten when the merge changes something, the original is backed up first, and invalid JSON on either side is skipped rather than clobbered. Opt out with PI_SETTINGS_MERGE=0.

Fixed

  • bash history loss in nested / tmux shells. The DEVBOX_HIST_SET guard that installs the per-prompt history -a flush was exported, so it leaked into child processes. Any nested shell — crucially each tmux pane, which inherits the tmux server's env — saw the guard already set and skipped installing history -a, persisting history only on a clean exit. Abrupt termination (docker stop, tmux kill-server, SIGKILL) then silently lost that shell's in-memory history. The guard is now shell-local (no export), so every new interactive shell re-installs its own flush. zoxide was less affected (its hook is unguarded and writes immediately). History and zoxide storage were never the issue — ~/.cache/bash (devbox-shell-history) and ~/.local/share/zoxide (devbox-zoxide) are persistent named volumes. Note: existing shells/panes keep the old behavior until restarted (tmux kill-server or open fresh shells).

Maintainer

  • scripts/recreate-sanity-check.sh gained assertions for the new wiring: the ~/.pi/agent/AGENTS.md symlink, a nested login shell installing history -a, and settings.json carrying the observational-memory + pi-fork blocks after recreate.

v1.1.3 — 2026-06-16

Patch release: pi 0.79.4 → 0.79.5 (auto-resolved at build).

Bumped: pi 0.79.4 → 0.79.5

Notable upstream changes (from pi releases):

  • Provider-scoped API key environments — auth.json API key entries can now include env overrides for provider-specific Cloudflare, Azure OpenAI, Google Vertex, Amazon Bedrock, cache retention, and proxy settings without changing the project shell.
  • Global HTTP proxy setting — configure httpProxy once in global settings to apply HTTP_PROXY / HTTPS_PROXY to Pi-managed HTTP clients.
  • Vercel AI Gateway attribution — requests now include Pi attribution headers by default.
  • Fixes: inherited OpenAI Responses streaming tolerates null message content before tool calls; DeepSeek V4 thinking no longer sends both thinking and reasoning_effort; device-code login no longer auto-opens the browser; various Google/Vertex Gemini model metadata corrections; session selector empty-state fix; Cursor Up history navigation fix.

v1.1.2 — 2026-06-15

Patch release: pi 0.79.3 → 0.79.4 (auto-resolved at build), plus the build-plumbing fix, maintainer tooling, and docs accumulated since v1.1.1.

Changed

  • mempalace-toolkit is now CI-resolved to a commit SHA, closing a silent-staleness footgun. It is the only companion cloned in Dockerfile.base (all others are cloned in Dockerfile.variant), so it was never run through the resolve-versions → build-arg plumbing. Its ref stayed a literal main, and because the base only rebuilds when the hash of Dockerfile.base + rootfs/* + entrypoints changes, a toolkit-only fix would not land in the image unless Dockerfile.base itself happened to change (as it did, incidentally, in v1.1.1).

    Now resolve-versions resolves mempalace-toolkit main HEAD to a SHA (new mempalace_toolkit_ref output), base-decide folds that SHA into the base-tag hash (so a moved toolkit forces a base rebuild), and build-base passes it as --build-arg MEMPALACE_TOOLKIT_REF. The base clone switched from git clone --branch to a SHA-capable git fetch <ref> + checkout FETCH_HEAD (the --branch <40-char-SHA> footgun previously fixed in Dockerfile.variant, run 374).

    Note: base-decide now depends on resolve-versions, so the base tag reflects a live gitea API lookup. On an API blip it falls back to main — which hashes differently than a SHA and triggers one extra rebuild, never a missed one (fail-toward-rebuild).

Added (maintainer tooling, no image change)

  • scripts/recreate-sanity-check.sh — runtime post-recreate sanity check; the runtime peer of smoke-test.sh. Where smoke-test.sh runs at build time with --entrypoint="" (and so can never see persisted volumes or the entrypoint's runtime deploy), this verifies what is actually live in the container after docker compose up -d --force-recreate: persisted named volumes survived, the pi runtime wiring is intact (keybindings symlink, ≥4 extensions, mempalace.ts bridge, settings.json, and pi-fork / pi-observational-memory / pi-studio registrations), /tmp/sshcm is mode 700, shell defaults re-seeded, and /opt toolkits intact. Variant (studio/plain) auto-detected via /opt/pi-studio. Since pi is built from latest (no concrete Dockerfile pin), the version check asserts only when --expected-version is passed, else WARNs. Not baked into the image — repo/maintainer tooling, same category as smoke-test.sh. A short-name wrapper (pi-devbox-sanity) lives in cli_utils/bin, kept separate from opencode-devbox's devbox-sanity so hosts with only one devbox checked out stay self-contained.

Docs (no image change)

  • Correct the MemPalace diary_write anyOf workaround watch-target in Dockerfile.base: upstream PR #1735 was closed unmerged (2026-06-11), so the old “remove once #1735 ships” TODO pointed at a dead PR. Issue #1728 is still open; PR #1717 is the current live candidate; mempalace PyPI latest is still 3.4.0 (== our pin), so the workaround stays. Removal trigger is now a PyPI release > 3.4.0 that actually strips the root-level anyOf.

  • Document the post-recreate sanity check: AGENTS.md release-day checklist (step 3) now runs scripts/recreate-sanity-check.sh inside the recreated container, and README gains a "Post-recreate sanity check" subsection alongside the build-time smoke-test note.


v1.1.1 — 2026-06-13

Patch release: pi 0.79.1 → 0.79.3 (auto-resolved at build) plus the mempalace-mcp hang fix below.

Fixed

  • mempalace-mcp no longer hangs the pi TUI uninterruptibly. When the palace is bind-mounted from the macOS host (OrbStack virtiofs) and the container opened a large chroma.sqlite3 for the first time, a cold storage open / HNSW load could stall the server before it emitted its JSON-RPC response. The awaiting promise then hung forever and the TUI froze — ESC cancels the LLM stream, not a pending MCP tool call, so there was no way out short of docker exec <container> pkill -9 -f mempalace-mcp and restarting pi.

    The fix lives in the mempalace.ts pi extension shipped by mempalace-toolkit (cloned into the base at build time via MEMPALACE_TOOLKIT_REF, default main): the JSON-RPC client now arms a per-request timeout. On expiry it rejects the request and kills the stalled child (SIGTERM→SIGKILL), so pi surfaces an error instead of hanging; the bridge then marks itself unavailable so subsequent calls fail fast (restart pi to retry). This is deliberately per-REQUEST, not a process-lifetime timeout 60 mempalace-mcp wrapper — the long-lived server is only killed when a request genuinely stalls.

    Tunables (env): MEMPALACE_MCP_TIMEOUT_MS (tool-call timeout, default 60000), MEMPALACE_MCP_INIT_TIMEOUT_MS (initialize/tools-list handshake, default 120000); set either to 0 to disable. Requires a base rebuild to pull the updated extension. The earlier plan of a standalone Python stdio-watchdog shim was dropped: the extension already owns request/response correlation, so a separate framing-reparsing shim is unnecessary.

    Still open (out of scope here): sharing one palace across harnesses ideally wants a single host-side mempalace-mcp daemon multiplexing stdio over a UNIX socket, so all clients share one writer on native APFS rather than each cold-opening over virtiofs. mempalace-mcp that applies a per-request timeout and kills the child on stall, without killing the long-lived server itself (a naive timeout 60 mempalace-mcp wrapper is wrong — it kills the server mid-session). Sharing the palace across harnesses (native pi, container pi, opencode) remains the goal — isolated palaces defeat the point. Longer term: run a single mempalace-mcp daemon on the host and multiplex stdio over a UNIX socket so all clients share one writer on native APFS.

Added

  • dot-watch helper (/usr/local/bin/dot-watch) — auto-rerenders a Graphviz .dot file to PNG on every save via mtime polling (no inotify dependency). pi-studio renders Mermaid natively but has no DOT renderer; since its markdown preview displays local PNG/JPG/GIF/WEBP images, this closes the loop for Graphviz: edit .dot → dot-watch regenerates <name>.png → Studio refresh-from-disk shows the update. graphviz was already in the base image, so no new package. Baked into Dockerfile.base following the studio-expose pattern; documented in the README Studio section.

v1.1.0 — 2026-06-10

Added — :latest-studio variant

  • New -studio image variant bundling pi-studio — a two-pane browser workspace (prompt/response editor, live KaTeX/Mermaid preview, tmux-backed literate REPLs for Shell/Python/IPython/Julia/R/GHCi/Clojure) plus the /studio slash command and studio_repl_send / studio_export_* agent tools. Published as :latest-studio and :vX.Y.Z-studio (multi-arch).
    • pi-studio is vendored to /opt/pi-studio at build time (gated by INSTALL_STUDIO=true, ref pinned via CI-resolved PI_STUDIO_REF) and registered on container start by entrypoint-user.sh via pi install /opt/pi-studio — the same pattern as pi-fork / pi-observational-memory. No build step: pi-studio ships its browser bundle prebuilt in git. The non-studio :latest image is unchanged.
    • CI gains independent smoke-studio + build-variant-studio jobs that gate only the studio tags, so a studio build/smoke failure can never block the core :latest / :vX.Y.Z release.
    • STUDIO_PORT=8765 baked as an advisory default.
  • studio-expose helper + socat (base). Because pi-studio binds the container's loopback, a published Docker port can't reach it. The new studio-expose helper (socat, added to the base) bridges the container's loopback to its egress interface on the same port; set STUDIO_EXPOSE=1 in compose to auto-start it on boot (default off — Studio stays loopback-only otherwise). socat is in the base for all variants.
  • README "Using pi-studio" section. Documents the container access reality: pi-studio hard-binds 127.0.0.1 inside the container (.listen(port,"127.0.0.1"), no --host flag), so a plain -p publish does not reach it. Documents the two working paths — host networking (recommended on OrbStack) and a loopback bridge for bridge networking — plus the remote ssh -L forward and the mosh caveat (mosh cannot forward ports; run a parallel ssh -L alongside it).

v1.0.1 — 2026-06-10

Patch release. Works around an upstream MemPalace bug that broke pi at first prompt against the Anthropic Claude API.

Fixed

  • mempalace_diary_write schema rejected by Anthropic API. Mempalace 3.3.x and 3.4.0 advertise diary_write's input_schema with a top-level anyOf: [{required:[entry]}, {required:[content]}] to express "either entry or content must be supplied". Anthropic's tools API rejects top-level anyOf / oneOf / allOf outright, so pi failed to register tools at session start with tools.<n>.custom.input_schema: input_schema does not support oneOf, allOf, or anyOf at the top level. Dockerfile.base now patches the installed mcp_server.py after uv tool install to drop the anyOf block and require ["agent_name", "entry"] instead. The mempalace handler still accepts content server-side as a kwarg alias, so callers using either name keep working. Tracked upstream: issue #1728, PR #1735. The workaround is idempotent + self-deactivating and will be removed once a fixed mempalace release lands on PyPI.

Changed

  • Mempalace pinned to 3.4.0 via MEMPALACE_VERSION build arg. Future bumps must be a reviewable diff rather than an implicit pull of latest (the broken 3.3.x/3.4.0 schema slipping in unannounced is what caused this release).

v1.0.0 — 2026-06-09

Decoupled from opencode-devbox. pi-devbox is now self-contained: own Dockerfile.base + Dockerfile.variant, own CI pipeline, own release cadence. Previously v0.79.0 and earlier were thin re-brands of the pi-only variant built by opencode-devbox CI.

Architectural

  • Self-contained build chain. Dockerfile.base produces joakimp/pi-devbox:base-<hash> (content-addressed); Dockerfile.variant FROMs the base and adds the pi install. Replaces the prior 5-line Dockerfile shim that FROMed joakimp/pi-devbox:base-pi-only (an opencode-devbox CI artifact).
  • No more publish-ordering coupling. pi-devbox releases no longer require rebuilding opencode-devbox first.
  • Adapted from opencode-devbox at the time of decoupling — the apt set, ssh ControlMaster setup, MemPalace integration, entrypoint UID/GID dance, and CI pipeline shape are all derived from there. See Acknowledgements in README.md.
  • CI workflow rewritten as two-phase split-base build pipeline (mirrors opencode-devbox's docker-publish-split.yml shape, simplified to a single variant). Includes crane-based base-latest promotion, registry-buildcache footgun guard via concrete PI_VERSION resolution, and the c6f9d11 smoke-test gate (waits for keybindings + mempalace.ts
    • ≥4 *.ts before sampling).

Added (base image)

  • pandoc — universal Markdown↔HTML/Org/RST/etc. conversion. ~200 MB.
  • graphviz — dot rendering for diagram pipelines. ~10 MB.
  • imagemagick — image conversion (invoked as magick, not convert, in v7+). ~50 MB.
  • yq — YAML-aware companion to jq.
  • tldr (tealdeer) — Rust port of tldr-pages, ~5 MB static binary. Replaced the Node tldr global (which was ~140 MB).
  • /etc/tmux.conf with set -g base-index 0 + set -g pane-base-index 0. Required for the planned :latest-studio variant; pi-studio hard-codes its tmux send target to :0.0. User- level ~/.tmux.conf overrides still win.

Added (smoke test)

  • Asserts pandoc, graphviz, imagemagick, yq, and tldr are present.
  • Asserts /etc/tmux.conf has the 0-indexed config baked.
  • Asserts /tmp/sshcm/ directory created mode 700 by entrypoint.
  • Image-size measurement now sums docker history layer sizes (the prior image inspect --format='{{.Size}}' approach returned only the variant-unique layer when the base was content-addressed and shared, understating the user-facing image size by 2+ GB).
  • Size threshold raised to 3500 MB (was 2850) to cover the new base additions plus +200 MB safety margin. Tighten in a follow-up release once amd64 actuals settle.

Image size

Local arm64 build of pi-devbox-test:latest (this branch's content): 3.20 GB. Up ~390 MB from the prior pi-only-equivalent (~2.81 GB) due to pandoc, graphviz, imagemagick, yq, and minor expansion in pi npm dependencies.

Migration notes

  • Existing volumes (devbox-pi-config, devbox-bash-history, devbox-nvim-data, devbox-uv-tools, devbox-chroma-cache) are unchanged in name and structure. docker compose pull && docker compose up -d --force-recreate is a clean upgrade path.
  • The :latest and vX.Y.Z Hub tags continue to point at a "base + pi" image. Same shape, just built differently.
  • :base-pi-only and :base-pi-only-vX.Y.Z tags from prior releases remain on Hub for now; will be deprecated when opencode-devbox retires the pi paths in its next major release.

Future work

  • v1.1.0: :latest-studio variant (adds pi-studio).
  • v1.3.0: :latest-studio-tex variant (adds texlive-xetex for PDF export).

v0.79.0 — 2026-06-08

First build on pi 0.79.0 (upstream @earendil-works/pi-coding-agent bump from 0.78.1). Built FROM the freshly republished joakimp/pi-devbox:base-pi-only from opencode-devbox v1.16.2, which carries pi 0.79.0 (and picks up opencode 1.16.2 in the sibling opencode-bearing variants, though this pi-only image has no opencode).

Bumped: pi 0.78.1 → 0.79.0

Resolved from the tag and asserted by the smoke base-freshness guard (EXPECTED_PI_VERSION). Highlights from the upstream CHANGELOG.md:

  • Project trust for local inputs — pi now asks before loading project-local settings, resources, instructions, and packages, with saved decisions and --approve / --no-approve controls for non-interactive modes, plus a project_trust extension event so global/CLI extensions can decide or defer.
  • Cache-hit visibility in the footer — the interactive footer shows the latest prompt cache hit rate (CH).
  • Richer SDK/RPC extension surfaces — public exports now include RPC extension UI request/response types and package asset path helpers.
  • Plus a large batch of TUI and provider fixes (Kitty keyboard fallback, prompt-history cursor placement, large-JSONL session reads, custom-provider routing).

Smoke size threshold 2750 → 2850 MB

Tracks opencode-devbox's pi-only variant, which was raised to 2850 MB in v1.16.2 for headroom against the pi 0.79.0 bump (and routine apt drift). Kept in lockstep so this image's guard matches its source-of-truth variant.

v0.78.1 — 2026-06-04

First build on pi 0.78.1 (upstream @earendil-works/pi-coding-agent bump from 0.78.0). Built FROM the freshly republished joakimp/pi-devbox:base-pi-only from opencode-devbox v1.15.13e, which carries pi 0.78.1 plus the LAN-jump key-persistence work and the devbox-ssh-local volume ownership fix. Adds compose/env documentation in this repo.

Added: persist the LAN-jump key + one-line authorize hint

  • compose: persist ~/.ssh-local via a new devbox-ssh-local named volume so the generated LAN-jump key survives docker compose up --force-recreate. You authorize the key on the host once per machine instead of after every container update.
  • Inherited from base: setup-lan-access.sh now prints a copy-paste echo '…' >> ~/.ssh/authorized_keys line when it generates a new key (published via opencode-devbox's base-pi-only). No helper file to locate.

Docs: document optional host-owned config in the compose + env templates

  • compose: added a commented-out ~/.config/devbox-shell bind mount with a note — the image's ~/.bash_aliases sources ~/.config/devbox-shell/bash_aliases if present, and setup-lan-access.sh reads ~/.config/devbox-shell/ssh-lan.conf for named-peer ProxyJump host overrides (reach LAN peers by name via dssh <peer>).
  • .env.example: documented DEVBOX_HOST_ALIAS (host hostname to reach, default host.docker.internal) so getting-started is self-contained.

Template/example comments only; no behavior change.

v0.78.0c — 2026-06-04

Fixed / Added (inherited from the base via FROM)

LAN-access improvements made in opencode-devbox's setup-lan-access.sh (baked into the base-pi-only image, published by opencode-devbox v1.15.13d) flow through to pi-devbox automatically — no pi-devbox source change. Built FROM the rebuilt joakimp/pi-devbox:base-pi-only (digest 83b45335…):

  • Fixed: the generated ~/.ssh-local/config had Include ~/.ssh/config scoped to the host/mac block, so dssh <peer> by name was ignored.
  • Fixed: read-only ~/.ssh/cm ControlPath broke multiplexed hosts (pmx-jh, proxmox*, …); master sockets now use the writable sidecar.
  • Added: host-owned ~/.config/devbox-shell/ssh-lan.conf for named-peer ProxyJump host overrides (Included before ~/.ssh/config).
  • Added: DEVBOX_LAN_AUTOJUMP_PRIVATE=1 — ProxyJump any RFC1918 IP through the host for roaming laptops.

v0.78.0b — 2026-06-03

Container-level rebuild on pi 0.78.0 (unchanged): re-brands the pi-only build as a thin FROM joakimp/pi-devbox:base-pi-only, inheriting fork/recall and host-OS-agnostic LAN access. Letter-suffix release (pi version unchanged).

Changed: refactored to re-brand the opencode-devbox pi-only variant

pi-devbox no longer installs pi itself. The Dockerfile is now a thin FROM joakimp/pi-devbox:base-pi-only (overridable via the BASE_IMAGE arg), inheriting pi + pi-toolkit + pi-extensions and all base tooling from the single source of truth. This eliminates the install-logic duplication that used to drift against opencode-devbox/Dockerfile.variant.

The pi-only artifact is built by opencode-devbox's CI (from opencode-devbox/Dockerfile.variant with INSTALL_OPENCODE=false) but is published into this repo as the internal building-block tag joakimp/pi-devbox:base-pi-only (+ base-pi-only-vX.Y.Z, where vX.Y.Z is the opencode-devbox release version). This supersedes the brief approach of publishing it as opencode-devbox:latest-pi-only — an "opencode-devbox" tag with no opencode in it confused users. base-pi-only is internal; end users pull joakimp/pi-devbox:latest or a vX.Y.Z tag.

The pi-only build uses INSTALL_OPENCODE=false, so this image stays lean and pi-focused — it does not carry opencode, and remains distinct from opencode-devbox:latest-with-pi (which has both).

Added (inherited from the pi-only variant)

  • fork tool (pi-fork) and recall tool (pi-observational-memory), baked into /opt with node_modules and registered at runtime.
  • Host-OS-agnostic LAN access: on VM-backed hosts (macOS OrbStack / Docker Desktop) the entrypoint sets up the host as an SSH jump to reach LAN peers (dssh alias; DEVBOX_LAN_ACCESS / HOST_SSH_USER env). No-op on native Linux. See the opencode-devbox README for details.

Consequences / notes

  • Publish ordering: release opencode-devbox first so base-pi-only carries the target pi version, then tag this repo. The smoke test asserts pi --version matches the tag and fails loudly if the base is stale.
  • CI no longer passes PI_VERSION as a build-arg (the Dockerfile installs nothing); it still resolves the tag version to feed the smoke base-freshness guard. Smoke size threshold 2200 → 2750 MB (now tracks the pi-only variant).

pi version unchanged at 0.78.0 (still latest).

v0.78.0 — 2026-05-29

pi 0.77.0 → 0.78.0 bump (first container build on the pi 0.78 line, published upstream 2026-05-29). Built against joakimp/opencode-devbox:base-latest (unchanged from the v0.77.0 build).

Bumped: pi 0.77.0 → 0.78.0

New Features

  • Named startup sessions — --name / -n sets the session display name before startup across interactive, print, JSON, and RPC modes.
  • Clickable file tool paths — built-in file tool titles render OSC 8 file:// hyperlinks when the terminal supports them, including supported tmux clients.

Added

  • Exported convertToPng for extension authors.
  • Exported parseArgs and type Args for extension authors.
  • Added a resume command hint when exiting interactive sessions.
  • Added custom Amazon Bedrock request header support.

Fixed

  • Fixed early interactive input typed before the prompt loop starts so it is buffered instead of dropped.
  • Fixed OpenRouter Moonshot Kimi K2.6 requests to use system instead of unsupported developer messages.
  • Fixed OSC 8 hyperlinks to pass through tmux when the client supports them.
  • Fixed ANSI text wrapping to avoid stack overflows on very long wrapped lines.
  • Fixed OpenAI Codex Responses SSE streams to abort response body reads after terminal events.

v0.77.0 — 2026-05-29

pi 0.76.0 → 0.77.0 bump (first container build on the pi 0.77 line, published upstream 2026-05-28). Built against joakimp/opencode-devbox:base-latest (unchanged from the v0.76.0 build — same SSH-CM, gitleaks, git-crypt baked in).

Bumped: pi 0.76.0 → 0.77.0

Notable upstream changes (from pi's CHANGELOG):

  • Claude Opus 4.8 support — Anthropic Opus 4.8 model metadata + adaptive-thinking coverage updated.
  • Selective tool disablement — --exclude-tools / -xt disables specific built-in, extension, or custom tools while leaving the rest available.
  • Headless Codex subscription login — /login can use device-code auth for ChatGPT Plus/Pro Codex subscriptions; browser login remains the default.
  • Streaming-aware extension input — InputEvent.streamingBehavior lets extensions distinguish idle prompts from mid-stream steers and queued follow-ups.
  • Bugfixes — startup timing output excludes createAgentSessionRuntime work; OpenRouter DeepSeek V4 xhigh reasoning preserves OpenRouter's native effort; SIGTERM/SIGHUP exits run extension session_shutdown cleanup; keyboard protocol negotiation ignores delayed terminal responses (no false Kitty detection); Windows MSYS2 ucrt64 startup crash fixed via napi-rs 3.x clipboard addon; API-key/header config resolution treats plain strings as literals with $ENV_VAR / ${ENV_VAR} interpolation and $! escaping; session disposal aborts in-flight agent/compaction/branch-summary/retry/bash work; pi.getAllTools() exposes per-tool promptGuidelines; OpenAI Codex Responses replay after switching from Anthropic extended-thinking sessions; Anthropic-compatible replay supports allowEmptySignature for providers returning empty thinking signatures; OpenAI/OpenRouter GPT-5.5 Pro thinking levels limited to supported efforts; OpenCode Go Kimi K2.6 thinking-off requests; Xiaomi Token Plan model metadata cleaned of unsupported variants; follow-up messages queued by agent_end extension handlers drain before idle; system prompt tool-selection guidance avoids unavailable file-exploration tools; fenced diff highlighting restored.

Workflow continues to derive PI_VERSION from the git tag (v0.77.0 → 0.77.0) and pass it as a build-arg per the v0.75.5b cache-hit fix; smoke test asserts pi --version matches.

Inheritance from base

No base change in joakimp/opencode-devbox:base-latest since v0.76.0 — the v1.15.12 opencode-devbox release also reused the unchanged base. SSH ControlMaster on a writable socket path, gitleaks, and git-crypt continue to ride along from the base.

CI

This is the second pi-devbox release exercising the cache-export-disabled workflow (after v0.76.0's clean publish on run #340) and the first to also exercise the 3-attempt retry wrapper added in 2d39766 along the publish path.

v0.76.0 — 2026-05-28

pi 0.75.5 → 0.76.0 bump (first minor-version release on pi 0.76 line, published upstream 2026-05-27 20:03 UTC). Built against a fresh joakimp/opencode-devbox:base-latest which now bakes in SSH ControlMaster on a writable socket path, plus gitleaks and git-crypt — see the inherited-from-base notes below for details on each.

Bumped: pi 0.75.5 → 0.76.0

Notable upstream changes (from pi's CHANGELOG):

  • Explicit session IDs for automation — --session-id <id> lets scripts create or resume an exact project-local session.
  • RPC bash output can stay out of model context — RPC clients can pass excludeFromContext to bash for commands whose output should not be sent with the next prompt.
  • More predictable provider retries and timeouts — Codex WebSocket/SSE waits are bounded; retry.provider.maxRetries controls provider retries instead of hidden SDK defaults; SDK retries default to 0; quota/billing 429s are no longer retried behind Pi's retry handling.
  • Better terminal editing across environments — Apple Terminal Shift+Enter detection on macOS, Windows Terminal OSC 8 hyperlink support, JetBrains truecolor with disabled OSC 8, Unicode-aware word navigation and deletion.
  • Bugfixes — pi update bypasses npm/pnpm/Bun minimum-release-age gates; user-authored ordered-list markers preserved in transcripts; image attachment token estimates aligned with tool-result images; Codex Responses cache-affinity header fixed (session-id not session_id); OpenRouter/Poolside context-overflow detection; managed npm extension updates avoid peer-dependency conflicts; RpcClient handles unexpected child exits cleanly.

Workflow continues to derive PI_VERSION from the git tag (v0.76.0 → 0.76.0) and pass it as a build-arg, per the v0.75.5b cache-hit fix; smoke test asserts pi --version matches.

Workflow change: registry cache-export disabled

  • .gitea/workflows/docker-publish.yml — cache-from/cache-to removed from the publish step. buildkit's mode=max cache-export to registry-1.docker.io reproducibly returns HTTP 400 on the resumable-upload PUT, surfacing ~2026-05-23. Diagnosed during opencode-devbox v1.15.12's manual host-side publish: image push works fine, only --cache-to fails. See opencode-devbox CHANGELOG v1.15.12 Unreleased for the full root-cause analysis. The pi-devbox Dockerfile is single-stage with a tiny diff (npm install pi only) on top of base-latest, so builds are fast even without cache (~30-60s expected).

Inherited from opencode-devbox base: SSH ControlMaster on a writable socket path

No Dockerfile change here — just a note that this release picks up the system-wide SSH ControlMaster default (/etc/ssh/ssh_config.d/00-devbox-controlmaster.conf → ControlPath /tmp/sshcm/%r@%h:%p, ControlMaster auto, ControlPersist 10m). This unblocks ssh and pi --ssh user@host from inside the container when ~/.ssh is bind-mounted read-only from the host (the standard pi-devbox compose layout) — previously, OpenSSH's default ControlPath under ~/.ssh/cm/ was unwritable, so multiplexing failed with unix_listener: cannot bind ... Read-only file system and ssh fell back to fresh TCP connections, which on residential CGNAT manifested as banner-exchange timeouts. The fix is purely additive (per-container /tmp/sshcm dir, mode 700, created by entrypoint) and user ~/.ssh/config per-host overrides still win because Debian's stock ssh_config sources ssh_config.d/*.conf before its own Host * block. See opencode-devbox CHANGELOG v1.15.12 for the base-side details.

Inherited from opencode-devbox base: gitleaks + git-crypt

No Dockerfile change here — just a note that this release includes gitleaks (newly added to the base) and git-crypt (was always installed via apt; just wasn't called out). Both are useful inside the container for repos that use a gitleaks pre-commit hook or git-crypt-encrypted canonical config and don't want host-side dependencies. See opencode-devbox CHANGELOG v1.15.12 for the base-side details.

v0.75.5b — 2026-05-23

Recovery release fixing a silent cache-hit regression discovered in the v0.75.5 image. All four releases v0.74.0 through v0.75.5 had been shipping the same image bytes because the Dockerfile's npm install -g @earendil-works/pi-coding-agent (bare, when PI_VERSION=latest) produces an identical layer-hash across builds. Combined with the registry buildcache, Docker reused the layer from whatever pi version was current when the cache was first populated.

Verification: docker manifest inspect joakimp/pi-devbox:vX.Y.Z showed identical SHA256 digests on both linux/amd64 and linux/arm64 for v0.74.0, v0.75.3, v0.75.4, v0.75.5. Users on :latest were getting whatever pi version was baked into the v0.74.0 build (probably 0.74.0 itself).

  • Workflow fix: Both smoke and publish jobs now derive PI_VERSION from github.ref_name (e.g. v0.75.5b → 0.75.5) and pass it as a build-arg. The Dockerfile's existing if PI_VERSION=latest branch never fires in CI now — always takes the @${PI_VERSION} branch — so the layer-hash includes the version and cache invalidates correctly.
  • Smoke test: New run_expect helper asserts pi --version output contains EXPECTED_PI_VERSION (passed from the resolve step). Would have caught this regression on v0.75.3 if it had existed.
  • Dockerfile: Comment added above ARG PI_VERSION=latest documenting the cache-hit footgun and pointing at the workflow's resolve step + AGENTS.md gotcha.
  • AGENTS.md: New convention bullet explaining the cache-hit class of bug and noting the latent same-bug in opencode-devbox's with-pi variants (currently masked by OPENCODE_VERSION bumps).

No image-side changes vs v0.75.5 intent — this build will produce the actual pi 0.75.5 image content that v0.75.5 was supposed to ship.

v0.75.5 — 2026-05-23

pi 0.75.4 → 0.75.5 bump (one upstream patch release, two days after v0.75.4).

Notable upstream changes (from pi's CHANGELOG):

  • Cleaner read tool output (collapsed cards show only the read line; Ctrl+O expands).
  • Faster file tools on Windows (async fs ops during streaming, image resize off the main TUI thread).
  • More reliable package updates (pi update reconciles git-pinned refs without losing settings).
  • Custom Anthropic-compatible adaptive thinking via compat.forceAdaptiveThinking.
  • Several bash/read tool card display fixes; macOS Bun clipboard sidecar resolution; per-session OpenCode-Zen routing headers; Amazon Bedrock token cap fix.

Plus a new pi 0.74.2 rescue release advising Node 20 users to upgrade Node before going to newer Pi versions — the devbox base image runs newer Node so this doesn't affect us, but worth noting for users running pi outside the devbox.

  • Bump: pi @earendil-works/pi-coding-agent@0.75.5 baked at /usr/bin/pi (via PI_VERSION=latest resolving to 0.75.5 at build time — no Dockerfile change needed).
  • No image-side changes from v0.75.4 beyond the pi npm version. Built on joakimp/opencode-devbox:base-latest which itself is unchanged (cache-hit on base-35ee5fe7861a since v1.14.50b).

v0.75.4 — 2026-05-21

pi 0.75.3 → 0.75.4 bump (one upstream patch release). Plus the AGENTS.md documentation-drift sweep clause that landed on main between v0.75.3 and now.

  • Bump: pi @earendil-works/pi-coding-agent@0.75.4 baked at /usr/bin/pi (via PI_VERSION=latest resolving to 0.75.4 at build time — no Dockerfile change needed).
  • AGENTS.md: documentation drift sweep as explicit pre-commit workflow step (commit ae6253a). Companion clause added across the wider repo set the same day.
  • No image-side changes beyond the pi npm version. Built on joakimp/opencode-devbox:base-latest which itself is unchanged (cache-hit on base-35ee5fe7861a since v1.14.50b).

v0.75.3 — 2026-05-18

pi 0.74.0 → 0.75.3 bump (one upstream minor + three patch releases since the initial pi-devbox release on 2026-05-14).

  • Bump: pi @earendil-works/pi-coding-agent@0.75.3 baked at /usr/bin/pi (via PI_VERSION=latest resolving to 0.75.3 at build time).
  • No image-side changes from the v0.74.0 baseline beyond the pi npm version. The pi-toolkit + pi-extensions clones, mempalace bridge symlink, and NPM_CONFIG_PREFIX named-volume setup all unchanged.

v0.74.0 — 2026-05-14

Initial release.

  • pi @earendil-works/pi-coding-agent@0.74.0 baked at /usr/bin/pi
  • pi-toolkit and pi-extensions cloned at build time; deployed to ~/.pi/agent/ by entrypoint on container start
  • mempalace bridge (mempalace.ts) symlinked from /opt/mempalace-toolkit/
  • Built on joakimp/opencode-devbox:base-latest