v1.10.1 (run 707) failed with ZERO failing steps: every step Success, `smoke`
printing "101 passed, 0 failed", `smoke-studio` "104 passed, 0 failed", both
jobs red. The clipboard fix from f3b3748 worked exactly as predicted; this is a
second, independent defect that run 704 had masked.
CHAIN, measured end to end rather than reasoned:
`ARG BASE_IMAGE` has no default on purpose (the two-phase build always supplies
it), which trips BuildKit's InvalidDefaultArgInFrom check. buildx attaches that
warning's source context to the build metadata as
buildx.build.warnings[].sourceInfo.data — the ENTIRE Dockerfile, base64, on ONE
line. docker/build-push-action writes that metadata to $GITHUB_OUTPUT as a
`name<<ghadelimiter_<uuid>` heredoc. Gitea's act_runner truncates any single
line at exactly 65536 chars, so the closing delimiter was cut off:
invalid format delimiter 'ghadelimiter_...' not found before end of file
and the runner failed the job while attributing it to no step at all.
v1.9.4 run 695 SUCCESS: longest metadata line 56355, 0 delimiter errors
v1.10.0 run 704 failed : longest metadata line 65536, 1 delimiter error
v1.10.1 run 707 failed : longest metadata line 65536, 1 delimiter error
65536 = 2^16: cut AT the cap, not merely long. Independent route via file size
rather than log parsing: 42251 B at v1.9.4 (base64 56335, matching the log) vs
50034 B now (base64 66712, truncated). The cap corresponds to a 49152 B
Dockerfile, so this cycle's comment growth crossed it by 882 B.
Located by asking which top-level metadata key CONTAINS the base64, instead of
assuming: it is buildx.build.warnings -> sourceInfo -> data in BOTH runs.
THIS IS THE SECOND FIX FOR THIS BUG. The first, `provenance: false` on the two
smoke build steps, was committed as 784fad7 and is REVERTED here: a real buildx
on another host showed provenance metadata present in both modes with a longest
string of 71 chars, i.e. the base64 was never in provenance. Shipping it would
have left the release broken a third time while looking like a fix.
Also rejected, each measured: floating action tags moving (act action-bundle
hashes byte-identical between runs 695 and 707, all four) and the build-check
annotation text changing (byte-identical).
FIX: `# check=skip=InvalidDefaultArgInFrom` as the FIRST line of
Dockerfile.variant — BuildKit parses `# check=` only before any other line, so
placement is load-bearing. No honest default exists for BASE_IMAGE: `scratch`
would satisfy the linter while being a lie, and would convert today's instant
"invalid reference format" into a failure deep in the build. Verified against
the actual edited file with `docker buildx build --check`: "Check complete, no
warnings found." Zero warnings means no sourceInfo, so the longest metadata line
drops 65536 -> ~1936 and the file's SIZE stops gating CI.
GUARD: scripts/check-dockerfile-directives.sh, wired into BOTH lint.yml and
docker-publish.yml's lint-gate, because a check that gates only `push` lets a
tag regress — the adoption slip that let doc-drift land 27 h after the v1.9.3
tag. It also fails if ARG BASE_IMAGE gains a default, so the directive cannot
rot into guarding a check that can no longer fire. Truth table, exit codes:
present 0, removed 1, demoted to line 2 → 1, ARG defaulted 1, file missing 2
(a gate that cannot run must not pass).
Release renamed v1.10.1 -> v1.10.2 with both failed tags left standing as
tombstones. Docs swept again (README "since" marker, Dockerfile decision
comments) because CI reads them from the TAG.
Gates: doc-drift 23 OK / 0 DRIFT; base-hash, workflow-shell, skill-floor,
lint-shell (17 files now), dockerfile-directives all rc=0.
v1.10.1 (run 707) failed with ZERO failing steps: every step reported Success,
`smoke` printed "Results: 101 passed, 0 failed", `smoke-studio` printed "104
passed, 0 failed", and both jobs went red anyway. The clipboard fix worked
exactly as predicted; this is a second, independent defect that run 704 hid.
MECHANISM, measured end to end:
buildx writes a provenance attestation into the metadata it returns, and that
metadata embeds the ENTIRE Dockerfile as a single base64 "data" field on ONE
line. docker/build-push-action writes that metadata to $GITHUB_OUTPUT as a
`name<<ghadelimiter_<uuid>` heredoc. Gitea's act_runner truncates any single
line at exactly 65536 chars, so once base64(Dockerfile.variant) crosses 64 KiB
the closing delimiter is cut off and the runner reports
invalid format delimiter 'ghadelimiter_...' not found before end of file
then fails the job while attributing it to no step at all.
v1.9.4 run 695 (SUCCESS): longest metadata line 56355 chars, 0 delimiter errors
v1.10.0 run 704 (failed) : longest metadata line 65536 chars, 1 delimiter error
v1.10.1 run 707 (failed) : longest metadata line 65536 chars, 1 delimiter error
65536 is exactly 2^16 — the line was TRUNCATED at the cap, not merely long.
Independent route via file size, not log parsing: Dockerfile.variant was 42251 B
at v1.9.4 (base64 56335, matching the log) and is 50034 B at HEAD (base64 66712,
cut to 65536). The cap corresponds to a 49152 B Dockerfile, so this cycle's
comment growth (+7783 B, mostly comments I wrote) crossed it by 882 B.
Hypotheses measured and REJECTED before landing this:
- floating action tags moved: the act action-bundle hashes are IDENTICAL
between run 695 and run 707 (all four), so no action changed version.
- build-check annotation changed: the ::warning text is byte-identical.
- stale clipboard assertion again: no, the suites print 0 failed.
WHY provenance: false and not shorter comments — trimming would restore the
margin and silently re-arm the trap for the next comment, with a failure mode of
"job red, no failing step, suite green", which cost most of this session to
diagnose once. Dropping the attestation removes the Dockerfile-size coupling
entirely.
BLAST RADIUS IS NIL for what ships: these two steps build with `load: true` and
the images are discarded after the suite runs, so the attestation has no
consumer. The PUBLISHED images are built by raw `docker buildx build --push` in
run: blocks (three call sites), which never write $GITHUB_OUTPUT metadata and
are therefore both unaffected by the bug and unchanged by this fix.
Verified: YAML parses; check-workflow-shell rc=0; doc-drift 23 OK / 0 DRIFT.
Next step is a `smoke_only=true` dispatch against main — the escape hatch this
workflow already documents — so the fix is proven before another tag is cut.
v1.10.0's build failed one assertion of 104 and published nothing. Rather than
re-point a tag CI had already consumed, the number moves: a version that failed
its build should stay failed and readable, not be quietly overwritten with
different bytes.
CHANGELOG gains a v1.10.1 section carrying the smoke fix (moved out of
v1.10.0's Fixed list), and states plainly that everything under v1.10.0 ships
here so the big section stays the content record. v1.10.0's heading is marked
"tagged, never published — superseded by v1.10.1" with a blockquote naming the
commit (12f99c4), the run (704), the four jobs that skipped, and the fact that
no image carries the tag.
Pre-tag doc sweep, because CI reads docs from the TAG and POSTs DOCKER_HUB.md
to the Hub as full_description:
- README "Terminal UI mode: fullscreen is the default (since v1.10.0)"
-> v1.10.1. A "since" marker pointing at an image nobody can pull is a
worse lie than no marker.
- Dockerfile.variant decision comments for the pi and pi-atelier bumps
-> v1.10.1 (these describe which release carries the change).
- scripts/smoke-test.sh deliberately KEEPS "the v1.10.0 tag build" — that
one is a historical event and v1.10.0 is its correct name.
- DOCKER_HUB.md needed no change: its fullscreen note is version-free.
No code changes. RELEASE_TAG is derived from github.ref_name, so nothing in the
build hardcodes the version.
Base layer does NOT rebuild: no base input (Dockerfile.base, rootfs/**,
entrypoint.sh, entrypoint-user.sh) has changed since 12f99c4, so base-decide
will cache-hit base-dad0f365ff24, which run 704's build-base already pushed at
07:52:38Z. This retry skips the expensive half.
Gates: doc-drift 23 OK / 0 DRIFT / 0 SKIP; base-hash, workflow-shell,
skill-floor, lint-shell all rc=0.
v1.10.0's first tag build (run 704) failed on ONE assertion out of 104, and
published nothing. The assertion was wrong, not the image.
scripts/smoke-test.sh required @mariozechner/clipboard to be installed and
require()-able at every install site, failing hard when the family was absent.
pi 0.86.0 (#9163) "Replaced the external native clipboard dependency with
bundled asynchronous macOS, Windows, and X11 helpers while preserving platform
command and OSC 52 fallbacks". Measured in the published tarballs rather than
inferred from the changelog:
pi 0.85.1 -> dependencies include @mariozechner/clipboard@0.3.9
pi 0.87.1 -> no @mariozechner/* at all
pi 1.0.0 -> no @mariozechner/* at all
So on a correct v1.10.0 image the package is legitimately gone, `if [ -z
"$sites" ]; then exit 1` fired, and both smoke jobs went red on the same line:
smoke 100 passed/1 failed, smoke-studio 103 passed/1 failed with every
studio-specific assertion green. build-variant, build-variant-studio,
promote-base-latest and update-description all skipped -> nothing published.
The gate worked; it was enforcing pi 0.85.1's dependency graph.
The guard is NOT deleted, because its "fail when absent" shape is the thing
that stops prune_foreign_natives() from deleting the binding instead of the
surplus platform copies. It now derives the expectation from what the image's
pi actually DECLARES:
declares + install loads -> pass
declares + no install -> FAIL (prune removed function)
does not declare + no install-> pass (pi bundles it since 0.86.0)
does not declare + install -> FAIL (orphan nothing depends on)
That re-arms automatically if a future pi re-adds the dependency, which a
skip-if-absent would not.
Verified as that four-quadrant truth table on EXIT CODES, not messages, since
run() keys on status: rc=0/1/0/1 as listed. The final check extracted the
command string verbatim from the committed file and ran it through `sh -c` the
way run() does, so the escaping was tested as shipped rather than as drafted.
Gates: doc-drift 23 OK / 0 DRIFT; base-hash, workflow-shell, skill-floor,
lint-shell all rc=0. smoke-test.sh is not a base-hash input, so the base layer
does not rebuild.
Decision: keep pi 1.0.0's default (tuiMode "fullscreen"). No tuiMode is baked,
so the image follows upstream rather than pinning the fleet to either mode -
but "we inherited a changed default" is only acceptable if the revert is
written down, so it now is.
README gains a "Terminal UI mode" section ahead of the pi-atelier section,
because the two interact. It gives all three scopes, taken from pi 1.0.0's own
docs/settings.md and docs/cli.md rather than from the changelog prose:
- permanent: "tuiMode": "regular" in ~/.pi/agent/settings.json
- one session: pi --tui-mode regular (the flag is real: cli.md:221)
- one project: the same key in .pi/settings.json, which overrides the agent
directory
It also documents the four related settings (fullscreenExitOutput,
fullscreenScrollbar, fullscreenCopyOnSelect, fullscreenWheelScrollLines) and
two things specific to this image:
- tmux: fullscreen uses the alternate screen, so tmux copy-mode shows the
pane history AROUND pi, not pi's transcript. That is the concrete reason
someone here would want "regular" back.
- fullscreenWheelScrollLines "auto" behaves differently over SSH (caps fast
wheel spins at 6 lines/event) than in a local macOS terminal (1 line) -
worth naming because this container is normally driven over SSH.
- pi-atelier works in BOTH modes (upstream has handled regular vs fullscreen
renderers separately since pi 0.84), so reverting costs nothing. Stated so
nobody assumes the sidebar is the price of the old scrollback.
The python3 merge snippet in that section was RUN before being documented:
against a realistic settings.json under an overridden HOME, twice, confirming
it is idempotent and preserves sibling keys including nested objects and the
_comment fields the seeded file uses. Documented code that has never been
executed is a guess.
DOCKER_HUB.md gets a short version of the same note, because CI reads it from
the TAG and POSTs it to Docker Hub as full_description - a behaviour change
this visible should not require reading the repo to undo.
Gates: doc-drift 23 OK / 0 DRIFT / 0 SKIP / 0 FAIL; base-hash, workflow-shell,
skill-floor, lint-shell all rc=0.
DOCKER_HUB.md is POSTed to Docker Hub as full_description by the
update-description job, and CI reads it from the TAG, not from main. Its
"Modern CLI tooling" list said "Data: jq, yq" while this release adds sqlite3,
bc, dc and column to the base - so the public page would have shipped
incomplete for the whole v1.10.0 cycle and the fix could not land until the
next tag.
Caught by the pre-tag doc sweep rather than by a gate: check-doc-drift.sh
verifies version CLAIMS (pins, Node major) against the build files, but it
cannot know that a hand-written feature list grew stale, because nothing
declares that list's contents. Same failure shape as the v1.9.0 Node 22/24
mismatch that shipped eight releases running.
README carries no equivalent list, so there is nothing to mirror.
Renames the unpublished v1.9.5 section to v1.10.0 and adopts the two bumps
that section had deliberately deferred. v1.9.5 was never tagged or published,
so nothing shipped under that name.
A minor, not a patch. v1.9.5 was numbered under the "patch - pi version bumps"
rule, but this release carries a pi MAJOR, a three-minor pi-atelier jump, four
new base packages and a new mailbox feature. The release carrying a 1.0.0
should not be the one numbered as a patch.
pi 0.85.1 -> 1.0.0. v1.9.5 held at 0.87.1 because obsmem's only compatibility
work names Pi 0.87 and no upstream issue mentions 0.99 or 1.0. That reasoning
had a hole worth naming: "no issue mentions 1.0" is an ABSENCE OF A STATEMENT,
not a measurement. So 1.0.0 was measured, against the published npm tarballs
for 0.87.1 and 1.0.0 unpacked side by side:
- 1.0.0 has NO "### Breaking Changes" section at all. Those belong to 0.87.0,
0.86.0, 0.84.3, 0.84.0, 0.83.0, 0.80.8, 0.80.7, 0.75.0 - none to anything
between 0.88 and 1.0.0. The major is a milestone (fullscreen default,
leaner codemode), not an API break.
- finishTurn is still in 1.0.0's dist, so the pinned obsmem SHA keeps the API
it migrated to.
- All 8 pi.* APIs our extensions call exist in 1.0.0's dist: registerTool,
registerCommand, registerFlag, getFlag, on, exec, sendMessage,
sendUserMessage.
- engines.node >=22.19.0 on both; the image ships 24.x.
The 0.99.2 change that looked fatal and is not: from 0.99.2 the DEFAULT MCP
exposure is `codemode`, so such tools are "neither declared to the model nor
listed" and must be found with searchTools(). That would gut the MemPalace
protocol if MemPalace were a builtin-MCP server. It is not - mempalace.ts and
mcp-loader.ts each run their own MCP client and register tools via
pi.registerTool(), which is why they are named mempalace_search and not
mcp__mempalace__search. Second route: settings.json has neither an mcpServers
block nor an mcp block. Filed upstream under "Changed", not "Breaking", so it
would have been easy to meet the hard way.
pi-atelier v0.10.3 (ed3837b) -> v0.13.0 (34d26f1), closing the gap v1.9.5
flagged as the pi bump's residual risk. peerDependencies are unchanged at pi
>=0.84.0 across v0.10.3/v0.12.1/v0.13.0 - a FLOOR, so not evidence of
anything. What was checked instead: v0.13.0 carries its OWN breaking change
(the entry point no longer exports the internal registry and layout helpers),
which reaches nothing of ours - grepping pi-atelier and registerSidebarPanel
across mempalace-toolkit, pi-devbox, pi-extensions, pi-fork,
pi-observational-memory and pi-studio returns ZERO matches in all six. v0.12.0
removed the showSessionActions setting (0 references here) and moved the
Control Center shortcut Alt+A -> F6 (0 references here, so no doc goes stale).
showSidebarAgent/showSidebarTodos, the only atelier keys README names, are
both still in v0.13.0's README. TuiMainScreen and renderLayoutFrame both still
exist in pi 1.0.0's dist.
PI_ATELIER_VERSION is bumped with PI_ATELIER_REF; it is a separate ARG and the
image label would otherwise have lied.
v0.13.0 is an ANNOTATED tag: `ls-remote --tags` reports the tag object
(dd06971), not the commit (34d26f1). check-doc-drift.sh resolves the commit, so
that is what the CHANGELOG names. The gate caught the wrong SHA on the first
attempt - noted in the CHANGELOG for future bumps.
NOT PROVEN, and stated so acceptance does not mistake it for cleared: grepping
dist shows the SYMBOLS survive, not that their SIGNATURES are unchanged -
necessary, not sufficient. 1.0.0 is four days old and neither obsmem nor
atelier has a commit naming it. Acceptance must prove the obsmem workers CAP
TURNS (peerDeps are *, so a mismatch is silent) and that the atelier sidebar
PAINTS.
User-visible behaviour change, deliberately NOT overridden: 1.0.0 makes the TUI
fullscreen by default, replacing the terminal's normal scrollback.
tuiMode: "regular" restores the old behaviour. No baked default is set, so the
image inherits upstream's choice rather than silently pinning the fleet.
Gates: check-doc-drift 23 OK / 0 DRIFT / 0 SKIP / 0 FAIL; base-hash,
workflow-shell, skill-floor, lint-shell all rc=0.
v1.9.4 held pi at 0.85.1 because 0.87.0 REMOVED `shouldStopAfterTurn`, which
pi-observational-memory 3.1.4 still used in all three workers. Its peerDeps are
`*`, so nothing refuses at install time and the breakage is silent at runtime --
turn caps ignored, workers losing their specialised prompts. PR #83 fixes that
and merged 2026-09-23.
Re-measured 2026-10-01 across src/agents/{observer,reflector,dropper}/agent.ts,
deliberately reusing the v1.9.4 audit's counting method so the numbers compare:
e7d77dc (3.1.4, baked in v1.9.4): shouldStopAfterTurn 3, finishTurn 0, systemPrompt 3
731c3d4 (pinned here) : shouldStopAfterTurn 0, finishTurn 3, systemPrompt 0
All three workers migrated, and the 0.86.0 AgentContext.systemPrompt reads are
gone too. The coupling is asymmetric and that is why these move in ONE commit:
3.1.4 + 0.87.1 silently ignores turn caps, and 731c3d4 + 0.85.1 breaks the
workers outright, because finishTurn does not exist before 0.87.0.
Pinned to a SHA rather than waiting for a tag, departing from the v1.9.4
instruction to wait for a release: #83 is merged but the newest obsmem tag is
still 3.1.4, cut 2026-09-20, BEFORE the merge. Upstream tags slowly and moves
master often (e7d77dc -> 1529e14 -> 731c3d4 in nine days), so waiting means
holding pi indefinitely. A pinned SHA keeps the property `master` lacks:
rebuilding this tag later produces the same image.
The 40-char form is load-bearing, not pedantry. check-doc-drift.sh recognises a
literal SHA only through a 40-char match, so a 7-char pin would fall through to
its branch-or-tag lookup, fail to resolve, and downgrade pi-obsmem's drift check
to a silent SKIP -- a pin that reads correctly and is no longer verified. Full
gate run after this change: 23 OK, 0 DRIFT, 0 SKIP, 0 FAIL.
0.99.0/0.99.1/0.99.2 and 1.0.0 all exist upstream and are deliberately skipped:
obsmem's only compatibility work names Pi 0.87 (2b1dc1c) and a repo-wide
issue/PR search for 0.99 or 1.0 returns zero matches. 1.0.0 is its own round.
pi-atelier stays at v0.10.3 and is the residual risk. Its peerDeps declare pi
>=0.84.0 -- a floor, satisfied -- on both v0.10.3 and the current v0.12.1, so it
spans this bump. But atelier hooks pi TUI internals that a declared floor does
not protect, and an under-declared floor is exactly what failed to warn anyone
at pi 0.84. Acceptance must confirm the sidebar PAINTS, using the two-sided
check from 0.84.4/0.85.1 that distinguishes "loaded" from "silently absent".
Also here:
- check-doc-drift.sh gains a pi-obsmem pin check. The ref-move check covers the
same component, but for a pinned SHA it can only answer "upstream did not
move", never "the table still says what we bake".
- scripts/lint-shell.sh mode 100644 -> 100755. Pre-existing since f25efa0 and
the only non-executable script in scripts/; it was latent because both CI
steps call it as `bash scripts/lint-shell.sh`, but it failed rc=126 "bad
interpreter" when invoked directly. Same dropped-exec-bit signature recorded
on 2026-09-22, found the same way: by RUNNING it, not by reading a diff.
- Unreleased section renamed to `## v1.9.5 — 2026-10-02`, satisfying the
release-gate rule that a tag's CHANGELOG must name its own version.
Both are floating refs that moved since v1.9.4, so the next tag would bake them
either way; the gate's complaint was that nothing in the repo said so. Naming
them is the fix the gate actually asks for.
mempalace-toolkit 2167a1b -> 975ab92 adds mailbox dormancy: an ask may declare a
`dormant_unless` predicate and is withheld from the ANNOUNCED owed-set while
every condition still matches its baseline. It is backward compatible by
construction -- an ask without the key behaves exactly as before -- and
fail-visible: any predicate that cannot be evaluated announces the ask rather
than hiding it, because the dangerous failure is work that disappears, not a
spurious nag.
pi-studio v0.9.60 -> v0.9.61 is adopted as-is. NOT pinned, deliberately, and the
reason is worth recording: check-doc-drift.sh gives pi-studio kind `studio`,
which resolves the highest semver TAG of the upstream repo and ignores
PI_STUDIO_REF entirely (resolve_studio, ~line 551). Setting
PI_STUDIO_REF=v0.9.61 would therefore not be tracked by the gate -- the moment
upstream tags v0.9.62 the gate would report DRIFT against a value we no longer
bake, i.e. a false positive by construction. Pinning pi-studio is a coherent
thing to want, but it requires moving it from kind `studio` to kind `ref` in the
gate at the same time, which changes release-gate semantics and is a separate
decision from adopting this release.
Still drifting, and left drifting on purpose: pi-obsmem e7d77dc -> 731c3d4.
PI_OBSMEM_REF=master is CI-resolved at build time, master is 4 commits ahead of
the 3.1.4 tag, and the newest tag remains 3.1.4 -- so PR #83 ("Fix Pi 0.87
memory worker compatibility", merged 2026-09-23) is still reachable only from
master. Whether to bake an unreleased master or stay on 3.1.4 and hold pi is a
pin decision with a silent failure mode behind it (obsmem's peerDeps are all
`*`), so it is not being made in a changelog commit.
sqlite3 is the only one of the four that was genuinely missing rather than
merely absent. MemPalace keeps both the palace and the logstream as SQLite
files (palace/chroma.sqlite3, logstream.sqlite3) and mempalace_status reports a
sqlite_integrity block, so every integrity or forensics check on this fleet has
so far gone through a python3 -c one-liner because the CLI was not in the image.
bc and dc are low value on their own and the CHANGELOG says so plainly: awk,
python3 and perl are all already baked and each is strictly more capable. They
are here because copy-pasted shell snippets assume bc exists. What actually went
wrong on 2026-09-28 was NOT bc's absence: printf '%6.2f' was handed the empty
output of the missing bc and rendered it as a confident "0.00 days" for a figure
that was really 4.81 days. No package fixes that failure mode -- only not taking
a formatted number on trust does.
bsdextrautils ships /usr/bin/column.
Cost MEASURED rather than estimated: 1311 KB total (587 + 236 + 149 + 339) and
ZERO transitive packages. libsqlite3-0, libreadline8t64, zlib1g, libsmartcols1
and libtinfo6 each already report "install ok installed", so nothing new is
pulled under --no-install-recommends.
Package names were verified with apt-cache and dpkg -S on a real trixie host
instead of being inferred, which caught two traps that would each have produced
either a build failure or a silently missing binary: dc is a SEPARATE binary
package from bc on Debian, and column ships in bsdextrautils, not in the
pre-bullseye bsdmainutils where it used to live.
Declined at the same time, recorded so the omission reads as a decision rather
than an oversight: datamash and xsv/csvkit. Python's stdlib csv module handled a
real 7-file Excel-export concatenation that day -- UTF-8 BOM, no trailing
newlines, and bare CR/LF inside quoted fields -- correctly and without them.
Dockerfile.base is in the base hash, so the next tag rebuilds the base (~64 min)
regardless of what else it carries.
~/.ssh is commonly bind-mounted READ-ONLY from the host, so a per-host
`ControlPath ~/.ssh/cm/%r@%h:%p` — the standard CGNAT multiplexing recipe, and
correct on the host — resolves inside an unwritable dir in the container. Every
push dies `unix_listener: cannot bind to path ...: Read-only file system`,
behind git's misleading "make sure you have the correct access rights".
setup-lan-access.sh already writes the fix: ~/.ssh-local/config overrides
ControlPath into the writable ~/.ssh-local/cm BEFORE `Include ~/.ssh/config`,
so -F repairs the socket path and keeps every per-host User/Port/IdentityFile.
entrypoint-user.sh now points git at it, guarded on the sidecar existing —
setup-lan-access.sh writes none on native Linux Docker, where -F at a missing
file would break every git-over-ssh call instead of fixing one. An existing
core.sshCommand is left alone (first-wins, as for the three git settings above).
Why code and not another doc line: the remedy was already in the global
AGENTS.md, in pi-devbox-environment SKILL.md §3, in 24 MemPalace drawers from
three devices, and printed verbatim by recreate-sanity-check.sh — and an agent
that had run that script two hours earlier still hit the failure and reinvented
a /tmp/sshcm workaround. A fifth copy was not the missing piece.
Assertions, each where it can actually pass:
- smoke-test.sh: two STATIC greps (wiring line + its [ -r ] guard). `run` uses
--entrypoint="", so asserting the runtime value there would repeat the v1.8.0
mistake of an assertion that cannot pass, unvalidated until the next tag.
- smoke-test.sh runtime phase: a BICONDITIONAL — sidecar present => must route
through it; absent => must be unset. The absent arm is the one CI exercises
(native Linux runner), so "is set" would have failed CI for a correct image.
- recreate-sanity-check.sh: the runtime assertion, plus an explicit fail for the
inverted state (set while the sidecar is missing). The permanent "default ssh
precedence" warning keeps its severity but now states that it is structural and
can never reach zero, and whether git is wired, unwired, or has no sidecar.
All five arms exercised against the real script before commit; that caught a
defect in the first draft, which reported "git IS wired ... unaffected" about a
state where the sidecar was gone and every git-over-ssh call failed.
Host ~/.ssh/config needs no change: the same line is right on the host and
unusable through a read-only mount, so the fix belongs in the container layer.
Unreleased; no pin moves except pi-extensions master 25c1265 -> 143a214.
1. entrypoint-user.sh: first-run sentinel is now config.json, which is what
`mempalace init` writes. The old test on palace/ (created by mining, not
init) re-fired "Initializing MemPalace" on every boot of a container that
never mined locally: v1.9.4 acceptance measured 1 line after one boot, 2
after a restart. Idempotent, so harmless; the comment lied. This file is a
base-hash input, so the next tag rebuilds the base (326fb7c03949 predicted
with the pinned sort below; f40c4b7b103d before).
2. Dockerfile.variant: `git config --system --add safe.directory` for the
seven root-owned /opt clones (listed, not `*`). Since git 2.35.2 any git
command in a repo owned by another user fails "dubious ownership"; as
`developer` that made `git -C /opt/pi-atelier rev-parse` print nothing
(one false FAIL in the v1.9.4 acceptance) and is the root cause of the
per-boot "WARN: pi-extensions install.sh failed" (fixed at the source in
pi-extensions 143a214; this is the belt to that suspender). Verified live
in a v1.9.3 container: add -> rev-parse prints 25c1265, unset -> fatal
again; /etc/gitconfig restored to its 5 lines afterwards.
3. docker-publish.yml: `find -print0 | LC_ALL=C sort -z` in base-decide.
sort collates per locale; identical rootfs hashed to base-f40c4b7b103d
under C/C.UTF-8 (== run 695) and base-d8df62216a81 under sv_SE/en_US.UTF-8.
The runner exports LANG=C.UTF-8, so the hash was stable by accident. Pin is
hash-neutral: pinned sort + HEAD inputs reproduces f40c4b7b103d exactly.
Per-command prefix only; nobody's locale changes.
CHANGELOG: new `## Unreleased` above v1.9.4 naming pi-extensions 143a214
(check 9 rc=0), the three fixes, and the synlig 3.10.0 hub upgrade that
v1.9.4 listed as still open. One invented URL org (gwpl) caught before commit;
upstream is elpapi42, as Dockerfile.variant:151 says.
Gates: check-doc-drift rc=0, check-base-hash rc=0, lint-shell rc=0 (16 files),
hadolint 2.15.1 rc=0, YAML parses (10 jobs), bash -n rc=0.
e3b38cd named "the feeder against the 3.9.0 hub" as the one path to exercise
before tagging, on the reasoning that 3.10.0's CLI writes now follow the daemon
write-routing policy. That was a reading of the mempalace changelog, not a
measurement of this image's feeder. Measured now:
mempalace-pi-session in remote mode (MEMPALACE_REMOTE_URL set, which is how
every fleet devbox runs) stages transcripts locally in python, rsyncs them to
the hub host, and calls the hub's own `mempalace_mine` MCP tool over HTTP
(run_remote_mine). The local `mempalace` CLI is required only in local mode
(line 475) and invoked only on the local branch (line 1029). With a PATH shim
logging every `mempalace` invocation, a 47-session `--dry-run` from this
container logged ZERO calls; the shim's positive control logged one. The pi
extension likewise speaks HTTP to the hub and spawns no local mempalace-mcp.
So the 3.10.0 CLI never runs against the 3.9.0 hub from this image. In remote
mode the client pin touches first-run `mempalace init` and the on-disk layout,
nothing else -- which is exactly what the ENV MEMPALACE_CONFIG_DIR change is
for. Both the Dockerfile.base comment and the CHANGELOG paragraph now say so,
and the "Still open" item becomes the thing that is actually unmeasured: a
first-boot acceptance of the built image from an empty ~/.mempalace volume.
Gates re-run on this tree: check-doc-drift.sh rc=0, lint-shell.sh rc=0,
hadolint 2.15.1 rc=0. lint.yml for e3b38cd: run 693, completed, success.
Three pins move or are deliberately held; the CHANGELOG's Unreleased section
becomes the v1.9.4 entry in this same commit, because since ee6cb9e the tag
build runs check-doc-drift.sh check 9 and every would-bake value must be named
above the last published heading.
mempalace 3.9.0 -> 3.10.0 (Dockerfile.base). Deferred at v1.9.3 for two
"Upgrade notes" items; both re-measured against the 3.10.0 wheel:
- New installs resolve ~/.config/mempalace. The order is $MEMPALACE_CONFIG_DIR,
then ~/.mempalace IF it holds config.json / people_map.json /
palace/chroma.sqlite3, then XDG. An EMPTY ~/.mempalace fails that test, and
an empty ~/.mempalace is what a freshly mounted devbox-palace volume (or
entrypoint.sh's mkdir on a volume-less container) looks like at first boot.
Measured with a fresh $HOME and `uvx mempalace==3.10.0 init`: config.json
landed in ~/.config/mempalace, ~/.mempalace stayed empty, and
entrypoint-user.sh's `[ ! -d ~/.mempalace/palace ]` would fire every start.
Same run with MEMPALACE_CONFIG_DIR set: everything in ~/.mempalace.
=> ENV MEMPALACE_CONFIG_DIR=/home/${USER_NAME}/.mempalace, placed in the
non-root-user section where USER_NAME is in scope, and entrypoint-user.sh
reads PALACE_DIR from the same variable (old path as fallback). Existing
volumes were safe either way; this is for first boots. MEMPALACE_PALACE_PATH
is still honoured (config.py:927), so smoke-test.sh's stage test holds.
- event_list defaults newest-first without a cursor. Server-side: the
extension talks to synlig over MEMPALACE_REMOTE_URL, so it lands when the
hub upgrades. mempalace-toolkit 2167a1b (floating main) makes every
cursor-less call say order:"desc" so the mailbox reads the same window on
either server version.
Corrected a stale sequencing comment while there: synlig serves the palace as
a `uv tool` under systemd (mempalace-serve.service; measured 2026-09-22:
mempalace 3.9.0, python 3.12.13, chromadb 1.5.9, no palace container), not via
docker-compose.mempalace.yml as the comment had said for two releases.
pi-atelier v0.10.1 -> v0.10.3 (Dockerfile.variant). Fixes only (both tags
released 2026-09-22); package.json at v0.10.3 still has zero runtime deps and no
build script, peerDeps still >=0.84.0. Annotated tags: the CHANGELOG names the
peeled SHA ed3837b, which is what resolve-versions bakes and check 9 compares.
pi HELD at 0.85.1 while 0.86.1 / 0.87.0 exist. 0.87.0 removed the
shouldStopAfterTurn agent option; pi-observational-memory 3.1.4 still uses it
and AgentContext.systemPrompt in observer/reflector/dropper (measured: 1 hit
per worker in src/agents on both the baked 3.1.3 tree and upstream master;
finishTurn 0). peerDeps are `*`, so it would break at runtime, not install.
Upstream #82, fix PR #83 open and unmerged. Bump pi and pi-obsmem together
once #83 ships. The other extensions are clean against the 0.86/0.87 lists.
Verified locally the way CI does: check-doc-drift.sh rc=0 with check 9
reporting all four moved components as named above the v1.9.3 heading;
lint-shell.sh rc=0; hadolint 2.15.1 (CI's pin, with .hadolint.yaml) rc=0 on
both Dockerfiles; PyPI 3.10.0 present, not yanked. Two claims caught by
re-measurement before commit and corrected in the text: a pi issue number
that does not exist on the 0.87.0 changelog line, and "8 shouldStopAfterTurn
hits" that was 3 (one per worker) in src/.
Still open before the tag: run the feeder from the built image against the
3.9.0 hub (dry-run first); recreate acceptance must show ~/.mempalace/palace
still passing and no ~/.config/mempalace appearing. After acceptance: upgrade
synlig's tool to 3.10.0.
lint-gate already existed to enforce "do not RELEASE a tree whose lint failed",
because lint.yml does not run on tag pushes. scripts/check-doc-drift.sh had the
same gap and it was never extended to cover it: check 9 ran only in lint.yml, so
no tag build has ever evaluated it. Add it to lint-gate, which resolve-versions
already needs, so it fails in ~8 s ahead of the 46-minute base build.
For this check the gap is strictly worse than it is for shellcheck. Shellcheck
judges the tree, so green on main is still green at the tag -- the bytes did not
move. Check 9 judges the tree against upstream NOW, and the floating refs it
watches move with no commit here at all, so a green reading on main carries no
information about tag time. v1.9.3 is the worked example: pi-observational-memory
moved cba0334 -> e7d77dc the day AFTER the tag, nothing went red, and it
surfaced only because someone ran the gate by hand.
Measured, not assumed:
- Same file lint.yml calls (one reference in each workflow), not a second copy.
- Works on a CI-shaped checkout: cloned --depth 1 --no-tags (0 tags, 1 commit),
rc=0. It needs no local tags because `last` comes from the Hub tags API, not
`git tag`, so the plain actions/checkout@v4 above is sufficient.
- Has teeth: deleting the Unreleased section from that clone gives rc=1 and
names the component; restoring it gives rc=0.
- ~8 s (7.8-8.3 s measured), vs ~1 s for lint-shell.sh.
Residual, accepted: the gate resolves the refs seconds before resolve-versions
resolves them again, so an upstream push inside that window still slips past.
Check 9 on the next release names it then.
check-doc-drift.sh was rc=1: pi-observational-memory moved cba0334 (3.1.3) ->
e7d77dc (3.1.4) upstream on 2026-09-20, the day after the v1.9.3 tag, and
nothing named it above the v1.9.3 heading.
No code change is needed or wanted here. PI_OBSMEM_REF defaults to master in
Dockerfile.variant and CI's resolve-versions turns that into a SHA at build
time, so the next rebuild adopts this whether or not anyone acts -- which is
precisely why it has to be written down. There is no pin in this repo to bump,
so the CHANGELOG is the only record a reader of the next tag would have.
v1.9.3 itself has no defect: its manifest, image labels and baked
/opt/pi-observational-memory clone all agree on cba0334. The drift is
post-tag, so a green drift reading taken on 2026-09-19 was correct when taken.
Measured, not assumed: the range cba0334...e7d77dc is 4 commits whose only
substance is one line in src/agents/worker-stream.ts (preserve registry
receiver when resolving stream, PR #78 / issue #77) plus a 19-line regression
test; the rest is the version bump and release merge. peerDependencies are
unchanged (all four @earendil-works/* peers still *) and there is no engines
block, so no pi or node floor to clear.
Gate now rc=0: "OK pi-obsmem cba0334 -> e7d77dc since v1.9.3, named above the
v1.9.3 heading".
CI checks out the tag, so the CHANGELOG heading has to name the version
before the tag exists (same convention as v1.9.1/v1.9.2). check-doc-drift
check 9 was sabotage-tested for exactly this rename and still passes:
everything above the published v1.9.2 heading is the pending text.
Contents of v1.9.3 (all measured, nothing inherited):
- f25efa0 recreate-sanity-check resolves the ControlPath ssh actually
uses (ssh -G); the ~/.pi/ssh/config hyphen typo deleted
- b8d818e Hub size claims corrected; check 8 gates them against full_size
- c7d369f promote-base-latest's conditional re-tag measured (run 669)
- 50153e6 skill floor -> pi-extensions 25c1265 (task tool + fork-gate)
- ea89605 changelog for the two floating-ref changes: pi-extensions
25c1265 and mempalace-toolkit 817b3a8 (mine deadline)
- cb6d9e5 check 9: what the next build bakes differently must be named
in the CHANGELOG; first run found pi-obsmem 7b397f4->cba0334
- b9057fd LABEL se.jordbo.pi-devbox.mempalace-version (inherited)
- 960aada dependency audit 2026-09-19; mempalace 3.10.0 deferred
Floating refs this tag will bake, named per check 9: pi-toolkit 9c87ee8,
pi-extensions 25c1265, mempalace-toolkit 817b3a8, pi-observational-memory
cba0334 (3.1.3). Base REBUILDS (Dockerfile.base + rootfs/ changed), so the
BASE_REBUILD_DATE marker is refreshed now, when it is free -- it had been
stale since 2026-07-13 through two rebuilds.
Deliberately NOT in this tag: mempalace 3.10.0 (see the audit section for
the three measured preconditions).
Every row by direct command: npm/PyPI/GitHub release APIs, git ls-remote,
the v1.9.2 config labels via the registry API, and binaries in a running
v1.9.2 container for the floating tools. Only pin with a newer upstream is
mempalace 3.9.0 -> 3.10.0, and it is NOT a small bump: new installs move to
~/.config/mempalace (entrypoint-user.sh tests ~/.mempalace/palace and
entrypoint.sh persists ~/.mempalace, so a fresh container would initialise
outside the volume), and mempalace_event_list flips to newest-first by
default (the toolkit's deriveClosed passes no order, deriveOwed on 2 of 3
calls — server-default-dependent). Deferred to its own release with the
three measured preconditions listed.
Closes the blind spot check 9 (cb6d9e5) named: no label recorded the palace
pin, so a MEMPALACE_VERSION bump — the one component whose skew against the
shared central palace is fleet-wide — could ship without a CHANGELOG line.
In Dockerfile.base, not Dockerfile.variant, deliberately: the value sits next
to the ARG that defines it (a copy in the variant is one more pin able to
drift); labels are inherited by every image built FROM the base, so no
build-arg to plumb through the variant's four call sites; and inheritance
means the label states the pin of the base the image ACTUALLY built on,
which is the question when base-decide cache-hits an older base. Both
mechanisms measured on the published v1.9.2 config blob rather than assumed:
maintainer + image.source (set only in Dockerfile.base) are present on the
variant image, and pi-version=0.85.1 is an ARG expanded inside a LABEL.
Intent, like every se.jordbo.pi-devbox.* label; the manifest's
mempalace_version (read from the installed binary) stays the ground truth,
and smoke-test.sh now asserts label == installed core — the one way they
diverge is a base built with INSTALL_MEMPALACE=false or an off-pin install,
both invisible to a label-only check.
check-doc-drift check 9 gains the component (literal, against
ARG MEMPALACE_VERSION in Dockerfile.base); the label-key rule generalises to
"names ending in -version are the label itself". Until a release carries the
label it reports a counted SKIP, not OK — measured: "v1.9.2 carries no
se.jordbo.pi-devbox.mempalace-version label", summary says 1 SKIPPED.
Costs nothing extra: this Unreleased already forces a base rebuild
(50153e6 rootfs/ skill floor). check-base-hash unchanged (no new *_REF).
The two most recent Unreleased entries (pi-extensions 25c1265: task tool +
fork-gate; mempalace-toolkit 817b3a8: the mine deadline that never reached
the transport) reached this image through floating *_REF=main ARGs. Neither
produced a diff here, so nothing asked for a CHANGELOG line, and neither had
one until a reader asked. Same class as check 8 — a claim with no in-repo
anchor rots — and the hand practice that covered it (the per-release
"Dependency audit" table, Baked-in vs Upstream-now) is a someone-remembers
mechanism that had lapsed.
Method, no docker/crane/token: last published vX.Y.Z from Hub's tag list
(one request, now shared with check 8); its amd64 config blob via the
anonymous registry API carries one se.jordbo.pi-devbox.<name>-ref label per
component with the SHA the build-args baked. "Would bake" is resolved the way
resolve-versions does it: 40-hex ARG is itself, branch/tag via git ls-remote
with the peeled ^{} form preferred, pi-studio = highest semver tag (label on
<tag>-studio), PI_VERSION as a literal against the pi-version label. Nine
components, ~7.5 s. Unchanged: OK. Moved: the new value's 7-char prefix (tag
for studio, version for pi) must appear above the last published version's
"## " heading — Unreleased plus any not-yet-published section, which is what
the release commit turns Unreleased into, so the tag build passes on the same
text. A published tag with no heading is a failure, not a skip.
First run found a move nobody had recorded: pi-observational-memory 7b397f4
-> cba0334 (6 commits, 3.1.1 -> 3.1.3; memory workers now routed by the
model's exact provider instead of the first registered provider with a
matching api — upstream #70). Named in the CHANGELOG with the honest scope:
no expected effect on the shipped bedrock-only configuration, unmeasured.
Sabotage-tested, expectation written before each run:
obsmem SHA removed from CHANGELOG ............ rc=1 DRIFT
Unreleased renamed to "## v1.9.3 — …" ........ rc=0 (release-commit shape)
"## v1.9.2" heading mangled to v1.9.2-typo .... rc=1 — this one FAILED first:
\b accepted "v1.9.2-typo" as v1.9.2's heading; now (\s|$)
bare "## v1.9.2" (no date) ................... rc=0
ARG PI_OBSMEM_REPO renamed ................... rc=2 (blind gate must not pass;
read_arg calls are bare top-level assignments on purpose — nested in a
heredoc's $(...) the exit 2 is swallowed by cat)
offline (proxy to a closed port) ............. rc=0, 2 SKIPs counted
SKIP_REF_CHECK=1 ............................. rc=0, 1 SKIP counted
restored ..................................... rc=0, files byte-identical
GIT_TERMINAL_PROMPT=0 on ls-remote: a repo flipped private fails in 0.3 s as
a SKIP instead of waiting for a username until the job times out.
Header, lint.yml job comment and AGENTS.md all still said this gate "needs
no network" — false since check 8 (2026-09-14). Now: two classes, hermetic
1-7 and published-state 8-9 that SKIP loudly and counted when offline.
Considered and not added: compose <-> .env.example variable cross-check (the
four mismatches are commented-out lines, the mempalace-server compose file's
own variables, and entrypoint-consumed ones — it would fire on nothing
wrong); "documented tag exists on Hub" (check 8 already SKIPs by name; a
hard fail would misreport the tag-to-publish window).
Both land in the image through floating refs (PI_EXTENSIONS_REF=main,
MEMPALACE_TOOLKIT_REF=main), so neither shows up as a diff in this repo — the
changelog is the only place a reader of a pi-devbox tag learns that the
container's delegation and feed behaviour changed. Recorded under Unreleased
with the measurements: 5/5 real fork briefs redirected, live pi -p proof for
gate and tool, 2-of-6 → 6-of-6 on the toolkit's new timeout test.
Dockerfile.base: the "Stall protection" comment listed MEMPALACE_MCP_TIMEOUT_MS
as THE tool-call deadline; it now says the feed's mine carries its own.
Comment-only, but Dockerfile.base is hashed into base_tag; the rebuild was
already forced by 50153e6 (rootfs/ skill floor).
check-skill-floor.sh: OK, tree_sha256 9b85a633… matches the package at main (25c1265).
Forces one base rebuild (base_tag hashes rootfs/), which is also what bakes
extensions/task.ts and fork-gate.ts into /opt/pi-extensions via the floating
PI_EXTENSIONS_REF=main; install.sh symlinks every extensions/*.ts on start.
recreate-sanity-check.sh asserted `/tmp/sshcm exists with mode 700` and printed
a green tick while every ssh in the container died rc=255. It was right about
what it checked: the breakage was in the directory a CONFIG named, not the one
the image creates, and the old check could not see the disagreement between them.
Found on the v1.9.2 first boot on emb-7kj4vr4g. A durable ~/.pi/ssh/config,
hand-written into the ~/.pi named volume by the previous session so it would
survive the recreate, declared `ControlPath /tmp/ssh-cm/%C` — with a hyphen.
Nothing here creates that path; the canonical spelling is /tmp/sshcm, identical
in Dockerfile.base, entrypoint-user.sh, recreate-sanity-check.sh and
smoke-test.sh. ControlMaster auto with an unusable ControlPath does not degrade
to an unmultiplexed connection — it fails hard:
unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory
rc=255, remote command never runs.
Now: resolve the ControlPath ssh itself would use via `ssh -G` and require its
parent to exist and be writable. -G applies real config precedence (first-value-
wins, the system drop-in, Include, an -F override), so it answers "which rule
captured this host" instead of re-implementing the guess, and it never opens a
connection — 0.116s for 48 hosts.
This also puts a check under a caveat documented in prose in Dockerfile.base
("SSH client defaults") and verified nowhere: a per-host ControlPath under a
read-only bind-mounted ~/.ssh gives the identical failure, "Read-only file
system". On the machine this was built on that is 16 of 50 hosts — freeipa-1..6,
gitea.egl.lan, runner-1..3, tor-ms22 — never reported by anything before.
Severity split is deliberate. Default ssh precedence legitimately lands in the
read-only ~/.ssh for any host whose own config pins it there, and the supported
workaround (ssh -F ~/.ssh-local/config, generated every start by
setup-lan-access.sh) already exists, so that is a warn. Failing it would paint
the script red on every run of every device, and a check that fires benignly
every time is one you learn to ignore — the same reasoning that keeps
lint-shell.sh at -S error. The sidecar route is prescribed, so there it is a
hard fail. Host lists cap at six names plus a count: unreadable output is
ignored output.
Teeth proven both directions, each sabotage confirmed by diff BEFORE the result
was believed: ControlPath -> nonexistent dir => rc=1; -> existing-but-read-only
dir => rc=1; restored => rc=0 with the sidecar byte-identical.
No config was added to fix the original problem — the fix was to DELETE
~/.pi/ssh/config, a third hand-maintained copy of what setup-lan-access.sh
already generates from version control on every container start, with fewer
features and one typo. The host-owned ~/.ssh/config is left alone on purpose:
those ~/.ssh/cm paths are correct on the host, where ~/.ssh is writable.
Also lint-shell.sh: the SC2088 count in the severity-choice rationale said 19
and is now 20, with the reproduce command recorded. That command needs a `$ `
prefix — a comment whose first word is "shellcheck" is parsed as a directive,
and the malformed one tripped SC1072/SC1073 at severity error. The gate caught
it on the very commit that introduced it.
Closes the open question from v1.9.2's release verification. crane copy DOES
bump Hub's last_updated when it runs (run 669: digests differed, copy ran
17:10:04->17:10:08, base-latest.last_updated moved to 17:10). The
identical-digest path stays unproven by construction -- the job prints
'base-latest already current; nothing to promote.' and never copies -- so a
freshness assertion on the alias is conditional on base-decide's need_build,
not a blanket upgrade. Operational rule lives in the ci-release-watcher skill
(skillset c8034af); the pipeline property is recorded here.
DOCKER_HUB.md claimed ~1.1 GB for :latest while Docker Hub served 1.37 GB --
20% wrong, drifting quietly across eight releases, and POSTed to Docker Hub by
update-description every time. It is the first number a stranger reads about
this image.
Root cause is structural, not carelessness: every OTHER claim check-doc-drift.sh
guards is anchored to a build file (a pin, an ARG, a placeholder), so it cannot
rot without someone editing the thing it describes. Nothing in this repo states
the image size, so nothing measured it.
Corrected against measured full_size after v1.9.2 published (amd64/arm64):
:latest ~1.1 -> ~1.23 GB (1.228 / 1.211)
:latest-studio ~1.15 -> ~1.25 GB (1.255 / 1.238)
:base-latest ~1.0 -> ~1.17 GB (1.167 / 1.151)
full_size tracks the FIRST manifest entry (amd64), not the sum across arches --
measured: full_size=1.228, amd64=1.228, arm64=1.211, sum=2.439 -- which is what
the table's per-arch "Size (compressed)" column claims.
New check 8 gates those claims against Hub. Fails on drift; SKIPS LOUDLY via a
new skip() helper (counted, named in the summary) when curl/python3 are absent,
the API is unreachable, or SKIP_SIZE_CHECK=1. A skip is deliberately neither OK
nor a failure: printing an unverified claim as OK is the habit this file exists
to break, and failing on Docker Hub's uptime would hold releases hostage to a
third party. Not wired into hooks/pre-push, so pushes stay offline.
Two bugs caught by writing the expected exit code down before running the check:
1. The tolerance would have missed its own motivating case. The percentage is
computed against the MEASURED size, but the 20% first chosen came from the
claim-relative figure; the real drift was |1.1-1.37|/1.37 = 19.7% and would
have passed. Now 15%, inside a window with both bounds measured: above the
largest legitimate skew (11.4%, a claim describing the published release
while the next tag changes the size) and below the rot it must catch.
2. A `|| true` on the python invocation made the gate FAIL OPEN -- it printed
DRIFT and exited 0. Removed; the outer `|| SIZE_RC=$?` satisfies set -e
without swallowing the code. Verified two-sided: 19% drift exits 1 at the
default tolerance and 0 at SIZE_TOLERANCE_PCT=25, so the threshold does the
work rather than the ordering.
NOT covered, and the script says so where a reader will see it: README.md's
~3.2 GB figures are UNCOMPRESSED and the registry exposes compressed sizes only,
so measuring them needs a real pull. A green check 8 says nothing about them.
Verified: check-doc-drift.sh rc=0 on the corrected tree, rc=1 on injected drift,
rc=0 with --warn-only, SKIP+rc=0 on the offline path (proxy to a dead port);
shellcheck clean at default severity (the sed single-quote SC2016 carries a
disable with its reason); lint-shell.sh 16 files clean.
CI checks out the tag, so the CHANGELOG heading has to name the version
before the tag exists, not after. v1.9.1's tagged tree carried no
"## Unreleased" either -- same convention, made explicit here.
Contents of v1.9.2 (all measured, nothing inherited):
- 1baba79 the +131 MB v1.9.1 residual: /root/.npm (110 MB) plus
@mariozechner/clipboard-* foreign natives at BOTH install
sites (global + /opt/pi-fork's nested pi-coding-agent copy)
- 852f900 smoke sentinels that name the residue instead of letting
the size gate's ~225 MB margin swallow it
- dab989b mempalace-toolkit: the owed-withdrawal suite gate
(node --check never type-strips)
- 9aaff26 vendored mempalace skill snapshot -> e9e45f7, which is what
turns the currently-red canary green
Deliberately NOT in this tag: DOCKER_HUB.md's "~1.1 GB" size claim is
low (Hub says v1.9.1 = 1.37 GB, v1.8.14 = 1.23 GB). This release
*changes* the size, so no number is right both before and after it --
writing the predicted ~1.24 GB would be an unmeasured claim in a
user-visible page. Measure post-build, fix on the next tag, and extend
check-doc-drift.sh to gate size claims against Hub's full_size so the
number cannot rot silently again.
Two things, and the second was found by doing the first.
v1.9.1's notes closed with an explicit open item: isWithdrawn had 17 assertions
and 4 mutation kills but had never been exercised on a released image against the
live logstream, which is a different state from untested and a worse one to leave
unlabelled. It has now been exercised, from a container recreated onto v1.9.1, and
the label is retired with the measurement rather than with an assurance.
What the entry records is the instrument as much as the result, because the result
is only as good as the thing that produced it: the shipped extractor was lifted
out of the toolkit's own test script and used to cut the four predicates out of
the BAKED mempalace.ts, and deriveOwed's three queries were replayed with their
exact shipped parameters. No second copy of the logic was written, which is the
same discipline the suite itself is built on. Baseline owed was established by
three independent routes BEFORE any probe was planted, because without a recorded
baseline a withdrawal that appears to work proves nothing. Predictions were
written down before each measurement; the table reports both columns.
The per-predicate attribution is in the entry deliberately. "The count went back
to baseline" is a claim about arithmetic and would have been satisfied by a
pre-filter or by an unrelated reply; isAnswered=false with isWithdrawn=true on a
candidate that was still in the raw set is a claim about mechanism. The negative
controls carry the same weight as the positive one -- a rule that lets a third
party or a prose sentence empty someone's mailbox would have passed a
count-only check.
Also recorded: the rule fires on the real incident it was written for (seq 112),
which is worth more than any synthetic probe, and the one path still unwitnessed
(the extension's own in-process poll, which only fires when the agent is idle) is
named as unwitnessed rather than folded into the pass.
Second: reaching for the suite as corroboration showed it cannot run on this image
at all -- its own precondition gate rejects the source it exists to check, because
`node --check` does not type-strip on any node and v1.9.1 moved to node 24. All 17
assertions had been skipped since the bump. Fixed in toolkit dab989b, named here
per the rule this repo adopted two commits ago after the withdrawal fix shipped
unnamed in v1.9.1. The gate now also distinguishes "cannot run" from "does not
parse", since collapsing those is how the defect pointed readers at the wrong
file.
The rebuild cost is stated and, having been measured, is smaller than the earlier
draft of this entry claimed: 9aaff26 already moved base_tag by refreshing the
vendored skill snapshot under rootfs/, so the base rebuild was pending before this
fix existed and the toolkit pickup rides along with it.
v1.9.1 bakes mempalace-toolkit e68ee20, which contains e2b060a. Requester-side
ask withdrawal has therefore been LIVE on every v1.9.1 device since 2026-09-10,
while the v1.9.0 section of CHANGELOG.md still read "not yet pinned ... this
image still pins e45f6b4", and the mempalace skill still told every agent at
session start that a withdrawal is impossible.
Measured, two independent routes, expectation recorded before looking:
- published image label, :v1.9.1-studio and :latest-studio (same digest):
se.jordbo.pi-devbox.mempalace-toolkit-ref = e68ee2071ca3ad39...
- ancestry: e2b060a is an ancestor of e68ee20
- baked mempalace.ts sha 7c16fe14 != v1.8.14's dfca71e9 (different bytes)
- grep -c isWithdrawn on this container's baked copy = 0 (v1.8.14, so
tor-ms22 cannot exercise the behaviour it is documenting)
The first label read came back EMPTY, and that empty was a claim about the
request rather than the image: Hub redirects blob fetches to a CDN and curl
without -L returns 0 bytes at exit 0. Recorded in the notes, because a registry
audit reporting "no labels" is missing -L until proven otherwise.
Changes:
- CHANGELOG Unreleased: the floating-ref mechanism, which is the reusable
part. ARG MEMPALACE_TOOLKIT_REF=main + CI resolving it to a SHA at build
time means a release absorbs whatever toolkit main holds, and "what
behaviour did this image gain" is a question nobody is forced to answer.
The fleet rule (name the pickup before tagging) was honoured for the
feed-tick commits v1.9.1 names, and missed for one in the same range.
- CHANGELOG v1.9.0: the stale caveat is ANNOTATED, not rewritten. The
sentence is the evidence for how the drift happened; deleting it would
destroy the only trace. Same discipline as fleet-ops d42412c.
- CHANGELOG Unreleased: corrected the size entry's "needs no base rebuild"
claim. True of that entry alone, false of the release now that the
vendored skill snapshot -- a base_tag input -- moves with it (~67 min).
- rootfs mempalace skill snapshot 44472 -> 46045 B, refreshed via
scripts/vendor-mempalace-skill.sh so the bytes and SKILLSET_SNAPSHOT_REF
(4d7c0ea -> e9e45f7) move together; --check confirms exact match.
- scripts/smoke-test.sh canary RE-PINNED. The retired pair was still green
against the new snapshot, i.e. blind to this refresh for the same reason
the pre-v1.8.13 pair was blind to that one. The replacement is stronger
than any predecessor here because BOTH witnesses come from the same
upstream commit: e9e45f7 added "Withdrawing an ask you sent" and deleted
"nothing anyone can do about it from the other end", the sentence the new
bullet contradicts. Directions measured against both files, then the
canary body EXECUTED against each: new -> rc=0 "ok", old -> rc=1 empty.
A canary whose negative witness was removed by the commit it pins fails
loudly on stale bytes instead of merely failing to notice them.
Gates: check-doc-drift OK, check-skill-floor OK, vendor --check OK,
hooks/pre-push OK (16 shell files clean at severity error).
Not fixable here: isWithdrawn is deployed and UNPROVEN. 17 assertions and 4
mutation kills, never once exercised on a released image against the live
logstream. Ask routed to a v1.9.1 device.
The post-boot check v1.9.1 left for the next machine was
node -e 'require("esbuild").transformSync("const x:number=1",{loader:"ts"})'
with "if this fails, the prune removed something needed on arm64 -> revert to
v1.8.14 and fix" attached. Run from /workspace on the freshly recreated v1.9.1
image it fails with MODULE_NOT_FOUND — on a perfectly healthy image. require()
resolves by walking up from the CURRENT DIRECTORY, esbuild is nested inside the
two pi trees, and a global install is not on node's require path (NODE_PATH is
unset). The version that "verified the prune" last night only passed because the
shell happened to sit inside the tree. So the check was CWD-dependent, its red
meant nothing, and the remedy it prescribed was reverting a good release.
Path-qualified, both sites compile TS today. Rather than leave a fragile command
in a note for a human to paste, CI now owns it:
- esbuild must transformSync at EVERY install site found in the image
- @mariozechner/clipboard must load with its native binding attached at every
site — for the clipboard prune that IS the proof, since napi-rs resolves the
platform package at require() time
Sites are discovered with find, so the studio variant's third site is covered
without naming it.
Verified as a four-way matrix, not just a green run: both assertions pass
against the real image FROM /workspace (the cwd that broke the old command), and
both fail (exit 1, "The package \"@esbuild/linux-arm64\" could not be found, and
is needed by esbuild") against copies of the same packages with the host
platform binary removed. An assertion never observed failing is not evidence.
Also documents both failure shapes in AGENTS.md next to the size gate: this one,
and the mode-700 /root one where `test ! -d` passes on a permission error.
v1.9.1's @esbuild prune fixed the 431 MB size-gate failure but still shipped
+131 MB compressed over v1.8.14, nearly all of it in the pi/extensions install
layer (87 -> 206 MB). That leftover was filed as an open item with an explicit
hypothesis — the @mariozechner/clipboard-* family, same npm 11 behaviour, a
different package — and an explicit warning that the hypothesis was not a
measured cause. Measured now, after recreating onto v1.9.1, and the hypothesis
accounted for one sixth of it:
+110 MB /root/.npm/_cacache (35.2 -> 145.3 MB) the build's npm cache
+ 21 MB clipboard foreign platform packages, both install sites
= 131 MB i.e. the whole delta, no unexplained remainder
Method, since there is no docker CLI inside the container: pulled both variant
layer blobs straight from the registry with a token + manifest + blob fetch and
listed the tarballs (29 789 vs 29 889 entries, 270.8 vs 401.9 MB uncompressed),
then aggregated per package. The file COUNT barely moved, which is what said
"few large files", not "npm installed more packages".
- purge_build_caches: npm cache clean --force + rm -rf /root/.npm, in the SAME
layer as the installs, in both the main RUN and the studio RUN. npm 11
caches every platform tarball it downloads, including the ones the prune
then deletes, so the cache grew faster than the tree. Nothing at runtime
reads it: build is root, container is developer with its own cache in $HOME.
- prune_foreign_esbuild -> prune_foreign_natives: now covers both MEASURED
families. Clipboard keeps linux-$arch-gnu AND -musl because its napi-rs
loader picks between them at runtime via its own isMusl() probe; the musl
package is a 420-byte stub. The bare @mariozechner/clipboard wrapper has no
hyphen suffix and cannot match the pattern.
Verified on arm64 before writing the glob — a widened rm -rf against a tree you
cannot inspect is the one change shape not to write blind, which is why the
order was update-then-patch. Exercised against a copy of the real trees with
foreign dirs fabricated back in (aix-ppc64, android-arm64, darwin-arm64,
win32-x64, linux-x64): all removed, host linux-arm64 kept at both sites, 21 MB
freed, require('@mariozechner/clipboard') still loads and exports all 18
functions, esbuild.transformSync still compiles TS at both sites. v1.9.1's own
arm64 validation also passed here; CI could only smoke amd64.
Two sentinel assertions, because the size gate did not catch this: it has
~225 MB of margin, so 131 MB of residue stayed green. Both were verified RED
against the running v1.9.1 image and GREEN against a pruned tree. The cache one
refuses to run as non-root: test ! -d /root/.npm on mode-700 /root would
otherwise pass for the wrong reason. Size-failure diagnostics now list cache
paths too — they previously enumerated only node_modules and /opt, where these
bytes were not.
Also fixed: -printf '%f\\n' reaches the shell with both backslashes (confirmed
from the published image's recorded created_by), so v1.9.1's progress line
printed a mangled "li ux-arm64" — find emitted a literal backslash and tr ate
the n out of the name. Single backslash now.
Deliberately not purged: /tmp/node-compile-cache (1.3 MB). The manifest RUN
calls pi --version again, so deleting it earlier only relocates those bytes into
that layer — today's manifest layer is 128 kB precisely because it finds the
cache warm.
Dockerfile.base untouched, so no base rebuild: this rides the next release.
v1.9.0 was tagged but never published: smoke failed 90-passed/3-failed and
build-variant needs smoke, so nothing reached the registry. All three are fixed.
1. Size, 431 MB over threshold. Node 24 brings npm 11, which installs EVERY
@esbuild/<platform> optional binary instead of the matching one: 26 dirs,
284 MB per pi-coding-agent copy. Measured on pi-fork: npm 10.9.8 -> 165 MB
(exactly what v1.8.14 shipped), npm 11.19.0 -> 449 MB. npm 11 ignores the
os/cpu constraints AND --os/--cpu AND an npmrc carrying them, all measured,
so Dockerfile.variant prunes explicitly, keeping linux-$(node -p
process.arch) so one line is right on both arches. Pruned in the SAME layer
as each install, or the bytes survive in the earlier layer. Three sites:
global pi, pi-fork, pi-studio. Verified esbuild still transforms TS after.
2. om's node_modules assertion tested an npm artefact. om has zero runtime deps;
its node_modules held ONE file (.package-lock.json) and 20 empty scope dirs.
npm 11 stopped creating it. Now asserts the entry point pi actually loads,
read from package.json -> pi.extensions.
3. The skill-source annotation added in v1.9.0 ("baked (package copy)") broke
the assertion matching "baked$". Pattern now allows an optional suffix; which
copy shipped stays authoritatively asserted against the manifest + tree hash.
Also: a failed size check now prints the largest layers, largest directories and
an @esbuild sentinel, so this class attributes itself next time instead of
costing a CI dig plus a local npm bisect.
Threshold stays 3800 MB: it caught a real regression and raising it would have
thrown the signal away. No Dockerfile.base/rootfs change, so base-0fb1256c7f99
is reused and build-base is skipped.
Five doc claims had rotted by v1.9.0, all the same shape: a value written once
by hand, in a file nothing verifies, about a number that lives elsewhere and
moved. README pin table wrong on all three rows; a "Planned" section describing
something already shipped; DOCKER_HUB.md claiming Node v22 against Node 24.
DOCKER_HUB.md is why this is a gate and not a resolution to be careful: it is
PUBLISHED (update-description POSTs it as Docker Hub full_description on every
tag), it had gone eight releases untouched, nothing generates it, and it is read
from the TAG -- so the stale page shipped with v1.9.0 regardless.
scripts/check-doc-drift.sh: seven checks, all repo-local (no network, token,
image, or sibling clone). Exit 0/1/2 matching lint-shell.sh; a renamed ARG is a
red 2, not a green tick. Wired as a fourth lint.yml job so "the docs lie" is its
own red name.
Verified with 15 controls, including two false-positive controls: the first
placeholder check flagged README.md:900, a Go template in a legitimate
`docker inspect --format` example. The gate was wrong, not the doc, so the
pattern is now anchored to the UPPER_SNAKE convention CI substitutes.
Not gated, deliberately: counts/sizes needing a running image (they belong in
smoke-test.sh -- a guessing gate is worse than none), and Dockerfile.base
BASE_REBUILD_DATE, because base_tag hashes that file content-wise and demanding
it be current would force a ~60 min rebuild on releases that touch no base
files. Free during a rebuild, expensive otherwise.
AGENTS.md step 3 rewritten around the mechanism: checkout@v4 with no ref: means
every job reads github.ref, the tag. Docs must be right BEFORE tagging.
DOCKER_HUB.md is hand-maintained (no generator: CI only substitutes
{{PI_VERSION}} and POSTs the file as Docker Hub full_description), and it was
last touched at v1.8.6 -- eight releases ago. The Node claim is now false as of
this release, and unlike README.md this file IS published, so a stale claim here
is user-visible rather than internal.
Verified countable claims rather than assuming: "7 user-facing extensions" is
correct (7 files in pi-extensions/extensions/). The "29 mempalace_* tools" claim
is suspect -- this session sees 45 -- but left alone because I cannot attribute
that count to the baked 3.9.0 server without measuring it.
The version-pin table existed precisely to be the reviewable record of what is
deliberately frozen, and it was wrong on every row: pi 0.84.4 -> 0.85.1,
pi-atelier v0.10.0 -> v0.10.1, mempalace 3.8.0 -> 3.9.0. Verified by parsing the
table and comparing against the ARGs it names rather than by eye.
Also: the "Planned for an upcoming minor release" section listed typst PDF
export, which shipped long ago and even carried a self-contradicting
"(shipped in Unreleased/base)" marker -- the fourth instance of the stale
in-repo Unreleased-pointer class this CHANGELOG already documents. typst 0.15.1
confirmed live in the running image, so the item is now stated as current fact.
The pi-devbox-version sample was v1.5.0-era and structurally outdated: it
predates the palace line the surrounding prose advertises, the pi-atelier
component, and the whole skills: block. Replaced with real observed output
rather than hand-written text.
README has no CI coupling (no workflow or gate reads it; the Docker Hub page
comes from DOCKER_HUB.md), so this cannot affect the in-flight v1.9.0 build.
Rename the Unreleased changelog section to its release heading, matching the
established `## vX.Y.Z — YYYY-MM-DD` format (em dash), per AGENTS.md release
step 3.
Minor rather than patch: CHANGELOG.md scopes minor to "significant base
additions", and this release bumps the Node runtime under every baked JS tool
(pi, agent-browser, playwright, mempalace) from 22 to 24 — the first
NODE_VERSION change since it was introduced at v1.0.0 — alongside five new
base packages (shellcheck, bind9-dnsutils, ldap-utils, xxd, python3-yaml).
Package additions alone have been patch here (v1.8.12, v1.8.13); the runtime
major is what lifts this one.
The recurring "[mempalace ext] feed (tick) failed: mine timed out after 30000ms"
message is fixed in mempalace-toolkit (309980b + e68ee20) and this image is what
delivers it, since CI resolves MEMPALACE_TOOLKIT_REF to a commit SHA at build
time. Worth a changelog entry rather than leaving it implicit in a ref bump: it
is the most visible symptom operators on this fleet have been living with, and
the entry records that it was a genuine defect (overlapping mines on a
single-writer palace) rather than the cosmetic annoyance it was parked as.
Closes the half deliberately left open by cac5e00's skill-floor gate, and the
more important half: "the floor is currently fresh" is a fact with a shelf
life, whereas "the image says which copy it got" keeps working.
The refresh in Dockerfile.variant is guarded by
`[ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates the
co-located skill keeps the vendored floor and still succeeds GREEN, with nothing
in the manifest, labels or logs separating that from a normal build. Afterwards
the two are indistinguishable by inspection -- same path, same filenames, same
permissions -- which is exactly how the floor went unnoticed from 2026-07-30 to
2026-09-10.
build-manifest.json gains pi_extensions_skill_source and
pi_extensions_skill_tree_sha256, MEASURED rather than passed as build-args, per
the ground-truth rule the surrounding block already follows -- and necessarily
so, since the outcome depends on the clone's contents and no ARG could express
it. Three values, because two would force a lie: package (served bytes equal
the clone's skill/), vendored-floor (clone had no skill/ at this ref), and
divergent (both exist but differ -- e.g. the clone ships SKILL.md but not
evaluate-extension-usage.py, so the served directory is a genuine MIX). No OCI
label mirrors these deliberately: LABEL cannot take a RUN-computed value, and a
label fed from an ARG would be the claim-not-measurement being removed here.
Two smoke assertions make the record a gate: the source must be named and be
`package` -- vendored-floor FAILS rather than warns, since these images track
main where the package has shipped skill/ since fa04d20, so a fallback means
the clone did not resolve as intended -- and the tree hash is recomputed over
the served directory, because a recorded hash never recompared is a claim.
pi-devbox-version annotates the line too: "baked (package copy)" normally, or a
yellow "(FALLBACK: vendored floor)". Its existing section reports which copy is
READ at runtime; this is the one fact decided at BUILD time and unrecoverable
later. Old images degrade cleanly -- field absent, jq // empty yields nothing,
line prints plain "baked" as before (verified against this v1.8.14 manifest).
Tested by running the exact logic against this container's real layout, with
the expected value written down before each: package (served == clone),
vendored-floor (clone path absent), divergent (clone lacking the .py while the
served dir has it), and null (empty served dir) -- all four as predicted. The
five pi-devbox-version render branches likewise, including the absent-field
case. Emitted JSON validated with jq for both the populated and null forms.
Gates green: lint-shell.sh (15 files), hadolint 2.15.1, actionlint 1.7.12,
check-base-hash.sh, check-skill-floor.sh, vendor-mempalace-skill.sh --check.
Audited every component the image obtains OUTSIDE debian/apt. Of ~23, the 19
that resolve `latest` at build time were already current or refresh themselves
on the next rebuild, and the hard pins for pi (0.85.1), mempalace (3.9.0) and
pi-atelier (v0.10.1) were already newest. Four needed a human.
NODE_VERSION 22 -> 24 (LTS "Krypton"). This was a latent defect rather than
housekeeping: agent-browser publishes engines.node ">=24.0.0", so the image sat
BELOW a declared requirement -- v1.8.14 shipped node 22.23.2 with agent-browser
0.37.1, so every build installed it with an npm EBADENGINE warning and ran the
baked browser automation outside its supported range. pi (">=22.19.0") and
playwright (">=20") are satisfied either way. Verified before bumping, since a
missing NodeSource suite breaks every arch at once: setup_24.x returns HTTP 200
and node_24.x advertises `Architectures: amd64 arm64 armhf x86_64`, covering
the arm64 fleet and the amd64 CI runners. Nothing else pinned the node major.
actionlint 1.7.7 -> 1.7.12 and hadolint 2.14.0 -> 2.15.1, each RUN AGAINST THIS
TREE at the new version before being pinned -- both clean, no new findings. A
linter bump is the one dependency update that can turn CI red on unchanged
code, so it is verified locally rather than discovered on a round trip.
SKILLSET_SNAPSHOT_REF e9e09d9 -> 4d7c0ea via scripts/vendor-mempalace-skill.sh,
never by hand: that script is the only thing permitted to write the ARG,
because a cp without a matching bump yields a manifest that confidently lies.
This proved PROVENANCE-ONLY -- the ref was 6 commits behind, but
skills/mempalace/SKILL.md is byte-identical at both (3675bfab), so the snapshot
was already correct and only its recorded origin was stale. No rootfs/ bytes
changed, the smoke-test phrase canary stays valid, and this ARG alone would not
force a base rebuild (the node bump does).
Two measurement traps worth recording, since both would have produced a wrong
answer: GitHub's releases/latest reports pi-atelier v0.10.0 as newest because
v0.10.1 is a TAG WITH NO RELEASE OBJECT -- the pin was already current, and
`git ls-remote --tags` is the instrument that shows it. And gitea-mcp is hosted
on gitea.com, not GitHub, so querying api.github.com returned nothing at all
rather than an error.
Verified with every gate this repo owns, all green, using the NEW linter pins:
lint-shell.sh (15 files), check-workflow-shell.sh, check-base-hash.sh,
actionlint 1.7.12, hadolint 2.15.1, check-skill-floor.sh, and
vendor-mempalace-skill.sh --check.
Follows ecfd2fc, which refreshed the stale floor by hand. A one-off refresh
fixes the symptom; this makes the drift impossible to reintroduce silently.
scripts/check-skill-floor.sh compares the repo floor
(rootfs/usr/local/share/pi-devbox/skills/pi-extensions/) against the package
repo it is a snapshot of, wired in as a new `skill-floor` job in lint.yml.
DIRECTORY hash, not `sha256sum SKILL.md`, using the same tree_sha256 pipeline
Dockerfile.variant uses for skillset_snapshot_tree_sha256 and for the reason
already documented there: a file-only compare answers "did this one file
change", not "is this the same skill". Verified by NEGATIVE CONTROL rather
than asserted -- with SKILL.md left byte-identical and only
evaluate-extension-usage.py edited, the directory check fails (rc=1) where a
file-only compare would have passed. Seven behaviour tests, each with its
expected rc written down before running: in-sync via local dir (0), in-sync
via anonymous remote clone (0), missing --package-dir (2), bad argument (2),
content drift (1), the sibling-file case (1), and --warn-only over drift (0).
Exit codes 0 in sync / 1 drift / 2 cannot-run, matching scripts/lint-shell.sh:
a gate that cannot run must not pass, so an unreachable package repo is a red
2 and never a green tick. A ref with no skill/ is NOT drift -- that is the
documented fallback -- but it emits ::warning:: because it is precisely the
condition under which the floor ships.
Gating on another repo is normally a smell. It is proportionate here because
the check can only fire when skill/ itself changed, which is exactly when the
floor has gone stale; pi-extensions commits that leave skill/ alone cannot
turn this red. It also needs no secret: pi-extensions is anonymously clonable
(verified with `git ls-remote` and no credentials), so it cannot start failing
when a token expires.
Also bakes python3-yaml (552 KB, zero extra deps) into Dockerfile.base. This
is the shellcheck story repeating exactly: scripts/check-workflow-shell.sh --
the guard against the Gitea sh/dash footgun that broke resolve-versions
(ed49b8d) and promote-base-latest (b7197e8) -- hard-exits with "python3 yaml
module missing", so a gate this repo already owns could not be run locally by
anyone. lint.yml installing it explicitly in CI was the evidence. Found while
wiring the job above: the guard could not be run before pushing.
CHANGELOG Unreleased updated for both this and ecfd2fc, including an explicit
note on what is NOT fixed -- the silent-fallback half still has no manifest
flag recording which copy was served.
Verified locally with every gate this repo owns, all green: lint-shell.sh (15
files clean), check-workflow-shell.sh, check-base-hash.sh, actionlint 1.7.7
(pinned, same version as CI), hadolint 2.14.0, and the new check itself.
Two changes that share one forced base rebuild, hence one commit.
1. THREE PACKAGES, each closing a capability gap measured during the
gitea.egl.lan/FreeIPA work on 2026-09-09..10 rather than a preference:
bind9-dnsutils (~6.1 MB measured) -- dig/host/nslookup were ALL absent,
so the container could resolve names but had no way to interrogate a
SPECIFIC nameserver. `getent hosts` only follows the resolver's default
path, so diagnosing "gateway 172.16.88.1 NXDOMAINs the egl.lan zone
while 10.20.253.1 is authoritative for it" had to be hand-rolled in
python3. Split-horizon DNS is a recurring class of bug on this fleet.
Note the package name: plain `dnsutils` is transitional in trixie.
ldap-utils (1244 KB, pulls nothing extra) -- the fleet authenticates
against FreeIPA, yet every LDAP probe had to be run by SSHing to an
already-enrolled host. Simple binds only; GSSAPI would additionally
need krb5-user + libsasl2-modules-gssapi-mit, deliberately not added
as that is a Kerberos-client decision, not a tool.
xxd (198 KB) -- convenience for verifying git-crypt blob magic in
myconfigs; `od -c` from coreutils already does the same job.
netcat-openbsd was in the original proposal and is deliberately NOT
here: measured redundant, because socat is already baked and bash's
/dev/tcp does reachability checks with zero packages (verified against
gitea.egl.lan:3000). Recorded in the Dockerfile so the omission reads
as a decision rather than an oversight.
2. ROOTFS FLOOR REFRESH: rootfs/.../pi-extensions/SKILL.md was 34284 B,
unchanged since fa04d20 (2026-07-30), while the canonical package copy
is 38973 B. Dockerfile.variant copies the fresh package copy over the
SERVED path at build time but never writes back to this floor, so the
floor is a silent fallback: if that build-time copy is ever absent it
ships the July skill with no log line or manifest flag to say which
version deployed. Refreshed from pi-extensions@c64c122, verified
byte-identical to both the canonical and the runtime-served copies.
Why one commit: the base_tag hash folds in `cat Dockerfile.base` AND
`find rootfs -type f | xargs cat` (.gitea/workflows/docker-publish.yml),
so either change alone forces the same full base rebuild -- and that
rebuild is precisely what re-bakes rootfs/ as it then stands. Emulating
the workflow hash with a fixed toolkit ref: f3d6462c7416 -> fc4edda03c54.
Verified: scripts/check-base-hash.sh passes (no new ARG *_REF added), and
no shell scripts are touched so the lint-shell gate is unaffected. Sizes
and dependency fan-out measured via apt-get --no-install-recommends
--dry-run on Debian 13 trixie.
v1.8.14 made shell lint a RELEASE gate (scripts/lint-shell.sh, shared by lint.yml
and the new lint-gate job that resolve-versions depends on), and that script
correctly exits 2 when shellcheck is absent -- "a gate that cannot run must not
pass". Measured on v1.8.14 on 2026-09-09 by three routes (command -v, dpkg -l, a
filesystem search): shellcheck was NOT IN THE IMAGE AT ALL. So the gate could not
be run by a developer in any container, only in CI, and the loop stayed
write-shell -> push -> wait for CI -> discover. That is the loop the gate was
added to shorten, after v1.8.14's first attempt burned ~46 min on a tree whose
lint had already been red for 24 hours.
shellcheck 0.10.0-1 added to the Dockerfile.base apt block: ~39 MB installed
(Installed-Size 40112 KB), measured to pull ZERO additional packages under
--no-install-recommends because libc6/libffi8/libgmp10 are already present.
NOTE this forces one full base rebuild -- base-decide hashes Dockerfile.base +
rootfs/, so unlike a scripts/ change it cannot reuse the existing base- layer.
hooks/pre-push is opt-in per clone (git config core.hooksPath hooks), bypassable
with --no-verify, and execs scripts/lint-shell.sh rather than reimplementing it
-- one copy, because a duplicated check that drifts is the failure this repo
keeps paying for. Matches the idiom skillset/ and myconfigs/ already use.
WHY THIS REPO HAD NO HOOKS, since it was reported as drift and is not: a peer
asked tor-ms22 for core.hooksPath per clone on the premise that unset meant the
gates were unverified there. Measured: pi-devbox unset, skillset hooks, myconfigs
common/hooks, pi-toolkit unset -- but `git ls-files | grep -i hook` is EMPTY in
both pi-devbox and pi-toolkit, so there was nothing to point at on any machine
and unset was the only correct value. This closes the real half for pi-devbox;
pi-toolkit still ships none.
Verified, expected result written down before each check:
* refusal paths -- shellcheck absent => rc 2 with the remedy named; linter
missing => rc 2. Never waved through on the assumption CI will catch it.
* the hook is IN the scan set -- "Checking 14 shell file(s)" with it present,
13 with it moved aside, so the extensionless file is found by the shebang
half of the linter's two-signal union. This check exists because the first
attempt was ambiguous: a planted `[ $UNSET_VAR = "x" ]` was not reported,
which could equally have meant "not scanned" or "below -S error". It was the
latter. A count that moves is unambiguous; a clean run is not.
* it catches the REAL v1.8.14 defect -- planting `echo 'the fleet\'s thing'`
in hooks/pre-push yields SC1073/SC1072 at severity error, rc=1.
And the gate earned its keep inside this commit: the first version of the
smoke-test assertion carried a comment beginning "# shellcheck is a GATE
DEPENDENCY", and a comment whose first word is the tool's name is parsed as a
DIRECTIVE, not a comment. The new gate failed it with SC1073/SC1072 before the
push -- same family as the v1.8.14 apostrophe, a line that reads as prose to a
human and as syntax to the parser.
v1.8.14's first attempt spent ~46 minutes building a base image for a tree whose
own lint had been failing for 24 hours. shellcheck had already flagged the
defect (SC2289, severity error) on the push that introduced it; the lint
workflow went red at run 186 and nobody read it.
lint.yml deliberately skips tag pushes and its reasoning is sound -- the tagged
tree was already linted on main, and a tag-ref lint run sorts above the publish
run, making a release look finished before anything ships. The missing invariant
was never "lint the tag". It was "do not RELEASE a tree whose lint failed", and
only a job inside the publish workflow can enforce that.
So: extract the shell-lint logic from lint.yml into scripts/lint-shell.sh and
call it from both places, then add a lint-gate job that resolve-versions depends
on. resolve-versions is the graph root, so gating it gates everything. Cost is
~40 s at the front of a release; the alternative already cost fifty minutes.
Extracted rather than copied on purpose. A second copy of a check is the drift
this repo keeps paying for -- the same evening produced a skillset mirror that
had sat 9579 B behind its upstream through two consecutive edits.
The script adds one behaviour the inline version lacked: if shellcheck is not
installed it exits 2 rather than silently finding nothing, inheriting the
existing "a gate that cannot run must not pass" rule from hooks/pre-commit in
the skillset repo. Without that, reordering the install step away would turn the
gate into a green tick over zero checks.
Verified locally with a stubbed shellcheck (the real binary is not in the
devbox), five cases, each with its expectation stated first: absent shellcheck
-> rc=2; stub pass -> rc=0 and a non-zero file count; stub fail -> rc=1; a
deliberately unterminated `if` planted in scripts/ -> rc=1 via the bash -n half,
naming the file; removal -> rc=0 again. Discovery cross-checks against CI's own
number: the inline version reported 12 files, the extracted one reports 13, the
difference being lint-shell.sh itself. YAML re-parsed (10 jobs, was 9) with an
assertion that resolve-versions needs lint-gate, and the repo's
check-workflow-shell.sh guard still passes.
The agent-browser execution guard added on 2026-09-07 carried its explanation
INSIDE the single-quoted script body, and the explanation contained an
apostrophe ("the fleet\'s only recurring amd64 runtime proof"). Inside '...'
bash treats a backslash as literal, so \' does not escape the quote -- it CLOSES
the string. The body truncated at that point and the remaining lines were parsed
by the calling shell.
Consequences, both measured rather than inferred:
- exec_test received 12 arguments instead of 2 (verified two-sided: the fixed
tree yields argc=2, HEAD yields argc=12).
- the leaked `v=$(agent-browser --version)` ran on the CI RUNNER instead of
inside the image. The runner has no agent-browser, so smoke and
smoke-studio both failed with "line 770: command not found" after
build-base had already spent ~46 minutes. Every downstream job was skipped.
- the truncated body still passed inside the container and printed its green
tick first, so the log shows a PASS immediately followed by the failure --
the tick was real, it just no longer covered the assertion.
The prose now sits above the exec_test call, where an apostrophe cannot
terminate anything, and a comment at that spot records why it must stay there.
Not a new failure class: shellcheck flagged it as SC2289 at severity error the
same day, so the lint job has been red since run 186 (2026-09-07 21:21) and was
not read. The gate did its job; nobody looked.
Converts the Unreleased section and records what this build carries beyond it:
the mempalace-toolkit bump that makes closing replies reach the mailbox
(deriveClosed, 21023e7 -> e45f6b4), and the L0-L4 subtask documentation landing
via pi-toolkit adfb553 + pi-extensions c64c122.
Notes the mechanism that makes the toolkit fix land at all — the resolved
toolkit SHA is folded into the content-addressed base tag, so the toolkit
moving forces a base rebuild rather than waiting for one — and the consequence
for the amd64 item already in this section: v1.8.13's base was cached, so
Dockerfile.base:607's agent-browser assertion never ran. This base is not
cached, so the native-amd64 proof is finally collected instead of discarded.
Unset means the mailbox is silent outside the pi TUI, and the reason is
structural: docker exec does not forward KITTY_WINDOW_ID/TERM_PROGRAM, so
'desktop' detection always falls through to OSC 777, which Kitty does not
implement — the notification then silently does nothing, the worst failure for a
feature whose only job is to break a silence. Documents the four modes, and that
MEMPALACE_MAILBOX_POLL_MS is a FLOOR BETWEEN activity-coupled polls rather than a
wall-clock interval (an idle session polls zero times) — the exact expectation
mismatch reported today.
Second instance of the same bug class as the node line, in the same file, found
the same way. The agent-browser guard captured the version inside an echo with
2>/dev/null:
echo "resolved=[$r] version=[$(agent-browser --version 2>/dev/null|head -n1)]" >&2
so the exit code was discarded and a binary that could not execute at all still
PASSED, printing version=[]. Verified two-sided: a stub exiting 127 passes the old
form and is caught by the new one.
Why this exit code matters more than most: smoke runs platforms: linux/amd64 on an
x86 runner, i.e. NATIVE amd64, so this line is the fleet's only recurring amd64
runtime proof for agent-browser's linux-x64 ELF.
NO DEVBOX CAN EVER SUPPLY THAT PROOF. Every machine in the pi fleet is an Apple
Silicon Mac: mbp-m1-2020; tor-ms22 = Mac Studio Mac13,1 M1 Max (fleet-ops
hosts/tor-ms22.md, verified 2026-08-17 with system_profiler); emb-7kj4vr4g =
Apple Silicon, verified 4 routes 2026-09-07. The open "amd64 runtime proof still
needed" ask sent to two devices was asking for the impossible, and emb's reply
naming tor-ms22 as "the only remaining candidate" is wrong for the same reason.
CI had the answer all along and was throwing it away.
Dockerfile.base:607 DOES assert it (`agent-browser --version && \`), but only when
the base rebuilds, and v1.8.13's base was cached — so smoke is where the recurring
gate belongs.
Summarises what changed since v1.8.13: the node-major assertion (a bump would
have passed the suite silently), the two-sided verification of the derivation,
and the v1.8.13 agent-browser 0.35.2 -> 0.36.0 correction. No image content
changes; NODE_VERSION still 22.
Two findings from a delegated read-only audit of this repo, both verified from the
filesystem before patching.
1. No test asserted the node major, so a node-24 bump would have passed the smoke
suite SILENTLY. scripts/smoke-test.sh:94 was a bare `run "node" "node --version"`
— exit-0 and non-empty output only, the printed version compared to nothing —
while the line above it uses run_expect against $EXPECTED_PI_VERSION for pi. A
reader skimming the suite would reasonably assume node regressions were covered.
Worse, this is where the "node v22.23.2 verified" line in the v1.8.13 recreate
notes came from: printed output, not an assertion.
Now gated on EXPECTED_NODE_MAJOR, which CI derives from Dockerfile.base's ARG
NODE_VERSION — the single source of truth (Dockerfile.base:557 is the ONLY hard
pin in the repo; Dockerfile.variant has no node install at all). That also
catches a stale cached layer whose node disagrees with the declared ARG.
Unset => previous behaviour, so this is backward compatible.
Verified two-sided rather than assumed: the sed derivation yields 22 (empty
would have silently disabled the assertion, reintroducing the bug); grep -Fq
"v22." matches v22.23.2; "v24." does NOT match, so a wrong major is caught; and
"v2." does not prefix-collide. Workflow YAML re-parsed after editing (9 jobs).
2. The v1.8.13 entry claimed "the image's own 0.35.2" for agent-browser. The image
ships 0.36.0: /usr/lib/node_modules/agent-browser/package.json says version
0.36.0, engines.node >=24.0.0, and no 0.35.2 exists anywhere in the image. The
claim was also internally incoherent, contrasting 0.36.0 against a version that
is not present. Corrected in place with a visible note, since the entry is
already released. The reasoning survives untouched: the engines floor really is
vestigial, because /usr/bin/agent-browser is a prebuilt aarch64 ELF invoked
directly and never through node — which is why 0.36.0 runs fine on 22.23.2.
Measured at the wrong layer during the v1.8.13 audit. I read `ARG
PI_STUDIO_REF=main` in Dockerfile.variant, concluded the release would adopt
main (= v0.9.60-rc.0), set PI_STUDIO_VERSION to that, and wrote a comment plus a
CHANGELOG entry describing deliberate RC adoption. A Dockerfile default cannot
answer "what will CI publish?" when CI overrides it, and it does: build-variant
passes PI_STUDIO_REF=studio_ref and PI_STUDIO_VERSION=studio_tag (lines 598-599
and 787-788), and resolve-versions picks the newest STABLE semver tag via
`^v?[0-9]+\.[0-9]+\.[0-9]+$`, which excludes pre-releases.
Caught by reading run 639's own resolve-versions output rather than the
Dockerfile: studio_tag=v0.9.59, studio_ref=9eed84f = refs/tags/v0.9.59^{}, while
main/v0.9.60-rc.0 is 658536f and never gets built. So published v1.8.13 studio
images carry pi-studio v0.9.59.
ARG restored to `none` rather than pinned to v0.9.59: the local-build default
should not hardcode a tag that goes stale as soon as main moves, which is how the
previous value came to lie. The comment now leads with the override so the next
reader starts at the layer that decides. Upstream's tag-over-main policy is
deliberate (Releases stopped at v0.5.55, main receives half-finished commits), so
adopting an RC from CI would mean changing that filter, not this ARG.
Consequence kept on purpose: the RC's opt-in Studio network binding is in NO
published v1.8.13 image, so it needs no audit this release.
Doc/label-only: base_tag hashes Dockerfile.base + rootfs/** + both entrypoints +
mempalace_toolkit_ref, none of which this touches, so the in-flight base build
(base-ad9faf00f2b2) stays valid and the tag run will reuse it. Verified with CI's
pinned linters: hadolint 2.14.0 exit 0 on both Dockerfiles, actionlint 1.7.7 exit
0, shellcheck 0.10.0 -S error exit 0.
Gitea (1.26.2) builds the "Run workflow" dialog from each input's `type:`.
With no type declared, the form renders a branch selector and NO input fields,
so a manual run silently takes every default -- and for release_tag: '' that
means env.RELEASE_TAG resolves EMPTY, the variant tag list becomes `<image>:`,
and the run dies on an invalid docker reference only AFTER paying the full base
+ smoke cost (~70 min). Net effect: the `smoke_only` escape hatch documented in
this file's own header has been unreachable from the UI for its entire
existence. Found 2026-09-06 while trying to use it to validate three new smoke
assertions before cutting v1.8.13.
Typed as `string`, deliberately, even though promote_latest/smoke_only read as
booleans: all six consumption sites compare strings against 'true'
(inputs.smoke_only != 'true' at both build-variant gates,
inputs.promote_latest == 'true' at both promote gates) or interpolate into
env.PROMOTE_LATEST. A boolean-typed input yields a real boolean, so `!= 'true'`
would compare across types and could invert a publish gate silently rather than
fail loudly. This keeps the change a pure rendering fix with zero semantic
delta; switching to boolean would require re-auditing all six call sites.
Validated locally with CI's own pinned tools before pushing, because lint is
the only gate on this file: actionlint 1.7.7 exit 0 (clean baseline before the
edit, clean after), shellcheck 0.10.0 -S error exit 0 across all 17 shell
files, and a pyyaml structural check confirming the three inputs still carry
string defaults, the `v*` tag trigger is intact, and all 9 jobs still parse.
Folded into v1.8.13 at zero marginal cost: the snapshot is hashed into
base_tag, but Dockerfile.base already changed this release, so the ~67 min base
rebuild was already being paid. vendor-mempalace-skill.sh --check reported exit
0 (stale-but-truthful) beforehand, so skipping was sanctioned -- this is the
deliberate call the release checklist asks for. Upstream content: the bare
project-name wing convention and the <harness>@<device> added_by rule, both
downstream of the attribution defect measured on this device 2026-09-06.
The canary re-pin matters more than the refresh. Its old pair ("Provenance is
stamped for you" present / "Attribute what you file yourself" absent) still
PASSED against the new snapshot, so leaving it would have yielded a canary
green on both old and new bytes -- blind to exactly the refresh it exists to
witness, the same false-green family as the pre-v1.8.5 canary. New pair chosen
by measuring direction against both files rather than reading the diff
("Diaries self-heal; plain drawers do not" new=1/old=0; "Agent diaries live in"
new=0/old=1), then tested two-sided: PASS on refreshed bytes, FAIL on the old
bytes recovered from git.
Gates after the change: smoke-test.sh parses, vendor --check exit 0,
check-base-hash exit 0.
Version audit for the next release. pi 0.84.4 -> 0.85.1, deliberately skipping
0.85.0 (it published internal experimental code and broke SDK imports,
upstream #9132). mempalace 3.8.0 -> 3.9.0. pi-atelier v0.10.0 -> v0.10.1.
PI_STUDIO_VERSION relabelled none -> v0.9.60-rc.0 so the floating main ref's
RC status is visible at docker-inspect time instead of discovered later.
PI_FORK_REF stays floating and adopts e69725c.
The pi bump was verified by running it under a pty in five combinations rather
than by reading the changelog, because this repo has already shipped a version
pair no changelog flagged (atelier < 0.7.1 hangs pi >= 0.84). CPU delta
0.00-0.01s over 5s against a ~5s sustained-CPU hang signature, two-sided via
the atelier sidebar painting identically to the 0.84.4 control.
NODE_VERSION stays 22 on purpose: node 24 is technically safe (pi's five
prebuilt addons are all NAPI, nothing declares a ceiling, agent-browser's
engines.node >=24 is vestigial for the shipped aarch64 ELF), but this release
already moves two minors and bakes an RC, and a node major would leave four
suspects if the image misbehaves. Own release, smoke suite as the gate.
Also corrects a stale claim at the mempalace ARG: synlig serves 3.8.0
server-side, not 3.7.1 (measured over ssh 2026-09-06).
agent-browser volume shadowing: the image has shipped 0.35.2, but every
session on mbp-m1-2020 ran 0.27.0 from a 2026-07-17 hand-install in
~/.pi/npm-global (a VOLUME, at PATH position 2 vs /usr/bin at 8). Third
package hit by this hazard after pi and pi-atelier, so the guard is now
generalised: entrypoint-user.sh retires the copy by moving it aside
(reversible, only when the image ships its own), recreate-sanity-check.sh
asserts resolution under /usr where the volume is real, smoke-test.sh carries
the build-time half and says in the source why it is weak. The real damage was
the stale BUNDLED SKILL (3 skillsets/17.6 KB vs 8/31.5 KB, ten subcommands
undocumented to the agent) - a stale tool errors, a stale skill quietly
teaches wrong commands.
pi-fork capability floor (extensions: []): forks were measured across four
dispatches ignoring their brief, answering in the user's voice, fabricating
self-referential measurements, and once filing a diary entry as agent_name=pi.
Cause is upstream by design - the child gets getHeader()+getBranch(), the
whole active session branch, with the brief as the final user message. Not a
model-capability problem: the same model as the fast profile obeyed the
identical brief perfectly with a fresh session and no inherited context.
extensions: [] runs children with --no-extensions, so the mempalace bridge is
absent and palace writes are impossible by construction (verified by asking a
child to enumerate its tools: read, bash, edit, write). Removes palace writes,
not filesystem writes.
§5 said embedding_metadata.string_value holds "metadata fields only". False,
measured directly: chroma also stores a copy of the document text there, under
key chroma:document. Confirmed with a disposable sentinel drawer (pi@tor-ms22,
2026-08-30): one row in fts_content AND one row in embedding_metadata for the
same drawer.
This was a real mistake in shipped guidance, not a nitpick: this section's own
scanning advice was written to guard against explaining a zero with a
mechanism nobody verified from source, and the section itself did exactly
that -- I downgraded a census to "a floor" on the strength of a metadata-blind
claim I never checked against chroma's actual storage layout. The practical
scan order is unchanged (fts_content is still the direct target, raw bytes are
still the backstop); only the stated REASON for a metadata zero changes: it
needs a different explanation now (key filter, query shape, escaping), not
"structurally absent".
§6's row-gone/bytes-gone claim is upgraded from asserted to measured, same
sentinel: delete_by_source took both fts_content and embedding_metadata 1->0,
raw bytes stayed 4->4 (freed pages persist until VACUUM). Also records the
method that unblocked the measurement: not a better instrument, a disposable
sentinel drawer instead of risking real fleet data.
No image behaviour changes.
v1.8.12's release retitled the previous Unreleased heading, leaving the file
with no place to put the next change — so the next contributor either invents a
heading or appends to a released section. The note under it points at the
release checklist step that renames it, so the convention is discoverable from
the file rather than only from AGENTS.md.
Also the first push after moving CI off synlig: lint.yml should now run on
runner-a1 (8 vCPU / 16 GB, Debian 13, upstream Docker CE) instead of the box
that hosts the palace.
pi 0.84.3 -> 0.84.4 (no Breaking Changes / Removed heading in that section,
grepped). Adopted for three fixes that land on machinery this fleet runs:
#6879 (large tool results crossing the auto-compaction threshold were sent to
the provider before compacting), #8345 (a resumed session corrupted its next
appended entry when the JSONL lacked a trailing newline -- that file is the
memory feeder's input; measured 49/49 clean here beforehand), and #8537
(triggerTurn:false messages sent mid-run were inserted between a tool call and
its result). The mempalace mailbox is outside #8537's precondition: it delivers
at agent_settled with deliverAs:"steer" and no triggerTurn, and 0.84.4 leaves
the documented steer semantics unchanged.
pi-atelier v0.8.2 -> v0.10.0: two minor releases, both UI-only, no BREAKING
notice. v0.9.0 raises its minimum pi to 0.84.0 and, unlike the
0.7.1-under-pi-0.84 startup-hang precedent, encodes it in peerDependencies
(>=0.84.0). Satisfied by PI_VERSION=0.84.4. Both executable floors compare with
sort -V, so 0.10.0 >= 0.7.1 evaluates correctly.
docs/observational-memory.md: pi's own compaction.md gained one paragraph in
0.84.4 -- autoCompact is now also checked mid-run, after a tool batch's results
are appended. Our text said compaction is checked only when pi goes idle and so
"never interrupts a turn"; that was only ever true of the OM trigger. The
section now states both entry points into session_before_compact and the
diagram carries the second edge (mermaid checker re-run: 6 blocks, 44 labels,
0 soft-wrapped, no cut glyphs at 1280px and 800px).
README: the version-pin table had been wrong since v1.8.6 -- 93f986e moved
ARG PI_VERSION to 0.84.3 and MEMPALACE_VERSION to 3.8.0 and neither table row,
so it advertised pi 0.84.2 / mempalace 3.7.1. Corrected, plus the
--expected-version example that would now fail against a 0.84.4 image.
CHANGELOG: Unreleased retitled v1.8.12 (2026-08-31) with the audits above and a
dependency-audit table -- every other component measured SAME (skillset
snapshot --check OK at a12fe5e, 0 commits since baked).
Docs only; no image behaviour changes.
WHY THIS AND NOT A PRIVATE NOTE. pi@emb-7kj4vr4g reported itself for printing
sha256[:8] fingerprints of GIT_USER_EMAIL, GIT_USER_NAME and HOST_SSH_USER, and
wrote a private rule forbidding it. It had not broken a rule. It followed §2 of
this skill as written, and §2 is incomplete: it says a fingerprint lets you
compare a credential "without ever materialising the secret" with no condition
attached. When two agents independently make the same mistake, the artifact that
taught them both is the bug.
§2 NOW CARRIES THE PRECONDITION. A fingerprint is 32 bits over its INPUT SPACE,
so publishing fp8(x) hands anyone a MEMBERSHIP ORACLE: they can test x == v for
every candidate v they can generate. Safe for a 40-char random token; a wordlist
for a hostname, username, e-mail, port, path, commit SHA or weak password. "High
entropy" is the usual sufficient condition, NOT the test — a commit SHA is
160-bit and still fully enumerable from the repo. Operationally: if you can
imagine writing the wordlist, you cannot publish the fingerprint. Also added:
candidate fingerprints are working memory and never output (an extractor hashes
hostnames and paths too, so the tempting "print what the scanner saw" debug step
leaks low-entropy fingerprints wholesale), and a plain statement that a
fingerprint register is a CONFIRMATION ORACLE for anyone already holding a
candidate corpus — which is exactly how a retired token is identified in old
transcripts, and works identically for someone else holding those same files.
NEW §6, "Proving absence: instrument strength, and four ways a scan lies clean",
placed next to §5 on purpose: §5 optimises against false POSITIVES, and every
failure in §6 is a false NEGATIVE. Triage optimises precision, a gate optimises
recall, and conflating them is what produced three clean reports over secrets
that were really there. Contents: instrument ranking (exact-byte value search >
class/structure pass > fingerprint census) with the instruction to state which
one produced your zero; census vs class passes as different questions, both
failure modes measured on this fleet; the tokenisation trap where quoting alone
decided detectability; scan the index or pushed tree, never the working tree;
git filters never run on symlinks while check-attr claims they do; two-sided
self-tests that abort, incl. the fixture-interaction artifact; row-gone is not
bytes-gone.
Attribution kept per finding: the census/class split and the instrument
ranking's provenance are pi@emb-7kj4vr4g's; exact-byte search over index blobs
is pi@tor-ms22's. The credential sense of "census" originated in this skill, not
with either agent.
TRAP FOR THE NEXT EDITOR, also in the CHANGELOG: the frontmatter description is
now 1022 of 1024 characters. Trim before adding, or the skill silently fails to
load. Verified by parsing the frontmatter (1022 chars, name intact, every prior
trigger phrase retained).
Deployment: baked skill -> needs an image rebuild AND a container recreate to
reach a running container.
v1.8.11 linked cli_utils' bin/ COMMANDS onto PATH and stopped there. Nothing ever
sourced cli_utils.sh, so its 14 FUNCTIONS were missing from every interactive shell
whose $HOME has no zsh rc -- which is the normal case, not an edge case: the
container's interactive shell is bash and zsh is not installed in the image. A
symlink cannot carry a shell function and a function cannot be reached from a
non-interactive shell, so the two mechanisms are disjoint and both are required.
The image was already paying this layer's dependency cost (fzf, bat, fd, rg, jq are
baked partly FOR these functions) while delivering none of its benefit.
The two changes ship together because they are coupled: portcheck is one of the 14,
and it was a hard stub in every image up to v1.8.11 -- neither ss nor ip nor lsof
nor netstat was present, so it printed "portcheck requires at least one of: ss,
lsof, netstat" and exited. Wiring the functions in without iproute2 would have
shipped a visibly broken one.
MEASURED, not assumed:
- the loader is bash-safe despite the *.zsh filenames: `bash --noprofile --norc`
exits 0, defines all 14, and they run (pathls, mkcd, up, extract, agents-sync,
fhist verified). The tree's one zsh-only construct (print -z in fzf/fhist.zsh)
is already guarded by [[ -n $ZSH_VERSION ]] with a bash fallback.
- fresh-$HOME seeding resolves 14/14; CLI_UTILS_SOURCE=0 is honoured; an absent
checkout is a genuinely silent no-op (no output, no leaked _cu).
- interactive shell startup 12 ms -> 17 ms.
- iproute2 is ~5.5 MB (4.2 MB itself + 6 libs under --no-install-recommends;
libpam-cap is a Recommends and correctly dropped). ss lands at /usr/bin/ss,
ip at /usr/sbin/ip, both already on the developer PATH, and `portcheck --all`
then correctly identifies the socat listener on 8765.
- hadolint clean on both Dockerfiles; repo-wide shellcheck -S error and bash -n
clean. .bash_aliases is outside CI's discovery (no shebang, not *.sh), so it
was checked by hand with -s bash at error AND warning level.
Named explicitly per this repo's floating-ref rule: /workspace/cli_utils is a HOST
BIND MOUNT, not a pinned ref, so the image now executes unpinned content in every
interactive shell. Errors are left visible rather than sent to /dev/null so that a
future zsh-only file in that repo is diagnosable rather than mysterious, and
CLI_UTILS_SOURCE=0 is the documented escape hatch. It is deliberately independent
of CLI_UTILS_LINK=0: the two disable independent mechanisms.
Deployment: needs a rebuild AND a recreate. $HOME is the container's writable layer
rather than a named volume (verified -- ~/.bash_aliases carries the container start
mtime while ~/.bashrc carries the image's), so the skel file is re-seeded on every
recreate; a host-bind-mounted ~/.bash_aliases is still never overwritten.
Claimed "## Unreleased is a new convention in this repo" after reading
CHANGELOG.md once — minutes after b615571 had renamed that very section to
`## v1.8.11`. 33 commits touch the heading; the convention is that new entries
land under Unreleased and the heading is renamed at tag time, exactly as it
was renamed out from under my snapshot.
Same root cause as the five instances the section above already records: an
absence observed in one frame, promoted to a fact about the world, with no
second measurement. `git log -S'## Unreleased' -- CHANGELOG.md` was the oracle
and costs one command.
The generalisation is worth more than the instance, so it goes in the habits
block: to learn a repeating PROCESS, read history, not the file. A file's
current content is one frame of a cycle, and the frame you catch may be the one
where the thing you are looking for has just been consumed.
Records the two skill commits ahead of tomorrow's build, and states the finding
that shaped them: the "a negative result is usually your own filter" rule was
already baked, already symlinked in at every container start, and already
survived every recreate — then was violated five times by a session that had it
available. The gap was activation, not persistence, which is why the
cross-cutting form went into the always-appended AGENTS block instead of into a
skill that only loads when a task description matches.
Also notes what the entry's own subject implies for the reader: neither change
reaches a running container until the image is rebuilt AND the container
recreated, since ~/.agents/skills and the global AGENTS.md both live in the
image rather than in a volume or a mount.
Carries the facts a two-day credential incident produced, not the discipline:
probe the issuer FIRST (11 of 13 "exposed" credentials were already dead at the
provider, which cost five HTTP requests to learn and was never checked), the
403-vs-401 trap that scoped tokens introduce into liveness probes, revocation
beats deletion for anything already replicated, the three places a secret hides
in a Chroma palace (FTS content, metadata, raw bytes) in coverage order, scope
derivation from measured consumers, and this fleet's age store with its
single-recipient weakness.
Facts transfer between sessions; exhortations do not — hence a separate skill
for the domain knowledge and a one-line pointer in the always-loaded block.
Authored here, so baked is canonical and it is NOT added to skillset-owned.txt.
Skill dirs are picked up by a glob in entrypoint-user.sh, so no registration is
needed — verified rather than assumed, since an enumerated list would have left
the skill inert, a fitting failure given its subject. Three smoke assertions
extended so a future rebuild cannot silently drop it.
pi-devbox-environment already warned that "a negative result is usually your
own filter" — baked, symlinked in at every container start, authored by an
earlier session. It survived every recreate, was available all of a later
session, and was violated five times anyway. So the gap was never persistence.
That section also closed with "a positive result needs no such scepticism — it
carries its own evidence." That is false, and it aimed the scepticism budget
one way only. Three of those five errors were positives:
- an SSH handshake SUCCEEDED and greeted me as joakimp while I believed I was
probing gitea.egl.lan — `Host gitea*` had rewritten HostName
- a 401 that was a genuine answer from an issuer which never minted the token
- a "regression" produced by diffing against a value my own -p 2222 flag set
Adds the three missing false-negative rows, replaces the false claim with the
"a positive result only proves what you actually asked" subsection, and records
the two habits that actually caught these: `ssh -G` to learn which rule
captured a hostname, and declaring the expected result before running a check.
The cross-cutting form goes in pi-global-AGENTS.append.md rather than in a
skill, because it has to fire without a task description matching it — being
loadable on demand is exactly what failed. Across all five errors, none was
caught by re-reading my reasoning; every one was caught by a second measurement
that disagreed.
Names three things riding in via floating refs, since the rule this
fleet adopted (v1.8.9) is to name a behaviour change before tagging,
not after:
- mempalace-toolkit b2b50af -> 21023e7: the rsync --checksum ship fix
(a real behaviour change to every remote-mode client) and RFC 003
§7.13 (docs only).
- skillset 6eb20af -> a12fe5e: mermaid-diagrams cutU normalisation and
the mempalace skill's from_agent identity rule, plus the vendored
fallback snapshot re-pinned to match (this pi-devbox commit).
Dependency audit run by direct command against every component, not
assumed: pi/mempalace/pi-atelier/pi-studio/pi-toolkit/pi-extensions/
pi-fork/pi-observational-memory all confirmed unchanged since v1.8.10.
pi-studio's apparent SHA drift on ls-remote turned out to be an
annotated-tag-object hash vs its underlying commit hash, not real
drift -- checked with ^{commit} before writing it down as 'none'.
scripts/vendor-mempalace-skill.sh, real refresh not --check: skillset
moved 6eb20af -> a12fe5e (mermaid-diagrams cutU normalisation, and the
from_agent identity rule this same release ships in RFC 003). --check
reported stale-but-truthful (exit 0, the sanctioned skip) but this
release's point is getting today's fixes live fleet-wide, and the base
rebuild is already forced by the entrypoint change and the floating
mempalace-toolkit ref moving -- so the incremental cost of also
bumping this pin is zero. Phrase canary in scripts/smoke-test.sh
unaffected: neither pinned phrase ('Provenance is stamped for you'
present, 'Attribute what you file yourself' absent) is in the section
that changed; both verified still correct in the new snapshot.
cli_utils' install.sh reaches PATH by symlinking bin/ into ~/.local/bin.
That is persistent on a host and ephemeral in a container, so the same
installer produced opposite durability and every --force-recreate sent the
human back to typing /workspace/cli_utils/bin/git-status-all. Re-link at
start instead, and add a per-device boot hook so the next question of this
shape needs no image change at all.
Symlinks rather than a PATH edit in an rc file, deliberately: ~/.local/bin
is already ahead of /usr/local/bin in ENV PATH, so links resolve in
NON-interactive shells too (docker exec, agent tool shells, scripts). An
rc-file PATH edit cannot reach those because ~/.bashrc returns early when
not interactive — measured, that asymmetry is exactly why `command -v
git-status-all` failed in one shell and worked in another on the same box.
Shell FUNCTIONS remain the sourced file's job; a symlink cannot carry them.
Guards, because ~/.local/bin is shared: a real file is never clobbered, a
symlink pointing elsewhere is never stolen, ours are refreshed, and links
into a cli_utils/bin whose target vanished are pruned — a dangling link on
PATH reads as a broken container rather than a removed script.
The hook (~/.config/devbox-shell/init.sh) adds no trust boundary: that dir
is already sourced into every interactive shell by the baked bash_aliases,
so it is already arbitrary code from the same owner. Only WHEN it runs is
new. bash <file>, never sourced, exit status ignored, output to a log.
Caught before commit, and the reason the loops use `if` bodies instead of
`&&` chains: under `set -euo pipefail` a for-loop whose last command is a
false test exits non-zero, and with no match the /workspace/*/cli_utils
glob stays literal — so the first draft would have failed to START a
container on every machine that does not have this repo, rather than merely
skipping the links. Re-tested with set -e in place: no-cli_utils/empty-HOME
no-op, guards, idempotence, CLI_UTILS_LINK=0, a hook that exits 7, and a
hook that tries to mutate CLI_UTILS_BIN — all exit 0 with intact state.
Not covered by CI: docker-publish.yml runs only on tags and lint.yml lints
workflow run: steps, so neither executes this file. Validated by extracting
both sections and running them against fixtures, then for real in a live
v1.8.10 container. No smoke assertion added on purpose — the positive path
needs a /workspace mount smoke does not have, and asserting it there would
repeat the v1.8.0 mistake of a smoke check written against a stage that
does not exist at run time.
Moves the base hash (base-decide folds `cat entrypoint.sh
entrypoint-user.sh`), so this rides along with the next tag's ~40-minute
base rebuild rather than justifying a tag of its own.
Audited every component against upstream rather than assuming. Only two moved
since v1.8.9:
- mempalace-toolkit 5b8d78f -> b2b50af (ours): feeder scrubber, the symlink fix
that stops it refusing to stage, hlc owed-set join, queued-delivery note,
explicit MEMPALACE_MAILBOX_NOTIFY protocol modes, RFC 003 + fleet-memory docs.
- pi-studio v0.9.48 -> v0.9.52 (upstream, studio variant only): 22 commits, all
additive — PDF previews, hideable header, contextual side questions. No removals
or renames. Its pi floor is >=0.84.3 against our exact 0.84.3 pin: satisfied,
zero headroom, named as a watch item because the next floor bump breaks the
studio job only, after core has already published.
Verified unchanged, with dates proving they predate the v1.8.9 bake: pi 0.84.3 (=
npm latest), mempalace 3.8.0 (= PyPI latest), pi-atelier v0.8.2, playwright 1.62.1
(no drift this cycle despite floating on latest), pi-fork bf702b4, pi-obsmem
ce9fc98, pi-toolkit 0e1369e, pi-extensions 2022887. No breaking changes anywhere,
so everything can ship in one tag.
The reason for tagging now is not features. The scrubber has been live on exactly
ONE device since this morning; every other device has kept staging unscrubbed
transcripts into a shared palace, and cleanup after the fact is manual redaction
(done twice today, with freed-page residue left behind by choice).
The entry leads with the acceptance check instead of burying it, because this
release's worst failure is silent: the feeder is fail-closed, so a packaging or
path mistake stops the fleet's memory feed and nothing complains — refusing to
stage looks exactly like a quiet session. f0bffd1 exists because that very bug was
real (BASH_SOURCE reports the symlink path, not the target). First client on the
new image must confirm a [scrub] summary line appears, the drawer count moves, and
exit 3 did not fire. Silence is the failure signal, not success.
Two problems reported against docs/observational-memory.md, one cosmetic and one
substantive. Both turned out to be worth more than the fix.
Clipping. Several boxes lost their bottom line of text in the viewer, and neither
the source nor my own render showed it. Cause: Mermaid measures a node label with
its own font metrics, commits to a box size, then renders that label as real HTML
in a <foreignObject> — so any host stylesheet touching line-height or font-size
inflates the text past a box that is already fixed, and the overflow is clipped.
The error accumulates per line, which is why it always eats the last line of the
tallest labels.
Rejected the obvious fix after testing it rather than assuming it:
%%{init: {'flowchart': {'htmlLabels': false}}}%% *is* honoured (labels switch from
16 foreignObject to 7 tspan) and clips identically, because inflated font-size
inherits into SVG text too. The fix that works is a hard limit of two short lines
per node, with the detail moved into prose under each diagram — one- and two-line
boxes have the vertical slack to absorb inflation, three- and four-line boxes do
not. The nodes were carrying paragraph-sized text; the diagrams are better for
losing it.
Regression harness: render every block with line-height: 1.7 !important forced
onto the label HTML, screenshot, read it. That caught two survivors of the rewrite
that looked fine in the clean render — a long unbreakable /opt path wrapping to a
third line, and a cylinder shape whose curved bottom leaves less room than a
rectangle for the same two lines.
New §4, because the document explained that compaction folds the ledger and never
said what that leaves in the context. Asked directly: is the session back to
knowing nothing? No. A verbatim tail survives, sized by keepRecentTokens (20k) and
cut only at turn boundaries; the system prompt and AGENTS.md were never in the
compacted region because they are rebuilt from disk each request; nothing is
deleted from disk, since compaction appends a compaction entry rather than
rewriting lines; and recall keeps resolving ids whose sources left the context
because it reads the full branch via sessionManager.getBranch() and never consults
the context window. Repeated compaction renders from live records, not from the
previous summary's prose, so there is no generation-loss spiral.
Also corrects this repo's own "compaction calls no model" to the steady-state
claim it actually is: with an empty ledger the hook returns nothing and explicitly
declines ownership, and pi's native model-based summariser runs. Snippet quoted in
the doc.
And a bug shipped in the first version: §9 said the ledger entries are
custom_message and specifically not custom. Exactly backwards, so the one grep
that section existed to get right was the one it got wrong. Verified empirically
against the live session file — 11 om.observations.recorded and 6
om.reflections.recorded, all "type":"custom", next to "type":"custom_message"
entries whose customType is mempalace-mailbox and mempalace-wakeup. That is where
the confusion came from, and the distinction is load-bearing rather than
cosmetic: the mailbox uses the context-visible append API, om's ledger uses the
invisible one. Which makes it a feature the doc now advertises — the ledger costs
zero context until it is folded.
MEMPALACE_TOOLKIT_REF=main floats and docker-publish.yml resolves it to a SHA at
build time, so toolkit main moving 5b8d78f -> f0bffd1 (10 commits) ships in the
next tagged image whether or not this repo has a commit. v1.8.9 adopted the rule
after that shape bit twice (553d865, 5b8d78f): name the behaviour change BEFORE
tagging. This is that rule obeyed rather than re-learned — the work was pushed to
toolkit main earlier today and this entry was missing, which is exactly the gap
that caused a cross-host misattribution in v1.8.7.
Contents: the feeder-side secret scrubber (three tiers, T3 report-only after a
measured 403 -> 29 false-positive calibration on 52 MB of real transcripts, fail
closed); the symlink near-miss that fail-closed would have turned into a
fleet-wide silent memory outage at bake time, caught before tagging; the mailbox
work (hlc owed-set join, queued-until-next-turn note, explicit notify protocol
modes, tmux path documented unverified); and the documentation set (RFC 003,
fleet-memory.md, secret-hygiene.md). Also states what remains unscrubbed: the
opencode bridge write path and the unbuilt server-side layer.
No tag pushed — per the release protocol, no tag means no build.
`pi-observational-memory` is baked, registered in the seeded settings.json, and
handed a cheaper model than the session it serves — and the only description in
this repo was five words in a feature list. Someone meeting `/om:status` or a
"compacted memory" block had nothing to read that said whether to leave any of it
on. New docs/observational-memory.md, 272 lines and five diagrams.
Scoped by who is authoritative, so there is one copy of each claim:
- Upstream (/opt/pi-observational-memory/docs/) already documents the mechanism
well — concepts.md, how-it-works.md, configuration.md, including a v3 lifecycle
diagram that matches the deployed code. Linked, not re-derived.
- This document takes the four facts pi-devbox owns and can change: the pinned
commit it bakes (v3.0.4 ce9fc98, matching build-manifest.json), the packages[]
entry that decides which copy loads, the Haiku-workers-vs-Opus-session split,
and the devbox-pi-config volume that makes the ledger outlive the container.
- Plus the confusion this image creates by shipping two things called memory: a
section contrasting it with MemPalace, on the line "observational memory keeps
a session coherent, the palace keeps the fleet coherent".
Placement follows the split fleet-ops states for itself — reusable mechanism is
not deployment data — so a "why is this in my container" document belongs in the
repo that pins and wires the component. Linked twice from the README, because
until this commit the README referenced docs/ zero times and the file already
sitting there was reachable only by listing the directory.
Numbers were read out of the live container and the baked tree rather than out of
release notes, which caught one thing the pi-extensions skill still has wrong:
the dropper is gated on a successful same-turn reflection, not on a token
threshold of its own.
Also fixes a README sentence that v1.8.9 made false. § Cross-machine agent
coordination ended with "Nothing in this image polls the log on the agent's
behalf"; the mailbox has shipped since aac4a1c. Replaced with the three knobs and
their defaults, derived-not-read owed-ness, and the queued-into-the-next-turn
delivery measured on two devices — and dated to "as baked in v1.8.9
(mempalace-toolkit 5b8d78f)", pointing at RFC 003 §7.11–§7.12 for the mechanism,
because toolkit main is already ahead (a92c75d pings the human who is not
looking) and describing that here would trade a stale-behind claim for a
stale-ahead one.
The cause is worth more than the fix: the behaviour arrived through the floating
MEMPALACE_TOOLKIT_REF, so no diff in this repo ever touched the paragraph making
the claim. v1.8.9's own rule fired for the CHANGELOG and nobody swept the README.
The CHANGELOG records what changed, the README asserts what is true, and only the
first is reviewed at release time — so the rule now extends to grepping the README
for absolute claims (nothing, never, does not, only) about a component whose SHA
moved.
Diagrams were verified by rendering, not by parsing. Both comparison diagrams
parsed clean and rendered with their meaning reversed: Mermaid laid the second
declared subgraph out first, putting "with observational memory" before "without"
and MemPalace before observational memory in the diagram whose entire job was
that contrast. Rebuilt as declaration-ordered chains and re-rendered at mermaid@11
— the version pi-studio pins — in the baked headless browser, then read back as an
image. mermaid.parse() proves syntax and says nothing about layout.
Two versions, two flags. `--expected-version` has only ever asserted
`pi --version`, but AGENTS.md step 4 spelled it `X.Y.Z` inside a checklist where
every other X.Y.Z is the pi-devbox tag. Run as documented for v1.8.8 the final
runtime gate of the release printed
✗ pi version mismatch: expected 1.8.8, got 0.84.3
and exited 1 — a red accusing the image of being the wrong version. Not one
reader's slip: the v1.8.8 release-readiness handoff from pi@emb-7kj4vr4g
propagated the same wrong spelling twice while correctly calling step 4 "not
ceremonial", so two independent readers converged on it. README.md had it right
all along, which means the two documents disagreed.
- new --expected-image-version asserts the pi-devbox release tag, read from
release_tag in /etc/pi-devbox/build-manifest.json (no checkout, no network);
leading `v` optional on either side
- both flags detect being handed the other one's value, and the test is exact
rather than heuristic: the value is compared against the other quantity the
image itself reports, so it can only fire on a real mix-up
- neither flag is required now. With none, live `pi --version` is asserted
against the manifest's pi_version — not a tautology, since a stale pi in the
~/.pi/npm-global volume can shadow the baked one, exactly as a stale
npm:pi-atelier can in packages[]
- the header note replaced was stale and load-bearing: it claimed pi is resolved
from 'latest' and cannot be self-derived, while Dockerfile.variant pins
ARG PI_VERSION=0.84.3 and docker-publish.yml reads that ARG as its source of
truth. The same withdrawn claim also sat in cli_utils' pi-devbox-sanity --help
- argument parsing: a missing value, or a value that is another flag, is a usage
error instead of silently consuming the next argument; --help works
All fourteen flag combinations exercised by execution, including the two
manifest-absent branches and the shadowing branch a healthy container cannot
reach — mutation-tested with a doctored manifest so each failure branch was
observed firing rather than assumed present.
CHANGELOG also names what no commit here causes: mempalace-toolkit main moved
e70bef2 -> 5b8d78f, so this tag ships the auto-delivered logstream mailbox
because base_tag folds the resolved toolkit SHA. It would have landed either
way; going unnamed is the 553d865 shape that already caused one cross-host
misattribution. Component audit found nothing else to bump — pi, mempalace,
pi-atelier all equal their upstream latest, and every other floating ref
resolves to the commit already baked.
Freezes the v1.8.8 section and clears the two non-code checklist items
pi@emb-7kj4vr4g handed over (evt_20260826T134919_a614ecfc2d4f), plus the two
carried nits from its round-2 verification (evt_20260826T133356_d56792791a49).
Every claim below was re-measured here rather than taken from the handoff.
THE STALENESS NOTICE ASSERTED A DIRECTION IT NEVER TESTED — Blocker 1's shape,
one layer down, in the message I added to replace the message that named the
wrong cause. The notice fired on "recorded != HEAD" and then announced HEAD as
the newer side without testing ancestry, so a clone that was merely BEHIND got
"has moved to 82a8d3c; the snapshot describes the older 5fd0d5c" when 82a8d3c is
5fd0d5c's ANCESTOR. Found by EMB against the real state of its own host, not a
fabrication. The verdict was never wrong (rc 0, nothing mis-verified) but the
remedy it implies is a ~67-minute base rebuild when the actual fix is `git pull`
— the only one of the two carried nits with a price tag, which is why it went
first. Now tests ancestry with the merge-base --is-ancestor primitive the refresh
path 60 lines below already used, and reports three verdicts: stale (refresh),
clone behind (pull, do NOT refresh), diverged (reconcile). All three verified by
execution; only the first was correct before. --help no longer errors on the one
script whose argument order was itself a landmine. The `-s "$VENDORED"` guard is
now commented as load-bearing: it makes the empty-stdin collision unreachable by
construction, which also means no test below exercises it any more, so deleting
it as "redundant with the probes" would silently restore the false OK.
CHANGELOG: retitled, and three stale spots fixed in what becomes the permanent
record. It cited the skill at 82a8d3c (twice superseded); it RE-ASSERTED the
retracted mailbox measurement as live evidence 200 lines after withdrawing it,
which is the exact non-contradiction failure this release exists to fix; and its
warning block still described the pre-e8ddeaf world ("still records c04cd15",
"now exits 1") while quoting as exemplary the very notice whose direction was
unverified. Per EMB's steer the conclusion was kept and only the evidence
replaced: the status filter does drop broadcast noise, it just never computed
owed-ness. Honest replacement, measured on both machines: raw filter returns 2
here and 1 there, EVERY ONE already answered, derivation returns 0 for both.
SNAPSHOT RESYNCED AGAIN, 5fd0d5c -> 6eb20af, because skillset 6eb20af adds the
limit of my own seq ordering test: seq is REPLICA-LOCAL, equal to origin_seq only
because one replica authors for all four machines, so use hlc once mesh_peers
reports a peer. Recorded as reasoning not measurement — a second replica cannot
be stood up here. The durable half is the asymmetry: seq skew makes an ANSWERED
item resurface (noise, visible, self-correcting) while created_at SUPPRESSES AN
UNANSWERED ask forever (silent, permanent), so the skill now says outright that
"fixing" a resurfacing item with a timestamp trades the safe failure for the
dangerous one. The resync was free: rootfs/ was already changing, so the base
rebuild was forced regardless — the ordering warning about accidental staleness
does not apply to a deliberate refresh before the tag.
--check is a clean OK at 6eb20af with no notice, canary re-verified bidirectionally
(present 3, withdrawn 0), baked snapshot 0644, tree hash recomputed at build time
and re-verified in-container. bash -n clean; shellcheck/hadolint/actionlint remain
absent locally, so CI is still the only evidence for those.
Fixes the three blockers and seven should-fixes from pi@emb-7kj4vr4g's review
(logstream correlation skills-provenance-review, full text in
drawer_pi-devbox_reviews_e43e766641c9ec85217bc6ce). Every finding was
reproduced by execution here before being fixed; two were refined by that
reproduction rather than taken as given.
BLOCKER 1 — the provenance gate could print OK and exit 0 without verifying.
`git show <ref>:<path> | sha256sum` hashes EMPTY STDIN when the ref does not
resolve, so at_ref was never empty and the UNKNOWN branch was dead code.
Measured: a bogus ref reported MISMATCH — accusing the snapshot of lying when
the real cause was an incomplete clone, and the operator's natural remedy for
MISMATCH is to re-run the refresh, which rewrites provenance to silence the
complaint; and with a 0-byte snapshot against a 0-byte upstream file it printed
"OK: exactly skillset@aaaaaaa" with exit 0 for a ref that does not exist. The
script already had the sha_empty idiom and had applied it to blob_sha but not
to at_ref. Existence is now PROVEN with git cat-file -e before anything is
hashed, at two levels (ref resolves / path exists at it) because those deserve
different messages. Same defect class as the canary it replaces: a check that
can succeed without checking. A second, unflagged instance of the same pipeline
shape in blob_sha was found and fixed too.
Exit codes split, because the old contract failed the sanctioned case: 0
truthful (including stale, with a NOTICE), 1 a lying record only, 2 cannot
determine. AGENTS.md step 2 promised "the message distinguishes the two" and
was the thing this branch was breaking; rewritten to state all three.
BLOCKER 3 — VENDORED.md contradicted itself in the release whose stated
invariant is non-contradiction: its hand-maintained provenance line named
skillset 670f7f1, seven commits behind the ARG and itself the commit that told
agents to hand-stamp added_by — the withdrawn instruction this work exists to
stop shipping — while its cp recipe contradicted the "not cp" rule 20 lines
above. Line removed (nothing forced it to move when the ARGs did); 670f7f1 kept
only as a labelled cautionary example. The pi-extensions half was verified
redundant (CI require_sha resolves PI_EXTENSIONS_REF) before removal.
SHOULD-FIXES: `<root> --check`, the spelling VENDORED.md documented, silently
ran a REFRESH because only $1 was parsed (both tools now parse all args and
reject unknown ones); refresh at a detached/older HEAD silently rewound ref and
bytes (now refused unless the recorded ref is an ancestor, --force to override);
upstream_dirty was computed and never used in check mode; --no-skills --json
printed human text and broke jq; --help was a hardcoded sed range this branch
had already made stale; the fingerprint hashed SKILL.md alone so a live skill
differing only in a sibling file reported "identical", and pi-extensions already
ships two files, so it is now a per-skill TREE hash with the manifest field
renamed skillset_snapshot_tree_sha256; the --no-skills smoke assertion was
negative-only and passed on a crashed binary. mktemp+mv left files 0600 — CI was
unaffected since the index records 100644, so the blast radius was local builds
only, narrower than the review inferred.
Snapshot resynced c04cd15 -> 5fd0d5c so the no-clone fallback carries the
CORRECTED coordination protocol rather than the withdrawn one; --check is now OK
with no staleness notice, and the bidirectional canary re-verified against the
new bytes. Local validation is bash -n only (shellcheck, hadolint and actionlint
are all absent in this container) — CI remains the shellcheck gate.
The logstream has carried cross-machine work since 2026-08-18 — patch handoff,
review, a v1->v2 supersede — and nothing in this repo said it existed. That gap
had a measurable cost this morning: another host addressed a retraction to
pi@tor-ms22 by name and it was read only because the human said "read the
logstream", while the agent was actively rebuilding the thing it warned about.
Split by what each document is authoritative for, so there is one copy of each
claim rather than three that drift:
- README § Cross-machine agent coordination — what the CONTAINER needs.
MEMPALACE_REMOTE_URL selects the shared palace; MEMPALACE_PI_DEVICE is what
makes this machine reachable, because where every host is a thin client of one
palace the stamped agent name is the only thing that distinguishes them. Stated
as a rule with teeth: set both or neither, since a container missing the device
var can read the log but is addressable by nobody.
- AGENTS.md release checklist step 2 — the vendored-snapshot refresh, as a
MECHANISM in the document a releasing agent actually reads, not a comment
hoping to be noticed. It says the refresh costs a base rebuild, that skipping
it is legitimate (every enrolled host reads its live clone), and that skipping
it silently is not.
- CHANGELOG — the three-way split itself, plus the measurement that shaped the
ack contract: unfiltered, the mailbox returned 5 events, 4 of them finished
broadcasts from eight days earlier; with status="open", exactly the 1 that
needed an answer.
Norms live in the skillset skill (82a8d3c, already live on every host that mounts
the skillset — no rebuild) and mechanism in mempalace-toolkit's
extensions/pi/README.md (e70bef2, which also documents the edge stamper that
553d8657 shipped undocumented). Deliberately NOT duplicated here.
Consequence recorded rather than hidden: the skill edit lands in the skillset, so
this repo's SKILLSET_SNAPSHOT_REF now honestly reports itself behind, and
--check exits 1 with "has moved to 82a8d3c; the snapshot describes the older
c04cd15". That message is also fixed in this commit — it previously blamed "the
working tree" even when the tree was clean and only the ref had moved, which is
the same defect class as a canary pinned to a phrase the release deleted: a
message that names the wrong cause. Now distinguishes moved-HEAD from dirty-tree,
verified against both plus the in-sync case.
Found while verifying v1.8.7 from inside a fresh container: the baked mempalace
snapshot is read by no host on this fleet. devbox-skill-reconcile repoints
~/.agents/skills/mempalace at the mounted live clone (the v1.8.5 fix working as
designed), and all four compose stacks mount a workspace containing the
skillset. So the phrase canary that blocked v1.8.7's first tag polices a file
nobody opens, while the drift that could actually mislead an agent — a git pull
nobody ran in /workspace/skillset — was invisible from inside the container and
is invisible to CI by construction.
Record provenance instead of policing it, and move the check to where the
skillset actually is:
- Dockerfile.variant: ARG SKILLSET_SNAPSHOT_REF (the claim) + a sha256 of the
shipped bytes measured in the manifest layer (the fact), as manifest siblings
rather than components{} members, plus an OCI label. An ARG default, not a
CI-resolved output: no credential for the private skillset, no change at any
of the four variant build call sites, and a local docker build records what CI
does. Variant-only, so no base rebuild — check-base-hash.sh scans
Dockerfile.base alone, verified by running it.
- pi-devbox-version: a skills: section naming baked vs live <repo> @ <sha> per
vendored skill, and for mempalace whether the live copy is identical to the
baked fingerprint, at the same commit with uncommitted edits, or divergent.
entrypoint-user.sh passes the new --no-skills, because the banner prints
before the links exist and long before the reconcile runs.
- scripts/vendor-mempalace-skill.sh: refresh the file and rewrite the ref
together (a cp without an ARG bump makes the manifest lie, which is worse than
anonymity); --check verifies the claim against a real clone.
- 5 new smoke assertions (78 -> 83), mutation-tested through the real sh -c
path: 6 fabricated manifests, where a well-formed hash of the wrong file
proves the two manifest assertions are not redundant; the all-baked reporting
test verified to FAIL against a live-skillset environment.
Reviewed mid-flight by pi@emb-7kj4vr4g over the logstream (correlation
skillset-vendor-drift), which retracted its own earlier recommendation of a
build-time byte-compare against skillset HEAD and supplied the better framing:
the invariant is NON-CONTRADICTION, not currency. Byte parity on a fallback
would have cost a resync commit plus a ~67-min base rebuild for each of the four
skillset commits pushed in one evening. Its warning also found a real bug here:
the script now CONSTRUCTS the snapshot from `git show HEAD:<path>` instead of
copying the working tree, because a clean `git diff` says nothing about an
untracked file — the one input the first draft would have recorded a false ref
for. Tested: untracked, unstaged and staged-but-uncommitted all refuse, atomically.
Also fixes three stale in-repo markers of the same class the canary belongs to
(true when written, silently false at release): two dangling "Unreleased"
pointers and a typst line still marked Unreleased five releases after v1.4.0.
c04cd15 ('the withdrawal only holds where the bridge is live') landed after the
v1.8.7 snapshot was taken, so the baked fallback was already 4 lines behind the
skillset within hours of publishing. It adds the caveat this fleet is currently
living in: the bridge is baked at image build, so a container on an image older
than the stamping commit satisfies both env gates while stamping nothing, and
hand-stamping is still the only signal a hand-filed drawer gets there. It also
gives the one-line test —
grep -c MEMPALACE_PI_DEVICE "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)"
which returns 0 on this v1.8.6 container, confirming the gap empirically.
Note what this instance proves about the canary fixed one commit ago: it still
PASSES on the refreshed copy, because both pinned phrases survived the edit. A
phrase canary cannot detect 'older than skillset main' — only a diff can. This is
the second drift in 24h and is the argument for the Still-open item (a CI job
diffing this file against the skillset repo, blocked on a clone credential for a
private repo). No re-pin was needed here.
Run 589 built the base cleanly and then failed both smoke jobs 81-passed/1-failed
on 'mempalace skill snapshot is current'. That canary greps a phrase from the
vendored mempalace skill to detect a stale snapshot, and the phrase it pinned was
'Attribute what you file yourself' — the heading of the hand-stamping instruction
that THIS release withdraws. So it fired correctly: the snapshot changed and the
expectation did not. Every publish job was skipped, so nothing reached the
registry and v1.8.7 was never consumed.
Rather than bump the string:
* the assertion is now BIDIRECTIONAL — the new phrase must be present AND the
withdrawn one absent. A one-way canary only catches half the drift: it cannot
notice a re-vendored stale snapshot that happens to contain the pinned phrase.
Verified against v1.8.6's snapshot, which now correctly fails.
* the comment records the structural limit rather than just the fix: a phrase
canary can only ever detect 'older than what I remembered to pin', never
'older than skillset main'. Only a diff against the skillset repo can do that,
which is now a Still-open item — it needs a CI clone credential for a private
repo, i.e. a policy decision, not a code change.
Changelog consolidated: the SSH sidecar multiplexing default moves from
Unreleased into v1.8.7, since the retag will sit on a commit that contains it,
and the v1.8.7 summary now records the failed first attempt rather than quietly
presenting the second one as the whole story.
A target whose ~/.ssh/config entry never mentioned ControlMaster got no
multiplexing from the sidecar (only ControlPath was supplied), so every ssh call
opened a fresh TCP connection. On 2026-08-25 that produced ~12 connections to
one host in 15 min and a fail2ban block that looked like an outage — the tell
being that HTTPS to the same estate stayed healthy.
The correctness of this depends entirely on WHERE the block goes. ssh_config is
first-value-wins:
ControlPath before the Include -> override (the user's value points at
read-only ~/.ssh and cannot work in the container)
ControlMaster after the Include -> default (an explicit per-host
'ControlMaster no' must keep winning)
Force what is broken, default what is merely absent. The first draft put both in
the leading block and would have silently overridden an explicit 'no'.
Verified with ssh -G rather than from the man page, including the counterfactual:
under the shipped layout an explicit 'no' resolves to controlmaster false while a
silent host resolves to auto; under the rejected layout the 'no' host flips to
auto. So the test discriminates position, not presence. Plus a sandbox render of
the real script, bash -n, and shellcheck -S error (the v1.8.7 gate) clean.
Effect measured on 41 real host aliases: 22 silent entries gain auto+10m, 0
overridden. Note the fleet's one deliberate opt-out is written as absence plus a
comment ('# No ControlMaster — VPN means direct route'), which ssh cannot
distinguish from no opinion; that host now multiplexes, which its own comment
says is unnecessary rather than harmful.
Skill documents the sidecar-vs-~/.ssh trap (the failure misleads: read-only
ControlPath makes multiplexing look impossible rather than misconfigured) and
the stale-master recovery, ssh -O check / -O exit.
The provenance fix's client half lives in mempalace-toolkit, which the image
clones at build time, so it only reaches the fleet through a tag. Records both
routes (extension via MEMPALACE_TOOLKIT_REF folded into base_tag; vendored skill
via rootfs), the three design points (stamp in the client not the agent; diary
marker in TEXT because metadata is invisible to readers; solitary devbox stamps
nothing), and why the allowlist is per tool (3.8.0 hard-fails -32602 on
undeclared args). Carries the CI-hardening work already sitting in Unreleased,
and a Still open block for the three known bounds.
VENDORED.md's freshness model for `mempalace` is "Option 2 only — refreshed
manually per release", and it had drifted since 2026-08-23. The stale snapshot
still carried the instruction to hand-stamp added_by="<harness>@<device>", which
skillset 73c7c8e withdrew: the pi bridge now stamps at the edge
(mempalace-toolkit 553d865), and RFC 001 §7.3.2 ranks agent-side stamping ❌
worst-possible.
That matters specifically for the fallback case this snapshot exists to serve — a
container started WITHOUT the private skillset mounted would otherwise be the
only kind of container still being taught to do it by hand.
lint.yml has shellchecked every workflow `run:` step since the dash-vs-bash
incidents, but nothing had ever pointed shellcheck at entrypoint.sh, scripts/*.sh
or the extensionless tools under rootfs/usr/local/bin/. That gap is not
hypothetical: the skillset repo's ci-release-watcher template shipped
`echo "$json" | python3 <<'EOF' ... json.load(sys.stdin)` for two months, where
the heredoc IS python's stdin (no script arg) so the load hit EOF and the
function silently returned nothing. shellcheck names exactly that at severity
ERROR — SC2259, "This redirection overrides piped input" — and could have named
it the whole time.
New step in the existing actionlint job, so no second container pull: shellcheck
-S error plus bash -n over every shell file, discovered as *.sh UNION a shebang
scan (the glob alone misses pi-devbox-version, devbox-skill-reconcile, dot-watch
and studio-expose; a shebang scan alone would miss a sourced fragment without
one). Fails loudly on a zero-file match, because a green tick over an empty set
is not a check.
Severity chosen by measurement, not taste: -S error is 0 findings across all 11
shell files today, so the gate is green on arrival with no cleanup, while
-S warning is NOT free (19x SC2088 tilde-in-quotes in recreate-sanity-check.sh
plus assorted SC2016, all intentional) and would train everyone to ignore the
job — the same reasoning as the SHELLCHECK_OPTS exclusions already on the
actionlint step.
Closes the item v1.8.6 (and v1.8.5 before it) listed as "Still open": the
palace pin was a literal string in Dockerfile.base with zero references in
docker-publish.yml, while PI_VERSION had a concreteness gate, a
published-on-registry check and a never-silently-adopt drift warning.
resolve-versions now applies all of those to MEMPALACE_VERSION, read from
Dockerfile.base so a local `docker build` and CI install the same version by
construction, plus one gate pi does not need: a YANKED release is refused,
because an exact pin installs one silently under PEP 592 and would have
shipped a withdrawn palace client to the whole fleet.
smoke gains `installed mempalace matches CI's audited pin` via a new
EXPECTED_MEMPALACE_VERSION threaded into both smoke jobs. It is not redundant
with `manifest mempalace_version matches the installed core`: that compares two
properties of one image and cannot notice that both are the wrong version. The
case this covers is a variant built FROM a cached base carrying an older pin —
internally consistent, silently stale.
Mutation-tested by extracting the shipped block out of the YAML and stubbing
curl: 9 cases covering every gate, then once end-to-end against live PyPI. That
found a real defect in the first draft — the yank message inlined a jq program
inside $(...) inside a double-quoted string, where the escaping broke the
filter (jq compile error) while the surrounding `exit 1` still fired: a gate
that looked correct and reported garbage.
Note: correcting Dockerfile.base's now-false "known gap, carried forward"
comment forces a base rebuild (~67 min) on the next tag. Leaving a comment
asserting the audit does not exist was the worse option.
Four build-provenance assertions grepped the manifest for a field name and
never looked at the value:
run_expect "manifest records pi_version" "cat …manifest.json" '"pi_version"'
which passes on {"pi_version": ""} and on {"pi_version": null}. The tell was
in its own passing output the whole time — `✅ manifest records pi_version (got
"pi_version")` echoes the key back as the thing it claims to have found. Found
while reading run 579's smoke log to confirm v1.8.6's new assertions had really
executed rather than merely gone green.
Now checked against values, and against ground truth where it exists:
- every required component key present, naming the one that vanished
- every component value a full 40-hex SHA (null allowed for pi-studio alone,
which is legitimately absent in the non-studio variant)
- pi_version equal to `pi --version`, mirroring the mempalace ground-truth check
- release_tag non-empty; source_revision 40-hex and build_date ISO-8601 *when
populated*, since both default empty on a plain local `docker build` and
demanding them would fail honest local smoke runs
- --json compared byte-for-byte with the file, which is assertable because that
mode is a verbatim cat; the old form grepped its output for "release_tag"
Key presence and value shape are deliberately SEPARATE assertions: a single
"all values are valid SHAs" loop passes vacuously on components:{}, because
jq's all() over an empty list is true. Combining them would reproduce the same
shape of hole as the three false greens already recorded in CHANGELOG.md.
Dropped `manifest has no unresolved ('unknown') components`: the 40-hex check
strictly subsumes it ("unknown" is not 40-hex, and only rev() emits it, feeding
components{} exclusively). Removed rather than kept, because a check that can
no longer fail independently is one more green tick that means nothing.
Mutation-tested twice rather than reasoned about: nine fabricated manifests
through the raw jq filters, then twelve through the shipped assertions using
the real `run` helper's `sh -c` quoting path — the quoting is load-bearing,
since a jq filter dying on a quoting error exits non-zero and looks exactly
like a caught defect. Measured on the same twelve defects: old caught 3, missed
9; new catches 12. Three legitimate variations stay green (empty
source_revision, empty build_date, null pi-studio).
Also corrects a factually wrong "Still open" bullet in the released v1.8.6
entry, which claimed pi-devbox-version's human output does not show
mempalace_version and that only --json surfaces it. Both halves are false: it
prints a `palace:` line with live-vs-baked drift annotation, verified against
fabricated manifests (match, skew, and pre-v1.8.6 absent-field cases). Left as
a struck-through correction rather than deleted, since v1.8.6 is published.
Three coupled pieces of work, all of which ride on the base rebuild that the
mempalace bump forces anyway.
DRIFT ADOPTED
- pi 0.84.2 -> 0.84.3. Its release notes carry a "Breaking Changes" line
(GoogleThinkingLevel -> GoogleApiThinkingLevel). Audited before adopting:
zero references across all four vendored companions (pi-fork,
pi-observational-memory, pi-atelier, pi-studio), so it is inert for us. The
reason to adopt is two skill-discovery fixes that land directly on v1.8.5's
vendored-skill work: nested Markdown skills inside grouping directories were
not discovered, and root README.md/AGENTS.md in skill dirs were reported as
broken skills.
- mempalace core 3.7.1 -> 3.8.0. Additive/reliability only. Its sync fix
(#2320/#2322) stops sync --apply deleting drawers whose source_file was
unreachable *at that moment* -- which does NOT relax the standing landmine
against sync on the shared palace, because that landmine is about paths
permanently absent from whichever host runs the sync. Different failure
shape; the caution stands.
DOCS -- three defects, one of them public
- DOCKER_HUB.md advertised "neovim (LazyVim defaults)". Nothing in the image
installs LazyVim; the only nvim config is a 19-line sysinit.vim. CI PATCHes
this file into the Docker Hub description on every release, so this was a
false claim published to the world. Removed.
- agent-browser + Playwright + Chromium is the single largest addition in the
image (~625 MB) and had zero mentions in README, DOCKER_HUB or THIRD_PARTY --
it was documented only to agents, in the AGENTS.md managed block. Now
documented to humans, including the Chromium licence dimension.
- typst and socat appeared in README prose but not in the "What's inside"
inventory. Added.
OBSERVABILITY -- the three gaps v1.8.5 listed as still open
- build-manifest.json now records mempalace core, read from the live binary
(ground truth, not the build ARG). Placed as a sibling of pi_version rather
than inside components{}, because pi-devbox-version renders that map through
[0:12] and would truncate a version string.
- smoke asserts the pi-observational-memory clone actually CONTAINS the ce9fc98
auth fix, pinned to src/runtime.ts. Deliberately not a repo-wide grep: two of
the three markers also live under tests/, so the repo-wide form stays green
with the fix site reverted. That is the third false-green of this exact family
in this repo (canary phrase in both snapshots; reconciler fixture using a
non-owned name; now this) -- pin containment checks to the fix site.
- smoke asserts the feeder's pi@<device> agent default behaviourally. The
earlier audit concluded this needed a --print-config added upstream; it does
not. AGENT is assigned before arg parsing, so `bash -x mempalace-pi-session
--help` observes the real resolution with no toolkit change. Two-sided:
device set => pi@<device>, unset => must not be pi@*.
- pi-devbox-version now prints a palace: line with the same live-vs-baked drift
detection pi already had. This matters more than it looks: mempalace is the
one component that is both client (here) and server (synlig), so skew between
them is a real failure mode. Degrades quietly on pre-v1.8.6 images.
Deferred deliberately: a native arm64 act_runner on tor-ms22 (the current
runner is on synlig, x86_64, so every arm64 layer ships QEMU-emulated).
Analysis and caveats filed to the palace rather than actioned here.
The only MemPalace variable the template never mentioned, and the one that
moves the feeders' stage as a side effect: the palace root resolves as
$MEMPALACE_PALACE_PATH -> $MEMPAL_PALACE_PATH -> ~/.mempalace/config.json ->
~/.mempalace/palace, and the stage is derived from it (<palace-root>/pi-stage).
The comment states that precedence, records why neither the image nor the
entrypoint exports it (pinning the palace without carrying the stage along
re-creates the split a shared root removed, v1.8.2), warns that a stage whose
persistence differs from the palace makes a scoped `mempalace sync` prune
conversation drawers whose dedup key is the staged path, and notes it is a
path INSIDE the container unlike the host-side WORKSPACE_PATH/SSH_KEY_PATH
above it.
Found while auditing a live host whose .env sets it redundantly to the
default value.
No pin moved except mempalace-toolkit fd8b15f5 -> 0fe64c4 (feeder defaults
--agent to pi@$MEMPALACE_PI_DEVICE, so hand-filed palace writes carry
provenance; $USER remains the fallback when the var is unset). pi 0.84.2 and
mempalace 3.7.1 are still upstream-latest, atelier v0.8.2 keeps the >=0.7.1
floor for pi >=0.84, and om master is still ce9fc982 -- nothing landed after the
merge that fixed the eight-week silent-observation bug.
Expect a full multi-arch base rebuild: entrypoint-user.sh, rootfs/** and the
resolved toolkit SHA all feed the base hash.
Baked skill links won over the live skillset clone for all three vendored
skills, so a pushed edit to skills/mempalace/SKILL.md was invisible in every
container until the next image build -- measured on two hosts (live md5
129bcc4752 vs baked 5236024fef). Cause was ordering, not intent: the baked links
are created early with a create-only-when-absent guard to close a smoke
readiness race, and the skillset deploy runs last and treats them as foreign.
The comment claimed the opposite of the behaviour.
The fix is not "skillset always wins". Ownership is per-skill: pi-extensions is
owned by its package repo and copied over the snapshot at build time, so the
skillset's lagging duplicate must keep losing; pi-devbox-environment is authored
here. Only mempalace is skillset-owned. devbox-skill-reconcile therefore runs
after the deploy and repoints only the names in skills/skillset-owned.txt,
replacing a link solely when it points into the baked tree, so a real directory
or a link pointing elsewhere is never disturbed. Precedence is now user override
-> live clone (owned names) -> baked snapshot, with the early links intact as
the fallback so the readiness race stays closed.
Reviewing that turned up a latent boot-abort in the pre-existing baked-link
block: `[ ! -e "$link" ]` is TRUE for a dangling symlink, so once a link can
point into /workspace/skillset, a vanished mount makes plain `ln -s` fail with
"File exists" -- and under `set -euo pipefail` that aborts container start
before `exec "$@"`. Reachable on `docker restart` or a host reboot, not on a
recreate, since ~/.agents is not a volume on any host. Now `ln -sfn`, which
heals the link back to the baked fallback.
Smoke additions cover what let this ship: the stale-snapshot canary grepped a
phrase present in BOTH the stale and fresh copies, so it passed throughout;
it now pins the newest section. Link targets are asserted, not just `test -L`;
the owned-list content is asserted both ways; and the reconciler's replace path
-- which no CI container exercises, since none mounts a skillset -- is covered by
fabricating one. A mutation test showed the obvious three assertions still pass
with the "is this link ours?" guard deleted, so a discriminating case was added:
an owned name whose link is a user override outside the baked tree.
Also refreshes the mempalace snapshot to skillset 670f7f1 (without it the fix
helps only hosts that mount skillset) and corrects README, which documented the
old, wrong precedence in three places.
Verified with 12 fixture cases plus 2 mutants: ownership respected against the
real trees, user overrides preserved, relative/trailing-slash/CRLF/space/glob
inputs handled, dangling link healed, read-only skills dir exits 0, idempotent.
~/.agents/skills is asymmetric: the three pi-devbox-specific skills resolve to
the baked copies while all others resolve to the live skillset clone. Root cause
is ordering in entrypoint-user.sh -- baked links are created early (line 65) with
a create-only-when-absent guard, and the skillset deploy runs last (line 387) and
leaves them alone as foreign links. The comment at line 61 states the intent as
protecting the skillset skill from being clobbered, but the effect is the
reverse.
Cost measured on two hosts: a pushed edit to the mempalace skill (live md5
129bcc4752) was invisible to both containers, which kept loading the baked copy
(md5 5236024fef). Editing those three skills appears to work and silently does
nothing until a rebuild.
Filed as a known issue with a proposed fix rather than fixed here: changing
symlink precedence is image behaviour and wants its own review plus a smoke
assertion, and the early-link ordering exists to close a readiness race that
must not regress.
Three false negatives in one session, all self-inflicted, all convincing
because the command "succeeded": a `| head -20` proved an SSH peer absent
that sits at line 454 of a ~500-line config; `ssh mac 'docker ps'` proved
the host had no Docker, when the non-interactive PATH simply lacks
/usr/local/bin; and `grep 'ssh '` proved no ControlMaster was running,
when those processes rename themselves to `ssh: <path> [mux]`. Same root
cause each time, so it goes in the skill rather than in a commit message:
a positive result carries its own evidence, absence has to be earned.
The skill (rootfs/, symlinked into ~/.agents/skills) is BAKED, so this is
an image change and is logged in CHANGELOG Unreleased accordingly. Its §3
also now records that a live ControlMaster socket makes later commands
authenticate not at all -- after editing a peer's authorized_keys, "it
still works" proves nothing; prove it with -o ControlPath=none, or the
breakage waits for a future session that has no memory of the edit.
AGENTS.md: corrected a stale CI claim while placing the pointer. It said a
tag push produces two runs including lint; lint.yml has since been scoped
to branches: ['**'], which excludes tag refs, and refs/tags/v1.8.4 duly
produced run 571 (publish) and nothing else. Kept the head_sha + workflow
path filter advice, which is cheap and guards against a future v*-triggered
workflow. Added a short section on verifying this repo from inside a
container, including that docker-compose.yml here is a TEMPLATE pinning
:latest while a real host runs its own per-machine file -- recreating from
the repo copy can silently move a host off :latest-studio.
Placement note: AGENTS.md is only auto-read when the cwd is this repo, so
the durable rule lives in the skill, which loads by description match in
any pi-devbox session.
Headline: pi-observational-memory 37986b6 -> ce9fc98. The ambient-credential
gate fix (6f694e6 + 699ccc7) was merged upstream as PR #52 on 2026-08-22,
closing issue #51, so this release picks it up through the ordinary
PI_OBSMEM_REF=master path with nothing carried locally. Every image up to
and including v1.8.3 silently recorded zero observations on a Bedrock host
using ambient AWS credentials; from this one on, /opt is the fix and the
settings.json packages[] workaround should be deleted (verify against
/etc/pi-devbox/build-manifest.json first). npm still ships the broken 3.0.4,
which does not matter here because the image clones the ref instead.
Pin bump: PI_ATELIER_REF / PI_ATELIER_VERSION v0.8.1 -> v0.8.2, audited per
the floor note above the ARG. The only version in between is 0.8.2 itself and
both its entries are Workspace-Pulse-internal (inspection coalescing and
serialization; fresh inspection guaranteed at Turn end, retired sessions can
no longer publish stale results). Nothing touches pi's private TUI renderer,
which is the coupling behind the 0.6.0/0.7.0-under-pi-0.84 startup hang, and
pi is unchanged at 0.84.2 - so the bump stays outside that risk class.
Also baked by this build, no pin needed: pi-extensions 98eb07b -> 2022887
(the todo edit action), pi-fork 4a09af4 -> f1ff808, pi-studio 0.9.44 ->
v0.9.48, mempalace-toolkit b609cf5 -> fd8b15f (docs only - the pi-session
false-success guard was already baked in v1.8.3, confirmed by ancestry),
aws-cli 2.36.24 -> 2.36.29 and the other *_VERSION=latest tools.
README: the pin table also said mempalace 3.6.0, stale since v1.8.3 bumped it
to 3.7.1. Fixed in passing.
Unchanged and verified current: pi 0.84.2 (npm latest, published 2026-08-14),
MEMPALACE_VERSION 3.7.1 (PyPI latest), pi-toolkit 0e1369e.
The commit itself lives in pi-extensions (2022887), but /opt/pi-extensions
is baked into this image, so "which todo behaviour does this image have" is
an image question. PI_EXTENSIONS_REF=main means the next build picks it up
with no pin to bump -- worth stating explicitly, since a reader who expects
a version bump will otherwise go looking for one.
Also recorded, because it confused a session today: pi-atelier's
tool_result hook replaces todo output with "N/M done - see sidebar"
whenever the sidebar panel is visible, so an agent sees the counter and not
the item text. Upstream's list returns every item; the terseness is
atelier's deliberate context saving, not a limitation of the tool.
mempalace 3.6.0 -> 3.7.1. Verified against the 3.7.1 source rather than its
changelog, because the risk lands on palaces users cannot reconstruct: legacy
drawers lack the new chunk_total marker and both decision sites trust them, so
no mass re-mine; NORMALIZE_VERSION is 2 in both; chromadb stays <2 so no
index-format migration; no auto-migration exists; logstream.sqlite3 is created
lazily. Downgrade remains possible (3.6.0 has zero references to chunk_total).
Two behaviour changes documented in the CHANGELOG: ALLOW_PEER_WRITER no longer
works on local/chroma palaces, and writer-lock setup failures fail closed.
Neither affects this image's MCP-server-plus-CLI-feeder pattern, which already
serialised on the same lock under 3.6.0 -- the upstream "process-lifetime
single-writer" entry describes tightened escape hatches, not a new lease.
The motivation is the shared central palace: 3.7.1 drops the stale chromadb
SharedSystemClient cache on reconnect (3.6.0 could let a stale in-memory HNSW
segment overwrite a peer's writes, "index count going backwards"), releases the
writer lease on SIGTERM/SIGHUP, and stops treating an interrupted mine as
complete. The fleet primary was upgraded to 3.7.1 and restarted before this tag,
because 3.7.1 refuses writes when the served library drifts and reconnect cannot
clear that. opencode-devbox still pins 3.6.0, so the lockstep is broken until it
cuts its own release.
Vendored mempalace skill snapshot refreshed to skillset 936fed8 (was 63f3bf5).
This closes a gap that had been invisible for two commits: ~/.agents/skills/
mempalace symlinks to the IMAGE-BAKED copy, entrypoint-user.sh creates that link
first, and the skillset deploy never clobbers an existing name -- so in a devbox
container the vendored snapshot always wins and editing skillset alone changes
nothing a container reads. Brings the multi-machine shared-palace section and
the hand-crafted-provenance guard.
pi-global-AGENTS.append.md already carried the three shared-palace damage rules
(a55f636); this tags them into a release.
mempalace-census gets the /usr/local/bin symlink its three siblings have had all
along, plus chmod and a build-time --help smoke check, so RFC-002 Phase A
censuses no longer need an absolute path.
Two new smoke assertions, both confirmed to FAIL against a v1.8.2 container so
they are not tautological: the vendored snapshot must contain the multi-machine
section (a stale manual snapshot is otherwise invisible), and mempalace-census
must be on PATH.
No other pins move: pi stays 0.84.2 (npm latest), pi-atelier v0.8.1, and every
git-ref component was checked against its upstream head and is unchanged.
The managed block already tells agents to load the mempalace skill, but said
nothing about the palace being shared with other machines. Three failure modes
observed tonight while onboarding tor-ms22's feeders, all of which do damage
rather than merely confuse:
- `mempalace sync` prunes drawers whose source files look gitignored, deleted
or moved. On a central palace that describes most of the content, including
every other machine's. Compounded by RFC-001 7.2: feeders now stage inside
the palace root, so a scoped sync can delete the drawers it just filed.
- A client-side timeout is not a failure. The palace is single-writer and one
large mine blocks every client for minutes, so
`[mempalace ext] feed (tick) failed: mine timed out after 30000ms` usually
means the mine COMPLETED. Verified: drawers from a timed-out tick were
present 35 s after the client gave up. A blind retry files a duplicate.
- The `mempalace` CLI has zero references to MEMPALACE_REMOTE_URL, so it always
opens a palace on local disk and can silently disagree with the MCP tools.
Orientation depth stays in the skill; only damage-prevention belongs here,
because this file is always read and the skill's later sections often are not.
First boot of the 2026-08-15 image on EMB-7KJ4VR4G shipped 7 transcripts to the
palace host and filed none. Two causes, neither visible in the log:
- MEMPALACE_PI_REMOTE_PATH was unset, so the feeder used its /data/feed default,
which assumes a CONTAINERIZED palace server. That fleet's primary runs
natively (systemd user unit + uv tool), so it only sees host paths and the
mine died with "source directory not found". rsync had already succeeded.
- The feeder decided success with `'"error"' in body`, but MCP escapes the
tool's JSON inside result.content[].text, so the check was blind and the
catch-up log said "Done. Wing updated." Fixed in mempalace-toolkit 6e1f4f3,
which ships `--self-test` with fixtures pinning that exact response body.
- smoke-test.sh: run `mempalace-pi-session --self-test` against the BAKED
toolkit, so a stale or reverted MEMPALACE_TOOLKIT_REF cannot reintroduce a
feeder that mines nothing while reporting success.
- .env.example: spell out that MEMPALACE_PI_REMOTE_PATH is the path the SERVER
PROCESS can open — container path for a dockerized server, and identical to
the ssh-target path for a native one — and that a mismatch fails quietly.
Follow-up to a2f0a4a, which documented the hazard; this removes it.
resolve-versions read three PUBLIC Gitea repos with `curl -sf -H "$AUTH_HEADER"`.
Gitea rejects an invalid token rather than ignoring it, so the token turned a
read that works anonymously into a hard failure:
no Authorization header 200
empty token (secret unset) 200 <- absent secret was always safe
garbage/revoked token 401 <- stale secret broke the release
A revoked GITEA_BUILD_TOKEN therefore failed resolve-versions via require_sha,
presenting as connectivity or an API fault, on data any anonymous client could
fetch. Hit exactly that failure mode today with an expired PAT.
New gitea_sha() helper tries authed, and on 401/403 retries anonymously with a
loud stderr warning naming the token as the cause. Deliberate choices:
- non-200 after the retry emits nothing and returns 0, so require_sha still
raises the explicit abort — the helper never invents a fallback ref, which is
the property the surrounding code exists to guarantee
- warnings go to stderr, NOT as ::warning:: annotations: the function's stdout
IS the SHA, so an annotation there would be captured into the ref
- the header is still sent first, so a private repo keeps working
Verified by extracting the function from the YAML step body (so the test ran the
committed text, not a copy) and calling it against live Gitea:
valid token -> 0e1369e6b496 (pi-toolkit)
REVOKED token -> warns, retries anon, f60cf9c73205 (mempalace-toolkit)
unset secret -> 98eb07bce60a (pi-extensions)
nonexistent repo -> empty + HTTP 404 warning, so require_sha aborts
Behaviour unchanged on the happy path: all three SHAs are byte-identical to the
ones the v1.8.1 release run resolved with the old curl code. `bash -n` clean on
the extracted step body; YAML re-parsed.
resolve-versions claimed "Gitea API requires auth even for public-repo commit
listing" above the pi-toolkit / pi-extensions curls. Measurably false for the
repos it guards. Verified 2026-08-15, unauthenticated vs authenticated GET of
/api/v1/repos/joakimp/<repo>/commits?limit=1&sha=main:
pi-toolkit private=false unauth=200 auth=200 sha 0e1369e6b496 identical
pi-extensions private=false unauth=200 auth=200 sha 98eb07bce60a identical
mempalace-toolkit private=false unauth=200 auth=200 sha f60cf9c73205 identical
Only /api/v1/repos/*/actions/* refuses anonymous reads with 401 — almost
certainly what the claim was over-generalised from. (Same over-generalisation I
nearly committed to opencode-devbox's AGENTS.md today; 69fc80a there narrowed it
to the actions endpoints for the same reason.)
Checked the three repos actually queried rather than reusing the pi-devbox
result — if any had been private the comment would have been TRUE, and the
correction wrong.
Behaviour deliberately unchanged: the header still gets passed. It survives a
repo being flipped private, and an unset secret degrades cleanly because Gitea
ignores an empty `token ` value and serves anonymously:
no header 200
empty token (secret unset) 200
garbage token 401
That last row is the fragility now documented: a REVOKED or malformed token
returns 401 where anonymous returns 200, so a stale GITEA_BUILD_TOKEN converts a
healthy public read into a require_sha failure that presents as an API or
network fault. Encountered exactly that today with an expired PAT on the actions
endpoints, so the note tells the next reader to suspect the token first.
Comment-only: no non-comment line changed, YAML re-parsed.
v1.8.0 never shipped: smoke (67/68) and smoke-studio (70/71) each failed the
same single assertion, so build-variant and everything downstream skipped and
latest stayed on v1.7.0.
The assertion was wrong, not the product. It grepped for a literal
stage=/home/developer/.mempalace/pi-stage/, but run() invokes
docker run --rm --entrypoint="" "$IMAGE" sh -c "$cmd"
and neither Dockerfile sets USER or ENV HOME — the published base image config
has no HOME at all; it is normally set by entrypoint-user.sh, which
--entrypoint="" skips on purpose. So the assertion executed as root with
HOME=/root, mempalace-pi-session correctly resolved
stage=/root/.mempalace/pi-stage/... (the stage is $HOME-relative by design), and
the literal grep could never match under any circumstances.
The tell was one line below in the log: the sibling assertion "pi stage follows
MEMPALACE_PALACE_PATH" PASSED, because it sets the variable explicitly and never
consults HOME. Default fails + explicit passes = wrong HOME, not broken staging.
Now asserts the invariant actually intended — the stage sits beside the resolved
palace, sharing its lifetime — which is user-independent:
case "$stage" in "stage=$HOME/.mempalace/pi-stage/"*) exit 0 ;; *) exit 1 ;; esac
$HOME is expanded by the container's own shell, so it holds as root, as
developer, or under any future user. Verified all four cases against the real
bin/mempalace-pi-session by extracting the committed assertion bodies and
running them under sh -c: virgin HOME -> exit 0; HOME=/home/developer -> exit 0;
MEMPALACE_PI_STAGE pinned to a .cache path -> exit 1 (the regression this
assertion exists to catch still fails it); developer-identity companion -> 0.
Added that companion assertion, "pi stage is palace-adjacent for the developer
user", which covers the deployment-specific path properly by SUPPLYING
HOME=/home/developer rather than assuming it.
Why this took a release to surface: docker-publish.yml triggers on push tags v*
only. The assertion was added on a push to main (7c00dd6), where only lint.yml
runs, so v1.8.0 was its first execution ever. Any smoke assertion written
outside a release was unvalidated until a release consumed it.
New workflow_dispatch input smoke_only probes/builds the base, runs both smoke
jobs against HEAD, and stops before publishing. Implemented as
`if: inputs.smoke_only != 'true'` on build-variant and build-variant-studio,
deliberately WITHOUT always() so the implicit needs-succeeded gate survives and
a red smoke still blocks a release. promote-base-latest and update-description
already require build-variant success, so they skip on their own. On a tag push
inputs is unset and null != 'true' is true, so releases are unaffected.
Finally, run() no longer discards output. A red ❌ carried zero diagnostic
weight: explaining this one-line failure needed a CI-log dig plus a registry
image-config inspection, when the container had already printed the answer.
Failures now show the last lines of output (guarded with `if`, not a trailing
`&&`, which would abort under set -e), and the stage assertions echo the
resolved stage and the HOME they saw to stderr — invisible while they pass.
pi 0.84.2 closes the Amazon Bedrock tool-argument poison pill that v1.6.4
recorded as "Not fixed upstream". pi-ai 0.84.2 adds a recursive
sanitizeBedrockDocument() and applies it at exactly the site that entry named
(dist/api/bedrock-converse-stream.js, line 692 -> 704; upstream PR #7882):
- toolUse: { ..., input: c.arguments },
+ toolUse: { ..., input: sanitizeBedrockDocument(c.arguments) },
It strips object members whose key is the empty string, recursing through arrays
and nested objects. It runs at request-build time, so it covers the live turn and
a resume alike: a session already bricked by an empty-key tool argument now
replays instead of dying on a Bedrock ValidationException. pi-session-repair is
therefore no longer the recovery path on this image -- it stays useful for older
images and for inspection, since the fix sanitises what is sent, not what was
recorded.
Bumping PI_VERSION is the ONLY way to get that fix: pi publishes an
npm-shrinkwrap.json, so pi 0.84.1 pins pi-ai to exactly 0.84.1 even though its
package.json range (^0.84.1) would admit 0.84.2. Transitive upstream fixes never
leak into this image.
pi-atelier v0.8.1 is the matching companion -- both sides changed fullscreen
input handling within three days. Its only code change (src/split-pane.ts) stops
atelier writing its own 1002h/1006h pair around a sidebar resize under pi's
fullscreen renderer, which had been tearing down the mouse reporting pi itself
enabled and leaving the wheel dead.
Audited against the surfaces the pin policies name:
- session .jsonl: identical migrateV1ToV2/migrateV2ToV3 ladder
- node engine floor: unchanged >=22.19.0 (image ships 22.23.2)
- atelier's three private couplings all intact in pi-tui 0.84.2 --
class TuiAltScreen extends TuiBase (detected by constructor NAME, so a
rename would fail silently), inputListeners still `new Set()` at the same
line 103, own render(width) descriptor still present
- pi's mouse sequences byte-identical between 0.84.1 and 0.84.2
- atelier metadata unchanged: engines >=22.19.0, peerDeps >=0.80.7, zero
runtime deps, so the "no npm install step" note holds
Not proven by execution: CI smoke does not drive the TUI, and 0.84.2 adds a
focused fullscreen search overlay that also participates in input handling. The
pairing is reasoned from the diffs. Worth an alt+a plus a sidebar resize and
wheel scroll in fullscreen on first use.
When MEMPALACE_REMOTE_URL is set but MEMPALACE_PI_SSH_TARGET is not, the feeder
has nowhere to ship staged transcripts, so skipping is correct. The problem was
that the branch was a bare `:` AND the skip happens before the subshell that
writes ~/.pi/agent/mempalace-catchup.log -- so a container in that state
contributed nothing to the palace and left no artifact at all, not even an empty
log, to explain why. It is indistinguishable from a healthy run that had nothing
to file, which is the worst property a memory system can have: the failure looks
exactly like success.
Surfaced while flipping the first client onto the shared palace, where this is
the single most likely way to end up quietly memory-less -- the palace is the
only thing that survives a container recreate.
The notice goes to both the container start output (docker logs) and the log
path anyone debugging looks at first. It names both variables, says what still
works (MCP tools read/write the shared palace; only this container's own
conversations go nowhere), and points at MEMPALACE_FEED=0 for anyone who meant
it -- "HTTPS first, mining later" is a documented interim state, so the notice
has to be silenceable without being ignorable.
Guarded against becoming a startup failure. mkdir -p in a branch that previously
touched no filesystem is a new risk: an unwritable ~/.pi (root-owned volume, a
classic Docker accident) fails under set -e and would abort the entire
entrypoint. It now degrades to stdout-only. Verified all five paths by executing
the extracted block under `set -euo pipefail`: the trap prints and writes the
log; unwritable ~/.pi still exits 0 and still prints; MEMPALACE_FEED=0 stays
completely silent (no message, no file); and both normal-remote and local mode
still background the feeder with no notice.
Two smoke assertions guard against a regression to the silent no-op. They test
the entrypoint as shipped in the image rather than behaviour, because this branch
only runs at container start and a `docker run` one-shot cannot reach it.
entrypoint-user.sh is COPY'd in Dockerfile.base, so this rides the base rebuild
the Unreleased feeder work already needs.
The feeder now defaults to <palace-root>/pi-stage upstream, so pinning
MEMPALACE_PI_STAGE into ~/.pi here is unnecessary -- and was actively wrong. It
created a second convention that could still diverge from the palace: keep the
devbox-palace volume, drop devbox-pi-config, and a scoped `mempalace sync`
prunes every conversation drawer, because dedup keys on the staged path. Both
the ENV and the entrypoint export are gone; a comment explains why adding one
back re-introduces the split it was meant to fix.
docker-compose.mempalace.yml was broken on mempalace 3.6.0 in both directions:
- `--host 0.0.0.0` with no token in the environment makes the server refuse
to start, crash-looping under `restart: unless-stopped`.
- Supply a token and the healthcheck's unauthenticated `tools/list` POST 401s,
marking a perfectly healthy server unhealthy forever.
Now the token is required via ${MEMPALACE_REMOTE_TOKEN:?...} so it fails fast at
`docker compose up` with a readable message, and the healthcheck probes the
deliberately token-free /healthz. The "no authentication of its own" security
note has been stale since 3.6.0 and is replaced with the actual posture
(bearer token + Host pin + Origin allowlist), including why browser-shaped auth
must not be put in front of it.
Dockerfile.base: mempalace-pi-session symlinked onto PATH, with a build-time
`--help` check so a broken feeder fails the image build rather than the first
session.
smoke-test: assert the stage resolves beside the palace (default, and following
$MEMPALACE_PALACE_PATH) instead of asserting the removed ENV pin. The two
behavioural guards -- a synthetic session that must be captured, an abandoned
one that must not be -- are unchanged.
.env.example: recommend `mempalace serve` on the docker0 gateway rather than
`mempalace-mcp --transport http --host 0.0.0.0`, with the two binds to avoid.
Both keys shipped empty with no explanation, so they read as optional. They
are not: ~/.gitconfig is not on a persistent mount, so with these unset every
repo in the container fails "Author identity unknown" on first commit after
every recreate — and an agent asked to commit then infers an identity from
git log and picks the wrong one. Observed today across three repos on one
machine (two wrong-address commits, plus older `pi <pi@devbox>` fossils from
earlier sessions guessing the same way).
Also records that the address is per-machine (corporate vs private), so it
belongs in the per-machine .env rather than a skill or a repo-local override.
`install -d -m 700 ~/.ssh` was correct but wrong for the audience. This snippet
gets pasted onto an arbitrary peer — a NAS, a router, a BSD box — and `install`
is not in POSIX, so it is not guaranteed to be there. `mkdir -p` + `chmod` is
POSIX, present everywhere, and self-evidently idempotent to a reader deciding
whether it is safe to run on a machine that already has keys.
Behaviour was checked rather than assumed, on both coreutils and BSD/macOS
`install`: on an existing ~/.ssh it exits 0 and leaves authorized_keys intact
in content and mode, but it also silently chmods the directory (755 -> 700).
Desirable here, yet invisible in a doc — which is the second reason to prefer
the explicit two-step form, and why the surrounding text now states outright
that re-running is safe on an already-configured peer: mkdir -p is a no-op, the
chmods only tighten, and appending never touches keys already listed.
Verified the whole block end-to-end against a fresh HOME and one with a
pre-existing 755 ~/.ssh and an older key: 700/600 in both cases, older key
preserved, new line appended.
Step 2 of "Giving the container its own key for a peer" showed a correctly
narrowed authorized_keys line but never named the file it belongs in, and never
said which account's — the only mention of authorized_keys was an aside 35
lines further down about revoking one key per machine. A reader following the
steps had a public key, a line to construct, and nowhere to put it.
Now explicit: append to ~/.ssh/authorized_keys of the account named as `User`
in step 3, via a heredoc that shows `>>` rather than `>` (the truncation that
revokes every other key on that account), with `install -d -m 700 ~/.ssh` and
`chmod 600` so the file is created correctly the first time.
Adds the three failure modes that make a good key look broken, none of which
announce themselves on the client: the options prefix must be on the same
physical line as the key (a wrapped paste is the usual culprit), `ssh-copy-id`
cannot add that prefix at all so the line must be appended by hand, and sshd's
StrictModes silently ignores authorized_keys when the home directory, ~/.ssh or
the file is group- or world-writable — reporting it only in the peer's own log.
Two changes that belong together, because the first is what makes the second
dangerous to get wrong.
pi-atelier (TUI sidebar + status rail) is now vendored to /opt/pi-atelier at
PI_ATELIER_REF=v0.8.0 and registered by entrypoint-user.sh — the pi-fork /
pi-observational-memory / pi-studio pattern, deliberately NOT
`pi install npm:pi-atelier`, which writes into ~/.pi/npm-global on the config
volume where it shadows the image and pins nothing. Unlike its siblings it gets
no `npm install`: atelier declares zero runtime deps (peerDeps only, satisfied by
the baked pi) and has no build step, so pi loads its TypeScript straight from the
checkout via package.json `pi.extensions`.
pi is no longer resolved to npm `latest` at build time. The pin lives in
Dockerfile.variant and CI reads it from there, so a local `docker build` and a CI
release ship the same versions by construction. The pin is a CHECKPOINT, NOT A
FREEZE: bumping stays a one-line change; what stops is *unreviewed* adoption of
whatever shipped that morning, in the same build that then gets tagged and
published. CI fails when a pin is not concrete or not actually published on npm,
and warns — never adopts — when npm latest moves ahead, naming what to re-check.
Why this pairing needed care: pi-atelier 0.6.0/0.7.0 wrap pi's PRIVATE TUI
renderer, and under pi 0.84 that wrapper recurses — pi hangs at startup burning
CPU with no error. Upstream fixed the recursion in 0.7.1 and restored the
non-overlapping split in 0.7.2; 0.8.0 is additive on top. atelier's own
peerDependencies still say >=0.80.7, which does not express that floor, so
nothing in npm metadata could have warned us. The floor is therefore encoded as
an executable rule — pi >= 0.84 => pi-atelier >= 0.7.1 — asserted in both
smoke-test.sh (build time) and recreate-sanity-check.sh (after a real recreate),
verified against a 4x4 version matrix.
Existing volumes needed migration, not just vendoring: a hand-installed
`npm:pi-atelier` entry is counted as already-registered by the entrypoint guard,
so every existing volume would have kept its unpinned npm copy — and a 0.6.x copy
next to pi 0.84 is exactly the startup hang. The entrypoint now drops that one
exact string (settings.json.bak.atelier.<ts> backup, distinct prefix so it cannot
clobber the template merge's backup in the same second) and lets the pinned /opt
copy register. Tested against a real settings.json: only that entry removed,
other packages and all keys intact, idempotent, and unparseable JSON leaves the
file untouched. DEVBOX_ATELIER=0 opts out entirely — in the entrypoint rather
than via `pi uninstall`, because this component's failure mode is "pi will not
start", which cannot be repaired from inside pi.
0.84.1 was audited for this release, not merely adopted: theme/TUI changes are
additive, the session format is unchanged (CURRENT_SESSION_VERSION = 3 in both
0.83.0 and 0.84.1 with an identical migrateV1ToV2/migrateV2ToV3 ladder, so
existing transcripts are neither migrated nor at risk and pi-session-repair stays
valid), and the Node engine floor is unmoved at >=22.19.0. CI resolves the
atelier tag to its PEELED commit SHA — atelier uses annotated tags, so the
unpeeled ref is a tag object, not a commit; pi-studio's lightweight tags never
exposed that distinction.
Also: docs for overriding the read-only ~/.ssh/config from the container —
container-only keys in ~/.ssh-local, hardened authorized_keys, the fact that
`from=` must allow the HOST's addresses because container egress is NAT'd through
it, and the macOS-only-keyword trap (`UseKeychain` is fatal to Linux OpenSSH and
takes out dssh/pi --ssh while the host keeps working). Corrects two claims in
"Naming LAN peers": ssh-lan.conf is not ProxyJump-only, and first-time creation
does need one restart because the Include is emitted only when the file already
exists at start.
The previous commit justified excluding tag pushes partly on runner contention:
that the duplicate lint run stole one of two self-hosted runners from the release
build. Measured, that is false for THIS repo — pi-devbox lint runs take 0.3-0.9
min (ids 529/531/532/533) against a 77.6 min release build (id=530). I imported
the claim from opencode-devbox, where actionlint apt-installs shellcheck inside
the container and takes 6-15 min, so contention there is real.
The change stands on its actual merits: duplicate lint of an identical tree, and
release-run discovery ambiguity (the substantive one — it is what made the naive
"first run matching refs/tags/<tag>" rule pick lint over the publish run).
No functional change; comment only.
`on: push:` with no filter also fires on refs/tags/v*, which is duplicate work:
the tagged tree was already linted when that same commit was pushed to main
(v1.6.4 sha e86e5df linted as id=529 on main, then again as id=531 on the tag).
Two costs beyond the wasted run. It consumed one of the two self-hosted runners
while the release pipeline wanted both for its parallel multi-arch variant
builds; and it made release-run discovery ambiguous, since the newest-first runs
listing puts the tag-ref lint run above the publish run.
`branches: ['**']` keeps the documented intent exactly — lint fires early on
every branch push and PR, rather than only at tag time — while excluding tag
refs. docker-publish.yml is untouched and still tag-scoped.
The release-day checklist said "Watch CI" without saying which run, and the
Gitea API example used limit=5. Both are traps, because a tag push produces
TWO runs here: lint.yml has a bare `push:` trigger so it fires on the tag ref
as well, and docker-publish.yml fires on v*. The runs listing is newest-first
and the lint run sorts ABOVE the publish run, so "first run matching
refs/tags/<tag>" picks lint reliably. Verified against the real API for v1.6.4:
id=531 #104 lint.yml@refs/tags/v1.6.4 <- picked by the naive rule
id=530 #103 docker-publish.yml@refs/tags/v1.6.4 <- the actual release build
id=529 #102 lint.yml@refs/heads/main <- same sha, already linted
Lint goes green in minutes while the image is still building, so watching it
makes a release look finished before anything is published. limit=5 compounds
it: the publish run is already at position 4 of 5 in the current listing.
Documents: head_sha-filtered discovery with limit=20; the jobs endpoint takes
the internal id, never the run_number (silently returns another run's jobs);
and the correct ci-release-watcher config for this repo — EXPECT_WORKFLOW,
the studio tag pair, base-latest as existence-only, and CRITICAL_JOBS with
build-variant-studio spelled out (job names are matched exactly, and the
skill's default omits it) while excluding promote-base-latest, which
legitimately skips on a base cache hit.
Smoke-gate detail in step 5 is retained.
3.6.0 (2026-07-17, PyPI latest) is additive/reliability only: secure
`mempalace serve` remote mode, optional Milvus backend, atomic KG
supersede(), conversation chronology, mining exclusions, plus recovery and
locking fixes.
Reviewed for MCP tool-schema changes before bumping — that being the exact
regression class this pin exists to catch, after an unpinned install once
swept in the broken 3.3.x/3.4.0 diary_write schema. There are none, and
nothing touches diary_write, so the perl workaround removed in v1.2.2 stays
removed.
Two fixes matter for how this image uses mempalace: read-only mode now covers
checkpoint + delete_by_source in _MUTATING_TOOLS (#1930), and agent
attribution is preserved in mempalace_checkpoint (#2023/#2034) — the latter
because the diary protocol relies on per-agent attribution.
Also adds a CHANGELOG Unreleased block that backfills the per-variant image
description labels (1fd524e), pushed after the v1.6.4 tag without an entry.
Not tagged: more changes are queued for the next release. Pushing to main
triggers only lint.yml (actionlint + hadolint) — the image build/publish
workflow is tag-only. Verified clean against the CI-pinned hadolint 2.14.0.
Both published variants inherited Dockerfile.base's
description="pi-devbox — base image (variant-independent)", so v1.6.4 and
v1.6.4-studio both advertised themselves on Docker Hub as the base image —
misleading, and useless for telling the two apart.
A LABEL cannot branch on INSTALL_STUDIO, so the text arrives as a build-arg:
CI passes a variant-specific string (interpolating RELEASE_TAG, PI_VERSION and,
for studio, STUDIO_TAG), and the Dockerfile default keeps a bare local
`docker build -f Dockerfile.variant` honest instead of misleading.
Also sets org.opencontainers.image.title/description alongside the legacy bare
`description` key, so Hub and OCI-aware tooling both see it. ARGs stay in the
last-declared block, so the label layer is still the only thing invalidated.
Verified: hadolint 2.14.0 (the CI-pinned version) clean on both Dockerfiles;
workflow YAML parses; check-workflow-shell.sh passes. Lands on the next release.
Headline is the fork-guard fix (the `fork` tool had never loaded since v1.0.0
because the registration guard grepped the whole settings.json and matched the
pi-fork *config block* the template merge itself plants).
pi 0.82.1 -> 0.83.0 is a clean hop with no intermediates. 0.83.0 carries an
upstream Breaking Change (bundled TypeBox 1.1.38 -> 1.3.7, deprecated APIs
removed) that cannot reach us: pi-fork vendors @sinclair/typebox (a different
package name), obsmem's use is type-only, studio and atelier don't use TypeBox.
Verified from source rather than from the audit narrative: extension-facing
dist/core/extensions/*.d.ts declarations diff clean between the two versions,
all six CLI flags pi-fork passes to child processes are present, and
SESSION_VERSION is 3 in both so transcript tooling is unaffected. No PI_VERSION
pin needed.
Recorded for the release notes: the Bedrock poison pill is NOT fixed in pi-ai
0.83.0 (unsanitised input: c.arguments moved 634 -> 644), so pi-session-repair
stays the recovery path.
Vendored floor snapshot re-synced from pi-extensions 98eb07b, which documents
the mechanism behind fork boundary violations (full parent-transcript
inheritance via index.ts:47) plus the corrected claims about tool restriction
and narrative invention. CI resolves PI_EXTENSIONS_REF from main HEAD, so a
normal build ships the package-owned copy; this keeps the committed floor
identical so scripts/smoke-test.sh's cmp assertion holds either way.
Closes the drift found in 4d4abd9's investigation: the Temporal grounding
guidance was authored straight into this vendored fallback (904fe85) and never
returned to skillset, the declared owner. skillset 63f3bf5 now carries it (with
the wording generalized from "a pi-devbox container" to "a devbox container
(pi-devbox or opencode-devbox)", since opencode-devbox vendors the same skill),
so this snapshot is a pure mirror of the owner again.
VENDORED.md: refresh commands now copy each snapshot FROM ITS OWNER. The
pi-extensions lines previously pointed at <skillset>/skills/pi-extensions/,
which has been a downstream duplicate since a7f3044 co-located the canonical
skill in the package repo — following the old instruction would have silently
regressed the snapshot (e.g. undoing pi-extensions e73cb9f). Provenance bumped
to skillset 63f3bf5 / pi-extensions pkg e73cb9f.
~/.agents/skills/ lives in the ephemeral container layer and is rebuilt by
entrypoint-user.sh on every start from two sources, so an edit made through the
symlink may vanish on the next recreate. Adds to §1 (persistence tiers):
- `readlink -f ~/.agents/skills/<name>` as the first move, with a tier table:
resolves under /workspace/skillset → edit in place; resolves under
/usr/local/share/pi-devbox/skills → image layer, edit the canonical repo and
`sudo cp` to activate for the running session.
- Canonical owner per baked skill (pi-devbox-environment → this repo;
pi-extensions → the package repo's skill/, plus this repo's floor snapshot;
mempalace → the private skillset repo), pointing at VENDORED.md as
authoritative.
- The shadowing gotcha: image-baked links are created first and only when
absent, and deploy-skills.sh --prune-stale leaves foreign links alone, so for
a name present in BOTH sources the image copy wins and a skillset edit has no
effect in the container. Documented with the live example found while writing
this: the baked mempalace snapshot carries a Temporal grounding section
(904fe85) that skillset at its snapshot point (8e8db64) lacks.
- Checklist gets a matching line.
Found while adding session findings to the pi-extensions skill (e73cb9f), where
the same resolve-first step was what kept the edit out of the image layer.
Sync the image-baked "floor" copy at
rootfs/usr/local/share/pi-devbox/skills/pi-extensions/SKILL.md with
pi-extensions e73cb9f, which documents package-registration forensics
(packages[] vs whole-file grep), that /reload suffices for a newly installed
package, and a fork-output caveat — the findings from the fork-guard bug fixed
in 8248688.
Dockerfile.variant copies the pinned package's skill/ over this snapshot at
build time, so the vendored copy is only the fallback floor; it was byte-
identical to the package copy before this change, and keeping it in sync
prevents a silent divergence for builds whose PI_EXTENSIONS_REF predates the
skill.
The `pi install /opt/<pkg>` loop in entrypoint-user.sh guarded on a
whole-file substring grep of ~/.pi/agent/settings.json. settings.example.json
ships a top-level "pi-fork" CONFIG block (fork effort profiles, pi-toolkit
adb6907, 2026-06-17), so `grep -q pi-fork settings.json` matched the config
key itself and `pi install /opt/pi-fork` never ran — on fresh or preserved
volumes. The `fork` tool has therefore been absent since v1.0.0.
The non-destructive template merge runs earlier in the same startup than the
install loop, so the mechanism that delivers new template keys to an old
volume is what plants the string that defeats the guard. pi-observational-
memory and pi-studio escaped only by luck: the template key is
"observational-memory" (no pi- prefix) and there is no studio block.
Guard now inspects the `packages` array via jq, with a grep fallback on the
stored `.../opt/<name>"` path form, which a config key can never produce.
Existing volumes self-heal on the next container start.
Both test suites asserted the bug as green — smoke-test.sh:244 and
recreate-sanity-check.sh:204 used the same whole-file grep, so "pi-fork
registered (fork tool)" passed on every build and recreate while the tool was
missing. Both now assert against packages[] with the entrypoint's predicate,
labels say packages[], and the smoke readiness wait loop uses the array check
plus `docker exec -u developer` + $HOME instead of a hard-coded
/home/developer path.
Evidence: zero `fork` tool calls across all 19 sessions on this volume; the
v1.6.3 session that tuned pi-fork.deep to opus-5 was configuring a tool that
never loaded.
Records the pi-toolkit 926f738 template change (defaultModel + pi-fork deep
tier -> eu.anthropic.claude-opus-5). No tag, no build: CI resolves the
pi-toolkit SHA from main at build time, so whichever release builds next
picks it up.
Pure pi version bump (variant-only rebuild). CI resolves pi@latest=0.82.1 at build. 0.82.0/0.82.1 audited: additive, no breaking changes to the extension execution API (pi-observational-memory) or pi-agent-core types (pi-fork).
Completes v1.6.1's studio publish. Run 512 shipped v1.6.1 non-studio
(3411 MB) cleanly, but smoke-studio failed the size gate at 3574 MB vs
the 3500 MB threshold. Threshold was set in v1.0.0 pre-agent-browser
(baseline was 3.20 GB local arm64 + 300 MB margin); v1.6.0 baked in
agent-browser + Chromium (+~291 MB net) but the threshold was never
lifted. v1.6.0 never got to smoke because of the network fault, so
nothing surfaced this until v1.6.1's smoke-studio.
Bump SIZE_THRESHOLD_MB to 3800 (~225 MB margin above observed studio
number, tight enough to still catch a genuine +GB regression). Refresh
the comment above the constant with the current baseline + run 512
actuals so future readers know where the number came from.
CI-only change; image bytes identical to v1.6.1 except for the manifest's
release_tag/source_revision. Not base-affecting.
The smoke workflow deliberately passes RELEASE_TAG=smoke to the variant
build so smoke images don't collide with real vX.Y.Z tags. The variant
bakes that into /etc/pi-devbox/build-manifest.json, and pi-devbox-version
prints 'pi-devbox smoke' — correct behaviour. But the smoke assertion
required substring 'pi-devbox v', which only holds for real releases.
Assertion never fired before because pi-devbox-version was added after
v1.5.0 (fb49828, 2026-07-15) and every CI attempt since was blocked
before smoke ran (v1.6.0 network flake, v1.6.1 first successful base
build hit this). CI run 509 (v1.6.1) surfaced it.
Fix: require the substring 'pi-devbox ' (space, no v). The two
neighbouring assertions on --json and --quiet already cover the value
of release_tag; this one just verifies the human line renders.
Everything else in 55/56 checks passed on run 509 including
pi 0.81.1 reported and base-229f04e5d021 pushed OK.
v1.6.0 was tagged 2026-07-17 but never reached Docker Hub (variant
publish blocked by a site-network SYN-drop fault, since fixed). Cut
v1.6.1 to land v1.6.0's content (agent-browser + pi-devbox-version)
alongside a first pi bump since v1.5.0.
0.81.0 is skipped deliberately: it removed the default stream fallback
for extensions using the pre-0.81 pi-agent-core API, which
pi-observational-memory relies on via agentLoop + stream.result().
0.81.1 restored the fallback (earendil-works/pi#6915), so 0.81.1 — but
not 0.81.0 — is a safe drop-in. pi-fork only imports types; unaffected.
Audit of 0.80.7–0.81.1 vs the two baked extensions and the base image:
no breaking changes affect pi-devbox. Node engine bumped to >=22.19.0
in 0.81.0 (nodesource 22.x currently 22.23.1, so no engine bump needed).
Base-affecting via the npm install line, so base-<hash> rebuilds.
v1.6.0's first build failed on amd64: playwright@latest fetches Chrome for
Testing, which extracts to chromium-<rev>/chrome-linux64/ on amd64 (vs
chrome-linux/ on arm64). The hardcoded chrome-linux/ glob matched only arm64,
so amd64 built no /usr/local/bin/agent-chrome symlink and 'test -x' failed
(agent-browser --version had already printed 0.32.1 — the tell).
Replace the glob with an arch-agnostic 'find -name chrome -path */chromium-*/*'
(skips the chrome-headless-shell binary and the chromium_headless_shell dir),
guarded by [ -n ] + the existing test -x. Verified locally: resolves the CfT
chrome and agent-browser drives it headless. hadolint clean.
Roll the unreleased changes into v1.6.0 (minor — significant base addition:
agent-browser + headless Chromium baked into every variant for browser
automation/front-end verification). Also includes the pi-devbox-version command
and the bundled pi-toolkit sonnet-5 template bump. No pi bump — CI resolves
latest pi from npm at build time (verified: 0.80.6→0.80.10 has no breaking
changes affecting pi-devbox).
Size concern ahead of the v1.6.0 base rebuild. agent-browser drives the full
chrome (verified headless: open/title/eval with the headless_shell removed), so
Playwright's chromium_headless_shell-* build is dead weight — drop it (~334MB),
and clean apt/npm caches in-layer. Net browser footprint ~625MB/arch.
CI disk is otherwise fine: build-base's 'Reclaim runner disk' step frees
~20-30GB (hostedtoolcache/dotnet/android/jvm) + docker prune before buildx.
hadolint clean.
Discoverability follow-up to baking agent-browser into the base: agents won't
reach for a browser they don't know they have (the exact gap that made this
capability easy to miss). Add a short pointer section to the pi-devbox managed
block — capability, the preset AGENT_BROWSER_EXECUTABLE_PATH, and
`agent-browser skills get core --full` for the command set. Pointer only; depth
stays in the skill. rootfs change → folds into the same base-<hash> rebuild.
CHANGELOG note updated.
Gives the agent a real browser it can drive so front-end work involving live
DOM/WebGL can be VERIFIED, not guessed. The agent-browser skill (from the
skillset repo) was a no-op without the binary; it now works out of the box.
- agent-browser CLI (standalone Rust, ships no browser) via npm, prefixed
NPM_CONFIG_PREFIX=/usr so it survives the ~/.pi/npm-global volume.
- Chromium fetched with 'playwright install --with-deps chromium' into
PLAYWRIGHT_BROWSERS_PATH=/usr/local/share/ms-playwright — a system path that
the /home/developer volume can't shadow (unlike agent-browser's own
~/.agent-browser/browsers default, which WOULD vanish on recreate).
- Stable /usr/local/bin/agent-chrome symlink, exported as
AGENT_BROWSER_EXECUTABLE_PATH, insulates the ENV from Playwright's
per-version chromium-<rev> dir name.
Verified end-to-end (this session): agent-browser drives the baked Chromium
headless (open + title + screenshot + eval into a WebGL SPA); doctor launch
test passes in ~0.5s. Debian trixie --with-deps resolution verified (exit 0;
t64 lib renames handled). hadolint clean; check-base-hash OK (only *_VERSION
args added). Cost ~960 MB (Chromium + headless shell). Base-affecting →
rebuilds base-<hash> on next release.
Accented filenames on a macOS host are stored decomposed (NFD), so a precomposed
(NFC) remote path in scp/dscp silently fails with 'No such file or directory'.
Document the wildcard / list-first workaround in the pi-devbox-environment skill,
next to the dssh/dscp alias table. (Hit while copying a screenshot named
'Skärmavbild ….png' from the host.)
Wraps /etc/pi-devbox/build-manifest.json (already written at docker-build
time in Dockerfile.variant) into a human-readable summary instead of
requiring users to know the manifest path and pipe it through jq
themselves.
- rootfs/usr/local/bin/pi-devbox-version: human (default) / --json /
--quiet output modes. Also flags live drift — compares the baked
pi_version against the actually-running `pi --version` and warns
on mismatch rather than trusting the manifest blindly (same
ground-truth philosophy as the manifest generation itself). Exits 1
with a short stderr notice on images built before the manifest
existed, instead of failing silently.
- entrypoint-user.sh: calls it as the very first line. CMD is
`bash -l` with tty:true in compose, so this banner lands directly
above the first prompt on container start — no separate motd/bashrc
hook needed (deliberately not wired into .bash_aliases, which would
reprint on every docker exec -it).
- Dockerfile.base: COPY + chmod, same pattern as dot-watch/studio-expose.
- scripts/smoke-test.sh: 4 new checks (binary present+executable, human
output has release tag, --json round-trips the manifest, --quiet is
a single line).
- README.md / AGENTS.md / CHANGELOG.md updated.
Roll the 7 unreleased commits since v1.4.0 into v1.5.0 (minor — significant
base additions: readable Neovim true-colour, broad terminal terminfo support).
Also includes: typst PDF template font fix, .claude gitignore seed, pi-studio
semver-tag pin + version label, and repo hygiene (LICENSE, THIRD_PARTY.md,
.dockerignore, hadolint lint, IDEAS backlog). No pi bump — CI resolves latest
pi from npm at build time as usual.
The base only shipped ncurses-base (xterm-256color, tmux), so SSHing in from
WezTerm/Alacritty/foot/Ghostty degraded to a dumb TERM fallback. Following the
maintainer's ansible common role:
- Dockerfile.base: install ncurses-term (terminfo for wezterm, alacritty,
foot, st, base ghostty entry, and many more).
- rootfs/.../terminfo-src/ghostty.terminfo: thin xterm-ghostty alias
(use=ghostty) — Ghostty connects as TERM=xterm-ghostty, which no distro
packages. Compiled into the system db with 'tic -x'; build asserts it
resolved via infocmp.
- kitty (xterm-kitty) already covered by kitty-terminfo; iTerm2 uses
xterm-256color (ncurses-base).
- smoke-test: assert ncurses-term emulators + xterm-ghostty alias resolve.
Validated end-to-end in a throwaway container: ncurses-term brings the
entries, the alias compiles and equals the ghostty capability set. hadolint
clean, bash -n OK. Base-affecting, rebuilds base-<hash>. No tag.
Repo/CI hygiene batch (none base-affecting; image contents unchanged):
- LICENSE: actual MIT file (repo previously declared MIT only in prose).
- THIRD_PARTY.md: notes bundled software + licenses (pi/pi-fork/pi-obsmem/
pi-studio MIT, gosu Apache-2.0, Debian packages under their own terms).
- .dockerignore: trims build context to what the Dockerfiles COPY (rootfs/ +
entrypoint*.sh); keeps .git/docs/scripts/compose out. Verified it excludes
none of the required COPY sources.
- lint.yml: new hadolint job (pinned v2.14.0) lints both Dockerfiles;
.hadolint.yaml grandfathers deliberate choices (DL3008/DL3016/DL4006/DL3003/
SC2086, mirroring the shellcheck excludes), fails on anything new at warning+.
Verified hadolint exit 0 and the repo shell-guard passes with the new job.
- IDEAS.md: parks deferred follow-ups (SHA-pin actions, trivy, buildx SBOM/
provenance, Makefile, renovate).
- README/DOCKER_HUB License sections now link LICENSE + THIRD_PARTY.md.
No tag.
Upstream omaclaren/pi-studio stopped publishing GitHub Releases at v0.5.55
but keeps tagging every version (v0.9.36 now) and pushing to main. Tracking
main HEAD risked baking half-finished commits that land after a tag.
resolve-versions now lists all tags via a single git ls-remote (the REST
tags API paginates at 100 and the repo has >140 tags, so a single page can
miss the newest), picks the highest X.Y.Z with sort -V (pre-releases
excluded), and peels it to a commit SHA for PI_STUDIO_REF. SHA (not moving
tag) keeps cache-busting + reproducibility; require_sha still enforced.
Studio-variant only, not base-affecting. No change to the resolved commit
today (v0.9.36 == current main HEAD). Validated: yq parse OK, bash -n OK,
live git ls-remote -> v0.9.36 -> 2ef38ef. No tag.
Vanilla Neovim fell back to a 256-colour palette over ssh/kitty and rendered
strings/comments in a muddy, low-contrast colour. Fix, all base-affecting:
- rootfs/etc/xdg/nvim/sysinit.vim: system vimrc enabling termguicolors. Loads
for every user before any personal ~/.config/nvim; overridable per-user.
- Dockerfile.base: install kitty-terminfo (TERM=xterm-kitty understood; lets
Neovim auto-detect true colour) and set ENV COLORTERM=truecolor.
- smoke-test.sh: assert xterm-kitty terminfo present and nvim tgc default on.
- CHANGELOG (Unreleased/Added) + README editor note.
No tag — lands on the next base-<hash> rebuild.
Claude Code's per-machine local settings can carry credentials and should never
be committed. Add the pattern to the seeded ~/.gitignore_global so fresh
containers match a host global that already ignores it. Existing containers are
unaffected (seed copied only when absent). Base-affecting; rebuilds base-<hash>.
pandoc's bundled typst template defaults the document font to an empty
tuple (font: ()), so a naked `pandoc --pdf-engine=typst` fails with
"font fallback list must not be empty". Patch the template default to
Libertinus Serif (typst's own bundled default) at build time so PDF
export works out of the box. Document usage in README and note the fix
in CHANGELOG (Unreleased). Base-affecting.
Finalize the Unreleased batch as v1.4.0 (minor — significant base
additions). Base rebuilds (Dockerfile.base for typst/xz-utils;
.bash_aliases for the SSH check). pi auto-resolves latest (0.80.3 ->
0.80.6); mempalace stays 3.5.0 (current).
pandoc has been in the base since v1.0.0 but only as a front-end; PDF
export (studio_export_pdf / pandoc -o out.pdf) failed with 'xelatex not
found' because no back-end engine was installed. Ship typst (~30 MB
static Rust binary, no LaTeX) as the default engine via
`pandoc --pdf-engine=typst`, chosen over a ~600 MB TeX Live install.
texlive-xetex remains the higher-fidelity install-on-demand fallback.
- Dockerfile.base: install typst (latest GitHub-release idiom, pin via
TYPST_VERSION); add xz-utils (typst ships .tar.xz); bump
BASE_REBUILD_DATE. Lands in base-<hash>.
- smoke-test.sh: verify typst + a real pandoc --pdf-engine=typst render.
- README/AGENTS/CHANGELOG: typst now shipped (supersedes the planned
:latest-studio-tex variant).
No tag pushed — CI build intentionally deferred.
Adds a _devbox_check_host_ssh() check to ~/.bash_aliases (baked into
the image). On first bash session of each container it tries a quick
SSH probe to the Mac host; if it fails it prints a clear one-time
warning with the exact two steps needed to fix it:
1. Enable Remote Login in macOS System Settings
2. echo '<public key>' >> ~/.ssh/authorized_keys
The check is guarded:
- only runs inside a container (/.dockerenv)
- only when the jump key exists (~/.ssh-local/devbox_jump_ed25519.pub)
- only once per container lifetime (/tmp flag, cleared on recreate)
After --force-recreate the key changes, the flag is gone, and the
check runs again on the first bash window. Subsequent windows are
silent.
PDF export from Studio/pandoc still isn't shipped. Record the engine decision
in the living docs (README + AGENTS): pandoc is in the image but has no PDF
back-end, so export fails with 'xelatex not found'. Prefer a lightweight engine
— typst (~30 MB static binary, 'pandoc --pdf-engine=typst') which is small
enough it could ship in base rather than needing a separate ':latest-studio-tex'
variant; texlive-xetex (~600 MB) kept as the higher-fidelity fallback. Also drop
stale 'v1.3.0' pins (v1.3.0 already shipped without PDF) in favour of 'a future
release'. CHANGELOG history left untouched.
Wire the shared-palace option (implemented in mempalace-toolkit's mempalace.ts)
into the container:
- .env.example: document MEMPALACE_REMOTE_URL / MEMPALACE_REMOTE_TOKEN (env_file-only,
per this repo's convention).
- docker-compose.mempalace.yml: optional shared server (mempalace-mcp --transport http),
loopback-bound by default.
- docker-compose.yml: local-vs-external note on the palace-volume comment.
- README + CHANGELOG (Unreleased).
The image shipped only nvim (EDITOR=nvim), a modal vi-style editor. Not
everyone is comfortable with vi keybindings, so add both a classic and a
modern non-modal option:
- nano (apt): ~2.8 MB installed; deps (libc6, libncursesw6, libtinfo6)
already present via nvim/less/htop/tmux, so no extra packages pulled in.
- micro: ~12 MB single static Go binary from GitHub releases (same pattern
as bat/eza/zoxide). Desktop-style keys (Ctrl+S/Ctrl+Q), mouse, syntax
highlighting. ARG MICRO_VERSION pins; defaults to latest.
Combined ~15 MB (<0.5% of the ~3.2 GB image). EDITOR stays nvim; both new
editors are opt-in (export EDITOR=micro | nano). Uses the canonical
micro-editor/micro URL because the old zyedidia/micro org rename makes
/releases/latest redirect to another /latest, defeating the tag-parsing
latest-resolution idiom.
Base-image change, so it lands on the next base-<hash> rebuild. Updates
README tool table + EDITOR note, CHANGELOG (Unreleased/Added), and
smoke-test.sh (nano + micro presence checks).
actionlint's no-arg project auto-detection looks for .github/workflows
and hard-fails (exit 3, 'no project was found') on this .gitea/workflows
layout — observed on run 420. Glob the workflow files explicitly. The
Gitea shell guard step already passed in that run; only the actionlint
invocation needed the path fix.
Root cause of the recurring 'Illegal option -o pipefail' failures
(ed49b8d resolve-versions; b7197e8 promote-base-latest, run 418):
docker-publish.yml had no workflow-level default shell, so Gitea's
sh/dash default applied and every bash-syntax step had to individually
remember 'shell: bash'.
- docker-publish.yml: add 'defaults: run: shell: bash' — fixes the whole
class; all pre-existing dash steps are POSIX so bash runs them unchanged.
- lint.yml: new workflow, runs on every push/PR (not just release tags):
* scripts/check-workflow-shell.sh — Gitea-accurate guard: fails if any
run: step doesn't resolve to bash. Catches the omit-shell+bash-syntax
case that actionlint MISSES (actionlint models GitHub, where the
default shell is bash, so a shell-less step is assumed bash).
* actionlint + shellcheck — catches explicit 'shell: sh' + bash syntax
(SC3040) and general workflow errors.
Verified locally: guard + actionlint pass current workflows; guard fails
a synthetic omit-shell+pipefail workflow; shellcheck clean.
b7197e8 moved the digest-compare into the re-tag step with 'set -euo
pipefail' but no 'shell: bash'; Gitea's default sh (dash) aborts on
-o pipefail, leaving base-latest un-promoted on the v1.2.4 release
(run 418). Same footgun as ed49b8d. Consumer tags unaffected (they
FROM base-<hash>, not base-latest).
Seed ~/.gitignore_global from /etc/skel-devbox (seed-if-absent, like
.bash_aliases/.inputrc, so user edits survive recreate) and wire it via
git config --global core.excludesFile, guarded so a user-set excludesFile
is never overridden. Ignores *.bak, *.bak.*, *~, *.orig, *.swp, *.tmp
across all repos without per-repo .gitignore entries.
Removes GITEA_ACCESS_TOKEN / GITEA_HOST / GITHUB_PERSONAL_ACCESS_TOKEN from
the compose environment: block. An environment: entry both overrides
env_file AND is interpolated from the host shell, so a stale shell export
(e.g. one auto-loaded by an opencode/dotenv hook) silently shadowed the
users .env — an updated token never reached the container. Secrets now flow
solely via env_file: .env; .env.example already documents every variable.
- docker-compose.yml: drop the 3 passthrough lines + explanatory comment
- README.md: sync the "basic shape" snippet
- CHANGELOG.md: note under Unreleased (no tag bump / unpublished)
The gate keyed off need_build=='true', assuming need_build==false meant
base-latest was already current. A dry-run dispatch (promote_latest=false)
that pre-builds base-<hash> falsifies that: the later tag run sees
need_build==false and skipped promotion, leaving base-latest one base
behind (observed 2026-06-27, v1.2.3 dry-run-first release).
Gate now runs on every tag release / promote dispatch; the no-op
optimization moved into the step as a crane digest compare so it re-tags
only when base-latest actually differs from the released base-<hash>.
Workflow-only change; base hash unaffected (no base rebuild).
Patch release. Headline: mempalace-mcp self-heals instead of latching
available=false permanently after a slow virtiofs cold-open. Base image
rebuilds via the mempalace-toolkit ref advancing to e12b624 (folded into
the base-decide hash). No pi/mempalace version change — pi npm latest is
still 0.80.2 (= v1.2.2). Also releases the queued yq (mikefarah Go yq) and
mempalace-skill temporal-grounding changes.
mempalace.ts now self-heals (respawn with capped backoff) instead of
latching unavailable, and the init-timeout default is 300000. Update the
explanatory comment + tunable list (MEMPALACE_MCP_MAX_RESPAWNS,
MEMPALACE_MCP_RESPAWN_BACKOFF_MS). Comment-only; no build/ENV change.
Baked mempalace SKILL.md now instructs agents to establish current date/time
and compute the delta against the actual diary/drawer timestamp before using
relative terms (yesterday/last week), and explicitly that a container recreate
or fresh session is NOT a day boundary (pi-devbox restarts several times a day).
Phase 1 wake-up section + anti-pattern bullet. CHANGELOG Unreleased.
Match the repo's latest-following convention (tealdeer/uv/etc.) and keep the
container in sync with the Mac's brew yq. smoke-test now asserts mikefarah AND
major v4, so a surprise yq v5 fails CI instead of silently breaking
provision.sh. Pin still available via --build-arg YQ_VERSION=vX.Y.Z.
Debian/Ubuntu `apt install yq` is kislyuk/yq (Python, v3.x), incompatible
with the mikefarah v4 syntax the cloud-init repo's provision.sh/deploy.sh
require. Replace the apt package with a pinned mikefarah Go binary, mirroring
the existing tealdeer ARG (latest-or-pin) pattern, multi-arch amd64/arm64.
smoke-test.sh now asserts `yq --version` reports mikefarah so CI catches a
regression. CHANGELOG: Unreleased entry.
- Bump mempalace pin 3.4.0 -> 3.5.0: 3.5.0 carries the upstream fix for the
top-level-anyOf diary_write schema (issue #1728 / PR #1717, merged
2026-06-14). Verified against the published 3.5.0 wheel that mcp_server.py
now advertises 'required: [agent_name]' with no root-level anyOf.
- Remove the Dockerfile.base perl workaround that stripped the anyOf from the
installed mcp_server.py — obsolete now the fix is upstream.
- pi auto-resolves to npm latest (0.79.10 -> 0.80.2) at build.
- CHANGELOG v1.2.2.
Bake pi-extensions + mempalace skills into the image (available without a
mounted skillset) and add the mempalace session-start proactive-load directive
so frequently-recreated containers actually pick the skill up. Closes the
fork/recall + mempalace under-utilisation gap.
CHANGELOG: [Unreleased] -> v1.2.1.
Baking the mempalace fallback skill fixed *availability*, but mempalace had
no proactive-load directive anywhere (pi-toolkit's global AGENTS.md only
points to pi-extensions), so a new container would still surface it only via
description-matching — the same under-utilisation the pi-extensions directive
was created to fix.
Add a session-start pointer to the pi-devbox managed AGENTS.md block
(pi-global-AGENTS.append.md): gated to pi-devbox containers and conditional on
the MemPalace MCP tools being present. Memory continuity matters most in a
frequently-recreated container — the palace is its only cross-recreate memory.
- pi-global-AGENTS.append.md: '## Session start: load the mempalace skill'.
- smoke-test: assert the pointer merges into the global AGENTS.md at build.
- docs: VENDORED.md, README, CHANGELOG [Unreleased].
Now both skills are complete in pi-devbox: directive + skill file.
pi-extensions = directive (pi-toolkit) + baked skill; mempalace = directive
(this block) + baked skill.
The pi-toolkit global AGENTS.md tells every pi session to read
~/.agents/skills/pi-extensions/SKILL.md at start (the fork/recall
under-utilisation fix), but that skill lived only in the private skillset
repo — so the pointer dangled in any container started without skillset
mounted. Bake fallbacks so the pointer always resolves.
- pi-extensions (Option 1 + Option 2, layered):
* Canonical skill promoted to the public pi-extensions package repo under
skill/ (separate commit there); co-located with the code it documents.
* rootfs/ carries a committed snapshot (the floor).
* Dockerfile.variant copies /opt/pi-extensions/skill/ over the snapshot
after the pinned clone, so a normal build ships the fresh package copy
(recorded via PI_EXTENSIONS_REF) and an old-ref/mirror build still ships
the snapshot. Helper evaluate-extension-usage.py travels with it.
- mempalace (Option 2 only): snapshot in rootfs/. Its consumer skill has no
public package home (mempalace-toolkit ships a different skill,
opencode-mempalace-bridge), so no build-time refresh.
- entrypoint links both (only-when-absent; mounted skillset still wins).
- smoke-test: build-time presence + package-match check + runtime symlink
assertions; readiness gate now waits on the last-linked skill.
- docs: skills/VENDORED.md (provenance + refresh), README, AGENTS.md,
CHANGELOG [Unreleased].
Note: shipped in the NEXT release; v1.2.0 (run 409) predates this.
The runtime 'pi-devbox-environment skill linked' smoke assertion failed in
CI run 408 (gating build-variant). Root cause: the skill-linking block ran
AFTER the pi-toolkit/extensions deploy, but the smoke readiness gate only
waits on pi-deploy markers (keybindings.json, mempalace.ts) — which land
before the skill symlink — so the assertion sampled too early.
- entrypoint-user.sh: move the image-baked-skills symlink loop to run early
(before the pi deploy block), so it completes before any readiness marker.
Still before the skillset deploy, so foreign-link semantics are unchanged.
- smoke-test.sh: add the skill symlink to the readiness gate as well.
Build-time checks (baked skill, append snippet, merged AGENTS marker) all
passed in 408; only the timing of the runtime check was wrong.
Ship skills inside the image (independent of any mounted skillset repo):
- rootfs/usr/local/share/pi-devbox/skills/<name>/ symlinked into
~/.agents/skills/ by entrypoint-user.sh (foreign-link, survives volume
recreate, never clobbers a skillset/user skill of the same name).
- New pi-devbox-environment skill: persistence model, host/LAN SSH
reachability, split-DNS mechanisms, interactive-vs-tool-shell alias
gotcha, tmux 0-index, uv-first Python, pi-studio reachability. Agnostic
to host OS / hostnames / domains / nameservers (discovered at runtime).
- Dockerfile.variant appends pi-global-AGENTS.append.md onto pi-toolkit's
pi-global-AGENTS.md (single global slot) so the skill is loaded
proactively; gated on /usr/local/lib/pi-devbox/. Idempotent.
- smoke-test: baked-skill + append-snippet + merged-marker presence and a
runtime symlink assertion.
- docs: README 'Agent skills' section, AGENTS.md layout, DOCKER_HUB.md;
moved studio-tex roadmap to v1.3.0.
pi 0.79.7 -> 0.79.10 (auto-resolved from npm latest at build).
The host-owned, bind-mounted ~/.config/devbox-shell/ssh-lan.conf is the
intended place to add `ProxyJump host` overrides for named LAN peers (so
`pi --ssh <peer>` / `dssh <peer>` route through the host), but it was only
documented in .env.example and the setup-lan-access.sh header — never in the
README, where someone hitting "can't reach LAN peers" actually looks.
- README: add a "Naming LAN peers" subsection under the macOS LAN-peers
troubleshooting block, with a ProxyJump example and the read-only ~/.ssh
caveat; add a pointer to it from the SSH and ControlMaster section.
- setup-lan-access.sh: correct the INCLUDE_BLOCK comment that suggested adding
ProxyJump to the read-only ~/.ssh/config; point at ssh-lan.conf instead.
- CHANGELOG: note under Unreleased.
Docs/comment only — no behavior change.
The default run shell is 'sh -e {0}' (dash on the act runner), which
rejects 'set -o pipefail' ('Illegal option -o pipefail') — failing the
resolve-versions job on line 2 and cascading every dependent job to
skipped (v1.1.6 run 401). The heavy build steps already declare
'shell: bash'; the resolve step did not. Added it.
Adds OCI labels + /etc/pi-devbox/build-manifest.json so a published tag is
self-describing and reconstructable after CI logs rotate (manifest is
written from the actual checked-out HEAD of each /opt clone + live
pi --version, not just the intended build-args).
Hardens the build plumbing:
- scripts/check-base-hash.sh guards the base-rebuild invariant: every
floating ARG *_REF in Dockerfile.base must be folded into the base_tag
hash, else a ref-only change silently fails to rebuild the base
(v1.1.2-class staleness footgun). Runs in base-decide and locally.
- resolve-versions now fails loud instead of falling back to a floating
main/master on a transient API failure — validates each ref is a 40-hex
SHA (and pi a real semver) and aborts the release otherwise.
- The three gitea companions (pi-toolkit, pi-extensions, mempalace-toolkit)
gained overridable *_REPO build-args (defaulting to the canonical gitea
origin) so a relocated/forked build can repoint them without editing the
Dockerfiles — matching the existing PI_FORK_REPO/PI_OBSMEM_REPO pattern.
README documents the forked/relocated build-arg trick and how to read the
labels + manifest. smoke-test asserts the manifest + labels. pi bumps
0.79.7 → 0.79.8 (auto-resolved at build).
Coordinated with the pi-extensions ssh-controlmaster fix (picked up at build via
PI_EXTENSIONS_REF=main), this makes `pi --ssh <host>` and `dssh`/`dscp` robust
to a user ~/.ssh/config whose per-host ControlPath points under the read-only
~/.ssh bind-mount (e.g. `ControlPath ~/.ssh/cm/%r@%h:%p`). A system default can
never override a user's per-host value, so the fix lives in two layers.
- setup-lan-access.sh: always render the writable ~/.ssh-local/config sidecar
(Host * ControlPath redirect into ~/.ssh-local/cm + Include ~/.ssh/config) on
EVERY host OS. Previously the script exited early (no-op) on native Linux,
leaving dssh/dscp broken when ~/.ssh was read-only there too. The host-jump
block, its key generation, and the authorize hints stay gated on VM-backed
detection / DEVBOX_LAN_ACCESS=jump (new NEED_JUMP flag).
- Dockerfile.base: document that the /etc/ssh drop-in default cannot override a
user per-host ControlPath; cross-ref the two handling layers.
- entrypoint-user.sh: correct the now-stale "no-op on native Linux" comment.
- README.md / DOCKER_HUB.md: document read-only-~/.ssh ControlPath handling.
CHANGELOG: v1.1.5 (Fixed + Changed + pi 0.79.6 -> 0.79.7 auto-resolved bump).
The settings.json bootstrap only fires when the file is ABSENT, so a
settings.json on a preserved named volume never picks up config added in a
later image (e.g. the observational-memory / pi-fork blocks, a newly-enabled
model). Users had to hand-merge after every upgrade.
On start, when settings.json already exists, deep-merge the template into it
with 'jq -s ".[0] * .[1]"' (template first, live second) so the user's values
always win and only MISSING keys are filled from the template. Arrays are
leaves (a model the user removed is not re-added). Rewrites only when the
merge changes something, backs up the original first, and skips safely (no
clobber) if either file is invalid JSON. Opt out with PI_SETTINGS_MERGE=0.
Add a recreate-sanity-check assertion that settings.json carries the
observational-memory + pi-fork blocks after recreate.
The history-flush guard was exported, so it leaked into child processes.
Any nested shell -- crucially each tmux pane (which inherits the tmux
server's env) -- then saw the guard already set and skipped installing
'history -a' in PROMPT_COMMAND. Those shells only persisted history on a
clean exit, so abrupt termination (docker stop, tmux kill-server, SIGKILL)
silently lost their in-memory history. zoxide was less affected (its hook
is installed unguarded and writes immediately).
Make the guard shell-local (drop 'export') so every new interactive shell
re-installs its own per-prompt flush. Add a recreate-sanity-check assertion
that a nested login shell still wires up 'history -a'.
Storage was never the issue: ~/.cache/bash (devbox-shell-history) and
~/.local/share/zoxide (devbox-zoxide) are both persistent named volumes.
pi-toolkit now symlinks pi-global-AGENTS.md -> ~/.pi/agent/AGENTS.md (pi's
global-instructions file, loaded at every start; directs the agent to read
the pi-extensions skill at session start). Add a recreate-sanity-check
assertion alongside the keybindings symlink check so a future image build
that bakes the new pi-toolkit verifies the wiring landed.
GITEA_ACCESS_TOKEN + GITEA_HOST (passed from host .env via compose,
primarily for gitea-mcp) are also usable for any direct Gitea API work —
run inspection, tag checks — not just ci-release-watcher. Prefer over a
PAT file when present; host-managed lifecycle, nothing to revoke. Release
checklist step 7 now notes the env-token alternative.
Patch release per AGENTS.md versioning scheme (pi version bump + smaller
fixes). Cut Unreleased → v1.1.2 (2026-06-15): pi 0.79.3→0.79.4
(CI-auto-resolved), the mempalace-toolkit SHA-resolution build fix,
recreate-sanity-check.sh maintainer tooling, and the AGENTS/README docs.
Bumped the README sanity-check version example to 0.79.4.
Step 3 now runs scripts/recreate-sanity-check.sh inside the running
container after a local recreate — the runtime peer of the build-time
smoke-test.sh gate. Verifies persisted volumes survived and pi wiring
re-deployed, not just that the container booted.
scripts/recreate-sanity-check.sh verifies what is actually live in a
recreated container — persisted volumes, pi runtime wiring (keybindings,
extensions, mempalace.ts bridge, settings.json, fork/obsmem/studio
registrations), /tmp/sshcm, skel defaults, /opt toolkits. smoke-test.sh
runs at build time with --entrypoint="" and cannot see any of this.
Variant (studio/plain) auto-detected via /opt/pi-studio. pi version is
asserted only with --expected-version (built from 'latest', no Dockerfile
pin to self-derive). Maintainer tooling, not baked into the image.
Documented in README and CHANGELOG.
Parity with opencode-devbox: PR #1735 (the diary_write root-anyOf fix) was
closed UNMERGED on 2026-06-11, so the old "remove once PR #1735 ships" TODO
pointed at a dead PR. Issue #1728 is still open; PR #1717 is the current live
candidate; mempalace PyPI latest is still 3.4.0 (== our pin), so the
workaround stays.
- Dockerfile.base: rewrite the upstream-tracking comment + TODO (#1735 dead,
watch #1717, removal trigger = a PyPI release > 3.4.0 stripping root anyOf).
- CHANGELOG: Unreleased Docs entry.
Docs-only; no behavior change.
mempalace-toolkit is the only companion cloned in Dockerfile.base (all
others live in Dockerfile.variant), so it bypassed the resolve-versions ->
build-arg plumbing and its ref stayed a literal `main`. Because the base
only rebuilds on a content hash of Dockerfile.base + rootfs/* + entrypoints,
a toolkit-only fix would silently fail to land unless Dockerfile.base itself
changed (as it incidentally did in v1.1.1).
Changes:
- resolve-versions: new mempalace_toolkit_ref output (gitea commits API,
mirrors pi-toolkit resolution; jq '.[0].sha // "main"' fallback).
- base-decide: needs resolve-versions; fold the resolved SHA into the
base-tag hash so a moved toolkit forces a base rebuild automatically.
- build-base: needs resolve-versions; pass --build-arg MEMPALACE_TOOLKIT_REF.
- Dockerfile.base: switch clone from `git clone --branch` to a SHA-capable
`git fetch <ref> + checkout FETCH_HEAD` (the --branch <SHA> footgun
already fixed in Dockerfile.variant, run 374).
base_tag now reflects a live gitea lookup; on API blip it falls back to
`main`, triggering one extra rebuild, never a missed one.
No new tag — lands on the next v* release or workflow_dispatch.
The per-request timeout + stall-kill landed in mempalace-toolkit's
mempalace.ts pi extension (commit a3b8829), which the base clones at
build via MEMPALACE_TOOLKIT_REF=main. A base rebuild picks it up.
- CHANGELOG: move from 'Known issues' to 'Fixed'; document the env knobs
(MEMPALACE_MCP_TIMEOUT_MS / MEMPALACE_MCP_INIT_TIMEOUT_MS) and why the
standalone stdio-watchdog shim was dropped.
- Dockerfile.base: replace the TODO with a note pointing at the fix.
Symptom: pi TUI blocks on a mempalace tool call, ESC does not abort.
Initial WAL-contention hypothesis ruled out (no other writer running).
Likely cause: virtiofs cold open of chroma.sqlite3 stalls the JSON-RPC
initialize handshake; pi has no per-call MCP timeout.
Recovery today: docker exec <ctr> pkill -9 -f mempalace-mcp, restart pi.
Planned fix (deferred until after opencode-devbox pi removal): stdio
watchdog shim with per-REQUEST timeout. A naive process-lifetime
timeout wrapper is wrong because mempalace-mcp is long-lived.
Sharing the palace across harnesses remains the goal.
pi-studio renders Mermaid natively but has no DOT renderer. Its markdown
preview displays local PNG/JPG/GIF/WEBP images, so dot-watch closes the
loop for Graphviz: edit .dot -> auto-render <name>.png -> Studio
refresh-from-disk shows the update. Uses mtime polling (no inotify dep).
- rootfs/usr/local/bin/dot-watch: the helper (executable)
- Dockerfile.base: COPY + chmod, following the studio-expose pattern
- README.md: 'Graphviz diagrams in Studio' subsection
- CHANGELOG.md: Unreleased entry
graphviz was already in the base image; no new package.
Audit found README/AGENTS carried a stale compose/volume set that
diverged from the shipped docker-compose.yml (DOCKER_HUB + compose +
.env.example were already consistent — README was the outlier):
- README compose block + 'Volumes and persistence' table: correct volume
names (devbox-shell-history not -bash-history; devbox-uv at
~/.local/share/uv not devbox-uv-tools at /opt/uv-tools — the latter
would SHADOW the baked mempalace install at UV_TOOL_DIR); add
devbox-ssh-local + devbox-zoxide; mark devbox-palace/-chroma-cache
optional; WORKSPACE_PATH/SSH_KEY_PATH (not HOST_WORKSPACE).
- README quickstart: 'compose exec -u developer' (no USER in image; bare
exec lands a root shell).
- README: pi-studio now 'shipped' not 'planned'; build-pipeline + tag
table cover -studio + smoke-studio/build-variant-studio.
- AGENTS: backward-compat volume names corrected; repo-layout bullets
cover pi-studio install + studio-expose + STUDIO_EXPOSE bridge.
- DOCKER_HUB: MemPalace source link -> upstream MemPalace/mempalace
(matches Dockerfile.base + CHANGELOG refs).
Note: the shipped v1.0.0 CHANGELOG migration note still lists the old
(incorrect) volume names; left as immutable released history.
pi-studio binds the container's 127.0.0.1, which a published Docker port
can't reach. Add a robust, portable bridge rather than a doc-only one-liner:
- Dockerfile.base: add socat (~1 MB, generally useful TCP relay).
- rootfs/usr/local/bin/studio-expose: socat TCP relay listening on the
container's egress IPv4 (not 0.0.0.0 — that would EADDRINUSE against
Studio's loopback listener) forwarding to 127.0.0.1:PORT on the SAME
port, so Studio's printed token URL works verbatim. Robust egress-IP
detection (hostname -I, loopback-filtered; ip route get fallback),
--help, port validation, foreground.
- entrypoint-user.sh: opt-in STUDIO_EXPOSE=1 auto-starts the bridge in the
background (studio variant only). Default OFF — Studio stays loopback-only
(its secure default) unless explicitly opted in.
- README: 'Using pi-studio' now documents host-networking (A) and the
studio-expose/STUDIO_EXPOSE bridge (B) with a security note; ssh -L for
remote, mosh caveat retained.
- smoke-test: assert socat + studio-expose present (base-level).
- CHANGELOG/AGENTS updated.
No tag — stopping for review.
Bundle pi-studio (omaclaren/pi-studio) as a new -studio image variant:
browser prompt editor, KaTeX/Mermaid preview, tmux-backed literate REPLs,
/studio command + studio_* agent tools.
- Dockerfile.variant: INSTALL_STUDIO + PI_STUDIO_REPO/REF args; vendor
pi-studio to /opt/pi-studio (no build step — prebuilt client in git;
npm install --omit=dev for 3 prod deps). STUDIO_PORT=8765 advisory.
- entrypoint-user.sh: register /opt/pi-studio via the existing pi install
local-path loop (auto-skips in non-studio variant).
- smoke-test.sh: auto-detected studio assertions (clone + prebuilt client
+ pi install registration).
- CI: resolve PI_STUDIO_REF to a SHA; independent smoke-studio +
build-variant-studio jobs that gate ONLY the -studio tags, so a studio
failure never blocks the core :latest release.
- README: 'Using pi-studio' section documenting the container access
reality — pi-studio hard-binds 127.0.0.1 (index.ts .listen(port,
'127.0.0.1'), no --host flag), so -p publish alone can't reach it.
Documents host-networking and loopback-bridge paths, the remote ssh -L
forward, and the mosh caveat (no port forwarding; run parallel ssh -L).
- CHANGELOG/AGENTS/DOCKER_HUB updated. Will tag as v1.1.0 (minor).
No tag created — stopping for review.
Anthropic's tools API rejects top-level anyOf/oneOf/allOf in
input_schema. Mempalace 3.3.x/3.4.0 advertise diary_write with
`anyOf: [{required:[entry]}, {required:[content]}]`, breaking pi
on first prompt with:
tools.<n>.custom.input_schema: input_schema does not support
oneOf, allOf, or anyOf at the top level
Patch the installed mcp_server.py after `uv tool install` to drop
the anyOf block and require ["agent_name", "entry"] instead. The
handler still accepts `content` server-side as a kwarg alias, so
callers using either name keep working.
The workaround is idempotent and self-deactivating: once upstream
ships the fix the regex no longer matches and the RUN is a silent
no-op.
Also pin mempalace to 3.4.0 via a new MEMPALACE_VERSION build arg
so future bumps are a deliberate, reviewable diff rather than an
implicit pull of latest (an implicit upgrade is what swept the
broken schema in unannounced).
Refs MemPalace/mempalace#1728, MemPalace/mempalace#1735
Run 376 published the Hub description with PI_VERSION=empty ("Current
:latest ships pi `` ") because the update-description job has
needs: [build-variant] but not resolve-versions. In Gitea Actions
needs.<job>.outputs.* only resolves for jobs in your needs: list.
The post-substitution sanity-check (grep -q '{{PI_VERSION}}') passed
because sed had successfully replaced the placeholder — with empty
string. Pre-empted by adding a non-empty assertion in this commit:
the step now fails loudly if PI_VERSION resolves to empty rather than
silently publishing a broken description.
The Hub description still described the pre-v1.0.0 reality (tags follow
pi npm version, builds FROM joakimp/pi-devbox:base-pi-only, opencode-
devbox lineage as source of truth) — none of that has been true since
v1.0.0 decoupled. End users on Hub got a misleading story.
DOCKER_HUB.md changes:
- Versioning section rewritten: semver from v1.0.0, with a deprecation
note for the pre-v1.0.0 v{pi_version}[letter] scheme.
- New 'Build pipeline' section briefly explains the two-phase
base/variant content-addressed structure so users understand what
base-<hash> and base-latest tags are for.
- New 'Document and image tooling' section (pandoc, graphviz,
imagemagick) added since these are new in v1.0.0 and broadly useful.
- Tealdeer noted (vs the old Node tldr).
- Tmux 0-indexing called out (relevant for future :latest-studio
variant).
- Removed all 'pi-only build' / 'FROM base-pi-only' / 'opencode-devbox
bakes the pi version' framing — pi-devbox is now self-contained.
- New {{PI_VERSION}} placeholder in 4 locations so the Hub description
always shows which pi is in :latest.
Workflow change:
- update-description step now substitutes {{PI_VERSION}} placeholders
in DOCKER_HUB.md before sending to Hub. PI_VERSION comes from the
resolve-versions output (same one baked into the image), so the page
and image can never disagree. Sanity-check fails the step if any
unsubstituted placeholder remains.
The previous clone helper for these two repos (git_clone_retry) used
`git clone --branch <ref>`, which only accepts branch names or tags,
NOT commit SHAs. Run 374 (the workflow_dispatch retry of v1.0.0)
failed at smoke because the workflow's resolve-versions step had been
extended to resolve PI_TOOLKIT_REF and PI_EXTENSIONS_REF to commit
SHAs (commit b55b44e), and `git clone --branch <40-char-SHA>` fails
with 'Remote branch not found'.
Switching all four clones to git_fetch_ref (`git fetch + checkout
FETCH_HEAD`) makes the build accept both branch names AND SHAs
uniformly. Both Gitea and GitHub allow fetching arbitrary commits by
default (uploadpack.allowReachableSHA1InWant).
The unused git_clone_retry helper is removed; comment explaining the
choice and the historical context is in its place.
Image was published successfully on run 373; this only affects the
v1.0.0-rerun path (description fix). Image bytes unchanged because the
SHAs being passed match what run 373 cloned by branch name.
The v1.0.0 release run failed at update-description because Docker Hub's
short-description field has a 100-byte limit and the previous string
was 151 bytes (the em dash is 3 bytes UTF-8). The image itself shipped
fine — only the cosmetic Hub description patch failed.
Changes:
- Short description: 'Linux container with the pi coding-agent, MemPalace,
and curated dev tooling.' (77 bytes, was 151)
- resolve-versions now also resolves pi-toolkit and pi-extensions main
HEADs to commit SHAs so workflow_dispatch re-runs produce byte-identical
images when those repos haven't moved. Fork+obsmem were already
SHA-resolved; toolkit+extensions were branch-named (drift risk on
re-runs that we got lucky on for v1.0.0).
Self-contained build chain — own Dockerfile.base + Dockerfile.variant
+ entrypoint scripts + rootfs + CI pipeline. Previously v0.79.0 and
earlier were thin re-brands of opencode-devbox's pi-only variant
(joakimp/pi-devbox:base-pi-only built by opencode-devbox CI).
Architectural changes:
- Replace 5-line Dockerfile shim with full base+variant pair.
- Adapt CI workflow from opencode-devbox/docker-publish-split.yml,
simplified to a single variant. Includes content-addressed base hash,
PI_VERSION concrete-resolution to defeat registry-buildcache footgun,
crane-based base-latest promotion, and the c6f9d11 smoke-test gate.
- pi-devbox releases no longer require rebuilding opencode-devbox first.
Base image additions:
- pandoc, graphviz, imagemagick, yq — broadly useful, ~260 MB total.
- tldr (tealdeer) — Rust port replaces Node tldr global, saves 135 MB.
- /etc/tmux.conf with base-index 0 + pane-base-index 0 — required for
the planned :latest-studio variant; pi-studio hard-codes :0.0 target.
Smoke test:
- New checks for pandoc, graphviz, imagemagick, yq, tldr, tmux config,
/tmp/sshcm directory.
- Image-size measurement now sums docker history layers (the prior
inspect --format='{{.Size}}' returned only the variant-unique layer
with the new base/variant split, understating by 2+ GB).
- Threshold 2850 → 3500 MB to absorb base additions + arch margin.
Image size:
- Local arm64 build: 3.20 GB. ~390 MB up from prior pi-only equivalent.
- Will tighten threshold once amd64 actuals settle in CI.
Pre-1.0 history preserved at tag pre-v1.0.0-decouple-backup.
Future work:
- v1.1.0: :latest-studio variant (adds pi-studio).
- v1.2.0: :latest-studio-tex variant (adds texlive-xetex for PDF).
- opencode-devbox v2.0.0 will retire INSTALL_PI / pi-only paths.
First build on pi 0.79.0. Built FROM the republished base-pi-only from
opencode-devbox v1.16.2 (carries pi 0.79.0). Bump smoke size threshold
2750 -> 2850 MB in lockstep with opencode-devbox's pi-only variant.
Promote CHANGELOG Unreleased -> v0.79.0.
Add the two missing doc entries (the ~/.config/devbox-shell compose mount,
3bfbafa; and DEVBOX_HOST_ALIAS in .env.example, 45f4488) and promote Unreleased
-> v0.78.1 (2026-06-04). v0.78.1 is a real pi version bump (0.78.0 -> 0.78.1);
builds FROM the republished base-pi-only carrying pi 0.78.1 from opencode-devbox
v1.15.13e. Docs only.
Persist ~/.ssh-local so the generated LAN-jump key survives container
recreation; authorize it on the host once per machine. Adds the volume
to the compose template and documents it in the README volumes table.
LAN-access mechanism/script changes are inherited from base-pi-only
(opencode-devbox).
opencode-devbox v1.15.13d published the rebuilt base-pi-only (digest 83b45335)
with the fixed setup-lan-access.sh (Include scope + ControlPath) and the new
ssh-lan.conf / RFC1918 autojump. Tagging now to build the thin pi-devbox on it.
setup-lan-access.sh fixes (Include scope, ControlPath) + ssh-lan.conf and
RFC1918 autojump flow in via FROM base-pi-only. Documents the knob and new
host-owned config. Tag v0.78.0c AFTER opencode-devbox v1.15.13d publishes the
rebuilt base-pi-only, so it doesn't build on the stale base.
The pi-only building block now lives in this repo as the internal
base-pi-only tag (produced by opencode-devbox CI from Dockerfile.variant,
INSTALL_OPENCODE=false) instead of opencode-devbox:latest-pi-only — so an
'opencode-devbox' tag never ships without opencode.
- Dockerfile: BASE_IMAGE default joakimp/opencode-devbox:latest-pi-only
-> joakimp/pi-devbox:base-pi-only.
- Updated README, AGENTS, DOCKER_HUB, docker-compose, CHANGELOG.
- Single source of truth unchanged (opencode-devbox/Dockerfile.variant);
publish ordering + EXPECTED_PI_VERSION smoke guard unchanged.
Re-point the re-brand at the new pi-only variant instead of with-pi, so
pi-devbox stays a lean pi-focused image (no opencode) while the pi install
logic still lives in one place upstream. This keeps pi-devbox meaningfully
distinct from opencode-devbox:latest-with-pi.
- Dockerfile: BASE_IMAGE default -> joakimp/opencode-devbox:latest-pi-only.
- smoke-test.sh: size threshold 2900 -> 2750 MB (pi-only = with-pi minus
opencode's ~145 MB binary).
- Docs (README/AGENTS/DOCKER_HUB/CHANGELOG/docker-compose): drop the
'also contains opencode' notes; describe pi-only basis and the distinction
from with-pi.
Publish ordering unchanged: release opencode-devbox first so latest-pi-only
carries the target pi version, then tag here (smoke asserts pi --version).
pi-devbox no longer installs pi itself. The Dockerfile is now a thin
FROM joakimp/opencode-devbox:latest-with-pi (overridable via BASE_IMAGE),
inheriting pi + pi-toolkit + pi-extensions + pi-fork (fork) +
pi-observational-memory (recall) + the LAN-access helper + all base tooling
from the single source of truth. Eliminates the install-logic duplication
that drifted against opencode-devbox/Dockerfile.variant (decision #3).
Consequences (documented in CHANGELOG/AGENTS):
- The image now ALSO contains opencode (with-pi has INSTALL_OPENCODE=true).
A leaner pi-only image would need a dedicated pi-only variant upstream.
- Publish ordering: release opencode-devbox first so latest-with-pi carries
the target pi version, THEN tag this repo. The smoke test asserts
pi --version matches the tag (EXPECTED_PI_VERSION) and fails loudly if the
base is stale — turning the version coupling into an enforced ordering guard.
CI: drop PI_VERSION build-arg (Dockerfile installs nothing); keep tag->version
resolution to feed the smoke base-freshness guard. Smoke adds fork/recall
clone + node_modules + settings.json registration checks; size threshold
2200 -> 2900 MB (now tracks with-pi). Docs updated across README, AGENTS,
DOCKER_HUB, .env.example, docker-compose.
First container build on pi 0.77 line (published upstream 2026-05-28).
Built against unchanged joakimp/opencode-devbox:base-latest (same as
v0.76.0 — SSH-CM, gitleaks, git-crypt all carry forward).
Notable pi 0.77.0 upstream:
- Claude Opus 4.8 support
- --exclude-tools / -xt for selective tool disablement
- Headless Codex subscription login (device-code auth)
- Streaming-aware extension input (InputEvent.streamingBehavior)
- Long bugfix list (startup timing, signal handling, terminal
protocol detection, Windows MSYS2 fixes, provider metadata
cleanups, session disposal abort, etc).
Also folds the previously-Unreleased CI retry-wrapper change
(2d39766) into this release block. Second publish exercising the
cache-export-disabled workflow; first to exercise the 3-attempt
retry wrapper through the publish path.
See CHANGELOG v0.77.0 for full notes.
Belt-and-braces against transient registry-1.docker.io blips (rate
limits, brief 5xx, CDN flap). Replaces docker/build-push-action@v7 with
a shell: bash step that runs docker buildx build --push in a for-loop
with backoff (15s, 30s).
Does NOT mask deterministic failures: a true regression (e.g. the
cache-export 400 we hit 2026-05-23..28) fails all 3 attempts
identically and the job still fails by design. Orthogonal layer to
both cache-export disablement and the ci-release-watcher skill's
transient-rerun heuristic.
No image-side change.
pi 0.75.5 → 0.76.0 (published upstream 2026-05-27 20:03 UTC). First
pi-devbox release built against opencode-devbox base-latest carrying the
SSH ControlMaster bake-in (commit 668592d) and gitleaks (73a7f96) — both
inherited transparently with no Dockerfile change here. PI_VERSION is
resolved from the git tag by the workflow (v0.75.5b cache-hit fix), so
no Dockerfile default bump needed.
Workflow change: registry cache-export removed from publish step. buildkit
mode=max cache-export to registry-1.docker.io reproducibly returns HTTP 400
(Hub-CDN protocol mismatch with buildx 0.34.x, surfaced ~2026-05-23).
Diagnosed during opencode-devbox v1.15.12 manual publish: image push works,
only --cache-to fails. Pi-devbox would hit the same regression on the next
tag push without this fix. See opencode-devbox CHANGELOG v1.15.12 for the
full root-cause analysis. Pi-devbox is single-stage with a tiny diff (npm
install pi only) on top of base-latest, so builds are fast even uncached.
Symmetric with the gitleaks/git-crypt inherit-note already present.
Cross-references opencode-devbox commit 668592d (Unreleased), which
bakes /etc/ssh/ssh_config.d/00-devbox-controlmaster.conf with a
writable /tmp/sshcm ControlPath. pi-devbox picks this up automatically
on its next build against base-latest; no Dockerfile change here.
Documents the symptom users see today inside pi-devbox <= v0.75.5b
(unix_listener Read-only file system on \~/.ssh/cm) and the fact
that pi --ssh user@host inside the container is currently silently
broken until the cascade lands.
No Dockerfile install change here — pi-devbox FROMs joakimp/opencode-
devbox:base-latest which gained gitleaks (and explicit acknowledgment
of git-crypt) in opencode-devbox commit adding both to the base layer.
The next pi-devbox release built against a fresh base-latest digest
inherits both with zero work on this side.
CHANGES
Dockerfile — comment block at top updated to name git-crypt + gitleaks
in the 'inherited from base' toolset enumeration. Helps future
readers: one less reason to think 'I need to install gitleaks here'.
CHANGELOG.md — new Unreleased entry pointing at the opencode-devbox
base-side change for full detail. Will be promoted whenever the next
pi-devbox release ships (probably alongside the next pi npm bump past
0.75.5).
Holding off on tagging — pi upstream still at 0.75.5, baseline release
v0.75.5b is already current with that. Will ride along with next pi
bump.
ALL FOUR releases v0.74.0 -> v0.75.5 had been shipping the same image
bytes due to a Docker layer-cache hit on the bare 'npm install -g
@earendil-works/pi-coding-agent' command (when PI_VERSION=latest).
The command string is identical across builds, so the layer-hash is
identical, so registry buildcache (cache-from/cache-to) silently
reuses the layer from whatever pi version was current when the cache
was first populated.
Verification: docker manifest inspect joakimp/pi-devbox:vX.Y.Z showed
identical SHA256 digests on both linux/amd64 and linux/arm64 for
v0.74.0, v0.75.3, v0.75.4, v0.75.5. Users on :latest were getting
whatever pi version was baked into the v0.74.0 build.
DISCOVERED 2026-05-23 by user trying to update pi-devbox on MBP-M1
and seeing pi 0.74.0 reported despite pulling v0.75.5.
CHANGES
.gitea/workflows/docker-publish.yml — both smoke and publish jobs
get a new 'Resolve PI_VERSION from tag' step that strips the leading
'v' and any trailing letter suffix from github.ref_name. Result is
passed as a build-arg to docker/build-push-action so the npm install
layer's hash includes the concrete version, forcing cache miss when
pi bumps.
scripts/smoke-test.sh — new run_expect helper that asserts pi
--version contains the EXPECTED_PI_VERSION env var. Smoke job sets
this from the resolve step output. Would have caught this regression
on v0.75.3.
Dockerfile — comment block above ARG PI_VERSION=latest documenting
the cache-hit footgun. The 'if latest' branch in the install RUN is
preserved for local dev convenience but never fires in CI now.
AGENTS.md — new convention bullet explaining the cache-hit class of
bug and noting the latent same-bug in opencode-devbox's with-pi
variants (currently masked by OPENCODE_VERSION bumps; will manifest
when cutting a vN.N.Nb-style opencode-version-unchanged release that
only bumps pi).
CHANGELOG.md — full entry under v0.75.5b describing the recovery,
the silent-failure mechanism, and the verification steps.
NO IMAGE-CONTENT CHANGES vs v0.75.5 INTENT. This build produces the
actual pi 0.75.5 image content that v0.75.5 was supposed to ship.
NEXT FOLLOWUP (parked, not in this commit)
opencode-devbox should get the same workflow change for its
build-variant-with-pi and build-variant-omos-with-pi jobs. Currently
masked because every release also bumps OPENCODE_VERSION which
invalidates the cache, but that masking would fail on a pi-only bump
release.
Companion to opencode-devbox's 'Upstream sources' section. Pi's npm
package ships a rich CHANGELOG.md with New Features / Added / Changed
/ Fixed sections — but the npm registry metadata ('npm view') doesn't
include the changelog body. Surface the 'npm pack + tar' recipe in
the release-day checklist so future-pi (and human-pi) doesn't try to
derive notes from npm view alone.
Doc-only, no CI implications.
One upstream patch release, two days after v0.75.4. PI_VERSION=latest
in Dockerfile resolves to 0.75.5 at build time, so no Dockerfile change
is needed; just a CHANGELOG promote.
Notable upstream changes (read tool card cleanup, faster Windows file
tools, more reliable pi update, custom adaptive-thinking knob, several
bash/Bedrock fixes) — see CHANGELOG.md for the full list.
Cache hit expected on opencode-devbox:base-latest (base-35ee5fe7861a).
Tagged together with opencode-devbox v1.15.10 — both releases go
through the queued CI runner overnight.
One upstream patch release. PI_VERSION=latest in Dockerfile resolves
to 0.75.4 at build time, so no Dockerfile change is needed; just a
CHANGELOG promote.
Tagged speculatively before opencode-devbox v1.15.6's omos-with-pi
smoke completes — pi 0.75.4 is a single patch on top of 0.75.3, low
risk on its own. If opencode-devbox v1.15.6 surfaces a pi 0.75.4
problem in the omos-with-pi smoke (3700 MB threshold trip, etc.),
both releases would fail in symmetric ways and recovery would be a
v0.75.4b/v1.15.6b pair. Same recovery muscle as v1.15.4 -> v1.15.4b
last week.
Built on opencode-devbox:base-latest, cache-hit on base-35ee5fe7861a
since v1.14.50b — base unchanged across both bumps.
Companion to the same addition in the cloud-init and ansible repos.
Caught real drift in those repos in a recent session only because
the user explicitly asked. Codify the sweep with concrete, repo-
specific drift hotspots rather than a vague 'watch for drift' rule
that gets ignored.
Each AGENTS.md addition lists the doc files most likely to fall
behind code changes here, plus a quick-triage one-liner using
'git diff --name-only HEAD | xargs grep -l ...' so the rule is
actionable not aspirational.
pi @earendil-works/pi-coding-agent@0.74.0 -> 0.75.3 (one upstream minor
+ three patch releases since the initial pi-devbox release on 2026-05-14).
Validated: opencode-devbox v1.15.4b's smoke-with-pi and smoke-omos-with-pi
both passed with pi 0.75.3 baked in. Node v22.22.2 is comfortably above
pi's new minimum requirement of 22.19.0.
Built on joakimp/opencode-devbox:base-latest (cache hit on
base-35ee5fe7861a from 2026-05-14). PI_VERSION=latest in Dockerfile
resolves to 0.75.3 at build time. Image-side unchanged from v0.74.0
beyond the pi npm version.
README rewrite:
- Two quick-start paths: 'no git clone' (curl docker-compose.yml +
.env.example) and 'with git clone' for hackers/forkers
- New 'Authentication' section with subsections per provider
(Anthropic, OpenAI, Gemini, AWS Bedrock static, AWS Bedrock SSO).
AWS SSO path documents the ~/.aws bind-mount.
- Persistent state expanded: 5-row volume table + optional volumes
table. Annotated what survives what.
- Configuration reference: full .env table.
- Versioning, building from source (with build args table),
troubleshooting FAQ, related projects, license.
- 11 kB total — comprehensive but readable.
DOCKER_HUB.md tweaks:
- Quick-start now has a 'no git clone' path (curl two files), pointing
users at the gitea README for the full setup guide. The git-clone
path was overkill for the 90% case (just want to docker run).
- Explicit link to gitea README at the end of the quick-start block.
Replace the 1-line placeholder with a proper Hub README:
image variants table, quick start (docker run + docker compose),
inherited-from-base + added-by-pi-devbox feature lists, versioning
scheme, persistent volumes table, user-installed pi packages note,
source links.
Already PATCH'd live on Docker Hub manually — this commit keeps the
in-repo file in sync so the next tag-triggered update-description job
won't roll it back to the stub.
2026-05-15 08:47:23 +02:00
49 changed files with 19331 additions and 165 deletions
Check release notes at https://github.com/earendil-works/pi/releases for
the upstream changelog to include in `CHANGELOG.md`.
2.**Refresh the vendored mempalace skill snapshot if the skillset moved:**
`scripts/vendor-mempalace-skill.sh --check` (reads a real skillset clone,
writes nothing). Three exit codes, not two — a stale-but-truthful record is
**not** a release blocker, so don't treat any non-zero exit as "must
refresh" without reading which one it was:
- **0** — the record is truthful. This includes stale-but-truthful
(upstream has moved past the recorded ref, or the local clone has
uncommitted changes) — a `NOTICE` is printed, but nothing is lying.
**Skipping the refresh in this case is the legitimate, sanctioned
outcome** — every enrolled host reads its own live skillset clone, so
the baked copy is only a no-mount fallback. What is not legitimate is
skipping it *silently*: the drift is visible here, in
`pi-devbox-version`, and in the manifest, so decide rather than forget.
- **1** — a confirmed problem: the vendored bytes provably do NOT match
the file at the recorded ref (a lying record), or the recorded ref
doesn't even resolve to that path in this clone. Refresh.
- **2** — cannot determine (the recorded ref itself isn't resolvable in
this clone — commonly a shallow checkout missing history). Fetch full
history and re-check before deciding; don't refresh blind.
Refresh with `scripts/vendor-mempalace-skill.sh`, which rewrites the file
**and** the ARG together so they cannot drift apart, and refuses (exit 1)
rather than silently rewinding provenance if the skillset clone's HEAD is
behind the already-recorded ref (detached HEAD, older checkout) — pass
`--force` only if that rewind is genuinely intended.
Two consequences to accept deliberately on an actual refresh: the snapshot
is hashed into `base_tag`, so it costs a base rebuild (~67 min); and if the
section the phrase canary names has changed, re-pin it in
`scripts/smoke-test.sh`.
3.**Update the docs this release makes stale — BEFORE you tag.** Rename
`CHANGELOG.md`'s `## Unreleased` to `## vX.Y.Z — YYYY-MM-DD` (em dash, as
every prior release heading uses), then run the gate:
## Key facts
```bash
bash scripts/check-doc-drift.sh # 0 in sync / 1 drift / 2 cannot run
```
- **Base image**: `joakimp/opencode-devbox:base-latest` — rebuilt whenever opencode-devbox cuts a new base
- **pi binary**: baked at `/usr/bin/pi` (system npm prefix); `NPM_CONFIG_PREFIX=/home/developer/.pi/npm-global` at runtime so user-installed pi/packages land on the named volume
- **Companion repos**: pi-toolkit and pi-extensions cloned to `/opt/` at build time; `entrypoint-user.sh` (inherited from base) deploys symlinks to `~/.pi/agent/` on container start
- **MemPalace**: fully operational — inherited from base image; bridge extension deployed by entrypoint
It compares README.md's version-pin table against the ARGs it names, and
DOCKER_HUB.md's Node claim against `ARG NODE_VERSION`, plus Hub's
25 000-char limit, unsubstituted `{{PLACEHOLDERS}}`, and stale `Unreleased`
pointers in user-facing docs. With network it also checks DOCKER_HUB.md's
size claims against Hub's measured sizes (check 8) and — check 9 — that
**every component the next build would bake differently from the last
published release is named in the CHANGELOG** above that release's heading:
it reads the `se.jordbo.pi-devbox.*-ref` labels off the published image and
`git ls-remote`s each floating `*_REF`. A red check 9 means an upstream
pi coding-agent container — built on opencode-devbox base. Includes pi, pi-toolkit, pi-extensions, mempalace, AWS CLI, neovim, and full dev toolchain. See https://gitea.jordbo.se/joakimp/pi-devbox for full docs.
# pi-devbox
A self-contained Docker container for the [pi coding-agent](https://github.com/earendil-works/pi) — pi + companion repos + MemPalace + a curated set of dev tooling, ready to run.
> **Current `:latest` ships pi `{{PI_VERSION}}`** (resolved at build time; see [Versioning](#versioning)).
## Image variants
| Tag | Architectures | Size (compressed) | What you get |
|---|---|---|---|
| `joakimp/pi-devbox:latest` | amd64, arm64 | ~1.23 GB | Self-contained: base + pi `{{PI_VERSION}}` + companions |
> **pi-studio (`-studio` tags):** launch with `/studio --no-browser --port 8765` inside a pi session. The server binds `127.0.0.1` **inside the container**, so reach it via host networking or a loopback bridge (and `ssh -L` for a remote host; mosh needs a parallel `ssh -L`). Full recipe: [README → Using pi-studio](https://gitea.jordbo.se/joakimp/pi-devbox#using-pi-studio--studio-variant).
## Quick start
One-shot, no persistence:
```bash
docker run -it --rm \
-v "$PWD":/workspace \
-v "$HOME/.ssh":/home/developer/.ssh:ro \
-e ANTHROPIC_API_KEY="$ANTHROPIC_API_KEY"\
joakimp/pi-devbox:latest pi
```
For a fully-configured environment with persistent settings, MemPalace memory, neovim plugins, and shell history surviving container recreation, use docker-compose. **You don't need to clone the repo** — just grab two template files:
# Edit .env — set WORKSPACE_PATH, an LLM API key (ANTHROPIC_API_KEY,
# OPENAI_API_KEY, GEMINI_API_KEY, or AWS_*), and your git identity.
docker compose run --rm devbox pi
```
Full setup guide — authentication for each provider (Anthropic, OpenAI, Gemini, AWS Bedrock SSO + static), persistence model, configuration reference, build args, troubleshooting: **<https://gitea.jordbo.se/joakimp/pi-devbox#readme>**
## What's inside
### pi and companions
- **pi `{{PI_VERSION}}`** ([`@earendil-works/pi-coding-agent`](https://www.npmjs.com/package/@earendil-works/pi-coding-agent)) — installed at `/usr/bin/pi`, pinned to an audited version (not npm `latest`)
- **pi-atelier** — TUI sidebar (ordered panels, split-pane, themes), vendored at `/opt/pi-atelier` and pinned to an audited tag; the exact tag is in the image labels (`se.jordbo.pi-devbox.pi-atelier-version`) and `/etc/pi-devbox/build-manifest.json`
- **`fork`** ([pi-fork](https://github.com/elpapi42/pi-fork)) and **`recall`** ([pi-observational-memory](https://github.com/elpapi42/pi-observational-memory)) tools
- **mempalace bridge** — MCP extension auto-symlinked so pi reads/writes the host-mounted palace
- **image-baked agent skills** — skills under `/usr/local/share/pi-devbox/skills/` (e.g. `pi-devbox-environment`, which teaches agents the container's persistence/networking/DNS/tmux/REPL specifics) are symlinked into `~/.agents/skills/` on start, available with or without a mounted skillset repo
The entrypoint deploys/registers all of these on first container start. Re-running is idempotent and preserves user edits.
**Terminal UI mode — fullscreen by default.** pi 1.0.0 made the TUI fullscreen, and this image adopts upstream's default. Fullscreen uses the terminal's alternate screen, so the transcript no longer lands in your terminal's (or tmux's) native scrollback. To get the previous behaviour back, set `"tuiMode": "regular"` in `~/.pi/agent/settings.json`, or pass `pi --tui-mode regular` for a single session. The bundled pi-atelier sidebar works in both modes. See the README for the related `fullscreenExitOutput` / `fullscreenScrollbar` / `fullscreenCopyOnSelect` / `fullscreenWheelScrollLines` settings — the last one behaves differently over SSH, which is how this container is usually driven.
### MemPalace (persistent agent memory)
- **MemPalace** + MCP server — semantic search over conversation history, knowledge graph, diary; queryable via 29 `mempalace_*` tools inside pi
- ChromaDB ONNX embedding model pre-warmed at build time (`all-MiniLM-L6-v2`)
- Bind-mount your host's `~/.mempalace` and the host-pi and container-pi share one brain
### Document and image tooling
- **pandoc** — universal Markdown↔HTML/Org/RST/etc. conversion. Useful well beyond pi: agent-driven doc exports, format conversion, etc.
- **Typst** — markup-based typesetting, used as pandoc's `--pdf-engine`
- **agent-browser** — CLI for driving a real browser (open pages, click/fill/`eval`, snapshot the DOM, screenshots) so agents can verify front-end work instead of guessing
- **Playwright** + a headless **Chromium** are pre-installed and pinned together; `AGENT_BROWSER_EXECUTABLE_PATH` is preset to the baked browser, so `agent-browser open <url>` works out of the box with no setup
- **socat** — TCP bridge used to expose the pi-studio server outside the container's loopback
### Modern CLI tooling
- **Editor**: neovim (system-wide `termguicolors` default; bring your own config/plugins), tmux (configured for 0-indexed sessions)
- **Search/nav**: ripgrep, fd, fzf, zoxide
- **Display**: bat, eza, htop, tree
- **Data**: jq, yq, sqlite3, bc/dc, column
- **Help**: tldr (tealdeer — Rust port; run `tldr --update` once to populate cache)
- **Python**: system Python 3 + **uv** (preferred) for fast Python package management. Run any Python REPL/notebook stack on demand without bloating the image:
```bash
uv run --with ipython ipython
uv run --with jupyterlab jupyter lab --no-browser --port 8888
uv run --with marimo marimo edit
```
- **Node.js** v24 LTS + npm (used by pi itself)
- **Rust** — `rustup-init` is on PATH; install toolchains on demand
- **Go** — opt-in via `--build-arg INSTALL_GO=true` if rebuilding from source
### Cloud + secrets
- **AWS CLI v2** — for SSO + Bedrock auth (pi's preferred LLM provider for the maintainer's setup)
- **Gitea MCP** server — for Gitea API access from inside pi
- **age**, **git-crypt** — encryption tooling
### SSH and networking
- OpenSSH client with **ControlMaster auto** preconfigured on a writable socket path (`/tmp/sshcm/`). Mitigates ssh banner-exchange failures behind CGNAT-restricted residential ISPs (~4-flow caps). A read-only `~/.ssh` carrying a per-host `ControlPath` (common CGNAT configs) is handled too — redirected to a writable socket dir for both `pi --ssh` and `dssh`/`dscp`.
- A **LAN-access helper** that auto-configures ssh jump-via-host on VM-backed hosts (OrbStack / Docker Desktop on macOS) so the container can reach the host's directly-attached LAN peers (`dssh <peer>` alias; `DEVBOX_LAN_ACCESS` / `HOST_SSH_USER`).
## Versioning
From v1.0.0 onward, pi-devbox uses **semver**:
- **Major** — architectural changes. v1.0.0 is the first decoupled release, where pi-devbox got its own self-contained build chain (previously it was a thin re-brand of opencode-devbox's `pi-only` variant).
- **Minor** — new image variants, significant base additions.
- **Patch** — pi version bumps, smaller fixes.
The pi binary version inside any given release is shown in this description (currently **`{{PI_VERSION}}`** for `:latest`) and asserted by smoke tests to match what's documented — version drift is caught at CI time, not on user pull.
> **Pre-v1.0.0 history.** Tags v0.74.0…v0.79.0 followed the pi npm version directly (`v{pi_version}[letter]`). Those images remain on Hub but are deprecated in favor of `:latest` / `:v1.X.Y`. The legacy `:base-pi-only*` tags were CI artifacts of the old opencode-devbox-based build pipeline; they will be removed in a future opencode-devbox v2.0.0.
### Build pipeline
pi-devbox is built in two phases:
1. **Base** (`Dockerfile.base`) → `base-<hash>` tag, content-addressed over `Dockerfile.base` + `rootfs/` + `entrypoint*.sh`. Rebuilt only when those change.
2. **Variant** (`Dockerfile.variant`) → `:latest` and `:vX.Y.Z`. FROMs the base, adds the pi install + companions.
`base-latest` is an alias of the most recent base.
## Persistent state
User edits and pi-installed packages survive container recreation when you mount these named volumes. Use the included `docker-compose.yml` and they're set up automatically.
| Volume | Mount point | What it holds |
|---|---|---|
| `devbox-pi-config` | `/home/developer/.pi/` | pi settings, extension toggles, sessions, user-installed pi packages (`npm install -g`, `pi install npm:…`) |
| `devbox-shell-history` | `/home/developer/.cache/bash` | bash history |
| `devbox-chroma-cache` | `/home/developer/.cache/chroma` | ChromaDB embedding model cache (~80 MB, can be rebuilt) |
## User-installed pi packages
`NPM_CONFIG_PREFIX` is set inside the container to `/home/developer/.pi/npm-global`. Anything you `pi install npm:<pkg>` or `npm install -g` lands on the `devbox-pi-config` named volume — survives container recreation **and** image rebuilds. A user-installed `pi` wins over the baked one via `PATH` order, so you can pin a different pi version without rebuilding the image.
RUN sed -i -E '/(en_US|en_GB|sv_SE|da_DK|nb_NO|fi_FI|de_DE|fr_FR|es_ES|it_IT|pt_BR|nl_NL|pl_PL|ja_JP|ko_KR|zh_CN)\.UTF-8/s/^# //g' /etc/locale.gen && locale-gen
ENVLANG=en_US.UTF-8
ENVLANGUAGE=en_US:en
ENVLC_ALL=en_US.UTF-8
ENVEDITOR=nvim
# Advertise 24-bit colour so colour-aware tools (Neovim's own auto-detect, bat,
# delta, ...) use true colour instead of a 256-colour fallback. Safe for the
# modern terminals this devbox targets; override by exporting `COLORTERM=`
# (empty) from a terminal that lacks true-colour support.
pi-devbox is distributed under the MIT License (see [`LICENSE`](LICENSE)), which
covers **this repository's own contents** — the Dockerfiles, entrypoint scripts,
`rootfs/` seeds, CI workflows, and docs.
The **published container images** (`joakimp/pi-devbox:*`) additionally *bundle*
third-party software, each of which remains under its own license. This file is
a good-faith summary; the authoritative sources are the upstream projects and,
for OS packages, the per-package copyright files inside the image at
`/usr/share/doc/<package>/copyright`.
## pi and its extensions (installed in the variant layer)
| Component | Upstream | License |
| --- | --- | --- |
| pi (`@earendil-works/pi-coding-agent`) | npm | MIT |
| pi-fork | github.com/elpapi42/pi-fork | MIT |
| pi-observational-memory | github.com/elpapi42/pi-observational-memory | MIT |
| pi-studio *(`-studio` variant only)* | github.com/omaclaren/pi-studio | MIT |
| pi-atelier | github.com/michaelmjhhhh/pi-atelier | MIT |
| pi-toolkit, pi-extensions, mempalace-toolkit | authored by the maintainer (Joakim Persson) | MIT |
## MemPalace (AI memory)
| Component | Upstream | License |
| --- | --- | --- |
| mempalace (core, MCP server) | github.com/MemPalace/mempalace (PyPI: `mempalace`) | MIT — the GitHub repo declares MIT; the PyPI package's own metadata omits a license classifier, so if you need clearance from the package artifact alone, verify against the repo's `LICENSE` file rather than the sdist/wheel metadata |
| Chromium | chromium.googlesource.com/chromium/src | BSD-3-Clause for Chromium's own code, plus a large set of bundled third-party components each under their own license (see Chromium's own `LICENSE`/`about:credits`). The binary in this image is **not compiled here** — it is the build Playwright downloads for its pinned version ("Chrome for Testing"), installed via `playwright install --with-deps chromium` at `/usr/local/share/ms-playwright/`. Treat Playwright's own distribution terms for that build as authoritative over any summary here. |
## Tooling baked into the base image
| Component | Upstream | License (best effort) |
| --- | --- | --- |
| gosu | github.com/tianon/gosu | Apache-2.0 |
| Node.js | nodejs.org | MIT (bundles components under their own licenses) |
| uv | github.com/astral-sh/uv | Apache-2.0 OR MIT |
| Model calls | up to **three** background calls per consolidation pass (observer, reflector, dropper), each capped at `agentMaxTurns` [16], on the configured memory model — not your session model |
| Latency in your turns | none by construction: workers run from `turn_end`, compaction runs when pi is idle, and the fold itself does no model work |
| Disk | negligible — JSON lines inside a session file that would exist anyway (measured here: `~/.pi/agent/sessions` = 30 MB total, tens of `om.*` entries per session) |
| Context window | **zero until compaction.**`custom` entries do not enter LLM context; only the folded summary does |
| Attention | none once configured; there is no protocol for you or the agent to remember |
If that is still more than you want on a given run, §9's `passive` switch turns
off all proactive work while keeping `recall` and `/om:*` usable.
## 9. Configuration
Global: `~/.pi/agent/settings.json` (persisted in the volume). Per project:
`<project>/.pi/settings.json`, which overrides global. Precedence is
project → global → environment, and the environment can only override `passive`.
| Key | Default | What it changes |
|---|---|---|
| `observeAfterTokens` | `10000` | observer cadence — lower means smaller chunks and more calls |
| "What did we decide 200 turns ago in *this* session?" | observational memory (and `recall` for the exact wording) |
| "What did we decide last month, or on another machine?" | MemPalace (`mempalace_search`, diaries) |
| "What is true *right now* about version X?" | MemPalace knowledge graph |
| "Does another machine need something from me?" | MemPalace coordination log — see [Cross-machine agent coordination](../README.md#cross-machine-agent-coordination) |
| "Why is compaction not losing my session?" | observational memory |
The crisp version: **observational memory keeps a session coherent; the palace
keeps the fleet coherent.** A container recreate wipes neither — but only because
`~/.pi` and the palace both live outside the container filesystem.
## 12. Gotchas
- **Branch-local means branch-local.** Resuming or forking changes which ledger
is folded. Memory that "disappeared" is usually on another branch.
- **`recall` needs an id, not a topic.** If you only have a topic, that is a
palace search, not a recall.
- **A `/workspace` clone is not evidence of what is loaded** — see §7.3.
- **`showWorkerNotifications: true` is not proof of work**; it reports runs, and
an observer that deliberately emits nothing writes no ledger entry and simply
retries after another `observeAfterTokens`.
- **A turn bigger than `keepRecentTokens` splits.** The cut then lands mid-turn at
an assistant message and pi merges two summaries — rare, but it is why a very
large single turn can lose more verbatim detail than you would expect.
- **`git log` in the baked tree needs `safe.directory`** (`/opt` is root-owned):
The pi-toolkit global `AGENTS.md` tells every pi session to read
`~/.agents/skills/pi-extensions/SKILL.md` at start (to fix fork/recall
under-utilisation). That pointer dangles in a container started **without** the
private `skillset` repo mounted. Baking the skill closes that *availability*
gap. `mempalace` is baked for the same reason (memory continuity); since
nothing in pi-toolkit's `AGENTS.md` points to it, the pi-devbox managed block
(`pi-global-AGENTS.append.md`) also adds the matching *proactive-load*
directive ("load the mempalace skill at session start") so a new container
actually picks it up rather than relying on description-matching.
`pi-extensions`'s directive already ships in pi-toolkit's `AGENTS.md`, so only
its skill file needed baking.
## Freshness model (layered — see Dockerfile.variant)
- **`pi-extensions`** — Option 1 + Option 2. The committed copy here is the
*floor*; at build time `Dockerfile.variant` copies `/opt/pi-extensions/skill/`
(the pinned, package-owned source) over it, so a normal build ships the fresh
package copy and a stale-ref / mirror build still ships the snapshot. Keep
`evaluate-extension-usage.py` alongside `SKILL.md` — the skill calls it via
`./`.
- **`mempalace`** — Option 2 only. The `mempalace`*consumer* skill lives only
in the private `skillset` repo (the `mempalace-toolkit` repo ships a
*different* skill, `opencode-mempalace-bridge`), so there is no public
package source to copy from. This snapshot is refreshed manually per release.
**Refresh it with `scripts/vendor-mempalace-skill.sh <skillset-root>`, not
`cp`.** Because the image cannot clone the private upstream, the snapshot used
to be *anonymous* — nothing recorded which skillset commit the bytes came
from, so the only staleness check possible was a hand-maintained phrase canary
in `scripts/smoke-test.sh`, which by construction detects "older than the
phrase I remembered to pin", never "older than skillset main". Two facts now
travel with the file:
| Fact | Where | Kind |
|---|---|---|
| `ARG SKILLSET_SNAPSHOT_REF` in `Dockerfile.variant` | manifest `skillset_snapshot_ref` + OCI label `se.jordbo.pi-devbox.skillset-snapshot-ref` | a **claim** about which commit these bytes are |
| `sha256sum` of this file, measured in the manifest layer | manifest `skillset_snapshot_sha256` | the bytes that **actually shipped** |
The script writes both together, refuses when the upstream file has
uncommitted modifications (no commit describes those bytes), and
`--check` verifies the claim against a real clone. Deliberately an `ARG`
default rather than a CI-resolved value: no credential for a private repo, no
change at any of the four `Dockerfile.variant` build call sites, and a local
`docker build` records the same thing CI does.
Verifying "is this snapshot current?" is **not** a CI job and was deliberately
not made one — see the Unreleased CHANGELOG entry for why (private repo;
another repo's branch must not be able to fail this build; and the artefact it
would guard is read by no host on this fleet). The check belongs where the
skillset actually is: `vendor-mempalace-skill.sh --check` for a maintainer,
and `pi-devbox-version`'s `skills:` section for an agent inside a container.
## Runtime precedence (v1.8.5+)
The baked links are created **early** in `entrypoint-user.sh` (before pi-deploy,
to close a smoke readiness race) with a create-only-when-absent guard, and the
skillset deploy runs **last** and treats them as foreign links. Through v1.8.4
that combination meant the baked snapshot always won: an edit pushed to
`skillset/skills/mempalace/SKILL.md` was invisible in every container until the
next image build (measured on two hosts — live `md5 129bcc4752` vs baked
`5236024fef`, new section absent). Editing those skills *appeared* to work.
`devbox-skill-reconcile` now runs immediately after the skillset deploy and
repoints the links for skills the **skillset owns**, listed one per line in
`skillset-owned.txt`. Precedence, highest first:
1.**user override** — a real directory, or a symlink pointing outside the baked
tree; never touched by anything
2.**live skillset clone** — but only for names in `skillset-owned.txt`
3.**baked snapshot** — everything else, and every skill when no skillset is
mounted
**Which one won is now reportable from inside the container:**
`pi-devbox-version` prints a `skills:` section naming, per vendored skill,
`baked` or `live <repo> @ <sha>` — and for `mempalace` whether that live copy is
identical to the baked fingerprint, at the same commit but with uncommitted
edits, or genuinely divergent. Before that, a stale baked snapshot and a current
live clone were indistinguishable from inside, which is how the freshness of
this file went unexamined for three releases. The section is suppressed with
`--no-skills` on the container-start banner, because `entrypoint-user.sh` prints
the version *before* the links exist and long before the reconcile below runs.
On this fleet, precedence 2 wins for `mempalace` on **every** host — all four
compose stacks mount a workspace containing the skillset — so the baked copy is
exercised only by CI and by a hypothetical no-mount container. Worth
remembering before spending effort on its freshness.
Ownership is per-skill on purpose: `pi-extensions`' authoritative source is the
package repo (copied over the snapshot at build), and `skillset` carries a
downstream copy that can lag, so handing it to the clone would *regress* the
skill. Only `mempalace` is skillset-owned today.
Verify with `readlink -f ~/.agents/skills/<skill>` — not by reading the
entrypoint. Smoke covers both directions (baked resolution with no skillset
mounted, plus a fabricated-skillset run of the reconciler).
Copy `pi-extensions`**from its owner in the table above** — the package
repo's `skill/` (since `a7f3044` co-located it there; `skillset` also carries a
copy, but it is a downstream duplicate and can lag). Copying `pi-extensions`
from `skillset` would regress the snapshot to whatever that repo last mirrored.
`mempalace` is **not** refreshed by `cp` — see the *Freshness model* section
above: `scripts/vendor-mempalace-skill.sh <skillset-root>` is the only thing
that should ever touch that snapshot, because a bare copy can update the bytes
without updating the ref that claims to describe them, which produces a
manifest that confidently lies.
Neither vendored skill has a hand-maintained "last refreshed at" line here on
purpose — one previously existed (skillset `670f7f1`, pi-extensions pkg
`e73cb9f`) and went stale within hours, because nothing forced it to move
when the ARGs did. `670f7f1` is now a cautionary example rather than a fact
worth recording: it is the commit that told agents to hand-stamp `added_by`,
which a later skillset commit (and the pi-devbox edge stamper) withdrew — so a
reader trusting that line would have been pointed at superseded guidance.
Both facts it tried to capture now live somewhere that cannot drift by hand:
| Fact | Where |
|---|---|
| which skillset commit `mempalace`'s bytes came from | `ARG SKILLSET_SNAPSHOT_REF` (Dockerfile.variant) + `skillset_snapshot_ref` in `build-manifest.json`, written *only* by `vendor-mempalace-skill.sh` |
| which pi-extensions package commit was vendored | `ARG PI_EXTENSIONS_REF` (Dockerfile.variant, CI-resolved to a 40-hex commit) → OCI label `se.jordbo.pi-devbox.pi-extensions-ref` and `build-manifest.json`'s `components.pi-extensions`, both read from the actual `/opt/pi-extensions` checkout, not from intent |
When you refresh the `mempalace` snapshot, also update the phrase asserted by
the "mempalace skill snapshot is current" smoke test — it deliberately pins the
**newest** section, because the previous canary grepped a phrase that survived
the very edit that made the snapshot stale, and so passed on stale content.
description: MemPalace agent memory protocol. Use on every session to maintain continuity across conversations — search before answering about past work, write diary entries before session ends, and mine new projects into the palace. Load this skill at session start.
---
# MemPalace Agent Memory Protocol
## Overview
MemPalace gives you persistent memory across sessions via an MCP server. It stores project knowledge (mined from files), conversation summaries (diary entries), and entity relationships (knowledge graph). Without this protocol, you have tools but no habits — and memory without habits is just storage.
**Core principle:** Storage is not memory. Storage + protocol = memory.
## When to Load This Skill
- At the **start of every session** (proactively, before the user asks)
- When the user mentions **past conversations, decisions, or work**
- When working on a **new project or repository** for the first time
- When the user asks about **people, projects, or relationships**
## Session Lifecycle
### Phase 1: Wake Up (session start)
Run these immediately when a session begins, before responding to the user:
1.**Load palace overview:**
```
mempalace_status
```
This returns wing/room counts, the AAAK spec, and the memory protocol reminder.
**Never guess about facts that might be in the palace.** Wrong is worse than slow. Say "let me check" and query.
#### Search Before You *Probe*
The rule above covers **questions**. This one covers **actions** — and it is the one
that actually gets skipped, because mid-task the impulse is to go and *look* rather
than to remember. The palace is a **fleet** record: another machine's agent has
usually already paid the cost of discovering how this environment is wired, and its
notes include the corrections that came afterwards, which a fresh probe cannot show
you.
**Before you SSH somewhere to find out how it is set up, enumerate infrastructure,
or derive a deployment — search.** Concrete triggers, all meaning *search first*:
- about to run `ssh <host> …`, `docker ps`, `systemctl list-units`, `ip addr` to
discover how something is deployed or connected
- about to establish topology: which hosts/runners/services exist, where they live,
which of them can reach which
- about to conclude "this isn't documented anywhere" or "there's no way to know"
- about to assert an environment fact you learned **earlier in this same session**
**That last trigger is the sharp edge.** A compacted session summary is lossy by
design, and a belief you formed 40 turns ago may already be *retracted* in the
palace by another machine. Trusting your own context over the shared record is how a
withdrawn claim gets re-published as fact.
Search broadly before narrowing — fleet knowledge often sits in another machine's
wing, or inside a mined conversation, not where you would file it yourself:
```
mempalace_search(query="<topic> <host> <mechanism>") # no wing filter first
mempalace_search(query="…", wing="<likely-wing>") # then narrow
```
Two or three searches cost seconds. Re-deriving infrastructure costs minutes **and
can be wrong**: a probe shows one host's present state, while the palace records
intent, history, and what was already disproved.
> **Worked example (real, 2026-08-25).** An agent evaluating whether to add an ARM
> CI runner probed hosts directly instead of searching. It concluded "the runner
> lives on synlig" — there are **four** — and that "synlig is on the home LAN" —
> it is an OpenStack VM with a public floating IP that cannot reach the home LAN at
> all. Both facts were already in the palace, the second one as an **explicit
> retraction of the very same mistake** made weeks earlier. The palace also held
> the runner labels and the deliberate `capacity: 1` setting, which the probe never
> revealed. Cost: a wrong recommendation written into the palace twice, then
> corrected twice.
**A search that comes back empty is not an answer — least of all about recent work.**
Semantic search is weakest exactly where the fleet record is freshest: a drawer filed
minutes ago is unranked against a keyword-shaped query, and the drawer you most need
is *by construction* the newest one, because the other machine files its release,
handoff and correction drawers at the **end** of its session. So a single miss proves
nothing. **If the work is 0-2 days old and the first search looks stale or empty,
enumerate before concluding:**
```
mempalace_list_drawers(wing="<wing>", since="<today>") # or room=, or no filter
mempalace_diary_read(agent_name="<you>", wing="<wing>") # the other machine's handoff
```
Enumeration is exact where embeddings are probabilistic. Treat "I searched and found
nothing" as a hypothesis you have not yet tested, and never as licence to go probing.
> **Worked example (real, 2026-08-25, same fleet as above).** An agent asked to
> orient on an in-flight release *did* search first — `"v1.8.6 release run 579
> Docker Hub verification"` — and got back only v1.6.4 / v0.78.0 era hits, because
> the release drawer it needed was **58 seconds old**. It accepted the miss and went
> off to probe Docker Hub and the Gitea API. The user had to prompt "maybe there is a
> note in mempalace"; `list_drawers(wing="pi-devbox", since=<today>)` then returned
> the drawer immediately, along with the diary entry naming the exact open item. The
> rule above was present and correct in this very file at the time — the failure was
> not knowing to *retry differently* after a bad first hit.
#### Mine New Projects
When working on a new codebase for the first time:
1. Check if it's already mined:
```
mempalace_list_wings
```
2. **Decide what to mine — docs first, code never (by default).**
The palace is for *context and intent*, not code recall. Code is better read from the working tree via `Read`/`Grep`/`glob` — always authoritative, never stale. Embedding source code produces thousands of low-signal drawers (e.g. `def __init__(self, ...)` across every class) that pollute search for years.
- `node_modules/`, `.venv/`, `__pycache__/`, `.mypy_cache/`, `.pytest_cache/`, `.ruff_cache/` (the miner respects `.gitignore` but double-check)
Exception: if a code file *is* the documentation (e.g. a heavily-commented reference script, or a protocol definition), file it manually via `mempalace_add_drawer`.
3. **Before mining**, inspect the repo to estimate drawer count:
The miner currently lacks a `--docs-only` or `--exclude-ext` flag (as of v3.3.3). Until it does, either:
- (a) Add a `mempalace.yaml` at the repo root with explicit include globs, OR
- (b) Mine everything, then surgically remove code-sourced drawers via SQL on `~/.mempalace/palace/chroma.sqlite3` (delete by `embedding_metadata.source_file LIKE '%.py'`), followed by `mempalace repair --yes`.
5. If the CLI miner misses a file you *do* want (e.g., `.zsh`, an undocumented extension), file it manually:
#### Feeding opencode session history (opencode + mempalace-toolkit only)
MemPalace has no upstream integration with [opencode](https://github.com/anomalyco/opencode) as of v3.3.3 — `hooks_cli.py` only supports `claude-code` and `codex` harnesses. Opencode persists every turn in a local SQLite DB at `~/.local/share/opencode/opencode.db`, but nothing moves that data into the palace automatically.
On a machine with opencode + the [`mempalace-toolkit`](https://gitea.jordbo.se/joakimp/mempalace-toolkit) installed, session history is fed into `wing_conversations` via `mempalace-session` — either manually, or on a weekly systemd user timer / cron schedule shipped in `mempalace-toolkit/contrib/`. If this is missing, opencode conversations exist only in the local SQLite DB and are invisible to `mempalace_search`.
**How to tell if it's set up:**
```
mempalace_list_wings
```
If `wing_conversations` exists and has a drawer count comparable to the user's opencode session count, session feeding is working. If it's empty or suspiciously small, suggest:
1. Check if the toolkit is installed: `which mempalace-session`.
2. If installed, suggest running `mempalace-session --dry-run` to preview and `mempalace-session` to file.
3. If not installed, point the user at `gitea.jordbo.se/joakimp/mempalace-toolkit` for setup.
**Don't try to paper over the gap by dumping turn-level content into the palace manually via `mempalace_add_drawer`** — that reinvents what `mempalace-session` does with normalization and dedup. Use the tool.
Full routine (triggers, cadence, automation) is in the [`opencode-mempalace-bridge`](https://gitea.jordbo.se/joakimp/mempalace-toolkit) skill and the toolkit's `ARCHITECTURE.md` §5. The two skills pair: this one (`mempalace`) covers using the palace; that one (`opencode-mempalace-bridge`) covers feeding it from opencode.
### Phase 3: Wind Down (session end)
**Always write a diary entry before the session ends.** This is the most important habit.
```
mempalace_diary_write(
agent_name="<your_agent_name>",
entry="<AAAK compressed summary>",
topic="session-summary"
)
```
#### Why still write diaries when sessions may be mined automatically?
On machines running opencode + `mempalace-toolkit`, every session is mined into `wing_conversations` on a weekly (or user-defined) schedule. A common and incorrect conclusion: *"since every turn is captured automatically, writing a diary entry is redundant."* It isn't.
Session mining captures **what was said** (every turn, verbatim). A diary captures **what the session meant** — editorial judgment by the agent who lived it:
- A compressed, recency-scannable summary for the *next* agent's wake-up
Mining raw turns cannot surface these because the words don't exist verbatim — they're the agent's reflection at wind-down. Think of the split as *release notes* (diary) vs. *git log with diffs* (session mine): a repo keeps both because they answer different questions. So does the palace.
**Practical rule:** automated mining does not replace Phase 3. Both systems cover each other's failure modes — a skipped diary is recovered from the raw turns; a missed mine is recovered from the diary summary. For the full treatment (comparison table, retrieval patterns, token economics), see [`mempalace-toolkit/ARCHITECTURE.md` §5 → "Diary vs session mine: why keep both?"](https://gitea.jordbo.se/joakimp/mempalace-toolkit/src/branch/main/ARCHITECTURE.md#diary-vs-session-mine-why-keep-both).
#### AAAK Diary Format
Write diary entries in compressed AAAK format for efficiency. Structure:
```
SESSION:<date>|<what.you.worked.on>|
TASKS:
1.<task.description>→<outcome>|
2.<task.description>→<outcome>|
DISCOVERED:<unexpected.findings>|
ENTITIES:<people.or.projects.encountered>|
<importance: one to five stars>
```
Example:
```
SESSION:2026-04-28|api.refactor+db.migration|
TASKS:
1.refactored.auth.endpoints→split.into.3.modules|
2.added.user.roles.migration→postgres.enum.type|
DISCOVERED:legacy.session.table.unused.since.v2|
ENTITIES:ProjectX;Alice(reviewer)|
***
```
Rules:
- Use dots instead of spaces within phrases
- Use pipes as field separators
- Use arrows for cause/effect or transitions
- Stars indicate session importance (one to five)
- Keep it tight — a future agent should get the gist in seconds
#### What to Capture
Prioritize recording:
- **Decisions made** and their rationale
- **Discoveries** — things that surprised you or that a future session needs to know
- **Unfinished work** — what's pending, what was deferred
- **User preferences** observed during the session
A candidate is **answered** when one of your own events
1. has a **higher `seq`** than the candidate, and
2. joins to it — `metadata.ack_of == candidate.id` (exact, written for you by
`event_ack`) or the same `correlation_id` (the fallback), and
3. carries a **terminal** status: `applied`, `superseded`, `failed`, `blocked`.
Everything else is still owed. Two calls, constant cost.
**Compare `seq`, never `created_at`** — the same reason you resume with
`since_event_id`. Without the ordering test, one terminal reply would suppress
every later ask on the same `correlation_id` for good; verified on a live thread
where a `ready` reply at `seq` 16 sits *before* the request at `seq` 17 that it
obviously cannot have answered.
**On a real mesh, compare `hlc` instead.** `seq` is *replica-local*: it equals
`origin_seq` today only because a single replica authors events for every
machine. Enrol a second replica and a late-syncing peer event gets a late local
`seq`, so two replicas can order the same pair differently and derive different
owed-sets from the same log. Every event already carries `hlc`
(`<millis>-<counter>-<replica_id>`), which is total and causally consistent.
So: compare `seq` while `mempalace_mesh_peers` reports no peers, `hlc` once it
reports any, and `created_at` never. (This is a legitimate use of `mesh_peers` —
choosing an ordering key — not the discredited gate on *whether* to read your
mailbox at all.)
**The failure directions are not symmetric, which is why this is safe to get
slightly wrong.** Local-`seq` skew can make an already-answered item *resurface*
as owed: noise, self-correcting, and visible. A timestamp comparison can
*suppress an unanswered ask forever*: silent and permanent. So if you ever see an
item you know you answered come back, do **not** "fix" it by reaching for
`created_at` — you would be trading the safe failure for the dangerous one.
This also supplies the "taken, not finished" state that looked missing:
`claimed` and `ready` are deliberately **not** terminal, so work you have picked
up keeps resurfacing until you close it out. No extra convention, no new field.
Two consequences worth internalising:
- **"Seen, not doing it" is a legitimate ack** — `status="blocked"` or
`"superseded"` plus the reason. Silence is not, and it is not merely rude:
with no terminal event of yours to join to, the ask stays in the owed set
indefinitely. The *original requester* — and nobody else — can release it from
the other end, but only by saying so explicitly: see **Withdrawing an ask you
sent** below. That is a release by the asker, not an escape for the answerer.
While the ask still stands, only *your* terminal event clears it.
- **Nothing expires, and it should not.** An `open` with no terminal reply is
still live by definition, and the finished threads are valuable history. If
content is genuinely perishable ("do not push to main for the next hour"), say
so in `metadata.expires_at` — metadata is stored verbatim — and honour it as a
hint when reading. An old `open` that the derivation still counts as owed is a
signal, not garbage: it means somebody asked and nobody answered.
### Writing to another machine
- **Address the stamped name you actually saw** in a `from_agent` field, e.g.
`pi@tor-ms22`. A bare `pi` reaches nobody's mailbox once stamping is live, and
older events in the log still carry bare names — do not copy them.
- **The rule runs in reverse too: what you put in YOUR OWN `from_agent` decides
where every reply to your event goes.** Nothing stops you writing a synthetic
or borrowed identity there, and a reply is always addressed back to exactly
that string — so if no live session ever runs as it, the reply is stored,
searchable, and delivered to no one. Measured cost: a directed ask sent under
a synthetic sender got two correct replies, one of them an urgent security
finding, and both sat unread for ~2h20m because nobody's mailbox was that
identity (RFC 003 §7.13). Authoring under a synthetic name is fine for a
deliberate control experiment — this fleet does it on purpose — but then
**name the real identity to reply to inside the body**, because the address
line is not a safe place to also carry provenance.
- **Use `status="open"` only when you truly need an answer.** It places an
obligation on another machine.
- **Never broadcast an ask.** `to_agent="*"` + `status="open"` obliges everyone
and therefore no one.
- **Always set a `correlation_id` on a directed `open`,** and reply with the
same one. It is not just for reconstructing a conversation later: it is the
join the owed-set derivation depends on. An uncorrelated ask can only ever be
closed by an `event_ack` (which sets `ack_of` for you) — a plain reply cannot
be matched to it at all.
- **Corrections are new events, never edits.** Say explicitly what you retract
and name the id — drawer or event — that carried the withdrawn claim.
- **Withdrawing an ask you sent: state it, never imply it.** Your release only
counts when the terminal event (a) comes from the same `from_agent` that sent
the ask, (b) is directed at that recipient exactly — never `*`, so a broadcast
can neither oblige nor release, (c) carries a terminal status (`claimed` and
`ready` are not terminal and do not release anything), (d) is strictly after
the ask, (e) joins it via `ack_of` or the same `correlation_id`, **and (f)
names that ask in `metadata.withdraws` or `metadata.closes`.** Prose in the
body does not count, and neither does a bare terminal event on the
correlation: inferring release from *any* terminal would let your own
bookkeeping silently delete a real obligation, so the release must be stated.
Needs toolkit ≥ `e2b060a` (image ≥ `v1.9.1`) — check with
`grep -c isWithdrawn /opt/mempalace-toolkit/extensions/pi/mempalace.ts` and
read `0` as "my withdrawal will have no effect on their mailbox". Measured
cost of getting it wrong: a `v1.8.13` rollout ask was withdrawn by its sender,
who recorded it as done; the recipient's derivation never saw the release and
still reported the ask owed **41 hours later**, for a release that device
never installed — and the asymmetry was invisible from the sender's side
(RFC 003 §3.3 clause 4).
- **Put a retraction where the reader will look.** A *directed open ask* reaches a
live agent's mailbox; a **terminal-status event reaches no mailbox at all**, and
a *drawer* is what a future semantic search finds. If you filed advice as a
drawer and later withdraw it, file the withdrawal as a drawer too — otherwise
the next agent finds your original confident advice and no trace of the
correction. (This is a real incident, not a hypothetical.)
- **Hand over exact content as an artifact**, not prose: `mempalace_artifact_put`
or `mempalace_patch_submit` store bytes with a sha256, and the event references
the id. Never paste a diff into a body and hope it survives.
- **Waiting on a specific reply?** `mempalace_event_wait` blocks with backoff —
do not poll `event_list` in a loop. A timeout there is a normal result, not an
error.
## Palace Structure
### Wings
Wings are top-level categories, typically one per project or domain.
**NAMING CONVENTION — decided 2026-09-06 by Joakim: bare project names, no `wing_`
prefix.** `home-network`, `pi-devbox`, `mempalace-toolkit` — *not* `wing_pi-devbox`. The
mass is already there (`pi-devbox` 2061 drawers vs `wing_pi-devbox` 25), and a prefix
present on some wings and absent on others turns every read into a guess about which
spelling holds the content.
- Named after the project directory or domain (e.g., `cli_utils`, `home-network`)
- **Always pass `wing` explicitly to `diary_write`.** Omitting it defaults to
`wing_{agent_name}`, which mints or feeds a *parallel* wing — this tool default, not
anyone's sloppiness, is the mechanism that produced the drift. Measured harm
(2026-09-06, `pi@mbp-m1-2020`): a diary entry written with `agent_name=pi` and no
`wing` landed in `wing_pi` while that agent's history lives in `pi-devbox`, so a
`diary_read` scoped to `pi-devbox` showed **no trace of it**. A wing-scoped read that
silently returns an incomplete history is the worst failure mode a memory store has.
- **Legacy `wing_*` wings are frozen and documented, not renamed.** `wing_conversations`
(written by the session feeders), `wing_pi`, `wing_pi-devbox`, `wing_pi-tor-ms22`,
`wing_pi-devbox-emb7kj`, `wing_mempalace`, `wing_orchestrator`, `wing_code` all still
hold real content. **When searching for history, check both spellings** — this is the
practical cost of the drift and it does not go away by decree.
- If a migration is ever done, the acceptance criterion must be at the **relationship**
level: chunk ids still resolve to their parent, and `diary_read` returns the same entry
set before and after. Per-wing drawer counts can look correct while the relationships
underneath are broken, because a count query never touches them.
#### Shared palace: multiple harnesses, and possibly multiple machines
A single palace can be fed by multiple coding-agent harnesses, and — when
`MEMPALACE_REMOTE_URL` points at a central palace — by multiple *machines*. On
this machine the palace is shared between **opencode** and **pi** (Mario
Zechner's pi-coding-agent). Implications:
- **`wing_conversations` mixes sources.** Both harnesses' session feeders write into the same wing. To tell them apart, look at the `source_file` metadata on each drawer:
- `pi_<uuid>.jsonl` → pi session
- `<slug>_ses_<id>.jsonl` → opencode session
- The first chunk of each session also carries a `| source: opencode` or `| source: pi` marker in the synthetic header line.
- **Other wings may belong to other harnesses.** For example `wing_pi` is pi's diary, not opencode's. Don't assume every diary entry was written by you — check `agent_name` on the entry.
- **Session feeders run on different schedules.** Pi sessions are fed Tue 03:00, opencode sessions Mon 03:00 (launchd `Weekday`: `0`/`7`=Sunday, `1`=Monday, `2`=Tuesday — misreading this by one day is easy). Recent sessions from either harness can lag the palace by up to a week, so absence-of-evidence in `wing_conversations` is not evidence-of-absence for recent work.
- **Reading another harness's diary is useful.** When orienting after a gap, `mempalace_diary_read agent_name=pi` (or whichever sibling agent has been active) often gives a fresher picture than waiting for the conversations feeder to catch up.
When the palace is **central** (shared across machines), these further things apply:
- **Check which machine a conversation came from.** Transcripts are fed per device, so `source_path` reads `…/mempalace-feed/<device>/pi_<uuid>.jsonl` while the displayed `source_file` is only the basename. One search can legitimately return hits from several machines at once — look at the device segment before attributing a decision to *this* project.
- **Provenance is stamped for you — leave it alone.** Drawers carry `device` and `agent_kind` metadata (plus `device_source`/`agent_kind_source` recording *how* each was determined, so an inference is never mistaken for a fact). You do **not** set these, and you no longer set `added_by` either: the pi bridge defaults the writer field to `<harness>@<device>` on `add_drawer`/`checkpoint`/`mine`/`event_append`/`artifact_put`, and prefixes diary entries with `HOST:<device>|`, from host-supplied `$MEMPALACE_PI_DEVICE`. RFC 001 §7.3.2 ranks "agent stamps it via a skill instruction" as the *worst possible* place for exactly the reason you would expect — it is per-call boilerplate that gets forgotten, and it did: the agent who wrote the previous version of this bullet then filed its own provenance drawer as `added_by=checkpoint`. **Confirm the bridge in your image actually stamps before trusting it:** the extension is baked at image build time, so a container on an image older than the stamping commit (pi-devbox < v1.8.7) stamps nothing while still satisfying both gates — the env vars are set and the code is simply absent. Check with `grep -c MEMPALACE_PI_DEVICE "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)"`; zero means keep passing `added_by="<harness>@<device>"` and a manual `HOST:<device>|` diary prefix until the container is recreated on a newer image. Two things remain yours: pass `source_drawer_id` on `kg_add` (triples have no provenance field, so that pointer is the only path back to a device), and pass an explicit `added_by` **only** when deliberately filing on behalf of another device — and when you do, it **must** be `<harness>@<device>`. A bare nickname (`pi-devbox-claude`) has no `@device` to parse, so `agent_at_device` cannot attribute it and the drawer is unattributable *by rule*, not by lag: it survives every future stamp run with no `device`, and on a shared palace a device-less drawer is one nobody can later scope, audit or clean up per machine. Measured 2026-09-06: 11 drawers on `tor-ms22` were filed this way — including the credential rows, i.e. exactly where "which machine measured this?" matters most — by an agent that had passed its own chosen nickname on every call. Its *diary* entries escaped, because `HOST:<device>|` in the AAAK text recovers the device. **Diaries self-heal; plain drawers do not.** The safest habit is the one above: pass nothing and let the bridge stamp. Never invent values for `device`/`agent_kind`/`origin_device` — a fabricated value is worse than a blank, because it silently corrupts a future merge.
- **Metadata is invisible to search — so check the text, not the fields.** `search` results are built from a fixed key list and `diary_read` returns content, so neither ever shows `device`/`added_by`. Only `mempalace_get_drawer` reveals them. This is why diary entries carry an in-text `HOST:<device>` marker: it is the only attribution a reader actually sees. **A diary entry with no `HOST:` marker predates the convention and may be from any machine — do not assume it is this one's history.**
- **Mined drawers carry the MINE date, not the session date.** When history is imported, or re-mined on the palace host, `filed_at`/`created_at` is the *import* time — so sorting by them does not give chronological order. Real session time is recoverable from the UUIDv7 in `pi_<uuid>.jsonl`: the first 12 hex digits are milliseconds since the epoch (and UUIDv7 sorts lexicographically in time order, so a plain filename sort is already chronological). Agent-authored drawers and diaries have no such backdoor — for those `filed_at` is the only chronology, which is why it must never be restamped.
- **Beware the timezone mismatch when you combine those.** Palace `filed_at`/`created_at` are naive timestamps in the palace host's local time, while a UUIDv7 decodes to UTC. Comparing them directly introduces a silent offset (2 h for a CEST host). Normalise before drawing conclusions about ordering.
- **`agent_name` is not device-scoped.** `mempalace_diary_read(agent_name="pi")` returns *every* machine's `pi` diary, interleaved. Read the entry before assuming it is your own history — and note that a container cannot tell you which machine it is on (`hostname` is a docker hash, `$DEVBOX_HOST_ALIAS` is generic). `$MEMPALACE_PI_DEVICE` is the cheap answer; `ssh -F ~/.ssh-local/config host hostname` is the independent one.
- **One writer, no queue.** A concurrent mine returns a structured `already-running` error rather than waiting its turn, and one large mine can make the palace unresponsive to every client for minutes. After another client's mine, call `mempalace_reconnect` to see the new drawers. A client-side timeout is not evidence of failure — verify before retrying, or you file a duplicate.
### Rooms
Rooms are aspects within a wing:
- `fzf`, `scripts`, `configuration`, `general` — whatever the miner detects
- Diary entries go into rooms by topic tag
### Drawers
Drawers hold verbatim content — never summarized, always searchable.
### Tunnels
Cross-wing connections linking related content across projects.
### Knowledge Graph
Entity-relationship triples with temporal validity. Query with `mempalace_kg_query`, browse with `mempalace_kg_timeline`.
## Troubleshooting
| Problem | Fix |
|---|---|
| "No palace found" | Run `mempalace init <dir>` then `mempalace mine <dir>` |
| "Error finding id" after mining | Run `mempalace repair --yes` then `mempalace_reconnect` |
| Search returns irrelevant results | Use `max_distance=1.0` for stricter matching; add `wing` filter |
| Miner skips file types | File manually with `mempalace_add_drawer` or use `--no-gitignore` |
| Stale results after external changes | Call `mempalace_reconnect` |
## Anti-Patterns
- **Don't guess when you can search.** If a question touches past work, search first.
- **Don't probe what the fleet already knows.** Before SSH-ing into a host, enumerating infrastructure, or deriving how something is deployed, search the palace. A probe reveals one host's present state; the palace holds intent, history and prior corrections — including the ones that contradict what you are about to conclude.
- **Don't trust this session's context over the palace.** A compacted summary is lossy, and another machine may have corrected the fact since. Verify load-bearing environment claims against the shared record before acting on them.
- **Don't take one empty search as proof the palace is silent.** Fresh drawers rank worst, and the drawer that matters is usually the newest one. For anything 0-2 days old, enumerate with `mempalace_list_drawers(since=…)` and read the other machine's diary before you go and probe.
- **Don't infer elapsed time from session or container boundaries.** A restart isn't a new day. Compare the actual timestamp (`timestamp` / `created_at`) against the current date/time before saying "yesterday", "last week", etc.
- **Don't skip the diary.** A session without a diary entry is a session forgotten.
- **Don't summarize drawer content.** File verbatim — the embedding model needs the original words.
- **Don't mine .git directories or node_modules.** The CLI miner respects .gitignore by default.
- **Don't create duplicate drawers.** Use `mempalace_check_duplicate` before adding manually.
- **Don't treat the palace as a task list.** It's for knowledge and context, not todos.
- **Don't broadcast an ask, and don't leave one unanswered.** On a shared palace, `to_agent="*"` + `status="open"` obliges every machine and therefore none of them. And don't expect acking to tidy your mailbox: `status` is immutable, so the event keeps matching either way — what a terminal reply buys you is that the *derived* owed set (see *What you actually owe*) stops counting it. Leave asks unanswered and that set only grows, until everyone learns to stop looking. "Seen, not doing it" is a complete answer — silence is not.
- **Don't assume you would have heard.** Nothing pushes another machine's message into your session. If you did not run the mailbox query at wake-up, a correction addressed to you by name can sit unread while you confidently rebuild the thing it warned you about.
- **Don't author an ask under an identity nobody runs as, including your own throwaway labels.** The failure is symmetric to the one above: it is not that you missed a message, it is that nothing could ever have delivered the reply to you, because you addressed it at a name instead of an agent. If you must use a synthetic sender for a control or an experiment, say inside the body who should actually receive the reply.
- **Don't invent provenance metadata, and don't hand-stamp it either.** An earlier version of this list told you to set `added_by="<harness>@<device>"` by hand; that instruction has been withdrawn, because RFC 001 §7.3.2 places provenance at the client/server boundary and the pi bridge now does it uniformly (see *Provenance is stamped for you* above) — but the withdrawal only holds where the bridge is live, so run the one-line check in that bullet first; on an older image hand-stamping is still the only signal a hand-filed drawer gets. DO NOT invent values for the palace's own metadata fields (`device`, `agent_kind`, `origin_device`): those are stamped by infrastructure that also records *how* each was determined, and a fabricated value is worse than none because it silently corrupts a future merge. DO pass `source_drawer_id` on `kg_add`. And never put a machine name in a diary's `agent_name` — it becomes the wing name and hides your entries from `diary_read`.
- **Durable work goes in `/workspace`** (it's the host filesystem, UID-aligned —
what you write appears with the user's normal ownership on the host).
- **Runtime-installed system packages and language toolchains are ephemeral.**
If a task needs them reproducibly, it belongs in the image (Dockerfile) or a
project manifest, not an ad-hoc `apt install`. Tell the user when you install
something that won't survive.
- **`~/.pi` is a named volume**, so things baked into the *image* under
`/home/<user>/...` are **shadowed** by the volume on existing containers and
only seen on a fresh volume. Image-owned content that must always be live
belongs under an image path like `/usr/local/...` or `/opt/...` and is linked
in by the entrypoint — not dropped into a home directory that a volume covers.
### Editing a skill: resolve the symlink before you touch it
`~/.agents/skills/` itself is in the **ephemeral container layer**, rebuilt by
`entrypoint-user.sh` on every start from two sources — so *where a skill really
lives* decides whether your edit survives:
```sh
readlink -f ~/.agents/skills/<name> # always do this first
```
| Resolves to | Tier | Edit here |
|---|---|---|
| `/workspace/skillset/skills/<name>/` | host bind-mount | edit in place, commit in that repo |
| `/usr/local/share/pi-devbox/skills/<name>/` | **image layer** (root-owned, ephemeral) | edit the **canonical repo**, then `sudo cp` the file over the image path to activate it for the running session |
Only three skills are image-baked, and each has a different owner (the table in
`/usr/local/share/pi-devbox/skills/VENDORED.md` is authoritative):
| `pi-extensions` | the `pi-extensions`**package** repo → `skill/`. `Dockerfile.variant` copies it over the vendored snapshot at build, so also refresh `pi-devbox`'s `rootfs/.../pi-extensions/` copy to keep the fallback floor from diverging |
| `mempalace` | the private `skillset` repo → `skills/mempalace/` (manual snapshot refresh per release) |
**Editing through the symlink into `/usr/local/...` is silently lost on the next
recreate** — and worse, it diverges from the canonical repo that every *other*
consumer (host pi, opencode) reads.
**Shadowing gotcha:** image-baked links are created **first** and only when the
name is absent, and the later `deploy-skills.sh --bootstrap --prune-stale` pass
treats them as foreign links and leaves them alone. So for a name present in
**both** the image and `skillset` — currently `mempalace` and `pi-extensions` —
**the image copy wins**, and a `skillset` edit to that skill has no effect in
the container. Verified 2026-07-29: the baked `mempalace` snapshot carries a
*Temporal grounding* section (`pi-devbox``904fe85`) that the `skillset` copy at
its snapshot point (`8e8db64`) lacks — containers load the richer baked text
while `skillset` consumers get the older one. When you change one of those two,
decide deliberately which copy is canonical and sync the other.
## 2. Interactive shell vs. your tool shell (a real footgun)
The conveniences below are defined in `~/.bash_aliases` and **only exist in an
interactive login shell.** Your `bash`*tool* runs non-interactively, so these
are "command not found" there — you must spell out the underlying command.
| Interactive alias | Non-interactive equivalent to actually run |
If a command "works in my terminal but not when the agent runs it," this alias
gap is the first thing to suspect.
### A negative result is usually your own filter
**When you are about to report that something is absent, unreachable, or not
running, the filter you wrote is the prime suspect — not the thing.** This
environment produces false negatives cheaply, and they are convincing because
the command "succeeded". Three real instances from one session, all wrong, all
mine:
| Claim I made | Why it was false |
|---|---|
| "`tor-ms22` is not in the SSH config" | `grep … \| head -20` — the entry was at **line 454**. `~/.ssh/config` here is ~500 lines. |
| "the Docker host has no `docker`" | non-interactive SSH `PATH` lacks `/usr/local/bin` (§2, §3). It was at `/usr/local/bin/docker`. |
| "no ControlMaster is running" | pattern `ssh ` (trailing space) cannot match a master: those processes **rename themselves** to `ssh: <controlpath> [mux]`. |
| "the credential is not in the palace" | scanned `embedding_metadata.string_value` only. Drawer **text** lives in `embedding_fulltext_search_content.c0`; 554k metadata rows proved nothing. |
| "this token is dead — 401" | probed it against the **wrong issuer**. A 401 from an instance that never issued the credential is not evidence about the credential. |
| "that host is unreachable, can't test" | tried ports 443 and 80. It was on **3000**, and the env var I already held (`GITEA_EGL_HOST`) stated the scheme and port. |
| "this repo has no `## Unreleased` convention" | read `CHANGELOG.md`**once**, minutes after a release commit had renamed that section to a version heading. 33 commits touch `## Unreleased`. A snapshot cannot show you a cycle. |
Habits that would have caught all three:
```sh
# don't cap the output of a search whose answer you don't already know
grep -n -i -A6 'tor-ms22' ~/.ssh/config # not | head -20
# on the host, resolve the binary instead of trusting PATH
ssh -F "$HOME/.ssh-local/config" mac 'command -v docker || ls /usr/local/bin/docker'
# match a process's ACTUAL argv, not the name you imagine
ps -eo pid,etime,args | grep -Ei 'mux|mosh|ssh'
# to learn a repeating PROCESS or convention, read history, not the file. A
# file's current content is one frame of a cycle, and the frame you happen to
# catch may be the one where the thing you are looking for was just consumed.
Absence has to be *earned*, so spend the extra command there.
### …and a positive result only proves what you *actually asked*
An earlier version of this section claimed "a positive result needs no such
scepticism — it carries its own evidence." **That is false, and believing it
cost a later session three more wrong findings.** A positive result is evidence
about the question your command really posed, which may not be the question you
meant. The failure is invisible precisely *because* the command succeeded.
| Claim | The command succeeded — at answering something else |
|---|---|
| "EGL git over SSH works" | `ssh git@gitea.egl.lan` greeted me as `joakimp`. `~/.ssh/config` had `Host gitea*` → `HostName gitea.jordbo.se`, so I authenticated **to the wrong instance**. The real EGL account is `ecsjper`. |
| "the port config regressed" | compared `ssh -G` output against `2222` — a value produced by **my own earlier `-p 2222` flag**, not by the config. I reported the user's edit as a regression it never caused. |
| "the CI runners authenticate with this token" | pure fabrication, contradicted by my own scan output already on screen. The runners use per-runner `REGISTRATION_TOKEN`. |
Two habits that actually catch this class, both cheap:
```sh
# 1. ask which RULE captured your hostname before trusting any ssh result.
# ssh_config is first-obtained-value-wins PER KEYWORD, not per block: a
# specific block only wins the keywords it declares, so a later `Host gitea*`
# still supplies HostName unless the specific block restates it.
Use the pi extensions (pi-fork, pi-observational-memory, ssh-controlmaster) effectively in the pi coding agent harness. Load this skill only when running inside pi (detection - `fork` and `recall` are present in your tool list, or `pi --ssh` was used to start the session). pi-fork dispatches focused subtasks to forked agents at fast/balanced/deep effort tiers; pi-observational-memory compacts long sessions into recallable observations + reflections; ssh-controlmaster rewires pi's read/write/edit/bash tools to execute on a remote host over a multiplexed SSH connection. Also covers the context ladder L0-L4 and the `task` tool (pi-task: isolated child, immutable spec, machine-checked envelope, write-boundary diff) that is the DEFAULT for delegated work, with `fork` reserved for read-only exploration and parallel opinions, plus the `fork-gate` hook that enforces the split. This skill covers rung and tier selection, task design, boundary discipline, when to use recall, and remote-pi mechanics.
---
# Pi Extensions: pi-fork + task/fork-gate, pi-observational-memory, ssh-controlmaster
## When to Load This Skill
Load only when **both** of these are true:
1. You are running inside the **pi coding agent harness** (not Claude Code, not opencode, not any other harness).
2. The `fork`, `task` and/or `recall` tools appear in your available tool list, **or** the session was started with `pi --ssh ...`.
If you do not see those tools, this skill does not apply — skip it. Other harnesses do not have these extensions and the patterns below will not work there.
This skill is most useful at the start of any non-trivial session where you may need to dispatch parallel subtasks, where the conversation is likely to compact (sessions running > ~80k tokens), or where pi is operating against a remote host.
## Pi extension landscape (where the wiring lives)
Pi has **two distinct extension locations** and it's easy to look in the wrong one:
| Location | Mechanism | Examples |
|---|---|---|
| `~/.pi/agent/extensions/*.ts` (or `.ts.off`) | **Local extensions** — TypeScript files, usually symlinks into `/opt/pi-extensions/extensions/` or similar. Toggled via `/ext` slash command. | `ssh-controlmaster`, `git-checkpoint`, `notify`, `todo`, `mempalace`, `mcp-loader`, `ext-toggle`, `confirm-destructive` |
| `~/.pi/agent/git/<host>/<owner>/<repo>/` | **Package extensions (git-installed)** — git-cloned npm packages registered via the `packages` array in `~/.pi/agent/settings.json`. | `pi-fork` (`github.com/elpapi42/pi-fork`), `pi-observational-memory` (`github.com/elpapi42/pi-observational-memory`, **default branch `master`** — a `main` branch does not exist, so `pi install git:...` resolves against `master`) |
| `~/.pi/agent/npm/node_modules/<pkg>/` | **Package extensions (npm-installed)** — `pi install npm:<pkg>`; recorded in `packages[]` as `npm:<pkg>`. | `pi-atelier` (status rail + sidebar TUI) |
| `/opt/<pkg>/` — **pi-devbox containers only** | **Vendored package extensions** — cloned into an image layer at build time with `node_modules` baked, then registered at container start by `entrypoint-user.sh` via `pi install /opt/<pkg>`. Recorded in `packages[]` as a **relative** path (`../../../../opt/pi-fork`) that resolves out of `~/.pi/agent` into the image layer, so it survives volume recreate. | `/opt/pi-fork`, `/opt/pi-observational-memory`, `/opt/pi-studio` |
When the user asks how to use "the X extension", **check all of these** — `find ~/.pi/agent -maxdepth 4 -name "*X*"` covers the first three, and `ls -d /opt/*X*` the fourth. The `/ext` slash command shows the local-extensions list with enable/disable state. There is also a distinct skill-bundled-script category (e.g. `ci-release-watcher`'s `ssh-control-master-setup.sh`) which is **not** a pi extension at all — it's a helper script inside a skill. Don't conflate the three.
**In a pi-devbox container, do not conclude "pi-fork isn't installed" because `~/.pi/agent/git/` is empty.** It is deliberately absent: `Dockerfile.variant` vendors to `/opt` and installs by local path, because a build-time `pi install git:...` would write into `~/.pi/agent`, which the named volume then shadows on first run.
### Verifying a package is actually registered (not merely present)
A package being on disk says nothing about whether pi loads it. Registration means an entry in the `packages` array of `~/.pi/agent/settings.json`. **Check the array, never grep the file:**
```bash
jq -e --arg n pi-fork \
'(.packages // []) | any((type == "string") and (. == "npm:" + $n or endswith("/" + $n)))'\
~/.pi/agent/settings.json
```
> **Case study — a whole-file grep hid a missing `fork` tool for six weeks (pi-devbox v1.0.0 → v1.6.3, found 2026-07-29).** `entrypoint-user.sh` guarded its `pi install /opt/<pkg>` loop with `grep -q "$_name" ~/.pi/agent/settings.json`. But `settings.example.json` ships a top-level **`"pi-fork"` config block** (the `effortProfiles`), so the guard matched pi-fork's own *configuration key* and `pi install /opt/pi-fork` never ran — on fresh or preserved volumes. Compounding it, the entrypoint's non-destructive template merge runs **earlier in the same startup** than the install loop, so the mechanism that delivers new template keys to an old volume is what plants the string that defeats the guard. `pi-observational-memory` and `pi-studio` escaped only by luck: the template key is `observational-memory` (no `pi-` prefix) and there is no studio block. Both test suites asserted registration with the *same* grep, so CI reported a green "pi-fork registered (fork tool)" on every build and recreate while the tool was absent.
>
> **Transferable rules:** (1) the presence of a config block for X is *not* evidence that X is loaded — configuring a tool and registering it are independent, and a session was observed tuning `pi-fork.effortProfiles.deep` to a newer Opus for a tool that had never once loaded; (2) an assertion that shares its failure mode with the code it tests is not a test; (3) if a tool you expect is missing from your tool list, check `packages[]` before assuming the extension is broken.
**Forensic check — did this tool *ever* run on this machine?** Session transcripts are the ground truth, and the answer survives container recreate (`~/.pi` is a named volume):
A tool that has never been called simply has **no line** — that absence is the proof. `evaluate-extension-usage.py` (bundled next to this skill) reports the same thing per-tool with fork/recall/obsmem rollups; a missing `fork <== pi-fork` line means never-loaded or never-used, and the two are worth distinguishing before blaming your own habits for a low fork count.
### `/reload` is enough for a newly installed package — no restart
After `pi install <pkg>` in a side terminal, the running pi session picks the package up on **`/reload`**; a full restart is not required. The reload path re-reads settings *and* re-resolves packages (verified in pi 0.82.1):
-`dist/core/agent-session.js` → `reload()` calls `settingsManager.reload()`, then `resourceLoader.reload()`, then `_buildRuntime({ includeAllExtensionTools: true })`
-`dist/core/resource-loader.js` → `reload()` calls `settingsManager.reload()` and then `packageManager.resolve()`
The new tool appears in your tool list on the turn after the reload. Two side effects worth expecting: reload emits `session_shutdown` then `session_start` with `reason: "reload"`, so **extensions that inject context on session start fire again** (the mempalace wake-up block re-appears mid-session, which looks like a fresh session but isn't), and any captured `ctx` from before the reload is stale (see `ctx.reload()` in pi's `docs/extensions.md`).
## Why These Extensions Belong Together
pi-fork and pi-observational-memory are symbiotic. **pi-fork burns context** (each fork dispatches a focused subtask whose detailed exploration would otherwise pollute your main thread). **pi-observational-memory preserves context** (when the main thread eventually compacts, observations + reflections survive the fold and can be recalled by ID). Aggressive forking only works long-term if the surviving summary is high-fidelity, and OM only earns its keep when it's preserving genuinely valuable distilled work.
ssh-controlmaster is orthogonal but composes cleanly: when pi is operating remotely, fork still spawns local sub-agents (each fork *itself* doesn't ssh), but their `bash`/`read`/`write`/`edit` calls do — see Part 3 caveats.
---
## Part 1: delegating work — `task` and `fork`
### Decide the rung BEFORE the brief (read this first)
Two tools run a child agent. They differ in one thing, and it decides the
quality of what comes back: **what the child sees.**
| tool | child sees | rung | gives you | use for |
|---|---|---|---|---|
| **`task`** (pi-extensions `task.ts`, wraps `pi-task`) | **only your spec** — goal, named files, curated facts | L0–L2 | immutable spec, PASS/FAIL envelope, per-root boundary diff, audit dir, budgets | **any delegated work that changes files or must obey a rule** — the default |
| **`fork`** (pi-fork) | **your entire branch**, brief appended last | L4 | prose report, effort tiers, parallel dispatch from one message | read-only exploration that needs this conversation; N independent opinions |
**Pre-flight before any `fork(...)` — one *yes* makes it a `task`:**
1. Does the brief say *do not / only / never / must not*?
2. Will the child write, edit, commit or push anything?
3. Do I want a PASS/FAIL I can check, rather than prose?
Why the text alone did not work (and why this is now enforced): the rule above
lived in this skill and in the global AGENTS.md for months and was still
violated by agents that had just read it — five fork briefs in one session on
2026-09-17, all carrying "do not", one of which returned confident verbatim
quotes that did not exist. `fork` is a *tool*: its self-recommending
description ("implementation, testing, review…") is in the tool list every turn
and survives compaction; this skill is gone after the first compaction, and
`pi-task` was a CLI to be remembered. Two structural fixes shipped 2026-09-19:
- **`task` is a tool** (`extensions/task.ts`), so both rungs sit in the tool
list with the decision rule in their descriptions, and `promptGuidelines`
puts the rule in the system prompt where compaction cannot remove it. It
rejects, before spending a model run, the two spec errors that make a
violation certain (see "roots" below) and serialises overlapping tasks.
- **`fork-gate`** (`extensions/fork-gate.ts`) is a `tool_call` hook that BLOCKS
a fork whose brief contains a prohibition, a write boundary or a clause-initial
file-changing imperative, and returns the `task(...)` call to make instead.
It matches wording, not intent: a genuinely read-only exploration brief that
trips it is rephrased, and one that cannot be rephrased without its
prohibition needed `task` all along. `PI_FORK_GATE=off` logs instead of
blocking; `/ext` disables it entirely.
Minimal call:
```
task(id="slug", goal="…verbatim; the child has NO other context…",
| `deep` | opus | architecture decisions, security analysis, concurrency reasoning, ambiguous debugging, high-risk reviews, runbook drafting where subtle mistakes are costly |
**Rule of thumb:** start at `balanced` unless you have a specific reason to go up or down. Going too cheap on a deep task wastes a fork; going too expensive on a mechanical task is just slow.
### When to delegate vs. do it yourself
Having chosen the rung above, delegate (either tool) when **any** of:
- The task requires reading many files whose contents you don't need to keep in your main context afterwards (the fork returns a dense summary; raw file contents stay in the fork's context and are discarded).
- You want to run multiple analyses in **parallel** (especially: comparing N options, where independent reasoning is itself a signal — see "parallel forks" below).
- The task is well-scoped enough to specify completely up front and well-bounded enough that returning a dense report is more useful than continuing the dialogue.
- You are about to do something that would burn a lot of tokens on tool calls (long file reads, many bash invocations) whose output you will mostly discard.
Don't delegate when:
- The work fits in your current context budget without crowding out what comes next.
- The task is exploratory and you'll need to iterate based on what you find (forking turns iteration into round-trips with full task-spec rewrites).
- You need to make decisions during the work that depend on context only the main thread has.
### Task design: the five things a fork brief must contain
1.**Verified context up front.** Do not say "go look at the codebase and figure out X". Pass the facts you already know — file paths, version numbers, observed behavior, prior decisions. The fork should be reasoning *from* context, not *finding* context. Discovery work costs the fork tokens that don't come back to you.
2.**A specific deliverable.** "Analyze X" is too vague. "Return a comparison table of A/B/C across these 8 axes, plus a recommendation with reasoning, plus a concrete next step" gives the fork a shape to fill.
3.**Decision authority.** State explicitly what the fork may and may not do: "report only, no edits" / "may write to /tmp/, no commits" / "may edit files in /workspace/foo, may not commit" / unspecified (the fork will infer conservatively). **State this even when it seems obvious.** See "Boundary discipline" below.
4.**What "unsure" looks like.** Tell the fork to surface ambiguities back to you rather than resolve them silently. "Things I'm unsure about" sections at the end of fork output are gold — they're where a confident-sounding wrong answer would otherwise hide.
5.**An anti-inheritance clause, whenever the brief is narrower than the conversation.** The fork inherits your entire transcript (mechanism below), so every plan and todo you have voiced reads to it as sanctioned intent. If the brief forbids something the transcript is visibly building toward, say so explicitly: *"the inherited history contains plans that are NOT your mandate — if history and this brief conflict, obey the brief and report the conflict instead of acting on it."* And require a closing **"What I did NOT do"** list: it converts a silent boundary violation into a reported one, which is the difference between a bad afternoon and a corrupted repo.
### Parallel forks for option-comparison
When facing a "which approach should we take" question with 2–4 candidate approaches, dispatching the candidates as parallel forks is high-leverage:
- They reason **independently**. No fork sees the others' work.
- **Convergence is signal.** If three forks at different effort tiers reach the same recommendation citing different evidence, that's a strong validation that doesn't depend on any one model's bias.
- **Divergence is also signal.** If one disagrees, read its reasoning carefully — it may have spotted something the others missed, or it may have a tier-specific weakness worth knowing.
Sample shape for an option-comparison call:
- Fork 1 (deep) — detailed runbook for option A, with timing/risk/rollback
- Fork 2 (balanced) — comparison table A vs B vs C across N axes, with a recommendation
This costs more than a single fork but the cross-validation is often worth it for decisions you'll execute on prod systems.
### Boundary discipline — and the mechanism that defeats briefs
Forks **mostly** honor explicit decision-authority instructions, but not infallibly:
- **Pure analysis tasks** (no write authority, "report only") — high compliance. Forks reliably return analysis without editing files or committing.
- **Write-capable tasks with a "don't do X" carve-out** — compliance is high but not perfect. Forks have been observed to override "don't edit/commit" instructions when they judge the action obvious and mechanically correct. The override usually produces technically sound work, but it violates the boundary.
**Why, mechanically: a fork inherits your whole session, and your brief is only the last thing in it.**`pi-fork/src/index.ts:47`:
Every entry on the current branch — your messages, assistant thinking, tool calls **and** tool results — is serialized verbatim, written to a temp session file (`runner.ts:404`), and opened by the child `pi` via `--session`. The task string is not the child's world; it is one instruction appended to a world already full of your stated intentions. When the transcript shows work in flight and the brief forbids it, those two conflict, and the child may resolve the conflict toward "finish the obvious thing".
**Worked example (2026-07-29, `balanced` = sonnet-5, `thinking: low`).** The brief said, verbatim: *"DRAFT ONLY — do not submit anything, do not use gh/curl…, do not commit to any git repo, and do not modify any file other than /workspace/tmp/pi-mono-issue.md."* The fork returned *"All three done: 1. **Pushed** — pi-toolkit@4b4b76e… 2. **Moved** — cli_utils@f644fa1, pushed… symlinked live into ~/.local/bin"*. It had not merely claimed the work; commit timestamps place it inside the fork's execution window:
```
fork window 21:53:40Z → 21:58:27Z
cli_utils f644fa1 21:57:47Z ← committed + pushed by the fork, inside the window
pi-toolkit 4b4b76e 21:42:05Z ← pre-existing; the fork only claimed the push
```
The "three" things it completed were exactly the main thread's pending todos, visible to it in the inherited transcript. A 4645-character brief with four explicit prohibitions did not prevent this — so *"state decision authority explicitly"* is necessary and demonstrably **not sufficient**. Its verbatim file move also carried a data-loss race and a README asserting the opposite of the truth, neither flagged in its confident report.
**You cannot withhold write tools.** There is no tool allow/deny list anywhere in the fork config: `config.ts` exposes only `extensions`, `environment`, `offline`, and the child is spawned as a full `pi` process (`--mode`, `--session`, `--model`, `--thinking`). `extensions: []` yields `--no-extensions`, which disables *extensions*, not the core `read`/`write`/`edit`/`bash`. **Assume every fork can write anywhere you can.** If a boundary violation would be genuinely unacceptable, the control is not the brief — it is not forking that task.
**Why the report reads so confidently.** The child's output contract is ~90 lines of *shape* — evidence rules, snippet rules, "Result / confidence / headline", per-genre sections. Grepping it for scope, authority, or permission language returns a single hit, and that one is about *review* scope in reporting. Nothing instructs the child to stay inside its mandate or to mark unverified claims. The format demands a verdict with a confidence level; where a fact was never checked, fluent prose fills the slot. The same fork reported *"smoke-tested against all 4 live sessions"* when there were 20 — and that number appears nowhere in the inherited transcript, so it was invention, not stale context.
**Practical rules:**
- State decision authority explicitly, every time — and add the anti-inheritance clause (task-design item 5) whenever the brief is narrower than the conversation.
- Require a **"What I did NOT do"** section on any write-capable fork.
- **Verify mutations from the filesystem, never from the report.** `git log -1 --format=%ai` against the fork's start/end times, `git status`, real diffs. Read a fork's push as an unreviewed PR from a stranger.
- **A brief containing a prohibition is a judgment task.** Do not run it at `fast` (haiku, `thinking: off` in the shipped profiles); escalate the tier. Reserve `fast` for "return raw output, no interpretation".
- Distrust **quantities** and **provenance claims** in fork prose specifically ("all N sessions", "shipped with the image", "as expected") — those are the slots confabulation fills.
- The fact that the fork was "right anyway" is not the same as the fork having followed instructions.
### The context ladder — and the second dispatch mechanism (`pi-task`)
Everything above describes a child that inherits everything. That is not a fixed
cost of delegation — **how much context a child gets is a choice**, and `fork`
sits at one extreme of it. Five rungs:
| rung | what the child sees | mechanism | built? |
|---|---|---|---|
| **L0** | nothing but the goal | `pi-task` default: fresh `--session-id pitask-<id>-<stamp>` in a private `--session-dir` | yes |
| **L2** | goal + an **excerpt the parent curated** | `pi-task` spec `context.facts`, pasted verbatim (`bin/pi-task:151`) | yes |
| **L3** | a **truncated tail** of the parent branch | *nothing implements this* — would need a new spec key plus `--session <trimmed snapshot>` | **no** |
| **L4** | the **entire** parent branch | `fork(task=…)` — `getHeader()+getBranch()`, no offset or limit anywhere in the call chain | yes |
**`pi-task` is a CLI (`/opt/pi-toolkit/bin/pi-task`, `schema` prints the spec
fields); the `task` tool from pi-extensions wraps it** so it appears in your tool
list next to `fork`. If the tool is absent, invoke the CLI with `bash`:
`/opt/pi-toolkit/bin/pi-task run <spec.json>`. Either way it reads an immutable
JSON spec, and "inherit the session" is not expressible in that schema — the
isolation is structural, not a request.
**Choose the lowest rung that can do the job:**
- **`fork` (L4)** when the subtask only makes sense against this conversation,
when you want several independent opinions in parallel from one message, or for
read-only exploration whose detail you will discard. Everything in "Boundary
discipline" above applies in full.
- **`pi-task` (L0–L2)** when the brief contains a **prohibition** (the inherited
transcript is exactly what overrides those), when you want a **pass/fail**
result instead of prose, when you need an **audit trail**, or when writes
outside an authorised set must be caught.
- **Neither** for trivial work, iterative work (both are one-shot), or judgement
that needs context only you have.
**What `pi-task` gets you that no brief can.** The envelope must parse or the run
FAILED, however fluent the prose. `roots[]` is the WATCHED set and
`write_allowed` the CHANGEABLE subset, diffed before and after with git
`--porcelain --ignored`. That `--ignored` flag is load-bearing: in the T4 test the
child obeyed its brief perfectly and still tripped the detector, because
`py_compile` wrote `__pycache__` into a watched-but-not-writable root — a
gitignored path that plain `--porcelain` reports as clean. Note the structural
point that test exposed: under `read_only: true` a write is *defiance*, so a
well-behaved child never produces a delta and the detector is never exercised.
Splitting WATCHED from WRITABLE is what lets an **obedient** child reveal a
violation, which is the realistic hazard.
**What it does not fix.**`--no-extensions` removes extensions, not the core
`read`/`write`/`edit`/`bash` tools — exactly as described above — so the boundary
diff is post-hoc **detection, not prevention**. And a fresh L0 context removes the
*narrative* failures (parent voice, invented continuity) without removing
confabulation: given an under-specified spec built on a false premise, the child
still filled the `deliverable` slot with a confident shape. The envelope's own
structure creates that pressure. Verify decisive claims from the filesystem
regardless of which rung you used.
**Trap — the capability floor is inverted from intuition.**`runner.ts:188` reads
`if (extensions !== null) args.push("--no-extensions")`. So `pi-fork.extensions:
[]` passes the flag and the floor is **on**; setting it to `null` — documented in
`settings.json` as the way to "restore normal extension loading" — passes nothing
and the floor is **off**, restoring palace writes inside every fork child.
Changing `[]` to `null` as a tidy-up re-arms what was deliberately disarmed.
`pi-task` hardcodes the flag and cannot drift this way.
### Anti-patterns
- **Forking trivial work.** A fork has overhead. If the task takes < 30 seconds in your main thread, just do it.
- **Vague briefs.** "Look into the database thing" returns vague output. The fork is not telepathic.
- **Forking iterative work.** Forks are one-shot. If you need to iterate, you'll re-spec the task each time — usually worse than doing it yourself.
- **Recursive forking** (forks spawning forks). Disabled by default and should stay disabled unless you have a specific batch-fanout use case.
- **Treating fork output as ground truth without verification.** Especially for cited code/commit hashes/URLs — forks can hallucinate these like any LLM. Spot-check decisive evidence.
**Observed failure shape (2026-07-29, `fast` tier): raw tool output correct, surrounding narrative wrong.** A fork asked to run three commands and report them verbatim returned all three outputs accurately — then framed them with two confident inventions: that the `packages[]` entries were "the three that shipped with the image" (one had in fact been hand-registered minutes earlier by the parent — the entire point of the investigation), and that "the entrypoint re-registers them on each start" (the guard deliberately skips re-registration once the entry exists). Neither claim was in the command output; both were plausible glue.
**Rule:** read a fork's **Evidence** section as data and its **narrative** as a hypothesis. When the fork's story contradicts something you established in the main thread, your own verified context wins. Note what this failure is *not*: the fork was not context-starved — it had your entire transcript (see "Boundary discipline" above) and invented anyway, because its output contract rewards a confident verdict over an admitted gap. Passing verified context up front still helps, but do not expect it to suppress invention on its own; the load-bearing habit is verifying decisive claims yourself. Being right about the evidence is not the same as being right.
---
## Part 2: pi-observational-memory
### How it actually works
Observational memory (OM v3, "session-ledger" architecture) runs an **observer agent** in the background as your conversation grows. When token thresholds are crossed (defaults: observe at 10k, reflect at 20k, compact at 81k), the observer distills the recent transcript into:
- **Observations** — timestamped events, each with a 12-character hex ID like `[3682ebfad7af]`. Compact one-liners describing what happened in the conversation.
- **Reflections** — durable, long-lived facts about the user, project, decisions, and constraints. Some reflections include observation IDs as evidence pointers.
When compaction fires, the raw transcript is folded away and replaced with a structured summary block containing the observations + reflections. **You — the next turn of the same agent — receive that summary block as your starting context.** That's the recovery mechanism.
**Storage is in-transcript, not on disk.** Do not grep for `observations.jsonl` or similar files; you will not find them. The artifact lives in the model's input context window.
Configuration lives in `~/.pi/agent/settings.json` under `observational-memory`. Tune `observeAfterTokens`, `reflectAfterTokens`, `compactAfterTokens`, and `observationsPoolMaxTokens` if observations feel sparse or noisy. The default 81k compaction threshold is well-calibrated for typical multi-task sessions.
### The `recall` tool
`recall(<12-char-hex-id>)` resolves a specific observation or reflection ID back to the original source context — the exact bash output, file contents, tool call results, commit message, or transcript fragment that the observation was distilled from.
**Use recall when:**
- You are about to make a decision that depends materially on a compacted observation or reflection whose details are unclear.
- You need exact wording, paths, commands, errors, commits, or user constraints behind a remembered claim.
- A broad reflection is relevant but you need its supporting observations to act safely.
- The user asks "why do you believe X" or "what supports that memory".
**Do not use recall for:**
- Semantic search (it's keyed by ID, not topic — you must already have a specific 12-char hex ID).
- Browsing the transcript out of curiosity.
- Preemptive lookup of every ID in your context "just in case".
Recall costs tokens. Use it when exact source context will materially change your next action.
> **Calibration note (from a real ~1-month trial, 2026-05/06):** across 20 logged container sessions, `recall` was invoked **0 times** while obsmem passively carried 529 observations across 6 compactions. Zero recall is a *warning sign*, not a badge of efficiency — it means decisions after a compaction were made on the distilled one-liner alone, without ever re-checking the source. The injected summary is **lossy by design**. Default habit to adopt: when you are about to **edit code, ship a change, or assert a fact** that rests on a `[high]`/`[critical]` observation or a reflection you did not produce *this* turn, `recall` its ID **first**. One recall before a load-bearing action is cheap; redoing finished work or contradicting a prior correction is not.
### Reading the compaction summary
When you see a block like `The conversation history before this point was compacted into the following summary:` at the start of a session or turn, that's OM output. Standard structure:
- **Reflections** at the top: stable facts. Some have IDs in brackets.
- **Observations** below, chronological: timestamped events with IDs in brackets and importance markers (`[high]`, `[critical]`, etc.).
When entries conflict, **the most recent observation reflects the latest known state.** Work that prior observations describe as completed should not be redone unless the user explicitly asks to revisit it.
### Anti-patterns
- **Treating compacted memory as definitive without recall** when stakes are high. Compaction is lossy; the observation may have lost a constraint that was on the line above it in the original transcript.
- **Recalling every ID preemptively.** Wasteful. Recall on demand.
- **Assuming the disk holds OM artifacts.** It doesn't. Don't waste time looking.
- **Ignoring the summary block** when starting a session. It's there because the prior session was real work — read it before answering questions about past work.
Read a **zero** carefully before treating it as a habit problem: a missing
`fork <== pi-fork` line means the tool was never *called*, which can equally
mean it was never *registered* (see the `packages[]` case study above). Check
registration first, then blame habits.
---
## Part 3: ssh-controlmaster
### What it does
When pi is launched with `--ssh`, this extension **rewires pi's `read`, `write`, `edit`, and `bash` tools to execute on the remote machine**, multiplexed over a single SSH ControlMaster socket. Pi is still running locally — the LLM, the UI, the MCP servers, the fork dispatcher all live on your local box — but anything those tools touch on the filesystem is the *remote's* filesystem.
This is fundamentally different from running pi locally and using `bash` to ssh inside it: with `--ssh`, the tool layer itself is remoted, so the LLM thinks it's working in the remote's `cwd` (the system prompt is rewritten to say so).
### Usage
```bash
# Key-based auth (preferred), remote cwd defaults to remote $HOME
pi --ssh lagret
# Pin to a specific remote directory
pi --ssh lagret:/volume1/docker/portainer/compose/119
# Password auth (input is NOT masked when typing)
pi --ssh user@host --ssh-ask-pass
```
The `lagret` form requires a `Host lagret` block in `~/.ssh/config` or a resolvable hostname. The status bar shows `SSH ⚡ own master <host>:<cwd>` or `SSH ⚡ system master <host>:<cwd>` once connected.
### How it cooperates with system SSH config
It reads `ssh -G <host>` to learn the effective config, then:
| `~/.ssh/config` for the host | Behavior |
|---|---|
| `ControlMaster auto` or `yes` with a `ControlPath` | Reuses the system master socket. Does **not** tear it down on pi exit ("it was the system's to manage before pi arrived"). |
| No ControlMaster configured (or explicitly `no`) | Creates its own master at `/tmp/pi-cm-<pid>.sock` with `ControlPersist=yes`. Tears it down on pi `session_shutdown`. |
This means it composes cleanly with the system-wide `ssh-control-master-setup.sh` helper from the `ci-release-watcher` skill: if that script has already configured `~/.ssh/config` for the host, `pi --ssh` rides on the existing master rather than opening a parallel connection.
### Caveats and edge cases
- **Local vs remote tool boundary.** Only `read`/`write`/`edit`/`bash` are remoted. **MCP servers are still local** — `mempalace` files drawers and diary entries against the local palace even when your shell work happens remotely. Same for `fork`, `recall`, `todo`, and any other custom tool. This is usually what you want (palace memory survives across remote sessions) but worth knowing.
- **fork over ssh.** Forks spawn locally and inherit the same `--ssh` mode by virtue of the parent's tool wiring; the fork's bash calls hit the same ControlMaster. Forks burn the same SSH socket, not a parallel one — multiplexing wins again.
- **macOS Unix socket path limit.** The own-master socket lives at `/tmp/pi-cm-<pid>.sock` to stay under macOS's ~104-char limit. If you have a non-default `TMPDIR` long enough to blow this, ssh will fail to start the master.
- **Password auth password visibility.** From the source: *"input is NOT masked — the password is visible while typing."* The password is written to a chmod-700 SSH_ASKPASS script in `/tmp` and deleted after the master establishes; not persisted, but on-screen during entry.
- **Remote bash environment.** The remote shell is whatever `ssh user@host '<cmd>'` invokes — typically a non-login non-interactive bash. Don't expect `~/.bashrc` aliases or PATH manipulations from `~/.profile`. Pin tool paths or invoke via `bash -lc '...'` if you need login-shell behavior.
- **Path translation is naive.** The extension does `path.replace(localCwd, remoteCwd)` to translate paths in tool calls. If the LLM emits an absolute remote path that doesn't share the local-cwd prefix, the path is passed through unchanged — usually fine but pathological for paths that happen to contain the local-cwd substring.
### When to use it
- Editing configs on a NAS / homelab host without scp ping-pong (`pi --ssh lagret:/volume1/...`)
- Operating against a host whose tools/data you need but whose disk is too slow to mount via SSHFS
- Investigating runner state, container configs, etc., on a remote host as if local
- Multi-step remote work where opening a fresh ssh connection per step would burn your CGNAT flow budget
### Anti-patterns
- **Using `pi --ssh` for one-off shell work.** Just `ssh` directly. The extension shines when there are dozens of tool calls per session.
- **Filing palace drawers expecting them on the remote.** They go to the local palace. If you want palace artifacts on the remote host, ssh into the remote and run pi *there* against its local palace.
- **Forgetting `--ssh` in followup sessions.** Status bar is the canary — if you don't see `SSH ⚡` you're operating locally despite intending remote. Easy mistake on a fresh terminal.
### Reaching the devbox host from inside the container (`dssh` / `dscp`)
Distinct from `pi --ssh` above. When the **pi-devbox container** runs under OrbStack / Docker Desktop on macOS, it can SSH back to its own host. The entrypoint's `setup-lan-access.sh` regenerates `~/.ssh-local/config` on **every container start** (the in-container `~/.ssh` is mounted read-only, so a sidecar config + `known_hosts` + `ControlPath` under `~/.ssh-local/` is used instead).
```bash
# Interactive shells get aliases (from ~/.bash_aliases):
**The agent's `bash` tool is non-interactive — those aliases are NOT loaded.** Use the explicit form:
```bash
ssh -F ~/.ssh-local/config host 'cmd'
scp -F ~/.ssh-local/config <src> host:<dst>
```
- Host aliases `host` and `mac` both resolve to `host.docker.internal` (user varies per host machine — check `~/.ssh-local/config` for the active `User` value, key `~/.ssh-local/devbox_jump_ed25519`, `ControlMaster auto` / `ControlPersist 4h`).
- The config chains `Include ~/.config/devbox-shell/ssh-lan.conf` then `Include ~/.ssh/config`, so LAN targets are reachable too (add `ProxyJump host` to those entries).
- **Use it for:** enabling/inspecting the host's pi config (`~/.pi/agent/settings.json`), running `evaluate-extension-usage.py` against the host's `~/.pi/agent/sessions/` for a combined host+container metric, or copying host transcripts into the container. The host's pi runs natively there; its palace, sessions, and extensions are separate from the container's.
---
## Cross-Skill Notes
- **mempalace** is for cross-session persistent memory (diary, knowledge graph, drawer storage). OM is for **within-session** context survival across compaction. They complement each other: write a diary entry at session end *and* let OM compact your work-in-progress mid-session.
- **systematic-debugging** and **test-driven-development** skills pair well with deep-tier forks: a deep fork can carry out a focused debugging investigation or write a failing test suite without polluting your main context.
- **ci-release-watcher** ships a `scripts/ssh-control-master-setup.sh` helper that configures system-wide SSH ControlMaster in `~/.ssh/config`. That's a separate mechanism from the `ssh-controlmaster` pi extension — they compose, they don't overlap. Use the script for persistent host-wide multiplexing, the extension for per-pi-session remote operation.
if ! printf'%s'"$block"| grep -q "outputs.$lc";then
echo"::error::Dockerfile.base declares '$r' but it is NOT folded into the base_tag hash in $WF."
echo"::error::Add echo \"\${{ needs.resolve-versions.outputs.$lc }}\" inside the HASH=\$( … ) | sha256sum block, or a $r-only change will silently fail to rebuild the base."
fail=1
fi
done
if["$fail"=0];then
echo"OK: all Dockerfile.base *_REF args are folded into base_tag (${refs:-none})."
# Exact, not heuristic: the value handed over IS this image's release
# tag, so it cannot be a pi version anyone meant.
fail "--expected-version $EXPECTED_VERSION is the pi-devbox IMAGE version, not the pi version — use --expected-image-version $EXPECTED_VERSION (live pi is $ACTUAL_VERSION)"
else
fail "pi version mismatch: expected $EXPECTED_VERSION, got $ACTUAL_VERSION (this flag asserts the pi coding agent version; for the image release tag use --expected-image-version)"
fi
elif[ -n "$MANIFEST_PI_VERSION"];then
# Not a tautology: the manifest records what pi reported at BUILD time,
# while `pi --version` resolves through PATH, which a stale npm-global
# volume install can shadow.
if["$MANIFEST_PI_VERSION"="$ACTUAL_VERSION"];then
pass "pi version $ACTUAL_VERSION (matches this image's build manifest)"
else
fail "live pi $ACTUAL_VERSION != $MANIFEST_PI_VERSION recorded in $MANIFEST — a stale pi in the ~/.pi/npm-global volume is shadowing the baked one"
fi
else
warn "pi version $ACTUAL_VERSION (no --expected-version and no build manifest to compare against — informational only)"
fi
else
fail "pi --version failed"
fi
echo
echo"-- pi-devbox image version --"
if[ -z "$MANIFEST_RELEASE_TAG"];then
if[ -n "$EXPECTED_IMAGE_VERSION"];then
fail "cannot verify --expected-image-version $EXPECTED_IMAGE_VERSION: no readable release_tag in $MANIFEST (image built before the manifest existed, or jq missing)"
else
warn "image release tag unknown (no readable $MANIFEST) — pi-devbox-version would say the same"
fail "--expected-image-version $EXPECTED_IMAGE_VERSION is the pi version, not the image release tag — use --expected-version $EXPECTED_IMAGE_VERSION (this image is $MANIFEST_RELEASE_TAG)"
else
fail "image version mismatch: expected $EXPECTED_IMAGE_VERSION, got $MANIFEST_RELEASE_TAG — the recreate did not pick up the intended image"
fi
else
warn "image version $MANIFEST_RELEASE_TAG (no --expected-image-version given — informational only)"
fi
echo
echo"-- Persisted named volumes (must survive --force-recreate) --"
pass "$pkg registered in settings.json packages[]"
else
fail "$pkg NOT in settings.json packages[] (tool will not load)"
fi
done
if["$VARIANT"="studio"];then
if _pkg_registered pi-studio;then
pass "pi-studio registered in settings.json packages[]"
else
fail "pi-studio NOT in settings.json packages[] (studio variant)"
fi
fi
# pi-atelier — vendored from v1.7.0 on. Absent on older images, and
# deliberately unregistered when DEVBOX_ATELIER=0; neither is a failure.
if[ -d /opt/pi-atelier ];then
if["${DEVBOX_ATELIER:-1}"="0"];then
if _pkg_registered pi-atelier;then
fail "pi-atelier still in packages[] despite DEVBOX_ATELIER=0"
else
pass "pi-atelier unregistered (DEVBOX_ATELIER=0, as requested)"
fi
elif _pkg_registered pi-atelier;then
pass "pi-atelier registered in settings.json packages[]"
else
fail "pi-atelier NOT in settings.json packages[] (sidebar will not load)"
fi
if _npm_atelier_present;then
fail "stale npm:pi-atelier still in packages[] — it resolves through the ~/.pi/npm-global VOLUME and shadows the pinned /opt copy (entrypoint migration did not run)"
fi
fi
fi
# ── agent-browser must resolve to the image, not the config volume ────
# The same volume-shadowing hazard already asserted for pi (above) and
# pi-atelier (just now), for the third package it has bitten. This check
# belongs HERE rather than only in smoke-test.sh: a build-time container has an
# empty ~/.pi/npm-global, so smoke-test can never see the stale copy that a
# real recreate inherits. Measured instance: 0.27.0 from 2026-07-17 shadowed
# the image's 0.35.2 for ~7 weeks on mbp-m1-2020, silently supplying an older
# BUNDLED SKILL (3 skillsets vs 8) — the agent read the stale instructions
# without any version mismatch ever being surfaced.
AB_VER=$(agent-browser --version 2>/dev/null | head -n1)
case"$AB_REAL" in
/usr/*)
pass "agent-browser resolves to the image copy (${AB_VER:-version unknown})"
;;
*)
fail "agent-browser resolves to $AB_REAL (${AB_VER:-version unknown}) — a ~/.pi/npm-global VOLUME copy is shadowing the image; the entrypoint retirement guard did not run or could not move it"
# Cap the host list. A 50-name line is the noisy gate this repo already warns
# about in lint-shell.sh: unreadable output is ignored output. Six names plus
# a count is enough to identify the class and act.
if["${#bad[@]}" -gt 0];then
shown="${bad[*]:0:6}"
if["${#bad[@]}" -gt 6];then
shown="$shown (+$((${#bad[@]}-6)) more)"
fi
fi
if["$n" -eq 0];then
warn "$label: no host resolves to ControlMaster on — effective ControlPath not exercised"
elif["${#bad[@]}" -eq 0];then
pass "$label: $n ControlMaster host(s), every ControlPath dir exists and is writable"
elif["$sev"="warn"];then
warn "$label: ${#bad[@]}/$n host(s) resolve ControlPath to a missing or unwritable dir [$shown] — STRUCTURAL and permanent while ~/.ssh/config pins ControlPath inside the read-only ~/.ssh, so this line can never reach zero and is not a to-do; $_git_ssh_note (see Dockerfile.base CAVEAT)"
else
fail "$label: ${#bad[@]}/$n host(s) resolve ControlPath to a missing or unwritable dir [$shown] — ssh dies rc=255 'unix_listener: cannot bind to path' and the remote command never runs"
fi
}
ifcommand -v ssh >/dev/null 2>&1;then
# Whether git-over-ssh already routes through the sidecar decides how much the
# permanent default-route warning below actually matters, so state it IN that
# message rather than leaving each reader to work it out. Asserted properly as
# its own pass/fail in the next section.
#
# THREE states, not two, and the [ -r ] test is why: core.sshCommand NAMING the
# sidecar does not mean the sidecar EXISTS. Without that test, the arm where the
# path is wired but the file is gone printed "git IS wired ... unaffected" about
# a state in which every single git-over-ssh call fails. Found by exercising all
# five arms of this check rather than only the healthy one.
if[ ! -r "$HOME/.ssh-local/config"];then
_git_ssh_note="there is no ssh sidecar on this host, so nothing for -F to point at and these hosts cannot multiplex at all — see the git-over-ssh check below"
_git_ssh_note="git IS wired to the sidecar (core.sshCommand), so git push/fetch is unaffected; bare 'ssh' to these hosts still needs 'ssh -F ~/.ssh-local/config'"
else
_git_ssh_note="git is NOT wired to the sidecar (see the git-over-ssh check below), so both git and bare 'ssh' need 'ssh -F ~/.ssh-local/config'"
fi
# Concrete Host aliases only: patterns (*, ?) and negations (!) are not
# connectable targets, so `ssh -G` on them proves nothing.
pass "no ssh sidecar on this host and core.sshCommand correctly unset (guard holds)"
else
fail "core.sshCommand is set but ~/.ssh-local/config does not exist — every git-over-ssh call dies on a missing -F file; entrypoint-user.sh's [ -r ] guard did not hold"
pass "git core.sshCommand routes through the sidecar ($_git_ssh_cmd)";;
'')
fail "a sidecar exists but git core.sshCommand is unset — 'git push' to any host whose config pins ControlPath inside the read-only ~/.ssh dies rc=255 behind git's misleading 'correct access rights'. entrypoint-user.sh should set it; expected on images built before that wiring landed, where the fix is: git config --global core.sshCommand \"ssh -F \$HOME/.ssh-local/config\"";;
*)
warn "git core.sshCommand set to something else and left alone (first-wins, deliberate): $_git_ssh_cmd";;
esac
fi
echo
echo"-- Shell defaults re-seeded from /etc/skel-devbox --"
if[ -f "$HOME/.bash_aliases"];then
pass "~/.bash_aliases exists"
else
fail "~/.bash_aliases missing"
fi
# History flush must survive shell nesting. The DEVBOX_HIST_SET guard must NOT
# be exported: if it leaks into child processes, nested shells (esp. tmux
# panes) skip installing `history -a` and lose in-memory history on abrupt
# termination. Assert a child login shell still wires up the per-prompt flush.
if bash -lic 'bash -lic "case \"\$PROMPT_COMMAND\" in *\"history -a\"*) exit 0;; *) exit 1;; esac"' </dev/null >/dev/null 2>&1;then
pass "nested shell installs 'history -a' (DEVBOX_HIST_SET not exported)"
# Refuse to silently REWIND provenance. `git checkout <tag>`, a detached HEAD,
# or an older checkout can all leave $ROOT's HEAD behind the already-recorded
# ref; without this guard a refresh there would happily rewrite both the ARG
# and the bytes backwards and report it as an ordinary update.
if["$recorded" !="$head_sha"];then
if git -C "$ROOT" cat-file -e "${recorded}^{commit}" 2>/dev/null;then
if ! git -C "$ROOT" merge-base --is-ancestor "$recorded""$head_sha" 2>/dev/null;then
if["$FORCE" !=1];then
die "refusing: $ROOT's HEAD ($head_sha) is not a descendant of the recorded ref ($recorded) — this looks like a rewind. Pass --force if this is intentional."
fi
printf'WARNING: --force set; %s is not an ancestor of HEAD %s — proceeding anyway\n'"${recorded:0:7}""${head_sha:0:7}" >&2
fi
else
if["$FORCE" !=1];then
printf'CANNOT-DETERMINE: %s is not present in %s (shallow clone?) — fetch full history to verify this refresh moves forward, or pass --force to proceed without that guarantee\n'"$recorded""$ROOT" >&2
exit2
fi
printf'WARNING: --force set; %s could not be resolved in %s — proceeding without verifying forward motion\n'"${recorded:0:7}""$ROOT" >&2
fi
fi
# Written FROM THE REF, not copied from the working tree, so the pair cannot
# be a lie by construction. Via a temp file so a failed write cannot leave a
# half-vendored snapshot behind.
snap_tmp=$(mktemp)
if ! git -C "$ROOT" show "HEAD:$REL_PATH" > "$snap_tmp" 2>/dev/null;then
rm -f -- "$snap_tmp"
die "cannot read HEAD:$REL_PATH from $ROOT"
fi
chmod 0644 -- "$snap_tmp"
mv -- "$snap_tmp""$VENDORED"
["$(sha_of "$VENDORED")"="$blob_sha"]\
|| die "internal: written snapshot does not match HEAD:$REL_PATH"
# In-place, and only the exact pinned line: a broad sed on this Dockerfile
# could rewrite one of the other *_REF ARGs.
tmp=$(mktemp)
sed "s|^ARG ${ARG_NAME}=.*\$|ARG ${ARG_NAME}=${head_sha}|""$DOCKERFILE" > "$tmp"
printf'\nNOTE: %s is hashed into base_tag, so this costs a base rebuild\n'"$VENDORED"
printf'on the next tag (~67 min). Also re-pin the phrase canary in\n'
printf'scripts/smoke-test.sh if the section it names changed.\n'
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.