162 Commits

Author SHA1 Message Date
Joakim Persson 12f99c49e3 docs(v1.10.0): adopt pi's fullscreen default and document the way back
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 15s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 14s
Publish Docker Image / lint-gate (push) Successful in 21s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 55m32s
Publish Docker Image / smoke (push) Failing after 8m50s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Publish Docker Image / smoke-studio (push) Failing after 9m55s
Publish Docker Image / build-variant-studio (push) Has been skipped
Decision: keep pi 1.0.0's default (tuiMode "fullscreen"). No tuiMode is baked,
so the image follows upstream rather than pinning the fleet to either mode -
but "we inherited a changed default" is only acceptable if the revert is
written down, so it now is.

README gains a "Terminal UI mode" section ahead of the pi-atelier section,
because the two interact. It gives all three scopes, taken from pi 1.0.0's own
docs/settings.md and docs/cli.md rather than from the changelog prose:

  - permanent: "tuiMode": "regular" in ~/.pi/agent/settings.json
  - one session: pi --tui-mode regular   (the flag is real: cli.md:221)
  - one project: the same key in .pi/settings.json, which overrides the agent
    directory

It also documents the four related settings (fullscreenExitOutput,
fullscreenScrollbar, fullscreenCopyOnSelect, fullscreenWheelScrollLines) and
two things specific to this image:

  - tmux: fullscreen uses the alternate screen, so tmux copy-mode shows the
    pane history AROUND pi, not pi's transcript. That is the concrete reason
    someone here would want "regular" back.
  - fullscreenWheelScrollLines "auto" behaves differently over SSH (caps fast
    wheel spins at 6 lines/event) than in a local macOS terminal (1 line) -
    worth naming because this container is normally driven over SSH.
  - pi-atelier works in BOTH modes (upstream has handled regular vs fullscreen
    renderers separately since pi 0.84), so reverting costs nothing. Stated so
    nobody assumes the sidebar is the price of the old scrollback.

The python3 merge snippet in that section was RUN before being documented:
against a realistic settings.json under an overridden HOME, twice, confirming
it is idempotent and preserves sibling keys including nested objects and the
_comment fields the seeded file uses. Documented code that has never been
executed is a guess.

DOCKER_HUB.md gets a short version of the same note, because CI reads it from
the TAG and POSTs it to Docker Hub as full_description - a behaviour change
this visible should not require reading the repo to undo.

Gates: doc-drift 23 OK / 0 DRIFT / 0 SKIP / 0 FAIL; base-hash, workflow-shell,
skill-floor, lint-shell all rc=0.
2026-10-02 09:00:51 +02:00
Joakim Persson e026f6f65e release(v1.10.0): pi 1.0.0 and pi-atelier v0.13.0, on measured evidence
Lint / hadolint (push) Successful in 10s
Lint / skill-floor (push) Successful in 12s
Lint / actionlint (push) Successful in 19s
Lint / doc-drift (push) Successful in 19s
Renames the unpublished v1.9.5 section to v1.10.0 and adopts the two bumps
that section had deliberately deferred. v1.9.5 was never tagged or published,
so nothing shipped under that name.

A minor, not a patch. v1.9.5 was numbered under the "patch - pi version bumps"
rule, but this release carries a pi MAJOR, a three-minor pi-atelier jump, four
new base packages and a new mailbox feature. The release carrying a 1.0.0
should not be the one numbered as a patch.

pi 0.85.1 -> 1.0.0. v1.9.5 held at 0.87.1 because obsmem's only compatibility
work names Pi 0.87 and no upstream issue mentions 0.99 or 1.0. That reasoning
had a hole worth naming: "no issue mentions 1.0" is an ABSENCE OF A STATEMENT,
not a measurement. So 1.0.0 was measured, against the published npm tarballs
for 0.87.1 and 1.0.0 unpacked side by side:

  - 1.0.0 has NO "### Breaking Changes" section at all. Those belong to 0.87.0,
    0.86.0, 0.84.3, 0.84.0, 0.83.0, 0.80.8, 0.80.7, 0.75.0 - none to anything
    between 0.88 and 1.0.0. The major is a milestone (fullscreen default,
    leaner codemode), not an API break.
  - finishTurn is still in 1.0.0's dist, so the pinned obsmem SHA keeps the API
    it migrated to.
  - All 8 pi.* APIs our extensions call exist in 1.0.0's dist: registerTool,
    registerCommand, registerFlag, getFlag, on, exec, sendMessage,
    sendUserMessage.
  - engines.node >=22.19.0 on both; the image ships 24.x.

The 0.99.2 change that looked fatal and is not: from 0.99.2 the DEFAULT MCP
exposure is `codemode`, so such tools are "neither declared to the model nor
listed" and must be found with searchTools(). That would gut the MemPalace
protocol if MemPalace were a builtin-MCP server. It is not - mempalace.ts and
mcp-loader.ts each run their own MCP client and register tools via
pi.registerTool(), which is why they are named mempalace_search and not
mcp__mempalace__search. Second route: settings.json has neither an mcpServers
block nor an mcp block. Filed upstream under "Changed", not "Breaking", so it
would have been easy to meet the hard way.

pi-atelier v0.10.3 (ed3837b) -> v0.13.0 (34d26f1), closing the gap v1.9.5
flagged as the pi bump's residual risk. peerDependencies are unchanged at pi
>=0.84.0 across v0.10.3/v0.12.1/v0.13.0 - a FLOOR, so not evidence of
anything. What was checked instead: v0.13.0 carries its OWN breaking change
(the entry point no longer exports the internal registry and layout helpers),
which reaches nothing of ours - grepping pi-atelier and registerSidebarPanel
across mempalace-toolkit, pi-devbox, pi-extensions, pi-fork,
pi-observational-memory and pi-studio returns ZERO matches in all six. v0.12.0
removed the showSessionActions setting (0 references here) and moved the
Control Center shortcut Alt+A -> F6 (0 references here, so no doc goes stale).
showSidebarAgent/showSidebarTodos, the only atelier keys README names, are
both still in v0.13.0's README. TuiMainScreen and renderLayoutFrame both still
exist in pi 1.0.0's dist.

PI_ATELIER_VERSION is bumped with PI_ATELIER_REF; it is a separate ARG and the
image label would otherwise have lied.

v0.13.0 is an ANNOTATED tag: `ls-remote --tags` reports the tag object
(dd06971), not the commit (34d26f1). check-doc-drift.sh resolves the commit, so
that is what the CHANGELOG names. The gate caught the wrong SHA on the first
attempt - noted in the CHANGELOG for future bumps.

NOT PROVEN, and stated so acceptance does not mistake it for cleared: grepping
dist shows the SYMBOLS survive, not that their SIGNATURES are unchanged -
necessary, not sufficient. 1.0.0 is four days old and neither obsmem nor
atelier has a commit naming it. Acceptance must prove the obsmem workers CAP
TURNS (peerDeps are *, so a mismatch is silent) and that the atelier sidebar
PAINTS.

User-visible behaviour change, deliberately NOT overridden: 1.0.0 makes the TUI
fullscreen by default, replacing the terminal's normal scrollback.
tuiMode: "regular" restores the old behaviour. No baked default is set, so the
image inherits upstream's choice rather than silently pinning the fleet.

Gates: check-doc-drift 23 OK / 0 DRIFT / 0 SKIP / 0 FAIL; base-hash,
workflow-shell, skill-floor, lint-shell all rc=0.
2026-10-02 08:16:37 +02:00
Joakim Persson 9d0b3dec0b release(v1.9.5): pi 0.87.1 with pi-obsmem pinned to the merged finishTurn fix
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 19s
Lint / doc-drift (push) Successful in 12s
v1.9.4 held pi at 0.85.1 because 0.87.0 REMOVED `shouldStopAfterTurn`, which
pi-observational-memory 3.1.4 still used in all three workers. Its peerDeps are
`*`, so nothing refuses at install time and the breakage is silent at runtime --
turn caps ignored, workers losing their specialised prompts. PR #83 fixes that
and merged 2026-09-23.

Re-measured 2026-10-01 across src/agents/{observer,reflector,dropper}/agent.ts,
deliberately reusing the v1.9.4 audit's counting method so the numbers compare:

  e7d77dc (3.1.4, baked in v1.9.4): shouldStopAfterTurn 3, finishTurn 0, systemPrompt 3
  731c3d4 (pinned here)           : shouldStopAfterTurn 0, finishTurn 3, systemPrompt 0

All three workers migrated, and the 0.86.0 AgentContext.systemPrompt reads are
gone too. The coupling is asymmetric and that is why these move in ONE commit:
3.1.4 + 0.87.1 silently ignores turn caps, and 731c3d4 + 0.85.1 breaks the
workers outright, because finishTurn does not exist before 0.87.0.

Pinned to a SHA rather than waiting for a tag, departing from the v1.9.4
instruction to wait for a release: #83 is merged but the newest obsmem tag is
still 3.1.4, cut 2026-09-20, BEFORE the merge. Upstream tags slowly and moves
master often (e7d77dc -> 1529e14 -> 731c3d4 in nine days), so waiting means
holding pi indefinitely. A pinned SHA keeps the property `master` lacks:
rebuilding this tag later produces the same image.

The 40-char form is load-bearing, not pedantry. check-doc-drift.sh recognises a
literal SHA only through a 40-char match, so a 7-char pin would fall through to
its branch-or-tag lookup, fail to resolve, and downgrade pi-obsmem's drift check
to a silent SKIP -- a pin that reads correctly and is no longer verified. Full
gate run after this change: 23 OK, 0 DRIFT, 0 SKIP, 0 FAIL.

0.99.0/0.99.1/0.99.2 and 1.0.0 all exist upstream and are deliberately skipped:
obsmem's only compatibility work names Pi 0.87 (2b1dc1c) and a repo-wide
issue/PR search for 0.99 or 1.0 returns zero matches. 1.0.0 is its own round.

pi-atelier stays at v0.10.3 and is the residual risk. Its peerDeps declare pi
>=0.84.0 -- a floor, satisfied -- on both v0.10.3 and the current v0.12.1, so it
spans this bump. But atelier hooks pi TUI internals that a declared floor does
not protect, and an under-declared floor is exactly what failed to warn anyone
at pi 0.84. Acceptance must confirm the sidebar PAINTS, using the two-sided
check from 0.84.4/0.85.1 that distinguishes "loaded" from "silently absent".

Also here:
- check-doc-drift.sh gains a pi-obsmem pin check. The ref-move check covers the
  same component, but for a pinned SHA it can only answer "upstream did not
  move", never "the table still says what we bake".
- scripts/lint-shell.sh mode 100644 -> 100755. Pre-existing since f25efa0 and
  the only non-executable script in scripts/; it was latent because both CI
  steps call it as `bash scripts/lint-shell.sh`, but it failed rc=126 "bad
  interpreter" when invoked directly. Same dropped-exec-bit signature recorded
  on 2026-09-22, found the same way: by RUNNING it, not by reading a diff.
- Unreleased section renamed to `## v1.9.5 — 2026-10-02`, satisfying the
  release-gate rule that a tag's CHANGELOG must name its own version.
2026-10-02 00:23:28 +02:00
Joakim Persson cb8969ef2c changelog: adopt pi-studio v0.9.61 and mempalace-toolkit 975ab92
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 18s
Lint / doc-drift (push) Failing after 17s
Lint / skill-floor (push) Has been cancelled
Both are floating refs that moved since v1.9.4, so the next tag would bake them
either way; the gate's complaint was that nothing in the repo said so. Naming
them is the fix the gate actually asks for.

mempalace-toolkit 2167a1b -> 975ab92 adds mailbox dormancy: an ask may declare a
`dormant_unless` predicate and is withheld from the ANNOUNCED owed-set while
every condition still matches its baseline. It is backward compatible by
construction -- an ask without the key behaves exactly as before -- and
fail-visible: any predicate that cannot be evaluated announces the ask rather
than hiding it, because the dangerous failure is work that disappears, not a
spurious nag.

pi-studio v0.9.60 -> v0.9.61 is adopted as-is. NOT pinned, deliberately, and the
reason is worth recording: check-doc-drift.sh gives pi-studio kind `studio`,
which resolves the highest semver TAG of the upstream repo and ignores
PI_STUDIO_REF entirely (resolve_studio, ~line 551). Setting
PI_STUDIO_REF=v0.9.61 would therefore not be tracked by the gate -- the moment
upstream tags v0.9.62 the gate would report DRIFT against a value we no longer
bake, i.e. a false positive by construction. Pinning pi-studio is a coherent
thing to want, but it requires moving it from kind `studio` to kind `ref` in the
gate at the same time, which changes release-gate semantics and is a separate
decision from adopting this release.

Still drifting, and left drifting on purpose: pi-obsmem e7d77dc -> 731c3d4.
PI_OBSMEM_REF=master is CI-resolved at build time, master is 4 commits ahead of
the 3.1.4 tag, and the newest tag remains 3.1.4 -- so PR #83 ("Fix Pi 0.87
memory worker compatibility", merged 2026-09-23) is still reachable only from
master. Whether to bake an unreleased master or stay on 3.1.4 and hold pi is a
pin decision with a silent failure mode behind it (obsmem's peerDeps are all
`*`), so it is not being made in a changelog commit.
2026-10-02 00:01:24 +02:00
Joakim Persson 31f36182bf feat(base): add sqlite3, bc, dc and bsdextrautils
Lint / skill-floor (push) Successful in 10s
Lint / doc-drift (push) Failing after 13s
Lint / actionlint (push) Successful in 20s
Lint / hadolint (push) Successful in 21s
sqlite3 is the only one of the four that was genuinely missing rather than
merely absent. MemPalace keeps both the palace and the logstream as SQLite
files (palace/chroma.sqlite3, logstream.sqlite3) and mempalace_status reports a
sqlite_integrity block, so every integrity or forensics check on this fleet has
so far gone through a python3 -c one-liner because the CLI was not in the image.

bc and dc are low value on their own and the CHANGELOG says so plainly: awk,
python3 and perl are all already baked and each is strictly more capable. They
are here because copy-pasted shell snippets assume bc exists. What actually went
wrong on 2026-09-28 was NOT bc's absence: printf '%6.2f' was handed the empty
output of the missing bc and rendered it as a confident "0.00 days" for a figure
that was really 4.81 days. No package fixes that failure mode -- only not taking
a formatted number on trust does.

bsdextrautils ships /usr/bin/column.

Cost MEASURED rather than estimated: 1311 KB total (587 + 236 + 149 + 339) and
ZERO transitive packages. libsqlite3-0, libreadline8t64, zlib1g, libsmartcols1
and libtinfo6 each already report "install ok installed", so nothing new is
pulled under --no-install-recommends.

Package names were verified with apt-cache and dpkg -S on a real trixie host
instead of being inferred, which caught two traps that would each have produced
either a build failure or a silently missing binary: dc is a SEPARATE binary
package from bc on Debian, and column ships in bsdextrautils, not in the
pre-bullseye bsdmainutils where it used to live.

Declined at the same time, recorded so the omission reads as a decision rather
than an oversight: datamash and xsv/csvkit. Python's stdlib csv module handled a
real 7-file Excel-export concatenation that day -- UTF-8 BOM, no trailing
newlines, and bare CR/LF inside quoted fields -- correctly and without them.

Dockerfile.base is in the base hash, so the next tag rebuilds the base (~64 min)
regardless of what else it carries.
2026-10-01 23:31:03 +02:00
joakimp 2278b22ba7 fix(ssh): wire git core.sshCommand to the writable sidecar, and assert it
Lint / hadolint (push) Successful in 13s
Lint / skill-floor (push) Successful in 14s
Lint / doc-drift (push) Successful in 18s
Lint / actionlint (push) Successful in 29s
~/.ssh is commonly bind-mounted READ-ONLY from the host, so a per-host
`ControlPath ~/.ssh/cm/%r@%h:%p` — the standard CGNAT multiplexing recipe, and
correct on the host — resolves inside an unwritable dir in the container. Every
push dies `unix_listener: cannot bind to path ...: Read-only file system`,
behind git's misleading "make sure you have the correct access rights".

setup-lan-access.sh already writes the fix: ~/.ssh-local/config overrides
ControlPath into the writable ~/.ssh-local/cm BEFORE `Include ~/.ssh/config`,
so -F repairs the socket path and keeps every per-host User/Port/IdentityFile.
entrypoint-user.sh now points git at it, guarded on the sidecar existing —
setup-lan-access.sh writes none on native Linux Docker, where -F at a missing
file would break every git-over-ssh call instead of fixing one. An existing
core.sshCommand is left alone (first-wins, as for the three git settings above).

Why code and not another doc line: the remedy was already in the global
AGENTS.md, in pi-devbox-environment SKILL.md §3, in 24 MemPalace drawers from
three devices, and printed verbatim by recreate-sanity-check.sh — and an agent
that had run that script two hours earlier still hit the failure and reinvented
a /tmp/sshcm workaround. A fifth copy was not the missing piece.

Assertions, each where it can actually pass:
- smoke-test.sh: two STATIC greps (wiring line + its [ -r ] guard). `run` uses
  --entrypoint="", so asserting the runtime value there would repeat the v1.8.0
  mistake of an assertion that cannot pass, unvalidated until the next tag.
- smoke-test.sh runtime phase: a BICONDITIONAL — sidecar present => must route
  through it; absent => must be unset. The absent arm is the one CI exercises
  (native Linux runner), so "is set" would have failed CI for a correct image.
- recreate-sanity-check.sh: the runtime assertion, plus an explicit fail for the
  inverted state (set while the sidecar is missing). The permanent "default ssh
  precedence" warning keeps its severity but now states that it is structural and
  can never reach zero, and whether git is wired, unwired, or has no sidecar.

All five arms exercised against the real script before commit; that caught a
defect in the first draft, which reported "git IS wired ... unaffected" about a
state where the sidecar was gone and every git-over-ssh call failed.

Host ~/.ssh/config needs no change: the same line is right on the host and
unusable through a read-only mount, so the fix belongs in the container layer.
2026-09-22 23:08:59 +02:00
joakimp 37fcbfcf04 fix: three v1.9.4-acceptance findings (installer WARN, init sentinel, hash collation)
Lint / hadolint (push) Successful in 10s
Lint / skill-floor (push) Successful in 9s
Lint / actionlint (push) Successful in 20s
Lint / doc-drift (push) Successful in 18s
Unreleased; no pin moves except pi-extensions master 25c1265 -> 143a214.

1. entrypoint-user.sh: first-run sentinel is now config.json, which is what
   `mempalace init` writes. The old test on palace/ (created by mining, not
   init) re-fired "Initializing MemPalace" on every boot of a container that
   never mined locally: v1.9.4 acceptance measured 1 line after one boot, 2
   after a restart. Idempotent, so harmless; the comment lied. This file is a
   base-hash input, so the next tag rebuilds the base (326fb7c03949 predicted
   with the pinned sort below; f40c4b7b103d before).

2. Dockerfile.variant: `git config --system --add safe.directory` for the
   seven root-owned /opt clones (listed, not `*`). Since git 2.35.2 any git
   command in a repo owned by another user fails "dubious ownership"; as
   `developer` that made `git -C /opt/pi-atelier rev-parse` print nothing
   (one false FAIL in the v1.9.4 acceptance) and is the root cause of the
   per-boot "WARN: pi-extensions install.sh failed" (fixed at the source in
   pi-extensions 143a214; this is the belt to that suspender). Verified live
   in a v1.9.3 container: add -> rev-parse prints 25c1265, unset -> fatal
   again; /etc/gitconfig restored to its 5 lines afterwards.

3. docker-publish.yml: `find -print0 | LC_ALL=C sort -z` in base-decide.
   sort collates per locale; identical rootfs hashed to base-f40c4b7b103d
   under C/C.UTF-8 (== run 695) and base-d8df62216a81 under sv_SE/en_US.UTF-8.
   The runner exports LANG=C.UTF-8, so the hash was stable by accident. Pin is
   hash-neutral: pinned sort + HEAD inputs reproduces f40c4b7b103d exactly.
   Per-command prefix only; nobody's locale changes.

CHANGELOG: new `## Unreleased` above v1.9.4 naming pi-extensions 143a214
(check 9 rc=0), the three fixes, and the synlig 3.10.0 hub upgrade that
v1.9.4 listed as still open. One invented URL org (gwpl) caught before commit;
upstream is elpapi42, as Dockerfile.variant:151 says.

Gates: check-doc-drift rc=0, check-base-hash rc=0, lint-shell rc=0 (16 files),
hadolint 2.15.1 rc=0, YAML parses (10 jobs), bash -n rc=0.
2026-09-22 17:20:10 +02:00
joakimp 16fddebd43 docs(v1.9.4): the client/server skew is narrower than the release commit said
Lint / hadolint (push) Successful in 8s
Lint / skill-floor (push) Successful in 11s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 21s
Publish Docker Image / resolve-versions (push) Successful in 15s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-base (push) Successful in 1h4m42s
Publish Docker Image / smoke (push) Successful in 6m4s
Publish Docker Image / smoke-studio (push) Successful in 9m18s
Publish Docker Image / build-variant (push) Successful in 19m20s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / build-variant-studio (push) Successful in 24m15s
e3b38cd named "the feeder against the 3.9.0 hub" as the one path to exercise
before tagging, on the reasoning that 3.10.0's CLI writes now follow the daemon
write-routing policy. That was a reading of the mempalace changelog, not a
measurement of this image's feeder. Measured now:

  mempalace-pi-session in remote mode (MEMPALACE_REMOTE_URL set, which is how
  every fleet devbox runs) stages transcripts locally in python, rsyncs them to
  the hub host, and calls the hub's own `mempalace_mine` MCP tool over HTTP
  (run_remote_mine). The local `mempalace` CLI is required only in local mode
  (line 475) and invoked only on the local branch (line 1029). With a PATH shim
  logging every `mempalace` invocation, a 47-session `--dry-run` from this
  container logged ZERO calls; the shim's positive control logged one. The pi
  extension likewise speaks HTTP to the hub and spawns no local mempalace-mcp.

So the 3.10.0 CLI never runs against the 3.9.0 hub from this image. In remote
mode the client pin touches first-run `mempalace init` and the on-disk layout,
nothing else -- which is exactly what the ENV MEMPALACE_CONFIG_DIR change is
for. Both the Dockerfile.base comment and the CHANGELOG paragraph now say so,
and the "Still open" item becomes the thing that is actually unmeasured: a
first-boot acceptance of the built image from an empty ~/.mempalace volume.

Gates re-run on this tree: check-doc-drift.sh rc=0, lint-shell.sh rc=0,
hadolint 2.15.1 rc=0. lint.yml for e3b38cd: run 693, completed, success.
2026-09-22 14:51:22 +02:00
joakimp e3b38cdb0b release(v1.9.4): mempalace 3.10.0 with a pinned palace root, pi-atelier v0.10.3, pi held at 0.85.1
Lint / skill-floor (push) Successful in 8s
Lint / hadolint (push) Successful in 10s
Lint / doc-drift (push) Successful in 13s
Lint / actionlint (push) Successful in 22s
Three pins move or are deliberately held; the CHANGELOG's Unreleased section
becomes the v1.9.4 entry in this same commit, because since ee6cb9e the tag
build runs check-doc-drift.sh check 9 and every would-bake value must be named
above the last published heading.

mempalace 3.9.0 -> 3.10.0 (Dockerfile.base). Deferred at v1.9.3 for two
"Upgrade notes" items; both re-measured against the 3.10.0 wheel:

  - New installs resolve ~/.config/mempalace. The order is $MEMPALACE_CONFIG_DIR,
    then ~/.mempalace IF it holds config.json / people_map.json /
    palace/chroma.sqlite3, then XDG. An EMPTY ~/.mempalace fails that test, and
    an empty ~/.mempalace is what a freshly mounted devbox-palace volume (or
    entrypoint.sh's mkdir on a volume-less container) looks like at first boot.
    Measured with a fresh $HOME and `uvx mempalace==3.10.0 init`: config.json
    landed in ~/.config/mempalace, ~/.mempalace stayed empty, and
    entrypoint-user.sh's `[ ! -d ~/.mempalace/palace ]` would fire every start.
    Same run with MEMPALACE_CONFIG_DIR set: everything in ~/.mempalace.
    => ENV MEMPALACE_CONFIG_DIR=/home/${USER_NAME}/.mempalace, placed in the
    non-root-user section where USER_NAME is in scope, and entrypoint-user.sh
    reads PALACE_DIR from the same variable (old path as fallback). Existing
    volumes were safe either way; this is for first boots. MEMPALACE_PALACE_PATH
    is still honoured (config.py:927), so smoke-test.sh's stage test holds.
  - event_list defaults newest-first without a cursor. Server-side: the
    extension talks to synlig over MEMPALACE_REMOTE_URL, so it lands when the
    hub upgrades. mempalace-toolkit 2167a1b (floating main) makes every
    cursor-less call say order:"desc" so the mailbox reads the same window on
    either server version.

Corrected a stale sequencing comment while there: synlig serves the palace as
a `uv tool` under systemd (mempalace-serve.service; measured 2026-09-22:
mempalace 3.9.0, python 3.12.13, chromadb 1.5.9, no palace container), not via
docker-compose.mempalace.yml as the comment had said for two releases.

pi-atelier v0.10.1 -> v0.10.3 (Dockerfile.variant). Fixes only (both tags
released 2026-09-22); package.json at v0.10.3 still has zero runtime deps and no
build script, peerDeps still >=0.84.0. Annotated tags: the CHANGELOG names the
peeled SHA ed3837b, which is what resolve-versions bakes and check 9 compares.

pi HELD at 0.85.1 while 0.86.1 / 0.87.0 exist. 0.87.0 removed the
shouldStopAfterTurn agent option; pi-observational-memory 3.1.4 still uses it
and AgentContext.systemPrompt in observer/reflector/dropper (measured: 1 hit
per worker in src/agents on both the baked 3.1.3 tree and upstream master;
finishTurn 0). peerDeps are `*`, so it would break at runtime, not install.
Upstream #82, fix PR #83 open and unmerged. Bump pi and pi-obsmem together
once #83 ships. The other extensions are clean against the 0.86/0.87 lists.

Verified locally the way CI does: check-doc-drift.sh rc=0 with check 9
reporting all four moved components as named above the v1.9.3 heading;
lint-shell.sh rc=0; hadolint 2.15.1 (CI's pin, with .hadolint.yaml) rc=0 on
both Dockerfiles; PyPI 3.10.0 present, not yanked. Two claims caught by
re-measurement before commit and corrected in the text: a pi issue number
that does not exist on the 0.87.0 changelog line, and "8 shouldStopAfterTurn
hits" that was 3 (one per worker) in src/.

Still open before the tag: run the feeder from the built image against the
3.9.0 hub (dry-run first); recreate acceptance must show ~/.mempalace/palace
still passing and no ~/.config/mempalace appearing. After acceptance: upgrade
synlig's tool to 3.10.0.
2026-09-22 14:26:10 +02:00
joakimp 0d324f1855 changelog: name pi-obsmem 3.1.4 (e7d77dc) as the next rebuild's implicit adoption
Lint / skill-floor (push) Successful in 13s
Lint / hadolint (push) Successful in 15s
Lint / doc-drift (push) Successful in 14s
Lint / actionlint (push) Successful in 23s
check-doc-drift.sh was rc=1: pi-observational-memory moved cba0334 (3.1.3) ->
e7d77dc (3.1.4) upstream on 2026-09-20, the day after the v1.9.3 tag, and
nothing named it above the v1.9.3 heading.

No code change is needed or wanted here. PI_OBSMEM_REF defaults to master in
Dockerfile.variant and CI's resolve-versions turns that into a SHA at build
time, so the next rebuild adopts this whether or not anyone acts -- which is
precisely why it has to be written down. There is no pin in this repo to bump,
so the CHANGELOG is the only record a reader of the next tag would have.

v1.9.3 itself has no defect: its manifest, image labels and baked
/opt/pi-observational-memory clone all agree on cba0334. The drift is
post-tag, so a green drift reading taken on 2026-09-19 was correct when taken.

Measured, not assumed: the range cba0334...e7d77dc is 4 commits whose only
substance is one line in src/agents/worker-stream.ts (preserve registry
receiver when resolving stream, PR #78 / issue #77) plus a 19-line regression
test; the rest is the version bump and release merge. peerDependencies are
unchanged (all four @earendil-works/* peers still *) and there is no engines
block, so no pi or node floor to clear.

Gate now rc=0: "OK pi-obsmem cba0334 -> e7d77dc since v1.9.3, named above the
v1.9.3 heading".
2026-09-21 23:15:11 +02:00
Joakim Persson 6c13f43ac8 release: date the v1.9.3 section for the tag
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 16s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 41m49s
Publish Docker Image / smoke (push) Successful in 8m45s
Publish Docker Image / smoke-studio (push) Successful in 9m8s
Publish Docker Image / build-variant (push) Successful in 17m42s
Publish Docker Image / update-description (push) Successful in 11s
Publish Docker Image / promote-base-latest (push) Successful in 26s
Publish Docker Image / build-variant-studio (push) Successful in 21m28s
CI checks out the tag, so the CHANGELOG heading has to name the version
before the tag exists (same convention as v1.9.1/v1.9.2). check-doc-drift
check 9 was sabotage-tested for exactly this rename and still passes:
everything above the published v1.9.2 heading is the pending text.

Contents of v1.9.3 (all measured, nothing inherited):
  - f25efa0  recreate-sanity-check resolves the ControlPath ssh actually
             uses (ssh -G); the ~/.pi/ssh/config hyphen typo deleted
  - b8d818e  Hub size claims corrected; check 8 gates them against full_size
  - c7d369f  promote-base-latest's conditional re-tag measured (run 669)
  - 50153e6  skill floor -> pi-extensions 25c1265 (task tool + fork-gate)
  - ea89605  changelog for the two floating-ref changes: pi-extensions
             25c1265 and mempalace-toolkit 817b3a8 (mine deadline)
  - cb6d9e5  check 9: what the next build bakes differently must be named
             in the CHANGELOG; first run found pi-obsmem 7b397f4->cba0334
  - b9057fd  LABEL se.jordbo.pi-devbox.mempalace-version (inherited)
  - 960aada  dependency audit 2026-09-19; mempalace 3.10.0 deferred

Floating refs this tag will bake, named per check 9: pi-toolkit 9c87ee8,
pi-extensions 25c1265, mempalace-toolkit 817b3a8, pi-observational-memory
cba0334 (3.1.3). Base REBUILDS (Dockerfile.base + rootfs/ changed), so the
BASE_REBUILD_DATE marker is refreshed now, when it is free -- it had been
stale since 2026-07-13 through two rebuilds.

Deliberately NOT in this tag: mempalace 3.10.0 (see the audit section for
the three measured preconditions).
2026-09-19 17:45:50 +02:00
Joakim Persson 960aada769 changelog: dependency audit 2026-09-19 — mempalace 3.10.0 measured and deliberately deferred
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 17s
Lint / skill-floor (push) Successful in 8s
Lint / doc-drift (push) Successful in 13s
Every row by direct command: npm/PyPI/GitHub release APIs, git ls-remote,
the v1.9.2 config labels via the registry API, and binaries in a running
v1.9.2 container for the floating tools. Only pin with a newer upstream is
mempalace 3.9.0 -> 3.10.0, and it is NOT a small bump: new installs move to
~/.config/mempalace (entrypoint-user.sh tests ~/.mempalace/palace and
entrypoint.sh persists ~/.mempalace, so a fresh container would initialise
outside the volume), and mempalace_event_list flips to newest-first by
default (the toolkit's deriveClosed passes no order, deriveOwed on 2 of 3
calls — server-default-dependent). Deferred to its own release with the
three measured preconditions listed.
2026-09-19 17:38:02 +02:00
Joakim Persson b9057fdc8c Dockerfile.base: LABEL se.jordbo.pi-devbox.mempalace-version, inherited by both variants
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 10s
Lint / doc-drift (push) Successful in 16s
Lint / actionlint (push) Successful in 21s
Closes the blind spot check 9 (cb6d9e5) named: no label recorded the palace
pin, so a MEMPALACE_VERSION bump — the one component whose skew against the
shared central palace is fleet-wide — could ship without a CHANGELOG line.

In Dockerfile.base, not Dockerfile.variant, deliberately: the value sits next
to the ARG that defines it (a copy in the variant is one more pin able to
drift); labels are inherited by every image built FROM the base, so no
build-arg to plumb through the variant's four call sites; and inheritance
means the label states the pin of the base the image ACTUALLY built on,
which is the question when base-decide cache-hits an older base. Both
mechanisms measured on the published v1.9.2 config blob rather than assumed:
maintainer + image.source (set only in Dockerfile.base) are present on the
variant image, and pi-version=0.85.1 is an ARG expanded inside a LABEL.

Intent, like every se.jordbo.pi-devbox.* label; the manifest's
mempalace_version (read from the installed binary) stays the ground truth,
and smoke-test.sh now asserts label == installed core — the one way they
diverge is a base built with INSTALL_MEMPALACE=false or an off-pin install,
both invisible to a label-only check.

check-doc-drift check 9 gains the component (literal, against
ARG MEMPALACE_VERSION in Dockerfile.base); the label-key rule generalises to
"names ending in -version are the label itself". Until a release carries the
label it reports a counted SKIP, not OK — measured: "v1.9.2 carries no
se.jordbo.pi-devbox.mempalace-version label", summary says 1 SKIPPED.

Costs nothing extra: this Unreleased already forces a base rebuild
(50153e6 rootfs/ skill floor). check-base-hash unchanged (no new *_REF).
2026-09-19 17:27:34 +02:00
Joakim Persson cb6d9e5dd0 check-doc-drift: check 9 — what the next build bakes differently must be named in the CHANGELOG
Lint / hadolint (push) Successful in 8s
Lint / skill-floor (push) Successful in 10s
Lint / doc-drift (push) Successful in 14s
Lint / actionlint (push) Successful in 19s
The two most recent Unreleased entries (pi-extensions 25c1265: task tool +
fork-gate; mempalace-toolkit 817b3a8: the mine deadline that never reached
the transport) reached this image through floating *_REF=main ARGs. Neither
produced a diff here, so nothing asked for a CHANGELOG line, and neither had
one until a reader asked. Same class as check 8 — a claim with no in-repo
anchor rots — and the hand practice that covered it (the per-release
"Dependency audit" table, Baked-in vs Upstream-now) is a someone-remembers
mechanism that had lapsed.

Method, no docker/crane/token: last published vX.Y.Z from Hub's tag list
(one request, now shared with check 8); its amd64 config blob via the
anonymous registry API carries one se.jordbo.pi-devbox.<name>-ref label per
component with the SHA the build-args baked. "Would bake" is resolved the way
resolve-versions does it: 40-hex ARG is itself, branch/tag via git ls-remote
with the peeled ^{} form preferred, pi-studio = highest semver tag (label on
<tag>-studio), PI_VERSION as a literal against the pi-version label. Nine
components, ~7.5 s. Unchanged: OK. Moved: the new value's 7-char prefix (tag
for studio, version for pi) must appear above the last published version's
"## " heading — Unreleased plus any not-yet-published section, which is what
the release commit turns Unreleased into, so the tag build passes on the same
text. A published tag with no heading is a failure, not a skip.

First run found a move nobody had recorded: pi-observational-memory 7b397f4
-> cba0334 (6 commits, 3.1.1 -> 3.1.3; memory workers now routed by the
model's exact provider instead of the first registered provider with a
matching api — upstream #70). Named in the CHANGELOG with the honest scope:
no expected effect on the shipped bedrock-only configuration, unmeasured.

Sabotage-tested, expectation written before each run:
  obsmem SHA removed from CHANGELOG ............ rc=1  DRIFT
  Unreleased renamed to "## v1.9.3 — …" ........ rc=0  (release-commit shape)
  "## v1.9.2" heading mangled to v1.9.2-typo .... rc=1  — this one FAILED first:
      \b accepted "v1.9.2-typo" as v1.9.2's heading; now (\s|$)
  bare "## v1.9.2" (no date) ................... rc=0
  ARG PI_OBSMEM_REPO renamed ................... rc=2  (blind gate must not pass;
      read_arg calls are bare top-level assignments on purpose — nested in a
      heredoc's $(...) the exit 2 is swallowed by cat)
  offline (proxy to a closed port) ............. rc=0, 2 SKIPs counted
  SKIP_REF_CHECK=1 ............................. rc=0, 1 SKIP counted
  restored ..................................... rc=0, files byte-identical
GIT_TERMINAL_PROMPT=0 on ls-remote: a repo flipped private fails in 0.3 s as
a SKIP instead of waiting for a username until the job times out.

Header, lint.yml job comment and AGENTS.md all still said this gate "needs
no network" — false since check 8 (2026-09-14). Now: two classes, hermetic
1-7 and published-state 8-9 that SKIP loudly and counted when offline.
Considered and not added: compose <-> .env.example variable cross-check (the
four mismatches are commented-out lines, the mempalace-server compose file's
own variables, and entrypoint-consumed ones — it would fire on nothing
wrong); "documented tag exists on Hub" (check 8 already SKIPs by name; a
hard fail would misreport the tag-to-publish window).
2026-09-19 17:15:12 +02:00
Joakim Persson ea896054df changelog: task tool + fork-gate (pi-extensions 25c1265), and the mine deadline that never reached the transport (mempalace-toolkit 817b3a8)
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 18s
Lint / doc-drift (push) Successful in 6s
Both land in the image through floating refs (PI_EXTENSIONS_REF=main,
MEMPALACE_TOOLKIT_REF=main), so neither shows up as a diff in this repo — the
changelog is the only place a reader of a pi-devbox tag learns that the
container's delegation and feed behaviour changed. Recorded under Unreleased
with the measurements: 5/5 real fork briefs redirected, live pi -p proof for
gate and tool, 2-of-6 → 6-of-6 on the toolkit's new timeout test.

Dockerfile.base: the "Stall protection" comment listed MEMPALACE_MCP_TIMEOUT_MS
as THE tool-call deadline; it now says the feed's mine carries its own.
Comment-only, but Dockerfile.base is hashed into base_tag; the rebuild was
already forced by 50153e6 (rootfs/ skill floor).
2026-09-19 17:00:59 +02:00
Joakim Persson f25efa074d fix(sanity): verify the ControlPath ssh RESOLVES, not the dir the image creates
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 13s
Lint / doc-drift (push) Successful in 12s
Lint / actionlint (push) Successful in 24s
recreate-sanity-check.sh asserted `/tmp/sshcm exists with mode 700` and printed
a green tick while every ssh in the container died rc=255. It was right about
what it checked: the breakage was in the directory a CONFIG named, not the one
the image creates, and the old check could not see the disagreement between them.

Found on the v1.9.2 first boot on emb-7kj4vr4g. A durable ~/.pi/ssh/config,
hand-written into the ~/.pi named volume by the previous session so it would
survive the recreate, declared `ControlPath /tmp/ssh-cm/%C` — with a hyphen.
Nothing here creates that path; the canonical spelling is /tmp/sshcm, identical
in Dockerfile.base, entrypoint-user.sh, recreate-sanity-check.sh and
smoke-test.sh. ControlMaster auto with an unusable ControlPath does not degrade
to an unmultiplexed connection — it fails hard:

    unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory

rc=255, remote command never runs.

Now: resolve the ControlPath ssh itself would use via `ssh -G` and require its
parent to exist and be writable. -G applies real config precedence (first-value-
wins, the system drop-in, Include, an -F override), so it answers "which rule
captured this host" instead of re-implementing the guess, and it never opens a
connection — 0.116s for 48 hosts.

This also puts a check under a caveat documented in prose in Dockerfile.base
("SSH client defaults") and verified nowhere: a per-host ControlPath under a
read-only bind-mounted ~/.ssh gives the identical failure, "Read-only file
system". On the machine this was built on that is 16 of 50 hosts — freeipa-1..6,
gitea.egl.lan, runner-1..3, tor-ms22 — never reported by anything before.

Severity split is deliberate. Default ssh precedence legitimately lands in the
read-only ~/.ssh for any host whose own config pins it there, and the supported
workaround (ssh -F ~/.ssh-local/config, generated every start by
setup-lan-access.sh) already exists, so that is a warn. Failing it would paint
the script red on every run of every device, and a check that fires benignly
every time is one you learn to ignore — the same reasoning that keeps
lint-shell.sh at -S error. The sidecar route is prescribed, so there it is a
hard fail. Host lists cap at six names plus a count: unreadable output is
ignored output.

Teeth proven both directions, each sabotage confirmed by diff BEFORE the result
was believed: ControlPath -> nonexistent dir => rc=1; -> existing-but-read-only
dir => rc=1; restored => rc=0 with the sidecar byte-identical.

No config was added to fix the original problem — the fix was to DELETE
~/.pi/ssh/config, a third hand-maintained copy of what setup-lan-access.sh
already generates from version control on every container start, with fewer
features and one typo. The host-owned ~/.ssh/config is left alone on purpose:
those ~/.ssh/cm paths are correct on the host, where ~/.ssh is writable.

Also lint-shell.sh: the SC2088 count in the severity-choice rationale said 19
and is now 20, with the reproduce command recorded. That command needs a `$ `
prefix — a comment whose first word is "shellcheck" is parsed as a directive,
and the malformed one tripped SC1072/SC1073 at severity error. The gate caught
it on the very commit that introduced it.
2026-09-15 22:59:43 +02:00
joakimp c7d369f28d docs(changelog): measure promote-base-latest's conditional re-tag
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 20s
Closes the open question from v1.9.2's release verification. crane copy DOES
bump Hub's last_updated when it runs (run 669: digests differed, copy ran
17:10:04->17:10:08, base-latest.last_updated moved to 17:10). The
identical-digest path stays unproven by construction -- the job prints
'base-latest already current; nothing to promote.' and never copies -- so a
freshness assertion on the alias is conditional on base-decide's need_build,
not a blanket upgrade. Operational rule lives in the ci-release-watcher skill
(skillset c8034af); the pipeline property is recorded here.
2026-09-14 20:49:23 +02:00
joakimp b8d818ed99 fix(docs): correct the Hub size claims, and gate them so they cannot rot again
Lint / hadolint (push) Successful in 9s
Lint / doc-drift (push) Successful in 13s
Lint / skill-floor (push) Successful in 15s
Lint / actionlint (push) Successful in 18s
DOCKER_HUB.md claimed ~1.1 GB for :latest while Docker Hub served 1.37 GB --
20% wrong, drifting quietly across eight releases, and POSTed to Docker Hub by
update-description every time. It is the first number a stranger reads about
this image.

Root cause is structural, not carelessness: every OTHER claim check-doc-drift.sh
guards is anchored to a build file (a pin, an ARG, a placeholder), so it cannot
rot without someone editing the thing it describes. Nothing in this repo states
the image size, so nothing measured it.

Corrected against measured full_size after v1.9.2 published (amd64/arm64):
  :latest         ~1.1  -> ~1.23 GB   (1.228 / 1.211)
  :latest-studio  ~1.15 -> ~1.25 GB   (1.255 / 1.238)
  :base-latest    ~1.0  -> ~1.17 GB   (1.167 / 1.151)
full_size tracks the FIRST manifest entry (amd64), not the sum across arches --
measured: full_size=1.228, amd64=1.228, arm64=1.211, sum=2.439 -- which is what
the table's per-arch "Size (compressed)" column claims.

New check 8 gates those claims against Hub. Fails on drift; SKIPS LOUDLY via a
new skip() helper (counted, named in the summary) when curl/python3 are absent,
the API is unreachable, or SKIP_SIZE_CHECK=1. A skip is deliberately neither OK
nor a failure: printing an unverified claim as OK is the habit this file exists
to break, and failing on Docker Hub's uptime would hold releases hostage to a
third party. Not wired into hooks/pre-push, so pushes stay offline.

Two bugs caught by writing the expected exit code down before running the check:

  1. The tolerance would have missed its own motivating case. The percentage is
     computed against the MEASURED size, but the 20% first chosen came from the
     claim-relative figure; the real drift was |1.1-1.37|/1.37 = 19.7% and would
     have passed. Now 15%, inside a window with both bounds measured: above the
     largest legitimate skew (11.4%, a claim describing the published release
     while the next tag changes the size) and below the rot it must catch.

  2. A `|| true` on the python invocation made the gate FAIL OPEN -- it printed
     DRIFT and exited 0. Removed; the outer `|| SIZE_RC=$?` satisfies set -e
     without swallowing the code. Verified two-sided: 19% drift exits 1 at the
     default tolerance and 0 at SIZE_TOLERANCE_PCT=25, so the threshold does the
     work rather than the ordering.

NOT covered, and the script says so where a reader will see it: README.md's
~3.2 GB figures are UNCOMPRESSED and the registry exposes compressed sizes only,
so measuring them needs a real pull. A green check 8 says nothing about them.

Verified: check-doc-drift.sh rc=0 on the corrected tree, rc=1 on injected drift,
rc=0 with --warn-only, SKIP+rc=0 on the offline path (proxy to a dead port);
shellcheck clean at default severity (the sed single-quote SC2016 carries a
disable with its reason); lint-shell.sh 16 files clean.
2026-09-14 20:33:40 +02:00
joakimp f5c53b8693 release: date the v1.9.2 section for the tag
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 27s
Publish Docker Image / base-decide (push) Successful in 13s
Publish Docker Image / build-base (push) Successful in 43m59s
Publish Docker Image / smoke (push) Successful in 5m29s
Publish Docker Image / smoke-studio (push) Successful in 5m50s
Publish Docker Image / build-variant (push) Successful in 17m29s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / update-description (push) Successful in 10s
Publish Docker Image / build-variant-studio (push) Successful in 21m50s
CI checks out the tag, so the CHANGELOG heading has to name the version
before the tag exists, not after. v1.9.1's tagged tree carried no
"## Unreleased" either -- same convention, made explicit here.

Contents of v1.9.2 (all measured, nothing inherited):
  - 1baba79  the +131 MB v1.9.1 residual: /root/.npm (110 MB) plus
             @mariozechner/clipboard-* foreign natives at BOTH install
             sites (global + /opt/pi-fork's nested pi-coding-agent copy)
  - 852f900  smoke sentinels that name the residue instead of letting
             the size gate's ~225 MB margin swallow it
  - dab989b  mempalace-toolkit: the owed-withdrawal suite gate
             (node --check never type-strips)
  - 9aaff26  vendored mempalace skill snapshot -> e9e45f7, which is what
             turns the currently-red canary green

Deliberately NOT in this tag: DOCKER_HUB.md's "~1.1 GB" size claim is
low (Hub says v1.9.1 = 1.37 GB, v1.8.14 = 1.23 GB). This release
*changes* the size, so no number is right both before and after it --
writing the predicted ~1.24 GB would be an unmeasured claim in a
user-visible page. Measure post-build, fix on the next tag, and extend
check-doc-drift.sh to gate size claims against Hub's full_size so the
number cannot rot silently again.
2026-09-14 17:55:16 +02:00
joakimp 735565b9be docs: retire the deployed-and-unproven label on isWithdrawn, and name dab989b
Lint / skill-floor (push) Successful in 8s
Lint / hadolint (push) Successful in 17s
Lint / doc-drift (push) Successful in 7s
Lint / actionlint (push) Successful in 20s
Two things, and the second was found by doing the first.

v1.9.1's notes closed with an explicit open item: isWithdrawn had 17 assertions
and 4 mutation kills but had never been exercised on a released image against the
live logstream, which is a different state from untested and a worse one to leave
unlabelled. It has now been exercised, from a container recreated onto v1.9.1, and
the label is retired with the measurement rather than with an assurance.

What the entry records is the instrument as much as the result, because the result
is only as good as the thing that produced it: the shipped extractor was lifted
out of the toolkit's own test script and used to cut the four predicates out of
the BAKED mempalace.ts, and deriveOwed's three queries were replayed with their
exact shipped parameters. No second copy of the logic was written, which is the
same discipline the suite itself is built on. Baseline owed was established by
three independent routes BEFORE any probe was planted, because without a recorded
baseline a withdrawal that appears to work proves nothing. Predictions were
written down before each measurement; the table reports both columns.

The per-predicate attribution is in the entry deliberately. "The count went back
to baseline" is a claim about arithmetic and would have been satisfied by a
pre-filter or by an unrelated reply; isAnswered=false with isWithdrawn=true on a
candidate that was still in the raw set is a claim about mechanism. The negative
controls carry the same weight as the positive one -- a rule that lets a third
party or a prose sentence empty someone's mailbox would have passed a
count-only check.

Also recorded: the rule fires on the real incident it was written for (seq 112),
which is worth more than any synthetic probe, and the one path still unwitnessed
(the extension's own in-process poll, which only fires when the agent is idle) is
named as unwitnessed rather than folded into the pass.

Second: reaching for the suite as corroboration showed it cannot run on this image
at all -- its own precondition gate rejects the source it exists to check, because
`node --check` does not type-strip on any node and v1.9.1 moved to node 24. All 17
assertions had been skipped since the bump. Fixed in toolkit dab989b, named here
per the rule this repo adopted two commits ago after the withdrawal fix shipped
unnamed in v1.9.1. The gate now also distinguishes "cannot run" from "does not
parse", since collapsing those is how the defect pointed readers at the wrong
file.

The rebuild cost is stated and, having been measured, is smaller than the earlier
draft of this entry claimed: 9aaff26 already moved base_tag by refreshing the
vendored skill snapshot under rootfs/, so the base rebuild was pending before this
fix existed and the toolkit pickup rides along with it.
2026-09-14 16:29:52 +02:00
joakimp 9aaff26e3a docs: the withdrawal fix shipped in v1.9.1 unnamed — say so, and re-pin the canary
Lint / hadolint (push) Successful in 11s
Lint / doc-drift (push) Successful in 12s
Lint / skill-floor (push) Successful in 12s
Lint / actionlint (push) Successful in 16s
v1.9.1 bakes mempalace-toolkit e68ee20, which contains e2b060a. Requester-side
ask withdrawal has therefore been LIVE on every v1.9.1 device since 2026-09-10,
while the v1.9.0 section of CHANGELOG.md still read "not yet pinned ... this
image still pins e45f6b4", and the mempalace skill still told every agent at
session start that a withdrawal is impossible.

Measured, two independent routes, expectation recorded before looking:

  - published image label, :v1.9.1-studio and :latest-studio (same digest):
      se.jordbo.pi-devbox.mempalace-toolkit-ref = e68ee2071ca3ad39...
  - ancestry: e2b060a is an ancestor of e68ee20
  - baked mempalace.ts sha 7c16fe14 != v1.8.14's dfca71e9 (different bytes)
  - grep -c isWithdrawn on this container's baked copy = 0 (v1.8.14, so
    tor-ms22 cannot exercise the behaviour it is documenting)

The first label read came back EMPTY, and that empty was a claim about the
request rather than the image: Hub redirects blob fetches to a CDN and curl
without -L returns 0 bytes at exit 0. Recorded in the notes, because a registry
audit reporting "no labels" is missing -L until proven otherwise.

Changes:

  - CHANGELOG Unreleased: the floating-ref mechanism, which is the reusable
    part. ARG MEMPALACE_TOOLKIT_REF=main + CI resolving it to a SHA at build
    time means a release absorbs whatever toolkit main holds, and "what
    behaviour did this image gain" is a question nobody is forced to answer.
    The fleet rule (name the pickup before tagging) was honoured for the
    feed-tick commits v1.9.1 names, and missed for one in the same range.
  - CHANGELOG v1.9.0: the stale caveat is ANNOTATED, not rewritten. The
    sentence is the evidence for how the drift happened; deleting it would
    destroy the only trace. Same discipline as fleet-ops d42412c.
  - CHANGELOG Unreleased: corrected the size entry's "needs no base rebuild"
    claim. True of that entry alone, false of the release now that the
    vendored skill snapshot -- a base_tag input -- moves with it (~67 min).
  - rootfs mempalace skill snapshot 44472 -> 46045 B, refreshed via
    scripts/vendor-mempalace-skill.sh so the bytes and SKILLSET_SNAPSHOT_REF
    (4d7c0ea -> e9e45f7) move together; --check confirms exact match.
  - scripts/smoke-test.sh canary RE-PINNED. The retired pair was still green
    against the new snapshot, i.e. blind to this refresh for the same reason
    the pre-v1.8.13 pair was blind to that one. The replacement is stronger
    than any predecessor here because BOTH witnesses come from the same
    upstream commit: e9e45f7 added "Withdrawing an ask you sent" and deleted
    "nothing anyone can do about it from the other end", the sentence the new
    bullet contradicts. Directions measured against both files, then the
    canary body EXECUTED against each: new -> rc=0 "ok", old -> rc=1 empty.
    A canary whose negative witness was removed by the commit it pins fails
    loudly on stale bytes instead of merely failing to notice them.

Gates: check-doc-drift OK, check-skill-floor OK, vendor --check OK,
hooks/pre-push OK (16 shell files clean at severity error).

Not fixable here: isWithdrawn is deployed and UNPROVEN. 17 assertions and 4
mutation kills, never once exercised on a released image against the live
logstream. Ask routed to a v1.9.1 device.
2026-09-14 15:07:26 +02:00
Joakim Persson 852f900b53 test(smoke): make CI prove the natives still work — the runbook check didn't
Lint / doc-drift (push) Successful in 6s
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 9s
Lint / actionlint (push) Successful in 18m11s
The post-boot check v1.9.1 left for the next machine was

  node -e 'require("esbuild").transformSync("const x:number=1",{loader:"ts"})'

with "if this fails, the prune removed something needed on arm64 -> revert to
v1.8.14 and fix" attached. Run from /workspace on the freshly recreated v1.9.1
image it fails with MODULE_NOT_FOUND — on a perfectly healthy image. require()
resolves by walking up from the CURRENT DIRECTORY, esbuild is nested inside the
two pi trees, and a global install is not on node's require path (NODE_PATH is
unset). The version that "verified the prune" last night only passed because the
shell happened to sit inside the tree. So the check was CWD-dependent, its red
meant nothing, and the remedy it prescribed was reverting a good release.

Path-qualified, both sites compile TS today. Rather than leave a fragile command
in a note for a human to paste, CI now owns it:

  - esbuild must transformSync at EVERY install site found in the image
  - @mariozechner/clipboard must load with its native binding attached at every
    site — for the clipboard prune that IS the proof, since napi-rs resolves the
    platform package at require() time

Sites are discovered with find, so the studio variant's third site is covered
without naming it.

Verified as a four-way matrix, not just a green run: both assertions pass
against the real image FROM /workspace (the cwd that broke the old command), and
both fail (exit 1, "The package \"@esbuild/linux-arm64\" could not be found, and
is needed by esbuild") against copies of the same packages with the host
platform binary removed. An assertion never observed failing is not evidence.

Also documents both failure shapes in AGENTS.md next to the size gate: this one,
and the mode-700 /root one where `test ! -d` passes on a permission error.
2026-09-11 11:06:18 +02:00
Joakim Persson 1baba79c96 fix(size): the v1.9.1 residual was npm's own cache, not the platform binaries
Lint / doc-drift (push) Successful in 6s
Lint / hadolint (push) Successful in 10s
Lint / skill-floor (push) Successful in 9s
Lint / actionlint (push) Successful in 4m26s
v1.9.1's @esbuild prune fixed the 431 MB size-gate failure but still shipped
+131 MB compressed over v1.8.14, nearly all of it in the pi/extensions install
layer (87 -> 206 MB). That leftover was filed as an open item with an explicit
hypothesis — the @mariozechner/clipboard-* family, same npm 11 behaviour, a
different package — and an explicit warning that the hypothesis was not a
measured cause. Measured now, after recreating onto v1.9.1, and the hypothesis
accounted for one sixth of it:

  +110 MB  /root/.npm/_cacache  (35.2 -> 145.3 MB)   the build's npm cache
  + 21 MB  clipboard foreign platform packages, both install sites
  = 131 MB  i.e. the whole delta, no unexplained remainder

Method, since there is no docker CLI inside the container: pulled both variant
layer blobs straight from the registry with a token + manifest + blob fetch and
listed the tarballs (29 789 vs 29 889 entries, 270.8 vs 401.9 MB uncompressed),
then aggregated per package. The file COUNT barely moved, which is what said
"few large files", not "npm installed more packages".

  - purge_build_caches: npm cache clean --force + rm -rf /root/.npm, in the SAME
    layer as the installs, in both the main RUN and the studio RUN. npm 11
    caches every platform tarball it downloads, including the ones the prune
    then deletes, so the cache grew faster than the tree. Nothing at runtime
    reads it: build is root, container is developer with its own cache in $HOME.
  - prune_foreign_esbuild -> prune_foreign_natives: now covers both MEASURED
    families. Clipboard keeps linux-$arch-gnu AND -musl because its napi-rs
    loader picks between them at runtime via its own isMusl() probe; the musl
    package is a 420-byte stub. The bare @mariozechner/clipboard wrapper has no
    hyphen suffix and cannot match the pattern.

Verified on arm64 before writing the glob — a widened rm -rf against a tree you
cannot inspect is the one change shape not to write blind, which is why the
order was update-then-patch. Exercised against a copy of the real trees with
foreign dirs fabricated back in (aix-ppc64, android-arm64, darwin-arm64,
win32-x64, linux-x64): all removed, host linux-arm64 kept at both sites, 21 MB
freed, require('@mariozechner/clipboard') still loads and exports all 18
functions, esbuild.transformSync still compiles TS at both sites. v1.9.1's own
arm64 validation also passed here; CI could only smoke amd64.

Two sentinel assertions, because the size gate did not catch this: it has
~225 MB of margin, so 131 MB of residue stayed green. Both were verified RED
against the running v1.9.1 image and GREEN against a pruned tree. The cache one
refuses to run as non-root: test ! -d /root/.npm on mode-700 /root would
otherwise pass for the wrong reason. Size-failure diagnostics now list cache
paths too — they previously enumerated only node_modules and /opt, where these
bytes were not.

Also fixed: -printf '%f\\n' reaches the shell with both backslashes (confirmed
from the published image's recorded created_by), so v1.9.1's progress line
printed a mangled "li ux-arm64" — find emitted a literal backslash and tr ate
the n out of the name. Single backslash now.

Deliberately not purged: /tmp/node-compile-cache (1.3 MB). The manifest RUN
calls pi --version again, so deleting it earlier only relocates those bytes into
that layer — today's manifest layer is 128 kB precisely because it finds the
cache warm.

Dockerfile.base untouched, so no base rebuild: this rides the next release.
2026-09-11 09:39:54 +02:00
Joakim Persson 42bd29d654 fix(ci): unblock the release — npm 11 esbuild bloat + two self-inflicted assertions
Lint / hadolint (push) Successful in 8s
Lint / skill-floor (push) Successful in 9s
Lint / doc-drift (push) Successful in 10s
Lint / actionlint (push) Successful in 19s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 14s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Successful in 5m4s
Publish Docker Image / smoke-studio (push) Successful in 13m17s
Publish Docker Image / build-variant (push) Successful in 33m49s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 14s
Publish Docker Image / build-variant-studio (push) Successful in 30m58s
v1.9.0 was tagged but never published: smoke failed 90-passed/3-failed and
build-variant needs smoke, so nothing reached the registry. All three are fixed.

1. Size, 431 MB over threshold. Node 24 brings npm 11, which installs EVERY
   @esbuild/<platform> optional binary instead of the matching one: 26 dirs,
   284 MB per pi-coding-agent copy. Measured on pi-fork: npm 10.9.8 -> 165 MB
   (exactly what v1.8.14 shipped), npm 11.19.0 -> 449 MB. npm 11 ignores the
   os/cpu constraints AND --os/--cpu AND an npmrc carrying them, all measured,
   so Dockerfile.variant prunes explicitly, keeping linux-$(node -p
   process.arch) so one line is right on both arches. Pruned in the SAME layer
   as each install, or the bytes survive in the earlier layer. Three sites:
   global pi, pi-fork, pi-studio. Verified esbuild still transforms TS after.

2. om's node_modules assertion tested an npm artefact. om has zero runtime deps;
   its node_modules held ONE file (.package-lock.json) and 20 empty scope dirs.
   npm 11 stopped creating it. Now asserts the entry point pi actually loads,
   read from package.json -> pi.extensions.

3. The skill-source annotation added in v1.9.0 ("baked (package copy)") broke
   the assertion matching "baked$". Pattern now allows an optional suffix; which
   copy shipped stays authoritatively asserted against the manifest + tree hash.

Also: a failed size check now prints the largest layers, largest directories and
an @esbuild sentinel, so this class attributes itself next time instead of
costing a CI dig plus a local npm bisect.

Threshold stays 3800 MB: it caught a real regression and raising it would have
thrown the signal away. No Dockerfile.base/rootfs change, so base-0fb1256c7f99
is reused and build-base is skipped.
2026-09-10 23:39:40 +02:00
Joakim Persson 3a44e81cad feat(ci): gate documentation drift, and make docs a pre-tag release step
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Five doc claims had rotted by v1.9.0, all the same shape: a value written once
by hand, in a file nothing verifies, about a number that lives elsewhere and
moved. README pin table wrong on all three rows; a "Planned" section describing
something already shipped; DOCKER_HUB.md claiming Node v22 against Node 24.

DOCKER_HUB.md is why this is a gate and not a resolution to be careful: it is
PUBLISHED (update-description POSTs it as Docker Hub full_description on every
tag), it had gone eight releases untouched, nothing generates it, and it is read
from the TAG -- so the stale page shipped with v1.9.0 regardless.

scripts/check-doc-drift.sh: seven checks, all repo-local (no network, token,
image, or sibling clone). Exit 0/1/2 matching lint-shell.sh; a renamed ARG is a
red 2, not a green tick. Wired as a fourth lint.yml job so "the docs lie" is its
own red name.

Verified with 15 controls, including two false-positive controls: the first
placeholder check flagged README.md:900, a Go template in a legitimate
`docker inspect --format` example. The gate was wrong, not the doc, so the
pattern is now anchored to the UPPER_SNAKE convention CI substitutes.

Not gated, deliberately: counts/sizes needing a running image (they belong in
smoke-test.sh -- a guessing gate is worse than none), and Dockerfile.base
BASE_REBUILD_DATE, because base_tag hashes that file content-wise and demanding
it be current would force a ~60 min rebuild on releases that touch no base
files. Free during a rebuild, expensive otherwise.

AGENTS.md step 3 rewritten around the mechanism: checkout@v4 with no ref: means
every job reads github.ref, the tag. Docs must be right BEFORE tagging.
2026-09-10 22:01:49 +02:00
Joakim Persson 8f0960e134 release: v1.9.0
Lint / skill-floor (push) Successful in 8s
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / lint-gate (push) Successful in 22s
Publish Docker Image / resolve-versions (push) Successful in 15s
Publish Docker Image / base-decide (push) Successful in 8s
Publish Docker Image / build-base (push) Successful in 42m19s
Publish Docker Image / smoke-studio (push) Failing after 6m17s
Publish Docker Image / build-variant-studio (push) Has been skipped
Publish Docker Image / smoke (push) Failing after 9m7s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Rename the Unreleased changelog section to its release heading, matching the
established `## vX.Y.Z — YYYY-MM-DD` format (em dash), per AGENTS.md release
step 3.

Minor rather than patch: CHANGELOG.md scopes minor to "significant base
additions", and this release bumps the Node runtime under every baked JS tool
(pi, agent-browser, playwright, mempalace) from 22 to 24 — the first
NODE_VERSION change since it was introduced at v1.0.0 — alongside five new
base packages (shellcheck, bind9-dnsutils, ldap-utils, xxd, python3-yaml).
Package additions alone have been patch here (v1.8.12, v1.8.13); the runtime
major is what lifts this one.
2026-09-10 21:21:09 +02:00
Joakim Persson 1c905480e3 docs(changelog): note the mempalace feed-tick fix this image carries
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 10s
Lint / actionlint (push) Successful in 19s
The recurring "[mempalace ext] feed (tick) failed: mine timed out after 30000ms"
message is fixed in mempalace-toolkit (309980b + e68ee20) and this image is what
delivers it, since CI resolves MEMPALACE_TOOLKIT_REF to a commit SHA at build
time. Worth a changelog entry rather than leaving it implicit in a ref bump: it
is the most visible symptom operators on this fleet have been living with, and
the entry records that it was a genuine defect (overlapping mines on a
single-writer palace) rather than the cosmetic annoyance it was parked as.
2026-09-10 20:58:42 +02:00
Joakim Persson ff6fd1492a feat(manifest): record WHICH pi-extensions skill copy shipped
Closes the half deliberately left open by cac5e00's skill-floor gate, and the
more important half: "the floor is currently fresh" is a fact with a shelf
life, whereas "the image says which copy it got" keeps working.

The refresh in Dockerfile.variant is guarded by
`[ -f /opt/pi-extensions/skill/SKILL.md ]`, so a build whose clone predates the
co-located skill keeps the vendored floor and still succeeds GREEN, with nothing
in the manifest, labels or logs separating that from a normal build. Afterwards
the two are indistinguishable by inspection -- same path, same filenames, same
permissions -- which is exactly how the floor went unnoticed from 2026-07-30 to
2026-09-10.

build-manifest.json gains pi_extensions_skill_source and
pi_extensions_skill_tree_sha256, MEASURED rather than passed as build-args, per
the ground-truth rule the surrounding block already follows -- and necessarily
so, since the outcome depends on the clone's contents and no ARG could express
it. Three values, because two would force a lie: package (served bytes equal
the clone's skill/), vendored-floor (clone had no skill/ at this ref), and
divergent (both exist but differ -- e.g. the clone ships SKILL.md but not
evaluate-extension-usage.py, so the served directory is a genuine MIX). No OCI
label mirrors these deliberately: LABEL cannot take a RUN-computed value, and a
label fed from an ARG would be the claim-not-measurement being removed here.

Two smoke assertions make the record a gate: the source must be named and be
`package` -- vendored-floor FAILS rather than warns, since these images track
main where the package has shipped skill/ since fa04d20, so a fallback means
the clone did not resolve as intended -- and the tree hash is recomputed over
the served directory, because a recorded hash never recompared is a claim.

pi-devbox-version annotates the line too: "baked (package copy)" normally, or a
yellow "(FALLBACK: vendored floor)". Its existing section reports which copy is
READ at runtime; this is the one fact decided at BUILD time and unrecoverable
later. Old images degrade cleanly -- field absent, jq // empty yields nothing,
line prints plain "baked" as before (verified against this v1.8.14 manifest).

Tested by running the exact logic against this container's real layout, with
the expected value written down before each: package (served == clone),
vendored-floor (clone path absent), divergent (clone lacking the .py while the
served dir has it), and null (empty served dir) -- all four as predicted. The
five pi-devbox-version render branches likewise, including the absent-field
case. Emitted JSON validated with jq for both the populated and null forms.

Gates green: lint-shell.sh (15 files), hadolint 2.15.1, actionlint 1.7.12,
check-base-hash.sh, check-skill-floor.sh, vendor-mempalace-skill.sh --check.
2026-09-10 20:30:16 +02:00
Joakim Persson edc7659add chore(deps): node 22->24, actionlint 1.7.12, hadolint 2.15.1, skillset ref
Audited every component the image obtains OUTSIDE debian/apt. Of ~23, the 19
that resolve `latest` at build time were already current or refresh themselves
on the next rebuild, and the hard pins for pi (0.85.1), mempalace (3.9.0) and
pi-atelier (v0.10.1) were already newest. Four needed a human.

NODE_VERSION 22 -> 24 (LTS "Krypton"). This was a latent defect rather than
housekeeping: agent-browser publishes engines.node ">=24.0.0", so the image sat
BELOW a declared requirement -- v1.8.14 shipped node 22.23.2 with agent-browser
0.37.1, so every build installed it with an npm EBADENGINE warning and ran the
baked browser automation outside its supported range. pi (">=22.19.0") and
playwright (">=20") are satisfied either way. Verified before bumping, since a
missing NodeSource suite breaks every arch at once: setup_24.x returns HTTP 200
and node_24.x advertises `Architectures: amd64 arm64 armhf x86_64`, covering
the arm64 fleet and the amd64 CI runners. Nothing else pinned the node major.

actionlint 1.7.7 -> 1.7.12 and hadolint 2.14.0 -> 2.15.1, each RUN AGAINST THIS
TREE at the new version before being pinned -- both clean, no new findings. A
linter bump is the one dependency update that can turn CI red on unchanged
code, so it is verified locally rather than discovered on a round trip.

SKILLSET_SNAPSHOT_REF e9e09d9 -> 4d7c0ea via scripts/vendor-mempalace-skill.sh,
never by hand: that script is the only thing permitted to write the ARG,
because a cp without a matching bump yields a manifest that confidently lies.
This proved PROVENANCE-ONLY -- the ref was 6 commits behind, but
skills/mempalace/SKILL.md is byte-identical at both (3675bfab), so the snapshot
was already correct and only its recorded origin was stale. No rootfs/ bytes
changed, the smoke-test phrase canary stays valid, and this ARG alone would not
force a base rebuild (the node bump does).

Two measurement traps worth recording, since both would have produced a wrong
answer: GitHub's releases/latest reports pi-atelier v0.10.0 as newest because
v0.10.1 is a TAG WITH NO RELEASE OBJECT -- the pin was already current, and
`git ls-remote --tags` is the instrument that shows it. And gitea-mcp is hosted
on gitea.com, not GitHub, so querying api.github.com returned nothing at all
rather than an error.

Verified with every gate this repo owns, all green, using the NEW linter pins:
lint-shell.sh (15 files), check-workflow-shell.sh, check-base-hash.sh,
actionlint 1.7.12, hadolint 2.15.1, check-skill-floor.sh, and
vendor-mempalace-skill.sh --check.
2026-09-10 20:17:32 +02:00
Joakim Persson cac5e00a31 feat(ci): gate the vendored pi-extensions skill floor, and bake python3-yaml
Follows ecfd2fc, which refreshed the stale floor by hand. A one-off refresh
fixes the symptom; this makes the drift impossible to reintroduce silently.

scripts/check-skill-floor.sh compares the repo floor
(rootfs/usr/local/share/pi-devbox/skills/pi-extensions/) against the package
repo it is a snapshot of, wired in as a new `skill-floor` job in lint.yml.

DIRECTORY hash, not `sha256sum SKILL.md`, using the same tree_sha256 pipeline
Dockerfile.variant uses for skillset_snapshot_tree_sha256 and for the reason
already documented there: a file-only compare answers "did this one file
change", not "is this the same skill". Verified by NEGATIVE CONTROL rather
than asserted -- with SKILL.md left byte-identical and only
evaluate-extension-usage.py edited, the directory check fails (rc=1) where a
file-only compare would have passed. Seven behaviour tests, each with its
expected rc written down before running: in-sync via local dir (0), in-sync
via anonymous remote clone (0), missing --package-dir (2), bad argument (2),
content drift (1), the sibling-file case (1), and --warn-only over drift (0).

Exit codes 0 in sync / 1 drift / 2 cannot-run, matching scripts/lint-shell.sh:
a gate that cannot run must not pass, so an unreachable package repo is a red
2 and never a green tick. A ref with no skill/ is NOT drift -- that is the
documented fallback -- but it emits ::warning:: because it is precisely the
condition under which the floor ships.

Gating on another repo is normally a smell. It is proportionate here because
the check can only fire when skill/ itself changed, which is exactly when the
floor has gone stale; pi-extensions commits that leave skill/ alone cannot
turn this red. It also needs no secret: pi-extensions is anonymously clonable
(verified with `git ls-remote` and no credentials), so it cannot start failing
when a token expires.

Also bakes python3-yaml (552 KB, zero extra deps) into Dockerfile.base. This
is the shellcheck story repeating exactly: scripts/check-workflow-shell.sh --
the guard against the Gitea sh/dash footgun that broke resolve-versions
(ed49b8d) and promote-base-latest (b7197e8) -- hard-exits with "python3 yaml
module missing", so a gate this repo already owns could not be run locally by
anyone. lint.yml installing it explicitly in CI was the evidence. Found while
wiring the job above: the guard could not be run before pushing.

CHANGELOG Unreleased updated for both this and ecfd2fc, including an explicit
note on what is NOT fixed -- the silent-fallback half still has no manifest
flag recording which copy was served.

Verified locally with every gate this repo owns, all green: lint-shell.sh (15
files clean), check-workflow-shell.sh, check-base-hash.sh, actionlint 1.7.7
(pinned, same version as CI), hadolint 2.14.0, and the new check itself.
2026-09-10 19:14:01 +02:00
joakimp 15a3728ae9 feat: bake shellcheck and add a client-side pre-push lint gate
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 19s
v1.8.14 made shell lint a RELEASE gate (scripts/lint-shell.sh, shared by lint.yml
and the new lint-gate job that resolve-versions depends on), and that script
correctly exits 2 when shellcheck is absent -- "a gate that cannot run must not
pass". Measured on v1.8.14 on 2026-09-09 by three routes (command -v, dpkg -l, a
filesystem search): shellcheck was NOT IN THE IMAGE AT ALL. So the gate could not
be run by a developer in any container, only in CI, and the loop stayed
write-shell -> push -> wait for CI -> discover. That is the loop the gate was
added to shorten, after v1.8.14's first attempt burned ~46 min on a tree whose
lint had already been red for 24 hours.

shellcheck 0.10.0-1 added to the Dockerfile.base apt block: ~39 MB installed
(Installed-Size 40112 KB), measured to pull ZERO additional packages under
--no-install-recommends because libc6/libffi8/libgmp10 are already present.
NOTE this forces one full base rebuild -- base-decide hashes Dockerfile.base +
rootfs/, so unlike a scripts/ change it cannot reuse the existing base- layer.

hooks/pre-push is opt-in per clone (git config core.hooksPath hooks), bypassable
with --no-verify, and execs scripts/lint-shell.sh rather than reimplementing it
-- one copy, because a duplicated check that drifts is the failure this repo
keeps paying for. Matches the idiom skillset/ and myconfigs/ already use.

WHY THIS REPO HAD NO HOOKS, since it was reported as drift and is not: a peer
asked tor-ms22 for core.hooksPath per clone on the premise that unset meant the
gates were unverified there. Measured: pi-devbox unset, skillset hooks, myconfigs
common/hooks, pi-toolkit unset -- but `git ls-files | grep -i hook` is EMPTY in
both pi-devbox and pi-toolkit, so there was nothing to point at on any machine
and unset was the only correct value. This closes the real half for pi-devbox;
pi-toolkit still ships none.

Verified, expected result written down before each check:
  * refusal paths -- shellcheck absent => rc 2 with the remedy named; linter
    missing => rc 2. Never waved through on the assumption CI will catch it.
  * the hook is IN the scan set -- "Checking 14 shell file(s)" with it present,
    13 with it moved aside, so the extensionless file is found by the shebang
    half of the linter's two-signal union. This check exists because the first
    attempt was ambiguous: a planted `[ $UNSET_VAR = "x" ]` was not reported,
    which could equally have meant "not scanned" or "below -S error". It was the
    latter. A count that moves is unambiguous; a clean run is not.
  * it catches the REAL v1.8.14 defect -- planting `echo 'the fleet\'s thing'`
    in hooks/pre-push yields SC1073/SC1072 at severity error, rc=1.

And the gate earned its keep inside this commit: the first version of the
smoke-test assertion carried a comment beginning "# shellcheck is a GATE
DEPENDENCY", and a comment whose first word is the tool's name is parsed as a
DIRECTIVE, not a comment. The new gate failed it with SC1073/SC1072 before the
push -- same family as the v1.8.14 apostrophe, a line that reads as prose to a
human and as syntax to the parser.
2026-09-09 08:57:34 +02:00
joakimp 361babd4fd ci: gate the release on shell lint, from one shared script
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Publish Docker Image / smoke (push) Successful in 7m39s
Publish Docker Image / build-variant-studio (push) Successful in 17m15s
Publish Docker Image / build-variant (push) Successful in 17m58s
Publish Docker Image / update-description (push) Successful in 8s
Publish Docker Image / promote-base-latest (push) Successful in 12s
v1.8.14's first attempt spent ~46 minutes building a base image for a tree whose
own lint had been failing for 24 hours. shellcheck had already flagged the
defect (SC2289, severity error) on the push that introduced it; the lint
workflow went red at run 186 and nobody read it.

lint.yml deliberately skips tag pushes and its reasoning is sound -- the tagged
tree was already linted on main, and a tag-ref lint run sorts above the publish
run, making a release look finished before anything ships. The missing invariant
was never "lint the tag". It was "do not RELEASE a tree whose lint failed", and
only a job inside the publish workflow can enforce that.

So: extract the shell-lint logic from lint.yml into scripts/lint-shell.sh and
call it from both places, then add a lint-gate job that resolve-versions depends
on. resolve-versions is the graph root, so gating it gates everything. Cost is
~40 s at the front of a release; the alternative already cost fifty minutes.

Extracted rather than copied on purpose. A second copy of a check is the drift
this repo keeps paying for -- the same evening produced a skillset mirror that
had sat 9579 B behind its upstream through two consecutive edits.

The script adds one behaviour the inline version lacked: if shellcheck is not
installed it exits 2 rather than silently finding nothing, inheriting the
existing "a gate that cannot run must not pass" rule from hooks/pre-commit in
the skillset repo. Without that, reordering the install step away would turn the
gate into a green tick over zero checks.

Verified locally with a stubbed shellcheck (the real binary is not in the
devbox), five cases, each with its expectation stated first: absent shellcheck
-> rc=2; stub pass -> rc=0 and a non-zero file count; stub fail -> rc=1; a
deliberately unterminated `if` planted in scripts/ -> rc=1 via the bash -n half,
naming the file; removal -> rc=0 again. Discovery cross-checks against CI's own
number: the inline version reported 12 files, the extracted one reports 13, the
difference being lint-shell.sh itself. YAML re-parsed (10 jobs, was 9) with an
assertion that resolve-versions needs lint-gate, and the repo's
check-workflow-shell.sh guard still passes.
2026-09-08 23:41:44 +02:00
joakimp 70e675afee fix(smoke): keep prose out of the single-quoted exec_test body
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 22s
The agent-browser execution guard added on 2026-09-07 carried its explanation
INSIDE the single-quoted script body, and the explanation contained an
apostrophe ("the fleet\'s only recurring amd64 runtime proof"). Inside '...'
bash treats a backslash as literal, so \' does not escape the quote -- it CLOSES
the string. The body truncated at that point and the remaining lines were parsed
by the calling shell.

Consequences, both measured rather than inferred:
  - exec_test received 12 arguments instead of 2 (verified two-sided: the fixed
    tree yields argc=2, HEAD yields argc=12).
  - the leaked `v=$(agent-browser --version)` ran on the CI RUNNER instead of
    inside the image. The runner has no agent-browser, so smoke and
    smoke-studio both failed with "line 770: command not found" after
    build-base had already spent ~46 minutes. Every downstream job was skipped.
  - the truncated body still passed inside the container and printed its green
    tick first, so the log shows a PASS immediately followed by the failure --
    the tick was real, it just no longer covered the assertion.

The prose now sits above the exec_test call, where an apostrophe cannot
terminate anything, and a comment at that spot records why it must stay there.

Not a new failure class: shellcheck flagged it as SC2289 at severity error the
same day, so the lint job has been red since run 186 (2026-09-07 21:21) and was
not read. The gate did its job; nobody looked.
2026-09-08 23:31:48 +02:00
joakimp 601fc98a49 docs(changelog): release v1.8.14
Lint / hadolint (push) Successful in 9s
Publish Docker Image / resolve-versions (push) Successful in 14s
Lint / actionlint (push) Failing after 24s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 50m32s
Publish Docker Image / smoke-studio (push) Failing after 5m13s
Publish Docker Image / build-variant-studio (push) Has been skipped
Publish Docker Image / smoke (push) Failing after 7m40s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Converts the Unreleased section and records what this build carries beyond it:
the mempalace-toolkit bump that makes closing replies reach the mailbox
(deriveClosed, 21023e7 -> e45f6b4), and the L0-L4 subtask documentation landing
via pi-toolkit adfb553 + pi-extensions c64c122.

Notes the mechanism that makes the toolkit fix land at all — the resolved
toolkit SHA is folded into the content-addressed base tag, so the toolkit
moving forces a base rebuild rather than waiting for one — and the consequence
for the amd64 item already in this section: v1.8.13's base was cached, so
Dockerfile.base:607's agent-browser assertion never ran. This base is not
cached, so the native-amd64 proof is finally collected instead of discarded.
2026-09-08 22:09:31 +02:00
joakimp 6bd8b79d3a test(smoke): assert agent-browser EXECUTES — it was the discarded amd64 proof
Lint / hadolint (push) Successful in 11s
Lint / actionlint (push) Failing after 23s
Second instance of the same bug class as the node line, in the same file, found
the same way. The agent-browser guard captured the version inside an echo with
2>/dev/null:

  echo "resolved=[$r] version=[$(agent-browser --version 2>/dev/null|head -n1)]" >&2

so the exit code was discarded and a binary that could not execute at all still
PASSED, printing version=[]. Verified two-sided: a stub exiting 127 passes the old
form and is caught by the new one.

Why this exit code matters more than most: smoke runs platforms: linux/amd64 on an
x86 runner, i.e. NATIVE amd64, so this line is the fleet's only recurring amd64
runtime proof for agent-browser's linux-x64 ELF.

NO DEVBOX CAN EVER SUPPLY THAT PROOF. Every machine in the pi fleet is an Apple
Silicon Mac: mbp-m1-2020; tor-ms22 = Mac Studio Mac13,1 M1 Max (fleet-ops
hosts/tor-ms22.md, verified 2026-08-17 with system_profiler); emb-7kj4vr4g =
Apple Silicon, verified 4 routes 2026-09-07. The open "amd64 runtime proof still
needed" ask sent to two devices was asking for the impossible, and emb's reply
naming tor-ms22 as "the only remaining candidate" is wrong for the same reason.
CI had the answer all along and was throwing it away.

Dockerfile.base:607 DOES assert it (`agent-browser --version && \`), but only when
the base rebuilds, and v1.8.13's base was cached — so smoke is where the recurring
gate belongs.
2026-09-07 21:21:16 +02:00
joakimp fabf1274aa docs(changelog): Unreleased section for the smoke node assertion + agent-browser correction
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 19s
Summarises what changed since v1.8.13: the node-major assertion (a bump would
have passed the suite silently), the two-sided verification of the derivation,
and the v1.8.13 agent-browser 0.35.2 -> 0.36.0 correction. No image content
changes; NODE_VERSION still 22.
2026-09-07 21:17:06 +02:00
joakimp 5972a2c535 test+docs: assert the node major in smoke, and correct v1.8.13's agent-browser version
Two findings from a delegated read-only audit of this repo, both verified from the
filesystem before patching.

1. No test asserted the node major, so a node-24 bump would have passed the smoke
   suite SILENTLY. scripts/smoke-test.sh:94 was a bare `run "node" "node --version"`
   — exit-0 and non-empty output only, the printed version compared to nothing —
   while the line above it uses run_expect against $EXPECTED_PI_VERSION for pi. A
   reader skimming the suite would reasonably assume node regressions were covered.
   Worse, this is where the "node v22.23.2 verified" line in the v1.8.13 recreate
   notes came from: printed output, not an assertion.

   Now gated on EXPECTED_NODE_MAJOR, which CI derives from Dockerfile.base's ARG
   NODE_VERSION — the single source of truth (Dockerfile.base:557 is the ONLY hard
   pin in the repo; Dockerfile.variant has no node install at all). That also
   catches a stale cached layer whose node disagrees with the declared ARG.
   Unset => previous behaviour, so this is backward compatible.

   Verified two-sided rather than assumed: the sed derivation yields 22 (empty
   would have silently disabled the assertion, reintroducing the bug); grep -Fq
   "v22." matches v22.23.2; "v24." does NOT match, so a wrong major is caught; and
   "v2." does not prefix-collide. Workflow YAML re-parsed after editing (9 jobs).

2. The v1.8.13 entry claimed "the image's own 0.35.2" for agent-browser. The image
   ships 0.36.0: /usr/lib/node_modules/agent-browser/package.json says version
   0.36.0, engines.node >=24.0.0, and no 0.35.2 exists anywhere in the image. The
   claim was also internally incoherent, contrasting 0.36.0 against a version that
   is not present. Corrected in place with a visible note, since the entry is
   already released. The reasoning survives untouched: the engines floor really is
   vestigial, because /usr/bin/agent-browser is a prebuilt aarch64 ELF invoked
   directly and never through node — which is why 0.36.0 runs fine on 22.23.2.
2026-09-07 21:05:24 +02:00
joakimp aa0fbc5ec0 fix: correct the pi-studio claim — CI publishes v0.9.59, not the v0.9.60-rc.0 label
Lint / actionlint (push) Successful in 17s
Lint / hadolint (push) Successful in 14s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Successful in 4m44s
Publish Docker Image / smoke-studio (push) Successful in 5m8s
Publish Docker Image / build-variant (push) Successful in 15m52s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / build-variant-studio (push) Successful in 21m22s
Measured at the wrong layer during the v1.8.13 audit. I read `ARG
PI_STUDIO_REF=main` in Dockerfile.variant, concluded the release would adopt
main (= v0.9.60-rc.0), set PI_STUDIO_VERSION to that, and wrote a comment plus a
CHANGELOG entry describing deliberate RC adoption. A Dockerfile default cannot
answer "what will CI publish?" when CI overrides it, and it does: build-variant
passes PI_STUDIO_REF=studio_ref and PI_STUDIO_VERSION=studio_tag (lines 598-599
and 787-788), and resolve-versions picks the newest STABLE semver tag via
`^v?[0-9]+\.[0-9]+\.[0-9]+$`, which excludes pre-releases.

Caught by reading run 639's own resolve-versions output rather than the
Dockerfile: studio_tag=v0.9.59, studio_ref=9eed84f = refs/tags/v0.9.59^{}, while
main/v0.9.60-rc.0 is 658536f and never gets built. So published v1.8.13 studio
images carry pi-studio v0.9.59.

ARG restored to `none` rather than pinned to v0.9.59: the local-build default
should not hardcode a tag that goes stale as soon as main moves, which is how the
previous value came to lie. The comment now leads with the override so the next
reader starts at the layer that decides. Upstream's tag-over-main policy is
deliberate (Releases stopped at v0.5.55, main receives half-finished commits), so
adopting an RC from CI would mean changing that filter, not this ARG.

Consequence kept on purpose: the RC's opt-in Studio network binding is in NO
published v1.8.13 image, so it needs no audit this release.

Doc/label-only: base_tag hashes Dockerfile.base + rootfs/** + both entrypoints +
mempalace_toolkit_ref, none of which this touches, so the in-flight base build
(base-ad9faf00f2b2) stays valid and the tag run will reuse it. Verified with CI's
pinned linters: hadolint 2.14.0 exit 0 on both Dockerfiles, actionlint 1.7.7 exit
0, shellcheck 0.10.0 -S error exit 0.
2026-09-06 23:51:00 +02:00
joakimp f561acc89a skills: refresh vendored mempalace snapshot a12fe5e -> e9e09d9, re-pin the canary
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Failing after 31m52s
Folded into v1.8.13 at zero marginal cost: the snapshot is hashed into
base_tag, but Dockerfile.base already changed this release, so the ~67 min base
rebuild was already being paid. vendor-mempalace-skill.sh --check reported exit
0 (stale-but-truthful) beforehand, so skipping was sanctioned -- this is the
deliberate call the release checklist asks for. Upstream content: the bare
project-name wing convention and the <harness>@<device> added_by rule, both
downstream of the attribution defect measured on this device 2026-09-06.

The canary re-pin matters more than the refresh. Its old pair ("Provenance is
stamped for you" present / "Attribute what you file yourself" absent) still
PASSED against the new snapshot, so leaving it would have yielded a canary
green on both old and new bytes -- blind to exactly the refresh it exists to
witness, the same false-green family as the pre-v1.8.5 canary. New pair chosen
by measuring direction against both files rather than reading the diff
("Diaries self-heal; plain drawers do not" new=1/old=0; "Agent diaries live in"
new=0/old=1), then tested two-sided: PASS on refreshed bytes, FAIL on the old
bytes recovered from git.

Gates after the change: smoke-test.sh parses, vendor --check exit 0,
check-base-hash exit 0.
2026-09-06 22:31:47 +02:00
joakimp 0d984b1414 changelog: cut v1.8.13 section
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 18s
2026-09-06 22:14:20 +02:00
joakimp adcf56f829 release: audited bumps (pi 0.85.1, mempalace 3.9.0, atelier v0.10.1) + two guards
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 17s
Version audit for the next release. pi 0.84.4 -> 0.85.1, deliberately skipping
0.85.0 (it published internal experimental code and broke SDK imports,
upstream #9132). mempalace 3.8.0 -> 3.9.0. pi-atelier v0.10.0 -> v0.10.1.
PI_STUDIO_VERSION relabelled none -> v0.9.60-rc.0 so the floating main ref's
RC status is visible at docker-inspect time instead of discovered later.
PI_FORK_REF stays floating and adopts e69725c.

The pi bump was verified by running it under a pty in five combinations rather
than by reading the changelog, because this repo has already shipped a version
pair no changelog flagged (atelier < 0.7.1 hangs pi >= 0.84). CPU delta
0.00-0.01s over 5s against a ~5s sustained-CPU hang signature, two-sided via
the atelier sidebar painting identically to the 0.84.4 control.

NODE_VERSION stays 22 on purpose: node 24 is technically safe (pi's five
prebuilt addons are all NAPI, nothing declares a ceiling, agent-browser's
engines.node >=24 is vestigial for the shipped aarch64 ELF), but this release
already moves two minors and bakes an RC, and a node major would leave four
suspects if the image misbehaves. Own release, smoke suite as the gate.

Also corrects a stale claim at the mempalace ARG: synlig serves 3.8.0
server-side, not 3.7.1 (measured over ssh 2026-09-06).

agent-browser volume shadowing: the image has shipped 0.35.2, but every
session on mbp-m1-2020 ran 0.27.0 from a 2026-07-17 hand-install in
~/.pi/npm-global (a VOLUME, at PATH position 2 vs /usr/bin at 8). Third
package hit by this hazard after pi and pi-atelier, so the guard is now
generalised: entrypoint-user.sh retires the copy by moving it aside
(reversible, only when the image ships its own), recreate-sanity-check.sh
asserts resolution under /usr where the volume is real, smoke-test.sh carries
the build-time half and says in the source why it is weak. The real damage was
the stale BUNDLED SKILL (3 skillsets/17.6 KB vs 8/31.5 KB, ten subcommands
undocumented to the agent) - a stale tool errors, a stale skill quietly
teaches wrong commands.

pi-fork capability floor (extensions: []): forks were measured across four
dispatches ignoring their brief, answering in the user's voice, fabricating
self-referential measurements, and once filing a diary entry as agent_name=pi.
Cause is upstream by design - the child gets getHeader()+getBranch(), the
whole active session branch, with the brief as the final user message. Not a
model-capability problem: the same model as the fast profile obeyed the
identical brief perfectly with a fresh session and no inherited context.
extensions: [] runs children with --no-extensions, so the mempalace bridge is
absent and palace writes are impossible by construction (verified by asking a
child to enumerate its tools: read, bash, edit, write). Removes palace writes,
not filesystem writes.
2026-09-06 20:40:02 +02:00
joakimp c8622ece9d skills: correct the credential-incident-response §5 premise about chroma metadata
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 17s
§5 said embedding_metadata.string_value holds "metadata fields only". False,
measured directly: chroma also stores a copy of the document text there, under
key chroma:document. Confirmed with a disposable sentinel drawer (pi@tor-ms22,
2026-08-30): one row in fts_content AND one row in embedding_metadata for the
same drawer.

This was a real mistake in shipped guidance, not a nitpick: this section's own
scanning advice was written to guard against explaining a zero with a
mechanism nobody verified from source, and the section itself did exactly
that -- I downgraded a census to "a floor" on the strength of a metadata-blind
claim I never checked against chroma's actual storage layout. The practical
scan order is unchanged (fts_content is still the direct target, raw bytes are
still the backstop); only the stated REASON for a metadata zero changes: it
needs a different explanation now (key filter, query shape, escaping), not
"structurally absent".

§6's row-gone/bytes-gone claim is upgraded from asserted to measured, same
sentinel: delete_by_source took both fts_content and embedding_metadata 1->0,
raw bytes stayed 4->4 (freed pages persist until VACUUM). Also records the
method that unblocked the measurement: not a better instrument, a disposable
sentinel drawer instead of risking real fleet data.

No image behaviour changes.
2026-09-01 22:37:46 +02:00
joakimp 05843ecfae changelog: reopen an Unreleased section after v1.8.12
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 16s
v1.8.12's release retitled the previous Unreleased heading, leaving the file
with no place to put the next change — so the next contributor either invents a
heading or appends to a released section. The note under it points at the
release checklist step that renames it, so the convention is discoverable from
the file rather than only from AGENTS.md.

Also the first push after moving CI off synlig: lint.yml should now run on
runner-a1 (8 vCPU / 16 GB, Debian 13, upstream Docker CE) instead of the box
that hosts the palace.
2026-09-01 00:29:29 +02:00
joakimp a2846a5f7e release: adopt pi 0.84.4 + pi-atelier v0.10.0, and fix the doc claim the pi bump invalidates
Lint / hadolint (push) Successful in 11s
Lint / actionlint (push) Successful in 21s
Publish Docker Image / resolve-versions (push) Successful in 19s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Successful in 59m39s
Publish Docker Image / smoke-studio (push) Successful in 5m22s
Publish Docker Image / smoke (push) Successful in 17m53s
Publish Docker Image / build-variant-studio (push) Successful in 17m14s
Publish Docker Image / build-variant (push) Successful in 18m22s
Publish Docker Image / promote-base-latest (push) Successful in 12s
Publish Docker Image / update-description (push) Successful in 20s
pi 0.84.3 -> 0.84.4 (no Breaking Changes / Removed heading in that section,
grepped). Adopted for three fixes that land on machinery this fleet runs:
#6879 (large tool results crossing the auto-compaction threshold were sent to
the provider before compacting), #8345 (a resumed session corrupted its next
appended entry when the JSONL lacked a trailing newline -- that file is the
memory feeder's input; measured 49/49 clean here beforehand), and #8537
(triggerTurn:false messages sent mid-run were inserted between a tool call and
its result). The mempalace mailbox is outside #8537's precondition: it delivers
at agent_settled with deliverAs:"steer" and no triggerTurn, and 0.84.4 leaves
the documented steer semantics unchanged.

pi-atelier v0.8.2 -> v0.10.0: two minor releases, both UI-only, no BREAKING
notice. v0.9.0 raises its minimum pi to 0.84.0 and, unlike the
0.7.1-under-pi-0.84 startup-hang precedent, encodes it in peerDependencies
(>=0.84.0). Satisfied by PI_VERSION=0.84.4. Both executable floors compare with
sort -V, so 0.10.0 >= 0.7.1 evaluates correctly.

docs/observational-memory.md: pi's own compaction.md gained one paragraph in
0.84.4 -- autoCompact is now also checked mid-run, after a tool batch's results
are appended. Our text said compaction is checked only when pi goes idle and so
"never interrupts a turn"; that was only ever true of the OM trigger. The
section now states both entry points into session_before_compact and the
diagram carries the second edge (mermaid checker re-run: 6 blocks, 44 labels,
0 soft-wrapped, no cut glyphs at 1280px and 800px).

README: the version-pin table had been wrong since v1.8.6 -- 93f986e moved
ARG PI_VERSION to 0.84.3 and MEMPALACE_VERSION to 3.8.0 and neither table row,
so it advertised pi 0.84.2 / mempalace 3.7.1. Corrected, plus the
--expected-version example that would now fail against a 0.84.4 image.

CHANGELOG: Unreleased retitled v1.8.12 (2026-08-31) with the audits above and a
dependency-audit table -- every other component measured SAME (skillset
snapshot --check OK at a12fe5e, 0 commits since baked).
2026-08-31 07:08:41 +02:00
joakimp 58c22afb04 skills: the fingerprint advice was missing its precondition, and the skill had no section on proving absence
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 23s
Docs only; no image behaviour changes.

WHY THIS AND NOT A PRIVATE NOTE. pi@emb-7kj4vr4g reported itself for printing
sha256[:8] fingerprints of GIT_USER_EMAIL, GIT_USER_NAME and HOST_SSH_USER, and
wrote a private rule forbidding it. It had not broken a rule. It followed §2 of
this skill as written, and §2 is incomplete: it says a fingerprint lets you
compare a credential "without ever materialising the secret" with no condition
attached. When two agents independently make the same mistake, the artifact that
taught them both is the bug.

§2 NOW CARRIES THE PRECONDITION. A fingerprint is 32 bits over its INPUT SPACE,
so publishing fp8(x) hands anyone a MEMBERSHIP ORACLE: they can test x == v for
every candidate v they can generate. Safe for a 40-char random token; a wordlist
for a hostname, username, e-mail, port, path, commit SHA or weak password. "High
entropy" is the usual sufficient condition, NOT the test — a commit SHA is
160-bit and still fully enumerable from the repo. Operationally: if you can
imagine writing the wordlist, you cannot publish the fingerprint. Also added:
candidate fingerprints are working memory and never output (an extractor hashes
hostnames and paths too, so the tempting "print what the scanner saw" debug step
leaks low-entropy fingerprints wholesale), and a plain statement that a
fingerprint register is a CONFIRMATION ORACLE for anyone already holding a
candidate corpus — which is exactly how a retired token is identified in old
transcripts, and works identically for someone else holding those same files.

NEW §6, "Proving absence: instrument strength, and four ways a scan lies clean",
placed next to §5 on purpose: §5 optimises against false POSITIVES, and every
failure in §6 is a false NEGATIVE. Triage optimises precision, a gate optimises
recall, and conflating them is what produced three clean reports over secrets
that were really there. Contents: instrument ranking (exact-byte value search >
class/structure pass > fingerprint census) with the instruction to state which
one produced your zero; census vs class passes as different questions, both
failure modes measured on this fleet; the tokenisation trap where quoting alone
decided detectability; scan the index or pushed tree, never the working tree;
git filters never run on symlinks while check-attr claims they do; two-sided
self-tests that abort, incl. the fixture-interaction artifact; row-gone is not
bytes-gone.

Attribution kept per finding: the census/class split and the instrument
ranking's provenance are pi@emb-7kj4vr4g's; exact-byte search over index blobs
is pi@tor-ms22's. The credential sense of "census" originated in this skill, not
with either agent.

TRAP FOR THE NEXT EDITOR, also in the CHANGELOG: the frontmatter description is
now 1022 of 1024 characters. Trim before adding, or the skill silently fails to
load. Verified by parsing the frontmatter (1022 chars, name intact, every prior
trigger phrase retained).

Deployment: baked skill -> needs an image rebuild AND a container recreate to
reach a running container.
2026-08-30 23:32:53 +02:00
joakimp 30094782df shell: source cli_utils' functions, and install the iproute2 that one of them needs
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 30s
v1.8.11 linked cli_utils' bin/ COMMANDS onto PATH and stopped there. Nothing ever
sourced cli_utils.sh, so its 14 FUNCTIONS were missing from every interactive shell
whose $HOME has no zsh rc -- which is the normal case, not an edge case: the
container's interactive shell is bash and zsh is not installed in the image. A
symlink cannot carry a shell function and a function cannot be reached from a
non-interactive shell, so the two mechanisms are disjoint and both are required.
The image was already paying this layer's dependency cost (fzf, bat, fd, rg, jq are
baked partly FOR these functions) while delivering none of its benefit.

The two changes ship together because they are coupled: portcheck is one of the 14,
and it was a hard stub in every image up to v1.8.11 -- neither ss nor ip nor lsof
nor netstat was present, so it printed "portcheck requires at least one of: ss,
lsof, netstat" and exited. Wiring the functions in without iproute2 would have
shipped a visibly broken one.

MEASURED, not assumed:
  - the loader is bash-safe despite the *.zsh filenames: `bash --noprofile --norc`
    exits 0, defines all 14, and they run (pathls, mkcd, up, extract, agents-sync,
    fhist verified). The tree's one zsh-only construct (print -z in fzf/fhist.zsh)
    is already guarded by [[ -n $ZSH_VERSION ]] with a bash fallback.
  - fresh-$HOME seeding resolves 14/14; CLI_UTILS_SOURCE=0 is honoured; an absent
    checkout is a genuinely silent no-op (no output, no leaked _cu).
  - interactive shell startup 12 ms -> 17 ms.
  - iproute2 is ~5.5 MB (4.2 MB itself + 6 libs under --no-install-recommends;
    libpam-cap is a Recommends and correctly dropped). ss lands at /usr/bin/ss,
    ip at /usr/sbin/ip, both already on the developer PATH, and `portcheck --all`
    then correctly identifies the socat listener on 8765.
  - hadolint clean on both Dockerfiles; repo-wide shellcheck -S error and bash -n
    clean. .bash_aliases is outside CI's discovery (no shebang, not *.sh), so it
    was checked by hand with -s bash at error AND warning level.

Named explicitly per this repo's floating-ref rule: /workspace/cli_utils is a HOST
BIND MOUNT, not a pinned ref, so the image now executes unpinned content in every
interactive shell. Errors are left visible rather than sent to /dev/null so that a
future zsh-only file in that repo is diagnosable rather than mysterious, and
CLI_UTILS_SOURCE=0 is the documented escape hatch. It is deliberately independent
of CLI_UTILS_LINK=0: the two disable independent mechanisms.

Deployment: needs a rebuild AND a recreate. $HOME is the container's writable layer
rather than a named volume (verified -- ~/.bash_aliases carries the container start
mtime while ~/.bashrc carries the image's), so the skel file is re-seeded on every
recreate; a host-bind-mounted ~/.bash_aliases is still never overwritten.
2026-08-30 11:48:03 +02:00
joakimp d9a7fe101b changelog: an Unreleased section for a rule that was already there
Lint / actionlint (push) Successful in 16s
Lint / hadolint (push) Successful in 1m22s
Records the two skill commits ahead of tomorrow's build, and states the finding
that shaped them: the "a negative result is usually your own filter" rule was
already baked, already symlinked in at every container start, and already
survived every recreate — then was violated five times by a session that had it
available. The gap was activation, not persistence, which is why the
cross-cutting form went into the always-appended AGENTS block instead of into a
skill that only loads when a task description matches.

Also notes what the entry's own subject implies for the reader: neither change
reaches a running container until the image is rebuilt AND the container
recreated, since ~/.agents/skills and the global AGENTS.md both live in the
image rather than in a volume or a mount.
2026-08-30 00:52:36 +02:00
joakimp b615571913 changelog: release v1.8.11 — the mail nobody could deliver, and the ship that skipped a file
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 19s
Publish Docker Image / base-decide (push) Successful in 10s
Publish Docker Image / build-base (push) Successful in 53m20s
Publish Docker Image / smoke (push) Successful in 4m48s
Publish Docker Image / smoke-studio (push) Successful in 4m58s
Publish Docker Image / build-variant-studio (push) Successful in 17m0s
Publish Docker Image / build-variant (push) Successful in 26m43s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 24s
Names three things riding in via floating refs, since the rule this
fleet adopted (v1.8.9) is to name a behaviour change before tagging,
not after:

- mempalace-toolkit b2b50af -> 21023e7: the rsync --checksum ship fix
  (a real behaviour change to every remote-mode client) and RFC 003
  §7.13 (docs only).
- skillset 6eb20af -> a12fe5e: mermaid-diagrams cutU normalisation and
  the mempalace skill's from_agent identity rule, plus the vendored
  fallback snapshot re-pinned to match (this pi-devbox commit).

Dependency audit run by direct command against every component, not
assumed: pi/mempalace/pi-atelier/pi-studio/pi-toolkit/pi-extensions/
pi-fork/pi-observational-memory all confirmed unchanged since v1.8.10.
pi-studio's apparent SHA drift on ls-remote turned out to be an
annotated-tag-object hash vs its underlying commit hash, not real
drift -- checked with ^{commit} before writing it down as 'none'.
2026-08-27 23:39:39 +02:00
joakimp 45850bc973 entrypoint: put back the shell state a recreate eats
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 25s
cli_utils' install.sh reaches PATH by symlinking bin/ into ~/.local/bin.
That is persistent on a host and ephemeral in a container, so the same
installer produced opposite durability and every --force-recreate sent the
human back to typing /workspace/cli_utils/bin/git-status-all. Re-link at
start instead, and add a per-device boot hook so the next question of this
shape needs no image change at all.

Symlinks rather than a PATH edit in an rc file, deliberately: ~/.local/bin
is already ahead of /usr/local/bin in ENV PATH, so links resolve in
NON-interactive shells too (docker exec, agent tool shells, scripts). An
rc-file PATH edit cannot reach those because ~/.bashrc returns early when
not interactive — measured, that asymmetry is exactly why `command -v
git-status-all` failed in one shell and worked in another on the same box.
Shell FUNCTIONS remain the sourced file's job; a symlink cannot carry them.

Guards, because ~/.local/bin is shared: a real file is never clobbered, a
symlink pointing elsewhere is never stolen, ours are refreshed, and links
into a cli_utils/bin whose target vanished are pruned — a dangling link on
PATH reads as a broken container rather than a removed script.

The hook (~/.config/devbox-shell/init.sh) adds no trust boundary: that dir
is already sourced into every interactive shell by the baked bash_aliases,
so it is already arbitrary code from the same owner. Only WHEN it runs is
new. bash <file>, never sourced, exit status ignored, output to a log.

Caught before commit, and the reason the loops use `if` bodies instead of
`&&` chains: under `set -euo pipefail` a for-loop whose last command is a
false test exits non-zero, and with no match the /workspace/*/cli_utils
glob stays literal — so the first draft would have failed to START a
container on every machine that does not have this repo, rather than merely
skipping the links. Re-tested with set -e in place: no-cli_utils/empty-HOME
no-op, guards, idempotence, CLI_UTILS_LINK=0, a hook that exits 7, and a
hook that tries to mutate CLI_UTILS_BIN — all exit 0 with intact state.

Not covered by CI: docker-publish.yml runs only on tags and lint.yml lints
workflow run: steps, so neither executes this file. Validated by extracting
both sections and running them against fixtures, then for real in a live
v1.8.10 container. No smoke assertion added on purpose — the positive path
needs a /workspace mount smoke does not have, and asserting it there would
repeat the v1.8.0 mistake of a smoke check written against a stage that
does not exist at run time.

Moves the base hash (base-decide folds `cat entrypoint.sh
entrypoint-user.sh`), so this rides along with the next tag's ~40-minute
base rebuild rather than justifying a tag of its own.
2026-08-27 21:52:51 +02:00
joakimp 6891dc32b8 changelog: release v1.8.10 — deploy the scrubber, and name the check that must not be skipped
Publish Docker Image / build-variant-studio (push) Successful in 17m4s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 14s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-variant (push) Successful in 18m45s
Publish Docker Image / build-base (push) Successful in 41m59s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / smoke (push) Successful in 4m55s
Lint / hadolint (push) Successful in 8s
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Audited every component against upstream rather than assuming. Only two moved
since v1.8.9:

- mempalace-toolkit 5b8d78f -> b2b50af (ours): feeder scrubber, the symlink fix
  that stops it refusing to stage, hlc owed-set join, queued-delivery note,
  explicit MEMPALACE_MAILBOX_NOTIFY protocol modes, RFC 003 + fleet-memory docs.
- pi-studio v0.9.48 -> v0.9.52 (upstream, studio variant only): 22 commits, all
  additive — PDF previews, hideable header, contextual side questions. No removals
  or renames. Its pi floor is >=0.84.3 against our exact 0.84.3 pin: satisfied,
  zero headroom, named as a watch item because the next floor bump breaks the
  studio job only, after core has already published.

Verified unchanged, with dates proving they predate the v1.8.9 bake: pi 0.84.3 (=
npm latest), mempalace 3.8.0 (= PyPI latest), pi-atelier v0.8.2, playwright 1.62.1
(no drift this cycle despite floating on latest), pi-fork bf702b4, pi-obsmem
ce9fc98, pi-toolkit 0e1369e, pi-extensions 2022887. No breaking changes anywhere,
so everything can ship in one tag.

The reason for tagging now is not features. The scrubber has been live on exactly
ONE device since this morning; every other device has kept staging unscrubbed
transcripts into a shared palace, and cleanup after the fact is manual redaction
(done twice today, with freed-page residue left behind by choice).

The entry leads with the acceptance check instead of burying it, because this
release's worst failure is silent: the feeder is fail-closed, so a packaging or
path mistake stops the fleet's memory feed and nothing complains — refusing to
stage looks exactly like a quiet session. f0bffd1 exists because that very bug was
real (BASH_SOURCE reports the symlink path, not the target). First client on the
new image must confirm a [scrub] summary line appears, the drawer count moves, and
exit 3 did not fire. Silence is the failure signal, not success.
2026-08-27 17:51:46 +02:00