Compare commits

..

9 Commits

Author SHA1 Message Date
joakimp fabf1274aa docs(changelog): Unreleased section for the smoke node assertion + agent-browser correction
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 19s
Summarises what changed since v1.8.13: the node-major assertion (a bump would
have passed the suite silently), the two-sided verification of the derivation,
and the v1.8.13 agent-browser 0.35.2 -> 0.36.0 correction. No image content
changes; NODE_VERSION still 22.
2026-09-07 21:17:06 +02:00
joakimp 5972a2c535 test+docs: assert the node major in smoke, and correct v1.8.13's agent-browser version
Two findings from a delegated read-only audit of this repo, both verified from the
filesystem before patching.

1. No test asserted the node major, so a node-24 bump would have passed the smoke
   suite SILENTLY. scripts/smoke-test.sh:94 was a bare `run "node" "node --version"`
   — exit-0 and non-empty output only, the printed version compared to nothing —
   while the line above it uses run_expect against $EXPECTED_PI_VERSION for pi. A
   reader skimming the suite would reasonably assume node regressions were covered.
   Worse, this is where the "node v22.23.2 verified" line in the v1.8.13 recreate
   notes came from: printed output, not an assertion.

   Now gated on EXPECTED_NODE_MAJOR, which CI derives from Dockerfile.base's ARG
   NODE_VERSION — the single source of truth (Dockerfile.base:557 is the ONLY hard
   pin in the repo; Dockerfile.variant has no node install at all). That also
   catches a stale cached layer whose node disagrees with the declared ARG.
   Unset => previous behaviour, so this is backward compatible.

   Verified two-sided rather than assumed: the sed derivation yields 22 (empty
   would have silently disabled the assertion, reintroducing the bug); grep -Fq
   "v22." matches v22.23.2; "v24." does NOT match, so a wrong major is caught; and
   "v2." does not prefix-collide. Workflow YAML re-parsed after editing (9 jobs).

2. The v1.8.13 entry claimed "the image's own 0.35.2" for agent-browser. The image
   ships 0.36.0: /usr/lib/node_modules/agent-browser/package.json says version
   0.36.0, engines.node >=24.0.0, and no 0.35.2 exists anywhere in the image. The
   claim was also internally incoherent, contrasting 0.36.0 against a version that
   is not present. Corrected in place with a visible note, since the entry is
   already released. The reasoning survives untouched: the engines floor really is
   vestigial, because /usr/bin/agent-browser is a prebuilt aarch64 ELF invoked
   directly and never through node — which is why 0.36.0 runs fine on 22.23.2.
2026-09-07 21:05:24 +02:00
joakimp aa0fbc5ec0 fix: correct the pi-studio claim — CI publishes v0.9.59, not the v0.9.60-rc.0 label
Lint / actionlint (push) Successful in 17s
Lint / hadolint (push) Successful in 14s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Successful in 4m44s
Publish Docker Image / smoke-studio (push) Successful in 5m8s
Publish Docker Image / build-variant (push) Successful in 15m52s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / build-variant-studio (push) Successful in 21m22s
Measured at the wrong layer during the v1.8.13 audit. I read `ARG
PI_STUDIO_REF=main` in Dockerfile.variant, concluded the release would adopt
main (= v0.9.60-rc.0), set PI_STUDIO_VERSION to that, and wrote a comment plus a
CHANGELOG entry describing deliberate RC adoption. A Dockerfile default cannot
answer "what will CI publish?" when CI overrides it, and it does: build-variant
passes PI_STUDIO_REF=studio_ref and PI_STUDIO_VERSION=studio_tag (lines 598-599
and 787-788), and resolve-versions picks the newest STABLE semver tag via
`^v?[0-9]+\.[0-9]+\.[0-9]+$`, which excludes pre-releases.

Caught by reading run 639's own resolve-versions output rather than the
Dockerfile: studio_tag=v0.9.59, studio_ref=9eed84f = refs/tags/v0.9.59^{}, while
main/v0.9.60-rc.0 is 658536f and never gets built. So published v1.8.13 studio
images carry pi-studio v0.9.59.

ARG restored to `none` rather than pinned to v0.9.59: the local-build default
should not hardcode a tag that goes stale as soon as main moves, which is how the
previous value came to lie. The comment now leads with the override so the next
reader starts at the layer that decides. Upstream's tag-over-main policy is
deliberate (Releases stopped at v0.5.55, main receives half-finished commits), so
adopting an RC from CI would mean changing that filter, not this ARG.

Consequence kept on purpose: the RC's opt-in Studio network binding is in NO
published v1.8.13 image, so it needs no audit this release.

Doc/label-only: base_tag hashes Dockerfile.base + rootfs/** + both entrypoints +
mempalace_toolkit_ref, none of which this touches, so the in-flight base build
(base-ad9faf00f2b2) stays valid and the tag run will reuse it. Verified with CI's
pinned linters: hadolint 2.14.0 exit 0 on both Dockerfiles, actionlint 1.7.7 exit
0, shellcheck 0.10.0 -S error exit 0.
2026-09-06 23:51:00 +02:00
joakimp 702dd71f4c ci: declare workflow_dispatch input types so Gitea renders the dispatch form
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 16s
Gitea (1.26.2) builds the "Run workflow" dialog from each input's `type:`.
With no type declared, the form renders a branch selector and NO input fields,
so a manual run silently takes every default -- and for release_tag: '' that
means env.RELEASE_TAG resolves EMPTY, the variant tag list becomes `<image>:`,
and the run dies on an invalid docker reference only AFTER paying the full base
+ smoke cost (~70 min). Net effect: the `smoke_only` escape hatch documented in
this file's own header has been unreachable from the UI for its entire
existence. Found 2026-09-06 while trying to use it to validate three new smoke
assertions before cutting v1.8.13.

Typed as `string`, deliberately, even though promote_latest/smoke_only read as
booleans: all six consumption sites compare strings against 'true'
(inputs.smoke_only != 'true' at both build-variant gates,
inputs.promote_latest == 'true' at both promote gates) or interpolate into
env.PROMOTE_LATEST. A boolean-typed input yields a real boolean, so `!= 'true'`
would compare across types and could invert a publish gate silently rather than
fail loudly. This keeps the change a pure rendering fix with zero semantic
delta; switching to boolean would require re-auditing all six call sites.

Validated locally with CI's own pinned tools before pushing, because lint is
the only gate on this file: actionlint 1.7.7 exit 0 (clean baseline before the
edit, clean after), shellcheck 0.10.0 -S error exit 0 across all 17 shell
files, and a pyyaml structural check confirming the three inputs still carry
string defaults, the `v*` tag trigger is intact, and all 9 jobs still parse.
2026-09-06 23:40:26 +02:00
joakimp f561acc89a skills: refresh vendored mempalace snapshot a12fe5e -> e9e09d9, re-pin the canary
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Failing after 31m52s
Folded into v1.8.13 at zero marginal cost: the snapshot is hashed into
base_tag, but Dockerfile.base already changed this release, so the ~67 min base
rebuild was already being paid. vendor-mempalace-skill.sh --check reported exit
0 (stale-but-truthful) beforehand, so skipping was sanctioned -- this is the
deliberate call the release checklist asks for. Upstream content: the bare
project-name wing convention and the <harness>@<device> added_by rule, both
downstream of the attribution defect measured on this device 2026-09-06.

The canary re-pin matters more than the refresh. Its old pair ("Provenance is
stamped for you" present / "Attribute what you file yourself" absent) still
PASSED against the new snapshot, so leaving it would have yielded a canary
green on both old and new bytes -- blind to exactly the refresh it exists to
witness, the same false-green family as the pre-v1.8.5 canary. New pair chosen
by measuring direction against both files rather than reading the diff
("Diaries self-heal; plain drawers do not" new=1/old=0; "Agent diaries live in"
new=0/old=1), then tested two-sided: PASS on refreshed bytes, FAIL on the old
bytes recovered from git.

Gates after the change: smoke-test.sh parses, vendor --check exit 0,
check-base-hash exit 0.
2026-09-06 22:31:47 +02:00
joakimp 0d984b1414 changelog: cut v1.8.13 section
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 18s
2026-09-06 22:14:20 +02:00
joakimp adcf56f829 release: audited bumps (pi 0.85.1, mempalace 3.9.0, atelier v0.10.1) + two guards
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 17s
Version audit for the next release. pi 0.84.4 -> 0.85.1, deliberately skipping
0.85.0 (it published internal experimental code and broke SDK imports,
upstream #9132). mempalace 3.8.0 -> 3.9.0. pi-atelier v0.10.0 -> v0.10.1.
PI_STUDIO_VERSION relabelled none -> v0.9.60-rc.0 so the floating main ref's
RC status is visible at docker-inspect time instead of discovered later.
PI_FORK_REF stays floating and adopts e69725c.

The pi bump was verified by running it under a pty in five combinations rather
than by reading the changelog, because this repo has already shipped a version
pair no changelog flagged (atelier < 0.7.1 hangs pi >= 0.84). CPU delta
0.00-0.01s over 5s against a ~5s sustained-CPU hang signature, two-sided via
the atelier sidebar painting identically to the 0.84.4 control.

NODE_VERSION stays 22 on purpose: node 24 is technically safe (pi's five
prebuilt addons are all NAPI, nothing declares a ceiling, agent-browser's
engines.node >=24 is vestigial for the shipped aarch64 ELF), but this release
already moves two minors and bakes an RC, and a node major would leave four
suspects if the image misbehaves. Own release, smoke suite as the gate.

Also corrects a stale claim at the mempalace ARG: synlig serves 3.8.0
server-side, not 3.7.1 (measured over ssh 2026-09-06).

agent-browser volume shadowing: the image has shipped 0.35.2, but every
session on mbp-m1-2020 ran 0.27.0 from a 2026-07-17 hand-install in
~/.pi/npm-global (a VOLUME, at PATH position 2 vs /usr/bin at 8). Third
package hit by this hazard after pi and pi-atelier, so the guard is now
generalised: entrypoint-user.sh retires the copy by moving it aside
(reversible, only when the image ships its own), recreate-sanity-check.sh
asserts resolution under /usr where the volume is real, smoke-test.sh carries
the build-time half and says in the source why it is weak. The real damage was
the stale BUNDLED SKILL (3 skillsets/17.6 KB vs 8/31.5 KB, ten subcommands
undocumented to the agent) - a stale tool errors, a stale skill quietly
teaches wrong commands.

pi-fork capability floor (extensions: []): forks were measured across four
dispatches ignoring their brief, answering in the user's voice, fabricating
self-referential measurements, and once filing a diary entry as agent_name=pi.
Cause is upstream by design - the child gets getHeader()+getBranch(), the
whole active session branch, with the brief as the final user message. Not a
model-capability problem: the same model as the fast profile obeyed the
identical brief perfectly with a fresh session and no inherited context.
extensions: [] runs children with --no-extensions, so the mempalace bridge is
absent and palace writes are impossible by construction (verified by asking a
child to enumerate its tools: read, bash, edit, write). Removes palace writes,
not filesystem writes.
2026-09-06 20:40:02 +02:00
joakimp c8622ece9d skills: correct the credential-incident-response §5 premise about chroma metadata
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 17s
§5 said embedding_metadata.string_value holds "metadata fields only". False,
measured directly: chroma also stores a copy of the document text there, under
key chroma:document. Confirmed with a disposable sentinel drawer (pi@tor-ms22,
2026-08-30): one row in fts_content AND one row in embedding_metadata for the
same drawer.

This was a real mistake in shipped guidance, not a nitpick: this section's own
scanning advice was written to guard against explaining a zero with a
mechanism nobody verified from source, and the section itself did exactly
that -- I downgraded a census to "a floor" on the strength of a metadata-blind
claim I never checked against chroma's actual storage layout. The practical
scan order is unchanged (fts_content is still the direct target, raw bytes are
still the backstop); only the stated REASON for a metadata zero changes: it
needs a different explanation now (key filter, query shape, escaping), not
"structurally absent".

§6's row-gone/bytes-gone claim is upgraded from asserted to measured, same
sentinel: delete_by_source took both fts_content and embedding_metadata 1->0,
raw bytes stayed 4->4 (freed pages persist until VACUUM). Also records the
method that unblocked the measurement: not a better instrument, a disposable
sentinel drawer instead of risking real fleet data.

No image behaviour changes.
2026-09-01 22:37:46 +02:00
joakimp 05843ecfae changelog: reopen an Unreleased section after v1.8.12
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 16s
v1.8.12's release retitled the previous Unreleased heading, leaving the file
with no place to put the next change — so the next contributor either invents a
heading or appends to a released section. The note under it points at the
release checklist step that renames it, so the convention is discoverable from
the file rather than only from AGENTS.md.

Also the first push after moving CI off synlig: lint.yml should now run on
runner-a1 (8 vCPU / 16 GB, Debian 13, upstream Docker CE) instead of the box
that hosts the palace.
2026-09-01 00:29:29 +02:00
9 changed files with 496 additions and 30 deletions
+34 -3
View File
@@ -33,18 +33,39 @@ on:
- 'v*'
workflow_dispatch:
inputs:
# `type:` is REQUIRED for Gitea to render these fields in the "Run
# workflow" dialog. Without it (Gitea 1.26.2) the dispatch form shows a
# branch selector and NO inputs at all, so a manual run silently uses
# every default — which for `release_tag: ''` means RELEASE_TAG resolves
# empty, the variant tag list becomes `<image>:`, and the run dies on an
# invalid reference AFTER paying the full base + smoke cost (~70 min).
# That made the documented `smoke_only` escape hatch below unreachable
# from the UI for its whole existence; found 2026-09-06 trying to use it.
#
# Deliberately `string` and not `boolean`, even though these two read as
# flags: every consumption is a STRING comparison against 'true'
# (`inputs.smoke_only != 'true'` at the build-variant gates,
# `inputs.promote_latest == 'true'` at the promote gates) plus string
# interpolation into env.PROMOTE_LATEST. A boolean-typed input yields a
# real boolean, so `!= 'true'` would compare across types and could
# invert a publish gate rather than fail loudly. Changing the type here
# would mean re-auditing all six call sites; keeping it string is a
# rendering fix with provably zero semantic change.
release_tag:
description: 'Release tag to publish (e.g. v1.0.0). Used only for workflow_dispatch runs.'
required: false
default: ''
type: string
promote_latest:
description: 'Update latest aliases (default true for tag-push, false for manual test runs)'
required: false
default: 'false'
type: string
smoke_only:
description: 'Build base + run both smoke jobs against HEAD, then stop. Publishes nothing. Use to validate smoke assertions without cutting a tag.'
description: 'Build base + run both smoke jobs against HEAD, then stop. Publishes nothing. Use to validate smoke assertions without cutting a tag. Set to the literal string true.'
required: false
default: 'false'
type: string
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
@@ -522,7 +543,12 @@ jobs:
env:
EXPECTED_PI_VERSION: ${{ needs.resolve-versions.outputs.pi_version }}
EXPECTED_MEMPALACE_VERSION: ${{ needs.resolve-versions.outputs.mempalace_version }}
run: bash scripts/smoke-test.sh pi-devbox:smoke
run: |
# Single source of truth for the node major is Dockerfile.base's ARG.
# Asserting the BUILT image matches it also catches a stale cached layer.
EXPECTED_NODE_MAJOR=$(sed -n 's/^ARG NODE_VERSION=\([0-9][0-9]*\).*/\1/p' Dockerfile.base)
export EXPECTED_NODE_MAJOR
bash scripts/smoke-test.sh pi-devbox:smoke
# ── Phase 3b: amd64 smoke for the studio variant ────────────────────
# Additive + independent of the core `smoke` job: gates ONLY
@@ -585,7 +611,12 @@ jobs:
env:
EXPECTED_PI_VERSION: ${{ needs.resolve-versions.outputs.pi_version }}
EXPECTED_MEMPALACE_VERSION: ${{ needs.resolve-versions.outputs.mempalace_version }}
run: bash scripts/smoke-test.sh pi-devbox:smoke-studio
run: |
# Single source of truth for the node major is Dockerfile.base's ARG.
# Asserting the BUILT image matches it also catches a stale cached layer.
EXPECTED_NODE_MAJOR=$(sed -n 's/^ARG NODE_VERSION=\([0-9][0-9]*\).*/\1/p' Dockerfile.base)
export EXPECTED_NODE_MAJOR
bash scripts/smoke-test.sh pi-devbox:smoke-studio
# ── Phase 4: multi-arch publish ─────────────────────────────────────
build-variant:
+223
View File
@@ -11,6 +11,229 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
---
## Unreleased
**A test that was quietly checking nothing, and a version number that was wrong.**
Both found by delegating a read-only audit of this repo to a headless worker
(`pi-toolkit` `bin/pi-task`) and then spot-checking its pointers from the
filesystem — 5 of 5 held, and it also corrected a false premise planted in its
own brief.
**The node major is now asserted, not merely printed.**
`scripts/smoke-test.sh` ran `run "node" "node --version"`, which asserts only
that the binary exists and exits 0 — the printed version was compared to
nothing. The line above it has always used `run_expect` against
`$EXPECTED_PI_VERSION` for `pi`, so the suite *looked* like it covered node.
**A node major bump would have passed the whole smoke suite silently.** Worse,
this is where the "node v22.23.2 verified" line in the v1.8.13 recreate notes
came from: printed output, not an assertion — an expectation stated up front and
then falsified by the check.
Now gated on `EXPECTED_NODE_MAJOR`, which CI derives from `Dockerfile.base`'s
`ARG NODE_VERSION` — the single source of truth, and the *only* hard node pin in
the repo (`Dockerfile.variant` has no node install at all, so the two Dockerfiles
cannot disagree). That also catches a stale cached layer whose node disagrees
with the declared ARG. Unset ⇒ previous behaviour, so nothing breaks for anyone
running the suite by hand.
Verified two-sided, because a silent failure here reintroduces the exact bug it
fixes: the `sed` derivation yields `22` (an empty result would disable the
assertion silently); `grep -Fq "v22."` matches `v22.23.2`; `"v24."` does **not**
match, so a wrong major is caught; `"v2."` does not prefix-collide. The workflow
YAML was re-parsed after editing (9 jobs).
**v1.8.13's agent-browser version was wrong.** That entry said "the image's own
0.35.2". The image ships **0.36.0** — `/usr/lib/node_modules/agent-browser` at
0.36.0 with `engines.node >=24.0.0`, and no 0.35.2 exists anywhere in the image.
The sentence was also internally incoherent, contrasting 0.36.0 against a version
that is not present. Corrected in place with a visible note, since that entry is
already released. **The reasoning survives untouched**: the engines floor really
is vestigial, because `/usr/bin/agent-browser` is a prebuilt aarch64 ELF invoked
directly and never through node — which is exactly why 0.36.0 runs fine on
22.23.2, consistent with the runtime proof collected on 2026-09-07 and with the
retraction of the earlier false "0.36.0 requires node >= 24" alert.
No image content changes: `NODE_VERSION` still 22, no pins moved. This is a test
and a docs correction only.
---
## v1.8.13 — 2026-09-06
**Version audit + three pins moved, one deliberately not moved.** `pi`
0.84.4 -> 0.85.1, `mempalace` 3.8.0 -> 3.9.0, `pi-atelier` v0.10.0 -> v0.10.1.
`PI_FORK_REF=master` stays floating and therefore adopts e69725c. Each rationale
is written at the ARG itself rather than only here, because that is where the
next person doing the audit will be standing.
**Correction, made mid-release while run 639 was building:** the audit
originally recorded a fourth change — "`PI_STUDIO_VERSION` relabelled `none` ->
`v0.9.60-rc.0`, RC adopted deliberately" — and that was wrong. It was measured
at the wrong layer. `resolve-versions` passes BOTH `PI_STUDIO_REF` and
`PI_STUDIO_VERSION` as build-args and selects the newest **stable** semver tag
(its filter `^v?[0-9]+\.[0-9]+\.[0-9]+$` excludes pre-releases), so a Dockerfile
default cannot answer "what will CI publish?". Measured from the run itself:
`studio_tag=v0.9.59`, `studio_ref=9eed84f` (= `refs/tags/v0.9.59^{}`), while
`main`/`v0.9.60-rc.0` is 658536f and is not built. **Published v1.8.13 studio
images therefore contain pi-studio v0.9.59, not the RC**, and the ARG is back at
`none` rather than pinned to a pre-release that goes stale the moment main
moves. Consequence kept deliberately: the RC's opt-in Studio network binding is
absent from every published v1.8.13 image, so it needs no audit for this
release. Adopting an RC from CI would require changing that tag filter, which
exists on purpose — upstream stopped publishing Releases at v0.5.55 but keeps
tagging and pushing to main, so pinning main risked baking half-finished commits.
0.85.0 is SKIPPED on purpose: it shipped internal experimental code and extra
subpaths that broke SDK imports (upstream #9132), and 0.85.1 exists to undo
exactly that. Neither release has a Breaking/Removed changelog heading, the
engine floor is unchanged (>=22.19.0 against the container's 22.23.2), and
runtime deps drop 20 -> 19.
The pi bump was verified by RUNNING it, not by reading about it, because this
repo has already been burned by a version pair that no changelog flagged
(pi-atelier < 0.7.1 hangs pi >= 0.84 at startup with no error). 0.85.1 was
side-installed and driven under a pty in five combinations — each companion
extension plus atelier v0.10.0 AND v0.10.1 — with a CPU delta of 0.00-0.01s
over a 5s window where the known hang signature is ~5s of sustained CPU. The
check was two-sided: the atelier sidebar painted ACTIVITY+WORKSPACE markers
identically to the 0.84.4 control, so "alive" could be distinguished from
"silently absent".
**NODE_VERSION stays 22 — audited, not overlooked.** node 24 is technically
safe: all five prebuilt native addons in pi use NAPI (ABI-stable, no
NODE_MODULE_VERSION lock, no binding.gyp), nothing in the image declares a node
CEILING, and the install is one token (`setup_${NODE_VERSION}.x`). agent-browser
0.36.0 declares `engines.node >=24.0.0`, but that field is vestigial for the
artifact actually shipped: `/usr/bin/agent-browser` is the prebuilt aarch64 ELF
`bin/agent-browser-linux-arm64`, invoked directly and never through node, so npm's
engines floor is never enforced at runtime — verified running under 22.23.2 in
this image. (Corrected 2026-09-07: this paragraph originally said "the image's own
0.35.2 declares the same floor". That was wrong and incoherent — it contrasted
0.36.0 against a 0.35.2 that does not exist in the image. There is exactly one
agent-browser present, `/usr/lib/node_modules/agent-browser` at 0.36.0. The
argument is unaffected; only the version was wrong.) The reason to wait is
attribution, not compatibility — this release already moves pi a minor,
mempalace a minor and bakes a Studio RC, so adding a node major would leave four
suspects if the image misbehaves. Worth doing as its own release with the smoke
suite as the gate. (v22 is in maintenance until 2027-04-30; v24 is Active LTS
to 2026-10-20 and maintained to 2028-04-30, so there is real headroom.)
mempalace's client bump carries a sequencing note that is now also CORRECT: the
comment at the ARG claimed synlig serves 3.7.1 server-side, which was stale.
Measured 2026-09-06 over ssh, synlig's uv tool entry last changed 2026-08-25
and serves 3.8.0. Client 3.9.0 against server 3.8.0 is accepted skew until
synlig's compose stack is redeployed; 3.9.0's headline additions (release
awareness, `task create`/`task launch`) are SERVER-side and stay dark until
then — a client bump alone cannot light them up.
**agent-browser was running 7 weeks stale, and the interesting part is why
nothing noticed.** The image has shipped 0.35.2 since the last base rebuild,
but every session on mbp-m1-2020 was executing 0.27.0 from a 2026-07-17
hand-install: `npm i -g` writes into `~/.pi/npm-global`, which is the
devbox-pi-config VOLUME, and PATH puts that at position 2 against /usr/bin at
position 8. This is the third package hit by that exact hazard (pi itself and
pi-atelier already have guards), so the guard is now generalised instead of
re-invented a fourth time.
The damage was not the binary. It was the BUNDLED SKILL, which is the part an
agent reads: 3 skillsets / 17.6 KB core in 0.27.0 versus 8 skillsets / 31.5 KB
core in 0.35.2, with ten subcommands present in the image and entirely
undocumented to the agent (a11y, browser, data, mcp, page, plugin, read,
selectors, to, webmcp). A stale tool announces itself with an error; a stale
skill just quietly teaches the wrong commands and everything looks fine.
Three changes, at the three places this can be caught:
- `entrypoint-user.sh` retires a volume copy by MOVING it aside (reversible,
same instinct as the settings backups) and only when the image ships its own
copy, so a machine that deliberately hand-installs on an image without one
keeps it. The `bin/` shim is removed too — a dangling symlink would be a
worse failure than a stale version.
- `scripts/recreate-sanity-check.sh` asserts `agent-browser` resolves under
/usr. This is the check that matters, because it runs where the volume is
real.
- `scripts/smoke-test.sh` gets the build-time half, labelled WEAK in the source
for an honest reason: a `docker run` container has an empty config volume, so
it can never see the shadowing it is nominally testing for.
**pi-fork gets a capability floor: `extensions: []`.** Forks were measured
twice (2026-09-01, 2026-09-06, four dispatches) ignoring their brief, answering
in the USER's voice, fabricating self-referential measurements, and once filing
a diary entry as `agent_name=pi` — which landed in `wing_pi`, where a
wing-scoped `diary_read` never sees it.
The cause is upstream and by design, so there is nothing to wait for: the child
is handed `getHeader()+getBranch()`, i.e. the WHOLE active session branch, with
the brief appended as the final user message and the system prompt untouched
(pi-fork `src/index.ts`). In a long session the parent narrative simply
outweighs the task, and the child does the statistically obvious thing — it
continues the story it finds itself inside. Config offers no context knob
(extensions, environment, offline, costFooter, effort profiles only).
Falsified the tempting explanation before acting on it: the failures are NOT a
too-small model. The same model as the `fast` profile (haiku, thinking off)
obeyed the identical brief perfectly when run as
`pi -p --mode json --session-id <fresh> --no-extensions` — correct values,
exact format, no session recap, 3 seconds, $0.012. Model held constant, context
inheritance removed, failure gone.
`extensions: []` is therefore a mechanical guarantee rather than an
instruction: the mempalace bridge is a pi EXTENSION, so a fork child now runs
with `--no-extensions` and cannot write to the shared palace under the parent's
identity. Verified by asking a child to enumerate its own tools: `read, bash,
edit, write` — no `mempalace_*`, no `recall`, no nested `fork`. Two honest
limits, stated so nobody over-trusts this: it removes PALACE writes, not
FILESYSTEM writes (`edit`/`write` remain), and it costs forks their palace
search and recall. Set the key to `null` to restore normal loading.
Smoke asserts the floor is `[]` specifically, not merely falsy — `null` is the
unguarded state, so a "truthy or not" test would pass on exactly the
configuration being guarded against.
**Vendored mempalace skill snapshot refreshed `a12fe5e` -> `e9e09d9`, and the
phrase canary re-pinned with it.** Folded in at zero marginal cost: the
snapshot is hashed into `base_tag`, but `Dockerfile.base` already changed this
release, so the ~67 min base rebuild was already being paid. `--check` reported
exit 0 (stale-but-truthful) beforehand, i.e. skipping was sanctioned — this is
the deliberate decision the checklist asks for, not a drive-by. Upstream content
is the fleet wing-naming convention (bare project names, no `wing_` prefix) and
the `<harness>@<device>` rule for `added_by`, both of which came out of the
attribution defect measured on this device on 2026-09-06.
The canary re-pin is the interesting half. Its old pair — "Provenance is
stamped for you" present, "Attribute what you file yourself" absent — STILL
PASSED against the new snapshot, so leaving it in place would have produced a
canary that is green on both the old and the new bytes: blind to precisely the
refresh it exists to witness, which is the same false-green family the
pre-v1.8.5 canary died of. The replacement pair was picked by MEASURING
direction against both files rather than by reading the diff ("Diaries
self-heal; plain drawers do not" new=1/old=0; "Agent diaries live in"
new=0/old=1) and then tested two-sided: PASS on the refreshed bytes, FAIL on the
old bytes recovered from git. A canary that cannot fail is decoration.
**`credential-incident-response` §5/§6 corrected — a stated mechanism was wrong,
and this is the second time in three days this section named a wrong reason
for a zero.** Docs only.
§5 said `embedding_metadata.string_value` holds "metadata fields only". Measured
false on chroma 1.5.9 with a disposable sentinel drawer (pi@tor-ms22,
2026-08-30): the document text is ALSO there, under key `chroma:document` — one
row in `fts_content` and one in `embedding_metadata` for the same drawer. The
scan order in §5 is unchanged (scan `fts_content` directly, raw bytes as
backstop) but the stated REASON is fixed: a zero from `string_value` needs a
different explanation (key filter, query shape, escaping), not "it's
structurally blind". §6 already warns against explaining a zero with an
unverified mechanism; this was exactly that failure, in the file that carries
the warning.
§6's row-gone/bytes-gone claim is now backed by the same sentinel measurement
rather than asserted: `delete_by_source` took both `fts_content` (1->0) and
`embedding_metadata` (1->0) to zero, while raw bytes stayed 4->4 until VACUUM.
Also records how the measurement got unblocked at all — not a better
instrument, a disposable sentinel drawer instead of testing deletion on real
data.
---
## v1.8.12 — 2026-08-31
**`pi` `0.84.3` → `0.84.4`, and `pi-atelier` `v0.8.2` → `v0.10.0`.** Both audited
+20 -7
View File
@@ -450,13 +450,26 @@ ARG INSTALL_MEMPALACE=true
# the part that should stay manual.
#
# Deployment sequencing note for whoever ships this bump: synlig (the shared
# central palace host) currently serves mempalace 3.7.1 SERVER-SIDE via
# docker-compose.mempalace.yml, which reuses this same devbox image. Bumping
# this ARG changes only the CLIENT version baked into pi-devbox images: it
# introduces client/server skew until synlig's compose stack is separately
# rebuilt/redeployed with the new pin. Not something to code around here —
# just sequence the redeploy.
ARG MEMPALACE_VERSION=3.8.0
# central palace host) serves mempalace 3.8.0 SERVER-SIDE via
# docker-compose.mempalace.yml, which reuses this same devbox image. (Measured
# 2026-09-06 over ssh: synlig's UV_TOOL_DIR mempalace entry last changed
# 2026-08-25 15:33 — this comment previously said 3.7.1, which was stale.)
# Bumping this ARG changes only the CLIENT version baked into pi-devbox
# images: it introduces client/server skew until synlig's compose stack is
# separately rebuilt/redeployed with the new pin. Not something to code around
# here — just sequence the redeploy.
#
# v1.8.13: 3.8.0 -> 3.9.0. Audited: no Breaking/Removed changelog headings.
# Adopted mainly for #2281 (`mempalace_mine` accepts a single conversation
# file again) — though note that does NOT unblock this image's own feeder,
# which was measured to mine DIRECTORIES, not files, so it was never hitting
# that bug. Four behaviour changes ride along and are skew-relevant while
# synlig stays on 3.8.0: hub-forward escaping, an HTTP lock split, similarity
# score semantics, and parsed-output compatibility. 3.9.0-only features
# (release awareness, `task create`/`task launch` MCP tools) are SERVER-side,
# so they stay dark until synlig is redeployed — a client bump alone cannot
# light them up.
ARG MEMPALACE_VERSION=3.9.0
ENV UV_TOOL_DIR=/opt/uv-tools
ENV UV_TOOL_BIN_DIR=/usr/local/bin
RUN if [ "${INSTALL_MEMPALACE}" = "true" ]; then \
+51 -4
View File
@@ -95,7 +95,25 @@ ARG USER_NAME=developer
# `.agents/skills/<group>/` directories were not discovered, and root Markdown
# files such as README.md / AGENTS.md inside a skill dir were reported as
# broken skills unless they declared valid skill frontmatter.
ARG PI_VERSION=0.84.4
#
# v1.8.13: 0.84.4 -> 0.85.1. SKIP 0.85.0 deliberately — it accidentally
# published internal experimental code and extra subpaths, breaking SDK
# imports (upstream #9132); 0.85.1 exists specifically to undo that, with the
# supported SDK and stdio RPC API unchanged. Audited: no Breaking/Removed
# changelog headings in either release, engine floor unchanged (>=22.19.0,
# container runs 22.23.2), runtime deps 20 -> 19. User-visible changes are the
# streaming indicator moving into the editor border and faster fullscreen
# transcript search; no deprecation language anywhere.
#
# Verified EMPIRICALLY rather than from the changelog, because a pi bump has
# hung the TUI before (pi-atelier < 0.7.1 + pi >= 0.84): 0.85.1 was
# side-installed and driven under a pty against all four companion extensions,
# with atelier v0.10.0 AND v0.10.1 — five combinations, each rendering alive
# with a CPU delta of 0.00-0.01s over a 5s window, where the known hang
# signature is ~5s of sustained CPU. Two-sided check: the atelier sidebar
# painted ACTIVITY+WORKSPACE identically to the 0.84.4 control, so the test
# could distinguish "loaded" from "silently absent".
ARG PI_VERSION=0.85.1
ARG PI_TOOLKIT_REF=main
ARG PI_EXTENSIONS_REF=main
# Repo URLs default to the canonical gitea origin but are overridable so a
@@ -147,9 +165,14 @@ ARG PI_OBSMEM_REF=master
# the /opt checkout. Adding an install here would be a no-op that only costs
# build time.
ARG PI_ATELIER_REPO=https://github.com/michaelmjhhhh/pi-atelier.git
ARG PI_ATELIER_REF=v0.10.0
# v1.8.13: v0.10.0 -> v0.10.1. Refactor-only upstream (formatters, tests,
# panel identity); peerDependencies declare pi >=0.84.0, so it spans both the
# old and new pin. Included because it was already exercised: the pty matrix
# for PI_VERSION above ran atelier v0.10.1 against pi 0.85.1 and painted the
# sidebar identically to v0.10.0.
ARG PI_ATELIER_REF=v0.10.1
# Human-readable tag PI_ATELIER_REF was resolved from; recorded as a label.
ARG PI_ATELIER_VERSION=v0.10.0
ARG PI_ATELIER_VERSION=v0.10.1
RUN set -e && \
# git_fetch_ref: clone-equivalent helper that accepts EITHER a branch name
@@ -259,6 +282,30 @@ ARG PI_STUDIO_REF=main
# PI_STUDIO_VERSION is the human-readable tag (e.g. v0.9.36) that PI_STUDIO_REF
# was resolved from; recorded as a label below for at-a-glance identification.
# Only meaningful for the studio variant (default `none` otherwise).
#
# v1.8.13 — READ THIS BEFORE REASONING ABOUT WHICH pi-studio SHIPS. Neither
# default below survives a CI build. `resolve-versions` in
# .gitea/workflows/docker-publish.yml passes BOTH as build-args (studio_ref and
# studio_tag), and it deliberately selects the newest STABLE semver tag: its
# filter is `^v?[0-9]+\.[0-9]+\.[0-9]+$`, which excludes pre-releases. So a
# PUBLISHED v1.8.13 studio image contains pi-studio v0.9.59 (commit 9eed84f,
# = refs/tags/v0.9.59^{}), NOT the v0.9.60-rc.0 that `main` currently points at
# (658536f). The `main` default here only applies to a local `docker build`
# that passes no studio args.
#
# That upstream-tag-over-main choice is intentional and documented at the
# resolve step: pi-studio keeps tagging every version but stopped publishing
# GitHub Releases at v0.5.55 and pushes freely to main, so pinning main risked
# baking half-finished commits that land after a tag.
#
# Corrected here on 2026-09-06 after reading the run-639 resolve-versions
# output: the v1.8.13 audit had recorded "RC adopted deliberately" and set this
# ARG to v0.9.60-rc.0, which was measured at the wrong layer — a Dockerfile
# default cannot answer "what will CI publish?" when CI overrides it. Left at
# `none` rather than pinned to a tag, because a hardcoded pre-release here goes
# stale the moment main moves and would re-tell the same lie to the next reader.
# Consequence worth keeping: the RC's opt-in Studio network binding is NOT in
# any published v1.8.13 image, so it needs no audit for this release.
ARG PI_STUDIO_VERSION=none
RUN if [ "${INSTALL_STUDIO}" = "true" ]; then \
set -e; \
@@ -345,7 +392,7 @@ ARG MEMPALACE_TOOLKIT_REF=main
# no ~67-minute base rebuild. (scripts/check-base-hash.sh scans only
# Dockerfile.base, so no folding into the base hash is required — nor would
# it be correct, since this ARG changes nothing about the base's contents.)
ARG SKILLSET_SNAPSHOT_REF=a12fe5ecc71e60feb24791e3e33571105f1afba7
ARG SKILLSET_SNAPSHOT_REF=e9e09d95f92670536a199fc986dfa24d787f18d1
# Dockerfile.base sets description="pi-devbox — base image (variant-independent)"
# and every variant INHERITS it, so both published images used to advertise
+32
View File
@@ -471,6 +471,38 @@ if command -v pi &>/dev/null; then
done
fi
# ── agent-browser: retire a stale volume copy that shadows the image ───
# Same hazard class as the pi-atelier retirement above, different delivery
# path — and this block exists because that guard did not generalise.
# ~/.pi/npm-global lives on the devbox-pi-config VOLUME, so anything ever
# installed there with `npm i -g` survives every image upgrade, and PATH puts
# it AHEAD of /usr/bin (position 2 vs 8).
#
# Measured on mbp-m1-2020, 2026-09-06: a 2026-07-17 hand-install pinned
# agent-browser 0.27.0 in the volume while the image shipped 0.35.2, so every
# session for ~7 weeks ran a stale CLI. The damaging part was not the binary
# but its BUNDLED SKILL, which is what the agent actually reads: 3 skillsets /
# 17.6 KB core in 0.27.0 vs 8 skillsets / 31.5 KB core in 0.35.2, with ten
# subcommands present in the image and undocumented to the agent (a11y,
# browser, data, mcp, page, plugin, read, selectors, to, webmcp). A stale tool
# announces itself; a stale skill quietly teaches the wrong commands.
#
# MOVE rather than delete (reversible, same instinct as the settings backups
# above), and only when the image ships its own copy — a machine that
# deliberately hand-installs agent-browser on an image WITHOUT one keeps it.
_ab_vol="$HOME/.pi/npm-global/lib/node_modules/agent-browser"
if [ -d "$_ab_vol" ] && [ -d /usr/lib/node_modules/agent-browser ]; then
_ab_park="$HOME/.pi/npm-global/.retired-agent-browser-$(date +%Y%m%d-%H%M%S)"
if mkdir -p "$_ab_park" 2>/dev/null && mv "$_ab_vol" "$_ab_park/" 2>/dev/null; then
# The bin shim is what PATH actually hits; leaving it behind would give a
# dangling symlink, which is a worse failure than a stale version.
rm -f "$HOME/.pi/npm-global/bin/agent-browser" 2>/dev/null || true
echo "agent-browser: retired stale volume copy -> ${_ab_park} (image copy now wins; delete the parked dir when satisfied)"
else
echo "WARN: agent-browser: stale volume copy at $_ab_vol shadows the image copy and could not be moved; retire it by hand"
fi
fi
# ── pi-studio: optional loopback bridge (opt-in) ──────────────────────
# pi-studio binds its server to 127.0.0.1 inside the container, which a
# published Docker port cannot reach. When STUDIO_EXPOSE is truthy (set in
@@ -116,14 +116,24 @@ shared palace as incident response. High blast radius, low actual benefit.
## 5. Finding a secret in a Chroma palace — three targets, in this order
1. `embedding_fulltext_search_content.c0` — **where document text actually is**
2. `embedding_metadata.string_value` — metadata fields only
1. `embedding_fulltext_search_content.c0` — document text
2. `embedding_metadata.string_value` — metadata fields, **and a second copy of
the document text** under key `chroma:document`
3. raw byte scan of every `*.sqlite3` — backstop, covers FTS pages and free space
Scanning only (2) is the classic false clean: hundreds of thousands of rows,
zero hits, and the secret sitting in (1) the whole time. Semantic search proves
nothing about absence — it returns top-k. For completeness, enumerate by filing
window (`list_drawers(since=T, before=T+1m)`), since one mine shares a minute.
**Correction, measured on chroma 1.5.9 with a sentinel drawer:** one row in (1)
AND one row in (2) for the same drawer, so **(2) is not structurally
content-blind** — an earlier version of this section said it held "metadata
fields only", and that was wrong. Scan (1) and (3) regardless: (1) is the direct
target. But if a `string_value` query returns zero for a value you know is in a
drawer, the cause is a key filter, a query shape or escaping — *not* structural
absence, and the difference matters because the false explanation is what makes
the zero feel safe. See §6: do not explain a zero with a mechanism you have not
read from source.
Semantic search proves nothing about absence — it returns top-k. For
completeness, enumerate by filing window (`list_drawers(since=T, before=T+1m)`),
since one mine shares a minute.
Value-agnostic sweeps (uuid / 40-hex / `NAME=VALUE`) drown in false positives at
fleet scale — 608 candidates, mostly session UUIDs and git SHAs. Name-anchoring
@@ -199,10 +209,16 @@ extractor and making it look as strong as the union — a self-test artifact tha
has already fooled an agent here. And never gate on `$?` when the tool has a
lock-skip or no-op path that also exits 0; judge the reported line.
**Row-gone is not bytes-gone.** A correct sqlite `DELETE` leaves the payload in
freelist pages until `VACUUM`, so deletion effectiveness is *two* numbers: rows
removed, and a raw byte scan of the `.sqlite3`. One aggregate figure reported as
"erased" has only measured "unretrievable".
**Row-gone is not bytes-gone.** Measured, same sentinel drawer: after
`delete_by_source` the row count went 1 -> 0 in *both* the FTS content table and
`embedding_metadata`, while the raw byte count stayed 4 -> 4 — sqlite does not
zero freed pages, so the payload sits in free space until `VACUUM`. Deletion
effectiveness is therefore *two* numbers, and each direction has a trap: one
aggregate figure reported as "erased" has only measured "unretrievable", while a
raw byte scan used as the acceptance gate reads a CORRECT, complete deletion as a
failure. (Note how this was measured: the blocker was never a better instrument,
it was the subject — file your own disposable sentinel and delete that, instead
of testing deletion on real data.)
## 7. Choosing scopes: derive them from measured consumers
@@ -587,9 +587,31 @@ Two consequences worth internalising:
### Wings
Wings are top-level categories, typically one per project or domain:
- Named after the project directory (e.g., `cli_utils`, `opencode_devbox`)
- Agent diaries live in `wing_<agent_name>` (e.g., `wing_orchestrator`, `wing_pi`)
Wings are top-level categories, typically one per project or domain.
**NAMING CONVENTION — decided 2026-09-06 by Joakim: bare project names, no `wing_`
prefix.** `home-network`, `pi-devbox`, `mempalace-toolkit` — *not* `wing_pi-devbox`. The
mass is already there (`pi-devbox` 2061 drawers vs `wing_pi-devbox` 25), and a prefix
present on some wings and absent on others turns every read into a guess about which
spelling holds the content.
- Named after the project directory or domain (e.g., `cli_utils`, `home-network`)
- **Always pass `wing` explicitly to `diary_write`.** Omitting it defaults to
`wing_{agent_name}`, which mints or feeds a *parallel* wing — this tool default, not
anyone's sloppiness, is the mechanism that produced the drift. Measured harm
(2026-09-06, `pi@mbp-m1-2020`): a diary entry written with `agent_name=pi` and no
`wing` landed in `wing_pi` while that agent's history lives in `pi-devbox`, so a
`diary_read` scoped to `pi-devbox` showed **no trace of it**. A wing-scoped read that
silently returns an incomplete history is the worst failure mode a memory store has.
- **Legacy `wing_*` wings are frozen and documented, not renamed.** `wing_conversations`
(written by the session feeders), `wing_pi`, `wing_pi-devbox`, `wing_pi-tor-ms22`,
`wing_pi-devbox-emb7kj`, `wing_mempalace`, `wing_orchestrator`, `wing_code` all still
hold real content. **When searching for history, check both spellings** — this is the
practical cost of the drift and it does not go away by decree.
- If a migration is ever done, the acceptance criterion must be at the **relationship**
level: chunk ids still resolve to their parent, and `diary_read` returns the same entry
set before and after. Per-wing drawer counts can look correct while the relationships
underneath are broken, because a count query never touches them.
#### Shared palace: multiple harnesses, and possibly multiple machines
@@ -609,7 +631,7 @@ Zechner's pi-coding-agent). Implications:
When the palace is **central** (shared across machines), these further things apply:
- **Check which machine a conversation came from.** Transcripts are fed per device, so `source_path` reads `…/mempalace-feed/<device>/pi_<uuid>.jsonl` while the displayed `source_file` is only the basename. One search can legitimately return hits from several machines at once — look at the device segment before attributing a decision to *this* project.
- **Provenance is stamped for you — leave it alone.** Drawers carry `device` and `agent_kind` metadata (plus `device_source`/`agent_kind_source` recording *how* each was determined, so an inference is never mistaken for a fact). You do **not** set these, and you no longer set `added_by` either: the pi bridge defaults the writer field to `<harness>@<device>` on `add_drawer`/`checkpoint`/`mine`/`event_append`/`artifact_put`, and prefixes diary entries with `HOST:<device>|`, from host-supplied `$MEMPALACE_PI_DEVICE`. RFC 001 §7.3.2 ranks "agent stamps it via a skill instruction" as the *worst possible* place for exactly the reason you would expect — it is per-call boilerplate that gets forgotten, and it did: the agent who wrote the previous version of this bullet then filed its own provenance drawer as `added_by=checkpoint`. **Confirm the bridge in your image actually stamps before trusting it:** the extension is baked at image build time, so a container on an image older than the stamping commit (pi-devbox < v1.8.7) stamps nothing while still satisfying both gates — the env vars are set and the code is simply absent. Check with `grep -c MEMPALACE_PI_DEVICE "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)"`; zero means keep passing `added_by="<harness>@<device>"` and a manual `HOST:<device>|` diary prefix until the container is recreated on a newer image. Two things remain yours: pass `source_drawer_id` on `kg_add` (triples have no provenance field, so that pointer is the only path back to a device), and pass an explicit `added_by` **only** when deliberately filing on behalf of another device. Never invent values for `device`/`agent_kind`/`origin_device` — a fabricated value is worse than a blank, because it silently corrupts a future merge.
- **Provenance is stamped for you — leave it alone.** Drawers carry `device` and `agent_kind` metadata (plus `device_source`/`agent_kind_source` recording *how* each was determined, so an inference is never mistaken for a fact). You do **not** set these, and you no longer set `added_by` either: the pi bridge defaults the writer field to `<harness>@<device>` on `add_drawer`/`checkpoint`/`mine`/`event_append`/`artifact_put`, and prefixes diary entries with `HOST:<device>|`, from host-supplied `$MEMPALACE_PI_DEVICE`. RFC 001 §7.3.2 ranks "agent stamps it via a skill instruction" as the *worst possible* place for exactly the reason you would expect — it is per-call boilerplate that gets forgotten, and it did: the agent who wrote the previous version of this bullet then filed its own provenance drawer as `added_by=checkpoint`. **Confirm the bridge in your image actually stamps before trusting it:** the extension is baked at image build time, so a container on an image older than the stamping commit (pi-devbox < v1.8.7) stamps nothing while still satisfying both gates — the env vars are set and the code is simply absent. Check with `grep -c MEMPALACE_PI_DEVICE "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)"`; zero means keep passing `added_by="<harness>@<device>"` and a manual `HOST:<device>|` diary prefix until the container is recreated on a newer image. Two things remain yours: pass `source_drawer_id` on `kg_add` (triples have no provenance field, so that pointer is the only path back to a device), and pass an explicit `added_by` **only** when deliberately filing on behalf of another device — and when you do, it **must** be `<harness>@<device>`. A bare nickname (`pi-devbox-claude`) has no `@device` to parse, so `agent_at_device` cannot attribute it and the drawer is unattributable *by rule*, not by lag: it survives every future stamp run with no `device`, and on a shared palace a device-less drawer is one nobody can later scope, audit or clean up per machine. Measured 2026-09-06: 11 drawers on `tor-ms22` were filed this way — including the credential rows, i.e. exactly where "which machine measured this?" matters most — by an agent that had passed its own chosen nickname on every call. Its *diary* entries escaped, because `HOST:<device>|` in the AAAK text recovers the device. **Diaries self-heal; plain drawers do not.** The safest habit is the one above: pass nothing and let the bridge stamp. Never invent values for `device`/`agent_kind`/`origin_device` — a fabricated value is worse than a blank, because it silently corrupts a future merge.
- **Metadata is invisible to search — so check the text, not the fields.** `search` results are built from a fixed key list and `diary_read` returns content, so neither ever shows `device`/`added_by`. Only `mempalace_get_drawer` reveals them. This is why diary entries carry an in-text `HOST:<device>` marker: it is the only attribution a reader actually sees. **A diary entry with no `HOST:` marker predates the convention and may be from any machine — do not assume it is this one's history.**
- **Mined drawers carry the MINE date, not the session date.** When history is imported, or re-mined on the palace host, `filed_at`/`created_at` is the *import* time — so sorting by them does not give chronological order. Real session time is recoverable from the UUIDv7 in `pi_<uuid>.jsonl`: the first 12 hex digits are milliseconds since the epoch (and UUIDv7 sorts lexicographically in time order, so a plain filename sort is already chronological). Agent-authored drawers and diaries have no such backdoor — for those `filed_at` is the only chronology, which is why it must never be restamped.
- **Beware the timezone mismatch when you combine those.** Palace `filed_at`/`created_at` are naive timestamps in the palace host's local time, while a UUIDv7 decodes to UTC. Comparing them directly introduces a silent offset (2 h for a CEST host). Normalise before drawing conclusions about ordering.
+28
View File
@@ -374,6 +374,34 @@ if [ -f "$HOME/.pi/agent/settings.json" ]; then
fi
fi
# ── agent-browser must resolve to the image, not the config volume ────
# The same volume-shadowing hazard already asserted for pi (above) and
# pi-atelier (just now), for the third package it has bitten. This check
# belongs HERE rather than only in smoke-test.sh: a build-time container has an
# empty ~/.pi/npm-global, so smoke-test can never see the stale copy that a
# real recreate inherits. Measured instance: 0.27.0 from 2026-07-17 shadowed
# the image's 0.35.2 for ~7 weeks on mbp-m1-2020, silently supplying an older
# BUNDLED SKILL (3 skillsets vs 8) — the agent read the stale instructions
# without any version mismatch ever being surfaced.
AB_PATH=$(command -v agent-browser 2>/dev/null || true)
if [ -z "$AB_PATH" ]; then
warn "agent-browser not on PATH (expected in v1.6.0+ images; skipping shadow check)"
else
AB_REAL=$(readlink -f "$AB_PATH" 2>/dev/null || echo "$AB_PATH")
AB_VER=$(agent-browser --version 2>/dev/null | head -n1)
case "$AB_REAL" in
/usr/*)
pass "agent-browser resolves to the image copy (${AB_VER:-version unknown})"
;;
*)
fail "agent-browser resolves to $AB_REAL (${AB_VER:-version unknown}) — a ~/.pi/npm-global VOLUME copy is shadowing the image; the entrypoint retirement guard did not run or could not move it"
;;
esac
if [ -d "$HOME/.pi/npm-global/lib/node_modules/agent-browser" ]; then
fail "stale agent-browser still present in the ~/.pi/npm-global volume (entrypoint guard did not retire it)"
fi
fi
# ── pi <-> pi-atelier compatibility floor ─────────────────────────────
# atelier < 0.7.1 wraps pi's private TUI renderer in a way that recurses under
# pi >= 0.84: pi hangs at startup burning CPU, with no error message. atelier's
+56 -2
View File
@@ -5,6 +5,7 @@
#
# Verifies:
# - pi binary present and (if EXPECTED_PI_VERSION set) matches CI's resolved version
# - node MAJOR matches Dockerfile.base's ARG NODE_VERSION (if EXPECTED_NODE_MAJOR set)
# - mempalace core matches the audited pin (if EXPECTED_MEMPALACE_VERSION set)
# - new v1.0.0 base additions (pandoc, graphviz, imagemagick, yq, tealdeer)
# - typst PDF engine for pandoc (v1.4.0) — `pandoc --pdf-engine=typst`
@@ -91,7 +92,18 @@ if [ -n "${EXPECTED_PI_VERSION:-}" ]; then
else
run "pi" "pi --version"
fi
run "node" "node --version"
# Until 2026-09-07 this was a bare `run "node" "node --version"`, which asserts
# only that the binary exists and exits 0 — the printed version was never
# compared to anything. A node major bump would therefore have passed this suite
# SILENTLY, while a reader skimming it would reasonably assume node regressions
# were covered. EXPECTED_NODE_MAJOR closes that: CI derives it from
# Dockerfile.base's ARG NODE_VERSION (the single source of truth), so this also
# catches a stale cached layer whose node does not match the declared ARG.
if [ -n "${EXPECTED_NODE_MAJOR:-}" ]; then
run_expect "node major matches Dockerfile ARG" "node --version" "v${EXPECTED_NODE_MAJOR}."
else
run "node" "node --version"
fi
run "git" "git --version"
run "aws" "aws --version"
run "uv" "uv --version"
@@ -596,7 +608,21 @@ exec_test "mempalace skill linked (fallback)" 'test -L $HOME/.agents/skills
# This assertion is kept because it is orthogonal and free: it pins content,
# not provenance, so it still catches a re-vendored snapshot whose ref was
# bumped correctly but whose bytes came from the wrong place.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Provenance is stamped for you" "$f" && ! grep -q "Attribute what you file yourself" "$f" && echo ok'
#
# v1.8.13: RE-PINNED on refresh a12fe5e -> e9e09d9, which is the whole point of
# the mechanism — the previous pair ("Provenance is stamped for you" present /
# "Attribute what you file yourself" absent) still passed against the NEW
# snapshot, so leaving it would have produced a canary that is green on both the
# old and the new bytes, i.e. blind to precisely the refresh it exists to
# witness. Same false-green family as the pre-v1.8.5 canary this comment warns
# about. The replacement pair was chosen by MEASURING direction against both
# files rather than by reading the diff: "Diaries self-heal; plain drawers do
# not" is new=1/old=0, "Agent diaries live in" is new=0/old=1 — so each string
# discriminates on its own and the pair still fails loudly in BOTH directions
# (forgotten bump AND re-vendored stale snapshot). Upstream content behind this
# refresh: the bare project-name wing convention and the <harness>@<device>
# added_by rule.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Diaries self-heal; plain drawers do not" "$f" && ! grep -q "Agent diaries live in" "$f" && echo ok'
# Link TARGETS, not just link existence: with no skillset mounted (as here) the
# baked tree must be what resolves, for all four vendored skills.
exec_test "vendored skills resolve to the baked tree (no skillset mounted)" \
@@ -718,6 +744,34 @@ exec_test "pi-atelier registered in packages[] (TUI sidebar)" \
exec_test "pi-atelier registered from /opt, not npm: (volume-shadowing guard)" \
'jq -e "((.packages // []) | any((type == \"string\") and endswith(\"/pi-atelier\"))) and (((.packages // []) | any(. == \"npm:pi-atelier\")) | not)" $HOME/.pi/agent/settings.json'
# agent-browser: the third package hit by ~/.pi/npm-global volume shadowing
# (after pi itself and pi-atelier). This build-time check is deliberately WEAK
# and says so: a `docker run` container has an EMPTY config volume, so it can
# only prove the image ships a sane copy and nothing in the image itself
# shadows it. The check that actually bites lives in
# recreate-sanity-check.sh, which runs where the volume is real — that is
# where a 7-week-old 0.27.0 was caught shadowing 0.35.2 on 2026-09-06.
exec_test "agent-browser resolves under /usr (volume-shadowing guard, build-time half)" '
p=$(command -v agent-browser) || { echo "agent-browser not on PATH" >&2; exit 1; }
r=$(readlink -f "$p")
echo "resolved=[$r] version=[$(agent-browser --version 2>/dev/null | head -n1)]" >&2
case "$r" in /usr/*) ;; *) exit 1 ;; esac
test ! -d "$HOME/.pi/npm-global/lib/node_modules/agent-browser" || exit 1
echo ok
'
# pi-fork capability floor. `extensions: []` makes a fork child run with
# --no-extensions, which is the only MECHANICAL guarantee that a fork cannot
# file drawers or diary entries under the parent's identity — the mempalace
# bridge is an extension, so removing extensions removes the write path.
# Asserted because it is a security-shaped default that a settings merge or a
# hand-edit could silently drop, and its absence is invisible until a fork
# writes to the shared palace as you (measured twice: 2026-09-01, 2026-09-06).
# Deliberately compares to [] and not "is falsy": null means "load normal
# extensions", i.e. exactly the unguarded state this asserts against.
exec_test "pi-fork extensions floor is [] (forks cannot write to the palace)" \
'jq -e ".[\"pi-fork\"].extensions == []" $HOME/.pi/agent/settings.json'
# ── /tmp/sshcm directory created by entrypoint ────────────────────────
exec_test "/tmp/sshcm dir mode 700 (ssh ControlMaster)" \
'test -d /tmp/sshcm && [ "$(stat -c %a /tmp/sshcm)" = "700" ] && echo ok'