Compare commits

...

16 Commits

Author SHA1 Message Date
joakimp aa0fbc5ec0 fix: correct the pi-studio claim — CI publishes v0.9.59, not the v0.9.60-rc.0 label
Lint / actionlint (push) Successful in 17s
Lint / hadolint (push) Successful in 14s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Successful in 4m44s
Publish Docker Image / smoke-studio (push) Successful in 5m8s
Publish Docker Image / build-variant (push) Successful in 15m52s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / build-variant-studio (push) Successful in 21m22s
Measured at the wrong layer during the v1.8.13 audit. I read `ARG
PI_STUDIO_REF=main` in Dockerfile.variant, concluded the release would adopt
main (= v0.9.60-rc.0), set PI_STUDIO_VERSION to that, and wrote a comment plus a
CHANGELOG entry describing deliberate RC adoption. A Dockerfile default cannot
answer "what will CI publish?" when CI overrides it, and it does: build-variant
passes PI_STUDIO_REF=studio_ref and PI_STUDIO_VERSION=studio_tag (lines 598-599
and 787-788), and resolve-versions picks the newest STABLE semver tag via
`^v?[0-9]+\.[0-9]+\.[0-9]+$`, which excludes pre-releases.

Caught by reading run 639's own resolve-versions output rather than the
Dockerfile: studio_tag=v0.9.59, studio_ref=9eed84f = refs/tags/v0.9.59^{}, while
main/v0.9.60-rc.0 is 658536f and never gets built. So published v1.8.13 studio
images carry pi-studio v0.9.59.

ARG restored to `none` rather than pinned to v0.9.59: the local-build default
should not hardcode a tag that goes stale as soon as main moves, which is how the
previous value came to lie. The comment now leads with the override so the next
reader starts at the layer that decides. Upstream's tag-over-main policy is
deliberate (Releases stopped at v0.5.55, main receives half-finished commits), so
adopting an RC from CI would mean changing that filter, not this ARG.

Consequence kept on purpose: the RC's opt-in Studio network binding is in NO
published v1.8.13 image, so it needs no audit this release.

Doc/label-only: base_tag hashes Dockerfile.base + rootfs/** + both entrypoints +
mempalace_toolkit_ref, none of which this touches, so the in-flight base build
(base-ad9faf00f2b2) stays valid and the tag run will reuse it. Verified with CI's
pinned linters: hadolint 2.14.0 exit 0 on both Dockerfiles, actionlint 1.7.7 exit
0, shellcheck 0.10.0 -S error exit 0.
2026-09-06 23:51:00 +02:00
joakimp 702dd71f4c ci: declare workflow_dispatch input types so Gitea renders the dispatch form
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 16s
Gitea (1.26.2) builds the "Run workflow" dialog from each input's `type:`.
With no type declared, the form renders a branch selector and NO input fields,
so a manual run silently takes every default -- and for release_tag: '' that
means env.RELEASE_TAG resolves EMPTY, the variant tag list becomes `<image>:`,
and the run dies on an invalid docker reference only AFTER paying the full base
+ smoke cost (~70 min). Net effect: the `smoke_only` escape hatch documented in
this file's own header has been unreachable from the UI for its entire
existence. Found 2026-09-06 while trying to use it to validate three new smoke
assertions before cutting v1.8.13.

Typed as `string`, deliberately, even though promote_latest/smoke_only read as
booleans: all six consumption sites compare strings against 'true'
(inputs.smoke_only != 'true' at both build-variant gates,
inputs.promote_latest == 'true' at both promote gates) or interpolate into
env.PROMOTE_LATEST. A boolean-typed input yields a real boolean, so `!= 'true'`
would compare across types and could invert a publish gate silently rather than
fail loudly. This keeps the change a pure rendering fix with zero semantic
delta; switching to boolean would require re-auditing all six call sites.

Validated locally with CI's own pinned tools before pushing, because lint is
the only gate on this file: actionlint 1.7.7 exit 0 (clean baseline before the
edit, clean after), shellcheck 0.10.0 -S error exit 0 across all 17 shell
files, and a pyyaml structural check confirming the three inputs still carry
string defaults, the `v*` tag trigger is intact, and all 9 jobs still parse.
2026-09-06 23:40:26 +02:00
joakimp f561acc89a skills: refresh vendored mempalace snapshot a12fe5e -> e9e09d9, re-pin the canary
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Failing after 31m52s
Folded into v1.8.13 at zero marginal cost: the snapshot is hashed into
base_tag, but Dockerfile.base already changed this release, so the ~67 min base
rebuild was already being paid. vendor-mempalace-skill.sh --check reported exit
0 (stale-but-truthful) beforehand, so skipping was sanctioned -- this is the
deliberate call the release checklist asks for. Upstream content: the bare
project-name wing convention and the <harness>@<device> added_by rule, both
downstream of the attribution defect measured on this device 2026-09-06.

The canary re-pin matters more than the refresh. Its old pair ("Provenance is
stamped for you" present / "Attribute what you file yourself" absent) still
PASSED against the new snapshot, so leaving it would have yielded a canary
green on both old and new bytes -- blind to exactly the refresh it exists to
witness, the same false-green family as the pre-v1.8.5 canary. New pair chosen
by measuring direction against both files rather than reading the diff
("Diaries self-heal; plain drawers do not" new=1/old=0; "Agent diaries live in"
new=0/old=1), then tested two-sided: PASS on refreshed bytes, FAIL on the old
bytes recovered from git.

Gates after the change: smoke-test.sh parses, vendor --check exit 0,
check-base-hash exit 0.
2026-09-06 22:31:47 +02:00
joakimp 0d984b1414 changelog: cut v1.8.13 section
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 18s
2026-09-06 22:14:20 +02:00
joakimp adcf56f829 release: audited bumps (pi 0.85.1, mempalace 3.9.0, atelier v0.10.1) + two guards
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 17s
Version audit for the next release. pi 0.84.4 -> 0.85.1, deliberately skipping
0.85.0 (it published internal experimental code and broke SDK imports,
upstream #9132). mempalace 3.8.0 -> 3.9.0. pi-atelier v0.10.0 -> v0.10.1.
PI_STUDIO_VERSION relabelled none -> v0.9.60-rc.0 so the floating main ref's
RC status is visible at docker-inspect time instead of discovered later.
PI_FORK_REF stays floating and adopts e69725c.

The pi bump was verified by running it under a pty in five combinations rather
than by reading the changelog, because this repo has already shipped a version
pair no changelog flagged (atelier < 0.7.1 hangs pi >= 0.84). CPU delta
0.00-0.01s over 5s against a ~5s sustained-CPU hang signature, two-sided via
the atelier sidebar painting identically to the 0.84.4 control.

NODE_VERSION stays 22 on purpose: node 24 is technically safe (pi's five
prebuilt addons are all NAPI, nothing declares a ceiling, agent-browser's
engines.node >=24 is vestigial for the shipped aarch64 ELF), but this release
already moves two minors and bakes an RC, and a node major would leave four
suspects if the image misbehaves. Own release, smoke suite as the gate.

Also corrects a stale claim at the mempalace ARG: synlig serves 3.8.0
server-side, not 3.7.1 (measured over ssh 2026-09-06).

agent-browser volume shadowing: the image has shipped 0.35.2, but every
session on mbp-m1-2020 ran 0.27.0 from a 2026-07-17 hand-install in
~/.pi/npm-global (a VOLUME, at PATH position 2 vs /usr/bin at 8). Third
package hit by this hazard after pi and pi-atelier, so the guard is now
generalised: entrypoint-user.sh retires the copy by moving it aside
(reversible, only when the image ships its own), recreate-sanity-check.sh
asserts resolution under /usr where the volume is real, smoke-test.sh carries
the build-time half and says in the source why it is weak. The real damage was
the stale BUNDLED SKILL (3 skillsets/17.6 KB vs 8/31.5 KB, ten subcommands
undocumented to the agent) - a stale tool errors, a stale skill quietly
teaches wrong commands.

pi-fork capability floor (extensions: []): forks were measured across four
dispatches ignoring their brief, answering in the user's voice, fabricating
self-referential measurements, and once filing a diary entry as agent_name=pi.
Cause is upstream by design - the child gets getHeader()+getBranch(), the
whole active session branch, with the brief as the final user message. Not a
model-capability problem: the same model as the fast profile obeyed the
identical brief perfectly with a fresh session and no inherited context.
extensions: [] runs children with --no-extensions, so the mempalace bridge is
absent and palace writes are impossible by construction (verified by asking a
child to enumerate its tools: read, bash, edit, write). Removes palace writes,
not filesystem writes.
2026-09-06 20:40:02 +02:00
joakimp c8622ece9d skills: correct the credential-incident-response §5 premise about chroma metadata
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 17s
§5 said embedding_metadata.string_value holds "metadata fields only". False,
measured directly: chroma also stores a copy of the document text there, under
key chroma:document. Confirmed with a disposable sentinel drawer (pi@tor-ms22,
2026-08-30): one row in fts_content AND one row in embedding_metadata for the
same drawer.

This was a real mistake in shipped guidance, not a nitpick: this section's own
scanning advice was written to guard against explaining a zero with a
mechanism nobody verified from source, and the section itself did exactly
that -- I downgraded a census to "a floor" on the strength of a metadata-blind
claim I never checked against chroma's actual storage layout. The practical
scan order is unchanged (fts_content is still the direct target, raw bytes are
still the backstop); only the stated REASON for a metadata zero changes: it
needs a different explanation now (key filter, query shape, escaping), not
"structurally absent".

§6's row-gone/bytes-gone claim is upgraded from asserted to measured, same
sentinel: delete_by_source took both fts_content and embedding_metadata 1->0,
raw bytes stayed 4->4 (freed pages persist until VACUUM). Also records the
method that unblocked the measurement: not a better instrument, a disposable
sentinel drawer instead of risking real fleet data.

No image behaviour changes.
2026-09-01 22:37:46 +02:00
joakimp 05843ecfae changelog: reopen an Unreleased section after v1.8.12
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 16s
v1.8.12's release retitled the previous Unreleased heading, leaving the file
with no place to put the next change — so the next contributor either invents a
heading or appends to a released section. The note under it points at the
release checklist step that renames it, so the convention is discoverable from
the file rather than only from AGENTS.md.

Also the first push after moving CI off synlig: lint.yml should now run on
runner-a1 (8 vCPU / 16 GB, Debian 13, upstream Docker CE) instead of the box
that hosts the palace.
2026-09-01 00:29:29 +02:00
joakimp a2846a5f7e release: adopt pi 0.84.4 + pi-atelier v0.10.0, and fix the doc claim the pi bump invalidates
Lint / hadolint (push) Successful in 11s
Lint / actionlint (push) Successful in 21s
Publish Docker Image / resolve-versions (push) Successful in 19s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Successful in 59m39s
Publish Docker Image / smoke-studio (push) Successful in 5m22s
Publish Docker Image / smoke (push) Successful in 17m53s
Publish Docker Image / build-variant-studio (push) Successful in 17m14s
Publish Docker Image / build-variant (push) Successful in 18m22s
Publish Docker Image / promote-base-latest (push) Successful in 12s
Publish Docker Image / update-description (push) Successful in 20s
pi 0.84.3 -> 0.84.4 (no Breaking Changes / Removed heading in that section,
grepped). Adopted for three fixes that land on machinery this fleet runs:
#6879 (large tool results crossing the auto-compaction threshold were sent to
the provider before compacting), #8345 (a resumed session corrupted its next
appended entry when the JSONL lacked a trailing newline -- that file is the
memory feeder's input; measured 49/49 clean here beforehand), and #8537
(triggerTurn:false messages sent mid-run were inserted between a tool call and
its result). The mempalace mailbox is outside #8537's precondition: it delivers
at agent_settled with deliverAs:"steer" and no triggerTurn, and 0.84.4 leaves
the documented steer semantics unchanged.

pi-atelier v0.8.2 -> v0.10.0: two minor releases, both UI-only, no BREAKING
notice. v0.9.0 raises its minimum pi to 0.84.0 and, unlike the
0.7.1-under-pi-0.84 startup-hang precedent, encodes it in peerDependencies
(>=0.84.0). Satisfied by PI_VERSION=0.84.4. Both executable floors compare with
sort -V, so 0.10.0 >= 0.7.1 evaluates correctly.

docs/observational-memory.md: pi's own compaction.md gained one paragraph in
0.84.4 -- autoCompact is now also checked mid-run, after a tool batch's results
are appended. Our text said compaction is checked only when pi goes idle and so
"never interrupts a turn"; that was only ever true of the OM trigger. The
section now states both entry points into session_before_compact and the
diagram carries the second edge (mermaid checker re-run: 6 blocks, 44 labels,
0 soft-wrapped, no cut glyphs at 1280px and 800px).

README: the version-pin table had been wrong since v1.8.6 -- 93f986e moved
ARG PI_VERSION to 0.84.3 and MEMPALACE_VERSION to 3.8.0 and neither table row,
so it advertised pi 0.84.2 / mempalace 3.7.1. Corrected, plus the
--expected-version example that would now fail against a 0.84.4 image.

CHANGELOG: Unreleased retitled v1.8.12 (2026-08-31) with the audits above and a
dependency-audit table -- every other component measured SAME (skillset
snapshot --check OK at a12fe5e, 0 commits since baked).
2026-08-31 07:08:41 +02:00
joakimp 58c22afb04 skills: the fingerprint advice was missing its precondition, and the skill had no section on proving absence
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 23s
Docs only; no image behaviour changes.

WHY THIS AND NOT A PRIVATE NOTE. pi@emb-7kj4vr4g reported itself for printing
sha256[:8] fingerprints of GIT_USER_EMAIL, GIT_USER_NAME and HOST_SSH_USER, and
wrote a private rule forbidding it. It had not broken a rule. It followed §2 of
this skill as written, and §2 is incomplete: it says a fingerprint lets you
compare a credential "without ever materialising the secret" with no condition
attached. When two agents independently make the same mistake, the artifact that
taught them both is the bug.

§2 NOW CARRIES THE PRECONDITION. A fingerprint is 32 bits over its INPUT SPACE,
so publishing fp8(x) hands anyone a MEMBERSHIP ORACLE: they can test x == v for
every candidate v they can generate. Safe for a 40-char random token; a wordlist
for a hostname, username, e-mail, port, path, commit SHA or weak password. "High
entropy" is the usual sufficient condition, NOT the test — a commit SHA is
160-bit and still fully enumerable from the repo. Operationally: if you can
imagine writing the wordlist, you cannot publish the fingerprint. Also added:
candidate fingerprints are working memory and never output (an extractor hashes
hostnames and paths too, so the tempting "print what the scanner saw" debug step
leaks low-entropy fingerprints wholesale), and a plain statement that a
fingerprint register is a CONFIRMATION ORACLE for anyone already holding a
candidate corpus — which is exactly how a retired token is identified in old
transcripts, and works identically for someone else holding those same files.

NEW §6, "Proving absence: instrument strength, and four ways a scan lies clean",
placed next to §5 on purpose: §5 optimises against false POSITIVES, and every
failure in §6 is a false NEGATIVE. Triage optimises precision, a gate optimises
recall, and conflating them is what produced three clean reports over secrets
that were really there. Contents: instrument ranking (exact-byte value search >
class/structure pass > fingerprint census) with the instruction to state which
one produced your zero; census vs class passes as different questions, both
failure modes measured on this fleet; the tokenisation trap where quoting alone
decided detectability; scan the index or pushed tree, never the working tree;
git filters never run on symlinks while check-attr claims they do; two-sided
self-tests that abort, incl. the fixture-interaction artifact; row-gone is not
bytes-gone.

Attribution kept per finding: the census/class split and the instrument
ranking's provenance are pi@emb-7kj4vr4g's; exact-byte search over index blobs
is pi@tor-ms22's. The credential sense of "census" originated in this skill, not
with either agent.

TRAP FOR THE NEXT EDITOR, also in the CHANGELOG: the frontmatter description is
now 1022 of 1024 characters. Trim before adding, or the skill silently fails to
load. Verified by parsing the frontmatter (1022 chars, name intact, every prior
trigger phrase retained).

Deployment: baked skill -> needs an image rebuild AND a container recreate to
reach a running container.
2026-08-30 23:32:53 +02:00
joakimp 30094782df shell: source cli_utils' functions, and install the iproute2 that one of them needs
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 30s
v1.8.11 linked cli_utils' bin/ COMMANDS onto PATH and stopped there. Nothing ever
sourced cli_utils.sh, so its 14 FUNCTIONS were missing from every interactive shell
whose $HOME has no zsh rc -- which is the normal case, not an edge case: the
container's interactive shell is bash and zsh is not installed in the image. A
symlink cannot carry a shell function and a function cannot be reached from a
non-interactive shell, so the two mechanisms are disjoint and both are required.
The image was already paying this layer's dependency cost (fzf, bat, fd, rg, jq are
baked partly FOR these functions) while delivering none of its benefit.

The two changes ship together because they are coupled: portcheck is one of the 14,
and it was a hard stub in every image up to v1.8.11 -- neither ss nor ip nor lsof
nor netstat was present, so it printed "portcheck requires at least one of: ss,
lsof, netstat" and exited. Wiring the functions in without iproute2 would have
shipped a visibly broken one.

MEASURED, not assumed:
  - the loader is bash-safe despite the *.zsh filenames: `bash --noprofile --norc`
    exits 0, defines all 14, and they run (pathls, mkcd, up, extract, agents-sync,
    fhist verified). The tree's one zsh-only construct (print -z in fzf/fhist.zsh)
    is already guarded by [[ -n $ZSH_VERSION ]] with a bash fallback.
  - fresh-$HOME seeding resolves 14/14; CLI_UTILS_SOURCE=0 is honoured; an absent
    checkout is a genuinely silent no-op (no output, no leaked _cu).
  - interactive shell startup 12 ms -> 17 ms.
  - iproute2 is ~5.5 MB (4.2 MB itself + 6 libs under --no-install-recommends;
    libpam-cap is a Recommends and correctly dropped). ss lands at /usr/bin/ss,
    ip at /usr/sbin/ip, both already on the developer PATH, and `portcheck --all`
    then correctly identifies the socat listener on 8765.
  - hadolint clean on both Dockerfiles; repo-wide shellcheck -S error and bash -n
    clean. .bash_aliases is outside CI's discovery (no shebang, not *.sh), so it
    was checked by hand with -s bash at error AND warning level.

Named explicitly per this repo's floating-ref rule: /workspace/cli_utils is a HOST
BIND MOUNT, not a pinned ref, so the image now executes unpinned content in every
interactive shell. Errors are left visible rather than sent to /dev/null so that a
future zsh-only file in that repo is diagnosable rather than mysterious, and
CLI_UTILS_SOURCE=0 is the documented escape hatch. It is deliberately independent
of CLI_UTILS_LINK=0: the two disable independent mechanisms.

Deployment: needs a rebuild AND a recreate. $HOME is the container's writable layer
rather than a named volume (verified -- ~/.bash_aliases carries the container start
mtime while ~/.bashrc carries the image's), so the skel file is re-seeded on every
recreate; a host-bind-mounted ~/.bash_aliases is still never overwritten.
2026-08-30 11:48:03 +02:00
joakimp 9b5783f9dd skills: a sixth instance, found by the repo owner within the hour
Lint / actionlint (push) Successful in 16s
Lint / hadolint (push) Successful in 56s
Claimed "## Unreleased is a new convention in this repo" after reading
CHANGELOG.md once — minutes after b615571 had renamed that very section to
`## v1.8.11`. 33 commits touch the heading; the convention is that new entries
land under Unreleased and the heading is renamed at tag time, exactly as it
was renamed out from under my snapshot.

Same root cause as the five instances the section above already records: an
absence observed in one frame, promoted to a fact about the world, with no
second measurement. `git log -S'## Unreleased' -- CHANGELOG.md` was the oracle
and costs one command.

The generalisation is worth more than the instance, so it goes in the habits
block: to learn a repeating PROCESS, read history, not the file. A file's
current content is one frame of a cycle, and the frame you catch may be the one
where the thing you are looking for has just been consumed.
2026-08-30 10:50:20 +02:00
joakimp d9a7fe101b changelog: an Unreleased section for a rule that was already there
Lint / actionlint (push) Successful in 16s
Lint / hadolint (push) Successful in 1m22s
Records the two skill commits ahead of tomorrow's build, and states the finding
that shaped them: the "a negative result is usually your own filter" rule was
already baked, already symlinked in at every container start, and already
survived every recreate — then was violated five times by a session that had it
available. The gap was activation, not persistence, which is why the
cross-cutting form went into the always-appended AGENTS block instead of into a
skill that only loads when a task description matches.

Also notes what the entry's own subject implies for the reader: neither change
reaches a running container until the image is rebuilt AND the container
recreated, since ~/.agents/skills and the global AGENTS.md both live in the
image rather than in a volume or a mount.
2026-08-30 00:52:36 +02:00
joakimp 36e65fe657 skills: add credential-incident-response, and assert it stays baked
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 18s
Carries the facts a two-day credential incident produced, not the discipline:
probe the issuer FIRST (11 of 13 "exposed" credentials were already dead at the
provider, which cost five HTTP requests to learn and was never checked), the
403-vs-401 trap that scoped tokens introduce into liveness probes, revocation
beats deletion for anything already replicated, the three places a secret hides
in a Chroma palace (FTS content, metadata, raw bytes) in coverage order, scope
derivation from measured consumers, and this fleet's age store with its
single-recipient weakness.

Facts transfer between sessions; exhortations do not — hence a separate skill
for the domain knowledge and a one-line pointer in the always-loaded block.

Authored here, so baked is canonical and it is NOT added to skillset-owned.txt.
Skill dirs are picked up by a glob in entrypoint-user.sh, so no registration is
needed — verified rather than assumed, since an enumerated list would have left
the skill inert, a fitting failure given its subject. Three smoke assertions
extended so a future rebuild cannot silently drop it.
2026-08-30 00:50:11 +02:00
joakimp f0ebea2d98 skills: fix the half of the negative-result rule that was wrong
pi-devbox-environment already warned that "a negative result is usually your
own filter" — baked, symlinked in at every container start, authored by an
earlier session. It survived every recreate, was available all of a later
session, and was violated five times anyway. So the gap was never persistence.

That section also closed with "a positive result needs no such scepticism — it
carries its own evidence." That is false, and it aimed the scepticism budget
one way only. Three of those five errors were positives:

  - an SSH handshake SUCCEEDED and greeted me as joakimp while I believed I was
    probing gitea.egl.lan — `Host gitea*` had rewritten HostName
  - a 401 that was a genuine answer from an issuer which never minted the token
  - a "regression" produced by diffing against a value my own -p 2222 flag set

Adds the three missing false-negative rows, replaces the false claim with the
"a positive result only proves what you actually asked" subsection, and records
the two habits that actually caught these: `ssh -G` to learn which rule
captured a hostname, and declaring the expected result before running a check.

The cross-cutting form goes in pi-global-AGENTS.append.md rather than in a
skill, because it has to fire without a task description matching it — being
loadable on demand is exactly what failed. Across all five errors, none was
caught by re-reading my reasoning; every one was caught by a second measurement
that disagreed.
2026-08-30 00:50:11 +02:00
joakimp b615571913 changelog: release v1.8.11 — the mail nobody could deliver, and the ship that skipped a file
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 19s
Publish Docker Image / base-decide (push) Successful in 10s
Publish Docker Image / build-base (push) Successful in 53m20s
Publish Docker Image / smoke (push) Successful in 4m48s
Publish Docker Image / smoke-studio (push) Successful in 4m58s
Publish Docker Image / build-variant-studio (push) Successful in 17m0s
Publish Docker Image / build-variant (push) Successful in 26m43s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 24s
Names three things riding in via floating refs, since the rule this
fleet adopted (v1.8.9) is to name a behaviour change before tagging,
not after:

- mempalace-toolkit b2b50af -> 21023e7: the rsync --checksum ship fix
  (a real behaviour change to every remote-mode client) and RFC 003
  §7.13 (docs only).
- skillset 6eb20af -> a12fe5e: mermaid-diagrams cutU normalisation and
  the mempalace skill's from_agent identity rule, plus the vendored
  fallback snapshot re-pinned to match (this pi-devbox commit).

Dependency audit run by direct command against every component, not
assumed: pi/mempalace/pi-atelier/pi-studio/pi-toolkit/pi-extensions/
pi-fork/pi-observational-memory all confirmed unchanged since v1.8.10.
pi-studio's apparent SHA drift on ls-remote turned out to be an
annotated-tag-object hash vs its underlying commit hash, not real
drift -- checked with ^{commit} before writing it down as 'none'.
2026-08-27 23:39:39 +02:00
joakimp 495b7e3859 vendor: resync mempalace skill snapshot to skillset a12fe5e
scripts/vendor-mempalace-skill.sh, real refresh not --check: skillset
moved 6eb20af -> a12fe5e (mermaid-diagrams cutU normalisation, and the
from_agent identity rule this same release ships in RFC 003). --check
reported stale-but-truthful (exit 0, the sanctioned skip) but this
release's point is getting today's fixes live fleet-wide, and the base
rebuild is already forced by the entrypoint change and the floating
mempalace-toolkit ref moving -- so the incremental cost of also
bumping this pin is zero. Phrase canary in scripts/smoke-test.sh
unaffected: neither pinned phrase ('Provenance is stamped for you'
present, 'Attribute what you file yourself' absent) is in the section
that changed; both verified still correct in the new snapshot.
2026-08-27 23:36:03 +02:00
15 changed files with 1342 additions and 36 deletions
+22 -1
View File
@@ -33,18 +33,39 @@ on:
- 'v*'
workflow_dispatch:
inputs:
# `type:` is REQUIRED for Gitea to render these fields in the "Run
# workflow" dialog. Without it (Gitea 1.26.2) the dispatch form shows a
# branch selector and NO inputs at all, so a manual run silently uses
# every default — which for `release_tag: ''` means RELEASE_TAG resolves
# empty, the variant tag list becomes `<image>:`, and the run dies on an
# invalid reference AFTER paying the full base + smoke cost (~70 min).
# That made the documented `smoke_only` escape hatch below unreachable
# from the UI for its whole existence; found 2026-09-06 trying to use it.
#
# Deliberately `string` and not `boolean`, even though these two read as
# flags: every consumption is a STRING comparison against 'true'
# (`inputs.smoke_only != 'true'` at the build-variant gates,
# `inputs.promote_latest == 'true'` at the promote gates) plus string
# interpolation into env.PROMOTE_LATEST. A boolean-typed input yields a
# real boolean, so `!= 'true'` would compare across types and could
# invert a publish gate rather than fail loudly. Changing the type here
# would mean re-auditing all six call sites; keeping it string is a
# rendering fix with provably zero semantic change.
release_tag:
description: 'Release tag to publish (e.g. v1.0.0). Used only for workflow_dispatch runs.'
required: false
default: ''
type: string
promote_latest:
description: 'Update latest aliases (default true for tag-push, false for manual test runs)'
required: false
default: 'false'
type: string
smoke_only:
description: 'Build base + run both smoke jobs against HEAD, then stop. Publishes nothing. Use to validate smoke assertions without cutting a tag.'
description: 'Build base + run both smoke jobs against HEAD, then stop. Publishes nothing. Use to validate smoke assertions without cutting a tag. Set to the literal string true.'
required: false
default: 'false'
type: string
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
+571 -1
View File
@@ -11,7 +11,479 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
---
## Unreleased
## v1.8.13 — 2026-09-06
**Version audit + three pins moved, one deliberately not moved.** `pi`
0.84.4 -> 0.85.1, `mempalace` 3.8.0 -> 3.9.0, `pi-atelier` v0.10.0 -> v0.10.1.
`PI_FORK_REF=master` stays floating and therefore adopts e69725c. Each rationale
is written at the ARG itself rather than only here, because that is where the
next person doing the audit will be standing.
**Correction, made mid-release while run 639 was building:** the audit
originally recorded a fourth change — "`PI_STUDIO_VERSION` relabelled `none` ->
`v0.9.60-rc.0`, RC adopted deliberately" — and that was wrong. It was measured
at the wrong layer. `resolve-versions` passes BOTH `PI_STUDIO_REF` and
`PI_STUDIO_VERSION` as build-args and selects the newest **stable** semver tag
(its filter `^v?[0-9]+\.[0-9]+\.[0-9]+$` excludes pre-releases), so a Dockerfile
default cannot answer "what will CI publish?". Measured from the run itself:
`studio_tag=v0.9.59`, `studio_ref=9eed84f` (= `refs/tags/v0.9.59^{}`), while
`main`/`v0.9.60-rc.0` is 658536f and is not built. **Published v1.8.13 studio
images therefore contain pi-studio v0.9.59, not the RC**, and the ARG is back at
`none` rather than pinned to a pre-release that goes stale the moment main
moves. Consequence kept deliberately: the RC's opt-in Studio network binding is
absent from every published v1.8.13 image, so it needs no audit for this
release. Adopting an RC from CI would require changing that tag filter, which
exists on purpose — upstream stopped publishing Releases at v0.5.55 but keeps
tagging and pushing to main, so pinning main risked baking half-finished commits.
0.85.0 is SKIPPED on purpose: it shipped internal experimental code and extra
subpaths that broke SDK imports (upstream #9132), and 0.85.1 exists to undo
exactly that. Neither release has a Breaking/Removed changelog heading, the
engine floor is unchanged (>=22.19.0 against the container's 22.23.2), and
runtime deps drop 20 -> 19.
The pi bump was verified by RUNNING it, not by reading about it, because this
repo has already been burned by a version pair that no changelog flagged
(pi-atelier < 0.7.1 hangs pi >= 0.84 at startup with no error). 0.85.1 was
side-installed and driven under a pty in five combinations — each companion
extension plus atelier v0.10.0 AND v0.10.1 — with a CPU delta of 0.00-0.01s
over a 5s window where the known hang signature is ~5s of sustained CPU. The
check was two-sided: the atelier sidebar painted ACTIVITY+WORKSPACE markers
identically to the 0.84.4 control, so "alive" could be distinguished from
"silently absent".
**NODE_VERSION stays 22 — audited, not overlooked.** node 24 is technically
safe: all five prebuilt native addons in pi use NAPI (ABI-stable, no
NODE_MODULE_VERSION lock, no binding.gyp), nothing in the image declares a node
CEILING, and the install is one token (`setup_${NODE_VERSION}.x`). agent-browser
0.36.0 declares `engines.node >=24`, but that field is vestigial for the
artifact actually shipped: the image's own 0.35.2 declares the same floor and
runs fine on 22.23.2 as a prebuilt aarch64 ELF. The reason to wait is
attribution, not compatibility — this release already moves pi a minor,
mempalace a minor and bakes a Studio RC, so adding a node major would leave four
suspects if the image misbehaves. Worth doing as its own release with the smoke
suite as the gate. (v22 is in maintenance until 2027-04-30; v24 is Active LTS
to 2026-10-20 and maintained to 2028-04-30, so there is real headroom.)
mempalace's client bump carries a sequencing note that is now also CORRECT: the
comment at the ARG claimed synlig serves 3.7.1 server-side, which was stale.
Measured 2026-09-06 over ssh, synlig's uv tool entry last changed 2026-08-25
and serves 3.8.0. Client 3.9.0 against server 3.8.0 is accepted skew until
synlig's compose stack is redeployed; 3.9.0's headline additions (release
awareness, `task create`/`task launch`) are SERVER-side and stay dark until
then — a client bump alone cannot light them up.
**agent-browser was running 7 weeks stale, and the interesting part is why
nothing noticed.** The image has shipped 0.35.2 since the last base rebuild,
but every session on mbp-m1-2020 was executing 0.27.0 from a 2026-07-17
hand-install: `npm i -g` writes into `~/.pi/npm-global`, which is the
devbox-pi-config VOLUME, and PATH puts that at position 2 against /usr/bin at
position 8. This is the third package hit by that exact hazard (pi itself and
pi-atelier already have guards), so the guard is now generalised instead of
re-invented a fourth time.
The damage was not the binary. It was the BUNDLED SKILL, which is the part an
agent reads: 3 skillsets / 17.6 KB core in 0.27.0 versus 8 skillsets / 31.5 KB
core in 0.35.2, with ten subcommands present in the image and entirely
undocumented to the agent (a11y, browser, data, mcp, page, plugin, read,
selectors, to, webmcp). A stale tool announces itself with an error; a stale
skill just quietly teaches the wrong commands and everything looks fine.
Three changes, at the three places this can be caught:
- `entrypoint-user.sh` retires a volume copy by MOVING it aside (reversible,
same instinct as the settings backups) and only when the image ships its own
copy, so a machine that deliberately hand-installs on an image without one
keeps it. The `bin/` shim is removed too — a dangling symlink would be a
worse failure than a stale version.
- `scripts/recreate-sanity-check.sh` asserts `agent-browser` resolves under
/usr. This is the check that matters, because it runs where the volume is
real.
- `scripts/smoke-test.sh` gets the build-time half, labelled WEAK in the source
for an honest reason: a `docker run` container has an empty config volume, so
it can never see the shadowing it is nominally testing for.
**pi-fork gets a capability floor: `extensions: []`.** Forks were measured
twice (2026-09-01, 2026-09-06, four dispatches) ignoring their brief, answering
in the USER's voice, fabricating self-referential measurements, and once filing
a diary entry as `agent_name=pi` — which landed in `wing_pi`, where a
wing-scoped `diary_read` never sees it.
The cause is upstream and by design, so there is nothing to wait for: the child
is handed `getHeader()+getBranch()`, i.e. the WHOLE active session branch, with
the brief appended as the final user message and the system prompt untouched
(pi-fork `src/index.ts`). In a long session the parent narrative simply
outweighs the task, and the child does the statistically obvious thing — it
continues the story it finds itself inside. Config offers no context knob
(extensions, environment, offline, costFooter, effort profiles only).
Falsified the tempting explanation before acting on it: the failures are NOT a
too-small model. The same model as the `fast` profile (haiku, thinking off)
obeyed the identical brief perfectly when run as
`pi -p --mode json --session-id <fresh> --no-extensions` — correct values,
exact format, no session recap, 3 seconds, $0.012. Model held constant, context
inheritance removed, failure gone.
`extensions: []` is therefore a mechanical guarantee rather than an
instruction: the mempalace bridge is a pi EXTENSION, so a fork child now runs
with `--no-extensions` and cannot write to the shared palace under the parent's
identity. Verified by asking a child to enumerate its own tools: `read, bash,
edit, write` — no `mempalace_*`, no `recall`, no nested `fork`. Two honest
limits, stated so nobody over-trusts this: it removes PALACE writes, not
FILESYSTEM writes (`edit`/`write` remain), and it costs forks their palace
search and recall. Set the key to `null` to restore normal loading.
Smoke asserts the floor is `[]` specifically, not merely falsy — `null` is the
unguarded state, so a "truthy or not" test would pass on exactly the
configuration being guarded against.
**Vendored mempalace skill snapshot refreshed `a12fe5e` -> `e9e09d9`, and the
phrase canary re-pinned with it.** Folded in at zero marginal cost: the
snapshot is hashed into `base_tag`, but `Dockerfile.base` already changed this
release, so the ~67 min base rebuild was already being paid. `--check` reported
exit 0 (stale-but-truthful) beforehand, i.e. skipping was sanctioned — this is
the deliberate decision the checklist asks for, not a drive-by. Upstream content
is the fleet wing-naming convention (bare project names, no `wing_` prefix) and
the `<harness>@<device>` rule for `added_by`, both of which came out of the
attribution defect measured on this device on 2026-09-06.
The canary re-pin is the interesting half. Its old pair — "Provenance is
stamped for you" present, "Attribute what you file yourself" absent — STILL
PASSED against the new snapshot, so leaving it in place would have produced a
canary that is green on both the old and the new bytes: blind to precisely the
refresh it exists to witness, which is the same false-green family the
pre-v1.8.5 canary died of. The replacement pair was picked by MEASURING
direction against both files rather than by reading the diff ("Diaries
self-heal; plain drawers do not" new=1/old=0; "Agent diaries live in"
new=0/old=1) and then tested two-sided: PASS on the refreshed bytes, FAIL on the
old bytes recovered from git. A canary that cannot fail is decoration.
**`credential-incident-response` §5/§6 corrected — a stated mechanism was wrong,
and this is the second time in three days this section named a wrong reason
for a zero.** Docs only.
§5 said `embedding_metadata.string_value` holds "metadata fields only". Measured
false on chroma 1.5.9 with a disposable sentinel drawer (pi@tor-ms22,
2026-08-30): the document text is ALSO there, under key `chroma:document` — one
row in `fts_content` and one in `embedding_metadata` for the same drawer. The
scan order in §5 is unchanged (scan `fts_content` directly, raw bytes as
backstop) but the stated REASON is fixed: a zero from `string_value` needs a
different explanation (key filter, query shape, escaping), not "it's
structurally blind". §6 already warns against explaining a zero with an
unverified mechanism; this was exactly that failure, in the file that carries
the warning.
§6's row-gone/bytes-gone claim is now backed by the same sentinel measurement
rather than asserted: `delete_by_source` took both `fts_content` (1->0) and
`embedding_metadata` (1->0) to zero, while raw bytes stayed 4->4 until VACUUM.
Also records how the measurement got unblocked at all — not a better
instrument, a disposable sentinel drawer instead of testing deletion on real
data.
---
## v1.8.12 — 2026-08-31
**`pi` `0.84.3` → `0.84.4`, and `pi-atelier` `v0.8.2` → `v0.10.0`.** Both audited
by the routine in `Dockerfile.variant` rather than adopted on sight, and the
audit notes live next to the pins where the next reader will meet them.
**pi 0.84.4 (published 2026-08-28) carries no `Breaking Changes` and no
`Removed` heading** — checked by grepping the section, 0 matches, which is worth
stating because 0.84.3 *did* have one. It was adopted for three fixes that land
on machinery this fleet runs every day, not for the feature list:
- **#6879** — a large tool result crossing the auto-compaction threshold used to
be sent to the provider *before* compaction. Pi now compacts between tool
execution and the next assistant response inside the same run. That is the
shape of nearly every session on these boxes, where a single `event_list` or
palace search returns hundreds of KB.
- **#8345** — a resumed session corrupted its next appended entry when the JSONL
file lacked a trailing newline. That file is the memory feeder's *input*, so
the failure would have surfaced as unexplained gaps in `wing_conversations`
rather than as an error. Measured on tor-ms22 before bumping: 49/49
transcripts end in a newline and 0 lines fail `json.loads` — this corpus was
never bitten, and we now know that rather than hope it.
- **#8537** — extension messages sent with `triggerTurn: false` *while the agent
is running* were inserted between a tool call and its result, so
order-validating providers rejected the replayed history. **The mempalace
mailbox is outside that precondition**: it delivers at `agent_settled`, when
no inference is in flight, with `{deliverAs: "steer"}` and deliberately no
`triggerTurn`. 0.84.4 also leaves the documented steer semantics untouched
("delivered after the current assistant turn finishes executing its tool
calls, before the next LLM call"), so RFC 003 §7.11 stands as written. Recorded
because this fix is precisely what would make a *mid-run* delivery safe, which
is the only reason we would ever change that call.
Also new and relevant, though nothing here uses them yet: `ui_prompt_start` /
`ui_prompt_end` extension events (the `docs/extensions.md` diff is add-only — no
steer or `triggerTurn` semantics moved), and an RPC `clear_queue` that returns
and removes queued steering messages. The second one can discard an
already-delivered but unconsumed mailbox steer; that is survivable because the
mailbox re-delivers on `MEMPALACE_MAILBOX_RESURFACE_MS` (default 3600000), and it
is written down here so a future "the mailbox lost a message" report has a
candidate cause. The three new `PI_HYPERLINKS` / `PI_IMAGE_PROTOCOL` /
`PI_TRUE_COLOR` environment variables were grepped against this whole repo: no
collisions with anything the image sets.
**The bump moved one documented mechanism, so `docs/observational-memory.md` §3
moved with it.** Pi's own `docs/compaction.md` gained exactly one paragraph in
0.84.4: the `autoCompact` threshold is now *also* checked mid-run, after a tool
batch's results are appended and before the next assistant response, skipped only
when that batch ends the run and no queued message needs another response. Our
doc said compaction is "checked when pi goes idle, so it never interrupts a
turn". That was only ever true of observational-memory's **own** trigger
(`compaction-trigger.ts` hooks `agent_settled`); read as a statement about pi it
is now false. `session_before_compact` (`compaction-hook.ts`) therefore has
**two** entry points and the second can fire inside a turn — harmless for the
ledger fold, which makes no model call, but a doc that ships a false promise
about when a hook runs is worse than one that admits two paths. The §3 mermaid
diagram gained the second edge, and the whole file re-passes the bundled mermaid
checker (6 blocks, 44 labels, 0 soft-wrapped, no cut glyphs at 1280px and
800px).
**pi-atelier `v0.8.2` → `v0.10.0` is two minor releases and both are UI-only** —
Sidebar kept calm during an active Turn, composer frame and Status Rail polish,
fullscreen-copy-safe Sidebar, Windows path normalisation, Workspace Pulse
deferred until pi trusts the project. Neither release carries a BREAKING notice.
The coupling that matters runs the *opposite* way to this pin's hard-earned
floor: v0.9.0 renders the Sidebar as a separate split-layout child and therefore
"raises the minimum supported Pi version to 0.84.0", and — unlike the
0.7.1-under-pi-0.84 startup-hang precedent, which its metadata never encoded —
this time `peerDependencies` says so (`>=0.84.0`, up from `>=0.80.7`). Satisfied
with room to spare by `PI_VERSION=0.84.4`. It also pairs deliberately with a
0.84.4 feature: atelier keeps Sidebar content out of the fullscreen transcript
selection while pi adds `fullscreenCopyOnSelect` and Ctrl+X for the selection
itself. Both executable floors (`scripts/smoke-test.sh`,
`scripts/recreate-sanity-check.sh`) compare with `sort -V`, so `0.10.0 >= 0.7.1`
is evaluated correctly — verified by running the comparison, because the string
form of that test reads `0.10.0` as *older* than `0.7.1`.
**While bumping the pins, the README's own pin table turned out to have been
wrong since v1.8.6.** It advertised pi `0.84.2` and mempalace `3.7.1` in the very
table whose purpose is to tell a reader what is pinned and where. Both rows went
stale in the *same* commit — `93f986e` (v1.8.6, "adopt pi 0.84.3 + mempalace
3.8.0") moved both `ARG`s and neither table row; the rows themselves date from
`29b6209` (v1.8.0) and `2ebf00d` (v1.8.4). Only atelier's row was still true.
All three corrected now, and the `--expected-version 0.84.3` example in the
recreate-sanity section updated too, since that one is a copy-pasteable command
that would now fail against a 0.84.4 image. Worth noting how it survived two
releases: nothing checks prose against the `ARG`s, so this table has to be
remembered by hand on every pin bump, and once it was not.
**`credential-incident-response` gained the section its own guidance had been
missing, and §2 gained a precondition it should always have carried.** Docs only;
no image behaviour moves. Both changes came out of a session where three separate
detectors reported *clean* over secrets that were really there — the skill was
the artifact that had taught two agents the pattern, so the fix belongs here
rather than in either operator's private notes.
**§2 previously said an 8-hex fingerprint lets you compare a credential "without
ever materialising the secret", with no condition attached.** That is true only
when the *input space* is unreachable. A fingerprint is 32 bits over whatever it
was computed from, so publishing `fp8(x)` hands anyone a **membership oracle**:
they can test `x == v` for every candidate `v` they can generate. For a 40-char
random token, fine. For a hostname, username, e-mail, port, path, commit SHA or
weak password, that candidate set is a wordlist — and note that "high entropy" is
the usual sufficient condition, not the test: a commit SHA is 160-bit and still
fully enumerable from the repo. Two agents on this fleet published fingerprints of
`GIT_USER_EMAIL`-class values while following this section as written; harmless in
that instance, because those values sit in every commit trailer already, but the
guidance licensed it. §2 now states the precondition, adds that candidate
fingerprints are working memory and never output (a scanner hashes hostnames and
paths too, so "print what it saw" leaks wholesale), and names what a fingerprint
register *is* — a confirmation oracle for anyone already holding a candidate
corpus, which is exactly how a retired token gets identified in old transcripts,
and works the same way for someone else holding those files.
**New §6, "Proving absence: instrument strength, and four ways a scan lies
clean".** Deliberately placed next to §5, because §5 optimises against false
*positives* (name-anchoring, provenance — what stops a triage sweep drowning in
session UUIDs) and every failure in §6 is a false *negative*. Triage optimises
precision; a gate optimises recall, and conflating the two is what produced the
clean reports. It carries: an instrument-strength ranking (exact-byte value search
> class/structure pass > fingerprint census) with the standing instruction to say
which one produced your zero; census and class passes answering different
questions, with both failure modes measured here — a class-only pre-commit hook
passed plaintext UUID API credentials to a shared repo twice because a UUID has no
key header, while a census-only gate reported 0 hits with freshly-synced SSH
private keys in the tree because no key is in the census; the tokenisation trap,
where maximal-run extraction swallows an unquoted `VAR=<uuid>` so the value is
never hashed alone while a *quoted* one is found, meaning quoting alone decided
detectability; scan the index or the pushed tree, never the working tree, plus why
a repo-only fix on an rsync-published mirror is temporary rather than weaker; git
filters never running on symlinks, where `check-attr` answers `git-crypt` for a
path it can never encrypt, so a coverage audit must join the attribute against the
file mode and verify the blob magic; two-sided self-tests that abort, including
the fixture-interaction artifact where a quoted and unquoted probe share one
buffer and make the weak extractor look as strong as the union; and row-gone is
not bytes-gone, since a correct sqlite DELETE leaves the payload in freelist pages
until VACUUM.
Findings contributed by `pi@emb-7kj4vr4g` (the census/class split, and the
instrument ranking's provenance) and `pi@tor-ms22` (exact-byte value search over
index blobs). The description's trigger list grew accordingly and is 1022/1024
characters — **it has almost no headroom, so trim before adding to it**, or the
skill silently fails to load.
**Deployment:** the skill is baked at
`/usr/local/share/pi-devbox/skills/credential-incident-response/`, so this needs
an image rebuild **and** a container recreate to reach any running container.
**Two vendored skills changed, and one of the changes is a correction rather than
an addition.** Nothing about the image's behaviour moves; this is entirely about
what the next agent reads before it acts.
**`pi-devbox-environment` §2 had a rule that was half wrong, and the wrong half
cost five findings in one session.** The section "A negative result is usually
your own filter" closed with *"a positive result needs no such scepticism — it
carries its own evidence."* That sentence is false. A positive result is evidence
about the question your command *actually posed*, which may not be the question
you meant — and the failure is invisible precisely because the command succeeded.
Three measured instances, all from 2026-08-29, all filed as fact before being
caught: an SSH handshake that succeeded and greeted the agent as `joakimp` while
it believed it was probing `gitea.egl.lan` (a `Host gitea*` block had rewritten
`HostName`, so it authenticated to the wrong Gitea instance); a `401` that was a
genuine answer from an issuer which had never minted the credential being tested;
and a "regression" produced by diffing `ssh -G` output against a `2222` that the
agent's own earlier `-p 2222` flag had supplied. The section now carries a
counterpart, *"…and a positive result only proves what you actually asked"*, plus
the three false-negative rows that session added (a palace scan that queried
`embedding_metadata` while documents live in `embedding_fulltext_search_content`;
a token declared dead on a 401 from the wrong issuer; a host declared unreachable
after trying two of its three open ports, with the port written in an environment
variable the agent already held).
**The cross-cutting form of that rule went into `pi-global-AGENTS.append.md`, not
into the skill — deliberately, and this is the whole point of the change.** The
rule *already existed* in the baked skill, authored by an earlier session,
symlinked into `~/.agents/skills/` at every container start. It survived every
recreate, was available for the entire session that broke it, and was violated
five times anyway. So the gap was never persistence; it was **activation**.
A reasoning rule that only loads when a task description happens to match it
cannot fire on the occasions that need it, because "I am about to state something
false" is not a recognisable task type. The always-appended block is read by every
agent in every container without being asked for, which is the only property that
matters here. Writing a sixth document restating the rule would have felt like
progress and changed nothing.
**New baked skill: `credential-incident-response`.** Authored here, so the baked
copy is canonical and it is *not* listed in `skillset-owned.txt`. It carries the
*facts* a two-day credential incident produced, on the theory that facts transfer
between sessions where exhortations do not: probe the issuing provider **first**
(11 of 13 "exposed" credentials in that sweep turned out to be already dead at the
provider — five HTTP requests would have established it, and nobody asked);
`sha256[:8]` fingerprints as leak-free credential identity; the `403`-vs-`401`
trap that scoped tokens introduce into liveness probes, where a live token looks
revoked on `/api/v1/user`; **revocation beats deletion** for anything already
replicated, because deletion is best-effort over an unbounded copy set (FTS shadow
rows, per-host feed inboxes, sqlite free pages, mesh replicas, backups) while
revocation invalidates copies nobody enumerated; the three places a secret hides
in a Chroma palace, in coverage order; deriving least-privilege scopes from
*measured* consumers; and the exposures rotation does not fix (cleartext channels,
git history, agent-authored drawers).
**Three smoke assertions extended** so a rebuild cannot silently drop the new
skill: baked-file existence, resolves-to-the-baked-tree, and reported as `baked`
by `pi-devbox-version`. Skill directories are picked up by a glob in
`entrypoint-user.sh`, so no registration was needed — verified rather than
assumed, since an enumerated list would have left the skill inert, which would
have been a fitting way for *this* skill to fail.
Neither skills change reaches a running container until the image is rebuilt **and**
the container recreated: `~/.agents/skills/` and the global `AGENTS.md` both live in
the image, not in a volume or a mount.
**`cli_utils`' shell *functions* are now sourced, closing the half of that wiring
the image never did.** v1.8.11 linked the repo's `bin/` **commands** into
`~/.local/bin` so they resolve in non-interactive shells; nothing ever sourced
`cli_utils.sh`, so its 14 **functions** (`fgit`, `fhist`, `fssh`, `fdocker`,
`fmark`, `fproc`, `fex`, `fenv`, `extract`, `mkcd`, `pathls`, `portcheck`,
`agents-sync`, `up`) were missing from every interactive shell whose `$HOME` had
no zsh rc. That is the normal case, not an edge case: the container's interactive
shell is bash and **zsh is not installed in the image**. A symlink cannot carry a
shell function and a function cannot be reached from a non-interactive shell, so
the two mechanisms are disjoint and both are required — the image had been paying
this layer's dependency cost (`fzf`, `bat`, `fd`, `rg`, `jq` are baked partly *for*
these functions) while delivering none of its benefit. Now sourced from
`/etc/skel-devbox/.bash_aliases`, with the same detection order as the symlink
block so commands and functions can never come from two different clones.
`CLI_UTILS_SOURCE=0` opts out, deliberately independent of `CLI_UTILS_LINK=0`
because the two disable independent mechanisms. Measured: all 14 resolve in a
freshly-seeded `$HOME`, the opt-out is honoured, an absent checkout is a genuinely
silent no-op (no output, no leaked `_cu` variable), and interactive shell startup
goes from 12 ms to 17 ms.
**Named explicitly, per this repo's own floating-ref rule: `/workspace/cli_utils`
is a host bind mount, not a pinned ref.** Sourcing it means the image now executes
content it does not pin, on every interactive shell, on every device. It is
bash-safe today and that was measured rather than assumed — sourcing under
`bash --noprofile --norc` exits 0 and defines all 14 despite the `*.zsh`
filenames, the functions run, and the tree's single zsh-only construct (`print -z`
in `fzf/fhist.zsh`) is already guarded by `[[ -n $ZSH_VERSION ]]` with a bash
fallback. The residual risk is future content: a cli_utils commit adding a
genuinely zsh-only file would surface as parse errors at every prompt, fleet-wide.
Errors are therefore left visible rather than sent to `/dev/null`, so the failure
is diagnosable, and `CLI_UTILS_SOURCE=0` is the one-line escape hatch.
**`iproute2` is installed, so the container can answer "what is listening in
here".** Neither `ss` nor `ip` was present in any image up to and including
v1.8.11 — nor `lsof`, nor `netstat` — which made `cli_utils`' `portcheck` a hard
stub that printed `portcheck requires at least one of: ss, lsof, netstat` and
exited. `ss` satisfies its preferred branch (`ss -tlnp`), which is also the only
branch that reports the owning PID. `net-tools` is deliberately **not** added
(`netstat` is deprecated and only a fallback path) and neither is `lsof` (~500 KB
for a third route to the same answer). Cost measured, not estimated: ~5.5 MB total
— `iproute2` is 4.2 MB and pulls six libs under `--no-install-recommends`
(`libbpf1`, `libmnl0`, `libtirpc-common`, `libtirpc3t64`, `libxtables12`,
`libcap2-bin`; `libpam-cap` is a Recommends and is correctly dropped). Verified in
a live container: `ss` at `/usr/bin/ss`, `ip` at `/usr/sbin/ip`, both already on
the developer `PATH`, and `portcheck --all` then correctly identifies the `socat`
listener on 8765.
The two changes above also need a rebuild **and** a recreate, for a different
reason than the skills: `$HOME` is the container's writable layer rather than a
named volume (verified — `~/.bash_aliases` carries the container's start mtime
while `~/.bashrc` carries the image's), so the skel file is re-seeded on every
recreate. A `$HOME/.bash_aliases` that is bind-mounted from the host is still
never overwritten, which is the existing contract.
### Dependency audit (2026-08-31)
Every component checked against upstream by direct command, not assumed:
| Component | Baked in v1.8.11 | Upstream now | Action |
|---|---|---|---|
| **pi** | `0.84.3` (pinned) | **`0.84.4`** is npm latest | bumped + audited (above) |
| **pi-atelier** | `v0.8.2` (pinned) | **`v0.10.0`** highest tag | bumped + audited (above) |
| mempalace | `3.8.0` (pinned) | `3.8.0` is PyPI latest | none |
| skillset (mempalace fallback snapshot) | `a12fe5e` | `a12fe5e` == `origin/main`, 0 commits since | none — `--check` reports OK, no NOTICE |
| mempalace-toolkit | `21023e7` | `21023e7` | none |
| pi-toolkit | `0e1369e` | `0e1369e` | none |
| pi-extensions | `2022887` | `2022887` | none |
| pi-fork | `bf702b4` | `bf702b4` | none |
| pi-observational-memory | `ce9fc98` | `ce9fc98` (v3.0.4, peerDeps `*` → no pi floor to clear) | none |
| pi-studio (studio variant) | `3328b3d` | `3328b3d` | none |
| floating `*_VERSION=latest` tools (16) | — | 14 already at latest; `git-lfs` `3.7.1`→`3.8.0` (feature, no breaking section), `uv` `0.12.6`→`0.12.7` (patch) | adopted implicitly by the rebuild; named here per this repo's floating-ref rule |
| node | major pin `22`, installed `v22.23.2` | `v22.23.2` is the newest 22.x | none — a newer LTS *line* (24.x) exists and is deliberately not tracked |
Two method notes, because both would have produced a confident wrong answer:
- **An annotated tag's `ls-remote` SHA is the tag object, not the commit.**
`refs/tags/v0.8.2` is `6e07bf85` while `refs/tags/v0.8.2^{}` is `159f34cf` —
the value actually baked. Comparing the un-dereferenced form reported
`pi-atelier` as *drifted from its own pin*, which would have been a false
integrity alarm about the one component whose pin is load-bearing. Always
deref with `^{}` before calling a pin broken.
- **`git ls-remote --tags | sort -V | tail` is not a "latest release" proxy.**
`typst/typst` carries date-style tags (`v23-03-28`) and `mikefarah/yq` carries
`vTestA`/`vTestB`; both sort *after* the real releases. `Dockerfile.base`
itself resolves `latest` by reading the `Location` of
`curl -sI …/releases/latest`, so replaying that exact step is both
noise-immune and the same source of truth the build will see.
---
## v1.8.11 — 2026-08-27
**Shell state that the writable layer eats on every recreate now gets rebuilt at
start.** Two additions to `entrypoint-user.sh`, both idempotent, both silent
@@ -78,6 +550,104 @@ the v1.8.0 mistake of a smoke assertion written against a stage that does not
exist at run time. `workflow_dispatch` with `smoke_only` remains the way to
exercise this against `HEAD` before a tag.
**Also carried by the floating `mempalace-toolkit` main ref** (resolved at build
time, not by a pi-devbox commit — `MEMPALACE_TOOLKIT_REF=main`):
**A scrubbed re-export of a dormant session could silently never reach the
palace host.** `bin/mempalace-pi-session` ships to the palace with
`rsync -a --update`, and the stage file's mtime is deliberately the SOURCE
transcript's mtime (`os.utime()`, "preserve session mtime for dedup
stability"). Re-exporting a session that has not been appended to since its
last ship therefore produces a mtime that is *not newer* than the receiver's —
exactly the case a redactor upgrade needs to ship, since content differs while
mtime does not. `--update` reported success and sent nothing. Found and
patched by `pi@mbp-m1-2020` (mempalace-toolkit `a361b71`): `--update` →
`--checksum`, which compares content and ignores size/mtime entirely.
Dropping `--update` outright was considered and rejected — rsync's default
quick check already transfers on a size difference alone, which would have
masked the *next* instance of this (a redaction whose placeholder happens to
match the secret's length) as fixed. `os.utime()` is untouched; its backdating
is a separate, load-bearing design call for dedup stability. New regression
test, `scripts/test-rsync-ship-idempotency.sh`, runs fully offline (a local
rsync destination exercises the same size/mtime/checksum comparison as the ssh
transfer) and is built to *discriminate*: it must fail against `--update` and
pass against `--checksum,` not merely exercise the code path — the first draft
of the test used fixture strings of different lengths and passed for the wrong
reason (rsync's quick check transfers on size difference alone regardless of
`--update`), which is the same trap the patch itself was written to avoid.
**Acceptance line for this class of change going forward:** "receiver sha256
matches sender for every staged file", not "local stage is clean" — a clean
local stage says nothing about what a dormant session already sent.
**An event addressed to an identity no session runs as is delivered to
nobody, and this fleet has now hit it three separate ways.** RFC 003 gains
§7.13 and open-decision 10 (mempalace-toolkit `21023e7`, docs only, no image
behaviour change): the owed-set derivation — the log's only push channel — is
keyed on `to_agent`, and a reply is always addressed back to whatever string
the *original writer* put in `from_agent`. Nothing validates that string
against a live session identity, so authoring under a synthetic or foreign
name makes every reply to that event write-only. Measured cost this cycle: a
directed ask planted under a synthetic sender drew a correct reply containing
an urgent security finding, and it sat unread for ~2h20m, found only because a
human asked whether mail had arrived. Permitted exception, unchanged: a
synthetic sender is fine for a deliberate control experiment, provided the
body names the real identity to reply to.
**Also carried by the live `skillset` mount** (each device's own clone, not
baked — except the `mempalace` skill's fallback snapshot, re-vendored below):
**The mermaid-diagrams checker's cut gate moved from client pixels to a
per-SVG user-space unit.** `CUT_PX` was calibrated against one live page at
one render scale; sweeping `--viewport` 500→1600 on an *unchanged* document
moved the worst overflow −1.0px → −3.0px, near-proportional to the viewport,
i.e. a constant geometric overflow viewed through a changing scale. `cutU =
cutPx / scale` (scale taken per-SVG, never a page average — one page mixes
scales 0.643–0.988) recovers that invariant: the sweep now collapses to
exactly −3.0u at every width. Re-deriving the threshold against the live host
surfaced a real false negative the old pixel gate had: a label at
`cutPx=0.4, scale=0.678` read as healthy under `CUT_PX=0.5` but is `0.59u` —
a genuine cut hiding behind a compressed render scale. `CUT_U` stays `0.5`;
`cutPx` and `scale` are still printed on every issue so a devtools ruler still
confirms the number on the actual page. A new, explicitly-deferred finding
from the same review: `cut` only measures vertically, so an unbreakable token
wider than its box (a long URL, a `snake_case` identifier) is invisible to
soft-wrap, tall, *and* cut simultaneously — filed as a backlog item, not
implemented, pending a fifth acceptance control.
**The `from_agent`-identity finding above is also now in the `mempalace`
skill itself** ("Writing to another machine", and Anti-Patterns), and the
baked fallback snapshot of that skill was refreshed to match
(`vendor-mempalace-skill.sh`, `6eb20af` → `a12fe5e`) — sanctioned to skip on
its own (`--check` reported stale-but-truthful), done anyway because this
release's point is getting today's fixes live, and the base rebuild below was
already forced regardless.
### Dependency audit (2026-08-27)
Every component checked against upstream by direct command, not assumed:
| Component | Baked in v1.8.10 | Upstream now | Action |
|---|---|---|---|
| **mempalace-toolkit** | `b2b50af` | **`21023e7`** | ships the rsync ship-fix + RFC 003 §7.13 (both above) |
| **skillset** (mempalace fallback snapshot) | `6eb20af` | **`a12fe5e`** | re-vendored (above); live-mounted devices already had it |
| pi | `0.84.3` (pinned) | `0.84.3` is npm latest | none |
| mempalace | `3.8.0` (pinned) | `3.8.0` is PyPI latest | none |
| pi-atelier | `v0.8.2` (pinned) | `v0.8.2` highest tag | none |
| pi-studio (studio variant) | `v0.9.52` | `v0.9.52` — `main`'s commit and the tag's commit are identical (0 either direction) | none |
| pi-toolkit | `0e1369e` | `0e1369e` (local clone HEAD == `origin/main`) | none |
| pi-extensions | `2022887` | `2022887` (local clone HEAD == `origin/main`) | none |
| pi-fork | `bf702b4` | `bf702b4` | none |
| pi-observational-memory | `ce9fc98` | `ce9fc98` | none |
pi-toolkit / pi-extensions checked against their actual Gitea origin (the
Dockerfile's `PI_TOOLKIT_REPO` / `PI_EXTENSIONS_REPO`), not a GitHub mirror —
querying `api.github.com` for those two returned nothing (rate-limited or
blocked; not investigated, the local clones are the source of truth anyway).
No SHA above is fork-supplied; a fork inventing plausible-looking upstream SHAs
is a recorded failure mode (v1.8.9), so every value here came from
`git ls-remote`, a local clone's own `origin/HEAD`, `npm view`/registry JSON,
or the PyPI JSON API, run directly.
---
## v1.8.10 — 2026-08-27
+39 -7
View File
@@ -83,6 +83,24 @@ ENV DEBIAN_FRONTEND=noninteractive
# above); TERM=xterm-ghostty is compiled from an alias further
# down (ncurses ships `ghostty`, not `xterm-ghostty`). iTerm2
# defaults to xterm-256color (ncurses-base), so needs nothing.
# iproute2 — `ss` (socket statistics) and `ip`. Measured 2026-08-30 on
# v1.8.11: NEITHER was present, so the container could not
# answer "what is listening in here" by any means, and
# cli_utils' `portcheck` was a hard stub — it prints
# "portcheck requires at least one of: ss, lsof, netstat" and
# all three were absent. `ss` satisfies its preferred branch
# (`ss -tlnp`), which is also the branch that reports the
# owning PID, so nothing further is needed: net-tools is
# deliberately NOT added (`netstat` is deprecated and only a
# fallback branch) and neither is lsof (~500 KB for a third
# path to the same answer). ~5.5 MB total: iproute2 itself is
# 4.2 MB and pulls 6 libs under --no-install-recommends
# (libbpf1, libmnl0, libtirpc-common, libtirpc3t64,
# libxtables12, libcap2-bin — libpam-cap is a Recommends and
# is correctly dropped). Verified end-to-end in a live
# container: `ss` lands at /usr/bin/ss, `ip` at /usr/sbin/ip
# (both already on the developer PATH), and `portcheck --all`
# then correctly identifies the socat listener on 8765.
RUN apt-get update && \
apt-get upgrade -y --no-install-recommends && \
apt-get install -y --no-install-recommends \
@@ -122,6 +140,7 @@ RUN apt-get update && \
nano \
kitty-terminfo \
ncurses-term \
iproute2 \
&& ln -s /usr/bin/fdfind /usr/local/bin/fd \
&& apt-get clean \
&& rm -rf /var/lib/apt/lists/*
@@ -431,13 +450,26 @@ ARG INSTALL_MEMPALACE=true
# the part that should stay manual.
#
# Deployment sequencing note for whoever ships this bump: synlig (the shared
# central palace host) currently serves mempalace 3.7.1 SERVER-SIDE via
# docker-compose.mempalace.yml, which reuses this same devbox image. Bumping
# this ARG changes only the CLIENT version baked into pi-devbox images: it
# introduces client/server skew until synlig's compose stack is separately
# rebuilt/redeployed with the new pin. Not something to code around here —
# just sequence the redeploy.
ARG MEMPALACE_VERSION=3.8.0
# central palace host) serves mempalace 3.8.0 SERVER-SIDE via
# docker-compose.mempalace.yml, which reuses this same devbox image. (Measured
# 2026-09-06 over ssh: synlig's UV_TOOL_DIR mempalace entry last changed
# 2026-08-25 15:33 — this comment previously said 3.7.1, which was stale.)
# Bumping this ARG changes only the CLIENT version baked into pi-devbox
# images: it introduces client/server skew until synlig's compose stack is
# separately rebuilt/redeployed with the new pin. Not something to code around
# here — just sequence the redeploy.
#
# v1.8.13: 3.8.0 -> 3.9.0. Audited: no Breaking/Removed changelog headings.
# Adopted mainly for #2281 (`mempalace_mine` accepts a single conversation
# file again) — though note that does NOT unblock this image's own feeder,
# which was measured to mine DIRECTORIES, not files, so it was never hitting
# that bug. Four behaviour changes ride along and are skew-relevant while
# synlig stays on 3.8.0: hub-forward escaping, an HTTP lock split, similarity
# score semantics, and parsed-output compatibility. 3.9.0-only features
# (release awareness, `task create`/`task launch` MCP tools) are SERVER-side,
# so they stay dark until synlig is redeployed — a client bump alone cannot
# light them up.
ARG MEMPALACE_VERSION=3.9.0
ENV UV_TOOL_DIR=/opt/uv-tools
ENV UV_TOOL_BIN_DIR=/usr/local/bin
RUN if [ "${INSTALL_MEMPALACE}" = "true" ]; then \
+93 -6
View File
@@ -57,6 +57,32 @@ ARG USER_NAME=developer
# v0.74.0..v0.75.5; discovered + fixed in v0.75.5b, 2026-05-23). The `latest`
# branch below is kept only for a deliberate local `docker build` override.
#
# AUDITED AT 0.84.4 (2026-08-31, was 0.84.3): NO "Breaking Changes" and no
# "Removed" heading in the 0.84.4 section (grepped, 0 matches) — unlike 0.84.3,
# whose heading is described in the paragraph below and stays audited. Adopted
# for three fixes that land on machinery this fleet actually runs:
# - #6879 large tool results crossing the auto-compaction threshold were sent
# to the provider BEFORE compacting; pi now compacts between tool execution
# and the next assistant response in the same run. This is the shape of
# nearly every session here (multi-hundred-KB logstream/palace tool output).
# - #8345 a resumed session corrupted its next appended entry when the JSONL
# lacked a trailing newline. That file is the memory feeder's own input.
# Measured on tor-ms22 before the bump: 49/49 transcripts end in a newline,
# 0 lines fail json.loads — the bug had not bitten this corpus.
# - #8537 extension messages sent with `triggerTurn: false` WHILE THE AGENT IS
# RUNNING were inserted between a tool call and its result, so
# order-validating providers rejected the replayed history. The mempalace
# mailbox is outside that precondition — it delivers at `agent_settled`
# (idle) with `{deliverAs:"steer"}` and deliberately no `triggerTurn` — and
# 0.84.4 leaves the documented steer semantics unchanged, so RFC 003 §7.11
# still holds. Recorded because the fix is what would make a future mid-run
# delivery safe, which is the only reason we would ever change that call.
# One doc consequence, fixed in this same release: pi's own docs/compaction.md
# gained exactly one paragraph — the autoCompact threshold is now ALSO checked
# mid-run, after a tool batch's results are appended. See
# docs/observational-memory.md §3, which had said compaction is only checked
# when pi goes idle.
#
# AUDITED AT 0.84.3 (2026-08-25, was 0.84.2): upstream's notes carry a
# "Breaking Changes" heading — `GoogleThinkingLevel` renamed to
# `GoogleApiThinkingLevel`. INERT FOR THIS IMAGE: all four vendored companions
@@ -69,9 +95,25 @@ ARG USER_NAME=developer
# `.agents/skills/<group>/` directories were not discovered, and root Markdown
# files such as README.md / AGENTS.md inside a skill dir were reported as
# broken skills unless they declared valid skill frontmatter.
# pi-atelier needs no companion bump: v0.8.2 clears the >=0.7.1 floor that
# pi >= 0.84 requires (see PI_ATELIER_REF below).
ARG PI_VERSION=0.84.3
#
# v1.8.13: 0.84.4 -> 0.85.1. SKIP 0.85.0 deliberately — it accidentally
# published internal experimental code and extra subpaths, breaking SDK
# imports (upstream #9132); 0.85.1 exists specifically to undo that, with the
# supported SDK and stdio RPC API unchanged. Audited: no Breaking/Removed
# changelog headings in either release, engine floor unchanged (>=22.19.0,
# container runs 22.23.2), runtime deps 20 -> 19. User-visible changes are the
# streaming indicator moving into the editor border and faster fullscreen
# transcript search; no deprecation language anywhere.
#
# Verified EMPIRICALLY rather than from the changelog, because a pi bump has
# hung the TUI before (pi-atelier < 0.7.1 + pi >= 0.84): 0.85.1 was
# side-installed and driven under a pty against all four companion extensions,
# with atelier v0.10.0 AND v0.10.1 — five combinations, each rendering alive
# with a CPU delta of 0.00-0.01s over a 5s window, where the known hang
# signature is ~5s of sustained CPU. Two-sided check: the atelier sidebar
# painted ACTIVITY+WORKSPACE identically to the 0.84.4 control, so the test
# could distinguish "loaded" from "silently absent".
ARG PI_VERSION=0.85.1
ARG PI_TOOLKIT_REF=main
ARG PI_EXTENSIONS_REF=main
# Repo URLs default to the canonical gitea origin but are overridable so a
@@ -101,15 +143,36 @@ ARG PI_OBSMEM_REF=master
# pin and PI_VERSION together, checking atelier's CHANGELOG for the pi
# version it claims to track.
#
# AUDITED AT v0.10.0 (2026-08-31, was v0.8.2 — two minor releases): no
# BREAKING notice in either release, and both are UI-only (Sidebar calm during
# an active Turn, composer frame + Status Rail, fullscreen-copy-safe Sidebar,
# Windows path normalisation, Workspace Pulse deferred until pi trusts the
# project). The one coupling that matters runs the OPPOSITE way to the floor
# above: v0.9.0 renders the Sidebar as a separate split-layout child and
# therefore "raises the minimum supported Pi version to 0.84.0", which its
# peerDependencies do encode this time (`>=0.84.0`, up from `>=0.80.7`).
# Satisfied with room to spare by PI_VERSION 0.84.4 above — and note that both
# executable floors (scripts/smoke-test.sh, scripts/recreate-sanity-check.sh)
# compare with `sort -V`, so 0.10.0 >= 0.7.1 is evaluated correctly rather than
# as the string comparison that would read 0.10.0 as older than 0.7.1.
# Pairs deliberately with pi 0.84.4's own fullscreen selection-copy controls:
# atelier keeps Sidebar content out of the transcript selection, pi adds
# `fullscreenCopyOnSelect` + Ctrl+X for the selection itself.
#
# No `npm install` step, unlike pi-fork/pi-observational-memory/pi-studio:
# pi-atelier declares ZERO runtime dependencies (only peerDeps, satisfied by
# the baked pi) and has no build step — pi loads its TypeScript directly from
# the /opt checkout. Adding an install here would be a no-op that only costs
# build time.
ARG PI_ATELIER_REPO=https://github.com/michaelmjhhhh/pi-atelier.git
ARG PI_ATELIER_REF=v0.8.2
# v1.8.13: v0.10.0 -> v0.10.1. Refactor-only upstream (formatters, tests,
# panel identity); peerDependencies declare pi >=0.84.0, so it spans both the
# old and new pin. Included because it was already exercised: the pty matrix
# for PI_VERSION above ran atelier v0.10.1 against pi 0.85.1 and painted the
# sidebar identically to v0.10.0.
ARG PI_ATELIER_REF=v0.10.1
# Human-readable tag PI_ATELIER_REF was resolved from; recorded as a label.
ARG PI_ATELIER_VERSION=v0.8.2
ARG PI_ATELIER_VERSION=v0.10.1
RUN set -e && \
# git_fetch_ref: clone-equivalent helper that accepts EITHER a branch name
@@ -219,6 +282,30 @@ ARG PI_STUDIO_REF=main
# PI_STUDIO_VERSION is the human-readable tag (e.g. v0.9.36) that PI_STUDIO_REF
# was resolved from; recorded as a label below for at-a-glance identification.
# Only meaningful for the studio variant (default `none` otherwise).
#
# v1.8.13 — READ THIS BEFORE REASONING ABOUT WHICH pi-studio SHIPS. Neither
# default below survives a CI build. `resolve-versions` in
# .gitea/workflows/docker-publish.yml passes BOTH as build-args (studio_ref and
# studio_tag), and it deliberately selects the newest STABLE semver tag: its
# filter is `^v?[0-9]+\.[0-9]+\.[0-9]+$`, which excludes pre-releases. So a
# PUBLISHED v1.8.13 studio image contains pi-studio v0.9.59 (commit 9eed84f,
# = refs/tags/v0.9.59^{}), NOT the v0.9.60-rc.0 that `main` currently points at
# (658536f). The `main` default here only applies to a local `docker build`
# that passes no studio args.
#
# That upstream-tag-over-main choice is intentional and documented at the
# resolve step: pi-studio keeps tagging every version but stopped publishing
# GitHub Releases at v0.5.55 and pushes freely to main, so pinning main risked
# baking half-finished commits that land after a tag.
#
# Corrected here on 2026-09-06 after reading the run-639 resolve-versions
# output: the v1.8.13 audit had recorded "RC adopted deliberately" and set this
# ARG to v0.9.60-rc.0, which was measured at the wrong layer — a Dockerfile
# default cannot answer "what will CI publish?" when CI overrides it. Left at
# `none` rather than pinned to a tag, because a hardcoded pre-release here goes
# stale the moment main moves and would re-tell the same lie to the next reader.
# Consequence worth keeping: the RC's opt-in Studio network binding is NOT in
# any published v1.8.13 image, so it needs no audit for this release.
ARG PI_STUDIO_VERSION=none
RUN if [ "${INSTALL_STUDIO}" = "true" ]; then \
set -e; \
@@ -305,7 +392,7 @@ ARG MEMPALACE_TOOLKIT_REF=main
# no ~67-minute base rebuild. (scripts/check-base-hash.sh scans only
# Dockerfile.base, so no folding into the base hash is required — nor would
# it be correct, since this ARG changes nothing about the base's contents.)
ARG SKILLSET_SNAPSHOT_REF=6eb20af181f0147cb8c1377f6e36a6a47a68e8e5
ARG SKILLSET_SNAPSHOT_REF=e9e09d95f92670536a199fc986dfa24d787f18d1
# Dockerfile.base sets description="pi-devbox — base image (variant-independent)"
# and every variant INHERITS it, so both published images used to advertise
+4 -4
View File
@@ -1093,7 +1093,7 @@ persisted volumes survived, and pi runtime wiring is intact:
```bash
./scripts/recreate-sanity-check.sh # auto-detects variant
./scripts/recreate-sanity-check.sh --expected-image-version 1.8.9 # assert the pi-devbox release tag
./scripts/recreate-sanity-check.sh --expected-version 0.84.3 # assert the pi coding agent version
./scripts/recreate-sanity-check.sh --expected-version 0.84.4 # assert the pi coding agent version
```
Those are **two different versions**, and the flags are not interchangeable:
@@ -1132,9 +1132,9 @@ resolved to `latest` at build time:
| Component | Pin | Where |
|---|---|---|
| pi | `0.84.2` | `ARG PI_VERSION` — `Dockerfile.variant` |
| pi-atelier | `v0.8.2` | `ARG PI_ATELIER_REF` — `Dockerfile.variant` |
| mempalace | `3.7.1` | `ARG MEMPALACE_VERSION` — `Dockerfile.base` |
| pi | `0.84.4` | `ARG PI_VERSION` — `Dockerfile.variant` |
| pi-atelier | `v0.10.0` | `ARG PI_ATELIER_REF` — `Dockerfile.variant` |
| mempalace | `3.8.0` | `ARG MEMPALACE_VERSION` — `Dockerfile.base` |
The objective is **not** to freeze versions. Bumping is routine — usually one
line plus a changelog note. The objective is that adopting a new upstream
+14 -6
View File
@@ -21,8 +21,9 @@ palace, see
> Verified on pi-devbox **v1.8.9** (`release_tag v1.8.9`, source `aac4a1c`),
> which bakes pi-observational-memory **v3.0.4** at commit `ce9fc98` — the value
> in `/etc/pi-devbox/build-manifest.json` → `components.pi-observational-memory`.
> Every number below was read from that tree, from pi 0.84.3's own docs, or from
> the live container.
> Every number below was read from that tree, from pi's own docs, or from the
> live container. The pi-side mechanics were first read at pi **0.84.3** and
> re-checked at **0.84.4** (v1.8.12), which moved one of them — see §3.
---
@@ -93,6 +94,7 @@ flowchart TD
S(["agent_settled"]) --> C{"81k tokens<br/>since compacting?"}
C -- yes --> CP["ctx.compact()"]
CP --> H(["session_before_compact"])
A(["pi autoCompact<br/>idle, or mid-run<br/>after a tool batch"]) --> H
H --> F["fold the ledger<br/>no model call"]
F --> VIS["compacted memory"]
```
@@ -105,10 +107,16 @@ flowchart TD
a *successful same-turn* reflection **and** an active pool above
`observationsPoolTargetTokens` [10000]. Not a third worker on a third
threshold.
- **compaction** — `compactAfterTokens` [81000], checked when pi goes idle, so it
never interrupts a turn. Pi will also compact on its own when the context is
nearly full (`contextTokens > contextWindow - reserveTokens`, `reserveTokens`
[16384]).
- **compaction** — `compactAfterTokens` [81000], checked at `agent_settled`, so
*this* trigger never interrupts a turn. Pi will also compact on its own when
the context is nearly full (`contextTokens > contextWindow - reserveTokens`,
`reserveTokens` [16384]), and **from pi 0.84.4 that check also runs mid-run** —
after a tool batch's results are appended, before the next assistant response,
skipped only when the batch ends the run and no queued message needs another
response. So `session_before_compact` has **two** entry points and the second
one can fire *inside* a turn. Harmless for the fold itself, which makes no
model call, but worth stating plainly: "never interrupts a turn" was only ever
true of the observational-memory trigger, and reads as a promise about pi's.
## 4. What compaction actually does to your context
+32
View File
@@ -471,6 +471,38 @@ if command -v pi &>/dev/null; then
done
fi
# ── agent-browser: retire a stale volume copy that shadows the image ───
# Same hazard class as the pi-atelier retirement above, different delivery
# path — and this block exists because that guard did not generalise.
# ~/.pi/npm-global lives on the devbox-pi-config VOLUME, so anything ever
# installed there with `npm i -g` survives every image upgrade, and PATH puts
# it AHEAD of /usr/bin (position 2 vs 8).
#
# Measured on mbp-m1-2020, 2026-09-06: a 2026-07-17 hand-install pinned
# agent-browser 0.27.0 in the volume while the image shipped 0.35.2, so every
# session for ~7 weeks ran a stale CLI. The damaging part was not the binary
# but its BUNDLED SKILL, which is what the agent actually reads: 3 skillsets /
# 17.6 KB core in 0.27.0 vs 8 skillsets / 31.5 KB core in 0.35.2, with ten
# subcommands present in the image and undocumented to the agent (a11y,
# browser, data, mcp, page, plugin, read, selectors, to, webmcp). A stale tool
# announces itself; a stale skill quietly teaches the wrong commands.
#
# MOVE rather than delete (reversible, same instinct as the settings backups
# above), and only when the image ships its own copy — a machine that
# deliberately hand-installs agent-browser on an image WITHOUT one keeps it.
_ab_vol="$HOME/.pi/npm-global/lib/node_modules/agent-browser"
if [ -d "$_ab_vol" ] && [ -d /usr/lib/node_modules/agent-browser ]; then
_ab_park="$HOME/.pi/npm-global/.retired-agent-browser-$(date +%Y%m%d-%H%M%S)"
if mkdir -p "$_ab_park" 2>/dev/null && mv "$_ab_vol" "$_ab_park/" 2>/dev/null; then
# The bin shim is what PATH actually hits; leaving it behind would give a
# dangling symlink, which is a worse failure than a stale version.
rm -f "$HOME/.pi/npm-global/bin/agent-browser" 2>/dev/null || true
echo "agent-browser: retired stale volume copy -> ${_ab_park} (image copy now wins; delete the parked dir when satisfied)"
else
echo "WARN: agent-browser: stale volume copy at $_ab_vol shadows the image copy and could not be moved; retire it by hand"
fi
fi
# ── pi-studio: optional loopback bridge (opt-in) ──────────────────────
# pi-studio binds its server to 127.0.0.1 inside the container, which a
# published Docker port cannot reach. When STUDIO_EXPOSE is truthy (set in
+44
View File
@@ -116,6 +116,50 @@ if command -v fzf >/dev/null 2>&1; then
eval "$(fzf --bash)" 2>/dev/null || true
fi
# cli_utils — shell FUNCTIONS (fgit, fhist, fssh, portcheck, up, mkcd, extract,
# agents-sync, …). This is the OTHER HALF of the cli_utils wiring, and until
# v1.8.11 the image shipped only one half. entrypoint-user.sh symlinks the repo's
# bin/ COMMANDS into ~/.local/bin, which is what makes them resolve in
# NON-interactive shells (docker exec, agent tool shells, scripts). A symlink
# cannot carry a shell function, and a function cannot be reached from a
# non-interactive shell, so the two mechanisms are disjoint and both are
# required. Nothing sourced the loader: measured 2026-08-30 on v1.8.11, all 14
# functions were simply missing on a device whose $HOME has no zsh rc — which is
# the normal case, since the container's interactive shell is bash and zsh is not
# installed in the image. The image was already paying this layer's dependency
# cost (fzf, bat, fd, rg, jq are all baked partly FOR these functions) while
# delivering none of its benefit.
#
# Detection order deliberately mirrors the symlink block in entrypoint-user.sh so
# that commands and functions can never come from two different clones.
# CLI_UTILS_SOURCE=0 opts out. That is independent of CLI_UTILS_LINK=0 on purpose:
# they disable independent mechanisms, and someone who wants PATH commands
# without 14 extra functions in every prompt (or vice versa) should be able to
# say so.
#
# THE LOADER IS BASH-SAFE, MEASURED, NOT ASSUMED: despite every function file
# being named *.zsh, sourcing cli_utils.sh under `bash --noprofile --norc` exits
# 0 with no errors and defines all 14, and they run (pathls, mkcd, up, extract,
# agents-sync, fhist all verified). The single zsh-only construct in the tree
# (`print -z` in fzf/fhist.zsh) is already guarded by [[ -n $ZSH_VERSION ]] with
# a bash fallback, and the loader's own header states "bash & zsh compatible".
# ACCEPTED RISK, stated plainly: /workspace/cli_utils is a HOST BIND MOUNT, so
# unlike a pinned git ref this content floats outside the image's control. A
# future cli_utils commit that adds a genuinely zsh-only file would surface as
# parse errors at every prompt on every device. Errors are left VISIBLE rather
# than sent to /dev/null so that failure is diagnosable instead of mysterious,
# and CLI_UTILS_SOURCE=0 is the documented one-line escape hatch.
if [ "${CLI_UTILS_SOURCE:-1}" != "0" ]; then
for _cu in "${CLI_UTILS_CONTAINER_PATH:-}" /workspace/cli_utils "$HOME/cli_utils" /workspace/*/cli_utils; do
[ -n "$_cu" ] || continue
if [ -r "$_cu/cli_utils.sh" ]; then
. "$_cu/cli_utils.sh" || true
break
fi
done
unset _cu
fi
# ── PROMPT_COMMAND: flush history every prompt ───────────────────────
# Installed AFTER zoxide init so zoxide's hook is already in place;
# we append with a newline separator to avoid the ';;' parse error
@@ -70,3 +70,41 @@ rather than merely confusing you:
local disk, so `mempalace search` can return older and different results than
the MCP tools while both look correct. Use the MCP tools for the central
palace; the CLI only for a local one.
## Before you file a finding: second measurement, different route
This is here rather than in a skill because it has to fire *without* a matching
task description, and because the version of it that lived only in a skill was
violated five times in one session by an agent that had the skill available.
**Any claim you are about to record as fact — in a drawer, a diary entry, a
coordination event, or a report to the user — needs a second measurement taken
by a different route.** Not a re-read of your reasoning: re-reading has caught
zero of these. A disagreeing measurement has caught all of them.
The two shapes that get filed as fact and are not:
- **A negative result** (`401`, connection refused, zero rows, "not found") is
first a claim about *your filter*, not about the world. Wrong host, wrong port,
wrong table, capped output.
- **A positive result** proves only what your command *actually asked*. An SSH
handshake can succeed against the wrong host (`ssh -G` tells you which rule
captured the name); a `401` can be a real answer from an issuer that never
minted the credential.
Cheapest habit that works: **write the expected result next to each check before
running it**, then diff. Expectations declared up front turn a silent wrong
assumption into a visible mismatch. And if you cannot think of a second route to
the same fact, you do not have a finding — you have a hypothesis, so label it as
one.
## Handling an exposed credential
If a task touches a leaked secret, a token rotation, "is this credential still
live?", whether to delete stored content, or which scopes a new token needs:
**read `~/.agents/skills/credential-incident-response/SKILL.md` first.** One rule
is load-bearing enough to state here: **probe the issuing provider before doing
anything else** — most "exposed" credentials in a long-lived fleet are already
dead, and the ones that are live are often far more privileged than assumed.
Severity first, cleanup second, and prefer **revocation over deletion** for
anything already replicated.
@@ -9,6 +9,7 @@ one", which was a bug).
| skill | owner | how it gets here |
|-------|-------|------------------|
| `pi-devbox-environment` | pi-devbox (this repo) | authored here; the canonical copy |
| `credential-incident-response` | pi-devbox (this repo) | authored here; the canonical copy |
| `pi-extensions` | the `pi-extensions` package repo (`skill/`) | **vendored fallback** + refreshed at build |
| `mempalace` | the `skillset` repo | **vendored fallback** (snapshot only) |
@@ -0,0 +1,271 @@
---
name: credential-incident-response
description: >-
Respond correctly when a live credential is found where it should not be —
in a chat transcript, a MemPalace drawer, a log, a git-tracked config, or an
agent-authored note. Load this whenever a task involves a leaked/exposed
secret, a token rotation, a "is this credential still live?" question, deciding
whether to delete or scrub stored content, proving a corpus is clean, or
choosing scopes for a new API token. Covers the mandatory order of operations
(probe the issuer FIRST — severity before cleanliness), leak-free identity via
sha256[:8] fingerprints and when publishing one is safe,
why revocation beats deletion for anything already replicated, scopes derived
from measured consumers, the three places a secret hides in a Chroma palace, how to prove ABSENCE rather than assume it (instrument strength,
census vs class passes, the tokenisation trap where quoting decides detectability, why git filters never run on symlinks, self-tests that abort),
where this fleet's secrets live, and what rotation does NOT fix.
---
# Credential incident response
A leaked credential is a **severity** question before it is a cleanliness
question. Two days of scrubbing, redaction plumbing and deletion planning were
once spent on a set of 13 credentials of which **11 were already dead at the
provider** — a fact that cost five HTTP requests to establish and was never
checked. Meanwhile the two live ones turned out to be instance-owner **admin**
tokens, which nobody had looked at either.
## 1. Order of operations — do not reorder this
1. **Is it still accepted?** Probe the issuing provider. Dead credential →
hygiene item, stop panicking. Live → incident, continue.
2. **What can it do?** Read the identity back. `is_admin`, `id=1`, scopes,
which account. A read-only repo token and an instance-owner admin token are
not the same finding.
3. **What consumes it?** Grep for real consumers before assuming breakage.
4. **Where does it live?** Enumerate copies (store, palace, transcripts, git).
5. **Then** rotate/revoke, and only then consider cleanup.
Doing 4→3→1 in reverse produces confident, wrong severity calls and wasted
cleanup. If you only have time for one step, do step 1.
## 2. Leak-free identity: fingerprint, never the value
Publishing an 8-hex fingerprint lets you compare a credential across machines,
files, drawers and peers without ever materialising the secret. Same formula as
`mempalace_redact.py`:
```sh
printf '%s' "$SECRET" | sha256sum | cut -c1-8 # printf, NOT echo (no newline)
printf '%s' 'test' | sha256sum | cut -c1-8 # self-test -> 9f86d081
```
Report as `(variable, fp, length)`. Equal fingerprints across hosts prove a
shared credential; that is usually the important part. **Never** paste a live
value into a search query, a palace drawer, an event body, or a chat message —
in an agent context your own tool output is itself captured and re-filed.
**Precondition — only fingerprint what an adversary cannot enumerate.** An 8-hex
fingerprint is 32 bits over its *input space*, so publishing `fp8(x)` hands
anyone a **membership oracle**: they can test `x == v` for every candidate `v`
they can generate. For a 40-char random token that space is unreachable. For a
hostname, username, e-mail, port, path, commit SHA or weak password it is a
wordlist. **If you can imagine writing the wordlist, you cannot publish the
fingerprint** — reference those by name and location instead. "High entropy" is
the usual *sufficient condition*, not the test: a commit SHA is 160-bit and still
fully enumerable from the repo. `sha256("")` = `e3b0c442` is the degenerate case,
recognisable on sight precisely because its input space has one member.
**Candidate fingerprints are working memory, never output.** A scanner that hashes
every token in a file also hashes hostnames, paths and e-mails. Print only
fingerprints that *matched* a known entry — the tempting debug step when a scan
returns zero ("print what it saw") publishes low-entropy fingerprints wholesale.
And say plainly what a fingerprint register *is*, so nobody rediscovers it later
as an alarm: even for an unguessable secret, a published fingerprint is a
**confirmation oracle** for anyone who already holds a candidate corpus. That is
exactly how a long-retired token gets identified in old transcripts — and it works
identically for someone else holding those same files. Net positive, since they
would already hold the value; state it rather than leaving it implicit.
## 3. Liveness probes, and the trap that scoping creates
```sh
# Gitea
curl -sS -m 10 -o /dev/null -w '%{http_code}\n' -H "Authorization: token $T" \
"$GITEA_HOST/api/v1/repos/<owner>/<repo>/actions/runs?limit=1"
# GitHub
curl -sS -m 10 -o /dev/null -w '%{http_code}\n' -H "Authorization: token $T" \
https://api.github.com/user
```
- `200` live · `401` revoked/invalid · **`403` = wrong question, not a dead token**
- **Probe the issuer that minted it.** A 401 from an unrelated instance says
nothing. Resolve the host from config (`GITEA_EGL_HOST` etc.), do not assume.
- **Under scoped tokens, `/api/v1/user` returns 403 for a perfectly live token**
unless `user` scope was granted. So it cannot distinguish *revoked* from
*merely scoped*. Use a **repository route the token is authorised for**.
- Verify **both directions** after a rotation: old → 401, new → 200. The second
check is what catches "deleted the wrong token".
- Port/scheme come from config, not habit: one instance here is
`http://gitea.egl.lan:3000` — plain HTTP, with 443 refused.
## 4. Revocation beats deletion — the load-bearing rule
Once revoked, stored copies are **inert**; you may leave them. Deleting them is
best-effort over an *unbounded* copy set: FTS shadow rows, feed inbox `.jsonl`
files on every host, sqlite free pages after the delete, mesh replicas that
already synced, and backups. **Revocation invalidates every copy everywhere at
once, including copies nobody enumerated.**
So: **rotate + revoke first.** Treat drawer deletion as optional hygiene, never
as the remedy. Then record the retired fingerprints as *known-dead* so the next
census recognises them instead of reopening the investigation.
Corollary: never reach for `mempalace_sync` or a bulk `delete_by_source` on a
shared palace as incident response. High blast radius, low actual benefit.
## 5. Finding a secret in a Chroma palace — three targets, in this order
1. `embedding_fulltext_search_content.c0` — document text
2. `embedding_metadata.string_value` — metadata fields, **and a second copy of
the document text** under key `chroma:document`
3. raw byte scan of every `*.sqlite3` — backstop, covers FTS pages and free space
**Correction, measured on chroma 1.5.9 with a sentinel drawer:** one row in (1)
AND one row in (2) for the same drawer, so **(2) is not structurally
content-blind** — an earlier version of this section said it held "metadata
fields only", and that was wrong. Scan (1) and (3) regardless: (1) is the direct
target. But if a `string_value` query returns zero for a value you know is in a
drawer, the cause is a key filter, a query shape or escaping — *not* structural
absence, and the difference matters because the false explanation is what makes
the zero feel safe. See §6: do not explain a zero with a mechanism you have not
read from source.
Semantic search proves nothing about absence — it returns top-k. For
completeness, enumerate by filing window (`list_drawers(since=T, before=T+1m)`),
since one mine shares a minute.
Value-agnostic sweeps (uuid / 40-hex / `NAME=VALUE`) drown in false positives at
fleet scale — 608 candidates, mostly session UUIDs and git SHAs. Name-anchoring
plus entropy plus provenance, applied to **document text**, is what works.
## 6. Proving absence: instrument strength, and four ways a scan lies clean
Section 5's warning is about false *positives* — name-anchoring and provenance are
what stop a triage sweep drowning in session UUIDs. **A gate is the opposite job.**
Triage optimises precision; proving absence optimises recall. Every failure below
reported a reassuring zero over a secret that was really there.
**Rank the instrument, and state which one produced your zero.**
| Instrument | Needs | Blind to |
|---|---|---|
| exact-byte value search | you hold the value | nothing — no tokeniser to fool |
| class/structure pass | a header pattern | anything without a recognisable shape |
| fingerprint census | a fingerprint list | any secret not listed; tokenisation |
A census is deliberately value-free, so it must *extract candidates and hash them*
— which makes its sensitivity a property of the tokeniser, not of the corpus. If
you hold the value, search the bytes instead, and search the value's JSON-escaped
rendering too when the corpus is `.jsonl`.
**1. Census and class answer different questions; neither substitutes.** A census
answers *"has a KNOWN secret leaked?"*, a class pass *"is there secret-SHAPED
material here?"* Both failure modes were measured on this fleet: a class-only
pre-commit hook passed plaintext UUID API credentials to a shared repo twice,
because a UUID carries no key header — while a census-only gate reported 0 hits
with freshly-synced SSH private keys and an age identity in the tree, because no
key is in the census. Run both passes.
**2. Tokenisation — quoting alone can decide detectability.** Maximal-run
extraction swallows the value of an *unquoted* assignment:
```
PROXMOX_SECRET=<uuid> # ONE run; the uuid is never hashed alone -> MISS
export SECRET="<uuid>" # the quote ends the run; bare uuid hashed -> HIT
```
Take the **union** of three strategies, because each fails in a different
direction — (2) is the one that recovers the unquoted case:
~~~python
runs = re.findall(r'[^\s"\'`]{12,}', text) # 1. maximal runs
split = [p for r in runs for p in re.split(r'[=!,;:@|()\[\]{}<>]', r) if len(p) >= 12]
shape = re.findall(UUID_RE, text) + re.findall(r'[0-9a-f]{32,64}', text)
candidates = set(runs) | set(split) | set(shape)
~~~
**3. Scan the index or the pushed tree, never the working tree.** The working tree
is not what gets published. And for an rsync-published mirror a repo-only fix is
not weaker, it is *temporary*: the next sync re-publishes the live disk. Fix the
live file first, verify it clean **by fingerprint**, then sync. Read blobs with
`git ls-tree -r <sha>` plus one `git cat-file --batch` (thousands of `git show`
calls is the slow way).
**4. Git filters never run on symlinks — and `check-attr` will not tell you.** A
symlink's blob is the *target path*, so `filter=git-crypt` can never encrypt it,
yet `git check-attr filter` cheerfully answers `git-crypt` for that path. **A
symlinked secret stays plaintext no matter what `.gitattributes` says.** Join the
attribute against the **file mode** (`git ls-files -s`, mode `120000`) and verify
the index blob really begins `\0GITCRYPT\0`. Report encrypted / symlinked /
scanned as three separate numbers and assert they sum — encrypted and symlinked
blobs are *skipped*, not certified clean.
**Self-test two-sided, and abort if it cannot discriminate.** Require a synthetic
positive to fire AND a negative to stay silent before trusting any zero. Keep the
fixtures in *structurally separate buffers*: put a quoted and an unquoted probe in
one buffer and the quote terminates the run, handing the bare token to the weak
extractor and making it look as strong as the union — a self-test artifact that
has already fooled an agent here. And never gate on `$?` when the tool has a
lock-skip or no-op path that also exits 0; judge the reported line.
**Row-gone is not bytes-gone.** Measured, same sentinel drawer: after
`delete_by_source` the row count went 1 -> 0 in *both* the FTS content table and
`embedding_metadata`, while the raw byte count stayed 4 -> 4 — sqlite does not
zero freed pages, so the payload sits in free space until `VACUUM`. Deletion
effectiveness is therefore *two* numbers, and each direction has a trap: one
aggregate figure reported as "erased" has only measured "unretrievable", while a
raw byte scan used as the acceptance gate reads a CORRECT, complete deletion as a
failure. (Note how this was measured: the blocker was never a better instrument,
it was the subject — file your own disposable sentinel and delete that, instead
of testing deletion on real data.)
## 7. Choosing scopes: derive them from measured consumers
Before creating a replacement token, find out what actually uses it:
```sh
git -C <repo> remote get-url origin # ssh:// ? then git needs NO token
git config --global --list | grep -iE 'credential|insteadof' # and no helper?
grep -rhoE 'api/v1/[A-Za-z0-9/{}$_.-]+' <consumers> | sort -u # exact routes
grep -rhoE '\-X [A-Z]+' <consumers> # any writes?
```
Real outcome here: git used SSH keys throughout, and the token's only consumer
read three CI-run routes with `GET`. So `repository: Read` and nothing else
replaced two admin tokens. **Scoping shrinks the blast radius of the next leak
far more than any redaction pipeline does** — a read-only token in a transcript
is a hygiene event, not an instance compromise.
Then prove the scope with an acceptance suite that declares expectations first:
must-work routes → `200`; `/admin/*`, `/user`, `/user/repos` → `403`.
## 8. What rotation does *not* fix
- **A cleartext channel.** If the endpoint is `http://`, the *new* token is
exposed identically from first use. Raise TLS separately.
- **Git history.** A secret committed and pushed cannot be fixed by any store or
palace operation — it needs rotation *and* history surgery.
- **Agent-authored content.** Stage-write redactors see transcripts only, never
`add_drawer` / `checkpoint` / `diary_write` output. Never type a secret into
the palace yourself; nothing downstream will catch it.
- **Plaintext/encrypted drift.** Gitignored plaintext `.env` files go stale while
`.env.age` moves on, so old values linger on disk (and in backups) long after
rotation. They are a common source of "mystery" fingerprints in a census.
## 9. This fleet's secret store (verify, do not assume)
- All `*.env.age` live in **one** repo: `joakimp/docker-compose-repo`. `myconfigs`
has none.
- Every `.age` file has **one X25519 recipient** — a single key tracked in
`myconfigs` under git-crypt. Unlocking git-crypt therefore decrypts the entire
fleet's secrets, including hosts you have no access to. The age layer adds no
isolation beyond git-crypt.
- Flow: `./fetch-secrets.sh <host>` (decrypt → `.env`) → edit → `./encrypt-secrets.sh <host>`
→ commit → push → `docker compose up -d --force-recreate`.
- **Always pass the host argument** to `encrypt-secrets.sh`. Bare, it walks the
whole tree and re-encrypts every `.env` it finds, re-nonced, including stale
ones — silently rolling back other hosts' secrets.
- After any re-encrypt, check the header still shows exactly **one X25519
recipient**; a hand-rolled `age -r` locks the rest of the fleet out, and the
failure only appears on another machine, later.
@@ -428,10 +428,56 @@ An obligation you never agreed to is noise, so the sender states it:
| `to_agent="*"` (any status) | broadcast FYI | nothing |
| any other status (`ready`, `applied`, `blocked`, …) | a statement of fact | nothing |
**That table says what you *owe*. Delivery is stricter, and the difference bites:
the mailbox is an obligation channel, not a news channel.** Mailbox candidates are
drawn with `status="open"`, so an event carrying any **terminal** status
(`applied`, `superseded`, `failed`, `blocked`) is never a candidate — *whoever it
is addressed to*. A `task.reply` written to a named machine to share a finding is
delivered to nobody, ever, and neither is any `event_ack`. It sits in the log
until somebody reads the log.
So the most natural inter-machine message — *"here is something you should
know"* — is exactly the shape that gets no delivery. Pick deliberately:
| You want the peer to… | Write |
|---|---|
| **do something**, and you need it tracked until done | directed `status="open"` ask, with a `correlation_id` |
| **know something**, no response needed | terminal-status event **plus a drawer** — the drawer is what actually reaches them, via search |
What does **not** work is a terminal report plus an expectation of attention.
Measured 2026-08-26: a detailed report addressed to `pi@<peer>` with
`status="applied"` went unread for two and a half hours until the operator quoted
the event id by hand, with the mailbox working correctly the whole time. Full
mechanism in the toolkit's `docs/rfc-003-coordination-log.md` §7.12.
One more timing fact, because it looks like negligence and is not: a delivered
ask is queued into the agent's **next turn** (`deliverAs: "steer"`, deliberately
no `triggerTurn`), and the poll fires when the agent is *idle*. Between delivery
and the next turn no inference runs, so **a human starting a turn is the
trigger** (§7.11). An agent that "has not reacted" has usually not been running.
Ack with `mempalace_event_ack(event_id=…, from_agent="<you>", status=…)`. It
**appends a new event** and never mutates the original; the correlation id is
copied for you, and `metadata.ack_of` is set to the event you answered.
**Claiming, and what it does not do.** `status="claimed"` announces that you have
picked work up. Nothing requires it — a directed open ask owes "an ack *or* a
reply", and finishing the work is a complete answer. Do it anyway when the work is
long or the machine is unreliable, because it is the only thing that later
distinguishes *nobody started this* from *someone started and their container
died mid-task*. Be clear about its limits, both of which follow from candidacy
requiring exactly `status="open"`:
- **It does not notify the requester.** `claimed` is not `open`, so a claim is no
more deliverable than a finished report is (see the delivery table above). Its
reader is whoever pulls the log.
- **It does not quiet your own mailbox.** The ask stays owed until a *terminal*
event of yours joins it, so a claimed-then-silent thread keeps resurfacing —
correctly.
Prefer a prompt terminal reply over a claim plus a long silence; claim *in
addition*, when the gap between pickup and finish is where a machine might die.
#### What you actually owe — derive it, do not read it off `status`
The log is append-only and `status` is written **once**, so it is an honest
@@ -502,6 +548,17 @@ Two consequences worth internalising:
- **Address the stamped name you actually saw** in a `from_agent` field, e.g.
`pi@tor-ms22`. A bare `pi` reaches nobody's mailbox once stamping is live, and
older events in the log still carry bare names — do not copy them.
- **The rule runs in reverse too: what you put in YOUR OWN `from_agent` decides
where every reply to your event goes.** Nothing stops you writing a synthetic
or borrowed identity there, and a reply is always addressed back to exactly
that string — so if no live session ever runs as it, the reply is stored,
searchable, and delivered to no one. Measured cost: a directed ask sent under
a synthetic sender got two correct replies, one of them an urgent security
finding, and both sat unread for ~2h20m because nobody's mailbox was that
identity (RFC 003 §7.13). Authoring under a synthetic name is fine for a
deliberate control experiment — this fleet does it on purpose — but then
**name the real identity to reply to inside the body**, because the address
line is not a safe place to also carry provenance.
- **Use `status="open"` only when you truly need an answer.** It places an
obligation on another machine.
- **Never broadcast an ask.** `to_agent="*"` + `status="open"` obliges everyone
@@ -513,7 +570,8 @@ Two consequences worth internalising:
be matched to it at all.
- **Corrections are new events, never edits.** Say explicitly what you retract
and name the id — drawer or event — that carried the withdrawn claim.
- **Put a retraction where the reader will look.** An event reaches a live agent;
- **Put a retraction where the reader will look.** A *directed open ask* reaches a
live agent's mailbox; a **terminal-status event reaches no mailbox at all**, and
a *drawer* is what a future semantic search finds. If you filed advice as a
drawer and later withdraw it, file the withdrawal as a drawer too — otherwise
the next agent finds your original confident advice and no trace of the
@@ -529,9 +587,31 @@ Two consequences worth internalising:
### Wings
Wings are top-level categories, typically one per project or domain:
- Named after the project directory (e.g., `cli_utils`, `opencode_devbox`)
- Agent diaries live in `wing_<agent_name>` (e.g., `wing_orchestrator`, `wing_pi`)
Wings are top-level categories, typically one per project or domain.
**NAMING CONVENTION — decided 2026-09-06 by Joakim: bare project names, no `wing_`
prefix.** `home-network`, `pi-devbox`, `mempalace-toolkit` — *not* `wing_pi-devbox`. The
mass is already there (`pi-devbox` 2061 drawers vs `wing_pi-devbox` 25), and a prefix
present on some wings and absent on others turns every read into a guess about which
spelling holds the content.
- Named after the project directory or domain (e.g., `cli_utils`, `home-network`)
- **Always pass `wing` explicitly to `diary_write`.** Omitting it defaults to
`wing_{agent_name}`, which mints or feeds a *parallel* wing — this tool default, not
anyone's sloppiness, is the mechanism that produced the drift. Measured harm
(2026-09-06, `pi@mbp-m1-2020`): a diary entry written with `agent_name=pi` and no
`wing` landed in `wing_pi` while that agent's history lives in `pi-devbox`, so a
`diary_read` scoped to `pi-devbox` showed **no trace of it**. A wing-scoped read that
silently returns an incomplete history is the worst failure mode a memory store has.
- **Legacy `wing_*` wings are frozen and documented, not renamed.** `wing_conversations`
(written by the session feeders), `wing_pi`, `wing_pi-devbox`, `wing_pi-tor-ms22`,
`wing_pi-devbox-emb7kj`, `wing_mempalace`, `wing_orchestrator`, `wing_code` all still
hold real content. **When searching for history, check both spellings** — this is the
practical cost of the drift and it does not go away by decree.
- If a migration is ever done, the acceptance criterion must be at the **relationship**
level: chunk ids still resolve to their parent, and `diary_read` returns the same entry
set before and after. Per-wing drawer counts can look correct while the relationships
underneath are broken, because a count query never touches them.
#### Shared palace: multiple harnesses, and possibly multiple machines
@@ -551,7 +631,7 @@ Zechner's pi-coding-agent). Implications:
When the palace is **central** (shared across machines), these further things apply:
- **Check which machine a conversation came from.** Transcripts are fed per device, so `source_path` reads `…/mempalace-feed/<device>/pi_<uuid>.jsonl` while the displayed `source_file` is only the basename. One search can legitimately return hits from several machines at once — look at the device segment before attributing a decision to *this* project.
- **Provenance is stamped for you — leave it alone.** Drawers carry `device` and `agent_kind` metadata (plus `device_source`/`agent_kind_source` recording *how* each was determined, so an inference is never mistaken for a fact). You do **not** set these, and you no longer set `added_by` either: the pi bridge defaults the writer field to `<harness>@<device>` on `add_drawer`/`checkpoint`/`mine`/`event_append`/`artifact_put`, and prefixes diary entries with `HOST:<device>|`, from host-supplied `$MEMPALACE_PI_DEVICE`. RFC 001 §7.3.2 ranks "agent stamps it via a skill instruction" as the *worst possible* place for exactly the reason you would expect — it is per-call boilerplate that gets forgotten, and it did: the agent who wrote the previous version of this bullet then filed its own provenance drawer as `added_by=checkpoint`. **Confirm the bridge in your image actually stamps before trusting it:** the extension is baked at image build time, so a container on an image older than the stamping commit (pi-devbox < v1.8.7) stamps nothing while still satisfying both gates — the env vars are set and the code is simply absent. Check with `grep -c MEMPALACE_PI_DEVICE "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)"`; zero means keep passing `added_by="<harness>@<device>"` and a manual `HOST:<device>|` diary prefix until the container is recreated on a newer image. Two things remain yours: pass `source_drawer_id` on `kg_add` (triples have no provenance field, so that pointer is the only path back to a device), and pass an explicit `added_by` **only** when deliberately filing on behalf of another device. Never invent values for `device`/`agent_kind`/`origin_device` — a fabricated value is worse than a blank, because it silently corrupts a future merge.
- **Provenance is stamped for you — leave it alone.** Drawers carry `device` and `agent_kind` metadata (plus `device_source`/`agent_kind_source` recording *how* each was determined, so an inference is never mistaken for a fact). You do **not** set these, and you no longer set `added_by` either: the pi bridge defaults the writer field to `<harness>@<device>` on `add_drawer`/`checkpoint`/`mine`/`event_append`/`artifact_put`, and prefixes diary entries with `HOST:<device>|`, from host-supplied `$MEMPALACE_PI_DEVICE`. RFC 001 §7.3.2 ranks "agent stamps it via a skill instruction" as the *worst possible* place for exactly the reason you would expect — it is per-call boilerplate that gets forgotten, and it did: the agent who wrote the previous version of this bullet then filed its own provenance drawer as `added_by=checkpoint`. **Confirm the bridge in your image actually stamps before trusting it:** the extension is baked at image build time, so a container on an image older than the stamping commit (pi-devbox < v1.8.7) stamps nothing while still satisfying both gates — the env vars are set and the code is simply absent. Check with `grep -c MEMPALACE_PI_DEVICE "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)"`; zero means keep passing `added_by="<harness>@<device>"` and a manual `HOST:<device>|` diary prefix until the container is recreated on a newer image. Two things remain yours: pass `source_drawer_id` on `kg_add` (triples have no provenance field, so that pointer is the only path back to a device), and pass an explicit `added_by` **only** when deliberately filing on behalf of another device — and when you do, it **must** be `<harness>@<device>`. A bare nickname (`pi-devbox-claude`) has no `@device` to parse, so `agent_at_device` cannot attribute it and the drawer is unattributable *by rule*, not by lag: it survives every future stamp run with no `device`, and on a shared palace a device-less drawer is one nobody can later scope, audit or clean up per machine. Measured 2026-09-06: 11 drawers on `tor-ms22` were filed this way — including the credential rows, i.e. exactly where "which machine measured this?" matters most — by an agent that had passed its own chosen nickname on every call. Its *diary* entries escaped, because `HOST:<device>|` in the AAAK text recovers the device. **Diaries self-heal; plain drawers do not.** The safest habit is the one above: pass nothing and let the bridge stamp. Never invent values for `device`/`agent_kind`/`origin_device` — a fabricated value is worse than a blank, because it silently corrupts a future merge.
- **Metadata is invisible to search — so check the text, not the fields.** `search` results are built from a fixed key list and `diary_read` returns content, so neither ever shows `device`/`added_by`. Only `mempalace_get_drawer` reveals them. This is why diary entries carry an in-text `HOST:<device>` marker: it is the only attribution a reader actually sees. **A diary entry with no `HOST:` marker predates the convention and may be from any machine — do not assume it is this one's history.**
- **Mined drawers carry the MINE date, not the session date.** When history is imported, or re-mined on the palace host, `filed_at`/`created_at` is the *import* time — so sorting by them does not give chronological order. Real session time is recoverable from the UUIDv7 in `pi_<uuid>.jsonl`: the first 12 hex digits are milliseconds since the epoch (and UUIDv7 sorts lexicographically in time order, so a plain filename sort is already chronological). Agent-authored drawers and diaries have no such backdoor — for those `filed_at` is the only chronology, which is why it must never be restamped.
- **Beware the timezone mismatch when you combine those.** Palace `filed_at`/`created_at` are naive timestamps in the palace host's local time, while a UUIDv7 decodes to UTC. Comparing them directly introduces a silent offset (2 h for a CEST host). Normalise before drawing conclusions about ordering.
@@ -600,4 +680,5 @@ Entity-relationship triples with temporal validity. Query with `mempalace_kg_que
- **Don't treat the palace as a task list.** It's for knowledge and context, not todos.
- **Don't broadcast an ask, and don't leave one unanswered.** On a shared palace, `to_agent="*"` + `status="open"` obliges every machine and therefore none of them. And don't expect acking to tidy your mailbox: `status` is immutable, so the event keeps matching either way — what a terminal reply buys you is that the *derived* owed set (see *What you actually owe*) stops counting it. Leave asks unanswered and that set only grows, until everyone learns to stop looking. "Seen, not doing it" is a complete answer — silence is not.
- **Don't assume you would have heard.** Nothing pushes another machine's message into your session. If you did not run the mailbox query at wake-up, a correction addressed to you by name can sit unread while you confidently rebuild the thing it warned you about.
- **Don't author an ask under an identity nobody runs as, including your own throwaway labels.** The failure is symmetric to the one above: it is not that you missed a message, it is that nothing could ever have delivered the reply to you, because you addressed it at a name instead of an agent. If you must use a synthetic sender for a control or an experiment, say inside the body who should actually receive the reply.
- **Don't invent provenance metadata, and don't hand-stamp it either.** An earlier version of this list told you to set `added_by="<harness>@<device>"` by hand; that instruction has been withdrawn, because RFC 001 §7.3.2 places provenance at the client/server boundary and the pi bridge now does it uniformly (see *Provenance is stamped for you* above) — but the withdrawal only holds where the bridge is live, so run the one-line check in that bullet first; on an older image hand-stamping is still the only signal a hand-filed drawer gets. DO NOT invent values for the palace's own metadata fields (`device`, `agent_kind`, `origin_device`): those are stamped by infrastructure that also records *how* each was determined, and a fabricated value is worse than none because it silently corrupts a future merge. DO pass `source_drawer_id` on `kg_add`. And never put a machine name in a diary's `agent_name` — it becomes the wing name and hides your entries from `diary_read`.
@@ -143,6 +143,10 @@ mine:
| "`tor-ms22` is not in the SSH config" | `grep … \| head -20` — the entry was at **line 454**. `~/.ssh/config` here is ~500 lines. |
| "the Docker host has no `docker`" | non-interactive SSH `PATH` lacks `/usr/local/bin` (§2, §3). It was at `/usr/local/bin/docker`. |
| "no ControlMaster is running" | pattern `ssh ` (trailing space) cannot match a master: those processes **rename themselves** to `ssh: <controlpath> [mux]`. |
| "the credential is not in the palace" | scanned `embedding_metadata.string_value` only. Drawer **text** lives in `embedding_fulltext_search_content.c0`; 554k metadata rows proved nothing. |
| "this token is dead — 401" | probed it against the **wrong issuer**. A 401 from an instance that never issued the credential is not evidence about the credential. |
| "that host is unreachable, can't test" | tried ports 443 and 80. It was on **3000**, and the env var I already held (`GITEA_EGL_HOST`) stated the scheme and port. |
| "this repo has no `## Unreleased` convention" | read `CHANGELOG.md` **once**, minutes after a release commit had renamed that section to a version heading. 33 commits touch `## Unreleased`. A snapshot cannot show you a cycle. |
Habits that would have caught all three:
@@ -155,10 +159,55 @@ ssh -F "$HOME/.ssh-local/config" mac 'command -v docker || ls /usr/local/bin/doc
# match a process's ACTUAL argv, not the name you imagine
ps -eo pid,etime,args | grep -Ei 'mux|mosh|ssh'
# to learn a repeating PROCESS or convention, read history, not the file. A
# file's current content is one frame of a cycle, and the frame you happen to
# catch may be the one where the thing you are looking for was just consumed.
git log -S'## Unreleased' -- CHANGELOG.md # not `head -60 CHANGELOG.md`
```
A positive result needs no such scepticism — it carries its own evidence. Only
absence has to be *earned*, so spend the extra command there.
Absence has to be *earned*, so spend the extra command there.
### …and a positive result only proves what you *actually asked*
An earlier version of this section claimed "a positive result needs no such
scepticism — it carries its own evidence." **That is false, and believing it
cost a later session three more wrong findings.** A positive result is evidence
about the question your command really posed, which may not be the question you
meant. The failure is invisible precisely *because* the command succeeded.
| Claim | The command succeeded — at answering something else |
|---|---|
| "EGL git over SSH works" | `ssh git@gitea.egl.lan` greeted me as `joakimp`. `~/.ssh/config` had `Host gitea*` → `HostName gitea.jordbo.se`, so I authenticated **to the wrong instance**. The real EGL account is `ecsjper`. |
| "the port config regressed" | compared `ssh -G` output against `2222` — a value produced by **my own earlier `-p 2222` flag**, not by the config. I reported the user's edit as a regression it never caused. |
| "the CI runners authenticate with this token" | pure fabrication, contradicted by my own scan output already on screen. The runners use per-runner `REGISTRATION_TOKEN`. |
Two habits that actually catch this class, both cheap:
```sh
# 1. ask which RULE captured your hostname before trusting any ssh result.
# ssh_config is first-obtained-value-wins PER KEYWORD, not per block: a
# specific block only wins the keywords it declares, so a later `Host gitea*`
# still supplies HostName unless the specific block restates it.
ssh -G git@thehost | grep -E '^(hostname|port|user|identityfile)'
# 2. state the expected result BEFORE running the check, and diff against it.
# This is the single technique that separated the one verification that went
# right (10/10, expectations declared per probe) from five that went wrong
# (results interpreted after the fact, each time in the direction I expected).
probe "/repos/.../actions/runs" 200 # must work
probe "/admin/users" 403 # must be denied
```
And the meta-observation, which is the reason this subsection exists: across all
five errors, **not one was caught by re-reading my own reasoning.** Every one was
caught by a second measurement that disagreed — the SSH lie surfaced only because
the greeting said `joakimp` while a token probe minutes earlier had said
`ecsjper`; the fabrication surfaced only because the user read my own output back
to me. So the operational rule is not "be careful". It is: **for a load-bearing
claim, produce a second measurement by a different route, and expect it to
disagree.** If you cannot think of a second route, you do not yet have a finding
— you have a hypothesis.
**`dscp`/`scp` with accented filenames on a macOS host.** macOS stores filenames
in Unicode **NFD** (decomposed — e.g. `ä` is `a` + combining U+0308), while the
+28
View File
@@ -374,6 +374,34 @@ if [ -f "$HOME/.pi/agent/settings.json" ]; then
fi
fi
# ── agent-browser must resolve to the image, not the config volume ────
# The same volume-shadowing hazard already asserted for pi (above) and
# pi-atelier (just now), for the third package it has bitten. This check
# belongs HERE rather than only in smoke-test.sh: a build-time container has an
# empty ~/.pi/npm-global, so smoke-test can never see the stale copy that a
# real recreate inherits. Measured instance: 0.27.0 from 2026-07-17 shadowed
# the image's 0.35.2 for ~7 weeks on mbp-m1-2020, silently supplying an older
# BUNDLED SKILL (3 skillsets vs 8) — the agent read the stale instructions
# without any version mismatch ever being surfaced.
AB_PATH=$(command -v agent-browser 2>/dev/null || true)
if [ -z "$AB_PATH" ]; then
warn "agent-browser not on PATH (expected in v1.6.0+ images; skipping shadow check)"
else
AB_REAL=$(readlink -f "$AB_PATH" 2>/dev/null || echo "$AB_PATH")
AB_VER=$(agent-browser --version 2>/dev/null | head -n1)
case "$AB_REAL" in
/usr/*)
pass "agent-browser resolves to the image copy (${AB_VER:-version unknown})"
;;
*)
fail "agent-browser resolves to $AB_REAL (${AB_VER:-version unknown}) — a ~/.pi/npm-global VOLUME copy is shadowing the image; the entrypoint retirement guard did not run or could not move it"
;;
esac
if [ -d "$HOME/.pi/npm-global/lib/node_modules/agent-browser" ]; then
fail "stale agent-browser still present in the ~/.pi/npm-global volume (entrypoint guard did not retire it)"
fi
fi
# ── pi <-> pi-atelier compatibility floor ─────────────────────────────
# atelier < 0.7.1 wraps pi's private TUI renderer in a way that recurses under
# pi >= 0.84: pi hangs at startup burning CPU, with no error message. atelier's
+48 -4
View File
@@ -245,6 +245,8 @@ run "socat" "socat -V"
run "studio-expose helper" "test -x /usr/local/bin/studio-expose"
run "image-baked pi-devbox-environment skill" \
"test -f /usr/local/share/pi-devbox/skills/pi-devbox-environment/SKILL.md"
run "image-baked credential-incident-response skill" \
"test -f /usr/local/share/pi-devbox/skills/credential-incident-response/SKILL.md"
run "global-AGENTS append snippet present" \
"test -f /usr/local/share/pi-devbox/pi-global-AGENTS.append.md"
run "pi-devbox block merged into pi-global-AGENTS.md" \
@@ -594,11 +596,25 @@ exec_test "mempalace skill linked (fallback)" 'test -L $HOME/.agents/skills
# This assertion is kept because it is orthogonal and free: it pins content,
# not provenance, so it still catches a re-vendored snapshot whose ref was
# bumped correctly but whose bytes came from the wrong place.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Provenance is stamped for you" "$f" && ! grep -q "Attribute what you file yourself" "$f" && echo ok'
#
# v1.8.13: RE-PINNED on refresh a12fe5e -> e9e09d9, which is the whole point of
# the mechanism — the previous pair ("Provenance is stamped for you" present /
# "Attribute what you file yourself" absent) still passed against the NEW
# snapshot, so leaving it would have produced a canary that is green on both the
# old and the new bytes, i.e. blind to precisely the refresh it exists to
# witness. Same false-green family as the pre-v1.8.5 canary this comment warns
# about. The replacement pair was chosen by MEASURING direction against both
# files rather than by reading the diff: "Diaries self-heal; plain drawers do
# not" is new=1/old=0, "Agent diaries live in" is new=0/old=1 — so each string
# discriminates on its own and the pair still fails loudly in BOTH directions
# (forgotten bump AND re-vendored stale snapshot). Upstream content behind this
# refresh: the bare project-name wing convention and the <harness>@<device>
# added_by rule.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Diaries self-heal; plain drawers do not" "$f" && ! grep -q "Agent diaries live in" "$f" && echo ok'
# Link TARGETS, not just link existence: with no skillset mounted (as here) the
# baked tree must be what resolves, for all three vendored skills.
# baked tree must be what resolves, for all four vendored skills.
exec_test "vendored skills resolve to the baked tree (no skillset mounted)" \
'for s in mempalace pi-extensions pi-devbox-environment; do
'for s in mempalace pi-extensions pi-devbox-environment credential-incident-response; do
case "$(readlink -f $HOME/.agents/skills/$s)" in
/usr/local/share/pi-devbox/skills/$s) ;;
*) echo "$s resolves to $(readlink -f $HOME/.agents/skills/$s)" >&2; exit 1 ;;
@@ -612,7 +628,7 @@ exec_test "vendored skills resolve to the baked tree (no skillset mounted)" \
exec_test "pi-devbox-version reports skill sources (all baked, no skillset here)" \
'out=$(pi-devbox-version)
echo "$out" | grep -q "skills:" || { echo "no skills section" >&2; exit 1; }
for s in mempalace pi-extensions pi-devbox-environment; do
for s in mempalace pi-extensions pi-devbox-environment credential-incident-response; do
echo "$out" | grep -qE "^ $s +baked$" \
|| { echo "$s not reported as baked" >&2; exit 1; }
done; echo ok'
@@ -716,6 +732,34 @@ exec_test "pi-atelier registered in packages[] (TUI sidebar)" \
exec_test "pi-atelier registered from /opt, not npm: (volume-shadowing guard)" \
'jq -e "((.packages // []) | any((type == \"string\") and endswith(\"/pi-atelier\"))) and (((.packages // []) | any(. == \"npm:pi-atelier\")) | not)" $HOME/.pi/agent/settings.json'
# agent-browser: the third package hit by ~/.pi/npm-global volume shadowing
# (after pi itself and pi-atelier). This build-time check is deliberately WEAK
# and says so: a `docker run` container has an EMPTY config volume, so it can
# only prove the image ships a sane copy and nothing in the image itself
# shadows it. The check that actually bites lives in
# recreate-sanity-check.sh, which runs where the volume is real — that is
# where a 7-week-old 0.27.0 was caught shadowing 0.35.2 on 2026-09-06.
exec_test "agent-browser resolves under /usr (volume-shadowing guard, build-time half)" '
p=$(command -v agent-browser) || { echo "agent-browser not on PATH" >&2; exit 1; }
r=$(readlink -f "$p")
echo "resolved=[$r] version=[$(agent-browser --version 2>/dev/null | head -n1)]" >&2
case "$r" in /usr/*) ;; *) exit 1 ;; esac
test ! -d "$HOME/.pi/npm-global/lib/node_modules/agent-browser" || exit 1
echo ok
'
# pi-fork capability floor. `extensions: []` makes a fork child run with
# --no-extensions, which is the only MECHANICAL guarantee that a fork cannot
# file drawers or diary entries under the parent's identity — the mempalace
# bridge is an extension, so removing extensions removes the write path.
# Asserted because it is a security-shaped default that a settings merge or a
# hand-edit could silently drop, and its absence is invisible until a fork
# writes to the shared palace as you (measured twice: 2026-09-01, 2026-09-06).
# Deliberately compares to [] and not "is falsy": null means "load normal
# extensions", i.e. exactly the unguarded state this asserts against.
exec_test "pi-fork extensions floor is [] (forks cannot write to the palace)" \
'jq -e ".[\"pi-fork\"].extensions == []" $HOME/.pi/agent/settings.json'
# ── /tmp/sshcm directory created by entrypoint ────────────────────────
exec_test "/tmp/sshcm dir mode 700 (ssh ControlMaster)" \
'test -d /tmp/sshcm && [ "$(stat -c %a /tmp/sshcm)" = "700" ] && echo ok'