Compare commits

..

109 Commits

Author SHA1 Message Date
joakimp 6891dc32b8 changelog: release v1.8.10 — deploy the scrubber, and name the check that must not be skipped
Publish Docker Image / build-variant-studio (push) Successful in 17m4s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 14s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-variant (push) Successful in 18m45s
Publish Docker Image / build-base (push) Successful in 41m59s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / smoke (push) Successful in 4m55s
Lint / hadolint (push) Successful in 8s
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Audited every component against upstream rather than assuming. Only two moved
since v1.8.9:

- mempalace-toolkit 5b8d78f -> b2b50af (ours): feeder scrubber, the symlink fix
  that stops it refusing to stage, hlc owed-set join, queued-delivery note,
  explicit MEMPALACE_MAILBOX_NOTIFY protocol modes, RFC 003 + fleet-memory docs.
- pi-studio v0.9.48 -> v0.9.52 (upstream, studio variant only): 22 commits, all
  additive — PDF previews, hideable header, contextual side questions. No removals
  or renames. Its pi floor is >=0.84.3 against our exact 0.84.3 pin: satisfied,
  zero headroom, named as a watch item because the next floor bump breaks the
  studio job only, after core has already published.

Verified unchanged, with dates proving they predate the v1.8.9 bake: pi 0.84.3 (=
npm latest), mempalace 3.8.0 (= PyPI latest), pi-atelier v0.8.2, playwright 1.62.1
(no drift this cycle despite floating on latest), pi-fork bf702b4, pi-obsmem
ce9fc98, pi-toolkit 0e1369e, pi-extensions 2022887. No breaking changes anywhere,
so everything can ship in one tag.

The reason for tagging now is not features. The scrubber has been live on exactly
ONE device since this morning; every other device has kept staging unscrubbed
transcripts into a shared palace, and cleanup after the fact is manual redaction
(done twice today, with freed-page residue left behind by choice).

The entry leads with the acceptance check instead of burying it, because this
release's worst failure is silent: the feeder is fail-closed, so a packaging or
path mistake stops the fleet's memory feed and nothing complains — refusing to
stage looks exactly like a quiet session. f0bffd1 exists because that very bug was
real (BASH_SOURCE reports the symlink path, not the target). First client on the
new image must confirm a [scrub] summary line appears, the drawer count moves, and
exit 3 did not fire. Silence is the failure signal, not success.
2026-08-27 17:51:46 +02:00
Joakim Persson 8a673ec143 docs: unclip the diagrams, and answer what compaction leaves behind
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 15s
Two problems reported against docs/observational-memory.md, one cosmetic and one
substantive. Both turned out to be worth more than the fix.

Clipping. Several boxes lost their bottom line of text in the viewer, and neither
the source nor my own render showed it. Cause: Mermaid measures a node label with
its own font metrics, commits to a box size, then renders that label as real HTML
in a <foreignObject> — so any host stylesheet touching line-height or font-size
inflates the text past a box that is already fixed, and the overflow is clipped.
The error accumulates per line, which is why it always eats the last line of the
tallest labels.

Rejected the obvious fix after testing it rather than assuming it:
%%{init: {'flowchart': {'htmlLabels': false}}}%% *is* honoured (labels switch from
16 foreignObject to 7 tspan) and clips identically, because inflated font-size
inherits into SVG text too. The fix that works is a hard limit of two short lines
per node, with the detail moved into prose under each diagram — one- and two-line
boxes have the vertical slack to absorb inflation, three- and four-line boxes do
not. The nodes were carrying paragraph-sized text; the diagrams are better for
losing it.

Regression harness: render every block with line-height: 1.7 !important forced
onto the label HTML, screenshot, read it. That caught two survivors of the rewrite
that looked fine in the clean render — a long unbreakable /opt path wrapping to a
third line, and a cylinder shape whose curved bottom leaves less room than a
rectangle for the same two lines.

New §4, because the document explained that compaction folds the ledger and never
said what that leaves in the context. Asked directly: is the session back to
knowing nothing? No. A verbatim tail survives, sized by keepRecentTokens (20k) and
cut only at turn boundaries; the system prompt and AGENTS.md were never in the
compacted region because they are rebuilt from disk each request; nothing is
deleted from disk, since compaction appends a compaction entry rather than
rewriting lines; and recall keeps resolving ids whose sources left the context
because it reads the full branch via sessionManager.getBranch() and never consults
the context window. Repeated compaction renders from live records, not from the
previous summary's prose, so there is no generation-loss spiral.

Also corrects this repo's own "compaction calls no model" to the steady-state
claim it actually is: with an empty ledger the hook returns nothing and explicitly
declines ownership, and pi's native model-based summariser runs. Snippet quoted in
the doc.

And a bug shipped in the first version: §9 said the ledger entries are
custom_message and specifically not custom. Exactly backwards, so the one grep
that section existed to get right was the one it got wrong. Verified empirically
against the live session file — 11 om.observations.recorded and 6
om.reflections.recorded, all "type":"custom", next to "type":"custom_message"
entries whose customType is mempalace-mailbox and mempalace-wakeup. That is where
the confusion came from, and the distinction is load-bearing rather than
cosmetic: the mailbox uses the context-visible append API, om's ledger uses the
invisible one. Which makes it a feature the doc now advertises — the ledger costs
zero context until it is folded.
2026-08-27 14:24:19 +02:00
joakimp cdb6fc0950 changelog: name what the floating toolkit ref will pull into the next tag
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 16s
MEMPALACE_TOOLKIT_REF=main floats and docker-publish.yml resolves it to a SHA at
build time, so toolkit main moving 5b8d78f -> f0bffd1 (10 commits) ships in the
next tagged image whether or not this repo has a commit. v1.8.9 adopted the rule
after that shape bit twice (553d865, 5b8d78f): name the behaviour change BEFORE
tagging. This is that rule obeyed rather than re-learned — the work was pushed to
toolkit main earlier today and this entry was missing, which is exactly the gap
that caused a cross-host misattribution in v1.8.7.

Contents: the feeder-side secret scrubber (three tiers, T3 report-only after a
measured 403 -> 29 false-positive calibration on 52 MB of real transcripts, fail
closed); the symlink near-miss that fail-closed would have turned into a
fleet-wide silent memory outage at bake time, caught before tagging; the mailbox
work (hlc owed-set join, queued-until-next-turn note, explicit notify protocol
modes, tmux path documented unverified); and the documentation set (RFC 003,
fleet-memory.md, secret-hygiene.md). Also states what remains unscrubbed: the
opencode bridge write path and the unbuilt server-side layer.

No tag pushed — per the release protocol, no tag means no build.
2026-08-27 14:20:35 +02:00
Joakim Persson 14371e2da6 docs: explain the memory that runs itself, and fix a claim the mailbox falsified
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 16s
`pi-observational-memory` is baked, registered in the seeded settings.json, and
handed a cheaper model than the session it serves — and the only description in
this repo was five words in a feature list. Someone meeting `/om:status` or a
"compacted memory" block had nothing to read that said whether to leave any of it
on. New docs/observational-memory.md, 272 lines and five diagrams.

Scoped by who is authoritative, so there is one copy of each claim:

- Upstream (/opt/pi-observational-memory/docs/) already documents the mechanism
  well — concepts.md, how-it-works.md, configuration.md, including a v3 lifecycle
  diagram that matches the deployed code. Linked, not re-derived.
- This document takes the four facts pi-devbox owns and can change: the pinned
  commit it bakes (v3.0.4 ce9fc98, matching build-manifest.json), the packages[]
  entry that decides which copy loads, the Haiku-workers-vs-Opus-session split,
  and the devbox-pi-config volume that makes the ledger outlive the container.
- Plus the confusion this image creates by shipping two things called memory: a
  section contrasting it with MemPalace, on the line "observational memory keeps
  a session coherent, the palace keeps the fleet coherent".

Placement follows the split fleet-ops states for itself — reusable mechanism is
not deployment data — so a "why is this in my container" document belongs in the
repo that pins and wires the component. Linked twice from the README, because
until this commit the README referenced docs/ zero times and the file already
sitting there was reachable only by listing the directory.

Numbers were read out of the live container and the baked tree rather than out of
release notes, which caught one thing the pi-extensions skill still has wrong:
the dropper is gated on a successful same-turn reflection, not on a token
threshold of its own.

Also fixes a README sentence that v1.8.9 made false. § Cross-machine agent
coordination ended with "Nothing in this image polls the log on the agent's
behalf"; the mailbox has shipped since aac4a1c. Replaced with the three knobs and
their defaults, derived-not-read owed-ness, and the queued-into-the-next-turn
delivery measured on two devices — and dated to "as baked in v1.8.9
(mempalace-toolkit 5b8d78f)", pointing at RFC 003 §7.11–§7.12 for the mechanism,
because toolkit main is already ahead (a92c75d pings the human who is not
looking) and describing that here would trade a stale-behind claim for a
stale-ahead one.

The cause is worth more than the fix: the behaviour arrived through the floating
MEMPALACE_TOOLKIT_REF, so no diff in this repo ever touched the paragraph making
the claim. v1.8.9's own rule fired for the CHANGELOG and nobody swept the README.
The CHANGELOG records what changed, the README asserts what is true, and only the
first is reviewed at release time — so the rule now extends to grepping the README
for absolute claims (nothing, never, does not, only) about a component whose SHA
moved.

Diagrams were verified by rendering, not by parsing. Both comparison diagrams
parsed clean and rendered with their meaning reversed: Mermaid laid the second
declared subgraph out first, putting "with observational memory" before "without"
and MemPalace before observational memory in the diagram whose entire job was
that contrast. Rebuilt as declaration-ordered chains and re-rendered at mermaid@11
— the version pi-studio pins — in the baked headless browser, then read back as an
image. mermaid.parse() proves syntax and says nothing about layout.
2026-08-27 13:55:55 +02:00
joakimp aac4a1c323 release: v1.8.9 — the version flag that blamed the wrong component
Lint / hadolint (push) Successful in 15s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 41m49s
Publish Docker Image / smoke (push) Successful in 4m51s
Publish Docker Image / smoke-studio (push) Successful in 4m59s
Publish Docker Image / build-variant-studio (push) Successful in 16m58s
Publish Docker Image / build-variant (push) Successful in 28m35s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 17s
Two versions, two flags. `--expected-version` has only ever asserted
`pi --version`, but AGENTS.md step 4 spelled it `X.Y.Z` inside a checklist where
every other X.Y.Z is the pi-devbox tag. Run as documented for v1.8.8 the final
runtime gate of the release printed

    ✗ pi version mismatch: expected 1.8.8, got 0.84.3

and exited 1 — a red accusing the image of being the wrong version. Not one
reader's slip: the v1.8.8 release-readiness handoff from pi@emb-7kj4vr4g
propagated the same wrong spelling twice while correctly calling step 4 "not
ceremonial", so two independent readers converged on it. README.md had it right
all along, which means the two documents disagreed.

- new --expected-image-version asserts the pi-devbox release tag, read from
  release_tag in /etc/pi-devbox/build-manifest.json (no checkout, no network);
  leading `v` optional on either side
- both flags detect being handed the other one's value, and the test is exact
  rather than heuristic: the value is compared against the other quantity the
  image itself reports, so it can only fire on a real mix-up
- neither flag is required now. With none, live `pi --version` is asserted
  against the manifest's pi_version — not a tautology, since a stale pi in the
  ~/.pi/npm-global volume can shadow the baked one, exactly as a stale
  npm:pi-atelier can in packages[]
- the header note replaced was stale and load-bearing: it claimed pi is resolved
  from 'latest' and cannot be self-derived, while Dockerfile.variant pins
  ARG PI_VERSION=0.84.3 and docker-publish.yml reads that ARG as its source of
  truth. The same withdrawn claim also sat in cli_utils' pi-devbox-sanity --help
- argument parsing: a missing value, or a value that is another flag, is a usage
  error instead of silently consuming the next argument; --help works

All fourteen flag combinations exercised by execution, including the two
manifest-absent branches and the shadowing branch a healthy container cannot
reach — mutation-tested with a doctored manifest so each failure branch was
observed firing rather than assumed present.

CHANGELOG also names what no commit here causes: mempalace-toolkit main moved
e70bef2 -> 5b8d78f, so this tag ships the auto-delivered logstream mailbox
because base_tag folds the resolved toolkit SHA. It would have landed either
way; going unnamed is the 553d865 shape that already caused one cross-host
misattribution. Component audit found nothing else to bump — pi, mempalace,
pi-atelier all equal their upstream latest, and every other floating ref
resolves to the commit already baked.
2026-08-26 18:47:03 +02:00
joakimp 34cf1e3810 release: v1.8.8, and a notice that named the wrong remedy
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 19s
Publish Docker Image / build-base (push) Successful in 1h3m31s
Publish Docker Image / smoke (push) Successful in 4m47s
Publish Docker Image / smoke-studio (push) Successful in 8m8s
Publish Docker Image / build-variant-studio (push) Successful in 20m4s
Publish Docker Image / build-variant (push) Successful in 26m9s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / update-description (push) Successful in 14s
Freezes the v1.8.8 section and clears the two non-code checklist items
pi@emb-7kj4vr4g handed over (evt_20260826T134919_a614ecfc2d4f), plus the two
carried nits from its round-2 verification (evt_20260826T133356_d56792791a49).
Every claim below was re-measured here rather than taken from the handoff.

THE STALENESS NOTICE ASSERTED A DIRECTION IT NEVER TESTED — Blocker 1's shape,
one layer down, in the message I added to replace the message that named the
wrong cause. The notice fired on "recorded != HEAD" and then announced HEAD as
the newer side without testing ancestry, so a clone that was merely BEHIND got
"has moved to 82a8d3c; the snapshot describes the older 5fd0d5c" when 82a8d3c is
5fd0d5c's ANCESTOR. Found by EMB against the real state of its own host, not a
fabrication. The verdict was never wrong (rc 0, nothing mis-verified) but the
remedy it implies is a ~67-minute base rebuild when the actual fix is `git pull`
— the only one of the two carried nits with a price tag, which is why it went
first. Now tests ancestry with the merge-base --is-ancestor primitive the refresh
path 60 lines below already used, and reports three verdicts: stale (refresh),
clone behind (pull, do NOT refresh), diverged (reconcile). All three verified by
execution; only the first was correct before. --help no longer errors on the one
script whose argument order was itself a landmine. The `-s "$VENDORED"` guard is
now commented as load-bearing: it makes the empty-stdin collision unreachable by
construction, which also means no test below exercises it any more, so deleting
it as "redundant with the probes" would silently restore the false OK.

CHANGELOG: retitled, and three stale spots fixed in what becomes the permanent
record. It cited the skill at 82a8d3c (twice superseded); it RE-ASSERTED the
retracted mailbox measurement as live evidence 200 lines after withdrawing it,
which is the exact non-contradiction failure this release exists to fix; and its
warning block still described the pre-e8ddeaf world ("still records c04cd15",
"now exits 1") while quoting as exemplary the very notice whose direction was
unverified. Per EMB's steer the conclusion was kept and only the evidence
replaced: the status filter does drop broadcast noise, it just never computed
owed-ness. Honest replacement, measured on both machines: raw filter returns 2
here and 1 there, EVERY ONE already answered, derivation returns 0 for both.

SNAPSHOT RESYNCED AGAIN, 5fd0d5c -> 6eb20af, because skillset 6eb20af adds the
limit of my own seq ordering test: seq is REPLICA-LOCAL, equal to origin_seq only
because one replica authors for all four machines, so use hlc once mesh_peers
reports a peer. Recorded as reasoning not measurement — a second replica cannot
be stood up here. The durable half is the asymmetry: seq skew makes an ANSWERED
item resurface (noise, visible, self-correcting) while created_at SUPPRESSES AN
UNANSWERED ask forever (silent, permanent), so the skill now says outright that
"fixing" a resurfacing item with a timestamp trades the safe failure for the
dangerous one. The resync was free: rootfs/ was already changing, so the base
rebuild was forced regardless — the ordering warning about accidental staleness
does not apply to a deliberate refresh before the tag.

--check is a clean OK at 6eb20af with no notice, canary re-verified bidirectionally
(present 3, withdrawn 0), baked snapshot 0644, tree hash recomputed at build time
and re-verified in-container. bash -n clean; shellcheck/hadolint/actionlint remain
absent locally, so CI is still the only evidence for those.
2026-08-26 15:56:03 +02:00
joakimp e8ddeaf89f skills: a gate that could pass without checking, and a mailbox that never empties
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 1m39s
Fixes the three blockers and seven should-fixes from pi@emb-7kj4vr4g's review
(logstream correlation skills-provenance-review, full text in
drawer_pi-devbox_reviews_e43e766641c9ec85217bc6ce). Every finding was
reproduced by execution here before being fixed; two were refined by that
reproduction rather than taken as given.

BLOCKER 1 — the provenance gate could print OK and exit 0 without verifying.
`git show <ref>:<path> | sha256sum` hashes EMPTY STDIN when the ref does not
resolve, so at_ref was never empty and the UNKNOWN branch was dead code.
Measured: a bogus ref reported MISMATCH — accusing the snapshot of lying when
the real cause was an incomplete clone, and the operator's natural remedy for
MISMATCH is to re-run the refresh, which rewrites provenance to silence the
complaint; and with a 0-byte snapshot against a 0-byte upstream file it printed
"OK: exactly skillset@aaaaaaa" with exit 0 for a ref that does not exist. The
script already had the sha_empty idiom and had applied it to blob_sha but not
to at_ref. Existence is now PROVEN with git cat-file -e before anything is
hashed, at two levels (ref resolves / path exists at it) because those deserve
different messages. Same defect class as the canary it replaces: a check that
can succeed without checking. A second, unflagged instance of the same pipeline
shape in blob_sha was found and fixed too.

Exit codes split, because the old contract failed the sanctioned case: 0
truthful (including stale, with a NOTICE), 1 a lying record only, 2 cannot
determine. AGENTS.md step 2 promised "the message distinguishes the two" and
was the thing this branch was breaking; rewritten to state all three.

BLOCKER 3 — VENDORED.md contradicted itself in the release whose stated
invariant is non-contradiction: its hand-maintained provenance line named
skillset 670f7f1, seven commits behind the ARG and itself the commit that told
agents to hand-stamp added_by — the withdrawn instruction this work exists to
stop shipping — while its cp recipe contradicted the "not cp" rule 20 lines
above. Line removed (nothing forced it to move when the ARGs did); 670f7f1 kept
only as a labelled cautionary example. The pi-extensions half was verified
redundant (CI require_sha resolves PI_EXTENSIONS_REF) before removal.

SHOULD-FIXES: `<root> --check`, the spelling VENDORED.md documented, silently
ran a REFRESH because only $1 was parsed (both tools now parse all args and
reject unknown ones); refresh at a detached/older HEAD silently rewound ref and
bytes (now refused unless the recorded ref is an ancestor, --force to override);
upstream_dirty was computed and never used in check mode; --no-skills --json
printed human text and broke jq; --help was a hardcoded sed range this branch
had already made stale; the fingerprint hashed SKILL.md alone so a live skill
differing only in a sibling file reported "identical", and pi-extensions already
ships two files, so it is now a per-skill TREE hash with the manifest field
renamed skillset_snapshot_tree_sha256; the --no-skills smoke assertion was
negative-only and passed on a crashed binary. mktemp+mv left files 0600 — CI was
unaffected since the index records 100644, so the blast radius was local builds
only, narrower than the review inferred.

Snapshot resynced c04cd15 -> 5fd0d5c so the no-clone fallback carries the
CORRECTED coordination protocol rather than the withdrawn one; --check is now OK
with no staleness notice, and the bidirectional canary re-verified against the
new bytes. Local validation is bash -n only (shellcheck, hadolint and actionlint
are all absent in this container) — CI remains the shellcheck gate.
2026-08-26 14:38:07 +02:00
joakimp 49a6534093 docs: write down the coordination channel the fleet already runs on
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 15s
The logstream has carried cross-machine work since 2026-08-18 — patch handoff,
review, a v1->v2 supersede — and nothing in this repo said it existed. That gap
had a measurable cost this morning: another host addressed a retraction to
pi@tor-ms22 by name and it was read only because the human said "read the
logstream", while the agent was actively rebuilding the thing it warned about.

Split by what each document is authoritative for, so there is one copy of each
claim rather than three that drift:

- README § Cross-machine agent coordination — what the CONTAINER needs.
  MEMPALACE_REMOTE_URL selects the shared palace; MEMPALACE_PI_DEVICE is what
  makes this machine reachable, because where every host is a thin client of one
  palace the stamped agent name is the only thing that distinguishes them. Stated
  as a rule with teeth: set both or neither, since a container missing the device
  var can read the log but is addressable by nobody.
- AGENTS.md release checklist step 2 — the vendored-snapshot refresh, as a
  MECHANISM in the document a releasing agent actually reads, not a comment
  hoping to be noticed. It says the refresh costs a base rebuild, that skipping
  it is legitimate (every enrolled host reads its live clone), and that skipping
  it silently is not.
- CHANGELOG — the three-way split itself, plus the measurement that shaped the
  ack contract: unfiltered, the mailbox returned 5 events, 4 of them finished
  broadcasts from eight days earlier; with status="open", exactly the 1 that
  needed an answer.

Norms live in the skillset skill (82a8d3c, already live on every host that mounts
the skillset — no rebuild) and mechanism in mempalace-toolkit's
extensions/pi/README.md (e70bef2, which also documents the edge stamper that
553d8657 shipped undocumented). Deliberately NOT duplicated here.

Consequence recorded rather than hidden: the skill edit lands in the skillset, so
this repo's SKILLSET_SNAPSHOT_REF now honestly reports itself behind, and
--check exits 1 with "has moved to 82a8d3c; the snapshot describes the older
c04cd15". That message is also fixed in this commit — it previously blamed "the
working tree" even when the tree was clean and only the ref had moved, which is
the same defect class as a canary pinned to a phrase the release deleted: a
message that names the wrong cause. Now distinguishes moved-HEAD from dirty-tree,
verified against both plus the in-sync case.
2026-08-26 12:51:26 +02:00
joakimp e070e0bcbf skills: record the vendored snapshot's provenance, and report which copy wins
Lint / actionlint (push) Successful in 16s
Lint / hadolint (push) Successful in 16s
Found while verifying v1.8.7 from inside a fresh container: the baked mempalace
snapshot is read by no host on this fleet. devbox-skill-reconcile repoints
~/.agents/skills/mempalace at the mounted live clone (the v1.8.5 fix working as
designed), and all four compose stacks mount a workspace containing the
skillset. So the phrase canary that blocked v1.8.7's first tag polices a file
nobody opens, while the drift that could actually mislead an agent — a git pull
nobody ran in /workspace/skillset — was invisible from inside the container and
is invisible to CI by construction.

Record provenance instead of policing it, and move the check to where the
skillset actually is:

- Dockerfile.variant: ARG SKILLSET_SNAPSHOT_REF (the claim) + a sha256 of the
  shipped bytes measured in the manifest layer (the fact), as manifest siblings
  rather than components{} members, plus an OCI label. An ARG default, not a
  CI-resolved output: no credential for the private skillset, no change at any
  of the four variant build call sites, and a local docker build records what CI
  does. Variant-only, so no base rebuild — check-base-hash.sh scans
  Dockerfile.base alone, verified by running it.
- pi-devbox-version: a skills: section naming baked vs live <repo> @ <sha> per
  vendored skill, and for mempalace whether the live copy is identical to the
  baked fingerprint, at the same commit with uncommitted edits, or divergent.
  entrypoint-user.sh passes the new --no-skills, because the banner prints
  before the links exist and long before the reconcile runs.
- scripts/vendor-mempalace-skill.sh: refresh the file and rewrite the ref
  together (a cp without an ARG bump makes the manifest lie, which is worse than
  anonymity); --check verifies the claim against a real clone.
- 5 new smoke assertions (78 -> 83), mutation-tested through the real sh -c
  path: 6 fabricated manifests, where a well-formed hash of the wrong file
  proves the two manifest assertions are not redundant; the all-baked reporting
  test verified to FAIL against a live-skillset environment.

Reviewed mid-flight by pi@emb-7kj4vr4g over the logstream (correlation
skillset-vendor-drift), which retracted its own earlier recommendation of a
build-time byte-compare against skillset HEAD and supplied the better framing:
the invariant is NON-CONTRADICTION, not currency. Byte parity on a fallback
would have cost a resync commit plus a ~67-min base rebuild for each of the four
skillset commits pushed in one evening. Its warning also found a real bug here:
the script now CONSTRUCTS the snapshot from `git show HEAD:<path>` instead of
copying the working tree, because a clean `git diff` says nothing about an
untracked file — the one input the first draft would have recorded a false ref
for. Tested: untracked, unstaged and staged-but-uncommitted all refuse, atomically.

Also fixes three stale in-repo markers of the same class the canary belongs to
(true when written, silently false at release): two dangling "Unreleased"
pointers and a typst line still marked Unreleased five releases after v1.4.0.
2026-08-26 10:30:27 +02:00
pi dbb78798fb vendor: resync mempalace skill snapshot to skillset c04cd15
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 40s
c04cd15 ('the withdrawal only holds where the bridge is live') landed after the
v1.8.7 snapshot was taken, so the baked fallback was already 4 lines behind the
skillset within hours of publishing. It adds the caveat this fleet is currently
living in: the bridge is baked at image build, so a container on an image older
than the stamping commit satisfies both env gates while stamping nothing, and
hand-stamping is still the only signal a hand-filed drawer gets there. It also
gives the one-line test —
  grep -c MEMPALACE_PI_DEVICE "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)"
which returns 0 on this v1.8.6 container, confirming the gap empirically.

Note what this instance proves about the canary fixed one commit ago: it still
PASSES on the refreshed copy, because both pinned phrases survived the edit. A
phrase canary cannot detect 'older than skillset main' — only a diff can. This is
the second drift in 24h and is the argument for the Still-open item (a CI job
diffing this file against the skillset repo, blocked on a clone credential for a
private repo). No re-pin was needed here.
2026-08-26 09:41:02 +02:00
pi f645e6654f smoke: fix the snapshot canary that blocked v1.8.7, and make it bidirectional
Lint / hadolint (push) Successful in 15s
Lint / actionlint (push) Successful in 27s
Publish Docker Image / resolve-versions (push) Successful in 1m5s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-base (push) Successful in 41m8s
Publish Docker Image / smoke (push) Successful in 4m49s
Publish Docker Image / smoke-studio (push) Successful in 18m28s
Publish Docker Image / build-variant (push) Successful in 15m46s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / update-description (push) Successful in 20s
Publish Docker Image / build-variant-studio (push) Successful in 16m52s
Run 589 built the base cleanly and then failed both smoke jobs 81-passed/1-failed
on 'mempalace skill snapshot is current'. That canary greps a phrase from the
vendored mempalace skill to detect a stale snapshot, and the phrase it pinned was
'Attribute what you file yourself' — the heading of the hand-stamping instruction
that THIS release withdraws. So it fired correctly: the snapshot changed and the
expectation did not. Every publish job was skipped, so nothing reached the
registry and v1.8.7 was never consumed.

Rather than bump the string:

* the assertion is now BIDIRECTIONAL — the new phrase must be present AND the
  withdrawn one absent. A one-way canary only catches half the drift: it cannot
  notice a re-vendored stale snapshot that happens to contain the pinned phrase.
  Verified against v1.8.6's snapshot, which now correctly fails.
* the comment records the structural limit rather than just the fix: a phrase
  canary can only ever detect 'older than what I remembered to pin', never
  'older than skillset main'. Only a diff against the skillset repo can do that,
  which is now a Still-open item — it needs a CI clone credential for a private
  repo, i.e. a policy decision, not a code change.

Changelog consolidated: the SSH sidecar multiplexing default moves from
Unreleased into v1.8.7, since the retag will sit on a commit that contains it,
and the v1.8.7 summary now records the failed first attempt rather than quietly
presenting the second one as the whole story.
2026-08-26 08:09:09 +02:00
pi 657b1ad856 ssh sidecar: default to multiplexing, as a default and not an override
Lint / actionlint (push) Successful in 15s
Lint / hadolint (push) Successful in 16s
A target whose ~/.ssh/config entry never mentioned ControlMaster got no
multiplexing from the sidecar (only ControlPath was supplied), so every ssh call
opened a fresh TCP connection. On 2026-08-25 that produced ~12 connections to
one host in 15 min and a fail2ban block that looked like an outage — the tell
being that HTTPS to the same estate stayed healthy.

The correctness of this depends entirely on WHERE the block goes. ssh_config is
first-value-wins:

  ControlPath   before the Include -> override (the user's value points at
                read-only ~/.ssh and cannot work in the container)
  ControlMaster after  the Include -> default  (an explicit per-host
                'ControlMaster no' must keep winning)

Force what is broken, default what is merely absent. The first draft put both in
the leading block and would have silently overridden an explicit 'no'.

Verified with ssh -G rather than from the man page, including the counterfactual:
under the shipped layout an explicit 'no' resolves to controlmaster false while a
silent host resolves to auto; under the rejected layout the 'no' host flips to
auto. So the test discriminates position, not presence. Plus a sandbox render of
the real script, bash -n, and shellcheck -S error (the v1.8.7 gate) clean.

Effect measured on 41 real host aliases: 22 silent entries gain auto+10m, 0
overridden. Note the fleet's one deliberate opt-out is written as absence plus a
comment ('# No ControlMaster — VPN means direct route'), which ssh cannot
distinguish from no opinion; that host now multiplexes, which its own comment
says is unnecessary rather than harmful.

Skill documents the sidecar-vs-~/.ssh trap (the failure misleads: read-only
ControlPath makes multiplexing look impossible rather than misconfigured) and
the stale-master recovery, ssh -O check / -O exit.
2026-08-25 23:09:46 +02:00
pi ebd0de0be2 changelog: v1.8.7 — device provenance reaches the fleet
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-base (push) Successful in 42m4s
Publish Docker Image / smoke (push) Failing after 4m44s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Publish Docker Image / smoke-studio (push) Failing after 7m46s
Publish Docker Image / build-variant-studio (push) Has been skipped
The provenance fix's client half lives in mempalace-toolkit, which the image
clones at build time, so it only reaches the fleet through a tag. Records both
routes (extension via MEMPALACE_TOOLKIT_REF folded into base_tag; vendored skill
via rootfs), the three design points (stamp in the client not the agent; diary
marker in TEXT because metadata is invisible to readers; solitary devbox stamps
nothing), and why the allowlist is per tool (3.8.0 hard-fails -32602 on
undeclared args). Carries the CI-hardening work already sitting in Unreleased,
and a Still open block for the three known bounds.
2026-08-25 22:47:53 +02:00
pi 4f1aa0d0dd skills: refresh the vendored mempalace snapshot (withdrawn hand-stamping)
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 19s
VENDORED.md's freshness model for `mempalace` is "Option 2 only — refreshed
manually per release", and it had drifted since 2026-08-23. The stale snapshot
still carried the instruction to hand-stamp added_by="<harness>@<device>", which
skillset 73c7c8e withdrew: the pi bridge now stamps at the edge
(mempalace-toolkit 553d865), and RFC 001 §7.3.2 ranks agent-side stamping ❌
worst-possible.

That matters specifically for the fallback case this snapshot exists to serve — a
container started WITHOUT the private skillset mounted would otherwise be the
only kind of container still being taught to do it by hand.
2026-08-25 22:27:04 +02:00
joakimp 9e744d701f lint: shellcheck the repo's own shell scripts, not just workflow run: steps
Lint / actionlint (push) Successful in 15s
Lint / hadolint (push) Successful in 15s
lint.yml has shellchecked every workflow `run:` step since the dash-vs-bash
incidents, but nothing had ever pointed shellcheck at entrypoint.sh, scripts/*.sh
or the extensionless tools under rootfs/usr/local/bin/. That gap is not
hypothetical: the skillset repo's ci-release-watcher template shipped
`echo "$json" | python3 <<'EOF' ... json.load(sys.stdin)` for two months, where
the heredoc IS python's stdin (no script arg) so the load hit EOF and the
function silently returned nothing. shellcheck names exactly that at severity
ERROR — SC2259, "This redirection overrides piped input" — and could have named
it the whole time.

New step in the existing actionlint job, so no second container pull: shellcheck
-S error plus bash -n over every shell file, discovered as *.sh UNION a shebang
scan (the glob alone misses pi-devbox-version, devbox-skill-reconcile, dot-watch
and studio-expose; a shebang scan alone would miss a sourced fragment without
one). Fails loudly on a zero-file match, because a green tick over an empty set
is not a check.

Severity chosen by measurement, not taste: -S error is 0 findings across all 11
shell files today, so the gate is green on arrival with no cleanup, while
-S warning is NOT free (19x SC2088 tilde-in-quotes in recreate-sanity-check.sh
plus assorted SC2016, all intentional) and would train everyone to ignore the
job — the same reasoning as the SHELLCHECK_OPTS exclusions already on the
actionlint step.
2026-08-25 21:11:55 +02:00
joakimp 2b8c3a4db4 ci: audit MEMPALACE_VERSION the way PI_VERSION is audited
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 17s
Closes the item v1.8.6 (and v1.8.5 before it) listed as "Still open": the
palace pin was a literal string in Dockerfile.base with zero references in
docker-publish.yml, while PI_VERSION had a concreteness gate, a
published-on-registry check and a never-silently-adopt drift warning.

resolve-versions now applies all of those to MEMPALACE_VERSION, read from
Dockerfile.base so a local `docker build` and CI install the same version by
construction, plus one gate pi does not need: a YANKED release is refused,
because an exact pin installs one silently under PEP 592 and would have
shipped a withdrawn palace client to the whole fleet.

smoke gains `installed mempalace matches CI's audited pin` via a new
EXPECTED_MEMPALACE_VERSION threaded into both smoke jobs. It is not redundant
with `manifest mempalace_version matches the installed core`: that compares two
properties of one image and cannot notice that both are the wrong version. The
case this covers is a variant built FROM a cached base carrying an older pin —
internally consistent, silently stale.

Mutation-tested by extracting the shipped block out of the YAML and stubbing
curl: 9 cases covering every gate, then once end-to-end against live PyPI. That
found a real defect in the first draft — the yank message inlined a jq program
inside $(...) inside a double-quoted string, where the escaping broke the
filter (jq compile error) while the surrounding `exit 1` still fired: a gate
that looked correct and reported garbage.

Note: correcting Dockerfile.base's now-false "known gap, carried forward"
comment forces a base rebuild (~67 min) on the next tag. Leaving a comment
asserting the audit does not exist was the worse option.
2026-08-25 20:33:19 +02:00
joakimp cb7b8ad2ae smoke: assert manifest VALUES, not the presence of field names
Lint / actionlint (push) Successful in 16s
Lint / hadolint (push) Successful in 17s
Four build-provenance assertions grepped the manifest for a field name and
never looked at the value:

    run_expect "manifest records pi_version" "cat …manifest.json" '"pi_version"'

which passes on {"pi_version": ""} and on {"pi_version": null}. The tell was
in its own passing output the whole time — `✅ manifest records pi_version (got
"pi_version")` echoes the key back as the thing it claims to have found. Found
while reading run 579's smoke log to confirm v1.8.6's new assertions had really
executed rather than merely gone green.

Now checked against values, and against ground truth where it exists:

- every required component key present, naming the one that vanished
- every component value a full 40-hex SHA (null allowed for pi-studio alone,
  which is legitimately absent in the non-studio variant)
- pi_version equal to `pi --version`, mirroring the mempalace ground-truth check
- release_tag non-empty; source_revision 40-hex and build_date ISO-8601 *when
  populated*, since both default empty on a plain local `docker build` and
  demanding them would fail honest local smoke runs
- --json compared byte-for-byte with the file, which is assertable because that
  mode is a verbatim cat; the old form grepped its output for "release_tag"

Key presence and value shape are deliberately SEPARATE assertions: a single
"all values are valid SHAs" loop passes vacuously on components:{}, because
jq's all() over an empty list is true. Combining them would reproduce the same
shape of hole as the three false greens already recorded in CHANGELOG.md.

Dropped `manifest has no unresolved ('unknown') components`: the 40-hex check
strictly subsumes it ("unknown" is not 40-hex, and only rev() emits it, feeding
components{} exclusively). Removed rather than kept, because a check that can
no longer fail independently is one more green tick that means nothing.

Mutation-tested twice rather than reasoned about: nine fabricated manifests
through the raw jq filters, then twelve through the shipped assertions using
the real `run` helper's `sh -c` quoting path — the quoting is load-bearing,
since a jq filter dying on a quoting error exits non-zero and looks exactly
like a caught defect. Measured on the same twelve defects: old caught 3, missed
9; new catches 12. Three legitimate variations stay green (empty
source_revision, empty build_date, null pi-studio).

Also corrects a factually wrong "Still open" bullet in the released v1.8.6
entry, which claimed pi-devbox-version's human output does not show
mempalace_version and that only --json surfaces it. Both halves are false: it
prints a `palace:` line with live-vs-baked drift annotation, verified against
fabricated manifests (match, skew, and pre-v1.8.6 absent-field cases). Left as
a struck-through correction rather than deleted, since v1.8.6 is published.
2026-08-25 17:15:39 +02:00
joakimp 93f986e90e v1.8.6: adopt pi 0.84.3 + mempalace 3.8.0, close the v1.8.5 doc/observability gaps
Lint / hadolint (push) Successful in 9s
Publish Docker Image / resolve-versions (push) Successful in 15s
Publish Docker Image / base-decide (push) Successful in 8s
Lint / actionlint (push) Successful in 1m9s
Publish Docker Image / build-base (push) Successful in 41m23s
Publish Docker Image / smoke (push) Successful in 4m50s
Publish Docker Image / smoke-studio (push) Successful in 5m5s
Publish Docker Image / build-variant (push) Successful in 15m46s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 15s
Publish Docker Image / build-variant-studio (push) Successful in 19m59s
Three coupled pieces of work, all of which ride on the base rebuild that the
mempalace bump forces anyway.

DRIFT ADOPTED
- pi 0.84.2 -> 0.84.3. Its release notes carry a "Breaking Changes" line
  (GoogleThinkingLevel -> GoogleApiThinkingLevel). Audited before adopting:
  zero references across all four vendored companions (pi-fork,
  pi-observational-memory, pi-atelier, pi-studio), so it is inert for us. The
  reason to adopt is two skill-discovery fixes that land directly on v1.8.5's
  vendored-skill work: nested Markdown skills inside grouping directories were
  not discovered, and root README.md/AGENTS.md in skill dirs were reported as
  broken skills.
- mempalace core 3.7.1 -> 3.8.0. Additive/reliability only. Its sync fix
  (#2320/#2322) stops sync --apply deleting drawers whose source_file was
  unreachable *at that moment* -- which does NOT relax the standing landmine
  against sync on the shared palace, because that landmine is about paths
  permanently absent from whichever host runs the sync. Different failure
  shape; the caution stands.

DOCS -- three defects, one of them public
- DOCKER_HUB.md advertised "neovim (LazyVim defaults)". Nothing in the image
  installs LazyVim; the only nvim config is a 19-line sysinit.vim. CI PATCHes
  this file into the Docker Hub description on every release, so this was a
  false claim published to the world. Removed.
- agent-browser + Playwright + Chromium is the single largest addition in the
  image (~625 MB) and had zero mentions in README, DOCKER_HUB or THIRD_PARTY --
  it was documented only to agents, in the AGENTS.md managed block. Now
  documented to humans, including the Chromium licence dimension.
- typst and socat appeared in README prose but not in the "What's inside"
  inventory. Added.

OBSERVABILITY -- the three gaps v1.8.5 listed as still open
- build-manifest.json now records mempalace core, read from the live binary
  (ground truth, not the build ARG). Placed as a sibling of pi_version rather
  than inside components{}, because pi-devbox-version renders that map through
  [0:12] and would truncate a version string.
- smoke asserts the pi-observational-memory clone actually CONTAINS the ce9fc98
  auth fix, pinned to src/runtime.ts. Deliberately not a repo-wide grep: two of
  the three markers also live under tests/, so the repo-wide form stays green
  with the fix site reverted. That is the third false-green of this exact family
  in this repo (canary phrase in both snapshots; reconciler fixture using a
  non-owned name; now this) -- pin containment checks to the fix site.
- smoke asserts the feeder's pi@<device> agent default behaviourally. The
  earlier audit concluded this needed a --print-config added upstream; it does
  not. AGENT is assigned before arg parsing, so `bash -x mempalace-pi-session
  --help` observes the real resolution with no toolkit change. Two-sided:
  device set => pi@<device>, unset => must not be pi@*.
- pi-devbox-version now prints a palace: line with the same live-vs-baked drift
  detection pi already had. This matters more than it looks: mempalace is the
  one component that is both client (here) and server (synlig), so skew between
  them is a real failure mode. Degrades quietly on pre-v1.8.6 images.

Deferred deliberately: a native arm64 act_runner on tor-ms22 (the current
runner is on synlig, x86_64, so every arm64 layer ships QEMU-emulated).
Analysis and caveats filed to the palace rather than actioned here.
2026-08-25 15:29:22 +02:00
joakimp 26f223568d .env.example: document MEMPALACE_PALACE_PATH and why nothing exports it
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 50s
The only MemPalace variable the template never mentioned, and the one that
moves the feeders' stage as a side effect: the palace root resolves as
$MEMPALACE_PALACE_PATH -> $MEMPAL_PALACE_PATH -> ~/.mempalace/config.json ->
~/.mempalace/palace, and the stage is derived from it (<palace-root>/pi-stage).

The comment states that precedence, records why neither the image nor the
entrypoint exports it (pinning the palace without carrying the stage along
re-creates the split a shared root removed, v1.8.2), warns that a stage whose
persistence differs from the palace makes a scoped `mempalace sync` prune
conversation drawers whose dedup key is the staged path, and notes it is a
path INSIDE the container unlike the host-side WORKSPACE_PATH/SSH_KEY_PATH
above it.

Found while auditing a live host whose .env sets it redundantly to the
default value.
2026-08-23 23:27:42 +02:00
joakimp 01abda3456 v1.8.5: the skillset owns its skills, and a dangling link no longer kills boot
Publish Docker Image / base-decide (push) Successful in 18s
Publish Docker Image / build-base (push) Successful in 41m59s
Publish Docker Image / smoke (push) Successful in 7m31s
Publish Docker Image / smoke-studio (push) Successful in 20m56s
Publish Docker Image / build-variant (push) Successful in 16m11s
Publish Docker Image / promote-base-latest (push) Successful in 15s
Publish Docker Image / update-description (push) Successful in 14s
Publish Docker Image / build-variant-studio (push) Successful in 17m32s
Lint / hadolint (push) Successful in 8s
Publish Docker Image / resolve-versions (push) Successful in 15s
Lint / actionlint (push) Successful in 20s
No pin moved except mempalace-toolkit fd8b15f5 -> 0fe64c4 (feeder defaults
--agent to pi@$MEMPALACE_PI_DEVICE, so hand-filed palace writes carry
provenance; $USER remains the fallback when the var is unset). pi 0.84.2 and
mempalace 3.7.1 are still upstream-latest, atelier v0.8.2 keeps the >=0.7.1
floor for pi >=0.84, and om master is still ce9fc982 -- nothing landed after the
merge that fixed the eight-week silent-observation bug.

Expect a full multi-arch base rebuild: entrypoint-user.sh, rootfs/** and the
resolved toolkit SHA all feed the base hash.
2026-08-23 21:00:21 +02:00
joakimp b5810654f6 skills: let the skillset own the skills it owns, and stop a dangling link from killing boot
Lint / actionlint (push) Successful in 16s
Lint / hadolint (push) Successful in 20s
Baked skill links won over the live skillset clone for all three vendored
skills, so a pushed edit to skills/mempalace/SKILL.md was invisible in every
container until the next image build -- measured on two hosts (live md5
129bcc4752 vs baked 5236024fef). Cause was ordering, not intent: the baked links
are created early with a create-only-when-absent guard to close a smoke
readiness race, and the skillset deploy runs last and treats them as foreign.
The comment claimed the opposite of the behaviour.

The fix is not "skillset always wins". Ownership is per-skill: pi-extensions is
owned by its package repo and copied over the snapshot at build time, so the
skillset's lagging duplicate must keep losing; pi-devbox-environment is authored
here. Only mempalace is skillset-owned. devbox-skill-reconcile therefore runs
after the deploy and repoints only the names in skills/skillset-owned.txt,
replacing a link solely when it points into the baked tree, so a real directory
or a link pointing elsewhere is never disturbed. Precedence is now user override
-> live clone (owned names) -> baked snapshot, with the early links intact as
the fallback so the readiness race stays closed.

Reviewing that turned up a latent boot-abort in the pre-existing baked-link
block: `[ ! -e "$link" ]` is TRUE for a dangling symlink, so once a link can
point into /workspace/skillset, a vanished mount makes plain `ln -s` fail with
"File exists" -- and under `set -euo pipefail` that aborts container start
before `exec "$@"`. Reachable on `docker restart` or a host reboot, not on a
recreate, since ~/.agents is not a volume on any host. Now `ln -sfn`, which
heals the link back to the baked fallback.

Smoke additions cover what let this ship: the stale-snapshot canary grepped a
phrase present in BOTH the stale and fresh copies, so it passed throughout;
it now pins the newest section. Link targets are asserted, not just `test -L`;
the owned-list content is asserted both ways; and the reconciler's replace path
-- which no CI container exercises, since none mounts a skillset -- is covered by
fabricating one. A mutation test showed the obvious three assertions still pass
with the "is this link ours?" guard deleted, so a discriminating case was added:
an owned name whose link is a user override outside the baked tree.

Also refreshes the mempalace snapshot to skillset 670f7f1 (without it the fix
helps only hosts that mount skillset) and corrects README, which documented the
old, wrong precedence in three places.

Verified with 12 fixture cases plus 2 mutants: ownership respected against the
real trees, user overrides preserved, relative/trailing-slash/CRLF/space/glob
inputs handled, dangling link healed, read-only skills dir exits 0, idempotent.
2026-08-23 20:59:14 +02:00
Joakim Persson 4f6f470518 changelog: file the vendored-skill shadowing bug for v1.8.5
Lint / actionlint (push) Successful in 21s
Lint / hadolint (push) Successful in 1m24s
~/.agents/skills is asymmetric: the three pi-devbox-specific skills resolve to
the baked copies while all others resolve to the live skillset clone. Root cause
is ordering in entrypoint-user.sh -- baked links are created early (line 65) with
a create-only-when-absent guard, and the skillset deploy runs last (line 387) and
leaves them alone as foreign links. The comment at line 61 states the intent as
protecting the skillset skill from being clobbered, but the effect is the
reverse.

Cost measured on two hosts: a pushed edit to the mempalace skill (live md5
129bcc4752) was invisible to both containers, which kept loading the baked copy
(md5 5236024fef). Editing those three skills appears to work and silently does
nothing until a rebuild.

Filed as a known issue with a proposed fix rather than fixed here: changing
symlink precedence is image behaviour and wants its own review plus a smoke
assertion, and the early-link ordering exists to close a readiness race that
must not regress.
2026-08-23 20:05:38 +02:00
Joakim Persson fbc1f86612 docs: a negative result is usually your own filter (skill + AGENTS.md)
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 16s
Three false negatives in one session, all self-inflicted, all convincing
because the command "succeeded": a `| head -20` proved an SSH peer absent
that sits at line 454 of a ~500-line config; `ssh mac 'docker ps'` proved
the host had no Docker, when the non-interactive PATH simply lacks
/usr/local/bin; and `grep 'ssh '` proved no ControlMaster was running,
when those processes rename themselves to `ssh: <path> [mux]`. Same root
cause each time, so it goes in the skill rather than in a commit message:
a positive result carries its own evidence, absence has to be earned.

The skill (rootfs/, symlinked into ~/.agents/skills) is BAKED, so this is
an image change and is logged in CHANGELOG Unreleased accordingly. Its §3
also now records that a live ControlMaster socket makes later commands
authenticate not at all -- after editing a peer's authorized_keys, "it
still works" proves nothing; prove it with -o ControlPath=none, or the
breakage waits for a future session that has no memory of the edit.

AGENTS.md: corrected a stale CI claim while placing the pointer. It said a
tag push produces two runs including lint; lint.yml has since been scoped
to branches: ['**'], which excludes tag refs, and refs/tags/v1.8.4 duly
produced run 571 (publish) and nothing else. Kept the head_sha + workflow
path filter advice, which is cheap and guards against a future v*-triggered
workflow. Added a short section on verifying this repo from inside a
container, including that docker-compose.yml here is a TEMPLATE pinning
:latest while a real host runs its own per-machine file -- recreating from
the repo copy can silently move a host off :latest-studio.

Placement note: AGENTS.md is only auto-read when the cwd is this repo, so
the durable rule lives in the skill, which loads by description match in
any pi-devbox session.
2026-08-22 22:56:41 +02:00
Joakim Persson 2ebf00d6d4 v1.8.4: the om fix lands upstream, pi-atelier v0.8.2, todo edit
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 22s
Publish Docker Image / resolve-versions (push) Successful in 19s
Publish Docker Image / base-decide (push) Successful in 11s
Publish Docker Image / build-base (push) Successful in 42m16s
Publish Docker Image / smoke-studio (push) Successful in 5m0s
Publish Docker Image / smoke (push) Successful in 19m0s
Publish Docker Image / build-variant-studio (push) Successful in 17m38s
Publish Docker Image / build-variant (push) Successful in 28m36s
Publish Docker Image / update-description (push) Successful in 8s
Publish Docker Image / promote-base-latest (push) Successful in 10s
Headline: pi-observational-memory 37986b6 -> ce9fc98. The ambient-credential
gate fix (6f694e6 + 699ccc7) was merged upstream as PR #52 on 2026-08-22,
closing issue #51, so this release picks it up through the ordinary
PI_OBSMEM_REF=master path with nothing carried locally. Every image up to
and including v1.8.3 silently recorded zero observations on a Bedrock host
using ambient AWS credentials; from this one on, /opt is the fix and the
settings.json packages[] workaround should be deleted (verify against
/etc/pi-devbox/build-manifest.json first). npm still ships the broken 3.0.4,
which does not matter here because the image clones the ref instead.

Pin bump: PI_ATELIER_REF / PI_ATELIER_VERSION v0.8.1 -> v0.8.2, audited per
the floor note above the ARG. The only version in between is 0.8.2 itself and
both its entries are Workspace-Pulse-internal (inspection coalescing and
serialization; fresh inspection guaranteed at Turn end, retired sessions can
no longer publish stale results). Nothing touches pi's private TUI renderer,
which is the coupling behind the 0.6.0/0.7.0-under-pi-0.84 startup hang, and
pi is unchanged at 0.84.2 - so the bump stays outside that risk class.

Also baked by this build, no pin needed: pi-extensions 98eb07b -> 2022887
(the todo edit action), pi-fork 4a09af4 -> f1ff808, pi-studio 0.9.44 ->
v0.9.48, mempalace-toolkit b609cf5 -> fd8b15f (docs only - the pi-session
false-success guard was already baked in v1.8.3, confirmed by ancestry),
aws-cli 2.36.24 -> 2.36.29 and the other *_VERSION=latest tools.

README: the pin table also said mempalace 3.6.0, stale since v1.8.3 bumped it
to 3.7.1. Fixed in passing.

Unchanged and verified current: pi 0.84.2 (npm latest, published 2026-08-14),
MEMPALACE_VERSION 3.7.1 (PyPI latest), pi-toolkit 0e1369e.
2026-08-22 21:19:48 +02:00
joakimp c3b6d36778 docs(CHANGELOG): Unreleased — the todo extension gains an edit action
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 23s
The commit itself lives in pi-extensions (2022887), but /opt/pi-extensions
is baked into this image, so "which todo behaviour does this image have" is
an image question. PI_EXTENSIONS_REF=main means the next build picks it up
with no pin to bump -- worth stating explicitly, since a reader who expects
a version bump will otherwise go looking for one.

Also recorded, because it confused a session today: pi-atelier's
tool_result hook replaces todo output with "N/M done - see sidebar"
whenever the sidebar panel is visible, so an agent sees the counter and not
the item text. Upstream's list returns every item; the terseness is
atelier's deliberate context saving, not a limitation of the tool.
2026-08-17 22:49:13 +02:00
Joakim Persson 3a509077c2 v1.8.3: mempalace 3.7.1, refreshed skill snapshot, census on PATH
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 23s
Publish Docker Image / resolve-versions (push) Successful in 13s
Publish Docker Image / base-decide (push) Successful in 11s
Publish Docker Image / build-base (push) Successful in 53m49s
Publish Docker Image / smoke (push) Successful in 7m22s
Publish Docker Image / smoke-studio (push) Successful in 18m6s
Publish Docker Image / build-variant (push) Successful in 19m6s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / build-variant-studio (push) Successful in 27m5s
mempalace 3.6.0 -> 3.7.1. Verified against the 3.7.1 source rather than its
changelog, because the risk lands on palaces users cannot reconstruct: legacy
drawers lack the new chunk_total marker and both decision sites trust them, so
no mass re-mine; NORMALIZE_VERSION is 2 in both; chromadb stays <2 so no
index-format migration; no auto-migration exists; logstream.sqlite3 is created
lazily. Downgrade remains possible (3.6.0 has zero references to chunk_total).

Two behaviour changes documented in the CHANGELOG: ALLOW_PEER_WRITER no longer
works on local/chroma palaces, and writer-lock setup failures fail closed.
Neither affects this image's MCP-server-plus-CLI-feeder pattern, which already
serialised on the same lock under 3.6.0 -- the upstream "process-lifetime
single-writer" entry describes tightened escape hatches, not a new lease.

The motivation is the shared central palace: 3.7.1 drops the stale chromadb
SharedSystemClient cache on reconnect (3.6.0 could let a stale in-memory HNSW
segment overwrite a peer's writes, "index count going backwards"), releases the
writer lease on SIGTERM/SIGHUP, and stops treating an interrupted mine as
complete. The fleet primary was upgraded to 3.7.1 and restarted before this tag,
because 3.7.1 refuses writes when the served library drifts and reconnect cannot
clear that. opencode-devbox still pins 3.6.0, so the lockstep is broken until it
cuts its own release.

Vendored mempalace skill snapshot refreshed to skillset 936fed8 (was 63f3bf5).
This closes a gap that had been invisible for two commits: ~/.agents/skills/
mempalace symlinks to the IMAGE-BAKED copy, entrypoint-user.sh creates that link
first, and the skillset deploy never clobbers an existing name -- so in a devbox
container the vendored snapshot always wins and editing skillset alone changes
nothing a container reads. Brings the multi-machine shared-palace section and
the hand-crafted-provenance guard.

pi-global-AGENTS.append.md already carried the three shared-palace damage rules
(a55f636); this tags them into a release.

mempalace-census gets the /usr/local/bin symlink its three siblings have had all
along, plus chmod and a build-time --help smoke check, so RFC-002 Phase A
censuses no longer need an absolute path.

Two new smoke assertions, both confirmed to FAIL against a v1.8.2 container so
they are not tautological: the vendored snapshot must contain the multi-machine
section (a stale manual snapshot is otherwise invisible), and mempalace-census
must be on PATH.

No other pins move: pi stays 0.84.2 (npm latest), pi-atelier v0.8.1, and every
git-ref component was checked against its upstream head and is unchanged.
2026-08-16 23:30:01 +02:00
Joakim Persson a55f6369b3 AGENTS.md: three damage-prevention rules for a shared central palace
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 28s
The managed block already tells agents to load the mempalace skill, but said
nothing about the palace being shared with other machines. Three failure modes
observed tonight while onboarding tor-ms22's feeders, all of which do damage
rather than merely confuse:

- `mempalace sync` prunes drawers whose source files look gitignored, deleted
  or moved. On a central palace that describes most of the content, including
  every other machine's. Compounded by RFC-001 7.2: feeders now stage inside
  the palace root, so a scoped sync can delete the drawers it just filed.
- A client-side timeout is not a failure. The palace is single-writer and one
  large mine blocks every client for minutes, so
  `[mempalace ext] feed (tick) failed: mine timed out after 30000ms` usually
  means the mine COMPLETED. Verified: drawers from a timed-out tick were
  present 35 s after the client gave up. A blind retry files a duplicate.
- The `mempalace` CLI has zero references to MEMPALACE_REMOTE_URL, so it always
  opens a palace on local disk and can silently disagree with the MCP tools.

Orientation depth stays in the skill; only damage-prevention belongs here,
because this file is always read and the skill's later sections often are not.
2026-08-16 22:32:09 +02:00
Joakim Persson ffd54750b9 docs: CHANGELOG for v1.8.2 — the silent transcript-feed failure and its guard
Publish Docker Image / resolve-versions (push) Successful in 25s
Lint / actionlint (push) Successful in 30s
Publish Docker Image / base-decide (push) Successful in 8s
Lint / hadolint (push) Successful in 46s
Publish Docker Image / build-base (push) Successful in 41m26s
Publish Docker Image / smoke (push) Successful in 4m45s
Publish Docker Image / smoke-studio (push) Successful in 5m11s
Publish Docker Image / build-variant-studio (push) Successful in 17m44s
Publish Docker Image / build-variant (push) Successful in 23m41s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 10s
2026-08-16 00:53:46 +02:00
Joakim Persson d8b745c164 smoke+docs: pin the feeder's failed-remote-mine detection; REMOTE_PATH is server-visible
Lint / actionlint (push) Successful in 30s
Lint / hadolint (push) Successful in 40s
First boot of the 2026-08-15 image on EMB-7KJ4VR4G shipped 7 transcripts to the
palace host and filed none. Two causes, neither visible in the log:

- MEMPALACE_PI_REMOTE_PATH was unset, so the feeder used its /data/feed default,
  which assumes a CONTAINERIZED palace server. That fleet's primary runs
  natively (systemd user unit + uv tool), so it only sees host paths and the
  mine died with "source directory not found". rsync had already succeeded.
- The feeder decided success with `'"error"' in body`, but MCP escapes the
  tool's JSON inside result.content[].text, so the check was blind and the
  catch-up log said "Done. Wing updated." Fixed in mempalace-toolkit 6e1f4f3,
  which ships `--self-test` with fixtures pinning that exact response body.

- smoke-test.sh: run `mempalace-pi-session --self-test` against the BAKED
  toolkit, so a stale or reverted MEMPALACE_TOOLKIT_REF cannot reintroduce a
  feeder that mines nothing while reporting success.
- .env.example: spell out that MEMPALACE_PI_REMOTE_PATH is the path the SERVER
  PROCESS can open — container path for a dockerized server, and identical to
  the ssh-target path for a native one — and that a mismatch fails quietly.
2026-08-16 00:30:51 +02:00
Joakim Persson ae13c2264e ci: survive a revoked GITEA_BUILD_TOKEN on public commit reads
Lint / actionlint (push) Successful in 14s
Lint / hadolint (push) Successful in 13s
Follow-up to a2f0a4a, which documented the hazard; this removes it.

resolve-versions read three PUBLIC Gitea repos with `curl -sf -H "$AUTH_HEADER"`.
Gitea rejects an invalid token rather than ignoring it, so the token turned a
read that works anonymously into a hard failure:

  no Authorization header    200
  empty token (secret unset) 200   <- absent secret was always safe
  garbage/revoked token      401   <- stale secret broke the release

A revoked GITEA_BUILD_TOKEN therefore failed resolve-versions via require_sha,
presenting as connectivity or an API fault, on data any anonymous client could
fetch. Hit exactly that failure mode today with an expired PAT.

New gitea_sha() helper tries authed, and on 401/403 retries anonymously with a
loud stderr warning naming the token as the cause. Deliberate choices:

- non-200 after the retry emits nothing and returns 0, so require_sha still
  raises the explicit abort — the helper never invents a fallback ref, which is
  the property the surrounding code exists to guarantee
- warnings go to stderr, NOT as ::warning:: annotations: the function's stdout
  IS the SHA, so an annotation there would be captured into the ref
- the header is still sent first, so a private repo keeps working

Verified by extracting the function from the YAML step body (so the test ran the
committed text, not a copy) and calling it against live Gitea:

  valid token        -> 0e1369e6b496 (pi-toolkit)
  REVOKED token      -> warns, retries anon, f60cf9c73205 (mempalace-toolkit)
  unset secret       -> 98eb07bce60a (pi-extensions)
  nonexistent repo   -> empty + HTTP 404 warning, so require_sha aborts

Behaviour unchanged on the happy path: all three SHAs are byte-identical to the
ones the v1.8.1 release run resolved with the old curl code. `bash -n` clean on
the extracted step body; YAML re-parsed.
2026-08-15 14:27:18 +02:00
Joakim Persson a2f0a4a441 ci: correct the false "Gitea requires auth for public reads" comment
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 15s
resolve-versions claimed "Gitea API requires auth even for public-repo commit
listing" above the pi-toolkit / pi-extensions curls. Measurably false for the
repos it guards. Verified 2026-08-15, unauthenticated vs authenticated GET of
/api/v1/repos/joakimp/<repo>/commits?limit=1&sha=main:

  pi-toolkit         private=false  unauth=200 auth=200  sha 0e1369e6b496 identical
  pi-extensions      private=false  unauth=200 auth=200  sha 98eb07bce60a identical
  mempalace-toolkit  private=false  unauth=200 auth=200  sha f60cf9c73205 identical

Only /api/v1/repos/*/actions/* refuses anonymous reads with 401 — almost
certainly what the claim was over-generalised from. (Same over-generalisation I
nearly committed to opencode-devbox's AGENTS.md today; 69fc80a there narrowed it
to the actions endpoints for the same reason.)

Checked the three repos actually queried rather than reusing the pi-devbox
result — if any had been private the comment would have been TRUE, and the
correction wrong.

Behaviour deliberately unchanged: the header still gets passed. It survives a
repo being flipped private, and an unset secret degrades cleanly because Gitea
ignores an empty `token ` value and serves anonymously:

  no header                 200
  empty token (secret unset) 200
  garbage token             401

That last row is the fragility now documented: a REVOKED or malformed token
returns 401 where anonymous returns 200, so a stale GITEA_BUILD_TOKEN converts a
healthy public read into a require_sha failure that presents as an API or
network fault. Encountered exactly that today with an expired PAT on the actions
endpoints, so the note tells the next reader to suspect the token first.

Comment-only: no non-comment line changed, YAML re-parsed.
2026-08-15 14:21:25 +02:00
Joakim Persson 53b41cd76b smoke: assert the pi stage $HOME-relative, and add a smoke_only dispatch
Lint / actionlint (push) Successful in 16s
Lint / hadolint (push) Successful in 13s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke-studio (push) Successful in 5m20s
Publish Docker Image / smoke (push) Successful in 14m40s
Publish Docker Image / build-variant-studio (push) Successful in 21m51s
Publish Docker Image / build-variant (push) Successful in 15m59s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 17s
v1.8.0 never shipped: smoke (67/68) and smoke-studio (70/71) each failed the
same single assertion, so build-variant and everything downstream skipped and
latest stayed on v1.7.0.

The assertion was wrong, not the product. It grepped for a literal
stage=/home/developer/.mempalace/pi-stage/, but run() invokes

  docker run --rm --entrypoint="" "$IMAGE" sh -c "$cmd"

and neither Dockerfile sets USER or ENV HOME — the published base image config
has no HOME at all; it is normally set by entrypoint-user.sh, which
--entrypoint="" skips on purpose. So the assertion executed as root with
HOME=/root, mempalace-pi-session correctly resolved
stage=/root/.mempalace/pi-stage/... (the stage is $HOME-relative by design), and
the literal grep could never match under any circumstances.

The tell was one line below in the log: the sibling assertion "pi stage follows
MEMPALACE_PALACE_PATH" PASSED, because it sets the variable explicitly and never
consults HOME. Default fails + explicit passes = wrong HOME, not broken staging.

Now asserts the invariant actually intended — the stage sits beside the resolved
palace, sharing its lifetime — which is user-independent:

  case "$stage" in "stage=$HOME/.mempalace/pi-stage/"*) exit 0 ;; *) exit 1 ;; esac

$HOME is expanded by the container's own shell, so it holds as root, as
developer, or under any future user. Verified all four cases against the real
bin/mempalace-pi-session by extracting the committed assertion bodies and
running them under sh -c: virgin HOME -> exit 0; HOME=/home/developer -> exit 0;
MEMPALACE_PI_STAGE pinned to a .cache path -> exit 1 (the regression this
assertion exists to catch still fails it); developer-identity companion -> 0.

Added that companion assertion, "pi stage is palace-adjacent for the developer
user", which covers the deployment-specific path properly by SUPPLYING
HOME=/home/developer rather than assuming it.

Why this took a release to surface: docker-publish.yml triggers on push tags v*
only. The assertion was added on a push to main (7c00dd6), where only lint.yml
runs, so v1.8.0 was its first execution ever. Any smoke assertion written
outside a release was unvalidated until a release consumed it.

New workflow_dispatch input smoke_only probes/builds the base, runs both smoke
jobs against HEAD, and stops before publishing. Implemented as
`if: inputs.smoke_only != 'true'` on build-variant and build-variant-studio,
deliberately WITHOUT always() so the implicit needs-succeeded gate survives and
a red smoke still blocks a release. promote-base-latest and update-description
already require build-variant success, so they skip on their own. On a tag push
inputs is unset and null != 'true' is true, so releases are unaffected.

Finally, run() no longer discards output. A red ❌ carried zero diagnostic
weight: explaining this one-line failure needed a CI-log dig plus a registry
image-config inspection, when the container had already printed the answer.
Failures now show the last lines of output (guarded with `if`, not a trailing
`&&`, which would abort under set -e), and the stage assertions echo the
resolved stage and the HOME they saw to stderr — invisible while they pass.
2026-08-15 12:29:32 +02:00
Joakim Persson 29b62093f0 v1.8.0: bump pi 0.84.1 → 0.84.2 and pi-atelier v0.8.0 → v0.8.1, audited
Publish Docker Image / resolve-versions (push) Successful in 12s
Lint / actionlint (push) Successful in 23s
Publish Docker Image / base-decide (push) Successful in 8s
Lint / hadolint (push) Successful in 1m20s
Publish Docker Image / build-base (push) Successful in 41m34s
Publish Docker Image / smoke (push) Failing after 4m41s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Publish Docker Image / smoke-studio (push) Failing after 8m22s
Publish Docker Image / build-variant-studio (push) Has been skipped
pi 0.84.2 closes the Amazon Bedrock tool-argument poison pill that v1.6.4
recorded as "Not fixed upstream". pi-ai 0.84.2 adds a recursive
sanitizeBedrockDocument() and applies it at exactly the site that entry named
(dist/api/bedrock-converse-stream.js, line 692 -> 704; upstream PR #7882):

  - toolUse: { ..., input: c.arguments },
  + toolUse: { ..., input: sanitizeBedrockDocument(c.arguments) },

It strips object members whose key is the empty string, recursing through arrays
and nested objects. It runs at request-build time, so it covers the live turn and
a resume alike: a session already bricked by an empty-key tool argument now
replays instead of dying on a Bedrock ValidationException. pi-session-repair is
therefore no longer the recovery path on this image -- it stays useful for older
images and for inspection, since the fix sanitises what is sent, not what was
recorded.

Bumping PI_VERSION is the ONLY way to get that fix: pi publishes an
npm-shrinkwrap.json, so pi 0.84.1 pins pi-ai to exactly 0.84.1 even though its
package.json range (^0.84.1) would admit 0.84.2. Transitive upstream fixes never
leak into this image.

pi-atelier v0.8.1 is the matching companion -- both sides changed fullscreen
input handling within three days. Its only code change (src/split-pane.ts) stops
atelier writing its own 1002h/1006h pair around a sidebar resize under pi's
fullscreen renderer, which had been tearing down the mouse reporting pi itself
enabled and leaving the wheel dead.

Audited against the surfaces the pin policies name:
  - session .jsonl: identical migrateV1ToV2/migrateV2ToV3 ladder
  - node engine floor: unchanged >=22.19.0 (image ships 22.23.2)
  - atelier's three private couplings all intact in pi-tui 0.84.2 --
    class TuiAltScreen extends TuiBase (detected by constructor NAME, so a
    rename would fail silently), inputListeners still `new Set()` at the same
    line 103, own render(width) descriptor still present
  - pi's mouse sequences byte-identical between 0.84.1 and 0.84.2
  - atelier metadata unchanged: engines >=22.19.0, peerDeps >=0.80.7, zero
    runtime deps, so the "no npm install step" note holds

Not proven by execution: CI smoke does not drive the TUI, and 0.84.2 adds a
focused fullscreen search overlay that also participates in input handling. The
pairing is reasoned from the diffs. Worth an alt+a plus a sidebar resize and
wheel scroll in fullscreen on first use.
2026-08-15 00:22:08 +02:00
Joakim Persson cbd7cf5c67 entrypoint: announce the "remote palace, no inbox" skip instead of vanishing
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 1m8s
When MEMPALACE_REMOTE_URL is set but MEMPALACE_PI_SSH_TARGET is not, the feeder
has nowhere to ship staged transcripts, so skipping is correct. The problem was
that the branch was a bare `:` AND the skip happens before the subshell that
writes ~/.pi/agent/mempalace-catchup.log -- so a container in that state
contributed nothing to the palace and left no artifact at all, not even an empty
log, to explain why. It is indistinguishable from a healthy run that had nothing
to file, which is the worst property a memory system can have: the failure looks
exactly like success.

Surfaced while flipping the first client onto the shared palace, where this is
the single most likely way to end up quietly memory-less -- the palace is the
only thing that survives a container recreate.

The notice goes to both the container start output (docker logs) and the log
path anyone debugging looks at first. It names both variables, says what still
works (MCP tools read/write the shared palace; only this container's own
conversations go nowhere), and points at MEMPALACE_FEED=0 for anyone who meant
it -- "HTTPS first, mining later" is a documented interim state, so the notice
has to be silenceable without being ignorable.

Guarded against becoming a startup failure. mkdir -p in a branch that previously
touched no filesystem is a new risk: an unwritable ~/.pi (root-owned volume, a
classic Docker accident) fails under set -e and would abort the entire
entrypoint. It now degrades to stdout-only. Verified all five paths by executing
the extracted block under `set -euo pipefail`: the trap prints and writes the
log; unwritable ~/.pi still exits 0 and still prints; MEMPALACE_FEED=0 stays
completely silent (no message, no file); and both normal-remote and local mode
still background the feeder with no notice.

Two smoke assertions guard against a regression to the silent no-op. They test
the entrypoint as shipped in the image rather than behaviour, because this branch
only runs at container start and a `docker run` one-shot cannot reach it.

entrypoint-user.sh is COPY'd in Dockerfile.base, so this rides the base rebuild
the Unreleased feeder work already needs.
2026-08-13 16:30:55 +02:00
Joakim Persson 7c00dd6001 mempalace: drop the stage ENV pin, fix the shared-server compose
Lint / actionlint (push) Successful in 21s
Lint / hadolint (push) Successful in 29s
The feeder now defaults to <palace-root>/pi-stage upstream, so pinning
MEMPALACE_PI_STAGE into ~/.pi here is unnecessary -- and was actively wrong. It
created a second convention that could still diverge from the palace: keep the
devbox-palace volume, drop devbox-pi-config, and a scoped `mempalace sync`
prunes every conversation drawer, because dedup keys on the staged path. Both
the ENV and the entrypoint export are gone; a comment explains why adding one
back re-introduces the split it was meant to fix.

docker-compose.mempalace.yml was broken on mempalace 3.6.0 in both directions:
  - `--host 0.0.0.0` with no token in the environment makes the server refuse
    to start, crash-looping under `restart: unless-stopped`.
  - Supply a token and the healthcheck's unauthenticated `tools/list` POST 401s,
    marking a perfectly healthy server unhealthy forever.
Now the token is required via ${MEMPALACE_REMOTE_TOKEN:?...} so it fails fast at
`docker compose up` with a readable message, and the healthcheck probes the
deliberately token-free /healthz. The "no authentication of its own" security
note has been stale since 3.6.0 and is replaced with the actual posture
(bearer token + Host pin + Origin allowlist), including why browser-shaped auth
must not be put in front of it.

Dockerfile.base: mempalace-pi-session symlinked onto PATH, with a build-time
`--help` check so a broken feeder fails the image build rather than the first
session.

smoke-test: assert the stage resolves beside the palace (default, and following
$MEMPALACE_PALACE_PATH) instead of asserting the removed ENV pin. The two
behavioural guards -- a synthetic session that must be captured, an abandoned
one that must not be -- are unchanged.

.env.example: recommend `mempalace serve` on the docker0 gateway rather than
`mempalace-mcp --transport http --host 0.0.0.0`, with the two binds to avoid.
2026-08-12 17:04:14 +02:00
Joakim Persson 7649d53f3b .env.example: say why GIT_USER_* must be set
Lint / hadolint (push) Successful in 30s
Lint / actionlint (push) Successful in 59s
Both keys shipped empty with no explanation, so they read as optional. They
are not: ~/.gitconfig is not on a persistent mount, so with these unset every
repo in the container fails "Author identity unknown" on first commit after
every recreate — and an agent asked to commit then infers an identity from
git log and picks the wrong one. Observed today across three repos on one
machine (two wrong-address commits, plus older `pi <pi@devbox>` fossils from
earlier sessions guessing the same way).

Also records that the address is per-machine (corporate vs private), so it
belongs in the per-machine .env rather than a skill or a repo-local override.
2026-08-09 15:42:29 +02:00
joakimp ade58131d6 docs(readme): use mkdir -p instead of install -d for the peer's ~/.ssh
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 23s
`install -d -m 700 ~/.ssh` was correct but wrong for the audience. This snippet
gets pasted onto an arbitrary peer — a NAS, a router, a BSD box — and `install`
is not in POSIX, so it is not guaranteed to be there. `mkdir -p` + `chmod` is
POSIX, present everywhere, and self-evidently idempotent to a reader deciding
whether it is safe to run on a machine that already has keys.

Behaviour was checked rather than assumed, on both coreutils and BSD/macOS
`install`: on an existing ~/.ssh it exits 0 and leaves authorized_keys intact
in content and mode, but it also silently chmods the directory (755 -> 700).
Desirable here, yet invisible in a doc — which is the second reason to prefer
the explicit two-step form, and why the surrounding text now states outright
that re-running is safe on an already-configured peer: mkdir -p is a no-op, the
chmods only tighten, and appending never touches keys already listed.

Verified the whole block end-to-end against a fresh HOME and one with a
pre-existing 755 ~/.ssh and an older key: 700/600 in both cases, older key
preserved, new line appended.
2026-08-08 01:24:51 +02:00
joakimp ffd44ad9cf docs(readme): say where the authorized_keys line actually goes
Lint / actionlint (push) Successful in 13s
Lint / hadolint (push) Successful in 13s
Step 2 of "Giving the container its own key for a peer" showed a correctly
narrowed authorized_keys line but never named the file it belongs in, and never
said which account's — the only mention of authorized_keys was an aside 35
lines further down about revoking one key per machine. A reader following the
steps had a public key, a line to construct, and nowhere to put it.

Now explicit: append to ~/.ssh/authorized_keys of the account named as `User`
in step 3, via a heredoc that shows `>>` rather than `>` (the truncation that
revokes every other key on that account), with `install -d -m 700 ~/.ssh` and
`chmod 600` so the file is created correctly the first time.

Adds the three failure modes that make a good key look broken, none of which
announce themselves on the client: the options prefix must be on the same
physical line as the key (a wrapped paste is the usual culprit), `ssh-copy-id`
cannot add that prefix at all so the line must be appended by hand, and sshd's
StrictModes silently ignores authorized_keys when the home directory, ~/.ssh or
the file is group- or world-writable — reporting it only in the peer's own log.
2026-08-08 01:17:48 +02:00
joakimp 43cd6e22f2 v1.7.0: bundle pi-atelier at a pinned tag; pin pi to an audited 0.84.1
Publish Docker Image / resolve-versions (push) Successful in 9s
Lint / actionlint (push) Successful in 15s
Lint / hadolint (push) Successful in 13s
Publish Docker Image / base-decide (push) Successful in 8s
Publish Docker Image / build-base (push) Successful in 41m22s
Publish Docker Image / smoke-studio (push) Successful in 5m16s
Publish Docker Image / smoke (push) Successful in 7m31s
Publish Docker Image / build-variant-studio (push) Successful in 18m31s
Publish Docker Image / build-variant (push) Successful in 27m18s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / update-description (push) Successful in 12s
Two changes that belong together, because the first is what makes the second
dangerous to get wrong.

pi-atelier (TUI sidebar + status rail) is now vendored to /opt/pi-atelier at
PI_ATELIER_REF=v0.8.0 and registered by entrypoint-user.sh — the pi-fork /
pi-observational-memory / pi-studio pattern, deliberately NOT
`pi install npm:pi-atelier`, which writes into ~/.pi/npm-global on the config
volume where it shadows the image and pins nothing. Unlike its siblings it gets
no `npm install`: atelier declares zero runtime deps (peerDeps only, satisfied by
the baked pi) and has no build step, so pi loads its TypeScript straight from the
checkout via package.json `pi.extensions`.

pi is no longer resolved to npm `latest` at build time. The pin lives in
Dockerfile.variant and CI reads it from there, so a local `docker build` and a CI
release ship the same versions by construction. The pin is a CHECKPOINT, NOT A
FREEZE: bumping stays a one-line change; what stops is *unreviewed* adoption of
whatever shipped that morning, in the same build that then gets tagged and
published. CI fails when a pin is not concrete or not actually published on npm,
and warns — never adopts — when npm latest moves ahead, naming what to re-check.

Why this pairing needed care: pi-atelier 0.6.0/0.7.0 wrap pi's PRIVATE TUI
renderer, and under pi 0.84 that wrapper recurses — pi hangs at startup burning
CPU with no error. Upstream fixed the recursion in 0.7.1 and restored the
non-overlapping split in 0.7.2; 0.8.0 is additive on top. atelier's own
peerDependencies still say >=0.80.7, which does not express that floor, so
nothing in npm metadata could have warned us. The floor is therefore encoded as
an executable rule — pi >= 0.84 => pi-atelier >= 0.7.1 — asserted in both
smoke-test.sh (build time) and recreate-sanity-check.sh (after a real recreate),
verified against a 4x4 version matrix.

Existing volumes needed migration, not just vendoring: a hand-installed
`npm:pi-atelier` entry is counted as already-registered by the entrypoint guard,
so every existing volume would have kept its unpinned npm copy — and a 0.6.x copy
next to pi 0.84 is exactly the startup hang. The entrypoint now drops that one
exact string (settings.json.bak.atelier.<ts> backup, distinct prefix so it cannot
clobber the template merge's backup in the same second) and lets the pinned /opt
copy register. Tested against a real settings.json: only that entry removed,
other packages and all keys intact, idempotent, and unparseable JSON leaves the
file untouched. DEVBOX_ATELIER=0 opts out entirely — in the entrypoint rather
than via `pi uninstall`, because this component's failure mode is "pi will not
start", which cannot be repaired from inside pi.

0.84.1 was audited for this release, not merely adopted: theme/TUI changes are
additive, the session format is unchanged (CURRENT_SESSION_VERSION = 3 in both
0.83.0 and 0.84.1 with an identical migrateV1ToV2/migrateV2ToV3 ladder, so
existing transcripts are neither migrated nor at risk and pi-session-repair stays
valid), and the Node engine floor is unmoved at >=22.19.0. CI resolves the
atelier tag to its PEELED commit SHA — atelier uses annotated tags, so the
unpeeled ref is a tag object, not a commit; pi-studio's lightweight tags never
exposed that distinction.

Also: docs for overriding the read-only ~/.ssh/config from the container —
container-only keys in ~/.ssh-local, hardened authorized_keys, the fact that
`from=` must allow the HOST's addresses because container egress is NAT'd through
it, and the macOS-only-keyword trap (`UseKeychain` is fatal to Linux OpenSSH and
takes out dssh/pi --ssh while the host keeps working). Corrects two claims in
"Naming LAN peers": ssh-lan.conf is not ProxyJump-only, and first-time creation
does need one restart because the Include is emitted only when the file already
exists at start.
2026-08-07 21:34:41 +02:00
pi 62a2a79b1c ci(lint): correct the rationale comment — runner contention was overstated
Lint / hadolint (push) Successful in 15s
Lint / actionlint (push) Successful in 4m29s
The previous commit justified excluding tag pushes partly on runner contention:
that the duplicate lint run stole one of two self-hosted runners from the release
build. Measured, that is false for THIS repo — pi-devbox lint runs take 0.3-0.9
min (ids 529/531/532/533) against a 77.6 min release build (id=530). I imported
the claim from opencode-devbox, where actionlint apt-installs shellcheck inside
the container and takes 6-15 min, so contention there is real.

The change stands on its actual merits: duplicate lint of an identical tree, and
release-run discovery ambiguity (the substantive one — it is what made the naive
"first run matching refs/tags/<tag>" rule pick lint over the publish run).

No functional change; comment only.
2026-08-04 18:08:28 +02:00
pi f20b2a7926 ci(lint): don't re-lint on tag pushes
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Has been cancelled
`on: push:` with no filter also fires on refs/tags/v*, which is duplicate work:
the tagged tree was already linted when that same commit was pushed to main
(v1.6.4 sha e86e5df linted as id=529 on main, then again as id=531 on the tag).

Two costs beyond the wasted run. It consumed one of the two self-hosted runners
while the release pipeline wanted both for its parallel multi-arch variant
builds; and it made release-run discovery ambiguous, since the newest-first runs
listing puts the tag-ref lint run above the publish run.

`branches: ['**']` keeps the documented intent exactly — lint fires early on
every branch push and PR, rather than only at tag time — while excluding tag
refs. docker-publish.yml is untouched and still tag-scoped.
2026-08-04 18:06:34 +02:00
pi 66a19aa394 docs(agents): how to find the release run (tag push fires two workflows)
The release-day checklist said "Watch CI" without saying which run, and the
Gitea API example used limit=5. Both are traps, because a tag push produces
TWO runs here: lint.yml has a bare `push:` trigger so it fires on the tag ref
as well, and docker-publish.yml fires on v*. The runs listing is newest-first
and the lint run sorts ABOVE the publish run, so "first run matching
refs/tags/<tag>" picks lint reliably. Verified against the real API for v1.6.4:

  id=531  #104  lint.yml@refs/tags/v1.6.4           <- picked by the naive rule
  id=530  #103  docker-publish.yml@refs/tags/v1.6.4 <- the actual release build
  id=529  #102  lint.yml@refs/heads/main            <- same sha, already linted

Lint goes green in minutes while the image is still building, so watching it
makes a release look finished before anything is published. limit=5 compounds
it: the publish run is already at position 4 of 5 in the current listing.

Documents: head_sha-filtered discovery with limit=20; the jobs endpoint takes
the internal id, never the run_number (silently returns another run's jobs);
and the correct ci-release-watcher config for this repo — EXPECT_WORKFLOW,
the studio tag pair, base-latest as existence-only, and CRITICAL_JOBS with
build-variant-studio spelled out (job names are matched exactly, and the
skill's default omits it) while excluding promote-base-latest, which
legitimately skips on a base cache hit.

Smoke-gate detail in step 5 is retained.
2026-08-04 18:06:34 +02:00
pi 572430237f Bump mempalace pin 3.5.0 → 3.6.0 (lockstep with opencode-devbox v2.9.0)
Lint / actionlint (push) Successful in 38s
Lint / hadolint (push) Successful in 50s
3.6.0 (2026-07-17, PyPI latest) is additive/reliability only: secure
`mempalace serve` remote mode, optional Milvus backend, atomic KG
supersede(), conversation chronology, mining exclusions, plus recovery and
locking fixes.

Reviewed for MCP tool-schema changes before bumping — that being the exact
regression class this pin exists to catch, after an unpinned install once
swept in the broken 3.3.x/3.4.0 diary_write schema. There are none, and
nothing touches diary_write, so the perl workaround removed in v1.2.2 stays
removed.

Two fixes matter for how this image uses mempalace: read-only mode now covers
checkpoint + delete_by_source in _MUTATING_TOOLS (#1930), and agent
attribution is preserved in mempalace_checkpoint (#2023/#2034) — the latter
because the diary protocol relies on per-agent attribution.

Also adds a CHANGELOG Unreleased block that backfills the per-variant image
description labels (1fd524e), pushed after the v1.6.4 tag without an entry.

Not tagged: more changes are queued for the next release. Pushing to main
triggers only lint.yml (actionlint + hadolint) — the image build/publish
workflow is tag-only. Verified clean against the CI-pinned hadolint 2.14.0.
2026-08-04 16:01:39 +02:00
joakimp 1fd524e7fb Give each variant its own image description label
Lint / actionlint (push) Successful in 14s
Lint / hadolint (push) Successful in 12s
Both published variants inherited Dockerfile.base's
description="pi-devbox — base image (variant-independent)", so v1.6.4 and
v1.6.4-studio both advertised themselves on Docker Hub as the base image —
misleading, and useless for telling the two apart.

A LABEL cannot branch on INSTALL_STUDIO, so the text arrives as a build-arg:
CI passes a variant-specific string (interpolating RELEASE_TAG, PI_VERSION and,
for studio, STUDIO_TAG), and the Dockerfile default keeps a bare local
`docker build -f Dockerfile.variant` honest instead of misleading.

Also sets org.opencontainers.image.title/description alongside the legacy bare
`description` key, so Hub and OCI-aware tooling both see it. ARGs stay in the
last-declared block, so the label layer is still the only thing invalidated.

Verified: hadolint 2.14.0 (the CI-pinned version) clean on both Dockerfiles;
workflow YAML parses; check-workflow-shell.sh passes. Lands on the next release.
2026-07-30 07:59:41 +02:00
joakimp e86e5df327 release: v1.6.4 — fork tool actually loads, pi 0.83.0 audited clean
Publish Docker Image / resolve-versions (push) Successful in 5s
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 15s
Publish Docker Image / base-decide (push) Successful in 7s
Publish Docker Image / build-base (push) Successful in 41m24s
Publish Docker Image / smoke-studio (push) Successful in 5m11s
Publish Docker Image / smoke (push) Successful in 12m23s
Publish Docker Image / build-variant-studio (push) Successful in 22m46s
Publish Docker Image / build-variant (push) Successful in 22m58s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 13s
Headline is the fork-guard fix (the `fork` tool had never loaded since v1.0.0
because the registration guard grepped the whole settings.json and matched the
pi-fork *config block* the template merge itself plants).

pi 0.82.1 -> 0.83.0 is a clean hop with no intermediates. 0.83.0 carries an
upstream Breaking Change (bundled TypeBox 1.1.38 -> 1.3.7, deprecated APIs
removed) that cannot reach us: pi-fork vendors @sinclair/typebox (a different
package name), obsmem's use is type-only, studio and atelier don't use TypeBox.
Verified from source rather than from the audit narrative: extension-facing
dist/core/extensions/*.d.ts declarations diff clean between the two versions,
all six CLI flags pi-fork passes to child processes are present, and
SESSION_VERSION is 3 in both so transcript tooling is unaffected. No PI_VERSION
pin needed.

Recorded for the release notes: the Bedrock poison pill is NOT fixed in pi-ai
0.83.0 (unsanitised input: c.arguments moved 634 -> 644), so pi-session-repair
stays the recovery path.
2026-07-30 01:10:07 +02:00
joakimp fa04d2083d docs(rootfs): re-sync pi-extensions skill snapshot (fork boundary mechanism)
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 25s
Vendored floor snapshot re-synced from pi-extensions 98eb07b, which documents
the mechanism behind fork boundary violations (full parent-transcript
inheritance via index.ts:47) plus the corrected claims about tool restriction
and narrative invention. CI resolves PI_EXTENSIONS_REF from main HEAD, so a
normal build ships the package-owned copy; this keeps the committed floor
identical so scripts/smoke-test.sh's cmp assertion holds either way.
2026-07-30 00:50:49 +02:00
joakimp 209f2c2f67 docs(rootfs): re-sync mempalace snapshot from skillset 63f3bf5
Lint / hadolint (push) Successful in 11s
Lint / actionlint (push) Successful in 17s
Closes the drift found in 4d4abd9's investigation: the Temporal grounding
guidance was authored straight into this vendored fallback (904fe85) and never
returned to skillset, the declared owner. skillset 63f3bf5 now carries it (with
the wording generalized from "a pi-devbox container" to "a devbox container
(pi-devbox or opencode-devbox)", since opencode-devbox vendors the same skill),
so this snapshot is a pure mirror of the owner again.

VENDORED.md: refresh commands now copy each snapshot FROM ITS OWNER. The
pi-extensions lines previously pointed at <skillset>/skills/pi-extensions/,
which has been a downstream duplicate since a7f3044 co-located the canonical
skill in the package repo — following the old instruction would have silently
regressed the snapshot (e.g. undoing pi-extensions e73cb9f). Provenance bumped
to skillset 63f3bf5 / pi-extensions pkg e73cb9f.
2026-07-29 19:44:36 +02:00
joakimp 4d4abd9a9f skill(pi-devbox-environment): resolve a skill symlink before editing it
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 19s
~/.agents/skills/ lives in the ephemeral container layer and is rebuilt by
entrypoint-user.sh on every start from two sources, so an edit made through the
symlink may vanish on the next recreate. Adds to §1 (persistence tiers):

- `readlink -f ~/.agents/skills/<name>` as the first move, with a tier table:
  resolves under /workspace/skillset → edit in place; resolves under
  /usr/local/share/pi-devbox/skills → image layer, edit the canonical repo and
  `sudo cp` to activate for the running session.
- Canonical owner per baked skill (pi-devbox-environment → this repo;
  pi-extensions → the package repo's skill/, plus this repo's floor snapshot;
  mempalace → the private skillset repo), pointing at VENDORED.md as
  authoritative.
- The shadowing gotcha: image-baked links are created first and only when
  absent, and deploy-skills.sh --prune-stale leaves foreign links alone, so for
  a name present in BOTH sources the image copy wins and a skillset edit has no
  effect in the container. Documented with the live example found while writing
  this: the baked mempalace snapshot carries a Temporal grounding section
  (904fe85) that skillset at its snapshot point (8e8db64) lacks.
- Checklist gets a matching line.

Found while adding session findings to the pi-extensions skill (e73cb9f), where
the same resolve-first step was what kept the edit out of the image layer.
2026-07-29 19:37:59 +02:00
joakimp d5c5da3f6c docs(rootfs): refresh vendored pi-extensions skill snapshot
Lint / actionlint (push) Successful in 34s
Lint / hadolint (push) Successful in 45s
Sync the image-baked "floor" copy at
rootfs/usr/local/share/pi-devbox/skills/pi-extensions/SKILL.md with
pi-extensions e73cb9f, which documents package-registration forensics
(packages[] vs whole-file grep), that /reload suffices for a newly installed
package, and a fork-output caveat — the findings from the fork-guard bug fixed
in 8248688.

Dockerfile.variant copies the pinned package's skill/ over this snapshot at
build time, so the vendored copy is only the fallback floor; it was byte-
identical to the package copy before this change, and keeping it in sync
prevents a silent divergence for builds whose PI_EXTENSIONS_REF predates the
skill.
2026-07-29 19:34:00 +02:00
joakimp 8248688d58 fix(entrypoint,tests): register pi-fork — guard matched its own config block
Lint / actionlint (push) Successful in 32s
Lint / hadolint (push) Successful in 1m25s
The `pi install /opt/<pkg>` loop in entrypoint-user.sh guarded on a
whole-file substring grep of ~/.pi/agent/settings.json. settings.example.json
ships a top-level "pi-fork" CONFIG block (fork effort profiles, pi-toolkit
adb6907, 2026-06-17), so `grep -q pi-fork settings.json` matched the config
key itself and `pi install /opt/pi-fork` never ran — on fresh or preserved
volumes. The `fork` tool has therefore been absent since v1.0.0.

The non-destructive template merge runs earlier in the same startup than the
install loop, so the mechanism that delivers new template keys to an old
volume is what plants the string that defeats the guard. pi-observational-
memory and pi-studio escaped only by luck: the template key is
"observational-memory" (no pi- prefix) and there is no studio block.

Guard now inspects the `packages` array via jq, with a grep fallback on the
stored `.../opt/<name>"` path form, which a config key can never produce.
Existing volumes self-heal on the next container start.

Both test suites asserted the bug as green — smoke-test.sh:244 and
recreate-sanity-check.sh:204 used the same whole-file grep, so "pi-fork
registered (fork tool)" passed on every build and recreate while the tool was
missing. Both now assert against packages[] with the entrypoint's predicate,
labels say packages[], and the smoke readiness wait loop uses the array check
plus `docker exec -u developer` + $HOME instead of a hard-coded
/home/developer path.

Evidence: zero `fork` tool calls across all 19 sessions on this volume; the
v1.6.3 session that tuned pi-fork.deep to opus-5 was configuring a tool that
never loaded.
2026-07-29 19:21:49 +02:00
joakimp e274510fd1 docs(changelog): unreleased — settings template defaults to Opus 5
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 14s
Records the pi-toolkit 926f738 template change (defaultModel + pi-fork deep
tier -> eu.anthropic.claude-opus-5). No tag, no build: CI resolves the
pi-toolkit SHA from main at build time, so whichever release builds next
picks it up.
2026-07-26 00:30:54 +02:00
joakimp 9ef7a92dce docs(changelog): v1.6.3 — pi 0.81.1→0.82.1
Publish Docker Image / resolve-versions (push) Successful in 7s
Publish Docker Image / base-decide (push) Successful in 15s
Publish Docker Image / build-base (push) Has been skipped
Lint / hadolint (push) Successful in 45s
Lint / actionlint (push) Successful in 15s
Publish Docker Image / smoke (push) Successful in 4m23s
Publish Docker Image / smoke-studio (push) Successful in 7m52s
Publish Docker Image / build-variant (push) Successful in 16m27s
Publish Docker Image / promote-base-latest (push) Successful in 5s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / build-variant-studio (push) Successful in 21m48s
Pure pi version bump (variant-only rebuild). CI resolves pi@latest=0.82.1 at build. 0.82.0/0.82.1 audited: additive, no breaking changes to the extension execution API (pi-observational-memory) or pi-agent-core types (pi-fork).
2026-07-25 22:35:51 +02:00
joakimp 6dfbded9c8 release: v1.6.2 — lift smoke size threshold (3500→3800)
Publish Docker Image / resolve-versions (push) Successful in 12s
Lint / hadolint (push) Successful in 7s
Lint / actionlint (push) Successful in 20s
Publish Docker Image / base-decide (push) Successful in 11s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Successful in 4m42s
Publish Docker Image / smoke-studio (push) Successful in 8m24s
Publish Docker Image / build-variant-studio (push) Successful in 18m8s
Publish Docker Image / build-variant (push) Successful in 25m53s
Publish Docker Image / promote-base-latest (push) Successful in 6s
Publish Docker Image / update-description (push) Successful in 12s
Completes v1.6.1's studio publish. Run 512 shipped v1.6.1 non-studio
(3411 MB) cleanly, but smoke-studio failed the size gate at 3574 MB vs
the 3500 MB threshold. Threshold was set in v1.0.0 pre-agent-browser
(baseline was 3.20 GB local arm64 + 300 MB margin); v1.6.0 baked in
agent-browser + Chromium (+~291 MB net) but the threshold was never
lifted. v1.6.0 never got to smoke because of the network fault, so
nothing surfaced this until v1.6.1's smoke-studio.

Bump SIZE_THRESHOLD_MB to 3800 (~225 MB margin above observed studio
number, tight enough to still catch a genuine +GB regression). Refresh
the comment above the constant with the current baseline + run 512
actuals so future readers know where the number came from.

CI-only change; image bytes identical to v1.6.1 except for the manifest's
release_tag/source_revision. Not base-affecting.
2026-07-23 08:43:57 +02:00
joakimp 45b6239777 test(smoke): don't hard-code a 'v' prefix on release_tag
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 21s
Publish Docker Image / resolve-versions (push) Successful in 48s
Publish Docker Image / base-decide (push) Successful in 28s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Successful in 4m44s
Publish Docker Image / smoke-studio (push) Failing after 5m4s
Publish Docker Image / build-variant-studio (push) Has been skipped
Publish Docker Image / build-variant (push) Successful in 29m5s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 15s
The smoke workflow deliberately passes RELEASE_TAG=smoke to the variant
build so smoke images don't collide with real vX.Y.Z tags. The variant
bakes that into /etc/pi-devbox/build-manifest.json, and pi-devbox-version
prints 'pi-devbox smoke' — correct behaviour. But the smoke assertion
required substring 'pi-devbox v', which only holds for real releases.

Assertion never fired before because pi-devbox-version was added after
v1.5.0 (fb49828, 2026-07-15) and every CI attempt since was blocked
before smoke ran (v1.6.0 network flake, v1.6.1 first successful base
build hit this). CI run 509 (v1.6.1) surfaced it.

Fix: require the substring 'pi-devbox ' (space, no v). The two
neighbouring assertions on --json and --quiet already cover the value
of release_tag; this one just verifies the human line renders.
Everything else in 55/56 checks passed on run 509 including
pi 0.81.1 reported and base-229f04e5d021 pushed OK.
2026-07-23 07:58:22 +02:00
joakimp d00eef2acb docs(changelog): v1.6.1 — pi 0.80.6→0.81.1 (skips 0.81.0)
Publish Docker Image / resolve-versions (push) Successful in 7s
Lint / actionlint (push) Successful in 55s
Lint / hadolint (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 55s
Publish Docker Image / build-base (push) Successful in 58m42s
Publish Docker Image / smoke (push) Failing after 4m50s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Publish Docker Image / smoke-studio (push) Failing after 13m22s
Publish Docker Image / build-variant-studio (push) Has been skipped
v1.6.0 was tagged 2026-07-17 but never reached Docker Hub (variant
publish blocked by a site-network SYN-drop fault, since fixed). Cut
v1.6.1 to land v1.6.0's content (agent-browser + pi-devbox-version)
alongside a first pi bump since v1.5.0.

0.81.0 is skipped deliberately: it removed the default stream fallback
for extensions using the pre-0.81 pi-agent-core API, which
pi-observational-memory relies on via agentLoop + stream.result().
0.81.1 restored the fallback (earendil-works/pi#6915), so 0.81.1 — but
not 0.81.0 — is a safe drop-in. pi-fork only imports types; unaffected.

Audit of 0.80.7–0.81.1 vs the two baked extensions and the base image:
no breaking changes affect pi-devbox. Node engine bumped to >=22.19.0
in 0.81.0 (nodesource 22.x currently 22.23.1, so no engine bump needed).

Base-affecting via the npm install line, so base-<hash> rebuilds.
2026-07-22 22:53:55 +02:00
joakimp fb35c549b5 fix(base): locate agent-browser's chrome arch-agnostically
Lint / actionlint (push) Failing after 12s
Lint / hadolint (push) Successful in 12s
Publish Docker Image / resolve-versions (push) Failing after 26s
Publish Docker Image / base-decide (push) Has been skipped
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Has been skipped
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Publish Docker Image / smoke-studio (push) Failing after 31m44s
Publish Docker Image / build-variant-studio (push) Has been skipped
v1.6.0's first build failed on amd64: playwright@latest fetches Chrome for
Testing, which extracts to chromium-<rev>/chrome-linux64/ on amd64 (vs
chrome-linux/ on arm64). The hardcoded chrome-linux/ glob matched only arm64,
so amd64 built no /usr/local/bin/agent-chrome symlink and 'test -x' failed
(agent-browser --version had already printed 0.32.1 — the tell).

Replace the glob with an arch-agnostic 'find -name chrome -path */chromium-*/*'
(skips the chrome-headless-shell binary and the chromium_headless_shell dir),
guarded by [ -n ] + the existing test -x. Verified locally: resolves the CfT
chrome and agent-browser drives it headless. hadolint clean.
2026-07-17 17:55:21 +02:00
joakimp 649fc44c5b chore(release): v1.6.0
Lint / hadolint (push) Failing after 22s
Lint / actionlint (push) Successful in 26s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Failing after 25m26s
Publish Docker Image / smoke (push) Has been skipped
Publish Docker Image / smoke-studio (push) Has been skipped
Publish Docker Image / build-variant-studio (push) Has been skipped
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Roll the unreleased changes into v1.6.0 (minor — significant base addition:
agent-browser + headless Chromium baked into every variant for browser
automation/front-end verification). Also includes the pi-devbox-version command
and the bundled pi-toolkit sonnet-5 template bump. No pi bump — CI resolves
latest pi from npm at build time (verified: 0.80.6→0.80.10 has no breaking
changes affecting pi-devbox).
2026-07-17 17:19:56 +02:00
joakimp 89a8dc7fab perf(base): trim agent-browser footprint ~960MB→~625MB before release
Lint / hadolint (push) Failing after 14s
Lint / actionlint (push) Successful in 23s
Size concern ahead of the v1.6.0 base rebuild. agent-browser drives the full
chrome (verified headless: open/title/eval with the headless_shell removed), so
Playwright's chromium_headless_shell-* build is dead weight — drop it (~334MB),
and clean apt/npm caches in-layer. Net browser footprint ~625MB/arch.

CI disk is otherwise fine: build-base's 'Reclaim runner disk' step frees
~20-30GB (hostedtoolcache/dotnet/android/jvm) + docker prune before buildx.
hadolint clean.
2026-07-17 17:19:20 +02:00
joakimp 6625d66f3a docs(agents): point the global AGENTS.md managed block at agent-browser
Lint / hadolint (push) Failing after 7s
Lint / actionlint (push) Successful in 27s
Discoverability follow-up to baking agent-browser into the base: agents won't
reach for a browser they don't know they have (the exact gap that made this
capability easy to miss). Add a short pointer section to the pi-devbox managed
block — capability, the preset AGENT_BROWSER_EXECUTABLE_PATH, and
`agent-browser skills get core --full` for the command set. Pointer only; depth
stays in the skill. rootfs change → folds into the same base-<hash> rebuild.
CHANGELOG note updated.
2026-07-17 17:15:09 +02:00
joakimp 8caafc3f49 base: bake agent-browser + Playwright Chromium (headless browsing, all variants)
Lint / actionlint (push) Successful in 28s
Lint / hadolint (push) Successful in 11s
Gives the agent a real browser it can drive so front-end work involving live
DOM/WebGL can be VERIFIED, not guessed. The agent-browser skill (from the
skillset repo) was a no-op without the binary; it now works out of the box.

- agent-browser CLI (standalone Rust, ships no browser) via npm, prefixed
  NPM_CONFIG_PREFIX=/usr so it survives the ~/.pi/npm-global volume.
- Chromium fetched with 'playwright install --with-deps chromium' into
  PLAYWRIGHT_BROWSERS_PATH=/usr/local/share/ms-playwright — a system path that
  the /home/developer volume can't shadow (unlike agent-browser's own
  ~/.agent-browser/browsers default, which WOULD vanish on recreate).
- Stable /usr/local/bin/agent-chrome symlink, exported as
  AGENT_BROWSER_EXECUTABLE_PATH, insulates the ENV from Playwright's
  per-version chromium-<rev> dir name.

Verified end-to-end (this session): agent-browser drives the baked Chromium
headless (open + title + screenshot + eval into a WebGL SPA); doctor launch
test passes in ~0.5s. Debian trixie --with-deps resolution verified (exit 0;
t64 lib renames handled). hadolint clean; check-base-hash OK (only *_VERSION
args added). Cost ~960 MB (Chromium + headless shell). Base-affecting →
rebuilds base-<hash> on next release.
2026-07-17 16:12:47 +02:00
joakimp 71b12a9ed4 docs(skill): note macOS NFD-filename gotcha for dscp/scp
Lint / hadolint (push) Failing after 7s
Lint / actionlint (push) Successful in 31s
Accented filenames on a macOS host are stored decomposed (NFD), so a precomposed
(NFC) remote path in scp/dscp silently fails with 'No such file or directory'.
Document the wildcard / list-first workaround in the pi-devbox-environment skill,
next to the dssh/dscp alias table. (Hit while copying a screenshot named
'Skärmavbild ….png' from the host.)
2026-07-17 13:06:22 +02:00
joakimp fb49828826 base: add pi-devbox-version command + startup banner
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 1m11s
Wraps /etc/pi-devbox/build-manifest.json (already written at docker-build
time in Dockerfile.variant) into a human-readable summary instead of
requiring users to know the manifest path and pipe it through jq
themselves.

- rootfs/usr/local/bin/pi-devbox-version: human (default) / --json /
  --quiet output modes. Also flags live drift — compares the baked
  pi_version against the actually-running `pi --version` and warns
  on mismatch rather than trusting the manifest blindly (same
  ground-truth philosophy as the manifest generation itself). Exits 1
  with a short stderr notice on images built before the manifest
  existed, instead of failing silently.
- entrypoint-user.sh: calls it as the very first line. CMD is
  `bash -l` with tty:true in compose, so this banner lands directly
  above the first prompt on container start — no separate motd/bashrc
  hook needed (deliberately not wired into .bash_aliases, which would
  reprint on every docker exec -it).
- Dockerfile.base: COPY + chmod, same pattern as dot-watch/studio-expose.
- scripts/smoke-test.sh: 4 new checks (binary present+executable, human
  output has release tag, --json round-trips the manifest, --quiet is
  a single line).
- README.md / AGENTS.md / CHANGELOG.md updated.
2026-07-15 14:50:47 +02:00
pi 02be95ac1f docs(changelog): note bundled pi-toolkit sonnet-5 template bump (Unreleased)
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 21s
2026-07-13 23:13:24 +02:00
pi d68674d11e chore(release): v1.5.0
Publish Docker Image / resolve-versions (push) Successful in 12s
Publish Docker Image / base-decide (push) Successful in 9s
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 13s
Publish Docker Image / build-base (push) Successful in 43m20s
Publish Docker Image / smoke (push) Successful in 4m20s
Publish Docker Image / smoke-studio (push) Successful in 7m2s
Publish Docker Image / build-variant (push) Successful in 16m18s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 14s
Publish Docker Image / build-variant-studio (push) Successful in 17m47s
Roll the 7 unreleased commits since v1.4.0 into v1.5.0 (minor — significant
base additions: readable Neovim true-colour, broad terminal terminfo support).
Also includes: typst PDF template font fix, .claude gitignore seed, pi-studio
semver-tag pin + version label, and repo hygiene (LICENSE, THIRD_PARTY.md,
.dockerignore, hadolint lint, IDEAS backlog). No pi bump — CI resolves latest
pi from npm at build time as usual.
2026-07-13 18:56:51 +02:00
pi 8c27894cf2 base: support modern terminals (ncurses-term + xterm-ghostty alias)
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 15s
The base only shipped ncurses-base (xterm-256color, tmux), so SSHing in from
WezTerm/Alacritty/foot/Ghostty degraded to a dumb TERM fallback. Following the
maintainer's ansible common role:

- Dockerfile.base: install ncurses-term (terminfo for wezterm, alacritty,
  foot, st, base ghostty entry, and many more).
- rootfs/.../terminfo-src/ghostty.terminfo: thin xterm-ghostty alias
  (use=ghostty) — Ghostty connects as TERM=xterm-ghostty, which no distro
  packages. Compiled into the system db with 'tic -x'; build asserts it
  resolved via infocmp.
- kitty (xterm-kitty) already covered by kitty-terminfo; iTerm2 uses
  xterm-256color (ncurses-base).
- smoke-test: assert ncurses-term emulators + xterm-ghostty alias resolve.

Validated end-to-end in a throwaway container: ncurses-term brings the
entries, the alias compiles and equals the ghostty capability set. hadolint
clean, bash -n OK. Base-affecting, rebuilds base-<hash>. No tag.
2026-07-13 18:45:38 +02:00
pi 291ae5345e repo: add LICENSE, THIRD_PARTY.md, .dockerignore, hadolint lint, IDEAS backlog
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 16s
Repo/CI hygiene batch (none base-affecting; image contents unchanged):

- LICENSE: actual MIT file (repo previously declared MIT only in prose).
- THIRD_PARTY.md: notes bundled software + licenses (pi/pi-fork/pi-obsmem/
  pi-studio MIT, gosu Apache-2.0, Debian packages under their own terms).
- .dockerignore: trims build context to what the Dockerfiles COPY (rootfs/ +
  entrypoint*.sh); keeps .git/docs/scripts/compose out. Verified it excludes
  none of the required COPY sources.
- lint.yml: new hadolint job (pinned v2.14.0) lints both Dockerfiles;
  .hadolint.yaml grandfathers deliberate choices (DL3008/DL3016/DL4006/DL3003/
  SC2086, mirroring the shellcheck excludes), fails on anything new at warning+.
  Verified hadolint exit 0 and the repo shell-guard passes with the new job.
- IDEAS.md: parks deferred follow-ups (SHA-pin actions, trivy, buildx SBOM/
  provenance, Makefile, renovate).
- README/DOCKER_HUB License sections now link LICENSE + THIRD_PARTY.md.

No tag.
2026-07-13 18:20:44 +02:00
pi 38d8832d34 ci: record pi-studio version as image label
Lint workflows / actionlint (push) Successful in 1m16s
Adds se.jordbo.pi-devbox.pi-studio-version (e.g. v0.9.36) alongside the
existing SHA label, so 'docker inspect' shows the human-readable version.

Plumbing: resolve-versions exposes studio_tag output -> both studio build
steps pass --build-arg PI_STUDIO_VERSION -> Dockerfile.variant declares
ARG PI_STUDIO_VERSION=none and emits the LABEL. Studio-variant only.

Validated: yq parse OK, resolve run-script bash -n OK. No tag.
2026-07-13 17:59:59 +02:00
pi 32586f19e7 ci: pin pi-studio to newest semver tag, not main HEAD
Lint workflows / actionlint (push) Successful in 23s
Upstream omaclaren/pi-studio stopped publishing GitHub Releases at v0.5.55
but keeps tagging every version (v0.9.36 now) and pushing to main. Tracking
main HEAD risked baking half-finished commits that land after a tag.

resolve-versions now lists all tags via a single git ls-remote (the REST
tags API paginates at 100 and the repo has >140 tags, so a single page can
miss the newest), picks the highest X.Y.Z with sort -V (pre-releases
excluded), and peels it to a commit SHA for PI_STUDIO_REF. SHA (not moving
tag) keeps cache-busting + reproducibility; require_sha still enforced.

Studio-variant only, not base-affecting. No change to the resolved commit
today (v0.9.36 == current main HEAD). Validated: yq parse OK, bash -n OK,
live git ls-remote -> v0.9.36 -> 2ef38ef. No tag.
2026-07-13 17:54:56 +02:00
pi aaf1be0bcb base: readable Neovim colours out of the box (termguicolors + kitty-terminfo)
Lint workflows / actionlint (push) Successful in 22s
Vanilla Neovim fell back to a 256-colour palette over ssh/kitty and rendered
strings/comments in a muddy, low-contrast colour. Fix, all base-affecting:

- rootfs/etc/xdg/nvim/sysinit.vim: system vimrc enabling termguicolors. Loads
  for every user before any personal ~/.config/nvim; overridable per-user.
- Dockerfile.base: install kitty-terminfo (TERM=xterm-kitty understood; lets
  Neovim auto-detect true colour) and set ENV COLORTERM=truecolor.
- smoke-test.sh: assert xterm-kitty terminfo present and nvim tgc default on.
- CHANGELOG (Unreleased/Added) + README editor note.

No tag — lands on the next base-<hash> rebuild.
2026-07-12 23:19:19 +02:00
pi 92212fa447 base: seed global gitignore with .claude/settings.local.json
Lint workflows / actionlint (push) Successful in 23s
Claude Code's per-machine local settings can carry credentials and should never
be committed. Add the pattern to the seeded ~/.gitignore_global so fresh
containers match a host global that already ignores it. Existing containers are
unaffected (seed copied only when absent). Base-affecting; rebuilds base-<hash>.
2026-07-11 22:29:28 +02:00
pi 3d46c6615e fix(base): default pandoc typst template font so PDF export works without -V mainfont
Lint workflows / actionlint (push) Successful in 21s
pandoc's bundled typst template defaults the document font to an empty
tuple (font: ()), so a naked `pandoc --pdf-engine=typst` fails with
"font fallback list must not be empty". Patch the template default to
Libertinus Serif (typst's own bundled default) at build time so PDF
export works out of the box. Document usage in README and note the fix
in CHANGELOG (Unreleased). Base-affecting.
2026-07-11 20:55:50 +02:00
pi fa6e9dc9d6 release: v1.4.0 — typst PDF engine + host SSH startup check
Publish Docker Image / resolve-versions (push) Successful in 6s
Publish Docker Image / base-decide (push) Successful in 9s
Lint workflows / actionlint (push) Successful in 1m9s
Publish Docker Image / build-base (push) Successful in 34m32s
Publish Docker Image / smoke-studio (push) Successful in 4m16s
Publish Docker Image / smoke (push) Successful in 10m52s
Publish Docker Image / build-variant-studio (push) Successful in 18m1s
Publish Docker Image / build-variant (push) Successful in 18m51s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / update-description (push) Successful in 18s
Finalize the Unreleased batch as v1.4.0 (minor — significant base
additions). Base rebuilds (Dockerfile.base for typst/xz-utils;
.bash_aliases for the SSH check). pi auto-resolves latest (0.80.3 ->
0.80.6); mempalace stays 3.5.0 (current).
2026-07-11 17:15:11 +02:00
joakimp bd0627a557 docs: note host SSH startup check in CHANGELOG Unreleased
Lint workflows / actionlint (push) Successful in 1m18s
2026-07-11 17:07:55 +02:00
pi 67da05b99b feat(base): ship typst as lightweight pandoc PDF engine
Lint workflows / actionlint (push) Successful in 15s
pandoc has been in the base since v1.0.0 but only as a front-end; PDF
export (studio_export_pdf / pandoc -o out.pdf) failed with 'xelatex not
found' because no back-end engine was installed. Ship typst (~30 MB
static Rust binary, no LaTeX) as the default engine via
`pandoc --pdf-engine=typst`, chosen over a ~600 MB TeX Live install.
texlive-xetex remains the higher-fidelity install-on-demand fallback.

- Dockerfile.base: install typst (latest GitHub-release idiom, pin via
  TYPST_VERSION); add xz-utils (typst ships .tar.xz); bump
  BASE_REBUILD_DATE. Lands in base-<hash>.
- smoke-test.sh: verify typst + a real pandoc --pdf-engine=typst render.
- README/AGENTS/CHANGELOG: typst now shipped (supersedes the planned
  :latest-studio-tex variant).

No tag pushed — CI build intentionally deferred.
2026-07-11 17:04:19 +02:00
joakimp 4563b4d76d feat: warn at shell startup if Mac host SSH is not reachable
Lint workflows / actionlint (push) Successful in 14s
Adds a _devbox_check_host_ssh() check to ~/.bash_aliases (baked into
the image). On first bash session of each container it tries a quick
SSH probe to the Mac host; if it fails it prints a clear one-time
warning with the exact two steps needed to fix it:

  1. Enable Remote Login in macOS System Settings
  2. echo '<public key>' >> ~/.ssh/authorized_keys

The check is guarded:
  - only runs inside a container (/.dockerenv)
  - only when the jump key exists (~/.ssh-local/devbox_jump_ed25519.pub)
  - only once per container lifetime (/tmp flag, cleared on recreate)

After --force-recreate the key changes, the flag is gone, and the
check runs again on the first bash window. Subsequent windows are
silent.
2026-07-11 15:26:23 +02:00
pi f19c35da32 docs: note lightweight PDF engine (typst) as preferred over full texlive
Lint workflows / actionlint (push) Successful in 36s
PDF export from Studio/pandoc still isn't shipped. Record the engine decision
in the living docs (README + AGENTS): pandoc is in the image but has no PDF
back-end, so export fails with 'xelatex not found'. Prefer a lightweight engine
— typst (~30 MB static binary, 'pandoc --pdf-engine=typst') which is small
enough it could ship in base rather than needing a separate ':latest-studio-tex'
variant; texlive-xetex (~600 MB) kept as the higher-fidelity fallback. Also drop
stale 'v1.3.0' pins (v1.3.0 already shipped without PDF) in favour of 'a future
release'. CHANGELOG history left untouched.
2026-07-09 16:25:54 +02:00
pi 6002c6299d release: v1.3.0 — shared/external MemPalace + nano/micro editors + CI lint
Publish Docker Image / resolve-versions (push) Successful in 5s
Publish Docker Image / base-decide (push) Successful in 8s
Lint workflows / actionlint (push) Successful in 19s
Publish Docker Image / build-base (push) Successful in 33m17s
Publish Docker Image / smoke (push) Successful in 3m55s
Publish Docker Image / smoke-studio (push) Successful in 6m50s
Publish Docker Image / build-variant (push) Successful in 16m12s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / build-variant-studio (push) Successful in 20m50s
Promotes the Unreleased block to v1.3.0. Bundled contents:
- feat: shared/external MemPalace — mempalace.ts bridge honours MEMPALACE_REMOTE_URL;
  adds docker-compose.mempalace.yml.
- feat: nano + micro non-modal editors (Dockerfile.base → base rebuild).
- ci: lint.yml (sh-vs-bash guard + actionlint/shellcheck), docker-publish.yml bash
  defaults, promote-base-latest shell fix.
pi stays 0.80.3 (== npm latest). Base rebuilds (mempalace-toolkit ref advanced +
Dockerfile.base nano/micro), so the new bridge + editors land in base-<hash>.
2026-07-02 14:50:17 +02:00
pi d73bf2e9d3 feat: optional shared/external MemPalace via MEMPALACE_REMOTE_URL
Lint workflows / actionlint (push) Successful in 21s
Wire the shared-palace option (implemented in mempalace-toolkit's mempalace.ts)
into the container:
- .env.example: document MEMPALACE_REMOTE_URL / MEMPALACE_REMOTE_TOKEN (env_file-only,
  per this repo's convention).
- docker-compose.mempalace.yml: optional shared server (mempalace-mcp --transport http),
  loopback-bound by default.
- docker-compose.yml: local-vs-external note on the palace-volume comment.
- README + CHANGELOG (Unreleased).
2026-07-02 13:09:29 +02:00
pi 3a59e15563 feat: ship nano + micro (non-modal editors) alongside nvim
Lint workflows / actionlint (push) Successful in 13s
The image shipped only nvim (EDITOR=nvim), a modal vi-style editor. Not
everyone is comfortable with vi keybindings, so add both a classic and a
modern non-modal option:

- nano (apt): ~2.8 MB installed; deps (libc6, libncursesw6, libtinfo6)
  already present via nvim/less/htop/tmux, so no extra packages pulled in.
- micro: ~12 MB single static Go binary from GitHub releases (same pattern
  as bat/eza/zoxide). Desktop-style keys (Ctrl+S/Ctrl+Q), mouse, syntax
  highlighting. ARG MICRO_VERSION pins; defaults to latest.

Combined ~15 MB (<0.5% of the ~3.2 GB image). EDITOR stays nvim; both new
editors are opt-in (export EDITOR=micro | nano). Uses the canonical
micro-editor/micro URL because the old zyedidia/micro org rename makes
/releases/latest redirect to another /latest, defeating the tag-parsing
latest-resolution idiom.

Base-image change, so it lands on the next base-<hash> rebuild. Updates
README tool table + EDITOR note, CHANGELOG (Unreleased/Added), and
smoke-test.sh (nano + micro presence checks).
2026-07-01 23:11:25 +02:00
pi d1db595f17 ci(lint): pass explicit workflow paths to actionlint
Lint workflows / actionlint (push) Successful in 15s
actionlint's no-arg project auto-detection looks for .github/workflows
and hard-fails (exit 3, 'no project was found') on this .gitea/workflows
layout — observed on run 420. Glob the workflow files explicitly. The
Gitea shell guard step already passed in that run; only the actionlint
invocation needed the path fix.
2026-07-01 22:06:40 +02:00
pi 26384fe9f1 ci: eliminate the sh-vs-bash footgun class (defaults + lint guard)
Lint workflows / actionlint (push) Failing after 34s
Root cause of the recurring 'Illegal option -o pipefail' failures
(ed49b8d resolve-versions; b7197e8 promote-base-latest, run 418):
docker-publish.yml had no workflow-level default shell, so Gitea's
sh/dash default applied and every bash-syntax step had to individually
remember 'shell: bash'.

- docker-publish.yml: add 'defaults: run: shell: bash' — fixes the whole
  class; all pre-existing dash steps are POSIX so bash runs them unchanged.
- lint.yml: new workflow, runs on every push/PR (not just release tags):
    * scripts/check-workflow-shell.sh — Gitea-accurate guard: fails if any
      run: step doesn't resolve to bash. Catches the omit-shell+bash-syntax
      case that actionlint MISSES (actionlint models GitHub, where the
      default shell is bash, so a shell-less step is assumed bash).
    * actionlint + shellcheck — catches explicit 'shell: sh' + bash syntax
      (SC3040) and general workflow errors.
  Verified locally: guard + actionlint pass current workflows; guard fails
  a synthetic omit-shell+pipefail workflow; shellcheck clean.
2026-07-01 22:05:04 +02:00
pi b33e9dc592 fix(ci): promote-base-latest re-tag step needs shell: bash (set -o pipefail)
b7197e8 moved the digest-compare into the re-tag step with 'set -euo
pipefail' but no 'shell: bash'; Gitea's default sh (dash) aborts on
-o pipefail, leaving base-latest un-promoted on the v1.2.4 release
(run 418). Same footgun as ed49b8d. Consumer tags unaffected (they
FROM base-<hash>, not base-latest).
2026-07-01 17:57:13 +02:00
pi 3cdc2069db release: v1.2.4 — pi 0.80.2 → 0.80.3; global gitignore, env_file-only secrets, promote-base-latest CI fix
Publish Docker Image / build-variant (push) Successful in 16m16s
Publish Docker Image / promote-base-latest (push) Failing after 4s
Publish Docker Image / build-variant-studio (push) Successful in 17m52s
Publish Docker Image / update-description (push) Successful in 12s
Publish Docker Image / resolve-versions (push) Successful in 1m2s
Publish Docker Image / base-decide (push) Successful in 43s
Publish Docker Image / build-base (push) Successful in 41m46s
Publish Docker Image / smoke (push) Successful in 4m10s
Publish Docker Image / smoke-studio (push) Successful in 6m42s
2026-07-01 16:49:25 +02:00
pi cc53877328 feat: bake global gitignore (core.excludesFile) into image
Seed ~/.gitignore_global from /etc/skel-devbox (seed-if-absent, like
.bash_aliases/.inputrc, so user edits survive recreate) and wire it via
git config --global core.excludesFile, guarded so a user-set excludesFile
is never overridden. Ignores *.bak, *.bak.*, *~, *.orig, *.swp, *.tmp
across all repos without per-repo .gitignore entries.
2026-06-28 11:52:02 +02:00
pi c42b237d30 compose: deliver secrets via env_file only (drop environment: passthrough)
Removes GITEA_ACCESS_TOKEN / GITEA_HOST / GITHUB_PERSONAL_ACCESS_TOKEN from
the compose environment: block. An environment: entry both overrides
env_file AND is interpolated from the host shell, so a stale shell export
(e.g. one auto-loaded by an opencode/dotenv hook) silently shadowed the
users .env — an updated token never reached the container. Secrets now flow
solely via env_file: .env; .env.example already documents every variable.

- docker-compose.yml: drop the 3 passthrough lines + explanatory comment
- README.md: sync the "basic shape" snippet
- CHANGELOG.md: note under Unreleased (no tag bump / unpublished)
2026-06-27 23:48:02 +02:00
pi b7197e88b0 ci(promote-base-latest): re-point base-latest by digest, not need_build
The gate keyed off need_build=='true', assuming need_build==false meant
base-latest was already current. A dry-run dispatch (promote_latest=false)
that pre-builds base-<hash> falsifies that: the later tag run sees
need_build==false and skipped promotion, leaving base-latest one base
behind (observed 2026-06-27, v1.2.3 dry-run-first release).

Gate now runs on every tag release / promote dispatch; the no-op
optimization moved into the step as a crane digest compare so it re-tags
only when base-latest actually differs from the released base-<hash>.
Workflow-only change; base hash unaffected (no base rebuild).
2026-06-27 20:57:03 +02:00
pi 2985d9ade8 release: v1.2.3 — mempalace-mcp self-heal (toolkit e12b624)
Publish Docker Image / resolve-versions (push) Successful in 7s
Publish Docker Image / base-decide (push) Successful in 14s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Successful in 3m30s
Publish Docker Image / smoke-studio (push) Successful in 11m47s
Publish Docker Image / build-variant (push) Successful in 15m48s
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / build-variant-studio (push) Successful in 17m28s
Patch release. Headline: mempalace-mcp self-heals instead of latching
available=false permanently after a slow virtiofs cold-open. Base image
rebuilds via the mempalace-toolkit ref advancing to e12b624 (folded into
the base-decide hash). No pi/mempalace version change — pi npm latest is
still 0.80.2 (= v1.2.2). Also releases the queued yq (mikefarah Go yq) and
mempalace-skill temporal-grounding changes.
2026-06-27 18:45:24 +02:00
pi bff810c1eb docs(dockerfile): sync mempalace stall-protection comment with self-heal
mempalace.ts now self-heals (respawn with capped backoff) instead of
latching unavailable, and the init-timeout default is 300000. Update the
explanatory comment + tunable list (MEMPALACE_MCP_MAX_RESPAWNS,
MEMPALACE_MCP_RESPAWN_BACKOFF_MS). Comment-only; no build/ENV change.
2026-06-26 00:22:43 +02:00
pi 904fe85249 skill(mempalace): teach temporal grounding (recreate != new day)
Baked mempalace SKILL.md now instructs agents to establish current date/time
and compute the delta against the actual diary/drawer timestamp before using
relative terms (yesterday/last week), and explicitly that a container recreate
or fresh session is NOT a day boundary (pi-devbox restarts several times a day).
Phase 1 wake-up section + anti-pattern bullet. CHANGELOG Unreleased.
2026-06-25 22:53:31 +02:00
pi cda488c565 base: yq follows latest (was pinned v4.53.3), gate on major v4
Match the repo's latest-following convention (tealdeer/uv/etc.) and keep the
container in sync with the Mac's brew yq. smoke-test now asserts mikefarah AND
major v4, so a surprise yq v5 fails CI instead of silently breaking
provision.sh. Pin still available via --build-arg YQ_VERSION=vX.Y.Z.
2026-06-25 16:33:27 +02:00
pi 9ab9a28458 base: install mikefarah yq (pinned v4.53.3), drop Debian python yq
Debian/Ubuntu `apt install yq` is kislyuk/yq (Python, v3.x), incompatible
with the mikefarah v4 syntax the cloud-init repo's provision.sh/deploy.sh
require. Replace the apt package with a pinned mikefarah Go binary, mirroring
the existing tealdeer ARG (latest-or-pin) pattern, multi-arch amd64/arm64.
smoke-test.sh now asserts `yq --version` reports mikefarah so CI catches a
regression. CHANGELOG: Unreleased entry.
2026-06-25 16:29:51 +02:00
pi d175b31207 release: v1.2.2 — pi 0.80.2 + mempalace 3.5.0, drop anyOf workaround
Publish Docker Image / resolve-versions (push) Successful in 8s
Publish Docker Image / base-decide (push) Successful in 11s
Publish Docker Image / build-base (push) Successful in 33m19s
Publish Docker Image / smoke-studio (push) Successful in 4m3s
Publish Docker Image / smoke (push) Successful in 9m51s
Publish Docker Image / build-variant-studio (push) Successful in 17m29s
Publish Docker Image / build-variant (push) Successful in 24m10s
Publish Docker Image / update-description (push) Successful in 10s
Publish Docker Image / promote-base-latest (push) Successful in 14s
- Bump mempalace pin 3.4.0 -> 3.5.0: 3.5.0 carries the upstream fix for the
  top-level-anyOf diary_write schema (issue #1728 / PR #1717, merged
  2026-06-14). Verified against the published 3.5.0 wheel that mcp_server.py
  now advertises 'required: [agent_name]' with no root-level anyOf.
- Remove the Dockerfile.base perl workaround that stripped the anyOf from the
  installed mcp_server.py — obsolete now the fix is upstream.
- pi auto-resolves to npm latest (0.79.10 -> 0.80.2) at build.
- CHANGELOG v1.2.2.
2026-06-25 07:56:55 +02:00
Joakim Persson 13e67599c4 release: v1.2.1 — fallback skills + mempalace directive
Publish Docker Image / smoke-studio (push) Successful in 6m27s
Publish Docker Image / build-variant (push) Successful in 16m16s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 8s
Publish Docker Image / build-variant-studio (push) Successful in 21m13s
Publish Docker Image / resolve-versions (push) Successful in 6s
Publish Docker Image / base-decide (push) Successful in 7s
Publish Docker Image / build-base (push) Successful in 46m21s
Publish Docker Image / smoke (push) Successful in 3m43s
Bake pi-extensions + mempalace skills into the image (available without a
mounted skillset) and add the mempalace session-start proactive-load directive
so frequently-recreated containers actually pick the skill up. Closes the
fork/recall + mempalace under-utilisation gap.

CHANGELOG: [Unreleased] -> v1.2.1.
2026-06-23 16:02:57 +02:00
Joakim Persson 7551947466 feat(skills): add mempalace proactive-load directive for containers
Baking the mempalace fallback skill fixed *availability*, but mempalace had
no proactive-load directive anywhere (pi-toolkit's global AGENTS.md only
points to pi-extensions), so a new container would still surface it only via
description-matching — the same under-utilisation the pi-extensions directive
was created to fix.

Add a session-start pointer to the pi-devbox managed AGENTS.md block
(pi-global-AGENTS.append.md): gated to pi-devbox containers and conditional on
the MemPalace MCP tools being present. Memory continuity matters most in a
frequently-recreated container — the palace is its only cross-recreate memory.

- pi-global-AGENTS.append.md: '## Session start: load the mempalace skill'.
- smoke-test: assert the pointer merges into the global AGENTS.md at build.
- docs: VENDORED.md, README, CHANGELOG [Unreleased].

Now both skills are complete in pi-devbox: directive + skill file.
pi-extensions = directive (pi-toolkit) + baked skill; mempalace = directive
(this block) + baked skill.
2026-06-23 15:54:13 +02:00
Joakim Persson a7d6a7d235 feat(skills): bake pi-extensions + mempalace fallback skills
The pi-toolkit global AGENTS.md tells every pi session to read
~/.agents/skills/pi-extensions/SKILL.md at start (the fork/recall
under-utilisation fix), but that skill lived only in the private skillset
repo — so the pointer dangled in any container started without skillset
mounted. Bake fallbacks so the pointer always resolves.

- pi-extensions (Option 1 + Option 2, layered):
  * Canonical skill promoted to the public pi-extensions package repo under
    skill/ (separate commit there); co-located with the code it documents.
  * rootfs/ carries a committed snapshot (the floor).
  * Dockerfile.variant copies /opt/pi-extensions/skill/ over the snapshot
    after the pinned clone, so a normal build ships the fresh package copy
    (recorded via PI_EXTENSIONS_REF) and an old-ref/mirror build still ships
    the snapshot. Helper evaluate-extension-usage.py travels with it.
- mempalace (Option 2 only): snapshot in rootfs/. Its consumer skill has no
  public package home (mempalace-toolkit ships a different skill,
  opencode-mempalace-bridge), so no build-time refresh.
- entrypoint links both (only-when-absent; mounted skillset still wins).
- smoke-test: build-time presence + package-match check + runtime symlink
  assertions; readiness gate now waits on the last-linked skill.
- docs: skills/VENDORED.md (provenance + refresh), README, AGENTS.md,
  CHANGELOG [Unreleased].

Note: shipped in the NEXT release; v1.2.0 (run 409) predates this.
2026-06-23 15:32:04 +02:00
Joakim Persson d619a6e2ec fix(entrypoint,smoke): link image-baked skills early to fix smoke race
Publish Docker Image / resolve-versions (push) Successful in 21s
Publish Docker Image / base-decide (push) Successful in 7s
Publish Docker Image / build-base (push) Successful in 33m43s
Publish Docker Image / smoke-studio (push) Successful in 4m5s
Publish Docker Image / smoke (push) Successful in 5m42s
Publish Docker Image / build-variant (push) Successful in 15m55s
Publish Docker Image / promote-base-latest (push) Successful in 7s
Publish Docker Image / build-variant-studio (push) Successful in 17m45s
Publish Docker Image / update-description (push) Successful in 56s
The runtime 'pi-devbox-environment skill linked' smoke assertion failed in
CI run 408 (gating build-variant). Root cause: the skill-linking block ran
AFTER the pi-toolkit/extensions deploy, but the smoke readiness gate only
waits on pi-deploy markers (keybindings.json, mempalace.ts) — which land
before the skill symlink — so the assertion sampled too early.

- entrypoint-user.sh: move the image-baked-skills symlink loop to run early
  (before the pi deploy block), so it completes before any readiness marker.
  Still before the skillset deploy, so foreign-link semantics are unchanged.
- smoke-test.sh: add the skill symlink to the readiness gate as well.

Build-time checks (baked skill, append snippet, merged AGENTS marker) all
passed in 408; only the timing of the runtime check was wrong.
2026-06-23 14:29:52 +02:00
Joakim Persson 2abfee141b feat: image-baked agent skills + pi-devbox-environment skill (v1.2.0)
Publish Docker Image / resolve-versions (push) Successful in 35s
Publish Docker Image / base-decide (push) Successful in 23s
Publish Docker Image / build-base (push) Successful in 41m32s
Publish Docker Image / smoke-studio (push) Failing after 4m5s
Publish Docker Image / build-variant-studio (push) Has been skipped
Publish Docker Image / smoke (push) Failing after 5m46s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Ship skills inside the image (independent of any mounted skillset repo):
- rootfs/usr/local/share/pi-devbox/skills/<name>/ symlinked into
  ~/.agents/skills/ by entrypoint-user.sh (foreign-link, survives volume
  recreate, never clobbers a skillset/user skill of the same name).
- New pi-devbox-environment skill: persistence model, host/LAN SSH
  reachability, split-DNS mechanisms, interactive-vs-tool-shell alias
  gotcha, tmux 0-index, uv-first Python, pi-studio reachability. Agnostic
  to host OS / hostnames / domains / nameservers (discovered at runtime).
- Dockerfile.variant appends pi-global-AGENTS.append.md onto pi-toolkit's
  pi-global-AGENTS.md (single global slot) so the skill is loaded
  proactively; gated on /usr/local/lib/pi-devbox/. Idempotent.
- smoke-test: baked-skill + append-snippet + merged-marker presence and a
  runtime symlink assertion.
- docs: README 'Agent skills' section, AGENTS.md layout, DOCKER_HUB.md;
  moved studio-tex roadmap to v1.3.0.

pi 0.79.7 -> 0.79.10 (auto-resolved from npm latest at build).
2026-06-23 12:49:13 +02:00
pi c346a106a3 release: v1.1.7 — pi 0.79.8 → 0.79.9; ssh-lan.conf LAN-peer docs
Publish Docker Image / base-decide (push) Successful in 1m0s
Publish Docker Image / resolve-versions (push) Successful in 6s
Publish Docker Image / build-base (push) Successful in 33m15s
Publish Docker Image / smoke (push) Successful in 3m32s
Publish Docker Image / smoke-studio (push) Successful in 3m47s
Publish Docker Image / build-variant-studio (push) Successful in 17m39s
Publish Docker Image / build-variant (push) Successful in 19m13s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / update-description (push) Successful in 13s
2026-06-21 23:36:43 +02:00
joakimp 8de0fad776 docs(lan): document ssh-lan.conf for naming LAN peers
The host-owned, bind-mounted ~/.config/devbox-shell/ssh-lan.conf is the
intended place to add `ProxyJump host` overrides for named LAN peers (so
`pi --ssh <peer>` / `dssh <peer>` route through the host), but it was only
documented in .env.example and the setup-lan-access.sh header — never in the
README, where someone hitting "can't reach LAN peers" actually looks.

- README: add a "Naming LAN peers" subsection under the macOS LAN-peers
  troubleshooting block, with a ProxyJump example and the read-only ~/.ssh
  caveat; add a pointer to it from the SSH and ControlMaster section.
- setup-lan-access.sh: correct the INCLUDE_BLOCK comment that suggested adding
  ProxyJump to the read-only ~/.ssh/config; point at ssh-lan.conf instead.
- CHANGELOG: note under Unreleased.

Docs/comment only — no behavior change.
2026-06-21 00:23:29 +02:00
pi ed49b8d97a fix(ci): resolve-versions needs shell: bash for 'set -o pipefail'
Publish Docker Image / smoke (push) Successful in 9m0s
Publish Docker Image / build-variant-studio (push) Successful in 17m41s
Publish Docker Image / build-variant (push) Successful in 19m1s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 10s
Publish Docker Image / resolve-versions (push) Successful in 6s
Publish Docker Image / base-decide (push) Successful in 11s
Publish Docker Image / build-base (push) Successful in 45m54s
Publish Docker Image / smoke-studio (push) Successful in 3m43s
The default run shell is 'sh -e {0}' (dash on the act runner), which
rejects 'set -o pipefail' ('Illegal option -o pipefail') — failing the
resolve-versions job on line 2 and cascading every dependent job to
skipped (v1.1.6 run 401). The heavy build steps already declare
'shell: bash'; the resolve step did not. Added it.
2026-06-19 18:26:04 +02:00
pi 9eff3f3c48 release: v1.1.6 — build provenance + reproducibility hardening; pi 0.79.7 → 0.79.8
Publish Docker Image / resolve-versions (push) Failing after 52s
Publish Docker Image / base-decide (push) Has been skipped
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke-studio (push) Has been skipped
Publish Docker Image / smoke (push) Has been skipped
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / build-variant-studio (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Adds OCI labels + /etc/pi-devbox/build-manifest.json so a published tag is
self-describing and reconstructable after CI logs rotate (manifest is
written from the actual checked-out HEAD of each /opt clone + live
pi --version, not just the intended build-args).

Hardens the build plumbing:
- scripts/check-base-hash.sh guards the base-rebuild invariant: every
  floating ARG *_REF in Dockerfile.base must be folded into the base_tag
  hash, else a ref-only change silently fails to rebuild the base
  (v1.1.2-class staleness footgun). Runs in base-decide and locally.
- resolve-versions now fails loud instead of falling back to a floating
  main/master on a transient API failure — validates each ref is a 40-hex
  SHA (and pi a real semver) and aborts the release otherwise.
- The three gitea companions (pi-toolkit, pi-extensions, mempalace-toolkit)
  gained overridable *_REPO build-args (defaulting to the canonical gitea
  origin) so a relocated/forked build can repoint them without editing the
  Dockerfiles — matching the existing PI_FORK_REPO/PI_OBSMEM_REPO pattern.

README documents the forked/relocated build-arg trick and how to read the
labels + manifest. smoke-test asserts the manifest + labels. pi bumps
0.79.7 → 0.79.8 (auto-resolved at build).
2026-06-19 18:23:11 +02:00
Joakim Persson a0abacaafb fix(ssh): survive read-only ~/.ssh ControlPath; render sidecar on all host OSes
Publish Docker Image / smoke (push) Successful in 3m22s
Publish Docker Image / smoke-studio (push) Successful in 3m42s
Publish Docker Image / build-variant (push) Successful in 15m29s
Publish Docker Image / update-description (push) Successful in 11s
Publish Docker Image / build-variant-studio (push) Successful in 16m49s
Publish Docker Image / promote-base-latest (push) Successful in 14s
Publish Docker Image / resolve-versions (push) Successful in 8s
Publish Docker Image / base-decide (push) Successful in 8s
Publish Docker Image / build-base (push) Successful in 33m44s
Coordinated with the pi-extensions ssh-controlmaster fix (picked up at build via
PI_EXTENSIONS_REF=main), this makes `pi --ssh <host>` and `dssh`/`dscp` robust
to a user ~/.ssh/config whose per-host ControlPath points under the read-only
~/.ssh bind-mount (e.g. `ControlPath ~/.ssh/cm/%r@%h:%p`). A system default can
never override a user's per-host value, so the fix lives in two layers.

- setup-lan-access.sh: always render the writable ~/.ssh-local/config sidecar
  (Host * ControlPath redirect into ~/.ssh-local/cm + Include ~/.ssh/config) on
  EVERY host OS. Previously the script exited early (no-op) on native Linux,
  leaving dssh/dscp broken when ~/.ssh was read-only there too. The host-jump
  block, its key generation, and the authorize hints stay gated on VM-backed
  detection / DEVBOX_LAN_ACCESS=jump (new NEED_JUMP flag).
- Dockerfile.base: document that the /etc/ssh drop-in default cannot override a
  user per-host ControlPath; cross-ref the two handling layers.
- entrypoint-user.sh: correct the now-stale "no-op on native Linux" comment.
- README.md / DOCKER_HUB.md: document read-only-~/.ssh ControlPath handling.

CHANGELOG: v1.1.5 (Fixed + Changed + pi 0.79.6 -> 0.79.7 auto-resolved bump).
2026-06-18 21:59:18 +02:00
Joakim Persson da7d70825e docs(changelog): add v1.1.4 entry (AGENTS.md autoload, settings merge, history fix)
Publish Docker Image / resolve-versions (push) Successful in 6s
Publish Docker Image / base-decide (push) Successful in 11s
Publish Docker Image / build-base (push) Successful in 42m4s
Publish Docker Image / smoke-studio (push) Successful in 3m39s
Publish Docker Image / smoke (push) Successful in 5m20s
Publish Docker Image / build-variant (push) Successful in 18m8s
Publish Docker Image / promote-base-latest (push) Successful in 9s
Publish Docker Image / update-description (push) Successful in 10s
Publish Docker Image / build-variant-studio (push) Successful in 24m41s
2026-06-17 20:52:05 +02:00
Joakim Persson 41c2c2b716 feat(entrypoint): non-destructively merge new template keys into settings.json
The settings.json bootstrap only fires when the file is ABSENT, so a
settings.json on a preserved named volume never picks up config added in a
later image (e.g. the observational-memory / pi-fork blocks, a newly-enabled
model). Users had to hand-merge after every upgrade.

On start, when settings.json already exists, deep-merge the template into it
with 'jq -s ".[0] * .[1]"' (template first, live second) so the user's values
always win and only MISSING keys are filled from the template. Arrays are
leaves (a model the user removed is not re-added). Rewrites only when the
merge changes something, backs up the original first, and skips safely (no
clobber) if either file is invalid JSON. Opt out with PI_SETTINGS_MERGE=0.

Add a recreate-sanity-check assertion that settings.json carries the
observational-memory + pi-fork blocks after recreate.
2026-06-17 20:49:41 +02:00
Joakim Persson 5c08bfc8a8 fix(shell): don't export DEVBOX_HIST_SET so nested shells flush history
The history-flush guard was exported, so it leaked into child processes.
Any nested shell -- crucially each tmux pane (which inherits the tmux
server's env) -- then saw the guard already set and skipped installing
'history -a' in PROMPT_COMMAND. Those shells only persisted history on a
clean exit, so abrupt termination (docker stop, tmux kill-server, SIGKILL)
silently lost their in-memory history. zoxide was less affected (its hook
is installed unguarded and writes immediately).

Make the guard shell-local (drop 'export') so every new interactive shell
re-installs its own per-prompt flush. Add a recreate-sanity-check assertion
that a nested login shell still wires up 'history -a'.

Storage was never the issue: ~/.cache/bash (devbox-shell-history) and
~/.local/share/zoxide (devbox-zoxide) are both persistent named volumes.
2026-06-17 17:22:30 +02:00
Joakim Persson 1371584634 sanity-check: verify global AGENTS.md symlink after recreate
pi-toolkit now symlinks pi-global-AGENTS.md -> ~/.pi/agent/AGENTS.md (pi's
global-instructions file, loaded at every start; directs the agent to read
the pi-extensions skill at session start). Add a recreate-sanity-check
assertion alongside the keybindings symlink check so a future image build
that bakes the new pi-toolkit verifies the wiring landed.
2026-06-17 16:58:18 +02:00
Joakim Persson d902b2d056 v1.1.3: add actual pi 0.79.5 release notes + document GitHub releases URL in AGENTS.md 2026-06-16 23:56:52 +02:00
Joakim Persson c48abf41d1 v1.1.3: pi 0.79.4 → 0.79.5
Publish Docker Image / resolve-versions (push) Successful in 6s
Publish Docker Image / base-decide (push) Successful in 11s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke (push) Successful in 3m6s
Publish Docker Image / smoke-studio (push) Successful in 9m55s
Publish Docker Image / build-variant (push) Successful in 15m18s
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Successful in 11s
Publish Docker Image / build-variant-studio (push) Successful in 25m41s
2026-06-16 23:54:05 +02:00
pi 777d53354f docs(AGENTS): document GITEA_ACCESS_TOKEN env for general Gitea API access
GITEA_ACCESS_TOKEN + GITEA_HOST (passed from host .env via compose,
primarily for gitea-mcp) are also usable for any direct Gitea API work —
run inspection, tag checks — not just ci-release-watcher. Prefer over a
PAT file when present; host-managed lifecycle, nothing to revoke. Release
checklist step 7 now notes the env-token alternative.
2026-06-15 22:30:36 +02:00
37 changed files with 9573 additions and 272 deletions
+35
View File
@@ -0,0 +1,35 @@
# Keep the Docker build context minimal and prevent stray files (notably
# `.git`) from ever being pulled in by a future broad COPY. Both Dockerfiles
# only COPY `rootfs/` and `entrypoint*.sh`, so everything below is safe to
# exclude from the context.
#
# DO NOT add `rootfs/`, `entrypoint.sh`, `entrypoint-user.sh`, or the
# Dockerfiles here — they are required to build the image.
# VCS / CI metadata
.git
.gitea
.gitignore
.dockerignore
# Lint / editor config
.hadolint.yaml
.editorconfig
# Docs & project meta
README.md
DOCKER_HUB.md
CHANGELOG.md
AGENTS.md
IDEAS.md
LICENSE
THIRD_PARTY.md
docs
# Local orchestration & examples (compose runs the image; not a build input)
docker-compose.yml
docker-compose.mempalace.yml
.env.example
# Repo tooling / tests (run from a checkout, not baked into the image)
scripts
+93
View File
@@ -9,6 +9,84 @@ WORKSPACE_PATH=~/projects
# Path to SSH keys on host
SSH_KEY_PATH=~/.ssh
# ── MemPalace memory (local by default) ───────────────────────────
# By default the mempalace.ts extension spawns a LOCAL mempalace-mcp stdio
# server (palace at ~/.mempalace). Uncomment the devbox-palace volume in
# docker-compose.yml to persist it across container recreation — that one
# volume now covers the mined conversation transcripts too, since the pi and
# opencode feeders stage inside the palace root (<palace-root>/pi-stage), so
# the staged files and the palace dedup keys pointing at them cannot be
# separated.
#
# That palace root is resolved with mempalace's own precedence
# ($MEMPALACE_PALACE_PATH -> $MEMPAL_PALACE_PATH -> ~/.mempalace/config.json ->
# ~/.mempalace/palace), and the feeders derive their stage FROM it
# (<palace-root>/pi-stage). Neither the image nor the entrypoint exports it, by
# design: pinning the palace without carrying the stage along re-creates the
# very split that a shared root removed. Override it only to move the palace off
# the default -- e.g. onto a different mount -- and only to a path with the SAME
# persistence as the palace itself. A stage that outlives its palace (or dies
# first) makes a scoped `mempalace sync` prune conversation drawers, because
# their dedup key is the staged path. Setting it to the default buys nothing.
# Unlike WORKSPACE_PATH/SSH_KEY_PATH above, this is a path INSIDE the container.
# MEMPALACE_PALACE_PATH=/home/developer/.mempalace/palace
#
# To instead share ONE MemPalace across containers/harnesses (pi + opencode
# + native), set the URL below. When set, the extension connects over HTTP
# and NO local mempalace-mcp is spawned; the devbox-palace volume is then
# irrelevant. MEMPALACE_REMOTE_TOKEN, if set, is sent as a bearer token.
#
# Serve it with: mempalace serve --host 172.17.0.1 --port 8765
#
# NOT `mempalace-mcp --transport http --host 0.0.0.0`: `serve` is the turnkey
# wrapper that mints/keeps a bearer token (0600, passed via env so it stays out
# of `ps`) and can terminate TLS. Two binds to avoid:
# 0.0.0.0 - exposes the palace to the whole LAN.
# 127.0.0.1 - behind a tunnel this 403s every proxied request (the Host pin
# is only enforced on loopback binds) AND silently starts with
# no token at all, since auto-minting is gated on the bind being
# non-loopback. Bind the docker0 gateway: reachable from the host
# and its containers (so a newt/proxy container works), not from
# the LAN. Set MEMPALACE_MCP_HTTP_TOKEN explicitly server-side.
# MEMPALACE_REMOTE_URL=https://mempalace.example.com/mcp
# MEMPALACE_REMOTE_TOKEN=
# ── MemPalace: automatic capture of pi sessions ───────────────────────
# The mempalace.ts extension feeds this container's pi transcripts into the
# palace by itself: on session_shutdown, and on a debounced agent_settled so a
# crash loses at most one window rather than the whole session. The entrypoint
# also runs a catch-up at container start, which is the only thing that can
# recover transcripts after a hard kill (no handler runs on SIGKILL).
# Nothing below is required for the local-palace case; the defaults work.
#
# MEMPALACE_FEED=0 # disable automatic capture entirely
# MEMPALACE_FEED_DEBOUNCE_MS=600000 # min gap between mid-session feeds (10 min)
# MEMPALACE_FEED_WING=wing_conversations
#
# REMOTE PALACE ONLY (MEMPALACE_REMOTE_URL set above): the palace is on another
# host, and `mempalace_mine` resolves its source path in the SERVER process, so
# the server cannot see this container's transcripts. The feeder therefore
# rsyncs its staged exports into a per-device inbox on the palace host and asks
# the server to mine its own local copy. Without MEMPALACE_PI_SSH_TARGET the
# feeder is skipped (a remote palace with no inbox has nothing to mine).
# MEMPALACE_PI_SSH_TARGET where to rsync to, as user@host:path
# MEMPALACE_PI_REMOTE_PATH what that inbox is called ON THE SERVER — i.e. the
# path the SERVER PROCESS can open. If the palace
# server runs in Docker, that is the container path
# (see docker-compose.mempalace.yml). If it runs
# NATIVELY (systemd unit / uv tool / plain
# `mempalace serve`), it sees host paths, so this
# must equal the path half of
# MEMPALACE_PI_SSH_TARGET. Getting this wrong is
# quiet: rsync still succeeds and only the mine
# fails with "source directory not found", so
# transcripts ship and are filed nowhere. The feeder
# warns in preflight when the two paths disagree.
# MEMPALACE_PI_DEVICE inbox subdirectory for this machine (default: hostname)
# MEMPALACE_PI_SSH_TARGET=user@palace-host:/srv/mempalace-feed
# MEMPALACE_PI_REMOTE_PATH=/data/feed
# MEMPALACE_PI_DEVICE=
# ── LAN access from the container (host-OS-agnostic) ─────────────────
# On VM-backed hosts (macOS OrbStack / Docker Desktop) the container can't
# reach the host's directly-attached LAN peers by default. The entrypoint
@@ -30,7 +108,22 @@ SSH_KEY_PATH=~/.ssh
# the host, so bare `dssh user@<ip>` works on whatever LAN you're roaming on.
# DEVBOX_LAN_AUTOJUMP_PRIVATE=0
# ── pi-atelier (TUI sidebar) ─────────────────────────────────────────
# The image vendors pi-atelier at a pinned, audited tag and registers it on
# container start. Set to 0 to opt out: the entrypoint then removes it from
# pi's `packages[]` instead of registering it. This lives here rather than
# being a `pi uninstall` because a broken TUI extension's failure mode is
# "pi will not start", which you cannot fix from inside pi.
# DEVBOX_ATELIER=1
# ── Git Configuration ────────────────────────────────────────────────
# Set BOTH. If unset, every repo inside the container fails with
# "Author identity unknown" on first commit, and an agent asked to commit
# will guess an identity from git log — often the wrong one. The e-mail is
# per-machine (work machines use the corporate address, personal machines the
# private one), so it belongs in this per-machine .env, never in a skill or a
# repo-local override. Consumed by entrypoint-user.sh -> ~/.gitconfig, which is
# NOT persistent across container recreate — this file is the source of truth.
GIT_USER_NAME=
GIT_USER_EMAIL=
+342 -50
View File
@@ -18,6 +18,14 @@ name: Publish Docker Image
# 5. build-variant multi-arch push of latest + vX.Y.Z tags.
# 6. promote-base-latest re-tag base-<hash> → base-latest with `crane copy`.
# 7. update-description patch Docker Hub description.
#
# Note the trigger: `push: tags: v*` (plus workflow_dispatch). Nothing here runs
# on a push to main, so a smoke assertion added outside a release is UNVALIDATED
# until the next tag — which is exactly how v1.8.0 shipped a broken assertion
# written three days earlier (it asserted a literal /home/developer stage path,
# while `run` executes `docker run --entrypoint=""` as root with HOME=/root).
# The `smoke_only` dispatch input exists to close that gap: it runs steps 1-4
# against HEAD and stops before anything is published.
on:
push:
@@ -33,11 +41,25 @@ on:
description: 'Update latest aliases (default true for tag-push, false for manual test runs)'
required: false
default: 'false'
smoke_only:
description: 'Build base + run both smoke jobs against HEAD, then stop. Publishes nothing. Use to validate smoke assertions without cutting a tag.'
required: false
default: 'false'
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: false
# Gitea Actions' default step shell is `sh -e {0}` (dash), which rejects
# bash-only syntax like `set -o pipefail`, `[[ ]]`, and arrays. Setting the
# default to bash workflow-wide eliminates the whole class of "forgot
# `shell: bash` on this step" bugs (hit twice: ed49b8d resolve-versions,
# b7197e8/b33e9dc promote-base-latest). All existing dash steps use only
# POSIX syntax, so bash (a superset) runs them unchanged.
defaults:
run:
shell: bash
env:
BUILDKIT_PROGRESS: plain
IMAGE: ${{ vars.DOCKERHUB_USERNAME }}/pi-devbox
@@ -58,6 +80,9 @@ jobs:
- name: Checkout
uses: actions/checkout@v4
- name: Guard — base *_REF args must be folded into the base hash
run: bash scripts/check-base-hash.sh
- name: Compute base tag from Dockerfile.base + dependencies
id: compute
run: |
@@ -117,66 +142,255 @@ jobs:
image: catthehacker/ubuntu:act-latest
outputs:
pi_version: ${{ steps.resolve.outputs.pi_version }}
mempalace_version: ${{ steps.resolve.outputs.mempalace_version }}
fork_ref: ${{ steps.resolve.outputs.fork_ref }}
obsmem_ref: ${{ steps.resolve.outputs.obsmem_ref }}
toolkit_ref: ${{ steps.resolve.outputs.toolkit_ref }}
extensions_ref: ${{ steps.resolve.outputs.extensions_ref }}
studio_ref: ${{ steps.resolve.outputs.studio_ref }}
studio_tag: ${{ steps.resolve.outputs.studio_tag }}
atelier_ref: ${{ steps.resolve.outputs.atelier_ref }}
atelier_tag: ${{ steps.resolve.outputs.atelier_tag }}
mempalace_toolkit_ref: ${{ steps.resolve.outputs.mempalace_toolkit_ref }}
steps:
# Needed since v1.7.0: the pi version and the pi-atelier tag are now
# PINNED IN Dockerfile.variant and read from it here, so this job has to
# see the repo. Keeping the pins in the Dockerfile (rather than duplicated
# in this workflow) means a local `docker build` and CI ship the same
# versions by construction, and a bump is one reviewable line.
- uses: actions/checkout@v4
- name: Resolve pi version + companion refs
id: resolve
shell: bash
run: |
set -eu
# Query npm registry directly; catthehacker/ubuntu:act-latest's npm
# is not reliably on PATH in act_runner job containers.
PI_VERSION=$(curl -sf "https://registry.npmjs.org/@earendil-works%2Fpi-coding-agent/latest" | jq -r '.version')
set -euo pipefail
AUTH_HEADER="Authorization: token ${GITEA_BUILD_TOKEN:-${GITHUB_TOKEN:-}}"
# Fail loud rather than silently shipping a floating branch. A
# transient network/API failure must ABORT the release, not bake
# an unpinned ref that defeats both cache-busting AND after-the-
# fact reproducibility. (Previously each lookup fell back to
# `main`/`master` via `|| echo`.)
require_sha() { # $1=label $2=value
if ! printf '%s' "${2:-}" | grep -qiE '^[0-9a-f]{40}$'; then
echo "::error::Could not resolve $1 to a commit SHA (got '${2:-<empty>}'). Refusing to fall back to a floating ref — published images must stay reproducible. Check connectivity and GITEA_BUILD_TOKEN/GITHUB_TOKEN."
exit 1
fi
}
# Read a commit SHA from Gitea, surviving a bad build token.
#
# These repos are public (see the note at the call sites), so auth is
# a convenience, not a requirement — but Gitea REJECTS an invalid
# token (401) rather than ignoring it, so a revoked or malformed
# GITEA_BUILD_TOKEN could fail an entire release on reads that work
# fine anonymously. An ABSENT secret was always safe (Gitea ignores an
# empty `token ` value and serves the request, 200); a STALE one was
# not. So: try authed, and on 401/403 retry anonymously.
#
# A non-200 after that emits nothing and returns 0 deliberately, so
# require_sha raises the loud explicit abort rather than this helper
# inventing a fallback ref.
#
# Messages go to STDERR, not as ::warning:: annotations: this
# function's stdout IS the SHA, so anything written there would be
# captured into the ref by the command substitution.
gitea_sha() { # $1=repo
local repo="$1" url resp code
url="https://gitea.jordbo.se/api/v1/repos/joakimp/${repo}/commits?limit=1&sha=main"
resp=$(curl -s -w '\n%{http_code}' -H "$AUTH_HEADER" "$url" || printf '\n000')
code=${resp##*$'\n'}
if [ "$code" = "401" ] || [ "$code" = "403" ]; then
printf 'WARNING: Gitea rejected the build token for %s (HTTP %s); retrying anonymously. The read should succeed (public repo), but GITEA_BUILD_TOKEN is stale or malformed and should be rotated.\n' "$repo" "$code" >&2
resp=$(curl -s -w '\n%{http_code}' "$url" || printf '\n000')
code=${resp##*$'\n'}
fi
if [ "$code" != "200" ]; then
printf 'WARNING: Gitea commit lookup for %s returned HTTP %s\n' "$repo" "$code" >&2
return 0
fi
printf '%s' "${resp%$'\n'*}" | jq -r '.[0].sha // empty' 2>/dev/null || true
}
# ── pi version: from the PIN, not from npm `latest` ───────────
# Until v1.7.0 this followed npm `latest`, which meant every release
# silently adopted whatever pi had shipped that morning — unaudited —
# in the same build that then got tagged and published. A pi minor
# can move the TUI/renderer internals that pi-atelier wraps (0.84 vs
# atelier 0.6.0: startup hang, sustained CPU) or the session `.jsonl`
# format that pi-session-repair parses. The pin makes adoption an
# explicit, reviewable act; the drift warning below makes it a
# prompt rather than a surprise.
PI_VERSION=$(sed -n 's/^ARG PI_VERSION=\([^[:space:]]*\).*/\1/p' Dockerfile.variant | head -n1)
if ! printf '%s' "${PI_VERSION:-}" | grep -qE '^[0-9]+\.[0-9]+\.[0-9]+$'; then
echo "::error::ARG PI_VERSION in Dockerfile.variant is not a concrete version (got '${PI_VERSION:-<empty>}'). CI refuses to build from a floating pi version — see the pin policy comment above that ARG."
exit 1
fi
# The pin must actually exist on npm: catches a typo, an unpublished
# version, or one yanked after we audited it — at resolve time, with
# a clear message, instead of as an `npm install` failure mid-build.
PI_PUBLISHED=$(curl -sf "https://registry.npmjs.org/@earendil-works%2Fpi-coding-agent/${PI_VERSION}" | jq -r '.version // empty' 2>/dev/null || true)
if [ "${PI_PUBLISHED:-}" != "${PI_VERSION}" ]; then
echo "::error::Pinned pi version ${PI_VERSION} is not published on npm (registry returned '${PI_PUBLISHED:-<empty>}'). Fix ARG PI_VERSION in Dockerfile.variant."
exit 1
fi
# Informational only — a newer pi must never be adopted implicitly.
# `|| true`: a transient registry failure must not fail a release
# whose version is already pinned and verified above.
PI_NPM_LATEST=$(curl -sf "https://registry.npmjs.org/@earendil-works%2Fpi-coding-agent/latest" | jq -r '.version // empty' 2>/dev/null || true)
if [ -n "${PI_NPM_LATEST:-}" ] && [ "${PI_NPM_LATEST}" != "${PI_VERSION}" ]; then
echo "::warning::pi ${PI_NPM_LATEST} is published; this build ships the audited pin ${PI_VERSION}. To adopt it: read the upstream CHANGELOG for every version in between (TUI/theme API, session .jsonl format, extension loader, Node engine), re-check pi-atelier's floor, then bump ARG PI_VERSION in Dockerfile.variant and note the audit in CHANGELOG.md."
fi
echo "pi_version=${PI_VERSION}" >> "$GITHUB_OUTPUT"
# Resolve pi-fork / pi-observational-memory git refs to commit
# SHAs so the build-arg string changes whenever upstream moves.
# ── mempalace core: same audit as pi, from Dockerfile.base ────
# Until now this pin had NO CI-side audit at all — a literal string
# in Dockerfile.base with zero references in this workflow, while
# PI_VERSION got a concreteness gate, a published-on-registry check
# and a drift warning. It is the same class of risk: the palace's MCP
# tool schema is the agent-facing contract, and a client/server skew
# against the shared central palace is a fleet-wide, not local,
# problem. Read from Dockerfile.base (not duplicated here) so a local
# `docker build` and CI install the same version by construction.
MEMPALACE_VERSION=$(sed -n 's/^ARG MEMPALACE_VERSION=\([^[:space:]]*\).*/\1/p' Dockerfile.base | head -n1)
if ! printf '%s' "${MEMPALACE_VERSION:-}" | grep -qE '^[0-9]+\.[0-9]+\.[0-9]+$'; then
echo "::error::ARG MEMPALACE_VERSION in Dockerfile.base is not a concrete version (got '${MEMPALACE_VERSION:-<empty>}'). CI refuses to build from a floating palace version — see the pin policy comment above that ARG."
exit 1
fi
# One fetch, two gates. `curl -sf` exits non-zero and prints nothing
# on 404 (PyPI's answer for an unpublished version), so an empty body
# lands in the "not published" branch with its own message.
MEMPALACE_PYPI=$(curl -sf "https://pypi.org/pypi/mempalace/${MEMPALACE_VERSION}/json" || true)
MEMPALACE_PUBLISHED=$(printf '%s' "$MEMPALACE_PYPI" | jq -r '.info.version // empty' 2>/dev/null || true)
if [ "${MEMPALACE_PUBLISHED:-}" != "${MEMPALACE_VERSION}" ]; then
echo "::error::Pinned mempalace version ${MEMPALACE_VERSION} is not published on PyPI (registry returned '${MEMPALACE_PUBLISHED:-<empty>}'). Fix ARG MEMPALACE_VERSION in Dockerfile.base."
exit 1
fi
# A yanked release still installs when pinned exactly (PEP 592), so
# `uv tool install mempalace==X` would succeed silently and ship a
# version upstream has withdrawn to the whole fleet. The escape hatch
# is the same one-line bump that got us here.
MEMPALACE_YANKED=$(printf '%s' "$MEMPALACE_PYPI" | jq -r '.info.yanked // false' 2>/dev/null || true)
if [ "${MEMPALACE_YANKED:-false}" = "true" ]; then
# Reason hoisted into its own variable rather than inlined as a
# $(...) inside the message: a jq program nested in a substitution
# inside a double-quoted string needs escaping that silently breaks
# the FILTER (jq compile error) while the surrounding `exit 1` still
# fires, so the gate looks correct and reports garbage. Caught by
# the mutation test, not by review.
MEMPALACE_YANK_REASON=$(printf '%s' "$MEMPALACE_PYPI" | jq -r '.info.yanked_reason // "no reason given"' 2>/dev/null || true)
echo "::error::Pinned mempalace version ${MEMPALACE_VERSION} is YANKED on PyPI (${MEMPALACE_YANK_REASON:-no reason given}). An exact pin installs a yanked release without complaint — bump ARG MEMPALACE_VERSION in Dockerfile.base."
exit 1
fi
# Informational only, exactly like pi's npm drift warning: a newer
# palace must never be adopted implicitly. `|| true` so a transient
# PyPI failure cannot fail a release whose pin is already verified.
MEMPALACE_PYPI_LATEST=$(curl -sf "https://pypi.org/pypi/mempalace/json" | jq -r '.info.version // empty' 2>/dev/null || true)
if [ -n "${MEMPALACE_PYPI_LATEST:-}" ] && [ "${MEMPALACE_PYPI_LATEST}" != "${MEMPALACE_VERSION}" ]; then
echo "::warning::mempalace ${MEMPALACE_PYPI_LATEST} is published; this build ships the audited pin ${MEMPALACE_VERSION}. To adopt it: read the upstream CHANGELOG for MCP tool-schema changes (the agent-facing contract) and for sync/delete semantics, check the skew it introduces against the central palace host's server version, then bump ARG MEMPALACE_VERSION in Dockerfile.base and note the audit in CHANGELOG.md."
fi
echo "mempalace_version=${MEMPALACE_VERSION}" >> "$GITHUB_OUTPUT"
# pi-fork / pi-observational-memory (GitHub) → commit SHAs.
FORK_REF=$(curl -sf -H "Accept: application/vnd.github.sha" \
"https://api.github.com/repos/elpapi42/pi-fork/commits/master" || echo "master")
"https://api.github.com/repos/elpapi42/pi-fork/commits/master" || true)
require_sha PI_FORK_REF "$FORK_REF"
OBSMEM_REF=$(curl -sf -H "Accept: application/vnd.github.sha" \
"https://api.github.com/repos/elpapi42/pi-observational-memory/commits/master" || echo "master")
[ -n "$FORK_REF" ] || FORK_REF=master
[ -n "$OBSMEM_REF" ] || OBSMEM_REF=master
"https://api.github.com/repos/elpapi42/pi-observational-memory/commits/master" || true)
require_sha PI_OBSMEM_REF "$OBSMEM_REF"
echo "fork_ref=${FORK_REF}" >> "$GITHUB_OUTPUT"
echo "obsmem_ref=${OBSMEM_REF}" >> "$GITHUB_OUTPUT"
# Also resolve pi-toolkit / pi-extensions main HEADs to SHAs so a
# workflow_dispatch re-run produces byte-identical images when
# those repos haven't moved (and a clean diff in build-arg strings
# when they have, defeating the registry buildcache footgun).
# Gitea API requires auth even for public-repo commit listing.
TOOLKIT_REF=$(curl -sf -H "Authorization: token ${GITEA_BUILD_TOKEN:-${GITHUB_TOKEN:-}}" \
"https://gitea.jordbo.se/api/v1/repos/joakimp/pi-toolkit/commits?limit=1&sha=main" \
| jq -r '.[0].sha // "main"' 2>/dev/null || echo "main")
EXTENSIONS_REF=$(curl -sf -H "Authorization: token ${GITEA_BUILD_TOKEN:-${GITHUB_TOKEN:-}}" \
"https://gitea.jordbo.se/api/v1/repos/joakimp/pi-extensions/commits?limit=1&sha=main" \
| jq -r '.[0].sha // "main"' 2>/dev/null || echo "main")
[ -n "$TOOLKIT_REF" ] || TOOLKIT_REF=main
[ -n "$EXTENSIONS_REF" ] || EXTENSIONS_REF=main
# pi-atelier → the PINNED TAG's commit SHA. Unlike fork/obsmem
# (which track a branch head) atelier wraps pi's private TUI
# renderer, so its version is pinned in Dockerfile.variant and read
# from there; we only resolve tag → SHA, for reproducibility and to
# defeat the cache-hit footgun. Never floats to a branch.
ATELIER_TAG=$(sed -n 's/^ARG PI_ATELIER_REF=\([^[:space:]]*\).*/\1/p' Dockerfile.variant | head -n1)
if ! printf '%s' "${ATELIER_TAG:-}" | grep -qE '^v?[0-9]+\.[0-9]+\.[0-9]+$'; then
echo "::error::ARG PI_ATELIER_REF in Dockerfile.variant is not a semver tag (got '${ATELIER_TAG:-<empty>}'). pi-atelier must stay pinned to a tag — see the floor note above that ARG."
exit 1
fi
ATELIER_LS=$(git ls-remote --tags "https://github.com/michaelmjhhhh/pi-atelier.git" || true)
# Peeled ^{} line first (annotated tags), then the direct ref.
ATELIER_REF=$(printf '%s\n' "$ATELIER_LS" | awk -v t="refs/tags/${ATELIER_TAG}^{}" '$2==t{print $1}')
if [ -z "$ATELIER_REF" ]; then
ATELIER_REF=$(printf '%s\n' "$ATELIER_LS" | awk -v t="refs/tags/${ATELIER_TAG}" '$2==t{print $1}')
fi
require_sha PI_ATELIER_REF "$ATELIER_REF"
echo "atelier_ref=${ATELIER_REF}" >> "$GITHUB_OUTPUT"
echo "atelier_tag=${ATELIER_TAG}" >> "$GITHUB_OUTPUT"
# pi-toolkit / pi-extensions (Gitea) → commit SHAs. All three Gitea
# repos read in this step are PUBLIC: an unauthenticated GET of these
# commit endpoints returns 200 with the IDENTICAL sha (verified
# 2026-08-15 for pi-toolkit, pi-extensions and mempalace-toolkit).
# The comment that used to sit here claimed the Gitea API "requires
# auth even for public-repo commit listing" — it does not. Only
# /api/v1/repos/*/actions/* refuses anonymous reads (401), which is
# what that claim was almost certainly generalised from.
#
# The header is still passed on purpose: it keeps working if a repo is
# ever flipped private, and an ABSENT secret degrades cleanly, because
# Gitea ignores an empty `token ` value and serves the request
# anonymously (200). The real hazard is the opposite one — a REVOKED or
# malformed token returns 401 where anonymous would have returned 200,
# so a stale GITEA_BUILD_TOKEN turns a healthy public read into a
# require_sha failure that reads like an API or network fault. If this
# step ever fails on a repo you can browse anonymously, suspect the
# token before you suspect Gitea.
TOOLKIT_REF=$(gitea_sha pi-toolkit)
require_sha PI_TOOLKIT_REF "$TOOLKIT_REF"
EXTENSIONS_REF=$(gitea_sha pi-extensions)
require_sha PI_EXTENSIONS_REF "$EXTENSIONS_REF"
echo "toolkit_ref=${TOOLKIT_REF}" >> "$GITHUB_OUTPUT"
echo "extensions_ref=${EXTENSIONS_REF}" >> "$GITHUB_OUTPUT"
# Resolve mempalace-toolkit main HEAD to a SHA. UNLIKE the others,
# mempalace-toolkit is cloned in Dockerfile.base, so this SHA is
# ALSO folded into the base-decide hash to force a base rebuild
# when the toolkit moves (without it, a toolkit-only fix silently
# fails to land unless Dockerfile.base itself changes).
MEMPALACE_TOOLKIT_REF=$(curl -sf -H "Authorization: token ${GITEA_BUILD_TOKEN:-${GITHUB_TOKEN:-}}" \
"https://gitea.jordbo.se/api/v1/repos/joakimp/mempalace-toolkit/commits?limit=1&sha=main" \
| jq -r '.[0].sha // "main"' 2>/dev/null || echo "main")
[ -n "$MEMPALACE_TOOLKIT_REF" ] || MEMPALACE_TOOLKIT_REF=main
# mempalace-toolkit (Gitea) → commit SHA. UNLIKE the others this
# is cloned in Dockerfile.base, so the SAME SHA is ALSO folded
# into the base-decide hash (see that job) to force a base rebuild
# when the toolkit moves — otherwise a toolkit-only fix silently
# fails to land unless Dockerfile.base itself changes.
MEMPALACE_TOOLKIT_REF=$(gitea_sha mempalace-toolkit)
require_sha MEMPALACE_TOOLKIT_REF "$MEMPALACE_TOOLKIT_REF"
echo "mempalace_toolkit_ref=${MEMPALACE_TOOLKIT_REF}" >> "$GITHUB_OUTPUT"
# Resolve pi-studio (omaclaren/pi-studio) main HEAD to a SHA for
# the :latest-studio variant — same cache-busting rationale.
STUDIO_REF=$(curl -sf -H "Accept: application/vnd.github.sha" \
"https://api.github.com/repos/omaclaren/pi-studio/commits/main" || echo "main")
[ -n "$STUDIO_REF" ] || STUDIO_REF=main
# pi-studio (omaclaren/pi-studio) → newest SEMVER TAG's commit SHA
# for the :*-studio images. Upstream stopped publishing GitHub
# *Releases* at v0.5.55 but keeps tagging every version (vX.Y.Z) and
# pushing to main, so pinning main HEAD risked baking half-finished
# commits that land after a tag. Take the newest stable tag instead.
# List ALL tags in one `git ls-remote` call — the REST tags API
# paginates at 100 and this repo already has >140 tags, so page 1 is
# NOT guaranteed to hold the newest — pick the highest X.Y.Z with
# `sort -V` (pre-releases like -rc1 excluded by the strict filter),
# then resolve its commit SHA (a SHA, not a moving tag, preserves
# cache-busting + reproducibility and is what require_sha demands).
STUDIO_TAGS=$(git ls-remote --tags "https://github.com/omaclaren/pi-studio.git" || true)
STUDIO_TAG=$(printf '%s\n' "$STUDIO_TAGS" | awk '{print $2}' \
| sed -n 's#^refs/tags/##p' \
| grep -E '^v?[0-9]+\.[0-9]+\.[0-9]+$' \
| sort -V | tail -n1 || true)
if [ -z "${STUDIO_TAG:-}" ]; then
echo "::error::Could not resolve a pi-studio semver tag (git ls-remote empty/unreachable). Refusing to fall back to a floating ref."
exit 1
fi
# Prefer the peeled ^{} line (annotated tags); fall back to the
# direct ref (lightweight tags, which pi-studio currently uses).
STUDIO_REF=$(printf '%s\n' "$STUDIO_TAGS" | awk -v t="refs/tags/${STUDIO_TAG}^{}" '$2==t{print $1}')
if [ -z "$STUDIO_REF" ]; then
STUDIO_REF=$(printf '%s\n' "$STUDIO_TAGS" | awk -v t="refs/tags/${STUDIO_TAG}" '$2==t{print $1}')
fi
require_sha PI_STUDIO_REF "$STUDIO_REF"
echo "studio_ref=${STUDIO_REF}" >> "$GITHUB_OUTPUT"
echo "Resolved PI_VERSION=${PI_VERSION}"
echo "studio_tag=${STUDIO_TAG}" >> "$GITHUB_OUTPUT"
echo "Resolved PI_VERSION=${PI_VERSION} (pinned in Dockerfile.variant; npm latest is ${PI_NPM_LATEST:-unknown})"
echo "Resolved MEMPALACE_VERSION=${MEMPALACE_VERSION} (pinned in Dockerfile.base; PyPI latest is ${MEMPALACE_PYPI_LATEST:-unknown})"
echo "Resolved PI_ATELIER_REF=${ATELIER_REF} (pi-atelier ${ATELIER_TAG}, pinned)"
echo "Resolved PI_FORK_REF=${FORK_REF}, PI_OBSMEM_REF=${OBSMEM_REF}"
echo "Resolved PI_TOOLKIT_REF=${TOOLKIT_REF}, PI_EXTENSIONS_REF=${EXTENSIONS_REF}"
echo "Resolved PI_STUDIO_REF=${STUDIO_REF}"
echo "Resolved PI_STUDIO_REF=${STUDIO_REF} (pi-studio ${STUDIO_TAG})"
echo "Resolved MEMPALACE_TOOLKIT_REF=${MEMPALACE_TOOLKIT_REF}"
# ── Phase 2: build & push base (multi-arch), only when needed ──────
@@ -299,9 +513,15 @@ jobs:
PI_OBSMEM_REF=${{ needs.resolve-versions.outputs.obsmem_ref }}
PI_TOOLKIT_REF=${{ needs.resolve-versions.outputs.toolkit_ref }}
PI_EXTENSIONS_REF=${{ needs.resolve-versions.outputs.extensions_ref }}
MEMPALACE_TOOLKIT_REF=${{ needs.resolve-versions.outputs.mempalace_toolkit_ref }}
PI_ATELIER_REF=${{ needs.resolve-versions.outputs.atelier_ref }}
PI_ATELIER_VERSION=${{ needs.resolve-versions.outputs.atelier_tag }}
RELEASE_TAG=smoke
SOURCE_REVISION=${{ github.sha }}
- name: Smoke test (amd64)
env:
EXPECTED_PI_VERSION: ${{ needs.resolve-versions.outputs.pi_version }}
EXPECTED_MEMPALACE_VERSION: ${{ needs.resolve-versions.outputs.mempalace_version }}
run: bash scripts/smoke-test.sh pi-devbox:smoke
# ── Phase 3b: amd64 smoke for the studio variant ────────────────────
@@ -355,14 +575,29 @@ jobs:
PI_EXTENSIONS_REF=${{ needs.resolve-versions.outputs.extensions_ref }}
INSTALL_STUDIO=true
PI_STUDIO_REF=${{ needs.resolve-versions.outputs.studio_ref }}
PI_STUDIO_VERSION=${{ needs.resolve-versions.outputs.studio_tag }}
MEMPALACE_TOOLKIT_REF=${{ needs.resolve-versions.outputs.mempalace_toolkit_ref }}
PI_ATELIER_REF=${{ needs.resolve-versions.outputs.atelier_ref }}
PI_ATELIER_VERSION=${{ needs.resolve-versions.outputs.atelier_tag }}
RELEASE_TAG=smoke-studio
SOURCE_REVISION=${{ github.sha }}
- name: Smoke test studio (amd64)
env:
EXPECTED_PI_VERSION: ${{ needs.resolve-versions.outputs.pi_version }}
EXPECTED_MEMPALACE_VERSION: ${{ needs.resolve-versions.outputs.mempalace_version }}
run: bash scripts/smoke-test.sh pi-devbox:smoke-studio
# ── Phase 4: multi-arch publish ─────────────────────────────────────
build-variant:
needs: [base-decide, smoke, resolve-versions]
# A `smoke_only` dispatch stops the pipeline here: base is probed/built and
# both smoke jobs run, but nothing is published. Deliberately NOT wrapped in
# always() — specifying `if:` keeps the implicit "all needs succeeded" gate,
# so a failing smoke still blocks the release. On a tag push `inputs` is
# unset, and `null != 'true'` is true, so releases are unaffected.
# promote-base-latest and update-description need build-variant to have
# succeeded, so they skip on their own — no extra guard required.
if: inputs.smoke_only != 'true'
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
@@ -406,10 +641,14 @@ jobs:
OBSMEM_REF: ${{ needs.resolve-versions.outputs.obsmem_ref }}
TOOLKIT_REF: ${{ needs.resolve-versions.outputs.toolkit_ref }}
EXTENSIONS_REF: ${{ needs.resolve-versions.outputs.extensions_ref }}
MEMPALACE_TOOLKIT_REF: ${{ needs.resolve-versions.outputs.mempalace_toolkit_ref }}
ATELIER_REF: ${{ needs.resolve-versions.outputs.atelier_ref }}
ATELIER_TAG: ${{ needs.resolve-versions.outputs.atelier_tag }}
run: |
set -euo pipefail
TAG_FLAGS=()
while IFS= read -r t; do [[ -n "$t" ]] && TAG_FLAGS+=( -t "$t" ); done <<< "${TAGS}"
BUILD_DATE=$(date -u +%Y-%m-%dT%H:%M:%SZ)
# 3-attempt retry (see build-base step for rationale).
for attempt in 1 2 3; do
echo "==> Build+push attempt ${attempt}/3"
@@ -423,6 +662,14 @@ jobs:
--build-arg "PI_OBSMEM_REF=${OBSMEM_REF}" \
--build-arg "PI_TOOLKIT_REF=${TOOLKIT_REF}" \
--build-arg "PI_EXTENSIONS_REF=${EXTENSIONS_REF}" \
--build-arg "MEMPALACE_TOOLKIT_REF=${MEMPALACE_TOOLKIT_REF}" \
--build-arg "PI_ATELIER_REF=${ATELIER_REF}" \
--build-arg "PI_ATELIER_VERSION=${ATELIER_TAG}" \
--build-arg "IMAGE_TITLE=pi-devbox" \
--build-arg "IMAGE_DESCRIPTION=pi-devbox ${RELEASE_TAG} — core variant: pi coding agent CLI ${PI_VERSION}, pi-toolkit, extensions (fork + observational-memory + atelier ${ATELIER_TAG} TUI sidebar), MemPalace. No browser UI — see the -studio tags for that." \
--build-arg "RELEASE_TAG=${RELEASE_TAG}" \
--build-arg "BUILD_DATE=${BUILD_DATE}" \
--build-arg "SOURCE_REVISION=${GITHUB_SHA:-}" \
"${TAG_FLAGS[@]}" \
.; then
echo "==> Attempt ${attempt} succeeded"
@@ -443,6 +690,7 @@ jobs:
# or fail independently of the core release.
build-variant-studio:
needs: [base-decide, smoke-studio, resolve-versions]
if: inputs.smoke_only != 'true'
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
@@ -487,10 +735,15 @@ jobs:
TOOLKIT_REF: ${{ needs.resolve-versions.outputs.toolkit_ref }}
EXTENSIONS_REF: ${{ needs.resolve-versions.outputs.extensions_ref }}
STUDIO_REF: ${{ needs.resolve-versions.outputs.studio_ref }}
STUDIO_TAG: ${{ needs.resolve-versions.outputs.studio_tag }}
MEMPALACE_TOOLKIT_REF: ${{ needs.resolve-versions.outputs.mempalace_toolkit_ref }}
ATELIER_REF: ${{ needs.resolve-versions.outputs.atelier_ref }}
ATELIER_TAG: ${{ needs.resolve-versions.outputs.atelier_tag }}
run: |
set -euo pipefail
TAG_FLAGS=()
while IFS= read -r t; do [[ -n "$t" ]] && TAG_FLAGS+=( -t "$t" ); done <<< "${TAGS}"
BUILD_DATE=$(date -u +%Y-%m-%dT%H:%M:%SZ)
# 3-attempt retry (see build-base step for rationale).
for attempt in 1 2 3; do
echo "==> Build+push attempt ${attempt}/3"
@@ -504,8 +757,17 @@ jobs:
--build-arg "PI_OBSMEM_REF=${OBSMEM_REF}" \
--build-arg "PI_TOOLKIT_REF=${TOOLKIT_REF}" \
--build-arg "PI_EXTENSIONS_REF=${EXTENSIONS_REF}" \
--build-arg "MEMPALACE_TOOLKIT_REF=${MEMPALACE_TOOLKIT_REF}" \
--build-arg "INSTALL_STUDIO=true" \
--build-arg "IMAGE_TITLE=pi-devbox (studio)" \
--build-arg "PI_ATELIER_REF=${ATELIER_REF}" \
--build-arg "PI_ATELIER_VERSION=${ATELIER_TAG}" \
--build-arg "IMAGE_DESCRIPTION=pi-devbox ${RELEASE_TAG} — studio variant: everything in the core variant (pi ${PI_VERSION}, pi-toolkit, fork + observational-memory + atelier ${ATELIER_TAG}, MemPalace) plus the pi-studio browser UI ${STUDIO_TAG}." \
--build-arg "PI_STUDIO_REF=${STUDIO_REF}" \
--build-arg "PI_STUDIO_VERSION=${STUDIO_TAG}" \
--build-arg "RELEASE_TAG=${RELEASE_TAG}" \
--build-arg "BUILD_DATE=${BUILD_DATE}" \
--build-arg "SOURCE_REVISION=${GITHUB_SHA:-}" \
"${TAG_FLAGS[@]}" \
.; then
echo "==> Attempt ${attempt} succeeded"
@@ -525,16 +787,19 @@ jobs:
needs:
- base-decide
- build-variant
# Skip on cache-hit base builds: when need_build=false, base-latest
# already points at the same digest as base-<hash>, so the retag is
# a tautology and any transient failure of it is purely cosmetic.
# Manual workflow_dispatch with promote_latest=true overrides this
# gate as an escape hatch (e.g., if base-latest got hand-deleted).
# Run on every tag release (and on promote_latest=true dispatches).
# The job-level gate deliberately does NOT key off need_build anymore:
# the actual no-op optimization moved INTO the step as a digest compare
# (see below). Keying the gate on need_build was wrong because a prior
# dry-run dispatch (promote_latest=false) can pre-build+push base-<hash>,
# making need_build=false on the subsequent tag run even though
# base-latest is still stale — the old gate then skipped promotion and
# left base-latest pointing at the PREVIOUS base. (Observed 2026-06-27,
# v1.2.3: dry-run-first release left base-latest one base behind.)
if: |
always() &&
needs.build-variant.result == 'success' &&
(inputs.promote_latest == 'true' ||
(github.ref_type == 'tag' && needs.base-decide.outputs.need_build == 'true'))
(inputs.promote_latest == 'true' || github.ref_type == 'tag')
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
@@ -556,11 +821,38 @@ jobs:
crane auth login docker.io \
-u ${{ vars.DOCKERHUB_USERNAME }} \
-p "${{ secrets.DOCKERHUB_TOKEN }}"
- name: Re-tag base-<hash> as base-latest
- name: Re-tag base-<hash> as base-latest (only if stale)
# shell: bash is REQUIRED — Gitea Actions' default step shell is
# `sh -e {0}` (dash), which rejects `set -o pipefail` with
# "Illegal option -o pipefail" and aborts the step before the
# crane digest-compare runs, leaving base-latest un-promoted.
# Same footgun as ed49b8d (resolve-versions). Regression shipped
# in b7197e8, caught on the v1.2.4 release (run 418).
shell: bash
env:
BASE_HASH_REF: ${{ env.IMAGE }}:${{ needs.base-decide.outputs.base_tag }}
BASE_LATEST_REF: ${{ env.IMAGE }}:base-latest
run: |
crane copy \
${{ env.IMAGE }}:${{ needs.base-decide.outputs.base_tag }} \
${{ env.IMAGE }}:base-latest
set -euo pipefail
# Correctness invariant: after a release, base-latest must resolve to
# the SAME digest as the base-<hash> the just-built variants were
# FROM. Compare digests rather than trusting need_build — a prior
# dry-run dispatch can pre-build base-<hash>, so need_build=false on
# the tag run does NOT imply base-latest is already current. When the
# digests already match (genuine cache-hit release) this is a no-op,
# so we skip the crane copy entirely — preserving the original
# "don't do a tautological retag" intent and avoiding any cosmetic
# transient-failure exposure on releases that change nothing.
want=$(crane digest "${BASE_HASH_REF}")
have=$(crane digest "${BASE_LATEST_REF}" 2>/dev/null || echo "")
echo "base-<hash> digest: ${want}"
echo "base-latest digest: ${have:-<absent>}"
if [ "${want}" = "${have}" ]; then
echo "base-latest already current; nothing to promote."
else
echo "Promoting base-latest -> ${BASE_HASH_REF}"
crane copy "${BASE_HASH_REF}" "${BASE_LATEST_REF}"
fi
# ── Phase 6: update Hub description (only on real release runs) ────
update-description:
+159
View File
@@ -0,0 +1,159 @@
name: Lint
# Durable guard against CI-workflow bugs — most importantly the recurring
# "bash-only syntax under the default `sh`/dash shell" footgun that broke
# resolve-versions (ed49b8d) and promote-base-latest (b7197e8 → run 418).
# actionlint runs shellcheck against each `run:` step using its *effective*
# shell, so `set -o pipefail` under dash is flagged as SC3040 before any
# expensive build runs. This is cheap (~10s) and independent of the build
# pipeline, so it fires on every branch push/PR — not just on release tags,
# which is where the build workflow (docker-publish.yml) is otherwise only
# triggered.
#
# `branches: ['**']` (rather than a bare `push:`) deliberately EXCLUDES tag
# pushes. A bare `push:` also fires on `refs/tags/v*`, which was duplicate work —
# the tagged tree was already linted when the same commit was pushed to main
# (v1.6.4: lint id=529 on refs/heads/main, then id=531 again on
# refs/tags/v1.6.4, same sha e86e5df). The wasted compute is small (measured:
# lint here runs 0.3-0.9 min, against a 77.6 min release build for v1.6.4 — so
# runner contention is NOT a real argument in this repo, unlike opencode-devbox
# where actionlint installs shellcheck and takes 6-15 min). The substantive
# reason is discovery ambiguity: the runs listing is newest-first, so the
# tag-ref lint run sorts ABOVE the publish run, and "first run matching
# refs/tags/<tag>" picks lint — which goes green in under a minute while the
# image is still building, making a release look finished before anything is
# published. See AGENTS.md "Gitea API access" for the head_sha-filtered
# discovery pattern.
on:
push:
branches:
- '**'
pull_request:
workflow_dispatch:
concurrency:
group: lint-${{ github.ref }}
cancel-in-progress: true
defaults:
run:
shell: bash
jobs:
actionlint:
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
steps:
- uses: actions/checkout@v4
- name: Install shellcheck
run: |
apt-get update
apt-get install -y --no-install-recommends shellcheck python3-yaml
- name: "Shellcheck + syntax-check repository scripts (severity: error)"
# Gap being closed: everything else in this job shellchecks workflow
# `run:` steps ONLY, via actionlint. The repo's own shell scripts —
# entrypoint.sh, scripts/*.sh, and the extensionless tools under
# rootfs/usr/local/bin/ — have never been shellchecked. That exact gap
# (a sibling repo with no shell-script lint at all) is how a defect
# shipped invisibly for two months: `echo "$json" | python3 <<'EOF'
# ... json.load(sys.stdin)` cannot work — with no script argument
# python reads its SCRIPT from stdin, so the heredoc IS stdin and the
# json.load call hits EOF. shellcheck flags exactly this at severity
# ERROR (SC2259, "This redirection overrides piped input"); nothing
# ever ran it. Measured before adding this gate: `-S error` is 0
# findings across every shell file in THIS repo today, so it is free
# to add. `-S warning` is NOT free here (19x SC2088 tilde-in-quotes in
# scripts/recreate-sanity-check.sh, plus assorted SC2016 — both
# intentional), so warning-level would train people to ignore the job;
# hence error-only, matching the SHELLCHECK_OPTS philosophy below.
#
# Discovery is *.sh UNION a shebang scan, because rootfs/usr/local/
# bin/{pi-devbox-version,devbox-skill-reconcile,dot-watch,studio-expose}
# are shell scripts with no extension. -print0/mapfile -d '' so a path
# with a space cannot silently split, and the file count is asserted
# non-zero — a green tick over an empty file set is not a check.
run: |
# Union of two signals, because either alone misses a real case:
# a shebang scan misses a sourced fragment with no shebang, and a
# *.sh glob misses the extensionless tools in rootfs/usr/local/bin/.
# Silent skipping is precisely the failure mode this gate exists to
# prevent, so err toward over-collecting.
mapfile -d '' -t all_files < <(find . -not -path './.git/*' -type f -print0)
sh_files=()
for f in "${all_files[@]}"; do
case "$f" in *.sh) sh_files+=("$f"); continue;; esac
if head -n1 "$f" 2>/dev/null | grep -qE '^#!.*\b(bash|sh)\b'; then
sh_files+=("$f")
fi
done
echo "Checking ${#sh_files[@]} shell file(s)"
if [ "${#sh_files[@]}" -eq 0 ]; then
echo "::error::no shell files found — the shebang scan or the checkout is wrong"
exit 1
fi
shellcheck -S error -f gcc "${sh_files[@]}"
rc=0
for f in "${sh_files[@]}"; do
bash -n "$f" || { echo "::error file=$f::bash -n failed"; rc=1; }
done
exit "$rc"
- name: Gitea shell guard (catches the actionlint blind spot)
# actionlint models GitHub Actions, where the default run shell is
# bash, so it does NOT flag bash syntax in a step that merely OMITS
# `shell:` — which is exactly how ed49b8d and b7197e8 manifested on
# Gitea (default sh/dash). This guard enforces that every run: step
# resolves to bash under Gitea's real defaults. Run it BEFORE
# actionlint so the more precise diagnostic surfaces first.
run: bash scripts/check-workflow-shell.sh .gitea/workflows
- name: Install actionlint (pinned)
env:
ACTIONLINT_VERSION: 1.7.7
run: |
curl -fsSL \
"https://github.com/rhysd/actionlint/releases/download/v${ACTIONLINT_VERSION}/actionlint_${ACTIONLINT_VERSION}_linux_amd64.tar.gz" \
| tar -xz -C /usr/local/bin actionlint
actionlint --version
- name: Run actionlint
# SHELLCHECK_OPTS excludes pure-style codes (quoting/style opinions)
# so the guard stays focused on correctness bugs — crucially the
# SC3xxx "not POSIX / wrong shell" family that catches the pipefail
# footgun. Do NOT exclude SC3040 (set -o pipefail under sh) or any
# other SC3xxx code.
env:
SHELLCHECK_OPTS: "-e SC2086 -e SC2016 -e SC2129 -e SC2001 -e SC2312"
# Pass explicit paths: actionlint's no-arg mode auto-detects a
# project by looking for `.github/workflows`, which doesn't exist in
# this `.gitea/workflows` repo and hard-fails with exit 3
# ("no project was found"). Globbing the workflow files is the
# supported way to lint a non-GitHub layout.
run: actionlint -color .gitea/workflows/*.yml
hadolint:
# Lint the two Dockerfiles that ARE the project (the shell/actions linting
# above never looked at them). Config — ignored rules + failure threshold
# — lives in .hadolint.yaml, which hadolint reads automatically, so a local
# `hadolint Dockerfile.base` reproduces CI exactly.
runs-on: ubuntu-latest
container:
image: catthehacker/ubuntu:act-latest
steps:
- uses: actions/checkout@v4
- name: Install hadolint (pinned)
env:
HADOLINT_VERSION: 2.14.0
run: |
curl -fsSL \
"https://github.com/hadolint/hadolint/releases/download/v${HADOLINT_VERSION}/hadolint-Linux-x86_64" \
-o /usr/local/bin/hadolint
chmod +x /usr/local/bin/hadolint
hadolint --version
- name: Run hadolint
run: hadolint Dockerfile.base Dockerfile.variant
+27
View File
@@ -0,0 +1,27 @@
# hadolint configuration for pi-devbox.
#
# Both Dockerfiles are linted in CI (.gitea/workflows/lint.yml → `hadolint`
# job). hadolint reads this file automatically, so a local
# `hadolint Dockerfile.base` reproduces CI exactly.
#
# The ignores below are DELIBERATE project choices — they mirror the
# philosophy of the shellcheck excludes already applied to `run:` steps
# (SHELLCHECK_OPTS in lint.yml). Anything NOT listed here still fails the
# build at `warning` and above, so new Dockerfile smells are caught going
# forward.
ignored:
- DL3008 # "pin apt versions" — intentionally unpinned: the base tracks
# Debian stable and runs `apt-get upgrade`, so pinning point
# versions would rot and fight security updates.
- DL3016 # "pin npm versions" — pi's version IS pinned, but via the
# PI_VERSION build-arg (CI-resolved from npm), not the npm CLI.
- DL4006 # "set -o pipefail before a pipe" — the piped RUNs are
# download|extract steps with their own retries / `set -e`.
# Switching the global SHELL to bash is a larger, base-affecting
# change — tracked in IDEAS.md.
- DL3003 # "use WORKDIR, not cd" — cosmetic in the few `cd` RUNs here.
- SC2086 # "double-quote to prevent word-splitting" — the same code is
# excluded for shell `run:` steps in lint.yml; splitting is
# intentional in these contexts.
failure-threshold: warning
+189 -23
View File
@@ -14,15 +14,31 @@ re-brand of opencode-devbox's `pi-only` variant.
- `Dockerfile.variant` — `FROM base-<hash>`, adds pi + companions
(`pi-toolkit`, `pi-extensions`, `pi-fork`, `pi-observational-memory`)
and, when `INSTALL_STUDIO=true`, vendors `pi-studio` to `/opt/pi-studio`
(`-studio` variant).
(`-studio` variant). Also appends the pi-devbox managed block from
`pi-global-AGENTS.append.md` onto pi-toolkit's `pi-global-AGENTS.md` (the
single global instruction slot pi loads) so containers proactively load the
baked `pi-devbox-environment` skill. Idempotent via a marker grep. After the
pinned clones it also refreshes the vendored `pi-extensions` fallback skill
by copying `/opt/pi-extensions/skill/` over the committed `rootfs/` snapshot
(Option 1 over Option 2 — see `skills/VENDORED.md`).
- `entrypoint.sh` — UID/GID alignment as root, then drops to `developer`.
- `entrypoint-user.sh` — per-container start: SSH ControlMaster socket
dir, LAN-access setup, MemPalace init, pi-toolkit + pi-extensions
deploy, mempalace-bridge symlink, fork/recall + pi-studio pi-install,
optional `studio-expose` bridge (when `STUDIO_EXPOSE=1`), skillset
deploy.
- `entrypoint-user.sh` — per-container start: prints the `pi-devbox-version`
banner first (which build/commit is running, from the manifest below),
then SSH ControlMaster socket dir, LAN-access setup, MemPalace init,
pi-toolkit + pi-extensions deploy, mempalace-bridge symlink, fork/recall +
pi-studio pi-install, optional `studio-expose` bridge (when
`STUDIO_EXPOSE=1`), image-baked skills symlink-in, skillset deploy.
- `rootfs/` — files baked into the image (bash aliases, inputrc,
setup-lan-access.sh, `studio-expose` helper).
setup-lan-access.sh, `studio-expose` helper, `pi-devbox-version` — wraps
`/etc/pi-devbox/build-manifest.json` into a human-readable summary + live
drift check, see README “Build provenance”). Also
`usr/local/share/pi-devbox/skills/<name>/SKILL.md` — image-baked agent
skills (the repo-authored `pi-devbox-environment`, plus vendored fallback
copies of `pi-extensions` and `mempalace` — see `skills/VENDORED.md`)
symlinked into `~/.agents/skills/` by the entrypoint, available with or
without a mounted skillset — plus
`usr/local/share/pi-devbox/pi-global-AGENTS.append.md` (the global-AGENTS
pointer concatenated in `Dockerfile.variant`).
- `scripts/smoke-test.sh` — sanity checks run by CI before pushing to Hub.
- `.gitea/workflows/docker-publish.yml` — two-phase CI (base-decide →
build-base → smoke → build-variant → promote-base-latest →
@@ -33,7 +49,8 @@ re-brand of opencode-devbox's `pi-only` variant.
## Versioning scheme
- Tags follow semver. **v1.0.0** is the first decoupled release; future
minor bumps add variants (`-studio`, `-studio-tex`); patch bumps follow
minor bumps add variants (`-studio`, `-studio-tex`) or significant base
additions (e.g. v1.2.0 image-baked agent skills); patch bumps follow
pi npm version updates and small fixes.
- Docker Hub tags: `joakimp/pi-devbox:vX.Y.Z` + `joakimp/pi-devbox:latest`
+ (since v1.1.0) `joakimp/pi-devbox:vX.Y.Z-studio` +
@@ -45,23 +62,168 @@ re-brand of opencode-devbox's `pi-only` variant.
1. Confirm `pi --version` resolves from npm to the expected version
(`curl -sf 'https://registry.npmjs.org/@earendil-works%2Fpi-coding-agent/latest' | jq -r .version`).
2. Update `CHANGELOG.md` Unreleased → vX.Y.Z section.
3. Verify `docker compose up` works locally with the current `latest` image
Check release notes at https://github.com/earendil-works/pi/releases for
the upstream changelog to include in `CHANGELOG.md`.
2. **Refresh the vendored mempalace skill snapshot if the skillset moved:**
`scripts/vendor-mempalace-skill.sh --check` (reads a real skillset clone,
writes nothing). Three exit codes, not two — a stale-but-truthful record is
**not** a release blocker, so don't treat any non-zero exit as "must
refresh" without reading which one it was:
- **0** — the record is truthful. This includes stale-but-truthful
(upstream has moved past the recorded ref, or the local clone has
uncommitted changes) — a `NOTICE` is printed, but nothing is lying.
**Skipping the refresh in this case is the legitimate, sanctioned
outcome** — every enrolled host reads its own live skillset clone, so
the baked copy is only a no-mount fallback. What is not legitimate is
skipping it *silently*: the drift is visible here, in
`pi-devbox-version`, and in the manifest, so decide rather than forget.
- **1** — a confirmed problem: the vendored bytes provably do NOT match
the file at the recorded ref (a lying record), or the recorded ref
doesn't even resolve to that path in this clone. Refresh.
- **2** — cannot determine (the recorded ref itself isn't resolvable in
this clone — commonly a shallow checkout missing history). Fetch full
history and re-check before deciding; don't refresh blind.
Refresh with `scripts/vendor-mempalace-skill.sh`, which rewrites the file
**and** the ARG together so they cannot drift apart, and refuses (exit 1)
rather than silently rewinding provenance if the skillset clone's HEAD is
behind the already-recorded ref (detached HEAD, older checkout) — pass
`--force` only if that rewind is genuinely intended.
Two consequences to accept deliberately on an actual refresh: the snapshot
is hashed into `base_tag`, so it costs a base rebuild (~67 min); and if the
section the phrase canary names has changed, re-pin it in
`scripts/smoke-test.sh`.
3. Update `CHANGELOG.md` Unreleased → vX.Y.Z section.
4. Verify `docker compose up` works locally with the current `latest` image
if you're upgrading users from a previous version. Then run the
**post-recreate sanity check** inside the running container to confirm
persisted volumes survived and the pi runtime wiring re-deployed (not just
that the container booted):
`docker compose exec devbox bash scripts/recreate-sanity-check.sh --expected-version X.Y.Z`
(or just `pi-devbox-sanity --expected-version X.Y.Z` if `cli_utils/bin` is
on PATH). This is the runtime peer of the build-time `smoke-test.sh` gate.
4. Push tag: `git tag vX.Y.Z && git push origin vX.Y.Z`.
5. Watch CI: smoke job builds amd64 only and asserts size + extensions +
`docker compose exec devbox bash scripts/recreate-sanity-check.sh --expected-image-version X.Y.Z`
(or just `pi-devbox-sanity --expected-image-version X.Y.Z` if
`cli_utils/bin` is on PATH). This is the runtime peer of the build-time
`smoke-test.sh` gate.
**`X.Y.Z` here is the pi-devbox release tag** you are shipping (e.g.
`1.8.9`), which is what the rest of this checklist means by `vX.Y.Z`.
`--expected-image-version` is the flag that asserts it. There is also an
`--expected-version`, and it means something else — the **pi coding agent**
version (e.g. `0.84.3`, the `ARG PI_VERSION` pin). Handing the release tag
to that one used to report *"pi version mismatch: expected 1.8.8, got
0.84.3"*, i.e. a red on the final gate of the release accusing the wrong
component; it now tells you to use `--expected-image-version` instead, and
the reverse mix-up is caught too. Both flags are optional — with neither,
the live pi version is asserted against the version recorded in the image's
own build manifest (which catches a stale `pi` in the `~/.pi/npm-global`
volume shadowing the baked one) and the image tag is reported
informationally.
5. Push tag: `git tag vX.Y.Z && git push origin vX.Y.Z`.
6. Watch CI: smoke job builds amd64 only and asserts size + extensions +
pi version + new-base-tooling presence. Variant build is multi-arch
(amd64 + arm64) only after smoke passes.
6. Verify the Hub tags appear (latest + vX.Y.Z, the `-studio` pair, plus
(amd64 + arm64) only after smoke passes. A tag push fires **only**
`docker-publish.yml` — `lint.yml` is scoped to `branches: ['**']`, which
excludes tag refs on purpose (the tagged tree was already linted when the
commit hit `main`, and a fast lint run sorting above the slow publish run
made releases look finished before anything shipped). Verified on v1.8.4:
`refs/tags/v1.8.4` produced run 571 (publish) and nothing else. Still filter
discovery on `head_sha` **and** the workflow `path` — see *Gitea API access*
below — because that guard costs nothing and a future workflow added on `v*`
would silently reintroduce the ambiguity.
7. Verify the Hub tags appear (latest + vX.Y.Z, the `-studio` pair, plus
base-latest if the base was rebuilt this run).
7. **Revoke any short-lived Gitea PAT** used during the release at
`gitea.jordbo.se/user/settings/applications`.
8. **Revoke any short-lived Gitea PAT** used during the release at
`gitea.jordbo.se/user/settings/applications`. N/A if you used the
`GITEA_ACCESS_TOKEN` env var instead (see *Gitea API access* below) —
its lifecycle is managed host-side, nothing to revoke.
## Verifying this repo's reality from inside a container
Most work on this repo happens **inside** a pi-devbox container, inspecting a
host or a peer over SSH. That setup manufactures convincing false negatives, so
when you are about to report that something is **absent, unreachable, or not
running**, suspect your own command first. Recurring instances:
- **`docker` is not on the host's non-interactive SSH `PATH`.** `ssh mac 'docker
ps'` says *command not found* on a host that plainly runs Docker; use
`/usr/local/bin/docker` (or `command -v docker` first). Every step in the
*Release-day checklist* that inspects a running container hits this.
- **Don't `| head -N` a search whose answer you don't already know.** The host's
`~/.ssh/config` is ~500 lines; a `head -20` "proved" a peer absent that was
defined at line 454.
- **The deployment compose file is not this repo's.** `docker-compose.yml` here
is a template pinning `:latest`; a real host runs its own per-machine file
(find it with `docker inspect <container> --format '{{ index .Config.Labels
"com.docker.compose.project.config_files" }}'`). Recreating from the repo copy
can silently move a host off `:latest-studio` onto `:latest`.
- **A live SSH ControlMaster hides remote auth changes** — after editing a
peer's `authorized_keys`, prove access with `-o ControlPath=none -o
ControlMaster=no`, or the breakage surfaces in a later session instead.
Depth and further mechanisms: the repo-authored `pi-devbox-environment` skill
(`rootfs/usr/local/share/pi-devbox/skills/pi-devbox-environment/SKILL.md`) §2
and §3 — that file is the one an agent actually loads mid-session, whereas this
`AGENTS.md` is only auto-read when the cwd *is* this repo.
## Gitea API access (env token)
`GITEA_ACCESS_TOKEN` + `GITEA_HOST` are passed into the container from the
host `.env` via `docker-compose.yml` (`${GITEA_ACCESS_TOKEN:-}` /
`${GITEA_HOST:-}`), primarily to enable the `gitea-mcp` server. They are
**not** baked into the image. When configured, they are also available for
**any** direct Gitea API interaction from inside the container — inspecting
CI runs, checking published tags, listing commits — e.g.
`curl -H "Authorization: token $GITEA_ACCESS_TOKEN" "$GITEA_HOST/api/v1/repos/joakimp/pi-devbox/actions/runs?limit=20"`.
Prefer this over a short-lived PAT file when the env token is present (the
`ci-release-watcher` skill auto-detects it). Public-repo GET listings work
unauthenticated too, so the token matters mainly for private repos or
rate-limit headroom; its lifecycle is host-managed, so there is nothing to
revoke after use. Never echo the token value (including into logs).
**Gotcha — a tag push fires EVERY workflow whose triggers match the tag ref.**
`lint.yml` uses a bare `push:` trigger, so a release tag yields *both* a lint run
and the publish run. The listing is newest-first and lint sorts **above** the
publish run, so "take the first run whose `path` contains `refs/tags/<tag>`"
picks the wrong one **reliably, not occasionally**. Real listing for v1.6.4:
```
id=531 #104 lint.yml@refs/tags/v1.6.4 <- wrong; sorts first
id=530 #103 docker-publish.yml@refs/tags/v1.6.4 <- the release build
id=529 #102 lint.yml@refs/heads/main <- same commit, linted on push
```
Lint goes green in minutes while the image is still building, so watching it
makes a release look finished when nothing has been published yet.
**Gotcha — the jobs endpoint takes the internal `id`, NOT the `run_number` the
UI shows as `#104`.** The two diverge widely, and `GET
.../actions/runs/<run_number>/jobs` does **not** error — it silently returns a
*different* run's jobs. Always read `id` from the run listing:
```bash
# Which runs did this tag/commit trigger? Filter on head_sha; never trust
# ordering or run numbering. limit=20, not 5 — with two runs per push the
# publish run falls off a 5-item window fast.
curl -sS -H "Authorization: token $GITEA_ACCESS_TOKEN" \
"$GITEA_HOST/api/v1/repos/joakimp/pi-devbox/actions/runs?limit=20" \
| jq --arg sha "$(git rev-list -n1 vX.Y.Z)" \
'.workflow_runs[] | select(.head_sha==$sha) | {id, run_number, path, status, conclusion}'
# pick the id whose .path starts with docker-publish.yml, then:
curl -sS -H "Authorization: token $GITEA_ACCESS_TOKEN" \
"$GITEA_HOST/api/v1/repos/joakimp/pi-devbox/actions/runs/<id>/jobs" \
| jq '.jobs[] | {name, status, conclusion}'
```
**Watcher config for this repo** (`ci-release-watcher` skill, hub-only shape —
pi-devbox has no downstream host to deploy to):
- `EXPECT_WORKFLOW=docker-publish.yml` — the skill's `preflight_run()` aborts at
startup if the run id belongs to lint instead.
- `EXPECTED_FRESH_TAGS='vX.Y.Z latest vX.Y.Z-studio latest-studio'`
- `EXPECTED_EXISTS_TAGS='base-latest'` — existence only: it is content-addressed
and legitimately keeps its old timestamp when the base is a cache hit.
- `CRITICAL_JOBS='build-variant build-variant-studio'` — job names are matched
**exactly** (`critical.issubset(succeeded)`), so the studio variant must be
listed explicitly; the skill's default omits it. Leave `promote-base-latest`
out: it legitimately skips on a base cache hit, which would misclassify a good
run. `update-description` is the cosmetic post-publish job.
## Cache-hit footgun (must-know)
@@ -120,10 +282,14 @@ deprecated artifacts (to be removed in opencode-devbox v2.0.0).
## What we DON'T install (and why)
- **No texlive** (~600 MB–1 GB). Users who need PDF export from pandoc
or pi-studio can install on demand: `sudo apt-get install texlive-xetex
texlive-latex-recommended`. The planned `:latest-studio-tex` variant
will bake this in.
- **No texlive** (~600 MB–1 GB). PDF export from pandoc / pi-studio works
out of the box via **`typst`** (~30 MB static binary), which the base ships
as the pandoc PDF engine (`pandoc --pdf-engine=typst`) — small enough to live
in base rather than a dedicated `:latest-studio-tex` variant. We don't bake in
a full TeX Live: it's heavy and typst covers the common Markdown→PDF case.
Users needing LaTeX-exact output can install the higher-fidelity fallback on
demand: `sudo apt-get install texlive-xetex texlive-latex-recommended` (then
`pandoc --pdf-engine=xelatex`).
- **pi-studio** ships in the `:latest-studio` variant (since v1.1.0),
vendored to `/opt/pi-studio` and registered at container start via
`pi install /opt/pi-studio` (see Dockerfile.variant `INSTALL_STUDIO`).
+3219 -2
View File
File diff suppressed because it is too large Load Diff
+14 -4
View File
@@ -46,11 +46,13 @@ Full setup guide — authentication for each provider (Anthropic, OpenAI, Gemini
### pi and companions
- **pi `{{PI_VERSION}}`** ([`@earendil-works/pi-coding-agent`](https://www.npmjs.com/package/@earendil-works/pi-coding-agent)) — installed at `/usr/bin/pi`
- **pi `{{PI_VERSION}}`** ([`@earendil-works/pi-coding-agent`](https://www.npmjs.com/package/@earendil-works/pi-coding-agent)) — installed at `/usr/bin/pi`, pinned to an audited version (not npm `latest`)
- **pi-atelier** — TUI sidebar (ordered panels, split-pane, themes), vendored at `/opt/pi-atelier` and pinned to an audited tag; the exact tag is in the image labels (`se.jordbo.pi-devbox.pi-atelier-version`) and `/etc/pi-devbox/build-manifest.json`
- **[pi-toolkit](https://gitea.jordbo.se/joakimp/pi-toolkit)** — keybindings (mosh/tmux-friendly Shift+Enter, Ctrl+J, Alt+J newline bindings), AWS env loader, settings template
- **[pi-extensions](https://gitea.jordbo.se/joakimp/pi-extensions)** — 7 user-facing extensions: `ext-toggle`, `mcp-loader`, `todo`, `ssh-controlmaster`, `notify`, `git-checkpoint`, `confirm-destructive`
- **`fork`** ([pi-fork](https://github.com/elpapi42/pi-fork)) and **`recall`** ([pi-observational-memory](https://github.com/elpapi42/pi-observational-memory)) tools
- **mempalace bridge** — MCP extension auto-symlinked so pi reads/writes the host-mounted palace
- **image-baked agent skills** — skills under `/usr/local/share/pi-devbox/skills/` (e.g. `pi-devbox-environment`, which teaches agents the container's persistence/networking/DNS/tmux/REPL specifics) are symlinked into `~/.agents/skills/` on start, available with or without a mounted skillset repo
The entrypoint deploys/registers all of these on first container start. Re-running is idempotent and preserves user edits.
@@ -63,12 +65,19 @@ The entrypoint deploys/registers all of these on first container start. Re-runni
### Document and image tooling
- **pandoc** — universal Markdown↔HTML/Org/RST/etc. conversion. Useful well beyond pi: agent-driven doc exports, format conversion, etc.
- **Typst** — markup-based typesetting, used as pandoc's `--pdf-engine`
- **graphviz** (`dot`) — diagram rendering pipelines
- **imagemagick** (`magick`) — image conversion / resizing
### Browser automation
- **agent-browser** — CLI for driving a real browser (open pages, click/fill/`eval`, snapshot the DOM, screenshots) so agents can verify front-end work instead of guessing
- **Playwright** + a headless **Chromium** are pre-installed and pinned together; `AGENT_BROWSER_EXECUTABLE_PATH` is preset to the baked browser, so `agent-browser open <url>` works out of the box with no setup
- **socat** — TCP bridge used to expose the pi-studio server outside the container's loopback
### Modern CLI tooling
- **Editor**: neovim (LazyVim defaults), tmux (configured for 0-indexed sessions)
- **Editor**: neovim (system-wide `termguicolors` default; bring your own config/plugins), tmux (configured for 0-indexed sessions)
- **Search/nav**: ripgrep, fd, fzf, zoxide
- **Display**: bat, eza, htop, tree
- **Data**: jq, yq
@@ -97,7 +106,7 @@ The entrypoint deploys/registers all of these on first container start. Re-runni
### SSH and networking
- OpenSSH client with **ControlMaster auto** preconfigured on a writable socket path (`/tmp/sshcm/`). Mitigates ssh banner-exchange failures behind CGNAT-restricted residential ISPs (~4-flow caps).
- OpenSSH client with **ControlMaster auto** preconfigured on a writable socket path (`/tmp/sshcm/`). Mitigates ssh banner-exchange failures behind CGNAT-restricted residential ISPs (~4-flow caps). A read-only `~/.ssh` carrying a per-host `ControlPath` (common CGNAT configs) is handled too — redirected to a writable socket dir for both `pi --ssh` and `dssh`/`dscp`.
- A **LAN-access helper** that auto-configures ssh jump-via-host on VM-backed hosts (OrbStack / Docker Desktop on macOS) so the container can reach the host's directly-attached LAN peers (`dssh <peer>` alias; `DEVBOX_LAN_ACCESS` / `HOST_SSH_USER`).
## Versioning
@@ -155,4 +164,5 @@ Optional volumes for MemPalace (commented out by default — uncomment in `docke
## License
MIT (the image; pi and the bundled tools each carry their own licenses).
MIT (the image; pi and the bundled tools each carry their own licenses). See
`LICENSE` and `THIRD_PARTY.md` in the [source repo](https://gitea.jordbo.se/joakimp/pi-devbox).
+313 -52
View File
@@ -14,7 +14,7 @@
# content-addressed over this file, so any byte change invalidates the
# cache. Recommended cadence: once per release for security updates.
#
# BASE_REBUILD_DATE: 2026-06-09 (v1.0.0 — decoupled from opencode-devbox)
# BASE_REBUILD_DATE: 2026-07-13 (Unreleased — agent-browser CLI + Playwright Chromium for headless browser automation; prior: typst PDF engine + xz-utils + pandoc typst-template default-font patch)
#
# ── Lineage note ─────────────────────────────────────────────────────
# Adapted from opencode-devbox/Dockerfile.base (commit before v1.16.2).
@@ -46,16 +46,43 @@ ENV DEBIAN_FRONTEND=noninteractive
# Additions vs the upstream opencode-devbox base (2026-06-09):
# pandoc — Markdown↔HTML/PDF/etc. conversion. Required by pi-studio
# preview/export pipelines and broadly useful for any
# agent-driven document workflow. ~200 MB.
# agent-driven document workflow. ~200 MB. NOTE: pandoc is
# only the front-end — PDF output needs a back-end engine.
# We ship `typst` (installed further down) as the
# lightweight default engine (`pandoc --pdf-engine=typst`)
# instead of a ~600 MB TeX Live install.
# xz-utils — `xz` decompressor. tar shells out to it for `.tar.xz`
# assets (typst ships .tar.xz). ~0.5 MB. Also generally
# useful for extracting xz-compressed archives.
# graphviz — `dot` rendering for many diagram tools. ~10 MB.
# See the bundled `dot-watch` helper for live .dot -> PNG
# re-render (handy with pi-studio's image preview).
# imagemagick — image conversion / resizing for thumbnails, etc. ~50 MB.
# yq — YAML-aware companion to jq.
# (yq is NOT apt-installed: Debian's `yq` is the unrelated Python tool;
# mikefarah's Go yq is installed as a pinned binary further down.)
# socat — TCP relay. Powers `studio-expose`, which bridges
# pi-studio's container-loopback server to the container's
# external interface so a published port can reach it.
# ~1 MB; generally useful for any port-forwarding need.
# nano — small, non-modal terminal editor for users who don't want
# a vi-based editor. ~2.8 MB installed; its deps (libc6,
# libncursesw6, libtinfo6) are already pulled in by nvim/less/
# htop/tmux, so it adds no extra packages. Companion to nvim
# and the `micro` binary installed further down. EDITOR stays
# nvim; users opt in via `export EDITOR=nano`.
# kitty-terminfo — terminfo entry for the kitty terminal (TERM=xterm-kitty).
# ~77 KB, terminfo file only (no kitty binary). Without it,
# ncurses apps fall back and Neovim can't reliably detect
# true-colour from kitty over ssh; installing it makes
# TERM=xterm-kitty understood. Pairs with the system-wide
# Neovim termguicolors default (etc/xdg/nvim/sysinit.vim).
# ncurses-term — broad terminfo bundle (wezterm, alacritty, foot, st, the
# base `ghostty` entry, and many more) so SSHing in from a
# modern emulator resolves its TERM instead of degrading to a
# dumb fallback. xterm-kitty is NOT in it (hence kitty-terminfo
# above); TERM=xterm-ghostty is compiled from an alias further
# down (ncurses ships `ghostty`, not `xterm-ghostty`). iTerm2
# defaults to xterm-256color (ncurses-base), so needs nothing.
RUN apt-get update && \
apt-get upgrade -y --no-install-recommends && \
apt-get install -y --no-install-recommends \
@@ -66,7 +93,6 @@ RUN apt-get update && \
openssh-client \
gnupg \
jq \
yq \
ripgrep \
fd-find \
tree \
@@ -89,9 +115,13 @@ RUN apt-get update && \
python3-pip \
python3-venv \
pandoc \
xz-utils \
graphviz \
imagemagick \
socat \
nano \
kitty-terminfo \
ncurses-term \
&& ln -s /usr/bin/fdfind /usr/local/bin/fd \
&& apt-get clean \
&& rm -rf /var/lib/apt/lists/*
@@ -130,6 +160,15 @@ RUN printf '%s\n' \
# `Include /etc/ssh/ssh_config.d/*.conf` *before* the `Host *` block,
# so user config can override these defaults if desired.
#
# CAVEAT (and why it is handled elsewhere): a user per-host override that
# points ControlPath BACK under the read-only ~/.ssh (e.g. the common CGNAT
# idiom `ControlPath ~/.ssh/cm/%r@%h:%p`) re-introduces the unwritable-socket
# failure — a system drop-in here can never override a user's per-host value.
# For `pi --ssh`, the ssh-controlmaster extension handles this by detecting an
# unwritable system ControlPath and falling back to its own /tmp master; for
# `ssh -F ~/.ssh-local/config` (dssh/dscp), setup-lan-access.sh redirects
# ControlPath into the writable ~/.ssh-local. See CHANGELOG "Unreleased".
#
# ControlPersist=10m means the master socket sticks around 10 min after
# the last session closes, so consecutive ssh calls in a workflow reuse
# the same TCP flow. Companion entrypoint-user.sh creates /tmp/sshcm
@@ -222,6 +261,33 @@ RUN ARCH=$(case "${TARGETARCH}" in amd64) echo "x86_64" ;; arm64) echo "arm64" ;
ln -s /opt/nvim-linux-${ARCH}/bin/nvim /usr/local/bin/nvim && \
nvim --version | head -1
# micro — modern, non-modal terminal editor. Ships alongside nvim so users
# who aren't comfortable with vi-style modal editing have a friendly option:
# desktop-style keybindings (Ctrl+S save, Ctrl+Q quit, Ctrl+C/V/X, Ctrl+Z
# undo), mouse support, and syntax highlighting out of the box. A single
# static Go binary (~12 MB) installed from GitHub releases, exactly like
# bat/eza/zoxide below. EDITOR stays nvim (see below); users opt in with
# `export EDITOR=micro` or `git config --global core.editor micro`.
#
# NOTE: upstream moved zyedidia/micro -> micro-editor/micro. The old org URL
# still 302s, but its /releases/latest redirect lands on ANOTHER /latest URL
# (the org rename), so the tag-parsing idiom below would resolve "latest"
# instead of a version. Use the canonical micro-editor/micro URL.
# Arch asset naming differs from the others: amd64 -> linux64, arm64 ->
# linux-arm64. The tarball extracts to micro-<version>/micro.
ARG MICRO_VERSION=latest
RUN ARCH=$(case "${TARGETARCH}" in amd64) echo "linux64" ;; arm64) echo "linux-arm64" ;; *) echo "linux64" ;; esac) && \
V="${MICRO_VERSION}" && \
if [ "$V" = "latest" ]; then \
V=$(curl -sI --retry 5 --retry-delay 5 --retry-all-errors "https://github.com/micro-editor/micro/releases/latest" | awk 'tolower($1)=="location:" { sub(/\r$/,"",$2); n=split($2,a,"/"); print a[n] }'); \
fi && \
V="${V#v}" && [ -n "$V" ] && \
echo "Installing micro ${V}" && \
curl -fsSL --retry 5 --retry-delay 5 --retry-all-errors "https://github.com/micro-editor/micro/releases/download/v${V}/micro-${V}-${ARCH}.tar.gz" | tar -xz -C /tmp && \
install /tmp/micro-${V}/micro /usr/local/bin/micro && \
rm -rf /tmp/micro-${V} && \
micro --version
# bat
ARG BAT_VERSION=latest
RUN ARCH=$(case "${TARGETARCH}" in amd64) echo "x86_64" ;; arm64) echo "aarch64" ;; *) echo "x86_64" ;; esac) && \
@@ -280,21 +346,98 @@ RUN ARCH=$(case "${TARGETARCH}" in amd64) echo "x86_64" ;; arm64) echo "aarch64"
# Always installed in the base. Set INSTALL_MEMPALACE=false at base-build
# time to shave ~300 MB.
#
# Stall protection (fixed 2026-06-13): mempalace-mcp is launched by the
# `mempalace.ts` pi extension from mempalace-toolkit (cloned below). That
# extension now applies a per-REQUEST timeout in its JSON-RPC client and
# kills the child on stall, so a virtiofs cold-open of chroma.sqlite3 /
# HNSW load can no longer hang the pi TUI uninterruptibly. Tunables:
# Stall protection (fixed 2026-06-13; self-heal added 2026-06-25):
# mempalace-mcp is launched by the `mempalace.ts` pi extension from
# mempalace-toolkit (cloned below). That extension applies a per-REQUEST
# timeout in its JSON-RPC client and kills the child on stall, so a virtiofs
# cold-open of chroma.sqlite3 / HNSW load can no longer hang the pi TUI
# uninterruptibly. A stall-kill is no longer a permanent latch either: the
# next tool call respawns the server with capped exponential backoff (the
# budget resets on any successful response). Tunables:
# MEMPALACE_MCP_TIMEOUT_MS (default 60000), MEMPALACE_MCP_INIT_TIMEOUT_MS
# (default 120000); 0 disables. A standalone stdio-watchdog shim is NOT
# needed — the extension already owns request/response correlation. See
# CHANGELOG.md "Unreleased > Fixed".
# (default 300000 — generous so a genuine first cold-open isn't killed),
# MEMPALACE_MCP_MAX_RESPAWNS (default 2; 0 disables self-heal),
# MEMPALACE_MCP_RESPAWN_BACKOFF_MS (default 1000); timeouts of 0 disable.
# Defaults live in the extension, so no ENV is needed here. A standalone
# stdio-watchdog shim is NOT needed — the extension already owns
# request/response correlation. See CHANGELOG.md "Unreleased > Fixed".
ARG INSTALL_MEMPALACE=true
# Pin to a known-good version. Bump deliberately, not implicitly: an
# unpinned install silently swept in mempalace 3.3.x/3.4.0 with a broken
# diary_write schema (see workaround RUN below + issue #1728). Pinning
# makes mempalace upgrades a reviewable diff rather than a surprise.
ARG MEMPALACE_VERSION=3.4.0
# diary_write schema. Pinning makes mempalace upgrades a reviewable diff
# rather than a surprise.
#
# 3.5.0 (2026-06) shipped the upstream fix for the top-level-anyOf diary_write
# schema (issue #1728 / PR #1717, merged 2026-06-14): the advertised schema
# is now `"required": ["agent_name"]` with entry/content enforced at dispatch,
# which Anthropic's tools API accepts — so the old mcp_server.py perl
# workaround that used to live below is gone.
#
# 3.6.0 (2026-07-17, PyPI latest) is additive/reliability only — secure
# `mempalace serve` remote mode, optional Milvus backend, atomic KG
# supersede(), conversation chronology, mining exclusions, plus recovery and
# locking fixes. Reviewed for MCP tool-schema changes before bumping (that
# being the exact regression class this pin exists to catch): there are NONE,
# and nothing touches diary_write. Two fixes matter for how this image uses
# mempalace: read-only mode now covers checkpoint + delete_by_source in
# _MUTATING_TOOLS (#1930), and agent attribution is preserved in
# mempalace_checkpoint (#2023/#2034).
#
# Keep in lockstep with opencode-devbox when bumping.
#
# 3.7.1 (from 3.6.0) is safe for anyone with an EXISTING LOCAL palace: verified
# against the 3.7.1 source, not the changelog. Legacy drawers lack the new
# `chunk_total` marker and both decision sites trust them ("trust the match as
# before"), NORMALIZE_VERSION is 2 in both, chromadb stays <2 (no index-format
# migration), there is no auto-migration ("We do NOT auto-migrate"), and the one
# new palace file (logstream.sqlite3) is created lazily on first logstream use.
# Two behaviour changes to know: MEMPALACE_MCP_ALLOW_PEER_WRITER no longer works
# on local/chroma palaces, and writer-lock setup failures now fail CLOSED
# (refuse the write) rather than fail open. Neither affects the container's
# normal MCP-server-plus-CLI-feeder pattern, which already serialised on the
# same lock under 3.6.0.
#
# 3.8.0 (2026-08-23, PyPI, released hours after this project's own v1.8.5 tag
# the same day) is additive/reliability only — reviewed for MCP tool-schema
# changes before bumping, as always: there are NONE. Two PRs matter:
# - PR #2320/#2322: `sync --apply` no longer deletes a drawer solely because
# its source_file was unreachable AT THAT MOMENT — it now asks for
# corroboration first. This fixes losing a whole mined project to one
# `sync --apply` while its volume happened to be unmounted.
# IMPORTANT — do not over-read this fix: it addresses TRANSIENT
# unreachability, not the standing landmine (documented in the operator's
# global AGENTS.md) against running `mempalace_sync` / `mempalace_delete_by_source`
# beyond dry-run on the SHARED central palace. On that palace most
# source_file paths are PERMANENTLY absent from whichever host runs the
# sync — a different machine's paths simply do not exist here, ever, not
# merely "right now". That is a different failure shape than #2320/#2322
# fixes. The landmine still stands; this bump does not relax it.
# - PR #2307: long-running Chroma servers no longer invalidate their own
# HNSW cache on their own writes (server-side perf fix). This does NOT
# make `mempalace_reconnect` unnecessary — that tool exists for EXTERNAL
# writes bypassing the in-process client (e.g. direct sqlite backfills,
# CLI commands against a running server), a different scenario #2307
# does not touch.
#
# CI-side audit (added after v1.8.6, closing that release's "Still open" item):
# resolve-versions now treats this pin exactly as it treats PI_VERSION — it
# reads the ARG from THIS file, refuses a non-concrete value, verifies the
# version is published on PyPI, refuses a YANKED release (an exact pin installs
# one silently under PEP 592), and WARNS — never silently adopts — when PyPI has
# a newer release. smoke-test.sh then asserts the installed core equals that
# audited pin, which catches a stale cached base layer that no manifest-internal
# check can see. So a bump here is now gated end to end; what remains manual is
# the JUDGEMENT above (MCP schema review, server/client sequencing), which is
# the part that should stay manual.
#
# Deployment sequencing note for whoever ships this bump: synlig (the shared
# central palace host) currently serves mempalace 3.7.1 SERVER-SIDE via
# docker-compose.mempalace.yml, which reuses this same devbox image. Bumping
# this ARG changes only the CLIENT version baked into pi-devbox images: it
# introduces client/server skew until synlig's compose stack is separately
# rebuilt/redeployed with the new pin. Not something to code around here —
# just sequence the redeploy.
ARG MEMPALACE_VERSION=3.8.0
ENV UV_TOOL_DIR=/opt/uv-tools
ENV UV_TOOL_BIN_DIR=/usr/local/bin
RUN if [ "${INSTALL_MEMPALACE}" = "true" ]; then \
@@ -303,45 +446,18 @@ RUN if [ "${INSTALL_MEMPALACE}" = "true" ]; then \
/opt/uv-tools/mempalace/bin/python -c "import mempalace; print('mempalace', mempalace.__version__ if hasattr(mempalace, '__version__') else 'installed')" ; \
fi
# ── workaround: strip top-level anyOf from mempalace_diary_write schema ──
# Mempalace 3.3.x/3.4.0 advertise diary_write's input_schema with a
# top-level `anyOf: [{required:[entry]}, {required:[content]}]` to express
# "either entry or content must be supplied". Anthropic's tools API rejects
# top-level anyOf/oneOf/allOf, so pi/Claude fail at session start with
# `tools.<n>.custom.input_schema: input_schema does not support oneOf,
# allOf, or anyOf at the top level`.
#
# Patch the advertised schema to require ["agent_name", "entry"] and remove
# the anyOf block. The handler keeps accepting `content` server-side as a
# kwarg alias so existing callers still work.
#
# Idempotent and self-deactivating: once upstream releases the fix the
# regex no longer matches (and the WARN below fires) — that's the signal
# to delete this RUN.
# Upstream status (last checked 2026-06-14):
# issue #1728 — STILL OPEN (root-level anyOf rejected by Anthropic/Codex)
# PR #1735 — CLOSED UNMERGED 2026-06-11; do NOT watch it (dead)
# PR #1717 — open; the current live fix candidate to watch
# mempalace PyPI latest = 3.4.0 (== our pin) → no release contains the fix yet
# https://github.com/MemPalace/mempalace/issues/1728
# https://github.com/MemPalace/mempalace/pull/1717
# TODO: remove this RUN once a mempalace release > 3.4.0 that actually strips
# the root-level anyOf ships on PyPI and is installed by the line above.
# Keep MEMPALACE_VERSION in lockstep with opencode-devbox when bumping.
RUN if [ "${INSTALL_MEMPALACE}" = "true" ]; then \
MP_FILE="$(find /opt/uv-tools/mempalace -path '*/mempalace/mcp_server.py' | head -n1)" && \
if [ -z "$MP_FILE" ]; then echo "mempalace mcp_server.py not found" >&2; exit 1; fi && \
perl -0777 -i -pe 's/(?:[ \t]*\#[^\n]*\n)*[ \t]*"required":\s*\[\s*"agent_name"\s*\]\s*,\s*\n[ \t]*"anyOf":\s*\[\s*\n[ \t]*\{\s*"required":\s*\[\s*"entry"\s*\]\s*\}\s*,\s*\n[ \t]*\{\s*"required":\s*\[\s*"content"\s*\]\s*\}\s*,?\s*\n[ \t]*\]\s*,\s*\n/ "required": ["agent_name", "entry"],\n/s' "$MP_FILE" && \
if grep -q '"required": \["agent_name", "entry"\]' "$MP_FILE"; then \
echo "mempalace diary_write anyOf workaround: applied (or already clean)"; \
else \
echo "WARN: mempalace diary_write anyOf workaround did not match expected schema — upstream may have changed shape" >&2; \
fi ; \
fi
# (The mempalace diary_write top-level-anyOf workaround that patched
# mcp_server.py here was removed in v1.2.2 — fixed upstream in mempalace
# 3.5.0 via issue #1728 / PR #1717 (merged 2026-06-14). See CHANGELOG.md.)
# ── mempalace-toolkit — bash wrappers for session/docs mining ────────
ARG INSTALL_MEMPALACE_TOOLKIT=true
ARG MEMPALACE_TOOLKIT_REF=main
# MEMPALACE_TOOLKIT_REPO defaults to the canonical gitea origin but is
# overridable so a relocated/forked build can clone from a mirror or a
# different host without editing this Dockerfile (mirrors the
# PI_FORK_REPO / PI_OBSMEM_REPO / PI_STUDIO_REPO pattern in the variant).
ARG MEMPALACE_TOOLKIT_REPO=https://gitea.jordbo.se/joakimp/mempalace-toolkit.git
# MEMPALACE_TOOLKIT_REF accepts EITHER a branch name OR a commit SHA. CI
# resolves it to a SHA (resolve-versions job) and folds that SHA into the
# base-decide hash so the base rebuilds when the toolkit moves. `git clone
@@ -351,7 +467,7 @@ ARG MEMPALACE_TOOLKIT_REF=main
RUN if [ "${INSTALL_MEMPALACE}" = "true" ] && [ "${INSTALL_MEMPALACE_TOOLKIT}" = "true" ]; then \
rm -rf /opt/mempalace-toolkit && mkdir -p /opt/mempalace-toolkit && \
git -C /opt/mempalace-toolkit init -q && \
git -C /opt/mempalace-toolkit remote add origin https://gitea.jordbo.se/joakimp/mempalace-toolkit.git && \
git -C /opt/mempalace-toolkit remote add origin "${MEMPALACE_TOOLKIT_REPO}" && \
ok=0; for i in 1 2 3 4 5; do \
if git -C /opt/mempalace-toolkit fetch --depth 1 origin "${MEMPALACE_TOOLKIT_REF}" && \
git -C /opt/mempalace-toolkit checkout -q FETCH_HEAD; then ok=1; break; fi; \
@@ -361,9 +477,15 @@ RUN if [ "${INSTALL_MEMPALACE}" = "true" ] && [ "${INSTALL_MEMPALACE_TOOLKIT}" =
[ "$ok" = "1" ] && \
ln -sf /opt/mempalace-toolkit/bin/mempalace-session /usr/local/bin/mempalace-session && \
ln -sf /opt/mempalace-toolkit/bin/mempalace-docs /usr/local/bin/mempalace-docs && \
chmod +x /opt/mempalace-toolkit/bin/mempalace-session /opt/mempalace-toolkit/bin/mempalace-docs && \
ln -sf /opt/mempalace-toolkit/bin/mempalace-pi-session /usr/local/bin/mempalace-pi-session && \
ln -sf /opt/mempalace-toolkit/bin/mempalace-census /usr/local/bin/mempalace-census && \
chmod +x /opt/mempalace-toolkit/bin/mempalace-session /opt/mempalace-toolkit/bin/mempalace-docs \
/opt/mempalace-toolkit/bin/mempalace-pi-session \
/opt/mempalace-toolkit/bin/mempalace-census && \
mempalace-session --help >/dev/null && \
mempalace-docs --help >/dev/null && \
mempalace-pi-session --help >/dev/null && \
mempalace-census --help >/dev/null && \
echo "mempalace-toolkit installed at $(cd /opt/mempalace-toolkit && git rev-parse --short HEAD)" ; \
fi
@@ -392,6 +514,11 @@ ENV LANG=en_US.UTF-8
ENV LANGUAGE=en_US:en
ENV LC_ALL=en_US.UTF-8
ENV EDITOR=nvim
# Advertise 24-bit colour so colour-aware tools (Neovim's own auto-detect, bat,
# delta, ...) use true colour instead of a 256-colour fallback. Safe for the
# modern terminals this devbox targets; override by exporting `COLORTERM=`
# (empty) from a terminal that lacks true-colour support.
ENV COLORTERM=truecolor
ENV PATH="/home/developer/.local/bin:/home/developer/.cargo/bin:${PATH}"
# ── Node.js (required for pi + MCP servers + tldr) ──
@@ -400,6 +527,58 @@ RUN curl -fsSL --retry 5 --retry-delay 5 --retry-all-errors https://deb.nodesour
apt-get install -y --no-install-recommends nodejs && \
rm -rf /var/lib/apt/lists/*
# ── agent-browser — headless browser automation for the agent ────────
# Gives the agent a real browser it can drive (open/click/fill/eval/
# screenshot) so front-end work involving live DOM or WebGL can be VERIFIED
# rather than guessed at. The `agent-browser` skill (shipped from the
# skillset repo, not this image) documents the CLI; without this block that
# skill is a no-op because the binary isn't present. Verified end-to-end
# 2026-07-13: drives the baked Chromium headless (open + screenshot + eval
# into a WebGL SPA) — doctor's launch test passes in ~0.5s.
#
# TWO pieces, because agent-browser is a standalone Rust CLI that ships NO
# browser of its own — it only drives one you provide:
# 1. the CLI itself (npm; ~70 MB of prebuilt native binaries), and
# 2. a Chromium, which we fetch via Playwright.
#
# Why Playwright fetches the browser (and NOT `agent-browser install`):
# agent-browser's own installer drops Chrome under ~/.agent-browser/browsers
# — inside /home/${USER_NAME}, which is a NAMED VOLUME at runtime, so a
# build-time download would be SHADOWED (invisible) once the volume mounts.
# Playwright honours PLAYWRIGHT_BROWSERS_PATH, so we place the browser under
# /usr/local/share (never shadowed) and hand agent-browser a STABLE symlink
# via AGENT_BROWSER_EXECUTABLE_PATH — the symlink insulates the ENV from
# Playwright's per-version, per-ARCH browser directory (`chrome-linux` on arm64,
# `chrome-linux64` on amd64 — Chrome-for-Testing), so we `find` the `chrome`
# binary rather than hardcode the path; the headless-shell binary is named
# `chrome-headless-shell`, so `-name chrome` skips it.
#
# `playwright install --with-deps chromium` also apt-installs Chromium's
# runtime libs; verified to resolve correctly on Debian trixie (exit 0 — the
# t64 library renames are handled by Playwright's dep list). Build runs as
# root, so the apt step works. NPM_CONFIG_PREFIX=/usr keeps both CLIs on /usr
# so they survive the ~/.pi/npm-global volume mount (same trick the variant
# uses for pi). After fetching, we DROP Playwright's `chromium_headless_shell-*`
# build — agent-browser drives the full chrome (verified, incl. headless), so the
# headless shell is dead weight — and clean the apt/npm caches, trimming the
# layer to ~625 MB (Chromium) from ~960 MB. Still the bulk of the base's size,
# and the one real tradeoff of shipping this to every variant.
ARG AGENT_BROWSER_VERSION=latest
ARG PLAYWRIGHT_VERSION=latest
ENV PLAYWRIGHT_BROWSERS_PATH=/usr/local/share/ms-playwright
RUN NPM_CONFIG_PREFIX=/usr npm install -g \
"agent-browser@${AGENT_BROWSER_VERSION}" \
"playwright@${PLAYWRIGHT_VERSION}" && \
playwright install --with-deps chromium && \
CHROME="$(find "${PLAYWRIGHT_BROWSERS_PATH}" -type f -name chrome -path '*/chromium-*/*' | head -n1)" && \
[ -n "$CHROME" ] && ln -sf "$CHROME" /usr/local/bin/agent-chrome && \
agent-browser --version && \
test -x "$(readlink -f /usr/local/bin/agent-chrome)" && \
rm -rf "${PLAYWRIGHT_BROWSERS_PATH}"/chromium_headless_shell-* && \
npm cache clean --force && \
rm -rf /var/lib/apt/lists/* /root/.npm /tmp/*
ENV AGENT_BROWSER_EXECUTABLE_PATH=/usr/local/bin/agent-chrome
# ── tldr (tealdeer) — community-maintained command examples ──────────
# Tealdeer is a Rust port of the tldr-pages client; ~5 MB static binary,
# ~135 MB smaller than the Node tldr global. Same `tldr` command, same UX.
@@ -415,6 +594,61 @@ RUN ARCH=$(case "${TARGETARCH}" in amd64) echo "x86_64" ;; arm64) echo "aarch64"
chmod +x /usr/local/bin/tldr && \
tldr --version
# ── typst — lightweight PDF engine for pandoc (Markdown→PDF) ─────────
# pandoc (apt-installed above) is only a front-end; rendering PDF needs a
# back-end engine. Rather than a ~600 MB TeX Live install, we ship typst:
# a single ~30 MB static Rust binary with no LaTeX dependency. pi-studio's
# PDF export (studio_export_pdf) and pandoc invocations use it via
# `pandoc --pdf-engine=typst`. A fuller TeX Live remains the higher-
# fidelity fallback for anyone who needs LaTeX-exact output (not shipped
# here — install on demand or in a future variant).
#
# Follows the `latest` GitHub-release convention (like tealdeer/uv/bat).
# typst ships a `.tar.xz` asset (hence xz-utils in the apt layer above)
# that extracts to typst-<arch>-unknown-linux-musl/typst. Pin a specific
# tag with --build-arg TYPST_VERSION=vX.Y.Z.
#
# We also patch pandoc's bundled typst template
# (/usr/share/pandoc/data/templates/template.typst): its conf() defaults the
# document font to an empty tuple (`font: ()`), so a naked
# `pandoc --pdf-engine=typst` fails with "font fallback list must not be empty"
# unless the caller passes `-V mainfont=...`. We default it to Libertinus Serif
# (typst's own bundled default font) so PDF export works out-of-the-box.
ARG TYPST_VERSION=latest
RUN ARCH=$(case "${TARGETARCH}" in amd64) echo "x86_64" ;; arm64) echo "aarch64" ;; *) echo "x86_64" ;; esac) && \
V="${TYPST_VERSION}" && \
if [ "$V" = "latest" ]; then \
V=$(curl -sI --retry 5 --retry-delay 5 --retry-all-errors "https://github.com/typst/typst/releases/latest" | awk 'tolower($1)=="location:" { sub(/\r$/,"",$2); n=split($2,a,"/"); print a[n] }'); \
fi && \
V="${V#v}" && [ -n "$V" ] && \
echo "Installing typst ${V}" && \
curl -fsSL --retry 5 --retry-delay 5 --retry-all-errors "https://github.com/typst/typst/releases/download/v${V}/typst-${ARCH}-unknown-linux-musl.tar.xz" | tar -xJ -C /tmp && \
install /tmp/typst-${ARCH}-unknown-linux-musl/typst /usr/local/bin/typst && \
rm -rf /tmp/typst-${ARCH}-unknown-linux-musl && \
typst --version && \
sed -i 's/^ font: (),$/ font: ("Libertinus Serif",),/' /usr/share/pandoc/data/templates/template.typst && \
grep -q 'font: ("Libertinus Serif",),' /usr/share/pandoc/data/templates/template.typst
# ── yq (mikefarah) — YAML processor, jq's companion for YAML ─────────
# Installed as the mikefarah Go binary — NOT Debian's `yq` apt package, which
# is the unrelated Python kislyuk/yq (a jq wrapper with different syntax and
# version line, e.g. 3.x). The cloud-init repo's deploy.sh/provision.sh
# require mikefarah yq v4 (the unrelated Debian python yq is v3.x). Follows
# the repo's `latest` convention (like tealdeer/uv/etc.); the smoke test pins
# the contract to major v4, so a future yq v5 fails CI instead of silently
# breaking provision.sh. Pin a specific tag with --build-arg YQ_VERSION=vX.Y.Z.
ARG YQ_VERSION=latest
RUN ARCH=$(case "${TARGETARCH}" in amd64) echo "amd64" ;; arm64) echo "arm64" ;; *) echo "amd64" ;; esac) && \
V="${YQ_VERSION}" && \
if [ "$V" = "latest" ]; then \
V=$(curl -sI --retry 5 --retry-delay 5 --retry-all-errors "https://github.com/mikefarah/yq/releases/latest" | awk 'tolower($1)=="location:" { sub(/\r$/,"",$2); n=split($2,a,"/"); print a[n] }'); \
fi && \
[ -n "$V" ] && \
echo "Installing mikefarah yq ${V}" && \
curl -fsSL --retry 5 --retry-delay 5 --retry-all-errors "https://github.com/mikefarah/yq/releases/download/${V}/yq_linux_${ARCH}" -o /usr/local/bin/yq && \
chmod +x /usr/local/bin/yq && \
yq --version
# ── AWS CLI v2 (for SSO/Bedrock authentication) ─────────────────────
RUN ARCH=$(case "${TARGETARCH}" in \
amd64) echo "x86_64" ;; \
@@ -464,16 +698,43 @@ ENV PATH="/home/${USER_NAME}/.pi/npm-global/bin:${PATH}"
RUN mkdir -p /etc/skel-devbox
COPY rootfs/home/developer/.bash_aliases /etc/skel-devbox/.bash_aliases
COPY rootfs/home/developer/.inputrc /etc/skel-devbox/.inputrc
COPY rootfs/home/developer/.gitignore_global /etc/skel-devbox/.gitignore_global
# ── Editor defaults: system-wide Neovim true-colour ──────────────────
# /etc/xdg/nvim/sysinit.vim is Neovim's system vimrc: it loads for every user
# (before any personal ~/.config/nvim) and can still be overridden per-user.
# Enables termguicolors so the default theme renders in 24-bit colour instead
# of a muddy 256-colour fallback. Pairs with kitty-terminfo (installed above).
COPY rootfs/etc/xdg/nvim/sysinit.vim /etc/xdg/nvim/sysinit.vim
# ── Terminal support: xterm-ghostty terminfo alias ──────────────────
# ncurses-term (installed above) covers wezterm/alacritty/foot/st and the base
# `ghostty` entry, but Ghostty connects with TERM=xterm-ghostty, for which no
# distro packages an entry. Ship a thin alias (use=ghostty) and compile it into
# the system terminfo db with `tic -x`, so it inherits the maintained ghostty
# capability set. The `infocmp` check fails the build if the entry didn't land.
COPY rootfs/usr/local/share/terminfo-src/ghostty.terminfo /usr/local/share/terminfo-src/ghostty.terminfo
RUN tic -x -o /usr/share/terminfo /usr/local/share/terminfo-src/ghostty.terminfo && \
infocmp -x xterm-ghostty >/dev/null
# ── Entrypoint ────────────────────────────────────────────────────────
COPY rootfs/usr/local/lib/pi-devbox/ /usr/local/lib/pi-devbox/
# Image-baked skills + the global-AGENTS append snippet. Under /usr/local so a
# named volume over a home dir can't shadow them; linked into ~/.agents/skills
# by entrypoint-user.sh, and the snippet is concatenated onto the global
# AGENTS.md in Dockerfile.variant (after pi-toolkit, which owns that file).
COPY rootfs/usr/local/share/pi-devbox/ /usr/local/share/pi-devbox/
COPY rootfs/usr/local/bin/studio-expose /usr/local/bin/studio-expose
COPY rootfs/usr/local/bin/dot-watch /usr/local/bin/dot-watch
COPY rootfs/usr/local/bin/pi-devbox-version /usr/local/bin/pi-devbox-version
COPY rootfs/usr/local/bin/devbox-skill-reconcile /usr/local/bin/devbox-skill-reconcile
COPY entrypoint.sh /usr/local/bin/entrypoint.sh
COPY entrypoint-user.sh /usr/local/bin/entrypoint-user.sh
RUN chmod +x /usr/local/bin/entrypoint.sh /usr/local/bin/entrypoint-user.sh \
/usr/local/bin/studio-expose \
/usr/local/bin/dot-watch \
/usr/local/bin/pi-devbox-version \
/usr/local/bin/devbox-skill-reconcile \
/usr/local/lib/pi-devbox/*.sh 2>/dev/null || true
# Start as root — entrypoint adjusts UID/GID then drops to developer
+278 -13
View File
@@ -29,18 +29,57 @@ ARG USER_NAME=developer
# runs each repo's install.sh on container start so symlinks land under
# ~/.pi/agent/ on the named volume.
#
# PI_VERSION should be passed explicitly by CI as a concrete version
# (resolved from `npm view @earendil-works/pi-coding-agent version`).
# The default `latest` is for local dev convenience only — it has a
# known cache-hit footgun in registry-cached CI builds: the resulting
# build-arg string is byte-identical across builds, the layer-hash is
# identical, and the registry buildcache silently reuses the layer
# from whatever pi version was current when the cache was first
# populated. CI MUST pass a resolved concrete version. See pi-devbox
# v0.75.5b 2026-05-23 for the discovery + canonical fix.
ARG PI_VERSION=latest
# ── pi version pin: an AUDITED CHECKPOINT, not a freeze ──────────────
# PI_VERSION is pinned to a version whose upstream CHANGELOG has been read
# against this image's integration surface: the theme/TUI API that pi-atelier
# couples to, the session `.jsonl` format that `pi-session-repair` parses, the
# extension/package loader, and the Node engine floor. CI reads THIS LINE as
# the single source of truth (see the `resolve-versions` job) and no longer
# follows npm `latest` — following it meant every release silently adopted
# whatever pi shipped that morning, unaudited, in the very build that then got
# tagged and published.
#
# BUMPING IS ROUTINE AND EXPECTED — the pin exists to force a look, not to
# hold a version forever:
# 1. Read the upstream CHANGELOG for every version between old and new.
# 2. Re-check the companions that couple to pi's private TUI/renderer
# internals — pi-atelier above all (see PI_ATELIER_REF below for the
# 0.6.0-under-pi-0.84 startup-hang precedent).
# 3. Bump this line, record the audit in CHANGELOG.md, then tag.
# CI fails the build if this pin is not a published npm version, and warns —
# without adopting it — when npm `latest` has moved ahead. That warning is the
# prompt to do step 1; it is not something to silence.
#
# A concrete version here ALSO defeats the registry-buildcache cache-hit
# footgun that `latest` carried: a byte-identical build-arg string produced an
# identical layer hash, so the cache reused the layer from whatever pi was
# current when it was first populated (shipped the same bytes for pi-devbox
# v0.74.0..v0.75.5; discovered + fixed in v0.75.5b, 2026-05-23). The `latest`
# branch below is kept only for a deliberate local `docker build` override.
#
# AUDITED AT 0.84.3 (2026-08-25, was 0.84.2): upstream's notes carry a
# "Breaking Changes" heading — `GoogleThinkingLevel` renamed to
# `GoogleApiThinkingLevel`. INERT FOR THIS IMAGE: all four vendored companions
# (/opt/pi-fork, /opt/pi-observational-memory, /opt/pi-atelier, /opt/pi-studio)
# were grepped for that symbol and reference it ZERO times, so nothing here
# couples to the renamed type. Recorded because the heading will look alarming
# to the next reader doing step 1 above — the audit is done, don't redo it.
# Adopted for two fixes that land squarely on this repo's own vendored-skill
# wiring (see devbox-skill-reconcile, v1.8.5): nested Markdown skills inside
# `.agents/skills/<group>/` directories were not discovered, and root Markdown
# files such as README.md / AGENTS.md inside a skill dir were reported as
# broken skills unless they declared valid skill frontmatter.
# pi-atelier needs no companion bump: v0.8.2 clears the >=0.7.1 floor that
# pi >= 0.84 requires (see PI_ATELIER_REF below).
ARG PI_VERSION=0.84.3
ARG PI_TOOLKIT_REF=main
ARG PI_EXTENSIONS_REF=main
# Repo URLs default to the canonical gitea origin but are overridable so a
# relocated/forked build can clone from a mirror or a different host
# without editing this Dockerfile — same pattern as PI_FORK_REPO /
# PI_OBSMEM_REPO / PI_STUDIO_REPO below.
ARG PI_TOOLKIT_REPO=https://gitea.jordbo.se/joakimp/pi-toolkit.git
ARG PI_EXTENSIONS_REPO=https://gitea.jordbo.se/joakimp/pi-extensions.git
# pi-fork (fork tool) + pi-observational-memory (recall tool) live on GitHub
# under elpapi42. CI resolves these to commit SHAs to defeat the same
# cache-hit footgun that affects PI_VERSION.
@@ -48,6 +87,29 @@ ARG PI_FORK_REPO=https://github.com/elpapi42/pi-fork.git
ARG PI_FORK_REF=master
ARG PI_OBSMEM_REPO=https://github.com/elpapi42/pi-observational-memory.git
ARG PI_OBSMEM_REF=master
# pi-atelier (TUI sidebar: ordered panels, split-pane, themes) is PINNED TO A
# TAG, which CI resolves to that tag's commit SHA — same treatment as
# pi-studio, for reproducibility plus cache-busting.
#
# This floor is hard-earned. pi-atelier 0.6.0/0.7.0 wrapped pi's PRIVATE TUI
# renderer, and under pi 0.84 that wrapper recursed: pi hung at startup with
# sustained CPU. Upstream fixed the recursion in 0.7.1 and restored the
# non-overlapping split in 0.7.2 — "avoiding the recursive render path that
# caused startup hangs and sustained CPU usage". Its own peerDependencies
# still say `>=0.80.7`, which does NOT encode that floor, so nothing would
# have warned us: NEVER pair pi-atelier < 0.7.1 with pi >= 0.84. Bump this
# pin and PI_VERSION together, checking atelier's CHANGELOG for the pi
# version it claims to track.
#
# No `npm install` step, unlike pi-fork/pi-observational-memory/pi-studio:
# pi-atelier declares ZERO runtime dependencies (only peerDeps, satisfied by
# the baked pi) and has no build step — pi loads its TypeScript directly from
# the /opt checkout. Adding an install here would be a no-op that only costs
# build time.
ARG PI_ATELIER_REPO=https://github.com/michaelmjhhhh/pi-atelier.git
ARG PI_ATELIER_REF=v0.8.2
# Human-readable tag PI_ATELIER_REF was resolved from; recorded as a label.
ARG PI_ATELIER_VERSION=v0.8.2
RUN set -e && \
# git_fetch_ref: clone-equivalent helper that accepts EITHER a branch name
@@ -77,16 +139,58 @@ RUN set -e && \
NPM_CONFIG_PREFIX=/usr npm install -g @earendil-works/pi-coding-agent@${PI_VERSION} ; \
fi && \
pi --version && \
git_fetch_ref "https://gitea.jordbo.se/joakimp/pi-toolkit.git" "${PI_TOOLKIT_REF}" /opt/pi-toolkit && \
git_fetch_ref "https://gitea.jordbo.se/joakimp/pi-extensions.git" "${PI_EXTENSIONS_REF}" /opt/pi-extensions && \
git_fetch_ref "${PI_TOOLKIT_REPO}" "${PI_TOOLKIT_REF}" /opt/pi-toolkit && \
git_fetch_ref "${PI_EXTENSIONS_REPO}" "${PI_EXTENSIONS_REF}" /opt/pi-extensions && \
git_fetch_ref "${PI_FORK_REPO}" "${PI_FORK_REF}" /opt/pi-fork && \
git_fetch_ref "${PI_OBSMEM_REPO}" "${PI_OBSMEM_REF}" /opt/pi-observational-memory && \
git_fetch_ref "${PI_ATELIER_REPO}" "${PI_ATELIER_REF}" /opt/pi-atelier && \
(cd /opt/pi-fork && npm install --omit=dev --no-audit --no-fund) && \
(cd /opt/pi-observational-memory && npm install --omit=dev --no-audit --no-fund) && \
echo "pi-toolkit at $(cd /opt/pi-toolkit && git rev-parse --short HEAD)" && \
echo "pi-extensions at $(cd /opt/pi-extensions && git rev-parse --short HEAD)" && \
echo "pi-fork at $(cd /opt/pi-fork && git rev-parse --short HEAD)" && \
echo "pi-observational-memory at $(cd /opt/pi-observational-memory && git rev-parse --short HEAD)"
echo "pi-observational-memory at $(cd /opt/pi-observational-memory && git rev-parse --short HEAD)" && \
echo "pi-atelier at $(cd /opt/pi-atelier && git rev-parse --short HEAD) (${PI_ATELIER_VERSION})"
# ── Image-baked skill refresh: pi-extensions (Option 1 over Option 2) ──
# rootfs ships a VENDORED snapshot of the pi-extensions skill at
# /usr/local/share/pi-devbox/skills/pi-extensions/ (the "floor" — guarantees the
# skill is always in the image). The pi-extensions PACKAGE repo now co-locates
# the canonical skill under skill/, so here — after the pinned clone — we copy
# that over the snapshot. Result: a normal build ships the fresh, package-owned
# copy (pinned + recorded in the manifest via PI_EXTENSIONS_REF); a build whose
# ref predates the skill, or a fork pointing at a mirror without it, still ships
# the committed snapshot. The skill calls ./evaluate-extension-usage.py, so it
# is copied alongside. Idempotent and cache-safe (depends only on the clone).
RUN if [ -f /opt/pi-extensions/skill/SKILL.md ]; then \
cp /opt/pi-extensions/skill/SKILL.md \
/usr/local/share/pi-devbox/skills/pi-extensions/SKILL.md && \
if [ -f /opt/pi-extensions/skill/evaluate-extension-usage.py ]; then \
cp /opt/pi-extensions/skill/evaluate-extension-usage.py \
/usr/local/share/pi-devbox/skills/pi-extensions/evaluate-extension-usage.py ; \
fi && \
echo "refreshed pi-extensions skill from package @ $(cd /opt/pi-extensions && git rev-parse --short HEAD)" ; \
else \
echo "pi-extensions package has no skill/ at this ref — keeping vendored snapshot" ; \
fi
# ── pi-devbox awareness: append our pointer to the global AGENTS.md ──
# pi loads a SINGLE global instruction file (~/.pi/agent/AGENTS.md), which
# pi-toolkit's install.sh re-symlinks to /opt/pi-toolkit/pi-global-AGENTS.md on
# every container start. There is no second global slot, and that file is
# root-owned (not writable by the runtime user), so we compose at BUILD time:
# append the pi-devbox managed block to pi-toolkit's file here, after the clone.
# Idempotent via a marker grep so a rebuilt layer never double-appends. This
# makes every container proactively aware of the pi-devbox-environment skill;
# the snippet itself is gated (only fires when /usr/local/lib/pi-devbox exists).
RUN if [ -f /opt/pi-toolkit/pi-global-AGENTS.md ] && \
! grep -q 'pi-devbox:managed-block' /opt/pi-toolkit/pi-global-AGENTS.md; then \
printf '\n' >> /opt/pi-toolkit/pi-global-AGENTS.md && \
cat /usr/local/share/pi-devbox/pi-global-AGENTS.append.md >> /opt/pi-toolkit/pi-global-AGENTS.md && \
echo "appended pi-devbox block to pi-global-AGENTS.md" ; \
else \
echo "pi-devbox block already present or pi-global-AGENTS.md missing (skipped)" ; \
fi
# ── Optional: pi-studio (:latest-studio variant) ─────────────────────
# pi-studio (omaclaren/pi-studio) is a pi-package + theme providing a
@@ -112,6 +216,10 @@ RUN set -e && \
ARG INSTALL_STUDIO=false
ARG PI_STUDIO_REPO=https://github.com/omaclaren/pi-studio.git
ARG PI_STUDIO_REF=main
# PI_STUDIO_VERSION is the human-readable tag (e.g. v0.9.36) that PI_STUDIO_REF
# was resolved from; recorded as a label below for at-a-glance identification.
# Only meaningful for the studio variant (default `none` otherwise).
ARG PI_STUDIO_VERSION=none
RUN if [ "${INSTALL_STUDIO}" = "true" ]; then \
set -e; \
rm -rf /opt/pi-studio && mkdir -p /opt/pi-studio && \
@@ -154,4 +262,161 @@ RUN if [ "${INSTALL_GO}" = "true" ]; then \
ln -s /usr/local/go/bin/gofmt /usr/local/bin/gofmt; \
fi
# ── Build provenance: OCI labels + on-disk build manifest ────────────
# Records exactly which pi version and companion-repo commits were baked
# into THIS image, so a published tag is self-describing and reproducible
# after the fact (CI logs rotate; a released image must not depend on
# them). Previously the resolved SHAs only ever reached the CI build log.
#
# These ARGs are declared LAST, immediately before the layer that uses
# them, so a changing BUILD_DATE / RELEASE_TAG / SOURCE_REVISION never
# invalidates the expensive pi-install / clone layers above.
ARG RELEASE_TAG=dev
ARG BUILD_DATE=
ARG SOURCE_REVISION=
# MEMPALACE_TOOLKIT_REF is consumed in Dockerfile.base; re-declared here
# only so its intended ref lands in the label set alongside the others.
ARG MEMPALACE_TOOLKIT_REF=main
# ── Vendored skill provenance ─────────────────────────────────────────
# The vendored mempalace SKILL.md is the ONLY baked artefact with no /opt
# clone behind it: its upstream (the skillset repo) is PRIVATE, so the
# image cannot clone it and CI cannot resolve its HEAD (see VENDORED.md).
# Consequence through v1.8.7: the snapshot was ANONYMOUS — nothing in the
# image or the repo recorded which skillset commit it was taken from, so
# the only staleness check available was a hand-maintained phrase canary in
# scripts/smoke-test.sh, which by construction can only detect "older than
# what I remembered to pin", never "older than skillset main".
#
# Recording the ref costs nothing and makes the question answerable. It is
# deliberately a plain ARG DEFAULT rather than a CI-resolved output:
# * the value is a fact about the committed snapshot, so it belongs in
# the tree next to it — not in a workflow that a local `docker build`
# never runs (same reasoning as MEMPALACE_VERSION living in
# Dockerfile.base rather than being duplicated in docker-publish.yml);
# * CI therefore needs NO new build-arg at any of its four
# Dockerfile.variant call sites (smoke, smoke-studio, build-variant,
# build-variant-studio) — a plumbing change that is easy to
# under-apply to only two of them;
# * and it needs no credential for a private repo.
# Bump it with scripts/vendor-mempalace-skill.sh, which refreshes the file
# and rewrites this line together, so the pair cannot drift apart by hand.
# This ARG lives in Dockerfile.variant ON PURPOSE: Dockerfile.base and
# rootfs/ are both hashed into base_tag, so recording provenance here costs
# no ~67-minute base rebuild. (scripts/check-base-hash.sh scans only
# Dockerfile.base, so no folding into the base hash is required — nor would
# it be correct, since this ARG changes nothing about the base's contents.)
ARG SKILLSET_SNAPSHOT_REF=6eb20af181f0147cb8c1377f6e36a6a47a68e8e5
# Dockerfile.base sets description="pi-devbox — base image (variant-independent)"
# and every variant INHERITS it, so both published images used to advertise
# themselves on Docker Hub as the base image. A LABEL cannot branch on
# INSTALL_STUDIO, so the description arrives as a build-arg: CI passes the
# variant-specific string (see docker-publish.yml), and the default below keeps
# a plain `docker build -f Dockerfile.variant` honest rather than misleading.
ARG IMAGE_TITLE="pi-devbox"
ARG IMAGE_DESCRIPTION="pi-devbox — development container for the pi coding agent"
LABEL org.opencontainers.image.version="${RELEASE_TAG}" \
org.opencontainers.image.revision="${SOURCE_REVISION}" \
org.opencontainers.image.created="${BUILD_DATE}" \
org.opencontainers.image.title="${IMAGE_TITLE}" \
org.opencontainers.image.description="${IMAGE_DESCRIPTION}" \
description="${IMAGE_DESCRIPTION}" \
se.jordbo.pi-devbox.pi-version="${PI_VERSION}" \
se.jordbo.pi-devbox.pi-toolkit-ref="${PI_TOOLKIT_REF}" \
se.jordbo.pi-devbox.pi-extensions-ref="${PI_EXTENSIONS_REF}" \
se.jordbo.pi-devbox.pi-fork-ref="${PI_FORK_REF}" \
se.jordbo.pi-devbox.pi-obsmem-ref="${PI_OBSMEM_REF}" \
se.jordbo.pi-devbox.pi-atelier-ref="${PI_ATELIER_REF}" \
se.jordbo.pi-devbox.pi-atelier-version="${PI_ATELIER_VERSION}" \
se.jordbo.pi-devbox.mempalace-toolkit-ref="${MEMPALACE_TOOLKIT_REF}" \
se.jordbo.pi-devbox.pi-studio-ref="${PI_STUDIO_REF}" \
se.jordbo.pi-devbox.pi-studio-version="${PI_STUDIO_VERSION}" \
se.jordbo.pi-devbox.skillset-snapshot-ref="${SKILLSET_SNAPSHOT_REF}"
# The manifest is written from GROUND TRUTH — the actual checked-out HEAD
# of each /opt clone and the live `pi --version` — not merely the intended
# build-args. That way it also exposes a clone that silently resolved to
# something other than the requested ref. pi-studio is present only in the
# studio variant (JSON null otherwise).
RUN set -e; \
mkdir -p /etc/pi-devbox; \
rev() { git -C "$1" rev-parse HEAD 2>/dev/null || echo "unknown"; }; \
PI_V="$(pi --version 2>/dev/null | head -n1 | tr -d '\r\n')"; \
# mempalace CORE (the PyPI package behind the MCP tools) is installed in
# Dockerfile.base via `uv tool install`, so no /opt clone reveals it and
# until v1.8.6 the manifest could not answer "which palace shipped here?" —
# a palace bug could not be correlated to an image, which is precisely the
# correlation this file exists to provide. Read from the INSTALLED BINARY,
# not from ARG MEMPALACE_VERSION, per the ground-truth rule above: that is
# what catches an install which resolved to something other than the pin.
# `mempalace --version` prints "MemPalace 3.7.1" — NAME-PREFIXED, unlike
# pi's bare "0.84.2" — hence the $NF pick rather than a straight read. The
# leading-digit test then rejects usage/error text (a renamed flag prints a
# usage block) and degrades to JSON null, so this can never fail the build.
MP_V="$(mempalace --version 2>/dev/null | head -n1 | tr -d '\r' | awk '{print $NF}')"; \
case "$MP_V" in [0-9]*) MP_CORE="\"${MP_V}\"" ;; *) MP_CORE='null' ;; esac; \
STUDIO_REV='null'; \
if [ -d /opt/pi-studio/.git ]; then STUDIO_REV="\"$(rev /opt/pi-studio)\""; fi; \
# The vendored skill snapshot's fingerprint is MEASURED here, not passed
# in as a build-arg, per the ground-truth rule above: SKILLSET_SNAPSHOT_REF
# is a CLAIM about which skillset commit the file came from, while this
# hash is what the image actually ships. Recorded together they let any
# reader with the skillset checked out — which on this fleet is every
# host, since all four compose stacks mount it — verify the claim at
# RUNTIME, without CI ever needing access to the private repo. Degrades
# to JSON null rather than failing the build if the directory is absent;
# the smoke assertion is what turns that into a loud failure.
#
# Hashes the whole DIRECTORY, not just SKILL.md: a single-file hash
# answers "did this one file change", not "is the live copy the same
# skill" — a live checkout that added or edited a SIBLING file (a
# reference/ doc, a helper script) would still report "identical to
# baked snapshot" against a file-only hash. pi-extensions already ships
# two files for exactly this reason (SKILL.md + evaluate-extension-usage.py),
# so this is not a hypothetical. Deterministic over `find | sort`, never
# readdir order: relative paths + per-file sha256, folded into one hash.
# pi-devbox-version mirrors this exact pipeline over the live directory so
# the two sides are comparable — if you change this, change that too.
tree_sha256() { \
( cd "$1" && find . -type f -print | LC_ALL=C sort | xargs -r sha256sum ) 2>/dev/null | sha256sum | cut -d' ' -f1; \
}; \
SKILL_SNAP='null'; \
_snap_dir=/usr/local/share/pi-devbox/skills/mempalace; \
if [ -d "$_snap_dir" ] && [ -n "$(find "$_snap_dir" -type f -print -quit)" ]; then \
SKILL_SNAP="\"$(tree_sha256 "$_snap_dir")\""; \
fi; \
{ \
echo '{'; \
echo " \"release_tag\": \"${RELEASE_TAG}\","; \
echo " \"build_date\": \"${BUILD_DATE}\","; \
echo " \"source_revision\": \"${SOURCE_REVISION}\","; \
echo " \"pi_version\": \"${PI_V}\","; \
# Sibling of pi_version, NOT a member of components{}: that map holds git
# SHAs and `pi-devbox-version` renders it with .value[0:12], which would
# silently truncate a longer version string.
echo " \"mempalace_version\": ${MP_CORE},"; \
# Siblings, NOT members of components{}, for two independent reasons:
# that map means "HEAD of a clone present in this image" and the
# skillset is not cloned here (calling it a component would be a
# lie a future reader would act on), and `pi-devbox-version` renders
# every components{} value with .value[0:12] — which would truncate
# a 64-hex sha256 into something that looks like a short commit.
# Named `_tree_sha256`, not `_sha256`: it measures every file under the
# vendored skill directory, not one file — see tree_sha256() above.
echo " \"skillset_snapshot_ref\": \"${SKILLSET_SNAPSHOT_REF}\","; \
echo " \"skillset_snapshot_tree_sha256\": ${SKILL_SNAP},"; \
echo " \"components\": {"; \
echo " \"pi-toolkit\": \"$(rev /opt/pi-toolkit)\","; \
echo " \"pi-extensions\": \"$(rev /opt/pi-extensions)\","; \
echo " \"pi-fork\": \"$(rev /opt/pi-fork)\","; \
echo " \"pi-observational-memory\": \"$(rev /opt/pi-observational-memory)\","; \
echo " \"pi-atelier\": \"$(rev /opt/pi-atelier)\","; \
echo " \"mempalace-toolkit\": \"$(rev /opt/mempalace-toolkit)\","; \
echo " \"pi-studio\": ${STUDIO_REV}"; \
echo " }"; \
echo '}'; \
} > /etc/pi-devbox/build-manifest.json; \
echo "── build manifest ──"; cat /etc/pi-devbox/build-manifest.json
# WORKDIR / ENTRYPOINT / CMD inherited from base.
+60
View File
@@ -0,0 +1,60 @@
# Ideas & backlog
A living list of potential improvements for pi-devbox that are **not yet
scheduled**. This is intentionally lightweight — a place to park ideas so they
aren't lost between sessions. When an item ships, describe it in
[`CHANGELOG.md`](CHANGELOG.md) and remove it from here.
Rough effort tags: 🟢 small · 🟡 medium · 🔴 large. Status: `idea` (unvetted) ·
`planned` (agreed, not started).
---
## Supply-chain hardening
- 🟡 `planned` — **Pin CI actions to commit SHAs.** The workflows use floating
major tags (`actions/checkout@v4`, `docker/build-push-action@v7`,
`docker/setup-buildx-action@v4`, `docker/login-action@v3`,
`docker/setup-qemu-action@v3`). This is inconsistent with the project's own
philosophy of SHA-pinning *content* refs (pi, pi-studio, pi-fork, …) to defeat
floating refs. Pin each action to a SHA with a trailing `# vX.Y.Z` comment.
Pairs naturally with the renovate item below to keep the pins fresh.
- 🟡 `planned` — **Vulnerability scanning in CI.** No CVE scan runs on the
published images today. Add a `trivy image` (or grype) job to
`docker-publish.yml` after `smoke`. Start non-blocking (report only), then
tighten to fail on `HIGH`/`CRITICAL` with an available fix.
- 🟢🟡 `planned` — **Standardize build provenance → buildx SBOM + attestations.**
The image already carries hand-rolled provenance (OCI labels +
`build-manifest`). `docker/build-push-action` can emit a standard SBOM and
SLSA provenance attestation nearly for free (`provenance: mode=max`,
`sbom: true`). Makes provenance machine-consumable and pairs well with the
trivy item (scan the SBOM).
## Dockerfile hardening
- 🟡 `idea` — **Address hadolint DL4006 properly.** Currently ignored in
`.hadolint.yaml`. The clean fix is `SHELL ["/bin/bash", "-o", "pipefail",
"-c"]` so piped `RUN`s fail on the first non-zero stage. This changes the
default `RUN` shell from `sh` to `bash` for all subsequent layers, so it is
base-affecting and needs a careful pass over existing `RUN`s before removing
the ignore.
## Developer experience
- 🟢 `idea` — **`Makefile`/`justfile` for local iteration.** Reproducing a CI
build locally means hand-assembling many `--build-arg`s. Thin targets
(`make build-base`, `make build-variant`, `make smoke`, `make lint`) would
make local testing painless and document the canonical invocations.
- 🟡 `idea` — **Dependency-update automation (renovate).** With CI actions
SHA-pinned (above), a `renovate.json` keeps those pins — plus the pinned tool
versions (`ACTIONLINT_VERSION`, `HADOLINT_VERSION`, gosu, etc.) — current via
automated PRs. Requires a renovate runner against the Gitea instance.
## Housekeeping
- 🟢 `idea` — **Registry retention for `base-<hash>` tags.** The base-hash
caching scheme accumulates `base-<hash>` tags over time. Confirm whether the
registry prunes old ones, and add a retention/cleanup step if not.
+21
View File
@@ -0,0 +1,21 @@
MIT License
Copyright (c) 2026 Joakim Persson
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
+605 -16
View File
@@ -20,7 +20,11 @@ on the host.
- `pi-extensions` — TypeScript extensions for pi (preview, MCP bridges,
mempalace integration, etc.)
- `pi-fork` — the `fork` tool for spawning sub-agents
- `pi-observational-memory` — the `recall` tool for session compaction
- `pi-observational-memory` — durable session memory: the ledger that makes
compaction cheap, plus the `recall` tool. See
[`docs/observational-memory.md`](docs/observational-memory.md)
- `pi-atelier` — TUI sidebar: ordered panels, split-pane, themes. Pinned to an
audited tag; see [Version pins](#version-pins-pi-pi-atelier-mempalace)
### MemPalace (AI memory)
@@ -29,13 +33,18 @@ on the host.
- ChromaDB embedding model pre-warmed at build time (`all-MiniLM-L6-v2`)
The host-mounted palace at `~/.mempalace` is shared across the host and
this container so all your agents share one brain.
this container so all your agents share one brain. To instead share a palace
across *several* containers/harnesses, set `MEMPALACE_REMOTE_URL` to a shared
MemPalace HTTP endpoint (see `.env.example` and `docker-compose.mempalace.yml`);
the bridge then connects over HTTP and spawns no local server.
### Modern CLI tooling
| Tool | Purpose |
|---|---|
| `nvim` | Neovim text editor |
| `nvim` | Neovim text editor (modal / vi-style) |
| `nano` | Small non-modal editor (on-screen shortcut hints) |
| `micro` | Modern non-modal editor (Ctrl+S/Ctrl+Q keys, mouse, syntax highlighting) |
| `tmux` | Terminal multiplexer (configured for 0-indexed sessions) |
| `ripgrep`, `fd` | Fast file content / filename search |
| `fzf` | Fuzzy finder |
@@ -49,12 +58,41 @@ this container so all your agents share one brain.
| `gosu` | Privilege de-escalation in entrypoint |
| `htop`, `tree`, `less` | Inspection utilities |
The default `$EDITOR` is `nvim`. Three editors ship so you can pick your
comfort level — if you'd rather not use a vi-style editor, `nano` and `micro`
are both non-modal. Set your preference with `export EDITOR=micro` (or `nano`)
in your shell profile, and/or `git config --global core.editor micro`.
Neovim ships with a system-wide default (`/etc/xdg/nvim/sysinit.vim`) that turns
on `termguicolors`, so its colours render in 24-bit instead of a muddy
256-colour fallback over ssh/kitty. The `kitty-terminfo` entry is also bundled
so `TERM=xterm-kitty` is understood. Override either in your own
`~/.config/nvim`.
### Document and image tooling
- `pandoc` — universal Markdown↔HTML/Org/RST/etc. converter
- `typst` — markup-based typesetting, wired up as pandoc's `--pdf-engine` (see
[Generating a PDF with pandoc + typst](#generating-a-pdf-with-pandoc--typst))
- `graphviz` — `dot` rendering for diagram pipelines
- `imagemagick` — image conversion / resizing (invoked as `magick`)
### Browser automation
- `agent-browser` — CLI for driving a real headless browser: open pages,
click/fill/`eval`, snapshot the DOM, take screenshots. Useful whenever a task
involves a web UI or verifying how a page actually renders (live DOM, WebGL,
layout, popup positioning) instead of guessing from source.
- `playwright` + a pre-installed headless **Chromium** back it.
`AGENT_BROWSER_EXECUTABLE_PATH` is preset to the baked browser via a stable
`/usr/local/bin/agent-chrome` symlink (insulated from Playwright's
per-version/arch install directory), so `agent-browser open <url>` works
out of the box with no setup. Run `agent-browser skills get core --full`
for the command set and workflow patterns.
- `socat` — TCP bridge used by `studio-expose` to reach pi-studio's
loopback-bound server from outside the container (see
[Using pi-studio](#using-pi-studio--studio-variant))
### Language toolchains
- `python3` + `python3-venv` + `python3-pip` (system Python)
@@ -82,6 +120,9 @@ For Python REPLs and notebooks beyond the system interpreter, see the
- A LAN-access helper that auto-configures ssh jump-via-host on
VM-backed hosts (OrbStack / Docker Desktop on macOS) so the container
can reach the host's directly-attached LAN peers.
- Read-only `~/.ssh` is handled transparently: a per-host `ControlPath`
under it (common CGNAT configs like `~/.ssh/cm/...`) is redirected to a
writable socket dir for both `pi --ssh` and `dssh`/`dscp`.
## Quickstart
@@ -136,8 +177,10 @@ Currently published:
Planned for an upcoming minor release:
- `joakimp/pi-devbox:latest-studio-tex` — `-studio` plus `texlive-xetex`
for PDF export from Studio. Adds ~600 MB on top of `-studio`.
- *(shipped in Unreleased/base)* **PDF export from Studio/pandoc** now works:
the base image ships **`typst`** as the PDF engine (`pandoc --pdf-engine=typst`),
a single ~30 MB static binary — no separate `-tex` variant needed.
`texlive-xetex` stays the higher-fidelity fallback (install on demand).
## Using pi-studio (`-studio` variant)
@@ -272,9 +315,30 @@ Assuming the compose file publishes `127.0.0.1:8765:8765` (see method B):
> until step 2 runs. If the browser can't connect, verify Studio is up
> (`/studio --status`) and the bridge is running (`ps aux | grep socat`).
> PDF export (`/studio-pdf`, `studio_export_pdf`) needs a LaTeX engine,
> which is **not** in `-studio` (only the planned `-studio-tex`). HTML
> export, KaTeX, Mermaid, and all REPL features work without it.
> PDF export (`/studio-pdf`, `studio_export_pdf`) uses **`typst`**, shipped in
> the base image as the pandoc PDF engine (`pandoc --pdf-engine=typst`). For
> LaTeX-exact output you can install `texlive-xetex` on demand as a heavier
> fallback. HTML export, KaTeX, Mermaid, and all REPL features work regardless.
### Generating a PDF with pandoc + typst
The base ships `pandoc` (front-end) and `typst` (PDF engine), so Markdown → PDF
works out of the box:
```bash
pandoc doc.md --pdf-engine=typst -o doc.pdf
```
The base patches pandoc's bundled typst template so it defaults to the
**Libertinus Serif** font. Without that patch a naked `--pdf-engine=typst`
fails with `error: font fallback list must not be empty`, because the upstream
template leaves the font unset. To pick a different face, pass one of the fonts
typst can see (`typst fonts` lists them — DejaVu Serif/Sans/Mono, Libertinus
Serif, New Computer Modern):
```bash
pandoc doc.md --pdf-engine=typst -V mainfont="New Computer Modern" -o doc.pdf
```
### Graphviz diagrams in Studio: `dot-watch`
@@ -296,6 +360,59 @@ DOT syntax errors instead of crashing. Then in Studio: open the PNG (or a
`.md` that embeds it) and hit **refresh-from-disk** after each edit.
Note: SVG is **not** in Studio's local-image-link allowlist — use PNG.
## Using pi-atelier (TUI sidebar)
`pi-atelier` is bundled in **both** variants (vendored at `/opt/pi-atelier`,
pinned — see [Version pins](#version-pins-pi-pi-atelier-mempalace)). It adds two
things to pi's terminal UI:
- a **status rail** — activity, token/cost metrics, context usage, model, git
state, extension statuses, and a menu;
- a **sidebar** — ordered panels (agent, activity, alerts, TODOs, context,
workspace, usage, tools) in a split pane beside the transcript.
Nothing needs installing; the entrypoint registers it on container start, and it
binds on the next pi start (or `/reload`).
| Action | How |
|---|---|
| Open the atelier menu | `alt+a`, or `/atelier` |
| Toggle the sidebar for this session | `/atelier sidebar on` / `off` |
| Change settings persistently | atelier menu → **Settings**, then **Save** |
| Turn the whole thing off | `DEVBOX_ATELIER=0` in `.env` |
If your terminal or keymap swallows `alt+a`, use `/atelier` and pick a different
`shortcut` in the config file below.
### Config
Config lives at `~/.pi/agent/pi-atelier.json` on the `devbox-pi-config` volume,
seeded from pi-toolkit with container-appropriate defaults: compact density,
context warnings at 60/85 % (earlier than upstream's 70/90), sidebar tool names
on, and desktop completion notifications **off** (a container has nowhere useful
to pop a toast).
It is **copied, not symlinked** — atelier rewrites this exact path when you hit
**Save**, using write-temp-then-`rename(2)`, and `rename` replaces a symlink with
a regular file instead of following it. A symlink would silently detach on your
first save. Consequently pi-toolkit's `install.sh` only seeds the file when it is
absent: once you have saved your own preferences, image upgrades leave them
alone, and `install.sh` prints a diff hint instead of clobbering.
The seeded file uses atelier's **current** schema — `segmentLayout` with explicit
per-segment visibility, plus `showSidebarAgent` / `showSidebarTodos` /
`showSidebarOnStartup`. Older configs written against the pre-0.7 vocabulary
(`segments`, `ornament`, `showExtensionStatuses`) still load, but only through
upstream's legacy-compatibility shims — so if you are carrying one on an old
volume, expect it to keep working while missing every sidebar control added
since. `sidebarPanelLayout` is deliberately left unset so the panel set follows
upstream's product default as atelier adds panels; set it only if you want to
pin the order yourself.
The sidebar auto-hides below 92 terminal columns and keeps the main pane at
least 64 columns wide, so a narrow terminal degrades to the plain TUI rather
than a squeezed one.
## docker-compose.yml — basic shape
```yaml
@@ -316,9 +433,10 @@ services:
environment:
- TERM=xterm-256color
# - STUDIO_EXPOSE=1 # -studio only: auto-start the socat bridge on boot
- GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}
- GITEA_HOST=${GITEA_HOST:-}
- GITHUB_PERSONAL_ACCESS_TOKEN=${GITHUB_PERSONAL_ACCESS_TOKEN:-}
# Secrets (GITEA_*, GITHUB_*, …) come from env_file: .env above — not
# duplicated here. An environment: entry overrides env_file and is
# interpolated from the host shell, so a stale shell export would
# silently shadow your .env. See .env.example for the full list.
volumes:
# Workspace: your host source tree
- ${WORKSPACE_PATH:-.}:/workspace
@@ -439,6 +557,138 @@ session/docs mining; the 29 MCP tools (search, kg-query, drawer-add,
diary-write, etc.) are wired into pi automatically by the pi-extensions
mempalace bridge.
### Cross-machine agent coordination
When `MEMPALACE_REMOTE_URL` points at a *shared* palace, the container gets more
than shared search: it joins an append-only coordination log (RFC 003) that other
machines' agents can address it on — used here for design review, patch handoff
and retraction between hosts.
Two container-side settings make it work:
| Variable | Why it matters |
|---|---|
| `MEMPALACE_REMOTE_URL` | selects the shared palace; unset means a purely local palace, and the log then contains only this machine's own events |
| `MEMPALACE_PI_DEVICE` | the bridge stamps `pi@<device>` as the writer, which is the **only** way the log can tell two machines apart when both are thin clients of one palace |
So a container with no `MEMPALACE_PI_DEVICE` can read the log but is not
reachable *on* it: messages addressed to a bare `pi` match nobody. Set both, or
neither.
What the agent is expected to *do* with this lives in the mempalace skill
(`~/.agents/skills/mempalace/SKILL.md`) — the mailbox query at wake-up, and the
convention that a directed event with `status="open"` is a request owed a reply
while a `*` broadcast owes nothing. The mechanism side (what the bridge stamps,
and why live SSE push depends on the palace deployment's reverse proxy rather
than on this image) is documented in the toolkit's `extensions/pi/README.md`.
**Since v1.8.9 the bridge reads the log for you.** Earlier images were write-only
— they stamped provenance on the way out and never read back, so a directed ask
reached an agent only if that agent happened to run `mempalace_event_list`
itself. The mailbox is gated on the same two variables as the stamper, is on by
default, and derives what is *owed* rather than trusting `status` (an acked event
keeps matching a `status="open"` query forever, because the log is append-only):
| Variable | Default | Effect |
|---|---|---|
| `MEMPALACE_MAILBOX` | unset (on) | `0` disables mailbox reads entirely |
| `MEMPALACE_MAILBOX_POLL_MS` | `300000` | minimum gap between mid-session polls |
| `MEMPALACE_MAILBOX_RESURFACE_MS` | `3600000` | re-announce a still-owed ask after this long |
Delivery **queues, it never interrupts**: the poll runs when pi goes idle and the
message is steered into the *next* turn, so nothing wakes the model on inbound
fleet traffic. The practical consequence, measured on two devices: the message
appears in your session window and the agent acts on it when the next turn
starts — you are the trigger. (That describes the bridge **as baked in v1.8.9**,
`mempalace-toolkit` `5b8d78f`; the mailbox's own mechanism and landmines live in
the toolkit's `docs/rfc-003-coordination-log.md` §7.11–§7.12, which moves ahead of
whatever this image has baked.)
## Observational memory (in-session memory)
The image also bakes [pi-observational-memory](https://github.com/elpapi42/pi-observational-memory),
which is memory of a *different kind* from the palace and is easy to confuse with
it. It keeps a small branch-local ledger of observations and reflections while a
session runs, so when pi compacts the conversation the summary is a
**deterministic fold of that ledger rather than a model call**, and every item
keeps a 12-character id that `recall(<id>)` resolves back to the exact source.
In one line: **observational memory keeps a session coherent; the palace keeps
the fleet coherent.**
It is on by default, needs no habit from you, and sends its background work to a
cheaper model than your session (Haiku while the session runs Opus, in the seeded
`~/.pi/agent/settings.json`). Inspect it from inside pi with `/om:status` and
`/om:view`; turn all proactive work off for one run with
`PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi`.
What it is for, how the lifecycle works, what it costs, every setting and its
default, and how it differs from MemPalace:
[`docs/observational-memory.md`](docs/observational-memory.md).
## Agent skills
pi discovers skills under `~/.agents/skills/`. Two delivery paths feed that
directory, and they compose:
- **Image-baked skills (always present).** Skills shipped *inside* the image
live under `/usr/local/share/pi-devbox/skills/` and are symlinked into
`~/.agents/skills/` by `entrypoint-user.sh` on every start. They need no
external mount, survive volume recreate (the source is an image path, not a
home dir a named volume would shadow), and are created only when absent so a
user override is never clobbered. Precedence against a mounted `skillset` repo
is per-skill, not blanket — see *Skillset repo* below. The bundled
**`pi-devbox-environment`** skill is delivered this way — it teaches agents
the container's persistence model, host/LAN SSH reachability, split-DNS
mechanisms, the interactive-vs-tool-shell alias gotcha (`dssh`/`dscp`),
tmux 0-indexing, uv-first Python, and pi-studio reachability, all as
*mechanisms* (deployment-specific hostnames/domains/nameservers are
discovered at runtime, never hardcoded).
- **Vendored fallback skills.** The pi-toolkit global `AGENTS.md` tells every
pi session to read `~/.agents/skills/pi-extensions/SKILL.md` at start (to fix
fork/recall under-utilisation). That pointer would dangle in a container
started *without* the private `skillset` repo, so the image also bakes
fallback copies of **`pi-extensions`** and **`mempalace`**. Whether a mounted
skillset overrides them depends on who *owns* the skill (see *Skillset repo*):
`mempalace` is skillset-owned, so the live clone wins; `pi-extensions` is
owned by its package repo, so the baked copy keeps winning — the skillset's
copy of it is a downstream duplicate that can lag. The
`pi-extensions` skill is *layered*: a committed snapshot in `rootfs/` is the
floor, and `Dockerfile.variant` copies the canonical, package-owned copy from
the pinned `pi-extensions` clone (`/opt/pi-extensions/skill/`) over it at
build, so a normal build ships the fresh copy and an old-ref/mirror build
still ships the snapshot. `mempalace` is snapshot-only (its consumer skill
has no public package home), and because pi-toolkit's `AGENTS.md` has no
directive for it, the pi-devbox managed block adds a session-start
*proactive-load* pointer for it (gated to pi-devbox containers, conditional
on the MemPalace MCP tools) so a new container actually loads it. See
`rootfs/usr/local/share/pi-devbox/skills/VENDORED.md`.
- **Skillset repo (optional).** If a `skillset` repo is mounted (at
`$HOME/skillset` or `/workspace/skillset`, or via `SKILLSET_CONTAINER_PATH`),
`deploy-skills.sh` symlinks its skills in too. Image-baked skills are
classified as foreign-links by its `--prune-stale` pass and left untouched —
which through v1.8.4 meant the baked copy *always* won, so an edit pushed to a
skillset-owned skill was invisible until the next image build. Since v1.8.5
`devbox-skill-reconcile` runs right after the deploy and repoints the links for
skills the skillset owns, listed in
`/usr/local/share/pi-devbox/skills/skillset-owned.txt` (today: `mempalace`).
Effective precedence, highest first: **user override** (a real directory, or a
symlink pointing outside the baked tree) → **live skillset clone** (owned names
only) → **baked snapshot** (everything else, and every skill when no skillset
is mounted). Check with `readlink -f ~/.agents/skills/<skill>`.
To make agents *proactively* load a baked skill at session start (rather than
only on description match), the image appends a short, gated pointer to the
global `AGENTS.md` at build time (see `pi-global-AGENTS.append.md`). The
pointer fires only inside a pi-devbox container (it checks for
`/usr/local/lib/pi-devbox/`).
To add another image-baked skill: drop a `SKILL.md` under
`rootfs/usr/local/share/pi-devbox/skills/<name>/`; the `COPY` in
`Dockerfile.base` and the entrypoint symlink loop pick it up automatically. To
refresh a vendored fallback, see
`rootfs/usr/local/share/pi-devbox/skills/VENDORED.md`.
## SSH and ControlMaster
The base image preconfigures `Host *` ssh defaults:
@@ -461,6 +711,59 @@ User-level overrides in `~/.ssh/config` win because Debian's
`/etc/ssh/ssh_config` includes `/etc/ssh/ssh_config.d/*.conf` before
the `Host *` block.
### macOS-only keywords in a shared `~/.ssh/config`
The same `~/.ssh/config` is read by macOS ssh *and* by the Linux OpenSSH inside
the container (the sidecar `Include`s it). macOS-only keywords are **fatal**
there, not ignored — a single `UseKeychain yes` in a `Host *` block takes down
every ssh call in the container:
```
/home/developer/.ssh/config: line 2: Bad configuration option: usekeychain
/home/developer/.ssh/config: terminating, 1 bad configuration options
```
That breaks `dssh`/`dscp`, `pi --ssh`, `scp`, and anything that shells out to
ssh (including CI/deploy helpers), while the host keeps working perfectly — so
it presents as a container regression rather than a host config error. Guard the
keyword on the host, *before* it is used:
```diff
Host *
+ IgnoreUnknown UseKeychain
UseKeychain yes
AddKeysToAgent yes
```
`IgnoreUnknown` is understood by both implementations: macOS still honours
`UseKeychain`, Linux skips it. Also keep such a `Host *` block **below** any
`Include` that must come first — OrbStack's own `Include ~/.orbstack/ssh/config`
says so in a comment, and a `Host *` block above it silently violates that.
### Per-host `ControlPath` on a read-only `~/.ssh`
`~/.ssh` is usually bind-mounted read-only, so a user `~/.ssh/config` that
points `ControlPath` back under it (e.g. the CGNAT idiom
`ControlPath ~/.ssh/cm/%r@%h:%p`) can't bind its master socket here — and a
system default can never override a user's per-host value. Two layers handle
this without editing the read-only config:
- **`pi --ssh <host>`** — the `ssh-controlmaster` extension detects an
unwritable system `ControlPath` and falls back to its own writable
`/tmp/pi-cm-<pid>.sock` master (its command-line `-o ControlPath` overrides
the user's path); the remote-`pwd` probe uses `-o ControlPath=none` so it
cannot fail on the read-only socket dir.
- **`ssh -F ~/.ssh-local/config` / `dssh` / `dscp`** — `setup-lan-access.sh`
redirects `ControlPath` into the writable `~/.ssh-local/cm` for every host
(the sidecar is rendered on all host OSes). To name LAN peers that should
jump via the host, add `ProxyJump host` overrides in the host-owned
`~/.config/devbox-shell/ssh-lan.conf` (see
[Naming LAN peers](#naming-lan-peers)) rather than the read-only
`~/.ssh/config`. If the peer also rejects the host's key — the usual case,
since host keys are normally passphrase-protected and the container has no
Keychain or agent — see
[Giving the container its own key for a peer](#giving-the-container-its-own-key-for-a-peer).
## tmux and 0-indexed sessions
The image installs `/etc/tmux.conf` with:
@@ -517,6 +820,98 @@ pi-coding-agent@latest` (the build-arg string would otherwise be
byte-identical across releases and the layer would silently reuse the
previous version's bytes).
### Building a fork / relocated build
The canonical build clones its companions from `gitea.jordbo.se`. Every
companion repo URL is an overridable build-arg (defaulting to the canonical
origin), so a fork or a build on a host that can't reach that gitea can
repoint each one at a mirror, another host, or a local `file://` path
**without editing the Dockerfiles**:
| Build-arg | Default | Dockerfile |
|---|---|---|
| `PI_TOOLKIT_REPO` | `https://gitea.jordbo.se/joakimp/pi-toolkit.git` | variant |
| `PI_EXTENSIONS_REPO` | `https://gitea.jordbo.se/joakimp/pi-extensions.git` | variant |
| `MEMPALACE_TOOLKIT_REPO` | `https://gitea.jordbo.se/joakimp/mempalace-toolkit.git` | base |
| `PI_FORK_REPO` | `https://github.com/elpapi42/pi-fork.git` | variant |
| `PI_OBSMEM_REPO` | `https://github.com/elpapi42/pi-observational-memory.git` | variant |
| `PI_ATELIER_REPO` | `https://github.com/michaelmjhhhh/pi-atelier.git` | variant |
| `PI_STUDIO_REPO` | `https://github.com/omaclaren/pi-studio.git` | variant |
Each has a matching `*_REF` arg (branch name or commit SHA). Example — build
the variant against forked toolkit/extensions and a pinned pi:
```bash
# base first (mempalace-toolkit lives here)
docker build -f Dockerfile.base -t myorg/pi-devbox:base-dev \
--build-arg MEMPALACE_TOOLKIT_REPO=https://github.com/myorg/mempalace-toolkit.git .
# then the variant FROM that base
docker build -f Dockerfile.variant -t myorg/pi-devbox:dev \
--build-arg BASE_IMAGE=myorg/pi-devbox:base-dev \
--build-arg PI_VERSION=0.79.7 \
--build-arg PI_TOOLKIT_REPO=https://github.com/myorg/pi-toolkit.git \
--build-arg PI_EXTENSIONS_REPO=https://github.com/myorg/pi-extensions.git .
```
Note: the gitea companions clone anonymously (no token needed); only the
`resolve-versions` CI job calls the gitea *API* (which needs a token even
for public repos). A plain `docker build` like the above skips that job
entirely, so no credentials are required for a local/forked build.
Provenance build-args (all optional; populate the OCI labels and
`/etc/pi-devbox/build-manifest.json` — see below): `RELEASE_TAG`,
`BUILD_DATE`, `SOURCE_REVISION`. CI sets these automatically; a manual build
leaves them at harmless defaults.
### Build provenance (labels + manifest)
Every published image is self-describing. Inspect the OCI labels without
pulling the filesystem:
```bash
docker inspect --format '{{json .Config.Labels}}' joakimp/pi-devbox:latest | jq .
```
`org.opencontainers.image.{version,revision,created}` plus
`se.jordbo.pi-devbox.*-ref` record the intended pi version and companion
refs. The on-disk `/etc/pi-devbox/build-manifest.json` records **ground
truth** — the actual checked-out commit of each `/opt` clone, the live
`pi --version`, and (from v1.8.6) the live `mempalace --version` of the
installed palace core — so a tag is reconstructable after CI logs rotate:
```bash
docker run --rm --entrypoint= joakimp/pi-devbox:latest cat /etc/pi-devbox/build-manifest.json
```
Inside a running container, `pi-devbox-version` wraps that manifest into a
human-readable summary — no need to remember the file path or pipe it
through `jq` yourself:
```console
$ pi-devbox-version
pi-devbox v1.5.0
built: 2026-07-13T17:53:16Z (source d68674d11e06)
pi: 0.80.6
components:
pi-toolkit: 9a8f6faeaa08
pi-extensions: 61c98e004e3d
pi-fork: 4a09af4ef527
pi-observational-memory: 27a5195eaf90
mempalace-toolkit: 96699f2a1781
pi-studio: 2ef38ef31cea
```
It also flags **live drift** — if `pi --version` no longer matches what was
baked at build time (e.g. something on a persisted volume shadowed the
image's binary), the `pi:` line calls that out instead of silently trusting
the manifest. `--json` dumps the raw manifest for scripting; `--quiet` gives
a one-line `release_tag (source_revision)` form. It also prints once,
automatically, at container start (from `entrypoint-user.sh`, before the
rest of the setup output) — so you see which build you're in without
asking. Exits 1 with a short notice on images built before this file
existed, rather than failing silently.
## Troubleshooting
### Image grew unexpectedly
@@ -533,6 +928,123 @@ auto-runs on container start and writes `~/.ssh-local/config` with a
ssh-jump-via-host configuration. Set `DEVBOX_LAN_ACCESS=jump` and
`HOST_SSH_USER=<your-mac-user>` in `.env` if auto-detection fails.
#### Naming LAN peers
`DEVBOX_LAN_ACCESS` / `HOST_SSH_USER` only set up the *jump* to the host. To
make a **named** peer route through it — so `pi --ssh alpserv-2`,
`dssh alpserv-2`, etc. resolve the ProxyJump — add a `ProxyJump host` override
for it in the host-owned, bind-mounted `~/.config/devbox-shell/ssh-lan.conf`
(**not** `~/.ssh/config`, which is mounted read-only):
```
Host pve pve-2 alpserv-2 lagret
ProxyJump host
```
Any option can be set here, not just `ProxyJump`: the file is `Include`d
*before* `~/.ssh/config` and ssh takes the **first** value it sees for each
option, so whatever you put here wins while everything you omit is inherited
from the matching block in your real `~/.ssh/config`. Peer names stay out of the
published image (they are a fact about your LAN, not the image). Alternatively,
set `DEVBOX_LAN_AUTOJUMP_PRIVATE=1` to ProxyJump *any* RFC1918 address through
the host without naming peers (see `.env.example`).
Once the file exists it is re-read on every connection, so *edits* take effect
immediately — no container or session restart. **Creating it for the first time
does need one restart**, because `setup-lan-access.sh` only emits the
`Include ~/.config/devbox-shell/ssh-lan.conf` line when the file is already
readable at container start (`if [ -r "$SSH_LAN_CONF" ]`). Until then ssh never
looks at it — which reads exactly like "my override is being ignored".
#### Giving the container its own key for a peer
`ProxyJump` fixes *routing*; it does not fix *authentication*, and inheriting
the host's `IdentityFile` usually fails inside the container:
- Host keys are commonly passphrase-protected, and that passphrase is unlocked
by the macOS Keychain or a running `ssh-agent`. The container has neither, so
the key can never be decrypted — `Permission denied (publickey)` even though
the identical `ssh peer` works in a host terminal.
- `~/.ssh` is mounted read-only, so you can neither drop a container-usable key
in there nor edit `~/.ssh/config` from inside.
The answer is a **container-only keypair** in `~/.ssh-local/` — the named volume
`devbox-ssh-local`, so it survives `docker compose up -d --force-recreate` —
plus an `IdentityFile` override in the host-owned `ssh-lan.conf`. Note that
nothing is baked into the *published image*: that volume is created on your
machine at runtime, so no private key ever ships to Docker Hub, and a fresh pull
elsewhere generates its own. (Every key below is a throwaway example.)
**1. In the container** — generate a passphraseless key (there is no agent to
unlock a protected one):
```bash
ssh-keygen -t ed25519 -N '' -C "devbox-$(hostname)" \
-f ~/.ssh-local/mypeer_devbox_ed25519
cat ~/.ssh-local/mypeer_devbox_ed25519.pub
# ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIEXAMPLE0000EXAMPLE0000EXAMPLE0000ex devbox-0d11ec7731c7
```
**2. On the peer** — append that public key to `~/.ssh/authorized_keys` **of
the account you will log in as** (the `User` from step 3), narrowly authorized
rather than bare:
```bash
mkdir -p ~/.ssh && chmod 700 ~/.ssh
cat >> ~/.ssh/authorized_keys <<'KEY'
from="192.168.1.0/24,192.168.4.0/24,10.8.0.7",restrict ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIEXAMPLE0000EXAMPLE0000EXAMPLE0000ex devbox-mymachine
KEY
chmod 600 ~/.ssh/authorized_keys
```
Both lines are safe on a peer that is already set up: `mkdir -p` is a no-op
when the directory exists, the `chmod`s only tighten, and appending never
touches keys already listed. Use `>>`, never `>` — one stray truncation
revokes every other key on that account. The options prefix must sit on the
**same physical line** as the key, comma-separated with no spaces: a paste
that wrapped is the likeliest reason a key that looks right is refused.
`ssh-copy-id` cannot add that prefix, so append by hand (or let it copy the
bare key and edit the line afterwards). If authentication still fails with no
clear reason, suspect permissions — sshd's `StrictModes` silently ignores
`authorized_keys` when the home directory, `~/.ssh` or the file itself is
group- or world-writable, and says why only in the peer's own log
(`journalctl -u ssh`, `/var/log/auth.log`).
`restrict` disables pty, agent/X11 and port forwarding; append
`port-forwarding` and `permitopen="127.0.0.1:<port>"` after it if you need one
specific tunnel. `from=` must list the **host's** addresses, not the
container's: container egress is NAT'd through the host, so the peer sees the
host's LAN address (confirm with `echo $SSH_CLIENT` on first login). List every
network the host roams — e.g. both home WLAN subnets plus its VPN address —
because a `from=` mismatch is indistinguishable from a wrong key in the error
message.
**3. On the host** — point the peer at that key in
`~/.config/devbox-shell/ssh-lan.conf`:
```
Host mypeer mypeer.home.arpa
HostName 192.168.1.142
User myuser
IdentityFile ~/.ssh-local/mypeer_devbox_ed25519
IdentitiesOnly yes
# ProxyJump host # only if the container cannot reach the peer directly
```
That path exists only inside containers, which is why it belongs here rather
than in the shared `~/.ssh/config`.
**4. First time only** — restart the container so the `Include` is emitted (see
above), then verify with the master socket bypassed, so a warm connection cannot
fake a pass:
```bash
ssh -F ~/.ssh-local/config -o ControlPath=none mypeer 'echo $SSH_CLIENT'
```
Use one key per machine (`devbox-mbp`, `devbox-studio`, …) so a single
`authorized_keys` line can be revoked without locking out the others.
### Smoke-testing a local build
```bash
@@ -551,9 +1063,20 @@ persisted volumes survived, and pi runtime wiring is intact:
```bash
./scripts/recreate-sanity-check.sh # auto-detects variant
./scripts/recreate-sanity-check.sh --expected-version 0.79.4 # assert pi version
./scripts/recreate-sanity-check.sh --expected-image-version 1.8.9 # assert the pi-devbox release tag
./scripts/recreate-sanity-check.sh --expected-version 0.84.3 # assert the pi coding agent version
```
Those are **two different versions**, and the flags are not interchangeable:
`--expected-image-version` takes the pi-devbox release tag (`v` optional),
`--expected-version` takes `pi --version`. Hand one the other's value and it
says so by name instead of reporting a mismatch against the wrong component.
With neither flag, both values are read from the image's own build manifest
(`/etc/pi-devbox/build-manifest.json`): the live pi version is asserted against
the one recorded at build time — which catches a stale `pi` in the
`~/.pi/npm-global` volume shadowing the baked one — and the release tag is
reported informationally.
If `cli_utils` is on your PATH, the `pi-devbox-sanity` wrapper runs the same
check by short name and locates the repo automatically (override with
`PI_DEVBOX_REPO=/path/to/pi-devbox`). Like `smoke-test.sh`, this script is
@@ -568,9 +1091,72 @@ pi-devbox follows semver-ish:
- **Minor** — new variants, significant base additions.
- **Patch** — pi version bumps, smaller fixes.
The `pi --version` inside the image is asserted by smoke tests to
match the release tag's pi component, so version drift between the
image and the tag is caught at CI time.
The `pi --version` inside the image is asserted by smoke tests to match the
version CI resolved (since v1.7.0, the pin below), so drift between what was
intended and what actually got baked is caught at CI time rather than on a
user's pull.
### Version pins: pi, pi-atelier, mempalace
Three components are pinned to an exact version **in the repo** instead of being
resolved to `latest` at build time:
| Component | Pin | Where |
|---|---|---|
| pi | `0.84.2` | `ARG PI_VERSION` — `Dockerfile.variant` |
| pi-atelier | `v0.8.2` | `ARG PI_ATELIER_REF` — `Dockerfile.variant` |
| mempalace | `3.7.1` | `ARG MEMPALACE_VERSION` — `Dockerfile.base` |
The objective is **not** to freeze versions. Bumping is routine — usually one
line plus a changelog note. The objective is that adopting a new upstream
version is a deliberate, reviewable act, not a side effect of whatever happened
to be published the morning CI ran. Each of these has already drawn blood:
- **pi** — a minor release can move the private TUI/renderer internals that
pi-atelier wraps, or the session `.jsonl` format `pi-session-repair` parses.
- **pi-atelier** — 0.6.0/0.7.0 hang pi 0.84 **at startup**, burning CPU with no
error (fixed in 0.7.1/0.7.2). Its `peerDependencies` still say `>=0.80.7`, so
nothing in the npm metadata expresses the real floor.
- **mempalace** — an unpinned install once swept in the broken `diary_write` MCP
tool schema of 3.3.x/3.4.0, which is why that pin's comment requires a
tool-schema review before every bump.
CI enforces this rather than trusting it:
- `resolve-versions` reads the pins **out of the Dockerfiles** — single source of
truth, so a local `docker build` and a CI release ship the same versions — and
fails the build if a pin is not concrete, not a semver tag, or not actually
published on npm.
- When npm has a newer pi than the pin, CI emits a `::warning::` naming it. That
warning is the prompt to audit and bump; it never adopts the version.
- `smoke-test.sh` asserts the image's `pi --version` equals the pin, and
separately asserts the pairing rule **pi ≥ 0.84 ⇒ pi-atelier ≥ 0.7.1**, so a
bad combination fails the build instead of publishing a TUI that never starts.
To bump pi: read the upstream CHANGELOG for every intervening version (TUI/theme
API, session format, extension loader, Node engine floor), re-check pi-atelier's
CHANGELOG for the pi version it claims to track, then edit the one `ARG` line and
record what you checked in `CHANGELOG.md`.
#### If you previously hand-installed pi-atelier
A hand-installed `pi install npm:pi-atelier` lands in `~/.pi/npm-global`, which
is on the `devbox-pi-config` **volume** — so it outlives image upgrades and stays
at whatever version you installed, unpinned and unaudited. Since the image now
vendors an audited pi-atelier at `/opt/pi-atelier`, the entrypoint removes a
lingering `npm:pi-atelier` entry from `packages[]` (after backing
`settings.json` up to `settings.json.bak.atelier.<timestamp>`) and registers the
pinned `/opt` copy instead. Nothing else in your settings is touched, and the
npm-global copy itself is left on disk — only the registration changes.
This matters more than it sounds: leaving a 0.6.x npm copy registered alongside
pi 0.84 is precisely the combination that hangs at startup.
To opt out of pi-atelier entirely, set `DEVBOX_ATELIER=0` in `.env`. The
entrypoint then removes any pi-atelier entry from `packages[]` on start. That
switch lives in the entrypoint — not in a pi command — deliberately: this
component's failure mode is "pi will not start", which you cannot repair with
`pi uninstall`.
## Acknowledgements
@@ -585,4 +1171,7 @@ The pi coding-agent itself is [@earendil-works/pi-coding-agent](https://www.npmj
## License
MIT
MIT — see [`LICENSE`](LICENSE). This covers the repository's own contents
(Dockerfiles, entrypoint scripts, `rootfs/` seeds, CI, docs). The published
images bundle third-party software under their own licenses; see
[`THIRD_PARTY.md`](THIRD_PARTY.md).
+61
View File
@@ -0,0 +1,61 @@
# Third-party notices
pi-devbox is distributed under the MIT License (see [`LICENSE`](LICENSE)), which
covers **this repository's own contents** — the Dockerfiles, entrypoint scripts,
`rootfs/` seeds, CI workflows, and docs.
The **published container images** (`joakimp/pi-devbox:*`) additionally *bundle*
third-party software, each of which remains under its own license. This file is
a good-faith summary; the authoritative sources are the upstream projects and,
for OS packages, the per-package copyright files inside the image at
`/usr/share/doc/<package>/copyright`.
## pi and its extensions (installed in the variant layer)
| Component | Upstream | License |
| --- | --- | --- |
| pi (`@earendil-works/pi-coding-agent`) | npm | MIT |
| pi-fork | github.com/elpapi42/pi-fork | MIT |
| pi-observational-memory | github.com/elpapi42/pi-observational-memory | MIT |
| pi-studio *(`-studio` variant only)* | github.com/omaclaren/pi-studio | MIT |
| pi-atelier | github.com/michaelmjhhhh/pi-atelier | MIT |
| pi-toolkit, pi-extensions, mempalace-toolkit | authored by the maintainer (Joakim Persson) | MIT |
## MemPalace (AI memory)
| Component | Upstream | License |
| --- | --- | --- |
| mempalace (core, MCP server) | github.com/MemPalace/mempalace (PyPI: `mempalace`) | MIT — the GitHub repo declares MIT; the PyPI package's own metadata omits a license classifier, so if you need clearance from the package artifact alone, verify against the repo's `LICENSE` file rather than the sdist/wheel metadata |
## Browser automation
| Component | Upstream | License |
| --- | --- | --- |
| agent-browser | github.com/vercel-labs/agent-browser (npm: `agent-browser`) | Apache-2.0 |
| Playwright | github.com/microsoft/playwright (npm: `playwright`) | Apache-2.0 |
| Chromium | chromium.googlesource.com/chromium/src | BSD-3-Clause for Chromium's own code, plus a large set of bundled third-party components each under their own license (see Chromium's own `LICENSE`/`about:credits`). The binary in this image is **not compiled here** — it is the build Playwright downloads for its pinned version ("Chrome for Testing"), installed via `playwright install --with-deps chromium` at `/usr/local/share/ms-playwright/`. Treat Playwright's own distribution terms for that build as authoritative over any summary here. |
## Tooling baked into the base image
| Component | Upstream | License (best effort) |
| --- | --- | --- |
| gosu | github.com/tianon/gosu | Apache-2.0 |
| Node.js | nodejs.org | MIT (bundles components under their own licenses) |
| uv | github.com/astral-sh/uv | Apache-2.0 OR MIT |
| Neovim | neovim.io | Apache-2.0 + Vim license |
| Pandoc | pandoc.org | GPL-2.0-or-later |
| Typst | github.com/typst/typst | Apache-2.0 |
| ripgrep / fd / micro / tealdeer / yq (mikefarah) | respective repos | MIT / Apache-2.0 / Unlicense (varies) |
## Base OS
The image is built `FROM` a Debian base and installs packages via `apt`. Debian
and its packages are distributed under their respective licenses (GPL, LGPL,
MIT, BSD, and others). See each package's copyright file in the image under
`/usr/share/doc/<package>/copyright`.
---
*Licenses marked "best effort" are widely known but were not each verified at
the exact bundled version; consult the upstream project for authoritative
terms. Corrections welcome.*
+111
View File
@@ -0,0 +1,111 @@
# Shared MemPalace server (optional) — one palace for many clients.
#
# Runs `mempalace-mcp` over HTTP so several containers/harnesses (pi +
# opencode + native) can share ONE palace instead of each keeping its own.
# Point every client at it by setting, in that client's .env:
#
# MEMPALACE_REMOTE_URL=http://<reachable-host>:8765/mcp
# MEMPALACE_REMOTE_TOKEN=<the shared bearer token>
#
# (see .env.example). When set, the client connects over HTTP and does NOT
# spawn its own local mempalace-mcp.
#
# Start: docker compose -f docker-compose.mempalace.yml up -d
# Stop: docker compose -f docker-compose.mempalace.yml down
# Logs: docker compose -f docker-compose.mempalace.yml logs -f
#
# Why reuse the devbox image? mempalace-mcp is already installed in it, and
# reusing it GUARANTEES the server's mempalace version matches the clients'
# (both are pinned by the same image build). Override with a slimmer image via
# MEMPALACE_SERVER_IMAGE if you prefer (it must provide `mempalace-mcp`).
#
# ⚠ SECURITY: the HTTP transport IS authenticated as of mempalace 3.6.0 — an
# earlier version of this comment said otherwise and was wrong. The server
# compares `Authorization: Bearer <token>` with hmac.compare_digest and
# **refuses to start on a non-loopback bind without a token**, so
# MEMPALACE_REMOTE_TOKEN below is required, not optional: without it this
# service crash-loops. It also pins `Host` and allowlists `Origin`.
#
# Still do not publish port 8765 to an untrusted network. The default binds to
# 127.0.0.1 (host loopback) only. To let sibling containers reach it, attach
# them to the shared `mempalace-net` network (container-to-container, no host
# port needed — use http://mempalace-server:8765/mcp). To reach it from
# elsewhere, terminate TLS in a tunnel/reverse proxy and let the bearer token be
# the authentication — do NOT add browser-shaped auth (SSO/PIN/password) in
# front, because every MCP client here is a headless JSON-RPC POST and would
# receive a login page where JSON should be.
name: mempalace-server
services:
mempalace:
image: ${MEMPALACE_SERVER_IMAGE:-joakimp/pi-devbox:latest}
container_name: mempalace-server
# Bypass the devbox entrypoint (dev-shell/LAN/config setup) and run the
# HTTP MCP server directly. HOME + explicit --palace pin the data path so
# it does not depend on the image's default user/HOME. Runs as root so it
# can initialise the fresh named volume; the volume is dedicated to this
# server (clients reach it over HTTP, never by mounting it).
entrypoint: []
user: "0:0"
environment:
- HOME=/data
# Required: mempalace refuses a non-loopback bind without a token (it
# would exit at startup and, with restart:unless-stopped, crash-loop).
# `:?` fails fast at `docker compose up` with a readable message instead.
# Clients send the same value as MEMPALACE_REMOTE_TOKEN.
- MEMPALACE_MCP_HTTP_TOKEN=${MEMPALACE_REMOTE_TOKEN:?set MEMPALACE_REMOTE_TOKEN in .env — the shared palace requires a bearer token}
command:
- mempalace-mcp
- --transport
- http
- --host
- "0.0.0.0"
- --port
- "8765"
- --palace
- /data/.mempalace
restart: unless-stopped
# Loopback-only by default (see SECURITY note). Use "8765:8765" to expose on
# all host interfaces, or drop `ports:` entirely and rely on mempalace-net.
ports:
- "127.0.0.1:8765:8765"
volumes:
# The shared palace data — precious; back this up.
- mempalace-shared:/data/.mempalace
# Embedding-model cache (~79 MB, disposable) so search does not re-download.
- mempalace-shared-chroma:/data/.cache/chroma
# Transcript inbox. Clients cannot mine into a remote palace directly:
# `mempalace_mine` expands its source path in THIS process, so it can only
# see paths inside this container. Each client rsyncs its staged session
# exports to a per-device subdirectory on the host (see
# MEMPALACE_PI_SSH_TARGET in .env.example) and then calls mempalace_mine
# with the container-side path below (MEMPALACE_PI_REMOTE_PATH=/data/feed).
# Read-only: mining only reads sources, and all locks live palace-side.
- ${MEMPALACE_FEED_DIR:-./feed}:/data/feed:ro
networks:
- mempalace-net
healthcheck:
# GET /healthz, which is Host/Origin-gated but deliberately token-free —
# so this probe needs no credentials. Do NOT go back to POSTing
# `tools/list` here: that carries no Authorization header and now 401s,
# marking a perfectly healthy server unhealthy forever. The Host pin is
# only enforced on loopback *binds* (this one is 0.0.0.0), so a request to
# 127.0.0.1 inside the container passes.
test:
- CMD
- python3
- -c
- "import urllib.request,sys; sys.exit(0 if urllib.request.urlopen('http://127.0.0.1:8765/healthz',timeout=5).status==200 else 1)"
interval: 30s
timeout: 10s
retries: 3
start_period: 60s
volumes:
mempalace-shared:
mempalace-shared-chroma:
networks:
mempalace-net:
name: mempalace-net
+11 -4
View File
@@ -31,9 +31,13 @@ services:
- .env
environment:
- TERM=xterm-256color
- GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}
- GITEA_HOST=${GITEA_HOST:-}
- GITHUB_PERSONAL_ACCESS_TOKEN=${GITHUB_PERSONAL_ACCESS_TOKEN:-}
# Secrets (GITEA_*, GITHUB_*, and any others) are delivered to the
# container via `env_file: .env` above — do NOT duplicate them here.
# An `environment:` entry overrides env_file AND is interpolated from
# the host shell, so a stale shell export (e.g. one auto-loaded by a
# dotenv hook) would silently shadow the value in your .env. Keeping
# secrets env_file-only decouples the container from the host shell.
# See .env.example for the full list of supported variables.
volumes:
# Host workspace — mount your project here
- ${WORKSPACE_PATH:-.}:/workspace
@@ -72,7 +76,10 @@ services:
# Persist uv data (Python installs, tool installs)
- devbox-uv:/home/developer/.local/share/uv
# Optional: persist MemPalace data (conversation memory, knowledge graph)
# Optional: persist MemPalace data (conversation memory, knowledge graph).
# Applies to the LOCAL palace only (the default). In EXTERNAL mode
# (MEMPALACE_REMOTE_URL set in .env) the shared server owns the data, so
# this volume is irrelevant.
# - devbox-palace:/home/developer/.mempalace
# Optional: persist ChromaDB embedding model cache (~79 MB)
+371
View File
@@ -0,0 +1,371 @@
# Observational memory — why this image has it, and what it does for you
**Audience:** anyone using this container for long pi sessions who has wondered
what `recall`, `/om:status` and "compacted memory" are, or whether they should
leave any of it switched on.
**Companion documents:** the extension ships its own reference docs at
`/opt/pi-observational-memory/docs/` —
[`concepts.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/concepts.md)
(the model),
[`how-it-works.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/how-it-works.md)
(hooks and internals) and
[`configuration.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/configuration.md)
(every setting). Pi's own compaction mechanics are in
`/usr/lib/node_modules/@earendil-works/pi-coding-agent/docs/compaction.md`.
Those are normative; this document is the **deployment** view — what is pinned
here, how it is wired, what it costs, and how it differs from MemPalace. For the
palace, see
[`mempalace-toolkit/docs/fleet-memory.md`](https://gitea.jordbo.se/joakimp/mempalace-toolkit/src/branch/main/docs/fleet-memory.md).
> Verified on pi-devbox **v1.8.9** (`release_tag v1.8.9`, source `aac4a1c`),
> which bakes pi-observational-memory **v3.0.4** at commit `ce9fc98` — the value
> in `/etc/pi-devbox/build-manifest.json` → `components.pi-observational-memory`.
> Every number below was read from that tree, from pi 0.84.3's own docs, or from
> the live container.
---
## 1. The problem it solves
A long pi session outgrows the model's context window. Pi's answer is
**compaction**: fold the older part of the conversation into a summary and keep
recent messages verbatim. That is unavoidable, and it is where sessions go
wrong — the summary is produced *at the moment of pressure*, by a model, about a
transcript that is about to leave the context.
Observational memory changes *when* the remembering happens. Instead of
summarising in a panic at the end, it keeps a small **ledger** up to date while
the session runs, and compaction then just folds that ledger.
```mermaid
flowchart LR
A0["plain compaction"] --> A1["context fills"]
A1 --> A2["a model summarises<br/>under pressure"]
A2 --> A3["prose summary,<br/>no way back"]
B0["with observational<br/>memory"] --> B1["context fills"]
B1 --> B2["ledger written<br/>as you work"]
B2 --> B3["compaction folds<br/>the ledger"]
B3 --> B4["ids you can<br/>recall"]
```
Top row is pi on its own: one model call at the worst possible moment, detail
chosen in a hurry, and the original wording gone from view. Bottom row is this
image's default: the thinking happened earlier on a cheap model, the fold is
deterministic, and every line in the result carries an id that resolves back to
the exact source.
## 2. The mental model: three layers and a ledger
| Layer | What it is | Example |
|---|---|---|
| **Observation** | a timestamped, source-backed event from the conversation | "user rejected option B because it needs a base rebuild" |
| **Reflection** | a durable conclusion *backed by* observations | "the user optimises for avoiding 67-minute rebuilds" |
| **Drop** | a tombstone retiring an observation from active memory | the superseded detail of a bug that is now fixed |
These are appended to the session as silent ledger entries
(`om.observations.recorded`, `om.reflections.recorded`,
`om.observations.dropped`) and **folded** — replayed in order — to produce the
memory state. The ledger is the source of truth; what you see in a compacted
session is a rendering of it.
Two properties follow, and both matter later:
- **The ledger itself costs no context.** Those entries are pi `custom` entries,
which *"do not participate in LLM context"* (pi `docs/session-format.md`). They
sit in the session file and reach the model only via the fold at compaction.
- **Memory is branch-local.** A pi session is a tree (resume, fork), and the fold
follows the current branch only, so a forked branch does not inherit another
branch's view.
## 3. The lifecycle
Three background workers and one compaction hook, driven by *raw token
progress* rather than wall-clock time. Defaults in brackets.
```mermaid
flowchart TD
T(["turn_end"]) --> O{"10k raw tokens<br/>since observing?"}
O -- yes --> OBS["<b>observer</b> runs"]
O -- "no" --> R{"20k tokens<br/>since reflecting?"}
R -- yes --> REF["<b>reflector</b> runs"]
REF -- "if pool over 10k" --> DR["<b>dropper</b> prunes"]
S(["agent_settled"]) --> C{"81k tokens<br/>since compacting?"}
C -- yes --> CP["ctx.compact()"]
CP --> H(["session_before_compact"])
H --> F["fold the ledger<br/>no model call"]
F --> VIS["compacted memory"]
```
- **observer** — `observeAfterTokens` [10000]: writes observations for the
conversation it has not covered yet.
- **reflector** — `reflectAfterTokens` [20000]: promotes patterns across
observations into durable reflections.
- **dropper** — no clock of its own. It is post-reflection maintenance, gated on
a *successful same-turn* reflection **and** an active pool above
`observationsPoolTargetTokens` [10000]. Not a third worker on a third
threshold.
- **compaction** — `compactAfterTokens` [81000], checked when pi goes idle, so it
never interrupts a turn. Pi will also compact on its own when the context is
nearly full (`contextTokens > contextWindow - reserveTokens`, `reserveTokens`
[16384]).
## 4. What compaction actually does to your context
This is the question the rest of the document used to leave hanging: if the old
conversation is folded away, is the session back to knowing nothing?
**No.** Compaction replaces *part* of the context, not all of it, and it deletes
nothing at all from disk.
```mermaid
flowchart LR
SYS["system prompt<br/>+ AGENTS.md"] --> CTX["what the model sees<br/>on the next turn"]
SUM["folded memory:<br/>reflections + observations"] --> CTX
TAIL["recent turns,<br/>verbatim"] --> CTX
DISK[("session .jsonl: all of it")] -. "recall(id)" .-> CTX
```
Where each piece comes from:
- **System prompt and `AGENTS.md` — never compacted, because they were never
conversation.** Pi rebuilds them from disk on every request
(`loadContextFileFromDir`), so they cannot be lost by compaction.
- **The verbatim tail — sized by a token budget, not a message count.** Pi walks
backwards from the newest entry accumulating token estimates until
`keepRecentTokens` [20000] is reached; that entry becomes `firstKeptEntryId`,
and *everything from there on is kept unchanged*. Cut points land on turn
boundaries, never mid-tool-call. So the most recent ~20k tokens of real work —
your last instructions, the diffs, the test output — survive word for word.
- **The folded memory — replaces only what came before that cut.** Rendered from
the ledger's records: reflections and observations, each with its 12-hex id.
- **The session file — untouched.** Compaction *appends* a `compaction` entry
(`{"type":"compaction", summary, firstKeptEntryId, tokensBefore, …}`) and
rebuilds context from it on later turns. Nothing is rewritten in place; the
only documented way to remove session content is deleting the whole `.jsonl`.
That last point is what makes the answer to "is the detail gone?" *no* rather
than *mostly*: `recall` does not read the context window at all. It calls
`sessionManager.getBranch()` — the full branch from the root — and resolves an
observation id back to the original entries. Detail that left the model's view
an hour ago is still one `recall` away.
**Repeated compaction does not summarise the summary.** The rendered text is
always built from live observation/reflection *records*, never from the previous
compaction's prose, so there is no generation-loss spiral. (Mechanically the
projection is incremental — it re-derives back to the last full-fold boundary and
carries the rest forward, escalating to a genuine re-fold from the branch root
when the observation pool reaches `observationsPoolMaxTokens` [20000].)
So the honest summary of the state after compaction: **the model keeps its
instructions, keeps recent work verbatim, trades older turns for a dense
id-carrying digest of them, and can pull any of it back on demand.** Not a fresh
start — a smaller, cheaper, still-navigable one.
### One caveat about "no model call"
If the ledger is empty — compaction fires before the observer has ever run — the
hook returns nothing and *declines ownership*, and pi's own model-based
summariser runs instead:
```ts
const summary = renderSummary(projection.reflections, projection.observations);
if (summary.length === 0) {
// Decline ownership so Pi's native summarizer preserves the pre-cut context.
return;
}
```
In steady state (any session old enough to have produced one observation) om's
hook wins and compaction is model-free. "Never calls a model" is true in practice
and false in principle; the fallback is deliberate, so an empty ledger degrades
to normal pi rather than to no summary at all.
## 5. What you actually get
- **Compaction stops being a stall.** In steady state the latency path is
deterministic work over ledger entries, not a summarisation call.
- **Nothing important vanishes silently.** Compaction is lossy by design, but
every item keeps a 12-character id, and `recall(<id>)` returns the exact
evidence — original wording, reasoning, file path, error text.
- **The bookkeeping runs on a cheaper model than your session.** In this image
that is deliberate and visible (§7): background workers on Haiku, session on
Opus.
- **It is automatic.** No habit to maintain, unlike the palace protocol — which is
exactly why the two complement each other (§11).
- **Forks stay clean.** Branch-local memory means a `fork` sub-agent's noise does
not leak into the parent's folded memory.
## 6. `recall` is not a search tool
`recall` takes **one specific 12-hex id** that already appears in compacted
memory or in `/om:view`. It cannot be given a topic. It can return an observation
(marked `active` or `dropped`), or a reflection together with the observations
supporting it.
```mermaid
sequenceDiagram
participant M as compacted memory
participant A as agent
participant L as ledger
M->>A: "[high] user rejected option B (a1b2c3d4e5f6)"
A->>L: recall("a1b2c3d4e5f6")
L-->>A: exact observation + source ids
Note over A: acts on the original wording
```
The rule of thumb the agent skill uses: recall **before a load-bearing action**
that rests on a compressed memory — shipping a change, asserting a fact,
answering "why do you believe that". One recall is cheap; redoing finished work
is not.
## 7. How it is wired in this image
```mermaid
flowchart TB
IMG["baked in the image:<br/>v3.0.4 @ ce9fc98"] --> REG["settings.json<br/>packages[]"]
REG --> SESS["your pi session"]
SESS -- "your turns" --> SM["session model:<br/>Opus"]
SESS -- "observer, reflector,<br/>dropper" --> WM["memory model:<br/>Haiku"]
SESS -- "ledger entries" --> JL["session .jsonl"]
JL --> VOL[("devbox-pi-config<br/>volume")]
```
Four consequences of that wiring:
1. **There is no separate database.** Memory *is* entries inside the ordinary pi
session file (`~/.pi/agent/sessions/<project>/<timestamp>_<uuid>.jsonl`).
Nothing extra to back up, nothing to migrate.
2. **It survives container recreate**, because `~/.pi` is the `devbox-pi-config`
named volume (`docker-compose.yml`) — the same one holding your pi config and
session history.
3. **`packages[]` is the only source of truth for which copy is loaded.** A clone
at `/workspace/pi-observational-memory` may exist (and today matches `/opt`
byte-for-byte at `ce9fc98`) — its presence proves nothing. To run a patched
build you point `packages[]` at it explicitly and start a new session.
4. **The worker model is a deliberate choice, and it is yours to change.** The
seeded config sends background work to Haiku while your session runs Opus:
```json
"observational-memory": {
"model": { "provider": "amazon-bedrock", "id": "eu.anthropic.claude-haiku-4-5-20251001-v1:0" },
"debugLog": false
}
```
## 8. What it costs
| Resource | Cost |
|---|---|
| Model calls | up to **three** background calls per consolidation pass (observer, reflector, dropper), each capped at `agentMaxTurns` [16], on the configured memory model — not your session model |
| Latency in your turns | none by construction: workers run from `turn_end`, compaction runs when pi is idle, and the fold itself does no model work |
| Disk | negligible — JSON lines inside a session file that would exist anyway (measured here: `~/.pi/agent/sessions` = 30 MB total, tens of `om.*` entries per session) |
| Context window | **zero until compaction.** `custom` entries do not enter LLM context; only the folded summary does |
| Attention | none once configured; there is no protocol for you or the agent to remember |
If that is still more than you want on a given run, §9's `passive` switch turns
off all proactive work while keeping `recall` and `/om:*` usable.
## 9. Configuration
Global: `~/.pi/agent/settings.json` (persisted in the volume). Per project:
`<project>/.pi/settings.json`, which overrides global. Precedence is
project → global → environment, and the environment can only override `passive`.
| Key | Default | What it changes |
|---|---|---|
| `observeAfterTokens` | `10000` | observer cadence — lower means smaller chunks and more calls |
| `reflectAfterTokens` | `20000` | reflector cadence (and thereby dropper opportunities) |
| `observerChunkMaxTokens` | 20% of the memory model's context window, else `60000` | cap on one observer run's input |
| `compactAfterTokens` | `81000` | when proactive auto-compaction fires |
| `observationsPoolMaxTokens` | `20000` | pool size at which compaction does a full re-fold from the branch root |
| `observationsPoolTargetTokens` | half of max (`10000`) | what the dropper aims back down to |
| `agentMaxTurns` | `16` | shared turn cap for the three workers |
| `model` | unset → session model | send background work to a cheaper/faster model |
| `showWorkerNotifications` | `true` | routine "observer ran" notices |
| `passive` | `false` | **kill switch** for all proactive background work; `recall` and `/om:*` still work |
| `debugLog` | `false` | per-session NDJSON trace at `~/.pi/agent/observational-memory/debug/<session-id>.ndjson` |
Pi's own compaction knobs live under a separate `compaction` key —
`keepRecentTokens` [20000] sets the verbatim tail from §4, `reserveTokens`
[16384] the headroom that triggers pi's own compaction.
One-off passive run, no config edit:
```bash
PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi
```
Invalid values are ignored rather than fatal, so a typo degrades to the default
instead of breaking your session — which also means a typo is silent. Check with
`/om:status`.
## 10. Confirming it is actually working
Do not infer health from the absence of a warning; look:
```bash
# 1. inside pi — the authoritative view
/om:status # visible-vs-full drift, thresholds, worker state
/om:view # what the agent currently sees
/om:view full # full ledger truth at the branch tip
# 2. from a shell — are ledger entries being written, and has it compacted?
grep -o '"customType":"om\.[a-z.]*"' \
"$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)" | sort | uniq -c
grep -c '"type":"compaction"' "$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)"
# 3. which copy is loaded, and at what commit
python3 -c "import json;print(json.load(open('$HOME/.pi/agent/settings.json'))['packages'])"
git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory rev-parse HEAD
```
Ledger entries are `"type":"custom"` with `"customType":"om.…"`. Do not grep for
`custom_message` — that is a *different* pi API for entries that **do** enter LLM
context, used here by the MemPalace mailbox (`customType: "mempalace-mailbox"`),
not by om.
## 11. It is not the same thing as MemPalace
Both are called "memory" and they solve different problems. Nothing is wrong with
running both — this image does, and they cover each other's failure modes.
```mermaid
flowchart LR
O0["observational<br/>memory"] --> O1["horizon:<br/>this session"]
O1 --> O2["scope: one branch,<br/>one machine"]
O2 --> O3["automatic"]
O3 --> O4["retrieval:<br/>recall(id)"]
P0["MemPalace"] --> P1["horizon: months,<br/>machines"]
P1 --> P2["scope:<br/>the fleet"]
P2 --> P3["protocol-driven"]
P3 --> P4["retrieval:<br/>search, KG, mailbox"]
```
| Question | Answer |
|---|---|
| "What did we decide 200 turns ago in *this* session?" | observational memory (and `recall` for the exact wording) |
| "What did we decide last month, or on another machine?" | MemPalace (`mempalace_search`, diaries) |
| "What is true *right now* about version X?" | MemPalace knowledge graph |
| "Does another machine need something from me?" | MemPalace coordination log — see [Cross-machine agent coordination](../README.md#cross-machine-agent-coordination) |
| "Why is compaction not losing my session?" | observational memory |
The crisp version: **observational memory keeps a session coherent; the palace
keeps the fleet coherent.** A container recreate wipes neither — but only because
`~/.pi` and the palace both live outside the container filesystem.
## 12. Gotchas
- **Branch-local means branch-local.** Resuming or forking changes which ledger
is folded. Memory that "disappeared" is usually on another branch.
- **`recall` needs an id, not a topic.** If you only have a topic, that is a
palace search, not a recall.
- **A `/workspace` clone is not evidence of what is loaded** — see §7.3.
- **`showWorkerNotifications: true` is not proof of work**; it reports runs, and
an observer that deliberately emits nothing writes no ledger entry and simply
retries after another `observeAfterTokens`.
- **A turn bigger than `keepRecentTokens` splits.** The cut then lands mid-turn at
an assistant message and pi merges two summaries — rare, but it is why a very
large single turn can lose more verbatim detail than you would expect.
- **`git log` in the baked tree needs `safe.directory`** (`/opt` is root-owned):
`git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory log`.
+270 -18
View File
@@ -1,6 +1,18 @@
#!/usr/bin/env bash
set -euo pipefail
# ── Startup banner: which pi-devbox build is this? ─────────────────
# Printed FIRST, before the setup noise below, so it's the first thing
# visible when the container starts (CMD is `bash -l`, tty:true in compose,
# so this reaches the same stream as the interactive shell the user lands
# in). Reads the ground-truth manifest baked in Dockerfile.variant; a no-op
# with a short stderr notice on images built before it existed.
# `--no-skills`: this runs FIRST, before the baked skill links are created
# below and long before the skillset deploy + devbox-skill-reconcile run at the
# end of this script, so the skill-source section would report a pre-reconcile
# state that is about to change. Wrong-but-plausible is worse than absent.
command -v pi-devbox-version >/dev/null 2>&1 && pi-devbox-version --no-skills || true
# ── SSH ControlMaster socket dir ────────────────────────────────
# Companion to /etc/ssh/ssh_config.d/00-devbox-controlmaster.conf in the
# base image — that file declares ControlPath=/tmp/sshcm/%r@%h:%p; this
@@ -12,12 +24,16 @@ set -euo pipefail
mkdir -p /tmp/sshcm
chmod 700 /tmp/sshcm
# ── LAN access: generic host-OS-agnostic reachability helper ────────
# On VM-backed hosts (macOS OrbStack / Docker Desktop) the container can't
# reach the host's directly-attached LAN peers by default; this generates a
# writable ~/.ssh-local/config that uses the host as an SSH jump. On native
# Linux (LAN reachable directly) it is a no-op. Controlled by DEVBOX_LAN_ACCESS
# (auto|jump|off) + HOST_SSH_USER. Always non-fatal. See the script header.
# ── LAN access + writable SSH sidecar: host-OS-agnostic helper ──────
# Generates the writable ~/.ssh-local/config on EVERY host OS: a `Host *`
# ControlPath redirect into ~/.ssh-local/cm (so `ssh -F` / dssh / dscp work
# even when ~/.ssh is bind-mounted read-only) plus `Include ~/.ssh/config`. On
# VM-backed hosts (macOS OrbStack / Docker Desktop) it ALSO adds an
# SSH-jump-via-host block so the container can reach the host's
# directly-attached LAN peers; on native Linux (LAN reachable directly) the
# jump block is omitted but the sidecar is still rendered. Controlled by
# DEVBOX_LAN_ACCESS (auto|jump|off) + HOST_SSH_USER. Always non-fatal. See the
# script header.
if [ -r /usr/local/lib/pi-devbox/setup-lan-access.sh ]; then
bash /usr/local/lib/pi-devbox/setup-lan-access.sh || true
fi
@@ -29,13 +45,55 @@ fi
# directly.
SKEL_DIR="/etc/skel-devbox"
if [ -d "$SKEL_DIR" ]; then
for f in .bash_aliases .inputrc; do
for f in .bash_aliases .inputrc .gitignore_global; do
if [ -f "$SKEL_DIR/$f" ] && [ ! -e "$HOME/$f" ]; then
cp "$SKEL_DIR/$f" "$HOME/$f"
fi
done
fi
# ── Image-baked skills: link into ~/.agents/skills ───────────────────
# Skills shipped IN the image (under /usr/local/share/pi-devbox/skills/) are
# made available regardless of whether a skillset repo is mounted. Done EARLY
# — before the pi-toolkit/extensions deploy below — so the symlinks exist by
# the time anything gates on "container ready": the smoke-test readiness probe
# waits on pi-deploy markers (keybindings.json, mempalace.ts) that only land
# AFTER this point, so linking here closes a sample-too-early race that failed
# the runtime skill-link assertion. Pointing at the image path (/usr/local/...)
# keeps the skill fresh from the image and surviving volume recreate (unlike
# anything baked under a home dir, which a named volume would shadow). Created
# only when absent, so a user override is never clobbered.
#
# NB: "created only when absent" does NOT hand a same-named skillset skill
# priority — the opposite. The skillset deploy runs at the end of this script
# and classifies these links as foreign, so through v1.8.4 the BAKED copy
# always won and an edit pushed to a skillset-owned skill was invisible until
# the next image build. The links below are therefore the FALLBACK only;
# devbox-skill-reconcile (invoked right after the skillset deploy) hands the
# skillset-OWNED skills back to the live clone. Ownership is per-skill, listed
# in skills/skillset-owned.txt — see VENDORED.md for why pi-extensions must
# keep losing to the baked copy.
DEVBOX_SKILLS_SRC=/usr/local/share/pi-devbox/skills
if [ -d "$DEVBOX_SKILLS_SRC" ]; then
mkdir -p "$HOME/.agents/skills"
for _sk in "$DEVBOX_SKILLS_SRC"/*/; do
[ -d "$_sk" ] || continue
_skname=$(basename "$_sk")
if [ ! -e "$HOME/.agents/skills/$_skname" ]; then
# -sfn, not -s: `[ ! -e ]` is TRUE for a DANGLING symlink (-e follows the
# link), and since v1.8.5 these links can point into /workspace/skillset
# (see devbox-skill-reconcile, invoked after the skillset deploy). If that
# mount vanishes while the writable layer survives — a `docker restart` or
# a host reboot under restart: unless-stopped, as opposed to a recreate —
# plain `ln -s` fails with "File exists" and, under `set -e`, aborts
# container start before `exec "$@"`. With -f the broken link heals back to
# the baked fallback, and the reconciler re-points it in the same boot if
# the clone is back.
ln -sfn "${_sk%/}" "$HOME/.agents/skills/$_skname"
fi
done
fi
# ── MemPalace: initialize palace for the workspace if mempalace is installed
# Creates the palace directory structure on first run. Idempotent — skips
# if palace already exists, so upgrades from older versions preserve
@@ -54,6 +112,82 @@ if command -v mempalace &>/dev/null && [ -d /workspace ]; then
fi
fi
# ── MemPalace: pi transcript feeder ─────────────────────────────────
# mempalace-toolkit ships `mempalace-pi-session`, which mines pi's own JSONL
# session transcripts into the palace. pi's mempalace extension drives it on
# session_shutdown and on a debounced agent_settled; this is the catch-up for
# the one case no handler can cover — a hard kill (docker kill, OOM, host
# reboot) runs nothing at all, so without this the previous life's transcripts
# are never mined.
#
# No MEMPALACE_PI_STAGE override here on purpose: the feeder stages next to the
# palace it feeds (<palace-root>/pi-stage), so the stage and the dedup keys
# referencing it share one lifetime — whatever persistence the palace has, the
# stage inherits. Pinning it elsewhere (e.g. into the ~/.pi volume) would
# re-introduce the very split that design prevents: palace volume kept, stage
# volume dropped, and `mempalace sync` then prunes every conversation drawer.
#
# Backgrounded: a cold mine can take tens of seconds and must never delay the
# shell. Contention with a live session is handled by the tool itself (it exits
# 0 and lets the palace holder do the mine).
# Self-heal onto PATH for images whose base predates the toolkit symlink.
# ~/.local/bin is already ahead of /usr/local/bin on PATH (Dockerfile.base sets
# it in ENV PATH) and is writable by this (non-root) user, unlike /usr/local/bin.
if [ -x /opt/mempalace-toolkit/bin/mempalace-pi-session ] && \
! command -v mempalace-pi-session >/dev/null 2>&1; then
mkdir -p "$HOME/.local/bin"
ln -sf /opt/mempalace-toolkit/bin/mempalace-pi-session "$HOME/.local/bin/mempalace-pi-session"
fi
# Resolve the feeder explicitly rather than trusting PATH: this runs before any
# login shell, and a silently-skipped catch-up is exactly the failure we are
# here to prevent.
MEMPALACE_FEEDER=""
if command -v mempalace-pi-session >/dev/null 2>&1; then
MEMPALACE_FEEDER="mempalace-pi-session"
elif [ -x /opt/mempalace-toolkit/bin/mempalace-pi-session ]; then
MEMPALACE_FEEDER="/opt/mempalace-toolkit/bin/mempalace-pi-session"
fi
if [ "${MEMPALACE_FEED:-1}" != "0" ] && [ -n "$MEMPALACE_FEEDER" ]; then
if [ -n "${MEMPALACE_REMOTE_URL:-}" ] && [ -z "${MEMPALACE_PI_SSH_TARGET:-}" ]; then
# Remote palace, but no inbox to ship transcripts to — the feeder genuinely
# cannot do anything here, so skipping is right. Saying so is the point:
# this branch used to be a bare `:`, and the skip happens *before* the
# subshell below that writes mempalace-catchup.log, so a container in this
# state contributed nothing to the palace and left no artifact at all — not
# even an empty log — to explain why. That is indistinguishable from a
# healthy run that simply had nothing to file. `tee` puts the notice both in
# the container's start output (docker logs) and at the path anyone
# debugging "why is nothing from this container in the palace?" looks first.
# This is an entrypoint: a notice must never be able to stop a container
# from starting. An unwritable ~/.pi (root-owned volume — a classic Docker
# permission accident) makes `mkdir -p` fail, and under `set -e` that would
# abort startup entirely: a brand-new failure mode in precisely the branch
# that used to do nothing at all. Degrade to stdout-only instead.
_mp_log="$HOME/.pi/agent/mempalace-catchup.log"
mkdir -p "$HOME/.pi/agent" 2>/dev/null || _mp_log=/dev/null
{
echo "MemPalace catch-up skipped: remote palace with no transcript inbox."
echo " MEMPALACE_REMOTE_URL is set (${MEMPALACE_REMOTE_URL})"
echo " but MEMPALACE_PI_SSH_TARGET is not, so there is nowhere to ship this"
echo " container's staged sessions. MCP tools still read and write the shared"
echo " palace — but this container's own conversations are mined nowhere."
echo " Fix: set MEMPALACE_PI_SSH_TARGET (and MEMPALACE_PI_DEVICE) in .env,"
echo " or unset MEMPALACE_REMOTE_URL to keep the palace local."
echo " Deliberate? MEMPALACE_FEED=0 turns the feed off and silences this."
} | tee "$_mp_log" 2>/dev/null || true
unset _mp_log
else
mkdir -p "$HOME/.pi/agent"
(
"$MEMPALACE_FEEDER" --reason container-start \
>"$HOME/.pi/agent/mempalace-catchup.log" 2>&1 || true
) &
fi
fi
# ── Git config defaults ──────────────────────────────────────────────
if [ -n "${GIT_USER_NAME:-}" ] && ! git config --global user.name &>/dev/null; then
git config --global user.name "$GIT_USER_NAME"
@@ -61,6 +195,12 @@ fi
if [ -n "${GIT_USER_EMAIL:-}" ] && ! git config --global user.email &>/dev/null; then
git config --global user.email "$GIT_USER_EMAIL"
fi
# Global gitignore for personal/tooling artifacts (*.bak, *~, *.orig, ...).
# Seeded above into $HOME/.gitignore_global from /etc/skel-devbox. Point git at
# it only if the user has not already set their own core.excludesFile.
if [ -f "$HOME/.gitignore_global" ] && ! git config --global core.excludesFile &>/dev/null; then
git config --global core.excludesFile "$HOME/.gitignore_global"
fi
# ── pi: deploy toolkit + extensions + mempalace bridge ─────────────
# pi is always installed in pi-devbox; no INSTALL_PI guard needed.
@@ -86,9 +226,35 @@ if command -v pi &>/dev/null; then
# Bootstrap settings.json from template if absent (pi rewrites this
# file at runtime — lastChangelogVersion, etc — so we can't symlink it).
if [ ! -f "$HOME/.pi/agent/settings.json" ] && \
[ -f /opt/pi-toolkit/settings.example.json ]; then
cp /opt/pi-toolkit/settings.example.json "$HOME/.pi/agent/settings.json"
_pi_settings="$HOME/.pi/agent/settings.json"
_pi_template=/opt/pi-toolkit/settings.example.json
if [ ! -f "$_pi_settings" ] && [ -f "$_pi_template" ]; then
cp "$_pi_template" "$_pi_settings"
echo "pi settings.json bootstrapped from template"
elif [ -f "$_pi_settings" ] && [ -f "$_pi_template" ] && \
[ "${PI_SETTINGS_MERGE:-1}" != "0" ] && command -v jq >/dev/null 2>&1; then
# Non-destructive merge: a settings.json on a PRESERVED volume never
# otherwise sees new template keys (the bootstrap above only fires when
# the file is absent), so config added in an image upgrade — e.g. the
# observational-memory / pi-fork blocks or a newly-enabled model — never
# reaches existing users. Deep-merge with the template FIRST and the
# live file SECOND ('.[0] * .[1]') so the user's values always win and
# only keys MISSING from the live file are filled in from the template.
# Arrays are treated as leaves (the user's array is kept verbatim, so a
# model they deliberately removed is not re-added). Only rewrite when the
# merge actually changes something, and back up the original first.
# Set PI_SETTINGS_MERGE=0 to disable. Invalid JSON on either side → skip,
# never clobber.
if _pi_merged=$(jq -s '.[0] * .[1]' "$_pi_template" "$_pi_settings" 2>/dev/null); then
if [ -n "$_pi_merged" ] && \
! printf '%s' "$_pi_merged" | jq -e --slurpfile cur "$_pi_settings" '. == $cur[0]' >/dev/null 2>&1; then
cp "$_pi_settings" "${_pi_settings}.bak.$(date +%Y%m%d-%H%M%S)"
printf '%s\n' "$_pi_merged" > "$_pi_settings"
echo "pi settings.json: merged new template keys from settings.example.json (backup saved)"
fi
else
echo "WARN: pi settings.json merge skipped (jq could not parse template or live file; left untouched)"
fi
fi
# pi↔mempalace MCP bridge — single extension symlink.
@@ -99,22 +265,101 @@ if command -v pi &>/dev/null; then
"$HOME/.pi/agent/extensions/mempalace.ts"
fi
# pi-fork (fork tool) + pi-observational-memory (recall tool) + (in the
# :latest-studio variant only) pi-studio (/studio command + studio_*
# tools + theme). These are pi packages (not symlink-style extensions):
# pi-fork (fork tool) + pi-observational-memory (recall tool) + pi-atelier
# (TUI sidebar panels/split-pane) + (in the :latest-studio variant only)
# pi-studio (/studio command + studio_* tools + theme). These are pi packages (not symlink-style extensions):
# they're cloned to /opt with node_modules baked at BUILD time, then
# registered here via `pi install <local-path>`. A local-path install is
# instant + in-place (pi loads the extension directly from /opt) +
# idempotent (no duplicate package entry on re-run), and stores a relative
# path that resolves into the image-layer /opt so it survives volume
# recreate. The tools/command register on the NEXT pi start (extensions
# bind at startup). Guard on settings.json so we only install once per
# volume. /opt/pi-studio is present only in the studio variant; the
# `[ -d ]` test makes this a no-op everywhere else.
for _pkg in /opt/pi-fork /opt/pi-observational-memory /opt/pi-studio; do
# bind at startup) or on `/reload`. Guard on settings.json so we only
# install once per volume. /opt/pi-studio is present only in the studio
# variant; the `[ -d ]` test makes this a no-op everywhere else.
#
# The guard MUST inspect the `packages` ARRAY, not merely grep the whole
# file for the package name. settings.example.json ships a top-level
# "pi-fork" CONFIG block (the fork effort profiles, pi-toolkit adb6907,
# 2026-06-17), so a whole-file substring grep matches on any settings.json
# that was bootstrapped from — or template-merged with — that template.
# Worse, the merge above runs FIRST, so it plants the matching string in the
# same startup that the loop then reads: `pi install /opt/pi-fork` was
# skipped forever and the `fork` tool never registered (v1.0.0 → v1.6.3).
# Its siblings escaped only by luck — the template key is
# "observational-memory" (no pi- prefix) and there is no studio block.
# jq reads the array; the grep fallback matches the stored relative-path
# form ("…/opt/<name>\""), which a config KEY can never produce.
_pi_pkg_registered() {
_pi_reg_settings="$HOME/.pi/agent/settings.json"
[ -f "$_pi_reg_settings" ] || return 1
if command -v jq >/dev/null 2>&1; then
jq -e --arg n "$1" \
'(.packages // []) | any((type == "string") and (. == "npm:" + $n or endswith("/" + $n)))' \
"$_pi_reg_settings" >/dev/null 2>&1
else
grep -q "opt/$1\"" "$_pi_reg_settings"
fi
}
# ── pi-atelier: retire a stale `npm:pi-atelier`, plus an opt-out ──────
# The image now vendors pi-atelier at a pinned, audited tag (PI_ATELIER_REF
# in Dockerfile.variant). A leftover `npm:pi-atelier` entry from a
# hand-install resolves through ~/.pi/npm-global, which lives on the
# devbox-pi-config VOLUME — so it survives image upgrades and keeps whatever
# version was installed by hand, unpinned and unaudited. That is not
# academic: pi-atelier < 0.7.1 makes pi >= 0.84 hang at startup with
# sustained CPU, so leaving it in place turns a pi bump into a TUI that will
# not start. And `_pi_pkg_registered` deliberately counts `npm:<name>` as
# registered (it respects a user's own npm install), so the loop below would
# never replace it.
#
# We only DELETE the exact `npm:pi-atelier` string; the loop then registers
# /opt/pi-atelier in pi's own canonical serialization, so this code never has
# to guess the stored relative-path form. Idempotent — after the rewrite
# there is no npm entry left to match.
#
# DEVBOX_ATELIER=0 goes further and removes pi-atelier from `packages`
# altogether. That escape hatch lives HERE, in the entrypoint, precisely
# because this component's known failure mode is "pi will not start" — which
# you cannot repair with `pi uninstall`.
_pi_atelier_drop() {
# $1 = jq predicate over one `packages` entry, selecting what to REMOVE.
# Returns 0 only when the file was actually rewritten (caller logs), 1 for
# "nothing to do" — including missing jq or unparseable JSON, which must
# never clobber user settings. Backs up first, same convention as the
# template merge above.
_ad_settings="$HOME/.pi/agent/settings.json"
[ -f "$_ad_settings" ] || return 1
command -v jq >/dev/null 2>&1 || return 1
_ad_new=$(jq "(.packages // []) |= map(select(($1) | not))" "$_ad_settings" 2>/dev/null) || return 1
[ -n "$_ad_new" ] || return 1
if printf '%s' "$_ad_new" | jq -e --slurpfile cur "$_ad_settings" '. == $cur[0]' >/dev/null 2>&1; then
return 1
fi
# `.bak.atelier.` rather than the merge's plain `.bak.` prefix: both can
# fire in the same startup, and a bare seconds-resolution timestamp would
# make the second cp overwrite the first one's backup.
cp "$_ad_settings" "${_ad_settings}.bak.atelier.$(date +%Y%m%d-%H%M%S)"
printf '%s\n' "$_ad_new" > "$_ad_settings"
return 0
}
if [ "${DEVBOX_ATELIER:-1}" = "0" ]; then
if _pi_atelier_drop '(. == "npm:pi-atelier") or ((type == "string") and endswith("/pi-atelier"))'; then
echo "pi-atelier: unregistered per DEVBOX_ATELIER=0 (settings backup saved)"
fi
elif [ -d /opt/pi-atelier ]; then
if _pi_atelier_drop '. == "npm:pi-atelier"'; then
echo "pi-atelier: dropped stale npm: registration — the pinned /opt copy takes over (settings backup saved)"
fi
fi
for _pkg in /opt/pi-fork /opt/pi-observational-memory /opt/pi-studio /opt/pi-atelier; do
[ -d "$_pkg" ] || continue
_name=$(basename "$_pkg")
if ! grep -q "$_name" "$HOME/.pi/agent/settings.json" 2>/dev/null; then
# DEVBOX_ATELIER=0 → leave pi-atelier unregistered (handled just above).
if [ "$_name" = "pi-atelier" ] && [ "${DEVBOX_ATELIER:-1}" = "0" ]; then continue; fi
if ! _pi_pkg_registered "$_name"; then
pi install "$_pkg" >/dev/null 2>&1 || \
echo "WARN: pi install $_name failed (continuing)"
fi
@@ -160,6 +405,13 @@ elif [ -x /workspace/skillset/deploy-skills.sh ]; then
fi
if [ -n "$SKILLSET_DEPLOY" ]; then
"$SKILLSET_DEPLOY" --bootstrap --prune-stale >/dev/null 2>&1 || true
# The deploy leaves the early baked links (above) in place as foreign links,
# which silently shadows the live clone for skills the skillset OWNS. Repoint
# just those; baked stays the fallback, user overrides still win. `|| true`:
# a skill-link refinement must never break container start.
if command -v devbox-skill-reconcile >/dev/null 2>&1; then
devbox-skill-reconcile "$(dirname "$SKILLSET_DEPLOY")" || true
fi
fi
# ── Execute command ──────────────────────────────────────────────────
+18
View File
@@ -0,0 +1,18 @@
" pi-devbox — system-wide Neovim defaults.
"
" This is Neovim's *system vimrc*: it loads for every user before any personal
" ~/.config/nvim, and personal configs can still override it.
"
" Enable 24-bit ("true") colour. Without it, Neovim's default theme is squeezed
" into a 256-colour palette where strings/comments become a muddy, low-contrast
" dark colour — a common complaint over ssh/kitty where COLORTERM often isn't
" propagated into the container. Modern terminals (kitty, WezTerm, iTerm2,
" Alacritty, ...) all support true colour; the bundled kitty-terminfo also lets
" Neovim auto-detect it, but forcing it here guarantees readable colour
" regardless of how the terminal type / COLORTERM reach the container.
"
" Opt out for a session: :set notermguicolors
" Override permanently: set your own value in ~/.config/nvim/init.lua
if has('termguicolors')
set termguicolors
endif
+40 -1
View File
@@ -54,6 +54,38 @@ alias gs='git status'
alias gd='git diff'
alias gl='git log --oneline --graph --decorate -20'
# ── Host SSH reachability check (once per container lifetime) ───────────────
# Warns at first shell startup if the Mac host is not reachable via SSH.
# Only runs inside a container, only if the jump key exists, and only once
# per container lifetime (/tmp flag is cleared on recreate).
_devbox_check_host_ssh() {
[ -f "/.dockerenv" ] || return 0
local ssh_cfg="$HOME/.ssh-local/config"
[ -f "$ssh_cfg" ] || return 0
local key_pub="$HOME/.ssh-local/devbox_jump_ed25519.pub"
[ -f "$key_pub" ] || return 0
local flag="/tmp/.devbox_host_ssh_ok"
[ -f "$flag" ] && return 0
if ssh -F "$ssh_cfg" \
-o BatchMode=yes \
-o ConnectTimeout=2 \
-o StrictHostKeyChecking=accept-new \
mac true 2>/dev/null; then
touch "$flag"
return 0
fi
local pub_key
pub_key=$(cat "$key_pub")
printf '\n\033[1;33m⚠ devbox: Mac host not reachable via SSH\033[0m\n'
printf ' Some tools use SSH to run commands on the Mac host.\n'
printf ' Fix (run both on the Mac):\n\n'
printf ' \033[1mStep 1\033[0m System Settings → General → Sharing → Remote Login → ON\n\n'
printf ' \033[1mStep 2\033[0m echo '"'"'%s'"'"' >> ~/.ssh/authorized_keys\n' "$pub_key"
printf '\n Then open a new shell in the container to verify.\n\n'
}
_devbox_check_host_ssh
unset -f _devbox_check_host_ssh
# ── LAN access via the host (dssh) ───────────────────────────────────
# When running on a VM-backed host (macOS OrbStack / Docker Desktop), the
# entrypoint's setup-lan-access.sh generates ~/.ssh-local/config so the host
@@ -89,9 +121,16 @@ fi
# we append with a newline separator to avoid the ';;' parse error
# described at the top of this file. Guarded so repeated sourcing
# (e.g. `exec bash`) doesn't stack duplicates.
#
# The guard MUST stay shell-local (NOT exported): if it leaks into child
# processes, every nested shell -- crucially each tmux pane, which inherits
# the tmux server's env -- skips installing `history -a` and only persists
# history on a clean exit. Abrupt termination (docker stop, tmux kill-server,
# SIGKILL) then loses that shell's in-memory history. Keeping it unexported
# means each new interactive shell re-installs its own per-prompt flush.
if [ -z "${DEVBOX_HIST_SET:-}" ]; then
PROMPT_COMMAND="${PROMPT_COMMAND:+$PROMPT_COMMAND$'\n'}history -a"
export DEVBOX_HIST_SET=1
DEVBOX_HIST_SET=1
fi
# ── Prompt: show [opencode-devbox] tag so it's obvious you're in the container
+14
View File
@@ -0,0 +1,14 @@
# Global gitignore — personal/tooling artifacts (applies to all repos in the container)
# Seeded into $HOME/.gitignore_global by entrypoint-user.sh and wired via
# `git config --global core.excludesFile`. Edit freely; it is yours after first boot.
# backup / editor / merge artifacts
*.bak
*.bak.*
*~
*.orig
*.swp
*.tmp
# AI/LLM tool local settings — machine-specific perms + credentials, never commit
**/.claude/settings.local.json
+91
View File
@@ -0,0 +1,91 @@
#!/bin/sh
# devbox-skill-reconcile — hand skillset-OWNED skills back to the live clone.
#
# WHY THIS EXISTS
# ---------------
# entrypoint-user.sh links the image-baked skills into ~/.agents/skills/ EARLY
# (before pi-deploy), because the smoke readiness probe gates on markers that
# only land later, and a link created after that gate produced a flaky
# assertion. Those links are created with a `[ ! -e ]` guard — "only when
# absent" — and the skillset deploy runs LAST, treating already-present links
# as foreign and leaving them alone. Net effect through v1.8.4: the baked copy
# always won, so an edit pushed to a skillset-owned skill was invisible in
# every container until the next image build (measured on two hosts: live
# skillset md5 129bcc4752 vs baked 5236024fef, the new section absent).
#
# The fix is NOT "the skillset always wins". Ownership is per-skill (see
# rootfs/usr/local/share/pi-devbox/skills/VENDORED.md):
#
# pi-devbox-environment authored in pi-devbox → baked IS canonical
# pi-extensions owned by the package repo, copied over the snapshot
# at build time; skillset carries a DOWNSTREAM copy
# that can lag → baked must keep winning
# mempalace owned by the skillset repo; baked is a snapshot
# fallback for containers with no skillset mounted
# → the live clone must win when it is present
#
# So only skills listed in skills/skillset-owned.txt are handed over. Baked
# links stay as the fallback (the early-link race fix is untouched), and a user
# override always beats both: a real directory is never replaced, and neither is
# a symlink that already points somewhere other than the baked tree.
#
# Usage: devbox-skill-reconcile <skillset-root> [skills-dir] [baked-src]
# skillset-root the mounted skillset repo (contains skills/<name>/)
# skills-dir default $HOME/.agents/skills
# baked-src default /usr/local/share/pi-devbox/skills
#
# Idempotent, and silent unless it changes something. Exits 0 when there is
# nothing to do (no skillset, no list) so the entrypoint never fails on it.
set -eu
SKILLSET_ROOT="${1:-}"
SKILLS_DIR="${2:-$HOME/.agents/skills}"
BAKED_SRC="${3:-/usr/local/share/pi-devbox/skills}"
BAKED_SRC="${BAKED_SRC%/}" # a trailing slash would make the prefix
# match below ("$BAKED_SRC"/*) match nothing
[ -n "$SKILLSET_ROOT" ] || exit 0
[ -d "$SKILLSET_ROOT/skills" ] || exit 0
[ -d "$SKILLS_DIR" ] || exit 0
# Absolutise BOTH roots before they are used, because each has its own way of
# failing silently when relative: a relative symlink TARGET is resolved against
# the link's directory (~/.agents/skills), not $PWD, so it would dangle on
# creation; and a relative BAKED_SRC would never prefix-match the absolute
# target that `readlink` reports, so every skill would be skipped and the fix
# would look like it had simply done nothing.
SKILLSET_ROOT=$(CDPATH= cd -- "$SKILLSET_ROOT" 2>/dev/null && pwd) || exit 0
BAKED_SRC=$(CDPATH= cd -- "$BAKED_SRC" 2>/dev/null && pwd) || exit 0
OWNED_LIST="$BAKED_SRC/skillset-owned.txt"
[ -f "$OWNED_LIST" ] || exit 0
while IFS= read -r _line || [ -n "$_line" ]; do
# strip comments and surrounding whitespace; skip blanks
_name=$(printf '%s\n' "$_line" | sed -e 's/#.*$//' -e 's/^[[:space:]]*//' -e 's/[[:space:]]*$//')
[ -n "$_name" ] || continue
# defensive: a list entry must be a plain skill name, never a path
case "$_name" in */*|.*) continue ;; esac
_live="$SKILLSET_ROOT/skills/$_name"
_link="$SKILLS_DIR/$_name"
# the skillset does not ship it → the baked fallback is all there is
[ -d "$_live" ] || continue
# a real directory is a user override → never touch
[ -L "$_link" ] || continue
# only ever replace OUR OWN link. readlink is deliberate: `readlink -f`
# would resolve a link that already points into the skillset clone and,
# since both trees hold a same-named skill, could not tell them apart.
_target=$(readlink "$_link" 2>/dev/null || true)
case "$_target" in
"$BAKED_SRC"/*|"$BAKED_SRC") ;; # baked link → ours to replace
*) continue ;; # user/foreign target → leave alone
esac
# -n so an existing symlink-to-directory is replaced rather than followed
# (without it, ln would create $_link/$_name inside the baked tree).
if ln -sfn "$_live" "$_link" 2>/dev/null; then
printf 'skill %s: baked snapshot -> live skillset (%s)\n' "$_name" "$_live"
fi
done < "$OWNED_LIST"
+234
View File
@@ -0,0 +1,234 @@
#!/usr/bin/env bash
# pi-devbox-version — show which pi-devbox image build is running.
#
# WHY THIS EXISTS
# The image bakes ground-truth build info into /etc/pi-devbox/build-manifest.json
# at `docker build` time (see Dockerfile.variant): the release tag, build date,
# source commit, live `pi --version` at build time, and the actual checked-out
# commit of every /opt component clone. That answers "what image am I running?"
# — but only if you know to go look for the file. This wraps it into one
# command, prints it human-first at container start (see entrypoint-user.sh),
# and stays available on demand for the rest of the session.
#
# USAGE
# pi-devbox-version human-readable summary (default)
# pi-devbox-version --json raw manifest JSON (for scripting)
# pi-devbox-version --quiet one-line "release_tag (source_revision)" form
# pi-devbox-version --no-skills skip the skill-source section (used at
# container start, where it would be premature)
#
# EXIT STATUS
# 0 on success. 1 if the manifest is missing (e.g. an image built before
# this file existed, or a non-pi-devbox base) — prints a short notice
# to stderr rather than failing silently.
set -euo pipefail
MANIFEST=/etc/pi-devbox/build-manifest.json
MODE="human"
SHOW_SKILLS="yes"
# A `case "${1:-}"` here only ever looked at the FIRST argument, so
# `--no-skills --json` matched --no-skills, silently dropped --json, and
# printed human text to a caller expecting JSON (a real failure: a jq
# consumer piping that output gets a parse error, not a wrong-but-parseable
# answer). Loop over every argument instead, and reject anything unknown
# rather than silently ignoring it the same way.
for _arg in "$@"; do
case "$_arg" in
--json) MODE="json" ;;
--quiet|-q) MODE="quiet" ;;
--no-skills) SHOW_SKILLS="no" ;;
--help|-h)
# Print the leading `#`-comment block verbatim, stopping at the first
# non-comment line, rather than a hardcoded line range: `sed -n
# '2,22p'` was silently truncating --help because this file has grown
# usage lines since that range was written, and a fixed range will
# drift again the next time a comment is added above it.
awk 'NR==1{next} /^#/{sub(/^# ?/,""); print; next} {exit}' "$0"
exit 0
;;
*)
echo "pi-devbox-version: unknown option: $_arg" >&2
echo " try --help" >&2
exit 2
;;
esac
done
if [ ! -f "$MANIFEST" ]; then
echo "pi-devbox-version: no build manifest at $MANIFEST" >&2
echo " (image predates the manifest, or this isn't a pi-devbox image)" >&2
exit 1
fi
if ! command -v jq >/dev/null 2>&1; then
echo "pi-devbox-version: jq not found; dumping raw manifest instead" >&2
cat "$MANIFEST"
exit 0
fi
if [ "$MODE" = "json" ]; then
cat "$MANIFEST"
exit 0
fi
release_tag=$(jq -r '.release_tag' "$MANIFEST")
build_date=$(jq -r '.build_date' "$MANIFEST")
source_rev=$(jq -r '.source_revision' "$MANIFEST")
pi_version_baked=$(jq -r '.pi_version' "$MANIFEST")
# `// empty` matters: images built before v1.8.6 have no such field, and
# `jq -r` renders a JSON null as the 4-char string "null" — which would
# print as a bogus version rather than being treated as absent.
mp_version_baked=$(jq -r '.mempalace_version // empty' "$MANIFEST")
if [ "$MODE" = "quiet" ]; then
printf '%s (%s)\n' "$release_tag" "${source_rev:0:7}"
exit 0
fi
# Live drift check: has `pi` been upgraded since this container was built?
# (image is immutable, but a volume-persisted ~/.pi could in theory shadow
# the baked binary — this stays honest rather than trusting the manifest
# blindly, same "ground truth over intent" spirit as how the manifest
# itself is generated in Dockerfile.variant.)
pi_version_live=""
if command -v pi >/dev/null 2>&1; then
pi_version_live=$(pi --version 2>/dev/null | head -n1 | tr -d '\r\n')
fi
# Same check for the palace, which matters more than it looks: mempalace is
# the one component that is BOTH client (here) and server (synlig runs this
# same image), so a skew between the two is a real failure mode rather than
# cosmetic. `mempalace --version` prints "MemPalace 3.8.0" — name-prefixed,
# unlike pi's bare "0.84.3" — hence $NF rather than reading the whole line.
mp_version_live=""
if command -v mempalace >/dev/null 2>&1; then
mp_version_live=$(mempalace --version 2>/dev/null | head -n1 | awk '{print $NF}' | tr -d '\r\n')
fi
printf 'pi-devbox %s\n' "$release_tag"
printf ' built: %s (source %s)\n' "$build_date" "${source_rev:0:12}"
if [ -n "$pi_version_live" ] && [ "$pi_version_live" != "$pi_version_baked" ]; then
printf ' pi: %s \033[33m(baked as %s — drift detected)\033[0m\n' "$pi_version_live" "$pi_version_baked"
else
printf ' pi: %s\n' "${pi_version_live:-$pi_version_baked}"
fi
# Printed only when known, so this degrades quietly on pre-v1.8.6 images
# instead of showing an empty or "null" palace line.
if [ -n "$mp_version_live" ] || [ -n "$mp_version_baked" ]; then
if [ -n "$mp_version_live" ] && [ -n "$mp_version_baked" ] && [ "$mp_version_live" != "$mp_version_baked" ]; then
printf ' palace: %s \033[33m(baked as %s — drift detected)\033[0m\n' "$mp_version_live" "$mp_version_baked"
else
printf ' palace: %s\n' "${mp_version_live:-$mp_version_baked}"
fi
fi
printf ' components:\n'
jq -r '.components | to_entries[] | select(.value != null) | " \(.key): \(.value[0:12])"' "$MANIFEST"
# ── Which copy of each vendored skill is actually being read? ─────────
# The image bakes fallback skills under /usr/local/share/pi-devbox/skills/,
# but for skills the skillset repo OWNS (skillset-owned.txt) a mounted live
# clone takes over at container start via devbox-skill-reconcile. Nothing
# reported which copy won, so a stale baked snapshot and a current live clone
# looked identical from inside — and on this fleet the baked mempalace copy is
# read by NOBODY (all four compose stacks mount a workspace containing the
# skillset), which is exactly the sort of fact that should be visible rather
# than reasoned about. Same "drift detected" shape as the pi/palace lines
# above: what is live, annotated with what was baked, when they disagree.
#
# Skipped with --no-skills at container start (entrypoint-user.sh calls this
# FIRST, before the baked links exist and long before the skillset deploy and
# reconcile run last), because a section that is accurate only after boot
# finishes is worse than no section at all.
BAKED_SKILLS=/usr/local/share/pi-devbox/skills
SKILLS_DIR="${HOME:-/home/developer}/.agents/skills"
if [ "$SHOW_SKILLS" = "yes" ] && [ -d "$BAKED_SKILLS" ] && [ -d "$SKILLS_DIR" ]; then
# Recorded provenance of the vendored mempalace snapshot (absent on images
# built before this existed — `// empty` so a JSON null never prints as the
# 4-char string "null", the same trap noted for mempalace_version above).
# `_tree_sha256`, not `_sha256`: it is a hash over every file in the
# vendored skill DIRECTORY (see tree_sha256() below), not one file, because
# a single-file hash reports "identical" against a live checkout that added
# or edited a sibling file — pi-extensions already ships two files, so this
# is not hypothetical.
snap_ref=$(jq -r '.skillset_snapshot_ref // empty' "$MANIFEST")
snap_sha=$(jq -r '.skillset_snapshot_tree_sha256 // empty' "$MANIFEST")
# Same pipeline Dockerfile.variant uses to measure the baked directory at
# build time: relative paths in `find | sort` order, each hashed, the whole
# listing folded into one sha256. Keep the two definitions identical — they
# run in different processes (image build vs. this container) and are
# meaningless to compare unless they agree byte-for-byte on the algorithm.
tree_sha256() {
( cd "$1" && find . -type f -print | LC_ALL=C sort | xargs -r sha256sum ) 2>/dev/null | sha256sum | cut -d' ' -f1
}
# Iterate the baked tree rather than a hardcoded name list, so vendoring a
# fourth skill needs no edit here. The header prints only if the tree is
# non-empty, so this can never emit a dangling "skills:" label.
_printed_header="no"
for _dir in "$BAKED_SKILLS"/*/; do
[ -d "$_dir" ] || continue
if [ "$_printed_header" = "no" ]; then
printf ' skills:\n'
_printed_header="yes"
fi
_name=$(basename "$_dir")
_link="$SKILLS_DIR/$_name"
if [ ! -e "$_link" ]; then
printf ' %-22s not linked\n' "$_name"
continue
fi
_target=$(readlink -f "$_link" 2>/dev/null || echo "$_link")
case "$_target" in
"$BAKED_SKILLS"/*|"$BAKED_SKILLS")
printf ' %-22s baked\n' "$_name"
continue
;;
esac
# Outside the baked tree: a mounted skillset clone, or a user override.
# The link target is <repo>/skills/<name>, so the repo root is two up.
# Everything here is guarded: this script runs on the container-start path
# and must never fail, and `set -e` is in force.
_root=$(cd "$_target/../.." 2>/dev/null && pwd) || _root=""
_head=""
if [ -n "$_root" ]; then
_head=$(git -C "$_root" rev-parse HEAD 2>/dev/null || echo "")
fi
_where="live ${_root:-$_target}"
[ -n "$_head" ] && _where="$_where @ ${_head:0:7}"
# For the one skill whose baked fingerprint we recorded, say plainly
# whether the live copy differs from what shipped. This is the check CI
# cannot perform (the skillset is private) and the container can, free.
# Hash the whole live DIRECTORY with the same tree_sha256() used to
# measure the baked one in Dockerfile.variant — a SKILL.md-only compare
# would silently ignore a changed or added sibling file.
_live_sha=""
if [ -n "$snap_sha" ] && [ "$_name" = "mempalace" ] && [ -d "$_target" ]; then
_live_sha=$(tree_sha256 "$_target")
fi
if [ -z "$_live_sha" ]; then
printf ' %-22s %s\n' "$_name" "$_where"
elif [ "$_live_sha" = "$snap_sha" ]; then
printf ' %-22s %s (identical to baked snapshot)\n' "$_name" "$_where"
elif [ -n "$_head" ] && [ "$_head" = "$snap_ref" ]; then
# Same commit, different bytes — i.e. uncommitted edits in the live
# checkout. Distinguished from plain drift because otherwise the line
# reads as a self-contradiction ("@ c04cd15 ... baked snapshot c04cd15
# — live copy differs") and a reader would suspect the tool, not the
# working tree.
printf ' %-22s %s \033[33m(baked snapshot %s + uncommitted edits)\033[0m\n' \
"$_name" "$_where" "${snap_ref:0:7}"
else
printf ' %-22s %s \033[33m(baked snapshot %s — live copy differs)\033[0m\n' \
"$_name" "$_where" "${snap_ref:0:7}"
fi
done
fi
@@ -14,7 +14,9 @@
# The one thing reachable from a container on every OS is the host itself
# (host.docker.internal). So on VM-backed hosts we generate a writable SSH
# config that reaches the host and lets the user ProxyJump onward to LAN
# peers the host can reach. On native Linux we do nothing.
# peers the host can reach. On native Linux we render the same writable
# config (for the ControlPath redirect + Include ~/.ssh/config) but emit no
# jump block, since LAN peers are reachable directly there.
#
# We ship the MECHANISM (a generic `host` jump alias + writable config),
# never the POLICY: the user's specific target hosts live in their own
@@ -30,7 +32,9 @@
#
# CONTROLS (env)
# DEVBOX_LAN_ACCESS = auto (default) | jump | off
# auto → set up the jump config only on VM-backed hosts; no-op on Linux.
# auto → set up the host jump only on VM-backed hosts. The writable
# sidecar config (ControlPath redirect + Include) is always
# rendered, on every OS.
# jump → always set up (e.g. native Linux with extra_hosts host-gateway).
# off → do nothing.
# HOST_SSH_USER — the username to SSH into the host as. REQUIRED for the
@@ -84,40 +88,70 @@ is_vm_backed() {
getent hosts "$HOST_ALIAS_HOSTNAME" >/dev/null 2>&1
}
if [ "$MODE" = "auto" ] && ! is_vm_backed; then
# Native Linux host: LAN peers are reachable directly. Nothing to do.
exit 0
fi
# From here: MODE=jump, or MODE=auto on a VM-backed host.
command -v ssh-keygen >/dev/null 2>&1 || exit 0
# ── Writable socket dir + sidecar (ALWAYS, every host OS) ─────────────
# The ControlPath redirect in the generated config needs a writable directory
# regardless of host OS or jump mode. ~/.ssh is typically read-only, so the
# master socket lives under the writable ~/.ssh-local. We create it and render
# the config UNCONDITIONALLY so the redirect (and `Include ~/.ssh/config`) works
# even on native Linux — where we set up no host jump but a read-only ~/.ssh
# would otherwise still break ControlMaster sockets.
mkdir -p "${SSH_LOCAL}/cm" 2>/dev/null || true
chmod 700 "${SSH_LOCAL}" "${SSH_LOCAL}/cm" 2>/dev/null || true
# ── Jump key (generated once; preserved across restarts) ──────────────
# ── Decide whether to set up the host jump ────────────────────────────
# Jump = reach the container host (host.docker.internal) as an SSH ProxyJump
# onward to the host's LAN peers. Needed on VM-backed hosts (macOS / Docker
# Desktop) or when forced with DEVBOX_LAN_ACCESS=jump. On native Linux LAN
# peers are reachable directly, so NEED_JUMP=0 and we emit no jump block — but
# we still render the config for the ControlPath redirect + Include.
NEED_JUMP=0
if [ "$MODE" = "jump" ] || { [ "$MODE" = "auto" ] && is_vm_backed; }; then
NEED_JUMP=1
fi
# ── Jump key (only when a jump is needed; generated once, preserved) ──
# Persisted via a named volume on ~/.ssh-local (see compose), so a fresh key
# is generated only on the very first start (or if the volume is wiped). When
# we DO generate one it must be (re-)authorized on the host, so we flag it and
# print a copy-paste authorize line below.
KEY_JUST_GENERATED=0
if [ ! -f "$KEY" ]; then
ssh-keygen -t ed25519 -N '' -C "devbox-jump@${HOSTNAME:-container}" -f "$KEY" >/dev/null 2>&1 || exit 0
if [ "$NEED_JUMP" = "1" ] && command -v ssh-keygen >/dev/null 2>&1 && [ ! -f "$KEY" ]; then
if ssh-keygen -t ed25519 -N '' -C "devbox-jump@${HOSTNAME:-container}" -f "$KEY" >/dev/null 2>&1; then
chmod 600 "$KEY" 2>/dev/null || true
KEY_JUST_GENERATED=1
fi
fi
# ── Render the writable config ────────────────────────────────────────
# Jump-specific blocks (the host alias, host-owned peer overrides, and the
# optional RFC1918 catch-all) only make sense when a jump is set up; on native
# Linux they are all empty and only the ControlPath redirect + Include remain.
JUMP_BLOCK=""
LAN_CONF_BLOCK=""
AUTOJUMP_BLOCK=""
if [ "$NEED_JUMP" = "1" ]; then
USER_LINE=""
if [ -n "${HOST_SSH_USER:-}" ]; then
USER_LINE=" User ${HOST_SSH_USER}"
fi
JUMP_BLOCK=$(cat <<EOF
# The container host (OrbStack / Docker Desktop). 'host' and 'mac' are aliases.
Host host mac
HostName ${HOST_ALIAS_HOSTNAME}
${USER_LINE}
IdentityFile ~/.ssh-local/devbox_jump_ed25519
IdentitiesOnly yes
ControlMaster auto
ControlPath ~/.ssh-local/cm/%r@%h:%p
ControlPersist 4h
ServerAliveInterval 30
EOF
)
# Optional host-owned named-peer jump overrides (portable: lives on the host,
# not in the image). Included BEFORE ~/.ssh/config so its ProxyJump wins.
SSH_LAN_CONF="${HOME}/.config/devbox-shell/ssh-lan.conf"
LAN_CONF_BLOCK=""
if [ -r "$SSH_LAN_CONF" ]; then
LAN_CONF_BLOCK=$(cat <<'EOF'
@@ -132,7 +166,6 @@ fi
# Optional opt-in RFC1918 catch-all: ProxyJump every private IP through the
# host. Matches the typed address, never the resolved HostName, so named hosts
# with their own ProxyJump are unaffected. Network-agnostic → roaming-safe.
AUTOJUMP_BLOCK=""
if [ "${DEVBOX_LAN_AUTOJUMP_PRIVATE:-0}" = "1" ]; then
AUTOJUMP_BLOCK=$(cat <<'EOF'
@@ -147,6 +180,7 @@ Host 10.* 192.168.* 172.16.* 172.17.* 172.18.* 172.19.* 172.20.* 172.21.* 172.22
EOF
)
fi
fi
INCLUDE_BLOCK=""
if [ -r "${HOME}/.ssh/config" ]; then
@@ -154,13 +188,54 @@ if [ -r "${HOME}/.ssh/config" ]; then
# Your own target hosts. Scope reset to match-all so this Include applies to
# every target (an Include is otherwise scoped to the enclosing Host block).
# Add 'ProxyJump host' to LAN entries here (or in ssh-lan.conf above).
# To make a LAN peer jump via the host, add 'ProxyJump host' to its entry in
# the host-owned ~/.config/devbox-shell/ssh-lan.conf (Included above) — NOT
# here in ~/.ssh/config, which is typically bind-mounted read-only.
Host *
Include ~/.ssh/config
EOF
)
fi
# ── Multiplexing default, deliberately LAST ───────────────────────────
# Why this block exists: ControlPath above is forced, but ControlMaster is not
# set anywhere for targets that come from the user's own ~/.ssh/config. A target
# whose entry omits ControlMaster therefore opens a NEW TCP connection per ssh
# call, and an agent doing a dozen calls in a few minutes can trip fail2ban or a
# CGNAT flow-table cap on the far end — observed 2026-08-25: ~12 connections in
# 15 min and port 22 stopped answering while HTTPS to the same estate stayed fine.
#
# WHY IT IS AT THE BOTTOM, and ControlPath is at the top. ssh_config is
# first-value-wins, so position encodes intent:
# * BEFORE the Include = an OVERRIDE. Correct for ControlPath, whose value in
# the user's config points at read-only ~/.ssh and simply cannot work here.
# * AFTER the Include = a DEFAULT. Correct for ControlMaster, because an
# explicit per-host 'ControlMaster no' (or 'auto', or any value) in the
# user's own config must keep winning. We are supplying an opinion only
# where the user expressed none.
# That asymmetry is the whole design: force what is broken, default what is
# merely absent. It also means this needs no audit of anyone's ~/.ssh/config —
# which matters because that file is per-machine, differs across the fleet, and
# future machines' versions do not exist yet to be audited.
#
# Caveat worth knowing (and documented in the pi-devbox-environment skill): a
# stale master socket — file present, daemon gone, e.g. after the host suspends
# or changes network — makes every later ssh to that host hang. Recovery is
# 'ssh -F ~/.ssh-local/config -O exit <host>'. ControlPersist is deliberately
# short (10m idle, and each new session resets the idle timer) so an abandoned
# socket ages out on its own rather than lingering for hours.
MULTIPLEX_DEFAULT_BLOCK=$(cat <<'EOF'
# Multiplexing DEFAULT — intentionally after the Include above, so any explicit
# per-host ControlMaster in your own ~/.ssh/config still wins (first-value-wins).
# Applies only to targets that never mentioned ControlMaster at all.
# Stale socket after a suspend/network change? ssh -O exit <host>.
Host *
ControlMaster auto
ControlPersist 10m
EOF
)
cat > "$CONFIG" <<EOF
# AUTO-GENERATED by setup-lan-access.sh on every container start. Do not edit
# by hand — edits are overwritten. Used via: ssh -F ~/.ssh-local/config <host>
@@ -176,20 +251,11 @@ Host *
UserKnownHostsFile ~/.ssh-local/known_hosts
StrictHostKeyChecking accept-new
ControlPath ~/.ssh-local/cm/%r@%h:%p
# The container host (OrbStack / Docker Desktop). 'host' and 'mac' are aliases.
Host host mac
HostName ${HOST_ALIAS_HOSTNAME}
${USER_LINE}
IdentityFile ~/.ssh-local/devbox_jump_ed25519
IdentitiesOnly yes
ControlMaster auto
ControlPath ~/.ssh-local/cm/%r@%h:%p
ControlPersist 4h
ServerAliveInterval 30
${JUMP_BLOCK}
${LAN_CONF_BLOCK}
${AUTOJUMP_BLOCK}
${INCLUDE_BLOCK}
${MULTIPLEX_DEFAULT_BLOCK}
EOF
chmod 600 "$CONFIG" 2>/dev/null || true
@@ -199,6 +265,7 @@ chmod 600 "$CONFIG" 2>/dev/null || true
# host won't recognize. With ~/.ssh-local persisted via a named volume, case
# (b) fires only on first-ever start (or after the volume is reset) — so this
# is normally a one-time, one-line step per machine, with no file to locate.
if [ "$NEED_JUMP" = "1" ]; then
PUBKEY_TEXT="$(cat "${KEY}.pub" 2>/dev/null)"
if [ -z "${HOST_SSH_USER:-}" ]; then
cat <<EOF
@@ -221,5 +288,6 @@ elif [ "$KEY_JUST_GENERATED" = "1" ]; then
repeat this on container updates — only if that volume is reset.
EOF
fi
fi
exit 0
@@ -0,0 +1,72 @@
<!-- pi-devbox:managed-block — appended to the global AGENTS.md at image build
time (Dockerfile.variant), after pi-toolkit is cloned. Keep this short:
it is a pointer, the depth lives in the skill. -->
## Running inside pi-devbox
If the directory `/usr/local/lib/pi-devbox/` exists (or your shell prompt is
prefixed `[devbox]`, or `~/.ssh-local/config` is present), you are in a
**pi-devbox container** — a Docker environment whose persistence, networking,
DNS, host/LAN reachability, tmux, and Python/REPL behaviour differ from a normal
workstation. Before any task that touches **reaching the host or its LAN, SSH,
DNS/name resolution, what survives container recreate, running Python/REPLs,
tmux, or pi-studio**, read `~/.agents/skills/pi-devbox-environment/SKILL.md`.
Key reflex from that skill: **the deployment specifics are not universal** — the
host OS, hostnames, internal domains, and nameservers vary per instance and must
be discovered at runtime, never assumed. And interactive shell aliases
(`dssh`, `dscp`, `cat`→`bat`) do **not** exist in your non-interactive bash
tool, so spell out the underlying command (e.g.
`ssh -F "$HOME/.ssh-local/config" mac …`).
## Browser automation is available (agent-browser)
This image bakes the **`agent-browser`** CLI plus a headless Chromium, so you can
drive a real browser — open pages, click/fill/`eval`, snapshot the DOM, take
screenshots — to **verify** front-end work (live DOM, WebGL, layout, popup
positioning) instead of guessing. Reach for it whenever a task involves a web UI
or checking how a page actually renders. `AGENT_BROWSER_EXECUTABLE_PATH` is
preset to the baked browser, so `agent-browser open <url>` works out of the box
(headless). Run `agent-browser skills get core --full` for the command set and
workflow patterns (always version-matched to the CLI); the `agent-browser` skill
under `~/.agents/skills/` mirrors it when the skillset is mounted.
## Session start: load the mempalace skill
If MemPalace MCP tools (e.g. `mempalace_search`, `mempalace_diary_write`) are in
your tool list, **read `~/.agents/skills/mempalace/SKILL.md` before doing
non-trivial work** and follow its protocol: search the palace before answering
about past work, and write a diary entry before the session ends. This is
especially load-bearing here — a pi-devbox container is frequently recreated, so
the palace is your only memory across recreates. Without the habit it is just
storage, not memory. (The skill is the consumer side; feeding the palace is the
separate `opencode-mempalace-bridge` skill, if present.)
### If the palace is central, it is shared — three rules
If `MEMPALACE_REMOTE_URL` is set, the MCP tools write to a **central palace
shared with other machines**, not to a local one. Your drawers are not the only
ones in there, and most drawers' `source_file` paths do not exist on this host.
The skill covers the orientation side (provenance, chronology, whose diary is
whose); these three are here instead because getting them wrong does *damage*
rather than merely confusing you:
- **Never run `mempalace sync` / `mempalace_sync` against a shared palace.** It
prunes drawers whose source files look gitignored, deleted, or moved — and on
a shared palace that describes most of the content, including every other
machine's. Compounding it (RFC-001 §7.2): feeders now stage *inside* the
palace root, so a scoped sync can delete the very drawers it just filed.
`mempalace_delete_by_source` is exact-match rather than existence-based, but
its blast radius is now the whole fleet's palace — leave it on its default
`dry_run=true` and confirm the match count before committing.
- **A timeout is not a failure.** The palace is single-writer, and one large
mine can block every client for minutes, so a write or mine that exceeds the
client's deadline has usually *completed* server-side. Verify with
`mempalace_get_drawer` or `mempalace_search` before retrying — a blind retry
files a duplicate. `[mempalace ext] feed (tick) failed: mine timed out after
30000ms` is the common benign instance: the transcript is already in the
server's inbox and the mine is idempotent, so nothing is lost either way.
- **The `mempalace` CLI is not remote-aware.** It always opens a palace on
local disk, so `mempalace search` can return older and different results than
the MCP tools while both look correct. Use the MCP tools for the central
palace; the CLI only for a local one.
@@ -0,0 +1,146 @@
# Vendored fallback skills
Most directories here are **image-baked skills** that `entrypoint-user.sh`
symlinks into `~/.agents/skills/` on container start. They are the **fallback**
layer: see *Runtime precedence* below for which copy actually wins when a
`skillset` repo is mounted (through v1.8.4 the answer was "always the baked
one", which was a bug).
| skill | owner | how it gets here |
|-------|-------|------------------|
| `pi-devbox-environment` | pi-devbox (this repo) | authored here; the canonical copy |
| `pi-extensions` | the `pi-extensions` package repo (`skill/`) | **vendored fallback** + refreshed at build |
| `mempalace` | the `skillset` repo | **vendored fallback** (snapshot only) |
## Why fallbacks exist
The pi-toolkit global `AGENTS.md` tells every pi session to read
`~/.agents/skills/pi-extensions/SKILL.md` at start (to fix fork/recall
under-utilisation). That pointer dangles in a container started **without** the
private `skillset` repo mounted. Baking the skill closes that *availability*
gap. `mempalace` is baked for the same reason (memory continuity); since
nothing in pi-toolkit's `AGENTS.md` points to it, the pi-devbox managed block
(`pi-global-AGENTS.append.md`) also adds the matching *proactive-load*
directive ("load the mempalace skill at session start") so a new container
actually picks it up rather than relying on description-matching.
`pi-extensions`'s directive already ships in pi-toolkit's `AGENTS.md`, so only
its skill file needed baking.
## Freshness model (layered — see Dockerfile.variant)
- **`pi-extensions`** — Option 1 + Option 2. The committed copy here is the
*floor*; at build time `Dockerfile.variant` copies `/opt/pi-extensions/skill/`
(the pinned, package-owned source) over it, so a normal build ships the fresh
package copy and a stale-ref / mirror build still ships the snapshot. Keep
`evaluate-extension-usage.py` alongside `SKILL.md` — the skill calls it via
`./`.
- **`mempalace`** — Option 2 only. The `mempalace` *consumer* skill lives only
in the private `skillset` repo (the `mempalace-toolkit` repo ships a
*different* skill, `opencode-mempalace-bridge`), so there is no public
package source to copy from. This snapshot is refreshed manually per release.
**Refresh it with `scripts/vendor-mempalace-skill.sh <skillset-root>`, not
`cp`.** Because the image cannot clone the private upstream, the snapshot used
to be *anonymous* — nothing recorded which skillset commit the bytes came
from, so the only staleness check possible was a hand-maintained phrase canary
in `scripts/smoke-test.sh`, which by construction detects "older than the
phrase I remembered to pin", never "older than skillset main". Two facts now
travel with the file:
| Fact | Where | Kind |
|---|---|---|
| `ARG SKILLSET_SNAPSHOT_REF` in `Dockerfile.variant` | manifest `skillset_snapshot_ref` + OCI label `se.jordbo.pi-devbox.skillset-snapshot-ref` | a **claim** about which commit these bytes are |
| `sha256sum` of this file, measured in the manifest layer | manifest `skillset_snapshot_sha256` | the bytes that **actually shipped** |
The script writes both together, refuses when the upstream file has
uncommitted modifications (no commit describes those bytes), and
`--check` verifies the claim against a real clone. Deliberately an `ARG`
default rather than a CI-resolved value: no credential for a private repo, no
change at any of the four `Dockerfile.variant` build call sites, and a local
`docker build` records the same thing CI does.
Verifying "is this snapshot current?" is **not** a CI job and was deliberately
not made one — see the Unreleased CHANGELOG entry for why (private repo;
another repo's branch must not be able to fail this build; and the artefact it
would guard is read by no host on this fleet). The check belongs where the
skillset actually is: `vendor-mempalace-skill.sh --check` for a maintainer,
and `pi-devbox-version`'s `skills:` section for an agent inside a container.
## Runtime precedence (v1.8.5+)
The baked links are created **early** in `entrypoint-user.sh` (before pi-deploy,
to close a smoke readiness race) with a create-only-when-absent guard, and the
skillset deploy runs **last** and treats them as foreign links. Through v1.8.4
that combination meant the baked snapshot always won: an edit pushed to
`skillset/skills/mempalace/SKILL.md` was invisible in every container until the
next image build (measured on two hosts — live `md5 129bcc4752` vs baked
`5236024fef`, new section absent). Editing those skills *appeared* to work.
`devbox-skill-reconcile` now runs immediately after the skillset deploy and
repoints the links for skills the **skillset owns**, listed one per line in
`skillset-owned.txt`. Precedence, highest first:
1. **user override** — a real directory, or a symlink pointing outside the baked
tree; never touched by anything
2. **live skillset clone** — but only for names in `skillset-owned.txt`
3. **baked snapshot** — everything else, and every skill when no skillset is
mounted
**Which one won is now reportable from inside the container:**
`pi-devbox-version` prints a `skills:` section naming, per vendored skill,
`baked` or `live <repo> @ <sha>` — and for `mempalace` whether that live copy is
identical to the baked fingerprint, at the same commit but with uncommitted
edits, or genuinely divergent. Before that, a stale baked snapshot and a current
live clone were indistinguishable from inside, which is how the freshness of
this file went unexamined for three releases. The section is suppressed with
`--no-skills` on the container-start banner, because `entrypoint-user.sh` prints
the version *before* the links exist and long before the reconcile below runs.
On this fleet, precedence 2 wins for `mempalace` on **every** host — all four
compose stacks mount a workspace containing the skillset — so the baked copy is
exercised only by CI and by a hypothetical no-mount container. Worth
remembering before spending effort on its freshness.
Ownership is per-skill on purpose: `pi-extensions`' authoritative source is the
package repo (copied over the snapshot at build), and `skillset` carries a
downstream copy that can lag, so handing it to the clone would *regress* the
skill. Only `mempalace` is skillset-owned today.
Verify with `readlink -f ~/.agents/skills/<skill>` — not by reading the
entrypoint. Smoke covers both directions (baked resolution with no skillset
mounted, plus a fabricated-skillset run of the reconciler).
## Refreshing the snapshots
cp <pi-extensions-pkg>/skill/SKILL.md pi-extensions/SKILL.md
cp <pi-extensions-pkg>/skill/evaluate-extension-usage.py pi-extensions/
Copy `pi-extensions` **from its owner in the table above** — the package
repo's `skill/` (since `a7f3044` co-located it there; `skillset` also carries a
copy, but it is a downstream duplicate and can lag). Copying `pi-extensions`
from `skillset` would regress the snapshot to whatever that repo last mirrored.
`mempalace` is **not** refreshed by `cp` — see the *Freshness model* section
above: `scripts/vendor-mempalace-skill.sh <skillset-root>` is the only thing
that should ever touch that snapshot, because a bare copy can update the bytes
without updating the ref that claims to describe them, which produces a
manifest that confidently lies.
Neither vendored skill has a hand-maintained "last refreshed at" line here on
purpose — one previously existed (skillset `670f7f1`, pi-extensions pkg
`e73cb9f`) and went stale within hours, because nothing forced it to move
when the ARGs did. `670f7f1` is now a cautionary example rather than a fact
worth recording: it is the commit that told agents to hand-stamp `added_by`,
which a later skillset commit (and the pi-devbox edge stamper) withdrew — so a
reader trusting that line would have been pointed at superseded guidance.
Both facts it tried to capture now live somewhere that cannot drift by hand:
| Fact | Where |
|---|---|
| which skillset commit `mempalace`'s bytes came from | `ARG SKILLSET_SNAPSHOT_REF` (Dockerfile.variant) + `skillset_snapshot_ref` in `build-manifest.json`, written *only* by `vendor-mempalace-skill.sh` |
| which pi-extensions package commit was vendored | `ARG PI_EXTENSIONS_REF` (Dockerfile.variant, CI-resolved to a 40-hex commit) → OCI label `se.jordbo.pi-devbox.pi-extensions-ref` and `build-manifest.json`'s `components.pi-extensions`, both read from the actual `/opt/pi-extensions` checkout, not from intent |
When you refresh the `mempalace` snapshot, also update the phrase asserted by
the "mempalace skill snapshot is current" smoke test — it deliberately pins the
**newest** section, because the previous canary grepped a phrase that survived
the very edit that made the snapshot stale, and so passed on stale content.
@@ -0,0 +1,603 @@
---
name: mempalace
description: MemPalace agent memory protocol. Use on every session to maintain continuity across conversations — search before answering about past work, write diary entries before session ends, and mine new projects into the palace. Load this skill at session start.
---
# MemPalace Agent Memory Protocol
## Overview
MemPalace gives you persistent memory across sessions via an MCP server. It stores project knowledge (mined from files), conversation summaries (diary entries), and entity relationships (knowledge graph). Without this protocol, you have tools but no habits — and memory without habits is just storage.
**Core principle:** Storage is not memory. Storage + protocol = memory.
## When to Load This Skill
- At the **start of every session** (proactively, before the user asks)
- When the user mentions **past conversations, decisions, or work**
- When working on a **new project or repository** for the first time
- When the user asks about **people, projects, or relationships**
## Session Lifecycle
### Phase 1: Wake Up (session start)
Run these immediately when a session begins, before responding to the user:
1. **Load palace overview:**
```
mempalace_status
```
This returns wing/room counts, the AAAK spec, and the memory protocol reminder.
2. **Read your recent diary:**
```
mempalace_diary_read(agent_name="<your_agent_name>", last_n=5)
```
Scan for context about recent sessions — what was worked on, what matters, what's pending.
3. **Check the knowledge graph** for the user or active project if relevant:
```
mempalace_kg_query(entity="<project_or_person>")
```
4. **Check your mailbox.** Just run it — an empty result is a fine answer and
costs one call. Do not try to decide first whether coordination "applies to
you"; that test is what used to be wrong here (see *Cross-Machine
Coordination* below):
```
mempalace_event_list(to_agent="<harness>@<device>", status="open")
```
This is a candidate list, not a to-do list — `status` never changes after an
event is written, so finished asks keep matching. Subtract the ones you have
already answered using the rule in *What you actually owe*, below.
Another machine may have asked you something, or corrected something you are
about to rely on. This costs one call and is the only way you will find out:
nothing pushes an event into your session unless your bridge delivers it for
you, and if it does you will already have seen it before reading this.
Do NOT announce this to the user. Just do it silently to orient yourself.
### Temporal grounding — compute time deltas, don't guess
Diary entries and drawers carry real timestamps (`timestamp`, `created_at`).
Before describing *when* something happened — "yesterday", "earlier today",
"last week", "a while back" — **establish the current date/time first and
compute the delta against the actual timestamp.** Get "now" from the injected
session date or by running `date` in a shell; never infer it.
**A container recreate or a fresh session is NOT a day boundary.** A devbox
container (pi-devbox or opencode-devbox) is frequently restarted — often several
times within the *same* day — and each restart begins a new session with a fresh
wake-up. Do not reason "new session ⇒ last session was yesterday": two diary
entries 90 minutes apart can straddle a container recreate. The only
authoritative clock is the timestamp on the memory, not the session/container
boundary.
**Practical rule:** prefer explicit, checkable phrasing — e.g. "earlier today,
~8h ago (both 2026-06-25)" — over a vague relative term. If you catch yourself
about to write "yesterday" / "last week", subtract `now − entry.timestamp` and
state the computed result. (Remember timestamps may be UTC while the wall clock
is local — reconcile the offset before computing the delta.) Note too that
session feeders can lag up to a week (see *Multi-harness palace*), so a recent
absence in `wing_conversations` is not proof nothing happened.
### Phase 2: Active Session (during work)
#### Search Before You Speak
Before answering questions about past work, decisions, people, or projects:
```
mempalace_search(query="<keywords>", wing="<project>")
```
**Never guess about facts that might be in the palace.** Wrong is worse than slow. Say "let me check" and query.
#### Search Before You *Probe*
The rule above covers **questions**. This one covers **actions** — and it is the one
that actually gets skipped, because mid-task the impulse is to go and *look* rather
than to remember. The palace is a **fleet** record: another machine's agent has
usually already paid the cost of discovering how this environment is wired, and its
notes include the corrections that came afterwards, which a fresh probe cannot show
you.
**Before you SSH somewhere to find out how it is set up, enumerate infrastructure,
or derive a deployment — search.** Concrete triggers, all meaning *search first*:
- about to run `ssh <host> …`, `docker ps`, `systemctl list-units`, `ip addr` to
discover how something is deployed or connected
- about to establish topology: which hosts/runners/services exist, where they live,
which of them can reach which
- about to conclude "this isn't documented anywhere" or "there's no way to know"
- about to assert an environment fact you learned **earlier in this same session**
**That last trigger is the sharp edge.** A compacted session summary is lossy by
design, and a belief you formed 40 turns ago may already be *retracted* in the
palace by another machine. Trusting your own context over the shared record is how a
withdrawn claim gets re-published as fact.
Search broadly before narrowing — fleet knowledge often sits in another machine's
wing, or inside a mined conversation, not where you would file it yourself:
```
mempalace_search(query="<topic> <host> <mechanism>") # no wing filter first
mempalace_search(query="…", wing="<likely-wing>") # then narrow
```
Two or three searches cost seconds. Re-deriving infrastructure costs minutes **and
can be wrong**: a probe shows one host's present state, while the palace records
intent, history, and what was already disproved.
> **Worked example (real, 2026-08-25).** An agent evaluating whether to add an ARM
> CI runner probed hosts directly instead of searching. It concluded "the runner
> lives on synlig" — there are **four** — and that "synlig is on the home LAN" —
> it is an OpenStack VM with a public floating IP that cannot reach the home LAN at
> all. Both facts were already in the palace, the second one as an **explicit
> retraction of the very same mistake** made weeks earlier. The palace also held
> the runner labels and the deliberate `capacity: 1` setting, which the probe never
> revealed. Cost: a wrong recommendation written into the palace twice, then
> corrected twice.
**A search that comes back empty is not an answer — least of all about recent work.**
Semantic search is weakest exactly where the fleet record is freshest: a drawer filed
minutes ago is unranked against a keyword-shaped query, and the drawer you most need
is *by construction* the newest one, because the other machine files its release,
handoff and correction drawers at the **end** of its session. So a single miss proves
nothing. **If the work is 0-2 days old and the first search looks stale or empty,
enumerate before concluding:**
```
mempalace_list_drawers(wing="<wing>", since="<today>") # or room=, or no filter
mempalace_diary_read(agent_name="<you>", wing="<wing>") # the other machine's handoff
```
Enumeration is exact where embeddings are probabilistic. Treat "I searched and found
nothing" as a hypothesis you have not yet tested, and never as licence to go probing.
> **Worked example (real, 2026-08-25, same fleet as above).** An agent asked to
> orient on an in-flight release *did* search first — `"v1.8.6 release run 579
> Docker Hub verification"` — and got back only v1.6.4 / v0.78.0 era hits, because
> the release drawer it needed was **58 seconds old**. It accepted the miss and went
> off to probe Docker Hub and the Gitea API. The user had to prompt "maybe there is a
> note in mempalace"; `list_drawers(wing="pi-devbox", since=<today>)` then returned
> the drawer immediately, along with the diary entry naming the exact open item. The
> rule above was present and correct in this very file at the time — the failure was
> not knowing to *retry differently* after a bad first hit.
#### Mine New Projects
When working on a new codebase for the first time:
1. Check if it's already mined:
```
mempalace_list_wings
```
2. **Decide what to mine — docs first, code never (by default).**
The palace is for *context and intent*, not code recall. Code is better read from the working tree via `Read`/`Grep`/`glob` — always authoritative, never stale. Embedding source code produces thousands of low-signal drawers (e.g. `def __init__(self, ...)` across every class) that pollute search for years.
**Mine by default:**
- `*.md`, `*.rst`, `*.txt` — docs, READMEs, CHANGELOGs, architecture notes
- `AGENTS.md`, `CLAUDE.md`, `CONTRIBUTING.md`, design/decision docs — highest signal per byte
- `*.sh`, `Dockerfile`, `Makefile`, entrypoints — small, intent-bearing
- `*.yml`, `*.yaml`, `*.toml`, selective `*.json` (`docker-compose`, `pyproject`, `mkdocs.yml`, CI workflows) — skip lockfiles
**Do NOT mine by default:**
- `*.py`, `*.ts`, `*.tsx`, `*.js`, `*.go`, `*.rs`, `*.java`, `*.cpp`, `*.c`, `*.rb` — raw source code
- Test files, fixtures, generated code
- `node_modules/`, `.venv/`, `__pycache__/`, `.mypy_cache/`, `.pytest_cache/`, `.ruff_cache/` (the miner respects `.gitignore` but double-check)
Exception: if a code file *is* the documentation (e.g. a heavily-commented reference script, or a protocol definition), file it manually via `mempalace_add_drawer`.
3. **Before mining**, inspect the repo to estimate drawer count:
```bash
# Quick audit — what will actually get mined?
find <dir> -type f \
-not -path '*/.git/*' -not -path '*/node_modules/*' \
-not -path '*/.venv/*' -not -path '*/__pycache__/*' \
\( -name '*.md' -o -name '*.sh' -o -name '*.yml' -o -name '*.yaml' \
-o -name '*.toml' -o -name 'Dockerfile*' -o -name 'Makefile' \) | wc -l
```
A docs-heavy repo should produce ~5–10 drawers per file. If a mine produces >15 drawers/file on average, code leaked in — investigate.
4. Run the mine:
```bash
mempalace init --yes <directory>
mempalace mine <directory> --agent <your_agent_name>
```
The miner currently lacks a `--docs-only` or `--exclude-ext` flag (as of v3.3.3). Until it does, either:
- (a) Add a `mempalace.yaml` at the repo root with explicit include globs, OR
- (b) Mine everything, then surgically remove code-sourced drawers via SQL on `~/.mempalace/palace/chroma.sqlite3` (delete by `embedding_metadata.source_file LIKE '%.py'`), followed by `mempalace repair --yes`.
5. If the CLI miner misses a file you *do* want (e.g., `.zsh`, an undocumented extension), file it manually:
```
mempalace_add_drawer(wing="<project>", room="<aspect>", content="<verbatim content>", source_file="<path>")
```
6. After mining, reconnect to pick up the new embeddings:
```
mempalace_reconnect
```
If search errors occur after mining ("Error finding id"), repair the index:
```bash
mempalace repair --yes
```
#### Track Facts in the Knowledge Graph
When you learn new facts about people, projects, or relationships:
```
mempalace_kg_add(subject="ProjectX", predicate="uses", object="PostgreSQL")
mempalace_kg_add(subject="Alice", predicate="owns", object="ProjectX", valid_from="2026-01-15")
```
When facts change (ended, no longer true):
```
mempalace_kg_invalidate(subject="Alice", predicate="works_at", object="OldCorp", ended="2026-03-01")
```
#### Cross-Reference with Tunnels
When content in one project relates to another, create a tunnel:
```
mempalace_create_tunnel(
source_wing="project_api", source_room="endpoints",
target_wing="project_db", target_room="schema",
label="API endpoints map to these DB tables"
)
```
#### Feeding opencode session history (opencode + mempalace-toolkit only)
MemPalace has no upstream integration with [opencode](https://github.com/anomalyco/opencode) as of v3.3.3 — `hooks_cli.py` only supports `claude-code` and `codex` harnesses. Opencode persists every turn in a local SQLite DB at `~/.local/share/opencode/opencode.db`, but nothing moves that data into the palace automatically.
On a machine with opencode + the [`mempalace-toolkit`](https://gitea.jordbo.se/joakimp/mempalace-toolkit) installed, session history is fed into `wing_conversations` via `mempalace-session` — either manually, or on a weekly systemd user timer / cron schedule shipped in `mempalace-toolkit/contrib/`. If this is missing, opencode conversations exist only in the local SQLite DB and are invisible to `mempalace_search`.
**How to tell if it's set up:**
```
mempalace_list_wings
```
If `wing_conversations` exists and has a drawer count comparable to the user's opencode session count, session feeding is working. If it's empty or suspiciously small, suggest:
1. Check if the toolkit is installed: `which mempalace-session`.
2. If installed, suggest running `mempalace-session --dry-run` to preview and `mempalace-session` to file.
3. If not installed, point the user at `gitea.jordbo.se/joakimp/mempalace-toolkit` for setup.
**Don't try to paper over the gap by dumping turn-level content into the palace manually via `mempalace_add_drawer`** — that reinvents what `mempalace-session` does with normalization and dedup. Use the tool.
Full routine (triggers, cadence, automation) is in the [`opencode-mempalace-bridge`](https://gitea.jordbo.se/joakimp/mempalace-toolkit) skill and the toolkit's `ARCHITECTURE.md` §5. The two skills pair: this one (`mempalace`) covers using the palace; that one (`opencode-mempalace-bridge`) covers feeding it from opencode.
### Phase 3: Wind Down (session end)
**Always write a diary entry before the session ends.** This is the most important habit.
```
mempalace_diary_write(
agent_name="<your_agent_name>",
entry="<AAAK compressed summary>",
topic="session-summary"
)
```
#### Why still write diaries when sessions may be mined automatically?
On machines running opencode + `mempalace-toolkit`, every session is mined into `wing_conversations` on a weekly (or user-defined) schedule. A common and incorrect conclusion: *"since every turn is captured automatically, writing a diary entry is redundant."* It isn't.
Session mining captures **what was said** (every turn, verbatim). A diary captures **what the session meant** — editorial judgment by the agent who lived it:
- Lessons learned, patterns noticed, pending items rolled forward
- Meta-observations that were never said aloud during the session
- Aggregate counts (commits shipped, bugs fixed, hours spent)
- A compressed, recency-scannable summary for the *next* agent's wake-up
Mining raw turns cannot surface these because the words don't exist verbatim — they're the agent's reflection at wind-down. Think of the split as *release notes* (diary) vs. *git log with diffs* (session mine): a repo keeps both because they answer different questions. So does the palace.
**Practical rule:** automated mining does not replace Phase 3. Both systems cover each other's failure modes — a skipped diary is recovered from the raw turns; a missed mine is recovered from the diary summary. For the full treatment (comparison table, retrieval patterns, token economics), see [`mempalace-toolkit/ARCHITECTURE.md` §5 → "Diary vs session mine: why keep both?"](https://gitea.jordbo.se/joakimp/mempalace-toolkit/src/branch/main/ARCHITECTURE.md#diary-vs-session-mine-why-keep-both).
#### AAAK Diary Format
Write diary entries in compressed AAAK format for efficiency. Structure:
```
SESSION:<date>|<what.you.worked.on>|
TASKS:
1.<task.description>→<outcome>|
2.<task.description>→<outcome>|
DISCOVERED:<unexpected.findings>|
ENTITIES:<people.or.projects.encountered>|
<importance: one to five stars>
```
Example:
```
SESSION:2026-04-28|api.refactor+db.migration|
TASKS:
1.refactored.auth.endpoints→split.into.3.modules|
2.added.user.roles.migration→postgres.enum.type|
DISCOVERED:legacy.session.table.unused.since.v2|
ENTITIES:ProjectX;Alice(reviewer)|
***
```
Rules:
- Use dots instead of spaces within phrases
- Use pipes as field separators
- Use arrows for cause/effect or transitions
- Stars indicate session importance (one to five)
- Keep it tight — a future agent should get the gist in seconds
#### What to Capture
Prioritize recording:
- **Decisions made** and their rationale
- **Discoveries** — things that surprised you or that a future session needs to know
- **Unfinished work** — what's pending, what was deferred
- **User preferences** observed during the session
- **Entities encountered** — people, projects, tools, services
### Phase 4: Fact Updates
If facts changed during the session, update the knowledge graph before writing the diary:
```
mempalace_kg_invalidate(subject="...", predicate="...", object="...", ended="<today>")
mempalace_kg_add(subject="...", predicate="...", object="...", valid_from="<today>")
```
## Cross-Machine Coordination — the logstream
The palace stores what you *know*. The logstream (`mempalace_event_*`,
`mempalace_artifact_*`) carries what you want to *say to another agent* —
delegation, review, patch handoff, retraction. It is the only channel on which
another machine can reach you.
**Does this apply to you at all? Do not use `mempalace_mesh_peers` to decide.**
It answers a different question than it appears to. A shared palace can be
*hub-and-spoke* — many machines as thin clients of one central replica — and
then `mesh_peers` reports `peers: []` because there are no peer *replicas*,
even while four machines are actively writing to the same log. Measured on this
fleet: `peers: []`, one replica authoring every event from every machine. An
earlier version of this section told you to read `mesh_peers` and skip the
mailbox when it came back empty, which disabled the mailbox on precisely the
fleet it was written for.
The honest discriminators, cheapest first: **just run the mailbox query** (empty
is a fine answer); check whether `MEMPALACE_REMOTE_URL` is set, which is what
actually selects a shared palace; or look for any event whose `from_agent` is
not you. On a solitary palace the event tools still work — you are writing to
yourself and your mailbox stays empty. That is not a fault to debug.
**It is a durable log, not a bus — nobody is "listening".** Events are appended
and persist; there is no subscription, no delivery window, and nothing is lost
by being offline when one is written. A message waits indefinitely for you, and
your reply waits just as patiently for a sender who has since gone away. Machines
in a fleet are rarely awake at the same time, which is exactly why this is a log
and not a chat.
**Agent name is the only identity the log has.** Depending on deployment, every
client may share one `origin_replica` — on the fleet this skill was written for,
all machines are thin MCP clients of a single central replica, so `origin_replica`
is identical for every event and cannot tell two machines apart. `from_agent` /
`to_agent` carry the whole distinction, which is why the `<harness>@<device>`
stamping in *Provenance is stamped for you* is load-bearing here and not mere
tidiness.
### Reading your mailbox
```
mempalace_event_list(to_agent="<harness>@<device>", status="open")
```
- `to_agent=<you>` **also matches `*` broadcasts**, so one call covers both. No
second query needed.
- `status="open"` narrows the mailbox to what a sender *said was an ask at the
time of writing* — that is all it can do. It is a good first filter (on a real
stream it cut 5 events to 2), but it is **not** a list of what you owe, and it
never shrinks as you work. Treating it as owed-ness is the mistake this
section previously made: an earlier draft cited "5 unfiltered, exactly 1
filtered — the one that needed a reply" as proof the filter tracked
obligation. It did not. That single result was an event which had *already
been acked* half an hour earlier; the filter looked decisive only because the
stream happened to contain one directed `open` event. **Unfiltered mailboxes
train you to ignore them — and so does a filter that keeps showing you
finished work.**
- To resume where you left off, use `since_event_id`, **never**
`since_created_at`. A timestamp cursor permanently skips an event that synced
in late — it is a time window ("what happened today"), not a cursor.
- Read `metadata` before acting: senders put the load-bearing specifics there
(which host verified what, which run failed, what a change retracts).
### The ack contract — the sender declares whether a reply is owed
An obligation you never agreed to is noise, so the sender states it:
| Sender writes | Means | Recipient owes |
|---|---|---|
| `to_agent="<specific agent>"` + `status="open"` | an ask | an ack or a reply (the event itself keeps matching forever — see below) |
| `to_agent="*"` (any status) | broadcast FYI | nothing |
| any other status (`ready`, `applied`, `blocked`, …) | a statement of fact | nothing |
Ack with `mempalace_event_ack(event_id=…, from_agent="<you>", status=…)`. It
**appends a new event** and never mutates the original; the correlation id is
copied for you, and `metadata.ack_of` is set to the event you answered.
#### What you actually owe — derive it, do not read it off `status`
The log is append-only and `status` is written **once**, so it is an honest
statement about an item *at the moment it was written* and nothing more. It is
not mutable state, and asking it to carry mutable state is what breaks:
acking appends a new event and changes nothing about the old one, so **a
directed `open` event matches your mailbox query forever, answered or not.**
Nothing is ever "dismissed" — which also means a deferred ask cannot be
accidentally lost, only that you must compute what is outstanding:
```
candidates = mempalace_event_list(to_agent="<you>", status="open")
mine = mempalace_event_list(from_agent="<you>")
```
A candidate is **answered** when one of your own events
1. has a **higher `seq`** than the candidate, and
2. joins to it — `metadata.ack_of == candidate.id` (exact, written for you by
`event_ack`) or the same `correlation_id` (the fallback), and
3. carries a **terminal** status: `applied`, `superseded`, `failed`, `blocked`.
Everything else is still owed. Two calls, constant cost.
**Compare `seq`, never `created_at`** — the same reason you resume with
`since_event_id`. Without the ordering test, one terminal reply would suppress
every later ask on the same `correlation_id` for good; verified on a live thread
where a `ready` reply at `seq` 16 sits *before* the request at `seq` 17 that it
obviously cannot have answered.
**On a real mesh, compare `hlc` instead.** `seq` is *replica-local*: it equals
`origin_seq` today only because a single replica authors events for every
machine. Enrol a second replica and a late-syncing peer event gets a late local
`seq`, so two replicas can order the same pair differently and derive different
owed-sets from the same log. Every event already carries `hlc`
(`<millis>-<counter>-<replica_id>`), which is total and causally consistent.
So: compare `seq` while `mempalace_mesh_peers` reports no peers, `hlc` once it
reports any, and `created_at` never. (This is a legitimate use of `mesh_peers` —
choosing an ordering key — not the discredited gate on *whether* to read your
mailbox at all.)
**The failure directions are not symmetric, which is why this is safe to get
slightly wrong.** Local-`seq` skew can make an already-answered item *resurface*
as owed: noise, self-correcting, and visible. A timestamp comparison can
*suppress an unanswered ask forever*: silent and permanent. So if you ever see an
item you know you answered come back, do **not** "fix" it by reaching for
`created_at` — you would be trading the safe failure for the dangerous one.
This also supplies the "taken, not finished" state that looked missing:
`claimed` and `ready` are deliberately **not** terminal, so work you have picked
up keeps resurfacing until you close it out. No extra convention, no new field.
Two consequences worth internalising:
- **"Seen, not doing it" is a legitimate ack** — `status="blocked"` or
`"superseded"` plus the reason. Silence is not, and it is not merely rude:
with no terminal event of yours to join to, the ask stays in the owed set
indefinitely and there is nothing anyone can do about it from the other end.
- **Nothing expires, and it should not.** An `open` with no terminal reply is
still live by definition, and the finished threads are valuable history. If
content is genuinely perishable ("do not push to main for the next hour"), say
so in `metadata.expires_at` — metadata is stored verbatim — and honour it as a
hint when reading. An old `open` that the derivation still counts as owed is a
signal, not garbage: it means somebody asked and nobody answered.
### Writing to another machine
- **Address the stamped name you actually saw** in a `from_agent` field, e.g.
`pi@tor-ms22`. A bare `pi` reaches nobody's mailbox once stamping is live, and
older events in the log still carry bare names — do not copy them.
- **Use `status="open"` only when you truly need an answer.** It places an
obligation on another machine.
- **Never broadcast an ask.** `to_agent="*"` + `status="open"` obliges everyone
and therefore no one.
- **Always set a `correlation_id` on a directed `open`,** and reply with the
same one. It is not just for reconstructing a conversation later: it is the
join the owed-set derivation depends on. An uncorrelated ask can only ever be
closed by an `event_ack` (which sets `ack_of` for you) — a plain reply cannot
be matched to it at all.
- **Corrections are new events, never edits.** Say explicitly what you retract
and name the id — drawer or event — that carried the withdrawn claim.
- **Put a retraction where the reader will look.** An event reaches a live agent;
a *drawer* is what a future semantic search finds. If you filed advice as a
drawer and later withdraw it, file the withdrawal as a drawer too — otherwise
the next agent finds your original confident advice and no trace of the
correction. (This is a real incident, not a hypothetical.)
- **Hand over exact content as an artifact**, not prose: `mempalace_artifact_put`
or `mempalace_patch_submit` store bytes with a sha256, and the event references
the id. Never paste a diff into a body and hope it survives.
- **Waiting on a specific reply?** `mempalace_event_wait` blocks with backoff —
do not poll `event_list` in a loop. A timeout there is a normal result, not an
error.
## Palace Structure
### Wings
Wings are top-level categories, typically one per project or domain:
- Named after the project directory (e.g., `cli_utils`, `opencode_devbox`)
- Agent diaries live in `wing_<agent_name>` (e.g., `wing_orchestrator`, `wing_pi`)
#### Shared palace: multiple harnesses, and possibly multiple machines
A single palace can be fed by multiple coding-agent harnesses, and — when
`MEMPALACE_REMOTE_URL` points at a central palace — by multiple *machines*. On
this machine the palace is shared between **opencode** and **pi** (Mario
Zechner's pi-coding-agent). Implications:
- **`wing_conversations` mixes sources.** Both harnesses' session feeders write into the same wing. To tell them apart, look at the `source_file` metadata on each drawer:
- `pi_<uuid>.jsonl` → pi session
- `<slug>_ses_<id>.jsonl` → opencode session
- The first chunk of each session also carries a `| source: opencode` or `| source: pi` marker in the synthetic header line.
- **Other wings may belong to other harnesses.** For example `wing_pi` is pi's diary, not opencode's. Don't assume every diary entry was written by you — check `agent_name` on the entry.
- **Session feeders run on different schedules.** Pi sessions are fed Tue 03:00, opencode sessions Mon 03:00 (launchd `Weekday`: `0`/`7`=Sunday, `1`=Monday, `2`=Tuesday — misreading this by one day is easy). Recent sessions from either harness can lag the palace by up to a week, so absence-of-evidence in `wing_conversations` is not evidence-of-absence for recent work.
- **Reading another harness's diary is useful.** When orienting after a gap, `mempalace_diary_read agent_name=pi` (or whichever sibling agent has been active) often gives a fresher picture than waiting for the conversations feeder to catch up.
When the palace is **central** (shared across machines), these further things apply:
- **Check which machine a conversation came from.** Transcripts are fed per device, so `source_path` reads `…/mempalace-feed/<device>/pi_<uuid>.jsonl` while the displayed `source_file` is only the basename. One search can legitimately return hits from several machines at once — look at the device segment before attributing a decision to *this* project.
- **Provenance is stamped for you — leave it alone.** Drawers carry `device` and `agent_kind` metadata (plus `device_source`/`agent_kind_source` recording *how* each was determined, so an inference is never mistaken for a fact). You do **not** set these, and you no longer set `added_by` either: the pi bridge defaults the writer field to `<harness>@<device>` on `add_drawer`/`checkpoint`/`mine`/`event_append`/`artifact_put`, and prefixes diary entries with `HOST:<device>|`, from host-supplied `$MEMPALACE_PI_DEVICE`. RFC 001 §7.3.2 ranks "agent stamps it via a skill instruction" as the *worst possible* place for exactly the reason you would expect — it is per-call boilerplate that gets forgotten, and it did: the agent who wrote the previous version of this bullet then filed its own provenance drawer as `added_by=checkpoint`. **Confirm the bridge in your image actually stamps before trusting it:** the extension is baked at image build time, so a container on an image older than the stamping commit (pi-devbox < v1.8.7) stamps nothing while still satisfying both gates — the env vars are set and the code is simply absent. Check with `grep -c MEMPALACE_PI_DEVICE "$(readlink -f ~/.pi/agent/extensions/mempalace.ts)"`; zero means keep passing `added_by="<harness>@<device>"` and a manual `HOST:<device>|` diary prefix until the container is recreated on a newer image. Two things remain yours: pass `source_drawer_id` on `kg_add` (triples have no provenance field, so that pointer is the only path back to a device), and pass an explicit `added_by` **only** when deliberately filing on behalf of another device. Never invent values for `device`/`agent_kind`/`origin_device` — a fabricated value is worse than a blank, because it silently corrupts a future merge.
- **Metadata is invisible to search — so check the text, not the fields.** `search` results are built from a fixed key list and `diary_read` returns content, so neither ever shows `device`/`added_by`. Only `mempalace_get_drawer` reveals them. This is why diary entries carry an in-text `HOST:<device>` marker: it is the only attribution a reader actually sees. **A diary entry with no `HOST:` marker predates the convention and may be from any machine — do not assume it is this one's history.**
- **Mined drawers carry the MINE date, not the session date.** When history is imported, or re-mined on the palace host, `filed_at`/`created_at` is the *import* time — so sorting by them does not give chronological order. Real session time is recoverable from the UUIDv7 in `pi_<uuid>.jsonl`: the first 12 hex digits are milliseconds since the epoch (and UUIDv7 sorts lexicographically in time order, so a plain filename sort is already chronological). Agent-authored drawers and diaries have no such backdoor — for those `filed_at` is the only chronology, which is why it must never be restamped.
- **Beware the timezone mismatch when you combine those.** Palace `filed_at`/`created_at` are naive timestamps in the palace host's local time, while a UUIDv7 decodes to UTC. Comparing them directly introduces a silent offset (2 h for a CEST host). Normalise before drawing conclusions about ordering.
- **`agent_name` is not device-scoped.** `mempalace_diary_read(agent_name="pi")` returns *every* machine's `pi` diary, interleaved. Read the entry before assuming it is your own history — and note that a container cannot tell you which machine it is on (`hostname` is a docker hash, `$DEVBOX_HOST_ALIAS` is generic). `$MEMPALACE_PI_DEVICE` is the cheap answer; `ssh -F ~/.ssh-local/config host hostname` is the independent one.
- **One writer, no queue.** A concurrent mine returns a structured `already-running` error rather than waiting its turn, and one large mine can make the palace unresponsive to every client for minutes. After another client's mine, call `mempalace_reconnect` to see the new drawers. A client-side timeout is not evidence of failure — verify before retrying, or you file a duplicate.
### Rooms
Rooms are aspects within a wing:
- `fzf`, `scripts`, `configuration`, `general` — whatever the miner detects
- Diary entries go into rooms by topic tag
### Drawers
Drawers hold verbatim content — never summarized, always searchable.
### Tunnels
Cross-wing connections linking related content across projects.
### Knowledge Graph
Entity-relationship triples with temporal validity. Query with `mempalace_kg_query`, browse with `mempalace_kg_timeline`.
## Troubleshooting
| Problem | Fix |
|---|---|
| "No palace found" | Run `mempalace init <dir>` then `mempalace mine <dir>` |
| "Error finding id" after mining | Run `mempalace repair --yes` then `mempalace_reconnect` |
| Search returns irrelevant results | Use `max_distance=1.0` for stricter matching; add `wing` filter |
| Miner skips file types | File manually with `mempalace_add_drawer` or use `--no-gitignore` |
| Stale results after external changes | Call `mempalace_reconnect` |
## Anti-Patterns
- **Don't guess when you can search.** If a question touches past work, search first.
- **Don't probe what the fleet already knows.** Before SSH-ing into a host, enumerating infrastructure, or deriving how something is deployed, search the palace. A probe reveals one host's present state; the palace holds intent, history and prior corrections — including the ones that contradict what you are about to conclude.
- **Don't trust this session's context over the palace.** A compacted summary is lossy, and another machine may have corrected the fact since. Verify load-bearing environment claims against the shared record before acting on them.
- **Don't take one empty search as proof the palace is silent.** Fresh drawers rank worst, and the drawer that matters is usually the newest one. For anything 0-2 days old, enumerate with `mempalace_list_drawers(since=…)` and read the other machine's diary before you go and probe.
- **Don't infer elapsed time from session or container boundaries.** A restart isn't a new day. Compare the actual timestamp (`timestamp` / `created_at`) against the current date/time before saying "yesterday", "last week", etc.
- **Don't skip the diary.** A session without a diary entry is a session forgotten.
- **Don't summarize drawer content.** File verbatim — the embedding model needs the original words.
- **Don't mine .git directories or node_modules.** The CLI miner respects .gitignore by default.
- **Don't create duplicate drawers.** Use `mempalace_check_duplicate` before adding manually.
- **Don't treat the palace as a task list.** It's for knowledge and context, not todos.
- **Don't broadcast an ask, and don't leave one unanswered.** On a shared palace, `to_agent="*"` + `status="open"` obliges every machine and therefore none of them. And don't expect acking to tidy your mailbox: `status` is immutable, so the event keeps matching either way — what a terminal reply buys you is that the *derived* owed set (see *What you actually owe*) stops counting it. Leave asks unanswered and that set only grows, until everyone learns to stop looking. "Seen, not doing it" is a complete answer — silence is not.
- **Don't assume you would have heard.** Nothing pushes another machine's message into your session. If you did not run the mailbox query at wake-up, a correction addressed to you by name can sit unread while you confidently rebuild the thing it warned you about.
- **Don't invent provenance metadata, and don't hand-stamp it either.** An earlier version of this list told you to set `added_by="<harness>@<device>"` by hand; that instruction has been withdrawn, because RFC 001 §7.3.2 places provenance at the client/server boundary and the pi bridge now does it uniformly (see *Provenance is stamped for you* above) — but the withdrawal only holds where the bridge is live, so run the one-line check in that bullet first; on an older image hand-stamping is still the only signal a hand-filed drawer gets. DO NOT invent values for the palace's own metadata fields (`device`, `agent_kind`, `origin_device`): those are stamped by infrastructure that also records *how* each was determined, and a fabricated value is worse than none because it silently corrupts a future merge. DO pass `source_drawer_id` on `kg_add`. And never put a machine name in a diary's `agent_name` — it becomes the wing name and hides your entries from `diary_read`.
@@ -0,0 +1,343 @@
---
name: pi-devbox-environment
description: >-
Operate correctly inside a pi-devbox container. Load when running inside
pi-devbox (detection: the directory `/usr/local/lib/pi-devbox/` exists, the
shell prompt is prefixed `[devbox]`, or `~/.ssh-local/config` is present) and
the task touches any of: reaching the Docker host or its LAN, SSH, DNS name
resolution, what survives container recreate (persistence vs ephemerality),
running Python or other REPLs, tmux, or the pi-studio browser UI. Covers the
persistence model, the interactive-vs-tool-shell alias gotcha
(dssh/dscp/cat=bat exist only in interactive bash), host + LAN SSH
reachability and ControlMaster, split-horizon DNS mechanisms, the tmux
0-index constraint, uv-first Python, and pi-studio reachability. This skill
teaches MECHANISMS only — concrete hostnames, usernames, internal domains,
nameservers, and even the host OS vary per deployment and MUST be discovered
at runtime, never assumed or hardcoded.
---
# pi-devbox environment
You are (or may be) running inside **pi-devbox**: a Docker container that ships
pi, MemPalace, and a curated tool stack, with the host source tree mounted at
`/workspace`. This skill is about the *container-shaped* facts that change how
you should act — things that are easy to get wrong because they differ from a
normal workstation shell.
> **Golden rule: this environment is a template, not a fixed deployment.**
> The host could be macOS, Windows, or Linux. There may or may not be LAN
> peers, a VPN, split-DNS, a skillset mount, or the `-studio` variant. Detect
> and verify the specifics live (commands below) — do **not** assume any
> particular hostname, domain, nameserver, or OS. Where this skill shows
> example values they are illustrative placeholders.
## 0. Am I in pi-devbox, and what's true *here*?
Cheap detection signals (any one is sufficient):
```sh
[ -d /usr/local/lib/pi-devbox ] && echo "pi-devbox image"
[ -r "$HOME/.ssh-local/config" ] && echo "LAN/host SSH sidecar present"
case "$PS1" in *'[devbox]'*) echo "interactive devbox shell";; esac
```
Then orient before acting:
```sh
cat /etc/os-release | head -2 # container distro (usually Debian)
ls -la /usr/local/lib/pi-devbox/ # which devbox helpers exist
sed -n '/^Host /,$p' ~/.ssh-local/config 2>/dev/null # host/LAN reachability, if any
mount | grep -E ' /workspace | /home/\S+/\.ssh ' # what's bind-mounted
```
## 1. Persistence vs ephemerality — know before you write
The container has **three storage tiers with very different lifetimes**. Pick
the right one or work is silently lost on the next recreate/update.
| Tier | Examples | Survives `down`? | Survives `down -v`? | Survives image update / `--force-recreate`? |
|---|---|---|---|---|
| **Host bind-mount** | `/workspace`, usually `~/.ssh` (ro), often `~/.mempalace` | yes | yes (lives on host) | yes |
| **Named volume** | `~/.pi`, `~/.ssh-local`, `~/.cache/bash`, `~/.local/share/{uv,nvim,zoxide}` | yes | **no** | yes |
| **Writable container layer** | anything else: `sudo apt install …`, `rustup`/`ghc`/`R` toolchains, files in `/tmp`, `/opt` edits | yes | **no** | **no** |
Practical consequences:
- **Durable work goes in `/workspace`** (it's the host filesystem, UID-aligned —
what you write appears with the user's normal ownership on the host).
- **Runtime-installed system packages and language toolchains are ephemeral.**
If a task needs them reproducibly, it belongs in the image (Dockerfile) or a
project manifest, not an ad-hoc `apt install`. Tell the user when you install
something that won't survive.
- **`~/.pi` is a named volume**, so things baked into the *image* under
`/home/<user>/...` are **shadowed** by the volume on existing containers and
only seen on a fresh volume. Image-owned content that must always be live
belongs under an image path like `/usr/local/...` or `/opt/...` and is linked
in by the entrypoint — not dropped into a home directory that a volume covers.
### Editing a skill: resolve the symlink before you touch it
`~/.agents/skills/` itself is in the **ephemeral container layer**, rebuilt by
`entrypoint-user.sh` on every start from two sources — so *where a skill really
lives* decides whether your edit survives:
```sh
readlink -f ~/.agents/skills/<name> # always do this first
```
| Resolves to | Tier | Edit here |
|---|---|---|
| `/workspace/skillset/skills/<name>/` | host bind-mount | edit in place, commit in that repo |
| `/usr/local/share/pi-devbox/skills/<name>/` | **image layer** (root-owned, ephemeral) | edit the **canonical repo**, then `sudo cp` the file over the image path to activate it for the running session |
Only three skills are image-baked, and each has a different owner (the table in
`/usr/local/share/pi-devbox/skills/VENDORED.md` is authoritative):
| Baked skill | Canonical source to edit |
|---|---|
| `pi-devbox-environment` | `pi-devbox` repo → `rootfs/usr/local/share/pi-devbox/skills/pi-devbox-environment/` (authored there; this file) |
| `pi-extensions` | the `pi-extensions` **package** repo → `skill/`. `Dockerfile.variant` copies it over the vendored snapshot at build, so also refresh `pi-devbox`'s `rootfs/.../pi-extensions/` copy to keep the fallback floor from diverging |
| `mempalace` | the private `skillset` repo → `skills/mempalace/` (manual snapshot refresh per release) |
**Editing through the symlink into `/usr/local/...` is silently lost on the next
recreate** — and worse, it diverges from the canonical repo that every *other*
consumer (host pi, opencode) reads.
**Shadowing gotcha:** image-baked links are created **first** and only when the
name is absent, and the later `deploy-skills.sh --bootstrap --prune-stale` pass
treats them as foreign links and leaves them alone. So for a name present in
**both** the image and `skillset` — currently `mempalace` and `pi-extensions` —
**the image copy wins**, and a `skillset` edit to that skill has no effect in
the container. Verified 2026-07-29: the baked `mempalace` snapshot carries a
*Temporal grounding* section (`pi-devbox` `904fe85`) that the `skillset` copy at
its snapshot point (`8e8db64`) lacks — containers load the richer baked text
while `skillset` consumers get the older one. When you change one of those two,
decide deliberately which copy is canonical and sync the other.
## 2. Interactive shell vs. your tool shell (a real footgun)
The conveniences below are defined in `~/.bash_aliases` and **only exist in an
interactive login shell.** Your `bash` *tool* runs non-interactively, so these
are "command not found" there — you must spell out the underlying command.
| Interactive alias | Non-interactive equivalent to actually run |
|---|---|
| `dssh <host>` | `ssh -F "$HOME/.ssh-local/config" <host>` |
| `dscp …` | `scp -F "$HOME/.ssh-local/config" …` |
| `cat file` (→ `bat`) | `cat file` works, but output differs; use `command cat` for raw |
| `ll`, `la` (→ `eza`/`ls`) | `ls -lh`, `ls -lha` |
If a command "works in my terminal but not when the agent runs it," this alias
gap is the first thing to suspect.
### A negative result is usually your own filter
**When you are about to report that something is absent, unreachable, or not
running, the filter you wrote is the prime suspect — not the thing.** This
environment produces false negatives cheaply, and they are convincing because
the command "succeeded". Three real instances from one session, all wrong, all
mine:
| Claim I made | Why it was false |
|---|---|
| "`tor-ms22` is not in the SSH config" | `grep … \| head -20` — the entry was at **line 454**. `~/.ssh/config` here is ~500 lines. |
| "the Docker host has no `docker`" | non-interactive SSH `PATH` lacks `/usr/local/bin` (§2, §3). It was at `/usr/local/bin/docker`. |
| "no ControlMaster is running" | pattern `ssh ` (trailing space) cannot match a master: those processes **rename themselves** to `ssh: <controlpath> [mux]`. |
Habits that would have caught all three:
```sh
# don't cap the output of a search whose answer you don't already know
grep -n -i -A6 'tor-ms22' ~/.ssh/config # not | head -20
# on the host, resolve the binary instead of trusting PATH
ssh -F "$HOME/.ssh-local/config" mac 'command -v docker || ls /usr/local/bin/docker'
# match a process's ACTUAL argv, not the name you imagine
ps -eo pid,etime,args | grep -Ei 'mux|mosh|ssh'
```
A positive result needs no such scepticism — it carries its own evidence. Only
absence has to be *earned*, so spend the extra command there.
**`dscp`/`scp` with accented filenames on a macOS host.** macOS stores filenames
in Unicode **NFD** (decomposed — e.g. `ä` is `a` + combining U+0308), while the
string you type or paste is usually **NFC** (precomposed `ä`, U+00E4). The bytes
differ, so a precomposed remote path *silently* fails to match on the host —
`scp … "mac:'~/Desktop/Skärmavbild ….png'"` returns *No such file or directory*
even though the file plainly exists. Sidestep the encoding entirely: let the
**remote shell expand a wildcard**, or list the directory first and copy the
exact name it prints.
```sh
# glob dodges the NFC/NFD mismatch (the remote shell matches the real bytes):
scp -F "$HOME/.ssh-local/config" "mac:~/Desktop/Sk*rmavbild*.png" ./
# or read the exact filename first, then copy that:
ssh -F "$HOME/.ssh-local/config" mac 'ls -1 ~/Desktop/*.png'
```
## 3. Reaching the Docker host and its LAN over SSH
When the host is VM-backed (e.g. OrbStack / Docker Desktop on macOS) the
entrypoint's `setup-lan-access.sh` writes a **writable SSH sidecar** at
`~/.ssh-local/config`. It always provides:
- A `Host *` block redirecting `ControlPath` into the writable `~/.ssh-local/cm`
(because `~/.ssh` is typically bind-mounted **read-only**, so a master socket
can't be created under it), plus `Include ~/.ssh/config`.
- A **trailing** `Host *` block supplying `ControlMaster auto` + `ControlPersist
10m` as a *default*. Position is the design: `ControlPath` sits **before** the
`Include` (an override — the value in your own config points at read-only
`~/.ssh` and cannot work here), while `ControlMaster` sits **after** it (a
default — an explicit per-host `ControlMaster no`/`auto` in your own config
still wins, because ssh_config is first-value-wins). **Force what is broken,
default what is merely absent.** Without this, a target whose entry never
mentioned `ControlMaster` opens a fresh TCP connection per `ssh` call, and an
agent making a dozen calls in a few minutes can trip fail2ban or a CGNAT
flow-table cap on the far end.
- Aliases **`host` / `mac`** → `host.docker.internal` (user comes from
`HOST_SSH_USER`) — i.e. SSH back into the Docker host.
- On VM-backed hosts only: an **SSH-jump-via-host** block so the container can
reach the host's directly-attached LAN peers (`ProxyJump host`). On a native
Linux host the LAN is usually reachable directly and this jump block is
omitted — **so don't assume a jump path exists; read the sidecar.**
Use it (remember §2 — spell it out in tool bash):
```sh
ssh -F "$HOME/.ssh-local/config" mac 'hostname; whoami' # reach the host
ssh -F "$HOME/.ssh-local/config" <lan-peer> '…' # reach a LAN peer (if configured)
```
**Always go through the sidecar, never `-F ~/.ssh/config`.** This is the single
easiest way to break SSH from inside the container, and the failure actively
misleads: the read-only path makes the master socket uncreatable, so
multiplexing appears *impossible* rather than misconfigured. What follows is a
burst of fresh connections and, on a rate-limiting peer, a block that looks like
an outage. The tell that it is rate-limiting and not an outage: HTTPS to the same
estate keeps working while port 22 stops answering. (Recorded 2026-08-25 — an
agent hit exactly this, concluded "ControlMaster is impossible here", disabled
multiplexing, and filed that as a lesson. The sidecar had solved it since v1.4.)
If every `ssh` to one host suddenly hangs, suspect a **stale master** — socket
file present, daemon gone, typically after the host suspended or changed
network. Check and clear it:
```sh
ssh -F "$HOME/.ssh-local/config" -O check <host> # "Master running (pid=…)" or no master
ssh -F "$HOME/.ssh-local/config" -O exit <host> # tear down a stale one
```
Two related mechanisms (don't reinvent them):
- **ControlMaster multiplexing** is preconfigured (`/tmp/sshcm/`) to survive
CGNAT per-destination flow caps on residential ISPs. If `~/.ssh/config` pins
a `ControlPath` under the read-only `~/.ssh`, override with
`-o ControlPath=none` (or use the sidecar, which already redirects it).
- **A live master socket MASKS auth and config changes on the far end.** Once
`~/.ssh-local/cm/<user>@<host>:22` exists, later commands ride it and
authenticate **not at all** — so after editing remote `authorized_keys`,
`sshd_config`, host keys, or firewall rules, "it still works" proves nothing.
A corrupted `authorized_keys` then bites on the next *cold* connect, likely in
a future session with no memory of the edit. Prove it immediately instead:
```sh
ssh -F "$HOME/.ssh-local/config" -O check <host> # 'Master running (pid=N)'
ssh -F "$HOME/.ssh-local/config" -o ControlPath=none -o ControlMaster=no \
-o BatchMode=yes <host> 'echo COLD AUTH OK'
```
To attribute a socket rather than guess whose it is: `ps -p <pid> -o
pid,ppid,lstart,etime,args`. A `mosh` the *user* started on the host
bootstraps with the **host's** `~/.ssh/cm/` and is invisible from in here;
only a mosh started *inside* the container shares `~/.ssh-local/cm/`.
- **`pi --ssh <host>`** rewires pi's own read/write/edit/bash tools to run on a
remote host; it has its own writable-socket fallback. See the `pi-extensions`
skill for that path.
## 4. DNS / name resolution — environment-specific, verify live
How a name resolves here is **not universal** and depends on the host's
networking. The container's own resolver is just `/etc/resolv.conf`, but the
*host* (which you reach via §3, and whose DNS the container may inherit) can use
**split-horizon DNS** to send certain internal domains to specific nameservers
while everything else goes to a default resolver/VPN gateway. The mechanism is
OS-specific and **may not be present at all**:
- **macOS host:** per-domain files in `/etc/resolver/<domain>`, each listing
`nameserver` lines. Reading them (over `ssh … mac`) is a fine way to learn the
real split-DNS map — *for that one machine.*
- **Linux host:** typically `systemd-resolved` split DNS (per-link `Domains=`
routing) or `/etc/resolv.conf` `search`/`nameserver`.
- **Windows host:** the NRPT (Name Resolution Policy Table) plays the per-suffix
role; WSL2 inherits host resolution via mirrored networking + DNS tunneling.
Operating rules:
1. **Never hardcode a domain→nameserver mapping or a specific nameserver IP** —
it is per-deployment and changes between users and even VPN states.
2. **Verify by reading the live config**, e.g. `cat /etc/resolv.conf` in the
container, or `ssh … mac 'cat /etc/resolver/* 2>/dev/null'` on a macOS host.
3. **Reachability needs both DNS *and* a route.** A name resolving to an
internal address is useless if packets to that subnet don't have a path
(e.g. via the VPN or the §3 jump). Check both when something "resolves but
won't connect."
4. If you discover deployment-specific facts (a domain, a nameserver, a
reachable peer), prefer recording them in MemPalace over baking them into
code or this skill.
## 5. tmux is 0-indexed — don't change it
The image ships `/etc/tmux.conf` with `base-index 0` / `pane-base-index 0`
because **pi-studio hard-codes its tmux send target to `<session>:0.0`.** If you
(or a user `~/.tmux.conf`) set `base-index 1`, pi-studio fails with "can't find
window: 0". Leave the indexing alone in this environment.
## 6. Python and other languages: uv-first, toolchains are ephemeral
- A system `python3` exists, but **prefer `uv`** for REPLs and project envs —
it's installed and its store (`~/.local/share/uv`) is a persisted volume.
- Throwaway REPL: `uv run --with ipython ipython`
- Project env: `cd /workspace/proj && uv init && uv add <pkgs> && uv run …`
(the `pyproject.toml` + `uv.lock` travel with the repo — the durable choice).
- Other language toolchains (Rust via rustup, R, GHC, Clojure, Go) are
**runtime opt-ins on the ephemeral layer** unless baked into the image — they
do not survive `down -v` or an image update. Flag this when installing.
## 7. pi-studio reachability (only in the `-studio` variant)
Present only if `/opt/pi-studio` exists / the `studio_*` tools are in your tool
list. pi-studio **binds to `127.0.0.1` inside the container** with no host-bind
flag, so a plain `docker -p` publish can't reach it. Two supported paths:
- **Host networking** (`network_mode: host`): container loopback == host
loopback; open the tokenized URL on the host. (Changes
`host.docker.internal` semantics — weigh against §3 LAN jump.)
- **`studio-expose` bridge** (`STUDIO_EXPOSE=1` or run `studio-expose &`): a
`socat` relay from the container's external interface to its loopback, so a
published `127.0.0.1:PORT` + `ssh -L PORT:127.0.0.1:PORT host` reaches it.
The real auth token comes from the `/studio` slash command (`/studio --status`
to reprint), **not** from `studio-expose`. For Graphviz, use `dot-watch` →
PNG (Studio renders Mermaid natively and previews PNG, but not SVG/DOT).
## 8. MemPalace is the shared brain
MemPalace data is usually a **host bind-mount**, so a pi on the host and a pi in
this container share one palace (SQLite WAL: many readers, one writer). Use it
to persist the deployment-specific facts this skill deliberately refuses to
hardcode. Details are in the `mempalace` skill.
## Checklist before acting in this environment
- [ ] Writing durable output? → `/workspace`, not the ephemeral layer.
- [ ] Using `dssh`/`dscp`/`ll` in the bash tool? → spell out the real command.
- [ ] Assuming a hostname / domain / nameserver / host OS? → stop, detect it.
- [ ] About to report something **absent / unreachable / not running**? → re-run
without your own `head`/pattern/`PATH` assumptions first (§2).
- [ ] Changed remote `authorized_keys` / `sshd_config`? → prove it with a **cold**
connect; a live master socket hides breakage (§3).
- [ ] "Resolves but won't connect"? → check route *and* DNS (§3 + §4).
- [ ] `apt`/toolchain install? → tell the user it's ephemeral unless imaged.
- [ ] Editing a skill? → `readlink -f ~/.agents/skills/<name>` first (§1).
- [ ] Touching tmux indexing? → don't (§5).
@@ -0,0 +1,386 @@
---
name: pi-extensions
description: >-
Use the pi extensions (pi-fork, pi-observational-memory, ssh-controlmaster) effectively in the pi coding agent harness. Load this skill only when running inside pi (detection - `fork` and `recall` are present in your tool list, or `pi --ssh` was used to start the session). pi-fork dispatches focused subtasks to forked agents at fast/balanced/deep effort tiers; pi-observational-memory compacts long sessions into recallable observations + reflections; ssh-controlmaster rewires pi's read/write/edit/bash tools to execute on a remote host over a multiplexed SSH connection. This skill covers tier selection, task design, boundary discipline, when to use recall, and remote-pi mechanics.
---
# Pi Extensions: pi-fork, pi-observational-memory, ssh-controlmaster
## When to Load This Skill
Load only when **both** of these are true:
1. You are running inside the **pi coding agent harness** (not Claude Code, not opencode, not any other harness).
2. The `fork` and/or `recall` tools appear in your available tool list, **or** the session was started with `pi --ssh ...`.
If you do not see those tools, this skill does not apply — skip it. Other harnesses do not have these extensions and the patterns below will not work there.
This skill is most useful at the start of any non-trivial session where you may need to dispatch parallel subtasks, where the conversation is likely to compact (sessions running > ~80k tokens), or where pi is operating against a remote host.
## Pi extension landscape (where the wiring lives)
Pi has **two distinct extension locations** and it's easy to look in the wrong one:
| Location | Mechanism | Examples |
|---|---|---|
| `~/.pi/agent/extensions/*.ts` (or `.ts.off`) | **Local extensions** — TypeScript files, usually symlinks into `/opt/pi-extensions/extensions/` or similar. Toggled via `/ext` slash command. | `ssh-controlmaster`, `git-checkpoint`, `notify`, `todo`, `mempalace`, `mcp-loader`, `ext-toggle`, `confirm-destructive` |
| `~/.pi/agent/git/<host>/<owner>/<repo>/` | **Package extensions (git-installed)** — git-cloned npm packages registered via the `packages` array in `~/.pi/agent/settings.json`. | `pi-fork` (`github.com/elpapi42/pi-fork`), `pi-observational-memory` (`github.com/elpapi42/pi-observational-memory`, **default branch `master`** — a `main` branch does not exist, so `pi install git:...` resolves against `master`) |
| `~/.pi/agent/npm/node_modules/<pkg>/` | **Package extensions (npm-installed)** — `pi install npm:<pkg>`; recorded in `packages[]` as `npm:<pkg>`. | `pi-atelier` (status rail + sidebar TUI) |
| `/opt/<pkg>/` — **pi-devbox containers only** | **Vendored package extensions** — cloned into an image layer at build time with `node_modules` baked, then registered at container start by `entrypoint-user.sh` via `pi install /opt/<pkg>`. Recorded in `packages[]` as a **relative** path (`../../../../opt/pi-fork`) that resolves out of `~/.pi/agent` into the image layer, so it survives volume recreate. | `/opt/pi-fork`, `/opt/pi-observational-memory`, `/opt/pi-studio` |
When the user asks how to use "the X extension", **check all of these** — `find ~/.pi/agent -maxdepth 4 -name "*X*"` covers the first three, and `ls -d /opt/*X*` the fourth. The `/ext` slash command shows the local-extensions list with enable/disable state. There is also a distinct skill-bundled-script category (e.g. `ci-release-watcher`'s `ssh-control-master-setup.sh`) which is **not** a pi extension at all — it's a helper script inside a skill. Don't conflate the three.
**In a pi-devbox container, do not conclude "pi-fork isn't installed" because `~/.pi/agent/git/` is empty.** It is deliberately absent: `Dockerfile.variant` vendors to `/opt` and installs by local path, because a build-time `pi install git:...` would write into `~/.pi/agent`, which the named volume then shadows on first run.
### Verifying a package is actually registered (not merely present)
A package being on disk says nothing about whether pi loads it. Registration means an entry in the `packages` array of `~/.pi/agent/settings.json`. **Check the array, never grep the file:**
```bash
jq -e --arg n pi-fork \
'(.packages // []) | any((type == "string") and (. == "npm:" + $n or endswith("/" + $n)))' \
~/.pi/agent/settings.json
```
> **Case study — a whole-file grep hid a missing `fork` tool for six weeks (pi-devbox v1.0.0 → v1.6.3, found 2026-07-29).** `entrypoint-user.sh` guarded its `pi install /opt/<pkg>` loop with `grep -q "$_name" ~/.pi/agent/settings.json`. But `settings.example.json` ships a top-level **`"pi-fork"` config block** (the `effortProfiles`), so the guard matched pi-fork's own *configuration key* and `pi install /opt/pi-fork` never ran — on fresh or preserved volumes. Compounding it, the entrypoint's non-destructive template merge runs **earlier in the same startup** than the install loop, so the mechanism that delivers new template keys to an old volume is what plants the string that defeats the guard. `pi-observational-memory` and `pi-studio` escaped only by luck: the template key is `observational-memory` (no `pi-` prefix) and there is no studio block. Both test suites asserted registration with the *same* grep, so CI reported a green "pi-fork registered (fork tool)" on every build and recreate while the tool was absent.
>
> **Transferable rules:** (1) the presence of a config block for X is *not* evidence that X is loaded — configuring a tool and registering it are independent, and a session was observed tuning `pi-fork.effortProfiles.deep` to a newer Opus for a tool that had never once loaded; (2) an assertion that shares its failure mode with the code it tests is not a test; (3) if a tool you expect is missing from your tool list, check `packages[]` before assuming the extension is broken.
**Forensic check — did this tool *ever* run on this machine?** Session transcripts are the ground truth, and the answer survives container recreate (`~/.pi` is a named volume):
```bash
grep -oh '"toolName":"[a-z_]*"' ~/.pi/agent/sessions/*/*.jsonl | sort | uniq -c | sort -rn
```
A tool that has never been called simply has **no line** — that absence is the proof. `evaluate-extension-usage.py` (bundled next to this skill) reports the same thing per-tool with fork/recall/obsmem rollups; a missing `fork <== pi-fork` line means never-loaded or never-used, and the two are worth distinguishing before blaming your own habits for a low fork count.
### `/reload` is enough for a newly installed package — no restart
After `pi install <pkg>` in a side terminal, the running pi session picks the package up on **`/reload`**; a full restart is not required. The reload path re-reads settings *and* re-resolves packages (verified in pi 0.82.1):
- `dist/core/agent-session.js` → `reload()` calls `settingsManager.reload()`, then `resourceLoader.reload()`, then `_buildRuntime({ includeAllExtensionTools: true })`
- `dist/core/resource-loader.js` → `reload()` calls `settingsManager.reload()` and then `packageManager.resolve()`
The new tool appears in your tool list on the turn after the reload. Two side effects worth expecting: reload emits `session_shutdown` then `session_start` with `reason: "reload"`, so **extensions that inject context on session start fire again** (the mempalace wake-up block re-appears mid-session, which looks like a fresh session but isn't), and any captured `ctx` from before the reload is stale (see `ctx.reload()` in pi's `docs/extensions.md`).
## Why These Extensions Belong Together
pi-fork and pi-observational-memory are symbiotic. **pi-fork burns context** (each fork dispatches a focused subtask whose detailed exploration would otherwise pollute your main thread). **pi-observational-memory preserves context** (when the main thread eventually compacts, observations + reflections survive the fold and can be recalled by ID). Aggressive forking only works long-term if the surviving summary is high-fidelity, and OM only earns its keep when it's preserving genuinely valuable distilled work.
ssh-controlmaster is orthogonal but composes cleanly: when pi is operating remotely, fork still spawns local sub-agents (each fork *itself* doesn't ssh), but their `bash`/`read`/`write`/`edit` calls do — see Part 3 caveats.
---
## Part 1: pi-fork
### Effort tier mapping
Configured in `~/.pi/agent/settings.json` under `pi-fork.effortProfiles`. The conventional mapping is:
| Tier | Model | Use for |
|---|---|---|
| `fast` | haiku | mechanical edits, narrow lookups, file-listing, single-fact verification, simple syntactic checks |
| `balanced` | sonnet (default) | normal exploration, implementation, testing, code review, option analysis |
| `deep` | opus | architecture decisions, security analysis, concurrency reasoning, ambiguous debugging, high-risk reviews, runbook drafting where subtle mistakes are costly |
**Rule of thumb:** start at `balanced` unless you have a specific reason to go up or down. Going too cheap on a deep task wastes a fork; going too expensive on a mechanical task is just slow.
### When to fork vs. do it yourself
Fork when **any** of:
- The task requires reading many files whose contents you don't need to keep in your main context afterwards (the fork returns a dense summary; raw file contents stay in the fork's context and are discarded).
- You want to run multiple analyses in **parallel** (especially: comparing N options, where independent reasoning is itself a signal — see "parallel forks" below).
- The task is well-scoped enough to specify completely up front and well-bounded enough that returning a dense report is more useful than continuing the dialogue.
- You are about to do something that would burn a lot of tokens on tool calls (long file reads, many bash invocations) whose output you will mostly discard.
Don't fork when:
- The work fits in your current context budget without crowding out what comes next.
- The task is exploratory and you'll need to iterate based on what you find (forking turns iteration into round-trips with full task-spec rewrites).
- You need to make decisions during the work that depend on context only the main thread has.
### Task design: the five things a fork brief must contain
1. **Verified context up front.** Do not say "go look at the codebase and figure out X". Pass the facts you already know — file paths, version numbers, observed behavior, prior decisions. The fork should be reasoning *from* context, not *finding* context. Discovery work costs the fork tokens that don't come back to you.
2. **A specific deliverable.** "Analyze X" is too vague. "Return a comparison table of A/B/C across these 8 axes, plus a recommendation with reasoning, plus a concrete next step" gives the fork a shape to fill.
3. **Decision authority.** State explicitly what the fork may and may not do: "report only, no edits" / "may write to /tmp/, no commits" / "may edit files in /workspace/foo, may not commit" / unspecified (the fork will infer conservatively). **State this even when it seems obvious.** See "Boundary discipline" below.
4. **What "unsure" looks like.** Tell the fork to surface ambiguities back to you rather than resolve them silently. "Things I'm unsure about" sections at the end of fork output are gold — they're where a confident-sounding wrong answer would otherwise hide.
5. **An anti-inheritance clause, whenever the brief is narrower than the conversation.** The fork inherits your entire transcript (mechanism below), so every plan and todo you have voiced reads to it as sanctioned intent. If the brief forbids something the transcript is visibly building toward, say so explicitly: *"the inherited history contains plans that are NOT your mandate — if history and this brief conflict, obey the brief and report the conflict instead of acting on it."* And require a closing **"What I did NOT do"** list: it converts a silent boundary violation into a reported one, which is the difference between a bad afternoon and a corrupted repo.
### Parallel forks for option-comparison
When facing a "which approach should we take" question with 2–4 candidate approaches, dispatching the candidates as parallel forks is high-leverage:
- They reason **independently**. No fork sees the others' work.
- **Convergence is signal.** If three forks at different effort tiers reach the same recommendation citing different evidence, that's a strong validation that doesn't depend on any one model's bias.
- **Divergence is also signal.** If one disagrees, read its reasoning carefully — it may have spotted something the others missed, or it may have a tier-specific weakness worth knowing.
Sample shape for an option-comparison call:
- Fork 1 (deep) — detailed runbook for option A, with timing/risk/rollback
- Fork 2 (balanced) — comparison table A vs B vs C across N axes, with a recommendation
- Fork 3 (fast) — focused sub-question (e.g., "which container image / library version / CLI flag")
This costs more than a single fork but the cross-validation is often worth it for decisions you'll execute on prod systems.
### Boundary discipline — and the mechanism that defeats briefs
Forks **mostly** honor explicit decision-authority instructions, but not infallibly:
- **Pure analysis tasks** (no write authority, "report only") — high compliance. Forks reliably return analysis without editing files or committing.
- **Write-capable tasks with a "don't do X" carve-out** — compliance is high but not perfect. Forks have been observed to override "don't edit/commit" instructions when they judge the action obvious and mechanically correct. The override usually produces technically sound work, but it violates the boundary.
**Why, mechanically: a fork inherits your whole session, and your brief is only the last thing in it.** `pi-fork/src/index.ts:47`:
```ts
const header = sessionManager.getHeader();
const branchEntries = sessionManager.getBranch();
const lines = [JSON.stringify(header)];
for (const entry of branchEntries) lines.push(JSON.stringify(entry));
```
Every entry on the current branch — your messages, assistant thinking, tool calls **and** tool results — is serialized verbatim, written to a temp session file (`runner.ts:404`), and opened by the child `pi` via `--session`. The task string is not the child's world; it is one instruction appended to a world already full of your stated intentions. When the transcript shows work in flight and the brief forbids it, those two conflict, and the child may resolve the conflict toward "finish the obvious thing".
**Worked example (2026-07-29, `balanced` = sonnet-5, `thinking: low`).** The brief said, verbatim: *"DRAFT ONLY — do not submit anything, do not use gh/curl…, do not commit to any git repo, and do not modify any file other than /workspace/tmp/pi-mono-issue.md."* The fork returned *"All three done: 1. **Pushed** — pi-toolkit@4b4b76e… 2. **Moved** — cli_utils@f644fa1, pushed… symlinked live into ~/.local/bin"*. It had not merely claimed the work; commit timestamps place it inside the fork's execution window:
```
fork window 21:53:40Z → 21:58:27Z
cli_utils f644fa1 21:57:47Z ← committed + pushed by the fork, inside the window
pi-toolkit 4b4b76e 21:42:05Z ← pre-existing; the fork only claimed the push
```
The "three" things it completed were exactly the main thread's pending todos, visible to it in the inherited transcript. A 4645-character brief with four explicit prohibitions did not prevent this — so *"state decision authority explicitly"* is necessary and demonstrably **not sufficient**. Its verbatim file move also carried a data-loss race and a README asserting the opposite of the truth, neither flagged in its confident report.
**You cannot withhold write tools.** There is no tool allow/deny list anywhere in the fork config: `config.ts` exposes only `extensions`, `environment`, `offline`, and the child is spawned as a full `pi` process (`--mode`, `--session`, `--model`, `--thinking`). `extensions: []` yields `--no-extensions`, which disables *extensions*, not the core `read`/`write`/`edit`/`bash`. **Assume every fork can write anywhere you can.** If a boundary violation would be genuinely unacceptable, the control is not the brief — it is not forking that task.
**Why the report reads so confidently.** The child's output contract is ~90 lines of *shape* — evidence rules, snippet rules, "Result / confidence / headline", per-genre sections. Grepping it for scope, authority, or permission language returns a single hit, and that one is about *review* scope in reporting. Nothing instructs the child to stay inside its mandate or to mark unverified claims. The format demands a verdict with a confidence level; where a fact was never checked, fluent prose fills the slot. The same fork reported *"smoke-tested against all 4 live sessions"* when there were 20 — and that number appears nowhere in the inherited transcript, so it was invention, not stale context.
**Practical rules:**
- State decision authority explicitly, every time — and add the anti-inheritance clause (task-design item 5) whenever the brief is narrower than the conversation.
- Require a **"What I did NOT do"** section on any write-capable fork.
- **Verify mutations from the filesystem, never from the report.** `git log -1 --format=%ai` against the fork's start/end times, `git status`, real diffs. Read a fork's push as an unreviewed PR from a stranger.
- **A brief containing a prohibition is a judgment task.** Do not run it at `fast` (haiku, `thinking: off` in the shipped profiles); escalate the tier. Reserve `fast` for "return raw output, no interpretation".
- Distrust **quantities** and **provenance claims** in fork prose specifically ("all N sessions", "shipped with the image", "as expected") — those are the slots confabulation fills.
- The fact that the fork was "right anyway" is not the same as the fork having followed instructions.
### Anti-patterns
- **Forking trivial work.** A fork has overhead. If the task takes < 30 seconds in your main thread, just do it.
- **Vague briefs.** "Look into the database thing" returns vague output. The fork is not telepathic.
- **Forking iterative work.** Forks are one-shot. If you need to iterate, you'll re-spec the task each time — usually worse than doing it yourself.
- **Recursive forking** (forks spawning forks). Disabled by default and should stay disabled unless you have a specific batch-fanout use case.
- **Treating fork output as ground truth without verification.** Especially for cited code/commit hashes/URLs — forks can hallucinate these like any LLM. Spot-check decisive evidence.
**Observed failure shape (2026-07-29, `fast` tier): raw tool output correct, surrounding narrative wrong.** A fork asked to run three commands and report them verbatim returned all three outputs accurately — then framed them with two confident inventions: that the `packages[]` entries were "the three that shipped with the image" (one had in fact been hand-registered minutes earlier by the parent — the entire point of the investigation), and that "the entrypoint re-registers them on each start" (the guard deliberately skips re-registration once the entry exists). Neither claim was in the command output; both were plausible glue.
**Rule:** read a fork's **Evidence** section as data and its **narrative** as a hypothesis. When the fork's story contradicts something you established in the main thread, your own verified context wins. Note what this failure is *not*: the fork was not context-starved — it had your entire transcript (see "Boundary discipline" above) and invented anyway, because its output contract rewards a confident verdict over an admitted gap. Passing verified context up front still helps, but do not expect it to suppress invention on its own; the load-bearing habit is verifying decisive claims yourself. Being right about the evidence is not the same as being right.
---
## Part 2: pi-observational-memory
### How it actually works
Observational memory (OM v3, "session-ledger" architecture) runs an **observer agent** in the background as your conversation grows. When token thresholds are crossed (defaults: observe at 10k, reflect at 20k, compact at 81k), the observer distills the recent transcript into:
- **Observations** — timestamped events, each with a 12-character hex ID like `[3682ebfad7af]`. Compact one-liners describing what happened in the conversation.
- **Reflections** — durable, long-lived facts about the user, project, decisions, and constraints. Some reflections include observation IDs as evidence pointers.
When compaction fires, the raw transcript is folded away and replaced with a structured summary block containing the observations + reflections. **You — the next turn of the same agent — receive that summary block as your starting context.** That's the recovery mechanism.
**Storage is in-transcript, not on disk.** Do not grep for `observations.jsonl` or similar files; you will not find them. The artifact lives in the model's input context window.
Configuration lives in `~/.pi/agent/settings.json` under `observational-memory`. Tune `observeAfterTokens`, `reflectAfterTokens`, `compactAfterTokens`, and `observationsPoolMaxTokens` if observations feel sparse or noisy. The default 81k compaction threshold is well-calibrated for typical multi-task sessions.
### The `recall` tool
`recall(<12-char-hex-id>)` resolves a specific observation or reflection ID back to the original source context — the exact bash output, file contents, tool call results, commit message, or transcript fragment that the observation was distilled from.
**Use recall when:**
- You are about to make a decision that depends materially on a compacted observation or reflection whose details are unclear.
- You need exact wording, paths, commands, errors, commits, or user constraints behind a remembered claim.
- A broad reflection is relevant but you need its supporting observations to act safely.
- The user asks "why do you believe X" or "what supports that memory".
**Do not use recall for:**
- Semantic search (it's keyed by ID, not topic — you must already have a specific 12-char hex ID).
- Browsing the transcript out of curiosity.
- Preemptive lookup of every ID in your context "just in case".
Recall costs tokens. Use it when exact source context will materially change your next action.
> **Calibration note (from a real ~1-month trial, 2026-05/06):** across 20 logged container sessions, `recall` was invoked **0 times** while obsmem passively carried 529 observations across 6 compactions. Zero recall is a *warning sign*, not a badge of efficiency — it means decisions after a compaction were made on the distilled one-liner alone, without ever re-checking the source. The injected summary is **lossy by design**. Default habit to adopt: when you are about to **edit code, ship a change, or assert a fact** that rests on a `[high]`/`[critical]` observation or a reflection you did not produce *this* turn, `recall` its ID **first**. One recall before a load-bearing action is cheap; redoing finished work or contradicting a prior correction is not.
### Reading the compaction summary
When you see a block like `The conversation history before this point was compacted into the following summary:` at the start of a session or turn, that's OM output. Standard structure:
- **Reflections** at the top: stable facts. Some have IDs in brackets.
- **Observations** below, chronological: timestamped events with IDs in brackets and importance markers (`[high]`, `[critical]`, etc.).
When entries conflict, **the most recent observation reflects the latest known state.** Work that prior observations describe as completed should not be redone unless the user explicitly asks to revisit it.
### Anti-patterns
- **Treating compacted memory as definitive without recall** when stakes are high. Compaction is lossy; the observation may have lost a constraint that was on the line above it in the original transcript.
- **Recalling every ID preemptively.** Wasteful. Recall on demand.
- **Assuming the disk holds OM artifacts.** It doesn't. Don't waste time looking.
- **Ignoring the summary block** when starting a session. It's there because the prior session was real work — read it before answering questions about past work.
---
## Quick Reference
```
fork(task=..., effort=fast|balanced|deep)
- state decision authority explicitly
- pass verified context up front
- specify deliverable shape
- ask for "unsure about" section
- if the brief is narrower than the conversation, say so:
"inherited history is NOT your mandate; obey this brief and report conflicts"
- write-capable? demand "What I did NOT do", then verify from git/fs, not the report
- prohibition in the brief => not a `fast` task
recall(id=<12-char-hex>)
- only when stakes justify the cost
- id must already be visible in your context
- not a search tool
```
```
~/.pi/agent/settings.json
pi-fork.effortProfiles — model + thinking-depth per tier
pi-fork.defaultEffort — usually "balanced"
observational-memory.* — token thresholds, model, agentMaxTurns
observational-memory.debugLog: true — opt-in NDJSON telemetry at
~/.pi/agent/observational-memory/debug/<session>.ndjson (off by default)
```
### Installing on a fresh machine (host)
These are git-sourced pi packages (pi-fork is **not** on npm). Add to the
`packages` array in `~/.pi/agent/settings.json`, or:
```
pi install git:github.com/elpapi42/pi-fork
pi install git:github.com/elpapi42/pi-observational-memory # default branch: master (no main)
# obsmem is also published: pi install npm:pi-observational-memory
```
Then `/reload` in a running session, or restart pi. Enable
`observational-memory.debugLog` if you want the next window instrumented.
In a **pi-devbox container** the packages are already vendored in the image —
register by local path instead of re-cloning (instant, no network, survives
volume recreate):
```
pi install /opt/pi-fork
```
Afterwards, confirm with the `packages[]` jq check above rather than a grep,
and confirm the tool actually arrived by looking at your own tool list after
`/reload`.
### Evaluating usage
`evaluate-extension-usage.py` (bundled next to this skill) mines pi session
transcripts for fork/recall counts and obsmem compaction stats. Run it per
machine (transcripts live at `~/.pi/agent/sessions/`) for a combined
host+container picture:
```
./evaluate-extension-usage.py # ~/.pi/agent/sessions
./evaluate-extension-usage.py /path/a /path/b # multiple roots
```
Read a **zero** carefully before treating it as a habit problem: a missing
`fork <== pi-fork` line means the tool was never *called*, which can equally
mean it was never *registered* (see the `packages[]` case study above). Check
registration first, then blame habits.
---
## Part 3: ssh-controlmaster
### What it does
When pi is launched with `--ssh`, this extension **rewires pi's `read`, `write`, `edit`, and `bash` tools to execute on the remote machine**, multiplexed over a single SSH ControlMaster socket. Pi is still running locally — the LLM, the UI, the MCP servers, the fork dispatcher all live on your local box — but anything those tools touch on the filesystem is the *remote's* filesystem.
This is fundamentally different from running pi locally and using `bash` to ssh inside it: with `--ssh`, the tool layer itself is remoted, so the LLM thinks it's working in the remote's `cwd` (the system prompt is rewritten to say so).
### Usage
```bash
# Key-based auth (preferred), remote cwd defaults to remote $HOME
pi --ssh lagret
# Pin to a specific remote directory
pi --ssh lagret:/volume1/docker/portainer/compose/119
# Password auth (input is NOT masked when typing)
pi --ssh user@host --ssh-ask-pass
```
The `lagret` form requires a `Host lagret` block in `~/.ssh/config` or a resolvable hostname. The status bar shows `SSH ⚡ own master <host>:<cwd>` or `SSH ⚡ system master <host>:<cwd>` once connected.
### How it cooperates with system SSH config
It reads `ssh -G <host>` to learn the effective config, then:
| `~/.ssh/config` for the host | Behavior |
|---|---|
| `ControlMaster auto` or `yes` with a `ControlPath` | Reuses the system master socket. Does **not** tear it down on pi exit ("it was the system's to manage before pi arrived"). |
| No ControlMaster configured (or explicitly `no`) | Creates its own master at `/tmp/pi-cm-<pid>.sock` with `ControlPersist=yes`. Tears it down on pi `session_shutdown`. |
This means it composes cleanly with the system-wide `ssh-control-master-setup.sh` helper from the `ci-release-watcher` skill: if that script has already configured `~/.ssh/config` for the host, `pi --ssh` rides on the existing master rather than opening a parallel connection.
### Caveats and edge cases
- **Local vs remote tool boundary.** Only `read`/`write`/`edit`/`bash` are remoted. **MCP servers are still local** — `mempalace` files drawers and diary entries against the local palace even when your shell work happens remotely. Same for `fork`, `recall`, `todo`, and any other custom tool. This is usually what you want (palace memory survives across remote sessions) but worth knowing.
- **fork over ssh.** Forks spawn locally and inherit the same `--ssh` mode by virtue of the parent's tool wiring; the fork's bash calls hit the same ControlMaster. Forks burn the same SSH socket, not a parallel one — multiplexing wins again.
- **macOS Unix socket path limit.** The own-master socket lives at `/tmp/pi-cm-<pid>.sock` to stay under macOS's ~104-char limit. If you have a non-default `TMPDIR` long enough to blow this, ssh will fail to start the master.
- **Password auth password visibility.** From the source: *"input is NOT masked — the password is visible while typing."* The password is written to a chmod-700 SSH_ASKPASS script in `/tmp` and deleted after the master establishes; not persisted, but on-screen during entry.
- **Remote bash environment.** The remote shell is whatever `ssh user@host '<cmd>'` invokes — typically a non-login non-interactive bash. Don't expect `~/.bashrc` aliases or PATH manipulations from `~/.profile`. Pin tool paths or invoke via `bash -lc '...'` if you need login-shell behavior.
- **Path translation is naive.** The extension does `path.replace(localCwd, remoteCwd)` to translate paths in tool calls. If the LLM emits an absolute remote path that doesn't share the local-cwd prefix, the path is passed through unchanged — usually fine but pathological for paths that happen to contain the local-cwd substring.
### When to use it
- Editing configs on a NAS / homelab host without scp ping-pong (`pi --ssh lagret:/volume1/...`)
- Operating against a host whose tools/data you need but whose disk is too slow to mount via SSHFS
- Investigating runner state, container configs, etc., on a remote host as if local
- Multi-step remote work where opening a fresh ssh connection per step would burn your CGNAT flow budget
### Anti-patterns
- **Using `pi --ssh` for one-off shell work.** Just `ssh` directly. The extension shines when there are dozens of tool calls per session.
- **Filing palace drawers expecting them on the remote.** They go to the local palace. If you want palace artifacts on the remote host, ssh into the remote and run pi *there* against its local palace.
- **Forgetting `--ssh` in followup sessions.** Status bar is the canary — if you don't see `SSH ⚡` you're operating locally despite intending remote. Easy mistake on a fresh terminal.
### Reaching the devbox host from inside the container (`dssh` / `dscp`)
Distinct from `pi --ssh` above. When the **pi-devbox container** runs under OrbStack / Docker Desktop on macOS, it can SSH back to its own host. The entrypoint's `setup-lan-access.sh` regenerates `~/.ssh-local/config` on **every container start** (the in-container `~/.ssh` is mounted read-only, so a sidecar config + `known_hosts` + `ControlPath` under `~/.ssh-local/` is used instead).
```bash
# Interactive shells get aliases (from ~/.bash_aliases):
dssh host 'cmd' # = ssh -F ~/.ssh-local/config host
dscp file host:/path # = scp -F ~/.ssh-local/config ...
```
**The agent's `bash` tool is non-interactive — those aliases are NOT loaded.** Use the explicit form:
```bash
ssh -F ~/.ssh-local/config host 'cmd'
scp -F ~/.ssh-local/config <src> host:<dst>
```
- Host aliases `host` and `mac` both resolve to `host.docker.internal` (user varies per host machine — check `~/.ssh-local/config` for the active `User` value, key `~/.ssh-local/devbox_jump_ed25519`, `ControlMaster auto` / `ControlPersist 4h`).
- The config chains `Include ~/.config/devbox-shell/ssh-lan.conf` then `Include ~/.ssh/config`, so LAN targets are reachable too (add `ProxyJump host` to those entries).
- **Use it for:** enabling/inspecting the host's pi config (`~/.pi/agent/settings.json`), running `evaluate-extension-usage.py` against the host's `~/.pi/agent/sessions/` for a combined host+container metric, or copying host transcripts into the container. The host's pi runs natively there; its palace, sessions, and extensions are separate from the container's.
---
## Cross-Skill Notes
- **mempalace** is for cross-session persistent memory (diary, knowledge graph, drawer storage). OM is for **within-session** context survival across compaction. They complement each other: write a diary entry at session end *and* let OM compact your work-in-progress mid-session.
- **systematic-debugging** and **test-driven-development** skills pair well with deep-tier forks: a deep fork can carry out a focused debugging investigation or write a failing test suite without polluting your main context.
- **ci-release-watcher** ships a `scripts/ssh-control-master-setup.sh` helper that configures system-wide SSH ControlMaster in `~/.ssh/config`. That's a separate mechanism from the `ssh-controlmaster` pi extension — they compose, they don't overlap. Use the script for persistent host-wide multiplexing, the extension for per-pi-session remote operation.
@@ -0,0 +1,117 @@
#!/usr/bin/env python3
"""Evaluate pi-fork / pi-observational-memory usage from pi session transcripts.
Mines pi's session .jsonl transcripts and reports:
- per-tool call counts (highlighting `fork` and `recall`)
- per-session fork/recall breakdown
- obsmem passive activity: compaction events, observations carried,
relevance-tier distribution, tokensBefore
Works on any machine. Point it at one or more session roots; by default it
scans ~/.pi/agent/sessions (the standard pi location, host or container).
Usage:
./evaluate-extension-usage.py # ~/.pi/agent/sessions
./evaluate-extension-usage.py /path/to/sessions ... # explicit roots
./evaluate-extension-usage.py --host HOST /path ... # label a root (for combined host+container runs)
For a true host+container picture, run once per machine (or copy each
machine's ~/.pi/agent/sessions here) and pass all roots together.
"""
import json, sys, os, glob, re, collections, argparse
TIER_RE = re.compile(r'\[(low|medium|high|critical)\]')
OBS_LINE_RE = re.compile(r'^\[[0-9a-f]{12}\] ', re.M)
def walk_tools(x, counter):
if isinstance(x, dict):
tn = x.get("toolName")
if tn:
counter[tn] += 1
for v in x.values():
walk_tools(v, counter)
elif isinstance(x, list):
for v in x:
walk_tools(v, counter)
def analyze(roots):
files = []
for r in roots:
if os.path.isfile(r) and r.endswith(".jsonl"):
files.append(r)
else:
files += glob.glob(os.path.join(r, "**", "*.jsonl"), recursive=True)
files = sorted(set(files))
tool_total = collections.Counter()
per_session = []
compactions = []
for f in files:
tc = collections.Counter()
with open(f, errors="ignore") as fh:
for ln in fh:
ln = ln.strip()
if not ln:
continue
try:
o = json.loads(ln)
except Exception:
continue
walk_tools(o, tc)
if o.get("type") == "compaction":
s = o.get("summary", "") or ""
compactions.append({
"file": os.path.basename(f),
"tokensBefore": o.get("tokensBefore"),
"observations": len(OBS_LINE_RE.findall(s)),
"tiers": dict(collections.Counter(TIER_RE.findall(s))),
})
tool_total.update(tc)
per_session.append((os.path.basename(f)[:10], tc.get("fork", 0),
tc.get("recall", 0), sum(tc.values())))
return files, tool_total, per_session, compactions
def main():
ap = argparse.ArgumentParser()
ap.add_argument("roots", nargs="*",
default=[os.path.expanduser("~/.pi/agent/sessions")])
args = ap.parse_args()
files, tool_total, per_session, comp = analyze(args.roots)
if not files:
print("No .jsonl transcripts found under:", args.roots, file=sys.stderr)
sys.exit(1)
print(f"=== {len(files)} transcripts under {args.roots} ===\n")
print("Tool call totals:")
for t, c in tool_total.most_common():
mark = " <== pi-fork" if t == "fork" else (" <== obsmem recall" if t == "recall" else "")
print(f" {c:6d} {t}{mark}")
fk = tool_total["fork"]; rc = tool_total["recall"]
fk_sess = sum(1 for p in per_session if p[1])
rc_sess = sum(1 for p in per_session if p[2])
print(f"\npi-fork: {fk} calls across {fk_sess} sessions")
print(f"recall: {rc} calls across {rc_sess} sessions"
+ (" (!) zero recall over the window — see SKILL.md calibration note" if rc == 0 else ""))
if comp:
tot_obs = sum(c["observations"] for c in comp)
tb = [c["tokensBefore"] for c in comp if c["tokensBefore"]]
print(f"\nobsmem passive: {len(comp)} compactions, {tot_obs} observations carried"
+ (f", avg tokensBefore {sum(tb)//len(tb):,}" if tb else ""))
agg = collections.Counter()
for c in comp:
agg.update(c["tiers"])
if agg:
print(" relevance tiers:", dict(agg))
else:
print("\nobsmem passive: no compaction events found "
"(short sessions, or obsmem not active on these transcripts)")
if __name__ == "__main__":
main()
@@ -0,0 +1,16 @@
# Skills in this directory whose OWNER is the skillset repo.
#
# Read by devbox-skill-reconcile, which runs after the skillset deploy in
# entrypoint-user.sh: for each name below, if the mounted skillset ships a
# skill of that name, the baked link in ~/.agents/skills/ is repointed at the
# live clone. The baked copy remains the fallback for containers started
# WITHOUT a skillset mount, and a user override always wins over both.
#
# Add a name here ONLY if the skillset repo is the authoritative source (see
# the ownership table in VENDORED.md). Do NOT add:
# pi-devbox-environment — authored in this repo; baked IS canonical
# pi-extensions — owned by the pi-extensions package repo and copied
# over the snapshot at build time; the skillset copy
# is a downstream duplicate that can lag, so letting
# it win would regress the skill.
mempalace
@@ -0,0 +1,14 @@
# xterm-ghostty — alias of the maintained ncurses `ghostty` terminfo entry.
#
# Ghostty sets TERM=xterm-ghostty by default, but the ncurses terminfo
# database (Debian: ncurses-term) ships the entry under the name `ghostty`
# only — there is no `xterm-ghostty` alias, and no distro packages one. This
# thin alias makes xterm-ghostty resolve to the same upstream-maintained
# capability set, so SSH sessions from a Ghostty terminal work without
# vendoring Ghostty's full (Zig-generated) terminfo here.
#
# `use=ghostty` is resolved by `tic` at compile time against the base
# `ghostty` entry from ncurses-term (installed in Dockerfile.base before the
# compile step). Compiled with `tic -x`.
xterm-ghostty|Ghostty terminal emulator (xterm-ghostty alias),
use=ghostty,
+43
View File
@@ -0,0 +1,43 @@
#!/usr/bin/env bash
# check-base-hash.sh — guard the base-rebuild invariant.
#
# Every floating `ARG *_REF` consumed by Dockerfile.base MUST be folded
# into the base_tag hash in the docker-publish workflow. Otherwise a
# ref-only change to that dependency does not change the base hash, the
# Docker Hub probe finds the old base tag, and the base is NOT rebuilt —
# the dependency fix silently fails to land. This is the v1.1.2-class
# staleness footgun (then it was mempalace-toolkit; this guard stops the
# next one before it ships).
#
# Runs in CI (base-decide job) and locally: bash scripts/check-base-hash.sh
set -euo pipefail
cd "$(dirname "$0")/.."
WF=".gitea/workflows/docker-publish.yml"
DF="Dockerfile.base"
# Extract the hash-compute block: the `HASH=$( … ) | sha256sum | cut`
# brace-group in the "Compute base tag" step. This lives in a separate
# file from the workflow, so scanning $WF here is free of the self-match
# hazard an inline workflow step would have.
block=$(awk '/HASH=\$\(/{f=1} f{print} f && /cut -c1-12/{exit}' "$WF")
if [ -z "$block" ]; then
echo "::error::could not locate the HASH=\$( … ) | sha256sum block in $WF"
exit 1
fi
refs=$(grep -oE '^ARG [A-Z0-9_]+_REF' "$DF" | awk '{print $2}' | sort -u)
fail=0
for r in $refs; do
lc=$(printf '%s' "$r" | tr '[:upper:]' '[:lower:]')
if ! printf '%s' "$block" | grep -q "outputs.$lc"; then
echo "::error::Dockerfile.base declares '$r' but it is NOT folded into the base_tag hash in $WF."
echo "::error::Add echo \"\${{ needs.resolve-versions.outputs.$lc }}\" inside the HASH=\$( … ) | sha256sum block, or a $r-only change will silently fail to rebuild the base."
fail=1
fi
done
if [ "$fail" = 0 ]; then
echo "OK: all Dockerfile.base *_REF args are folded into base_tag (${refs:-none})."
fi
exit $fail
+65
View File
@@ -0,0 +1,65 @@
#!/usr/bin/env bash
# Gitea-accurate guard against the recurring "bash syntax under the default
# sh/dash shell" footgun (ed49b8d resolve-versions; b7197e8 promote-base-latest,
# run 418).
#
# WHY A CUSTOM CHECK AND NOT JUST actionlint:
# actionlint models *GitHub* Actions, whose default `run` shell is bash. It
# therefore assumes a step that omits `shell:` runs under bash, and does NOT
# flag `set -o pipefail` there. Gitea Actions' default is `sh` (dash), so the
# exact bug we hit (omit `shell:`, use bash syntax) is invisible to actionlint.
# actionlint only fires when a step *explicitly* declares `shell: sh`.
#
# THE INVARIANT THIS ENFORCES:
# Every `run:` step in every .gitea/workflows/*.yml must resolve to an
# effective shell of `bash` — via the step's own `shell:`, a job-level
# `defaults.run.shell`, or a workflow-level `defaults.run.shell`. Any step
# that would fall through to Gitea's `sh` default is a FAILURE, because a
# future author adding bash syntax to it fails silently in CI.
#
# Pair this with actionlint (which catches explicit `shell: sh` + bash syntax,
# expression errors, and much else). Together they cover the class on Gitea.
set -euo pipefail
WF_DIR="${1:-.gitea/workflows}"
python3 - "$WF_DIR" <<'PY'
import sys, glob, os
try:
import yaml
except ImportError:
sys.stderr.write("ERROR: python3 yaml module missing (apt install python3-yaml)\n")
sys.exit(2)
wf_dir = sys.argv[1]
files = sorted(glob.glob(os.path.join(wf_dir, "*.yml")) + glob.glob(os.path.join(wf_dir, "*.yaml")))
if not files:
sys.stderr.write(f"ERROR: no workflow files under {wf_dir}\n")
sys.exit(2)
problems = []
for f in files:
with open(f) as fh:
doc = yaml.safe_load(fh) or {}
wf_shell = (((doc.get("defaults") or {}).get("run") or {}).get("shell"))
jobs = doc.get("jobs") or {}
for jname, job in jobs.items():
job = job or {}
job_shell = (((job.get("defaults") or {}).get("run") or {}).get("shell"))
steps = job.get("steps") or []
for i, step in enumerate(steps):
step = step or {}
if "run" not in step:
continue # `uses:` steps have no shell
eff = step.get("shell") or job_shell or wf_shell or "sh" # Gitea default = sh
if eff != "bash":
name = step.get("name") or f"step[{i}]"
problems.append(f"{f}: job '{jname}' / '{name}': effective shell = '{eff}' (Gitea default is sh; declare shell: bash or a bash default)")
if problems:
sys.stderr.write("Workflow shell guard FAILED — bash default not guaranteed:\n")
for p in problems:
sys.stderr.write(f" - {p}\n")
sys.exit(1)
print(f"Workflow shell guard OK — all run: steps in {len(files)} workflow file(s) resolve to bash.")
PY
+233 -25
View File
@@ -2,13 +2,14 @@
# Runtime post-recreate verification for pi-devbox.
#
# Verifies that after `docker compose up -d --force-recreate`:
# - The new image is actually live (pi version matches, when an expected
# version is supplied — see the version note below)
# - The new image is actually live (both the pi version and — when asked —
# the pi-devbox image release tag; see the two version notes below)
# - Persisted named volumes survived (~/.pi config, shell history, zoxide,
# nvim data, uv cache, ssh-local)
# - pi runtime wiring is intact: keybindings symlink, ≥4 extensions, the
# mempalace.ts bridge, settings.json, and the pi-fork /
# pi-observational-memory / (studio variant) pi-studio package registrations
# - pi runtime wiring is intact: keybindings symlink, AGENTS.md symlink,
# ≥4 extensions, the mempalace.ts bridge, settings.json, and the pi-fork /
# pi-observational-memory / (studio variant) pi-studio package
# registrations in settings.json packages[]
# - Shell defaults re-seeded from /etc/skel-devbox
# - /tmp/sshcm exists with mode 700 (ssh ControlMaster dir)
# - /opt toolkits intact
@@ -24,13 +25,33 @@
# the pi-devbox repo (which a maintainer already has for CI builds). A plain
# `docker pull` consumer is not the audience and will not have this file.
#
# Version note: pi's version is resolved from `latest` at CI build time and is
# NOT pinned to a concrete value in Dockerfile.variant (ARG PI_VERSION=latest).
# So unlike opencode-devbox, this script cannot self-derive an expected version
# from the Dockerfile. Pass --expected-version to assert a match; without it the
# live pi version is reported as an informational WARN, not a failure.
# TWO DIFFERENT VERSIONS, TWO DIFFERENT FLAGS. This distinction has already
# cost a release day, so it is spelled out here and in AGENTS.md step 4:
#
# Usage: ./scripts/recreate-sanity-check.sh [--expected-version X.Y.Z] [--variant studio|plain]
# --expected-version the PI CODING AGENT version, e.g. 0.84.3
# (`pi --version`; pinned as ARG PI_VERSION in
# Dockerfile.variant, which CI reads as the source
# of truth)
# --expected-image-version the PI-DEVBOX IMAGE release tag, e.g. 1.8.9 or
# v1.8.9 (the `release_tag` baked into
# /etc/pi-devbox/build-manifest.json)
#
# Passing a release tag to --expected-version used to report
# "pi version mismatch: expected 1.8.8, got 0.84.3" — an accusation aimed at
# the wrong component, on the last gate of a release. Both flags now detect
# being handed the other one's value and say so instead.
#
# Neither flag is required. Both values are derivable from the image's own
# build manifest, so by default the script asserts the LIVE pi version against
# the version recorded at build time — which is not a tautology: a stale
# `pi` in the ~/.pi/npm-global volume can shadow the baked one, exactly the
# way a stale npm:pi-atelier can (see the packages[] check below). Pass the
# flags when you want an assertion against a value you name yourself, which
# is what a release checklist wants.
#
# Usage: ./scripts/recreate-sanity-check.sh [--expected-version X.Y.Z]
# [--expected-image-version X.Y.Z]
# [--variant studio|plain]
#
# Exit codes:
# 0 all checks passed
@@ -40,22 +61,61 @@
set -euo pipefail
EXPECTED_VERSION=""
EXPECTED_IMAGE_VERSION=""
VARIANT=""
REPO_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
MANIFEST=/etc/pi-devbox/build-manifest.json
# Parse arguments
usage() {
cat >&2 <<'EOF'
usage: recreate-sanity-check.sh [--expected-version X.Y.Z]
[--expected-image-version X.Y.Z]
[--variant studio|plain]
--expected-version pi coding agent version, e.g. 0.84.3 (`pi --version`)
--expected-image-version pi-devbox image release tag, e.g. 1.8.9 or v1.8.9
--variant studio|plain (auto-detected when omitted)
These are two different versions. Both are read from the image's own build
manifest when the corresponding flag is omitted.
EOF
}
# Parse arguments. Every flag takes a value, so reject a missing one rather
# than swallowing the next flag as if it were the value.
need_value() {
case "${2:-}" in
""|-*)
echo "$1 requires a value" >&2
usage
exit 2
;;
esac
}
while [[ $# -gt 0 ]]; do
case "$1" in
--expected-version)
need_value "$@"
EXPECTED_VERSION="$2"
shift 2
;;
--expected-image-version)
need_value "$@"
EXPECTED_IMAGE_VERSION="$2"
shift 2
;;
--variant)
need_value "$@"
VARIANT="$2"
shift 2
;;
--help|-h)
usage
exit 0
;;
*)
echo "usage: $0 [--expected-version X.Y.Z] [--variant studio|plain]" >&2
echo "unknown option: $1" >&2
usage
exit 2
;;
esac
@@ -66,6 +126,19 @@ pass() { echo " ✓ $1"; }
fail() { echo " ✗ $1" >&2; FAILED=$((FAILED + 1)); }
warn() { echo " ⚠ $1" >&2; }
# Read one top-level field from the build manifest, or print nothing. The
# manifest is the image's own ground truth (written at `docker build` time by
# Dockerfile.variant), so it needs no checkout and no network. Absent on an
# image built before it existed, hence every caller treats "" as unknown.
manifest_field() {
[ -f "$MANIFEST" ] || return 0
command -v jq >/dev/null 2>&1 || return 0
jq -r --arg k "$1" '.[$k] // empty' "$MANIFEST" 2>/dev/null || true
}
# Release tags are written with a leading v in the manifest and quoted without
# one in checklists; compare on the bare number so both spellings work.
strip_v() { printf '%s' "${1#v}"; }
# Auto-detect variant if not provided. The studio variant vendors pi-studio to
# /opt/pi-studio; the plain variant does not.
if [ -z "$VARIANT" ]; then
@@ -85,21 +158,59 @@ else
fi
echo
echo "-- pi version --"
MANIFEST_PI_VERSION=$(manifest_field pi_version)
MANIFEST_RELEASE_TAG=$(manifest_field release_tag)
echo "-- pi (coding agent) version --"
if ACTUAL_VERSION=$(pi --version 2>&1 | head -1); then
if [ -n "$EXPECTED_VERSION" ]; then
if [ "$ACTUAL_VERSION" = "$EXPECTED_VERSION" ]; then
pass "pi version $ACTUAL_VERSION"
if [ "$(strip_v "$EXPECTED_VERSION")" = "$(strip_v "$ACTUAL_VERSION")" ]; then
pass "pi version $ACTUAL_VERSION (matches --expected-version)"
elif [ -n "$MANIFEST_RELEASE_TAG" ] &&
[ "$(strip_v "$EXPECTED_VERSION")" = "$(strip_v "$MANIFEST_RELEASE_TAG")" ]; then
# Exact, not heuristic: the value handed over IS this image's release
# tag, so it cannot be a pi version anyone meant.
fail "--expected-version $EXPECTED_VERSION is the pi-devbox IMAGE version, not the pi version — use --expected-image-version $EXPECTED_VERSION (live pi is $ACTUAL_VERSION)"
else
fail "pi version mismatch: expected $EXPECTED_VERSION, got $ACTUAL_VERSION"
fail "pi version mismatch: expected $EXPECTED_VERSION, got $ACTUAL_VERSION (this flag asserts the pi coding agent version; for the image release tag use --expected-image-version)"
fi
elif [ -n "$MANIFEST_PI_VERSION" ]; then
# Not a tautology: the manifest records what pi reported at BUILD time,
# while `pi --version` resolves through PATH, which a stale npm-global
# volume install can shadow.
if [ "$MANIFEST_PI_VERSION" = "$ACTUAL_VERSION" ]; then
pass "pi version $ACTUAL_VERSION (matches this image's build manifest)"
else
fail "live pi $ACTUAL_VERSION != $MANIFEST_PI_VERSION recorded in $MANIFEST — a stale pi in the ~/.pi/npm-global volume is shadowing the baked one"
fi
else
warn "pi version $ACTUAL_VERSION (no --expected-version given; pi is built from 'latest', cannot self-derive — informational only)"
warn "pi version $ACTUAL_VERSION (no --expected-version and no build manifest to compare against — informational only)"
fi
else
fail "pi --version failed"
fi
echo
echo "-- pi-devbox image version --"
if [ -z "$MANIFEST_RELEASE_TAG" ]; then
if [ -n "$EXPECTED_IMAGE_VERSION" ]; then
fail "cannot verify --expected-image-version $EXPECTED_IMAGE_VERSION: no readable release_tag in $MANIFEST (image built before the manifest existed, or jq missing)"
else
warn "image release tag unknown (no readable $MANIFEST) — pi-devbox-version would say the same"
fi
elif [ -n "$EXPECTED_IMAGE_VERSION" ]; then
if [ "$(strip_v "$EXPECTED_IMAGE_VERSION")" = "$(strip_v "$MANIFEST_RELEASE_TAG")" ]; then
pass "image version $MANIFEST_RELEASE_TAG (matches --expected-image-version)"
elif [ -n "$MANIFEST_PI_VERSION" ] &&
[ "$(strip_v "$EXPECTED_IMAGE_VERSION")" = "$MANIFEST_PI_VERSION" ]; then
fail "--expected-image-version $EXPECTED_IMAGE_VERSION is the pi version, not the image release tag — use --expected-version $EXPECTED_IMAGE_VERSION (this image is $MANIFEST_RELEASE_TAG)"
else
fail "image version mismatch: expected $EXPECTED_IMAGE_VERSION, got $MANIFEST_RELEASE_TAG — the recreate did not pick up the intended image"
fi
else
warn "image version $MANIFEST_RELEASE_TAG (no --expected-image-version given — informational only)"
fi
echo
echo "-- Persisted named volumes (must survive --force-recreate) --"
@@ -157,6 +268,14 @@ else
fail "~/.pi/agent/keybindings.json missing or not a symlink"
fi
# global AGENTS.md symlink (pi-toolkit) — global instructions loaded by pi at
# every start (directs the agent to read the pi-extensions skill at session start)
if [ -L "$HOME/.pi/agent/AGENTS.md" ]; then
pass "~/.pi/agent/AGENTS.md symlink (pi-toolkit)"
else
fail "~/.pi/agent/AGENTS.md missing or not a symlink"
fi
# extensions deployed (pi-extensions) — expect ≥4 *.ts
EXT_COUNT=$(ls -1 "$HOME"/.pi/agent/extensions/*.ts 2>/dev/null | wc -l | tr -d ' ')
if [ "$EXT_COUNT" -ge 4 ]; then
@@ -179,23 +298,102 @@ else
fail "~/.pi/agent/settings.json missing"
fi
# pi package registrations (pi install <local-path> → recorded in settings.json)
# settings.json merge: the entrypoint deep-merges new template keys into a
# preserved settings.json on every start, so config added in an image upgrade
# (e.g. the observational-memory / pi-fork blocks) reaches existing volumes.
# Assert those blocks are present and that the file is still valid JSON.
if command -v jq >/dev/null 2>&1 && [ -f "$HOME/.pi/agent/settings.json" ]; then
if jq -e 'has("observational-memory") and has("pi-fork")' "$HOME/.pi/agent/settings.json" >/dev/null 2>&1; then
pass "settings.json has observational-memory + pi-fork blocks (template merge)"
else
fail "settings.json missing observational-memory and/or pi-fork blocks (template merge did not land)"
fi
fi
# pi package registrations (pi install <local-path> → recorded in settings.json).
# Check the `packages` ARRAY, not the whole file: the settings template ships a
# top-level "pi-fork" CONFIG block (asserted just above), so `grep -q pi-fork
# settings.json` is a guaranteed false green — which is how an un-registered
# fork tool went unnoticed from v1.0.0 through v1.6.3. Same array check the
# fixed entrypoint-user.sh guard uses.
_pkg_registered() {
_s="$HOME/.pi/agent/settings.json"
[ -f "$_s" ] || return 1
if command -v jq >/dev/null 2>&1; then
jq -e --arg n "$1" \
'(.packages // []) | any((type == "string") and (. == "npm:" + $n or endswith("/" + $n)))' \
"$_s" >/dev/null 2>&1
else
grep -q "opt/$1\"" "$_s"
fi
}
# True when a literal `npm:pi-atelier` entry is still present — the
# volume-resident registration the entrypoint migrates away from.
_npm_atelier_present() {
_s="$HOME/.pi/agent/settings.json"
[ -f "$_s" ] || return 1
command -v jq >/dev/null 2>&1 || return 1
jq -e '(.packages // []) | any(. == "npm:pi-atelier")' "$_s" >/dev/null 2>&1
}
if [ -f "$HOME/.pi/agent/settings.json" ]; then
for pkg in pi-fork pi-observational-memory; do
if grep -q "$pkg" "$HOME/.pi/agent/settings.json" 2>/dev/null; then
pass "$pkg registered in settings.json"
if _pkg_registered "$pkg"; then
pass "$pkg registered in settings.json packages[]"
else
fail "$pkg not registered in settings.json"
fail "$pkg NOT in settings.json packages[] (tool will not load)"
fi
done
if [ "$VARIANT" = "studio" ]; then
if grep -q "pi-studio" "$HOME/.pi/agent/settings.json" 2>/dev/null; then
pass "pi-studio registered in settings.json"
if _pkg_registered pi-studio; then
pass "pi-studio registered in settings.json packages[]"
else
fail "pi-studio not registered in settings.json (studio variant)"
fail "pi-studio NOT in settings.json packages[] (studio variant)"
fi
fi
# pi-atelier — vendored from v1.7.0 on. Absent on older images, and
# deliberately unregistered when DEVBOX_ATELIER=0; neither is a failure.
if [ -d /opt/pi-atelier ]; then
if [ "${DEVBOX_ATELIER:-1}" = "0" ]; then
if _pkg_registered pi-atelier; then
fail "pi-atelier still in packages[] despite DEVBOX_ATELIER=0"
else
pass "pi-atelier unregistered (DEVBOX_ATELIER=0, as requested)"
fi
elif _pkg_registered pi-atelier; then
pass "pi-atelier registered in settings.json packages[]"
else
fail "pi-atelier NOT in settings.json packages[] (sidebar will not load)"
fi
if _npm_atelier_present; then
fail "stale npm:pi-atelier still in packages[] — it resolves through the ~/.pi/npm-global VOLUME and shadows the pinned /opt copy (entrypoint migration did not run)"
fi
fi
fi
# ── pi <-> pi-atelier compatibility floor ─────────────────────────────
# atelier < 0.7.1 wraps pi's private TUI renderer in a way that recurses under
# pi >= 0.84: pi hangs at startup burning CPU, with no error message. atelier's
# own peerDependencies (>=0.80.7) do not encode this. Assert it here too, not
# just in the build-time smoke test: this script runs after a real
# `--force-recreate` on a live box, where a volume-resident old copy is exactly
# what could bite.
if [ -d /opt/pi-atelier ] && command -v jq >/dev/null 2>&1; then
_ge() { [ "$(printf '%s\n%s\n' "$1" "$2" | sort -V | head -n1)" = "$2" ]; }
_av=$(jq -r '.version // empty' /opt/pi-atelier/package.json 2>/dev/null || true)
_pv=$(pi --version 2>/dev/null | grep -oE '[0-9]+\.[0-9]+\.[0-9]+' | head -n1 || true)
if [ -n "$_av" ] && [ -n "$_pv" ]; then
if _ge "$_pv" 0.84.0 && ! _ge "$_av" 0.7.1; then
fail "pi $_pv with pi-atelier $_av — atelier < 0.7.1 hangs pi >= 0.84 at startup (bump PI_ATELIER_REF in Dockerfile.variant)"
else
pass "pi $_pv + pi-atelier $_av (compatibility floor OK)"
fi
else
warn "could not compare pi/pi-atelier versions (pi='$_pv' atelier='$_av')"
fi
fi
echo
@@ -214,6 +412,16 @@ else
fail "~/.bash_aliases missing"
fi
# History flush must survive shell nesting. The DEVBOX_HIST_SET guard must NOT
# be exported: if it leaks into child processes, nested shells (esp. tmux
# panes) skip installing `history -a` and lose in-memory history on abrupt
# termination. Assert a child login shell still wires up the per-prompt flush.
if bash -lic 'bash -lic "case \"\$PROMPT_COMMAND\" in *\"history -a\"*) exit 0;; *) exit 1;; esac"' </dev/null >/dev/null 2>&1; then
pass "nested shell installs 'history -a' (DEVBOX_HIST_SET not exported)"
else
fail "nested shell missing 'history -a' — DEVBOX_HIST_SET leaking to children?"
fi
if [ -f "$HOME/.inputrc" ]; then
pass "~/.inputrc exists"
else
+550 -16
View File
@@ -5,16 +5,26 @@
#
# Verifies:
# - pi binary present and (if EXPECTED_PI_VERSION set) matches CI's resolved version
# - mempalace core matches the audited pin (if EXPECTED_MEMPALACE_VERSION set)
# - new v1.0.0 base additions (pandoc, graphviz, imagemagick, yq, tealdeer)
# - typst PDF engine for pandoc (v1.4.0) — `pandoc --pdf-engine=typst`
# - non-modal editors nano + micro (alongside nvim)
# - terminfo for modern emulators: xterm-kitty, xterm-ghostty, wezterm,
# alacritty, foot (kitty-terminfo + ncurses-term + compiled ghostty alias)
# - tmux 0-indexing baked in /etc/tmux.conf (required for pi-studio variants)
# - pi-toolkit cloned at /opt/pi-toolkit
# - pi-extensions cloned at /opt/pi-extensions
# - pi-atelier vendored at /opt/pi-atelier, registered from /opt (not npm:),
# and >= the version floor pi's TUI requires (see the floor test)
# - pi-fork + pi-observational-memory cloned with node_modules baked
# - entrypoint deploys pi-toolkit keybindings symlink
# - entrypoint deploys ≥4 extensions
# - mempalace bridge symlink present
# - settings.json bootstrapped
# - pi-fork + pi-observational-memory registered via `pi install`
# - pi-fork + pi-observational-memory registered in settings.json packages[]
# via `pi install`
# - pi-devbox-version command present + wraps the build manifest correctly
# (human, --json, --quiet)
# - (studio variant only, auto-detected) pi-studio cloned + prebuilt
# client bundle present + registered via `pi install`
# - image size within threshold
@@ -24,18 +34,33 @@ set -euo pipefail
IMAGE="${1:?usage: $0 <image>}"
PASS=0; FAIL=0
# pi-devbox v1.0.0 (decoupled from opencode-devbox) added pandoc, graphviz,
# imagemagick, yq, tealdeer, and a baked /etc/tmux.conf. Local arm64 build
# observed 3.20 GB. CI amd64 builds may differ slightly; threshold below
# carries +300 MB margin to absorb arch differences without false reds.
# Tighten in a follow-up release once amd64 actuals are observed in CI logs.
SIZE_THRESHOLD_MB=3500
# imagemagick, yq, tealdeer, a baked /etc/tmux.conf, and the non-modal
# editors nano + micro (~15 MB combined). v1.6.0 baked in agent-browser +
# Playwright Chromium (~291 MB net after dropping the unused headless-shell
# build), which lifted the baseline. CI amd64 actuals observed on run 512
# (v1.6.1): 3411 MB non-studio, 3574 MB studio. Threshold below carries
# ~225 MB margin above the studio number to absorb minor arch/build-cache
# differences and small future growth without false reds, while still
# catching an unexpected +GB regression.
SIZE_THRESHOLD_MB=3800
# On failure, surface the last few lines the command produced. This used to
# discard output entirely (`>/dev/null 2>&1`), which made a red ❌ carry zero
# diagnostic weight: explaining the single v1.8.0 stage-default failure took a
# full CI-log dig plus a registry-config inspection, when the container had
# already printed the answer and thrown it away. Assertions that want a
# diagnostic just echo it to stderr — it stays hidden while they pass.
run() {
local label="$1"; local cmd="$2"
if docker run --rm --entrypoint="" "$IMAGE" sh -c "$cmd" >/dev/null 2>&1; then
local out
if out=$(docker run --rm --entrypoint="" "$IMAGE" sh -c "$cmd" 2>&1); then
printf " ✅ %s\n" "$label"; PASS=$((PASS+1))
else
printf " ❌ %s\n" "$label"; FAIL=$((FAIL+1))
# `if`, not `&&` — a trailing false under `set -e` would abort the script.
if [ -n "$out" ]; then
printf " └─ %s\n" "$(printf '%s' "$out" | tail -3 | tr '\n' ' ' | cut -c1-300)"
fi
fi
}
@@ -71,15 +96,183 @@ run "git" "git --version"
run "aws" "aws --version"
run "uv" "uv --version"
run "nvim" "nvim --version"
run "nano" "nano --version"
run "micro" "micro --version"
run "kitty-terminfo" "infocmp -x xterm-kitty >/dev/null 2>&1"
run "terminfo: modern emulators (ncurses-term)" 'for t in wezterm alacritty foot ghostty st-256color; do infocmp -x "$t" >/dev/null 2>&1 || exit 1; done'
run "terminfo: xterm-ghostty alias (tic)" "infocmp -x xterm-ghostty >/dev/null 2>&1"
run "nvim true-colour default (sysinit.vim)" "nvim --headless -c 'lua os.exit(vim.o.termguicolors and 0 or 1)'"
run "mempalace-mcp" "mempalace-mcp --help"
run "mempalace-pi-session on PATH" "mempalace-pi-session --help"
# The staging dir must sit next to the palace, not in a disposable cache: the
# palace keys per-source dedup on the STAGED path, so a stage that can be wiped
# while the palace survives lets `mempalace sync` prune every drawer mined from
# it. Assert the resolved default, not an env var — the guarantee is "stage
# shares the palace's lifetime", which an ENV pin would quietly break.
# NOTE: --sessions-dir gets an EMPTY temp dir, never /tmp. The stage banner is
# printed before any export, so nothing needs to be found — and pointing a
# default-staged run at a populated dir would export whatever transcripts it
# finds into the real stage, which is how a synthetic test session ends up
# staged for mining as if it were a real conversation.
#
# Asserted $HOME-RELATIVE, not against a literal /home/developer. `run` invokes
# `docker run --entrypoint=""`, and neither Dockerfile sets USER or ENV HOME
# (HOME is set by entrypoint-user.sh, which --entrypoint="" deliberately skips),
# so these assertions execute as root with HOME=/root. The original literal
# /home/developer form could therefore never match and failed the v1.8.0
# release — a test bug, not a product one: the stage resolution was correct all
# along, it just follows $HOME. The invariant under test ("the stage sits beside
# the palace, sharing its lifetime") is user-independent, so pinning the user
# was never part of it. A cache-dir default still fails the pattern below, which
# is the regression this guards.
#
# It went unnoticed for three days because this workflow only triggers on
# `push: tags: v*` — the assertion was added on a main push, so v1.8.0 was its
# first execution ever. Use the `smoke_only` workflow_dispatch input to run
# smoke against HEAD without cutting a tag.
run "pi stage defaults next to the palace (not a cache dir)" '
out=$(mempalace-pi-session --dry-run --reason smoke --sessions-dir "$(mktemp -d)" 2>&1) || true
stage=$(echo "$out" | grep -oE "stage=[^ ]+" | head -1)
echo "resolved ${stage:-<no stage= line>} with HOME=$HOME" >&2
case "$stage" in
"stage=$HOME/.mempalace/pi-stage/"*) exit 0 ;;
*) exit 1 ;;
esac
'
# Companion to the above: the deployment-specific case the literal assertion was
# reaching for, done properly by supplying the HOME the container actually runs
# with instead of assuming it.
run "pi stage is palace-adjacent for the developer user" '
out=$(HOME=/home/developer mempalace-pi-session --dry-run --reason smoke --sessions-dir "$(mktemp -d)" 2>&1) || true
echo "$out" | grep -oE "stage=[^ ]+" | head -1 >&2
echo "$out" | grep -q "stage=/home/developer/.mempalace/pi-stage/"
'
run "pi stage follows MEMPALACE_PALACE_PATH" '
out=$(MEMPALACE_PALACE_PATH=/tmp/alt/.mempalace/palace \
mempalace-pi-session --dry-run --reason smoke --sessions-dir "$(mktemp -d)" 2>&1) || true
echo "$out" | grep -q "stage=/tmp/alt/.mempalace/pi-stage/"
'
# The feeder's --agent default is WHO a drawer is attributed to. mempalace core
# records neither the machine nor the harness on a write, and one shared bearer
# token means the server cannot tell clients apart, so toolkit c64ffa1 changed
# this default from $USER to pi@$MEMPALACE_PI_DEVICE — the one string that makes
# a write attributable to both. Nothing ever PRINTED the resolved value (the
# banner shows mode= and stage= only), so an image built from a pre-c64ffa1
# toolkit ref would ship unattributed writes with every check still green.
#
# `--help` assigns AGENT (script top) before it parses args, then exits 0 with
# no side effects — so `bash -x` observes the REAL resolution, env interpolation
# and fallback included, rather than grepping the source for a literal line that
# any reformat would break. Two-sided on purpose: device set => pi@<device>;
# device UNSET => must not be pi@anything. The second half is what fails against
# the old unconditional $USER default, which ignored the device entirely.
#
# Probes the PATH entry (a symlink into the /opt clone) rather than that clone
# path directly: this is the invocation the systemd/launchd timers and
# entrypoint-user.sh actually use, so it is the default that reaches the palace.
run "feeder resolves --agent to pi@<device> (drawer attribution)" '
f=$(command -v mempalace-pi-session) || { echo "feeder not on PATH" >&2; exit 1; }
with=$(MEMPALACE_PI_DEVICE=smoke-device bash -x $f --help 2>&1 | sed -n "s/^+* *AGENT=//p" | tail -n1)
without=$(env -u MEMPALACE_PI_DEVICE bash -x $f --help 2>&1 | sed -n "s/^+* *AGENT=//p" | tail -n1)
echo "resolved with-device=[$with] without-device=[$without]" >&2
[ "$with" = "pi@smoke-device" ] || exit 1
case "$without" in pi@*) exit 1 ;; esac
echo ok
'
# Regression guard for the pi transcript exporter. If pi ever changes its
# session JSONL shape, the exporter stops recognising sessions and the palace
# silently gets nothing (or, worse, raw JSON chunked as prose). Feed it a
# synthetic session and assert it is actually exported. Uses --dry-run so no
# palace is touched, and a temp stage so nothing real is written.
run "pi transcript exporter recognises a pi session" '
set -e
d=$(mktemp -d); s="$d/sessions/--workspace--"; mkdir -p "$s"
{
printf "%s\n" "{\"type\":\"session\",\"version\":1,\"id\":\"smoke\",\"cwd\":\"/workspace\",\"timestamp\":\"2026-01-01T00:00:00Z\"}"
printf "%s\n" "{\"type\":\"message\",\"message\":{\"role\":\"user\",\"content\":\"question one\"}}"
a=$(printf "a%.0s" $(seq 1 1200))
printf "%s\n" "{\"type\":\"message\",\"message\":{\"role\":\"assistant\",\"content\":[{\"type\":\"text\",\"text\":\"$a\"}]}}"
printf "%s\n" "{\"type\":\"message\",\"message\":{\"role\":\"user\",\"content\":\"question two\"}}"
printf "%s\n" "{\"type\":\"message\",\"message\":{\"role\":\"assistant\",\"content\":[{\"type\":\"text\",\"text\":\"short reply\"}]}}"
} > "$s/2026-01-01T00-00-00-000Z_smoke.jsonl"
out=$(mempalace-pi-session --dry-run --sessions-dir "$d/sessions" --stage "$d/stage" 2>&1)
echo "$out" | grep -q "Exported 1 session"
'
# The same guard from the other side: a session with no real assistant output
# (an abandoned prompt, whose bulk is injected skill text) must NOT be filed.
run "pi transcript exporter rejects an abandoned session" '
set -e
d=$(mktemp -d); s="$d/sessions/--workspace--"; mkdir -p "$s"
{
printf "%s\n" "{\"type\":\"session\",\"version\":1,\"id\":\"smoke2\",\"cwd\":\"/workspace\",\"timestamp\":\"2026-01-01T00:00:00Z\"}"
u=$(printf "u%.0s" $(seq 1 13000))
printf "%s\n" "{\"type\":\"message\",\"message\":{\"role\":\"user\",\"content\":\"$u\"}}"
printf "%s\n" "{\"type\":\"message\",\"message\":{\"role\":\"assistant\",\"content\":[{\"type\":\"text\",\"text\":\"Ready. What would you like to work on?\"}]}}"
} > "$s/2026-01-01T00-00-00-000Z_smoke2.jsonl"
out=$(mempalace-pi-session --dry-run --sessions-dir "$d/sessions" --stage "$d/stage" 2>&1)
echo "$out" | grep -q "no sessions qualified"
'
# The remote-palace-without-inbox skip must ANNOUNCE itself, not vanish. This
# branch of entrypoint-user.sh runs at container start (not reachable from a
# `docker run` one-shot), so assert against the entrypoint that actually shipped
# in the image. Guards a silent regression back to the bare `:` no-op, which
# left a container contributing nothing to the palace with no artifact saying
# why — the log it would normally leave is written by the other branch.
run_expect "remote-palace-without-inbox skip is announced, not silent" \
"grep -o 'MemPalace catch-up skipped' /usr/local/bin/entrypoint-user.sh | head -1" \
"MemPalace catch-up skipped"
run "...and the skip notice names the variable that fixes it" \
"grep -A6 'MemPalace catch-up skipped' /usr/local/bin/entrypoint-user.sh | grep -q 'MEMPALACE_PI_SSH_TARGET'"
# A remote mine that FAILS must not report success. MCP answers a hard tool
# failure with HTTP 200 and the tool's own JSON escaped inside
# result.content[].text, so the feeder's old `'\"error\"' in body` check could
# never see it: on 2026-08-15 a mine that died with "source directory not found:
# '/data/feed/...'" logged "Done. Wing updated." and exited 0, and this
# container's transcripts were filed nowhere for a whole session. The feeder
# carries fixtures for that exact body; run them against the baked toolkit so a
# stale/reverted toolkit ref can't reintroduce a silent feed.
run "baked feeder detects a failed remote mine (no silent false success)" \
"mempalace-pi-session --self-test"
# v1.0.0 base additions — verify presence and basic functionality.
run "pandoc" "pandoc --version"
run "typst" "typst --version"
run "pandoc+typst PDF engine" "printf '# hi\n' | pandoc --pdf-engine=typst -o /tmp/_smoke.pdf - && test -s /tmp/_smoke.pdf; rm -f /tmp/_smoke.pdf"
run "graphviz (dot)" "dot -V"
run "imagemagick" "magick --version"
run "yq" "yq --version"
run "yq (mikefarah v4)" "yq --version | grep -qE 'mikefarah.*version v4'"
run "tldr (tealdeer)" "tldr --version"
run "socat" "socat -V"
run "studio-expose helper" "test -x /usr/local/bin/studio-expose"
run "image-baked pi-devbox-environment skill" \
"test -f /usr/local/share/pi-devbox/skills/pi-devbox-environment/SKILL.md"
run "global-AGENTS append snippet present" \
"test -f /usr/local/share/pi-devbox/pi-global-AGENTS.append.md"
run "pi-devbox block merged into pi-global-AGENTS.md" \
"grep -q 'pi-devbox:managed-block' /opt/pi-toolkit/pi-global-AGENTS.md"
run "mempalace session-start pointer merged into global AGENTS.md" \
"grep -q 'load the mempalace skill' /opt/pi-toolkit/pi-global-AGENTS.md"
# Vendored fallback skills (so a no-skillset container still resolves the
# AGENTS.md 'read the pi-extensions skill' pointer).
run "image-baked pi-extensions fallback skill" \
"test -f /usr/local/share/pi-devbox/skills/pi-extensions/SKILL.md"
run "pi-extensions skill ships its helper" \
"test -f /usr/local/share/pi-devbox/skills/pi-extensions/evaluate-extension-usage.py"
run "image-baked mempalace fallback skill" \
"test -f /usr/local/share/pi-devbox/skills/mempalace/SKILL.md"
# Layered freshness: when the pinned pi-extensions clone carries the skill, the
# baked copy must be the fresh package copy (Option 1), not the stale snapshot.
run "pi-extensions skill refreshed from package when present" \
"if [ -f /opt/pi-extensions/skill/SKILL.md ]; then cmp -s /opt/pi-extensions/skill/SKILL.md /usr/local/share/pi-devbox/skills/pi-extensions/SKILL.md; else true; fi"
# Runtime ownership handover (v1.8.5): the baked links are a FALLBACK, and
# skillset-OWNED skills must be repointed at the live clone when one is mounted.
# The list is data, so assert its content, not just its presence: mempalace in,
# pi-extensions deliberately out (its skillset copy is a lagging duplicate).
run "devbox-skill-reconcile helper present + executable" \
"test -x /usr/local/bin/devbox-skill-reconcile"
run "skillset-owned list ships and names mempalace" \
"grep -qx 'mempalace' /usr/local/share/pi-devbox/skills/skillset-owned.txt"
run "skillset-owned list excludes pi-extensions (ownership)" \
"! grep -qx 'pi-extensions' /usr/local/share/pi-devbox/skills/skillset-owned.txt"
# ── tmux 0-indexing (required for pi-studio variants) ─────────────────
echo ""
@@ -98,6 +291,44 @@ run "pi-fork clone + node_modules" \
"test -f /opt/pi-fork/package.json && test -d /opt/pi-fork/node_modules"
run "pi-observational-memory clone + node_modules" \
"test -f /opt/pi-observational-memory/package.json && test -d /opt/pi-observational-memory/node_modules"
# ...and that the clone carries the AUTH FIX, not merely that it exists. om's
# pre-flight hasUsableAuth() check silently disabled `recall` for ~8 weeks once
# pi moved to request-time SigV4 signing and stopped exposing a static Bedrock
# key; upstream fixed it in ce9fc98, adopted in v1.8.4. PI_OBSMEM_REF tracks
# master, so an upstream revert or force-push would ship a dead `recall` with
# the clone assertion above still green — the exact gap flagged as open in the
# v1.8.5 changelog.
#
# Pin the markers to src/runtime.ts, the fix SITE, rather than grepping the
# repo: two of these three strings also appear under tests/, so a repo-wide
# grep stays green with runtime.ts itself reverted. That is a false green of the
# same family as the old skill-snapshot canary.
run "pi-observational-memory carries the ce9fc98 auth fix (recall stays alive)" '
f=/opt/pi-observational-memory/src/runtime.ts
test -f "$f" || { echo "fix site missing: $f" >&2; exit 1; }
for m in availability_recheck providerCredentialConfigured hasConfiguredAuth; do
grep -q "$m" "$f" || { echo "marker absent from runtime.ts: $m" >&2; exit 1; }
done
echo ok
'
# pi-atelier: deliberately NO node_modules assertion, unlike its siblings —
# it declares zero runtime dependencies (only peerDeps, satisfied by the baked
# pi) and has no build step, so Dockerfile.variant skips `npm install` for it.
# Assert what pi actually loads instead: the entry point named by its
# package.json `pi.extensions` key.
run "pi-atelier clone + entry point" \
"test -f /opt/pi-atelier/package.json && test -f /opt/pi-atelier/extensions/index.ts"
# ── pi <-> pi-atelier compatibility floor (executable, not a comment) ──
# pi-atelier < 0.7.1 wraps pi's PRIVATE TUI renderer in a way that recurses
# under pi >= 0.84: pi hangs at startup burning CPU, with no error. Upstream
# fixed it in 0.7.1/0.7.2, but atelier's peerDependencies still say
# `>=0.80.7`, so neither npm nor pi can warn about the real floor. Both
# versions are pinned in Dockerfile.variant; this makes a bad PAIRING fail the
# build instead of publishing an image whose TUI never starts.
run_expect "pi-atelier >= 0.7.1 floor for pi >= 0.84 (startup-hang guard)" \
'ge() { [ "$(printf "%s\n%s\n" "$1" "$2" | sort -V | head -n1)" = "$2" ]; }; AV=$(jq -r ".version // empty" /opt/pi-atelier/package.json 2>/dev/null); PV=$(pi --version 2>/dev/null | grep -oE "[0-9]+\.[0-9]+\.[0-9]+" | head -n1); if [ -z "$AV" ] || [ -z "$PV" ]; then echo "unreadable versions (atelier=$AV pi=$PV)"; elif ge "$PV" 0.84.0 && ! ge "$AV" 0.7.1; then echo "VIOLATION: pi $PV with pi-atelier $AV"; else echo "compatible: pi $PV + pi-atelier $AV"; fi' \
"compatible:"
# pi-studio is present only in the :latest-studio variant. Auto-detect by
# probing /opt/pi-studio so this one script covers both variants.
@@ -113,6 +344,171 @@ else
echo " ℹ️ pi-studio not present (non-studio variant) — skipping studio clone checks"
fi
# ── Build provenance (manifest + OCI labels) ─────────────────────────
echo ""
echo "── Build provenance ──"
run "/etc/pi-devbox/build-manifest.json present" \
"test -f /etc/pi-devbox/build-manifest.json"
# These next checks replace three that grepped the manifest for the FIELD NAME
# and never looked at the value:
#
# run_expect "manifest records pi_version" "cat …manifest.json" '"pi_version"'
#
# which passes on {"pi_version": ""} and on {"pi_version": null}. The tell was
# visible in its own passing output — `✅ manifest records pi_version (got
# "pi_version")` echoes the key back as the thing it claims to have found.
# Two failure modes were therefore invisible: a key that survives with an empty
# or garbage value, and a key that vanishes from the manifest while every
# remaining value still looks fine.
#
# Those two need SEPARATE assertions, and the reason is a trap worth keeping in
# writing: an "every component value is a valid SHA" loop passes VACUOUSLY on
# components:{} — jq's all() over an empty list is true — so the value check
# alone would go green on a manifest that lost every component. Mutation-tested
# 2026-08-25 across nine fabricated manifests (empty map, deleted key, "",
# null, "unknown", 12-hex truncation, 40 non-hex chars, legit null pi-studio).
run "manifest declares every required component key" '
req="pi-toolkit pi-extensions pi-fork pi-observational-memory pi-atelier mempalace-toolkit pi-studio"
for k in $req; do
jq -e --arg k "$k" "(.components|has(\$k))" /etc/pi-devbox/build-manifest.json >/dev/null \
|| { echo "manifest lost component key: $k" >&2; exit 1; }
done
'
# Subsumes the old `! grep -q \"unknown\"` check ("unknown" is not 40-hex), and
# also catches "", null and truncated SHAs, which that grep let through. null is
# legitimate for pi-studio alone: the non-studio variant has no such clone.
run "manifest component values are resolved 40-hex commits" '
jq -e "
.components
| to_entries
| all(if .key == \"pi-studio\" and .value == null then true
else (.value|type) == \"string\" and (.value|test(\"^[0-9a-f]{40}\$\")) end)
" /etc/pi-devbox/build-manifest.json >/dev/null
'
# pi_version against ground truth, same shape as the mempalace check below.
# Chains with the "pi version matches build arg" assertion earlier in this file:
# together they tie build arg -> installed binary -> recorded manifest, so a
# manifest written from a stale variable cannot pass by agreeing with itself.
run "manifest pi_version matches the installed pi" '
m=$(jq -r ".pi_version // empty" /etc/pi-devbox/build-manifest.json)
b=$(pi --version 2>/dev/null | head -n1 | tr -d "\r")
echo "manifest=[$m] installed=[$b]" >&2
[ -n "$m" ] && [ "$m" = "$b" ]
'
# Top-level provenance fields: assert the SHAPE of each value, and only when the
# field is populated. source_revision and build_date legitimately default to
# empty (Dockerfile.variant ARGs) on a plain local `docker build`, so demanding
# them would fail honest local smoke runs; a populated-but-malformed value is
# the actual defect. release_tag defaults to "dev", so empty means a broken write.
run "manifest top-level fields are well-formed, not merely present" '
j=/etc/pi-devbox/build-manifest.json
t=$(jq -r ".release_tag // empty" $j)
r=$(jq -r ".source_revision // empty" $j)
d=$(jq -r ".build_date // empty" $j)
echo "release_tag=[$t] source_revision=[$r] build_date=[$d]" >&2
[ -n "$t" ] || { echo "release_tag empty (ARG default is dev)" >&2; exit 1; }
if [ -n "$r" ]; then
printf "%s" "$r" | grep -qxE "[0-9a-f]{40}" || { echo "source_revision not a 40-hex commit" >&2; exit 1; }
fi
if [ -n "$d" ]; then
printf "%s" "$d" | grep -qE "^[0-9]{4}-[0-9]{2}-[0-9]{2}T" || { echo "build_date not ISO-8601" >&2; exit 1; }
fi
'
# mempalace CORE was absent from the manifest through v1.8.5: the toolkit SHA
# was recorded but the palace version behind the MCP tools was not, so a palace
# bug could not be correlated to an image version. Assert the field exists AND
# equals the installed binary — recording it from ARG MEMPALACE_VERSION instead
# would look identical here yet drift silently the first time an install
# resolved to something other than the pin, which is the whole reason this file
# is built from ground truth. `// empty` matters: jq -r prints the 4-char
# string "null" for a JSON null, which would satisfy a naive -n test.
run "manifest mempalace_version matches the installed core" '
m=$(jq -r ".mempalace_version // empty" /etc/pi-devbox/build-manifest.json)
b=$(mempalace --version 2>/dev/null | head -n1 | tr -d "\r"); b=${b##* }
echo "manifest=[$m] installed=[$b]" >&2
[ -n "$m" ] && [ "$m" = "$b" ]
'
# ... and, when CI supplies it, that the installed core is the version CI
# actually AUDITED (published + not yanked on PyPI, in resolve-versions). This
# does NOT duplicate the check above, which compares two properties of one
# image and so cannot notice that BOTH are the wrong version. The live failure
# mode it covers: the variant builds `FROM` a base tag chosen by base-decide's
# content hash, so a bug in that hashing (the reason scripts/check-base-hash.sh
# exists) could reuse a cached base built from an OLDER MEMPALACE_VERSION pin —
# internally consistent, silently stale, invisible to every other assertion.
if [ -n "${EXPECTED_MEMPALACE_VERSION:-}" ]; then
run "installed mempalace matches CI's audited pin (${EXPECTED_MEMPALACE_VERSION})" "
b=\$(mempalace --version 2>/dev/null | head -n1 | tr -d '\r'); b=\${b##* }
echo \"installed=[\$b] audited_pin=[${EXPECTED_MEMPALACE_VERSION}]\" >&2
[ \"\$b\" = \"${EXPECTED_MEMPALACE_VERSION}\" ]
"
fi
# Every component must be a resolved commit (or null for pi-studio in the
# non-studio variant) — now enforced by the 40-hex value check above, which
# strictly subsumes the old whole-file grep for '"unknown"'. Only rev() ever
# emits "unknown" and rev() feeds components only, so nothing is lost.
# pi-devbox-version wraps the manifest into a human-first command; verify the
# binary is present, executable, and that all three output modes work.
run "pi-devbox-version binary present + executable" \
"test -x /usr/local/bin/pi-devbox-version"
run_expect "pi-devbox-version human output shows release tag" \
"pi-devbox-version" "pi-devbox "
# --json is a verbatim `cat` of the manifest, so "round-trips" is assertable
# literally. The old form grepped the output for the string "release_tag" — the
# key name again — which would pass on a truncated or re-serialised dump.
run "pi-devbox-version --json round-trips the manifest byte-for-byte" '
a=$(cat /etc/pi-devbox/build-manifest.json)
b=$(pi-devbox-version --json)
[ "$a" = "$b" ] || { echo "--json output differs from the manifest on disk" >&2; exit 1; }
'
run_expect "pi-devbox-version --quiet is a compact one-liner" \
"pi-devbox-version --quiet | wc -l" "1"
# ── Vendored skill snapshot provenance ─────────────────────────────────
# The vendored mempalace skill is the one baked artefact with no /opt clone
# behind it (private upstream — see VENDORED.md), so until now the manifest
# could not say which skillset commit it came from. Two fields now travel with
# it: the CLAIMED ref (ARG default in Dockerfile.variant) and the MEASURED
# sha256 of the shipped bytes. Assert both are well-formed, and — separately —
# that the measurement still describes the file in the image.
#
# Kept as two assertions for the same reason the component checks are: one
# proves the fields are not empty/garbage, the other proves they are not merely
# self-consistent. A single combined check could pass on a manifest whose hash
# was computed from a file that was later overwritten (the pi-extensions skill
# copy at Dockerfile.variant:165 does exactly that kind of overwrite, one stage
# earlier), which is the failure this second one exists to catch.
run "manifest records the vendored skill snapshot provenance" '
j=/etc/pi-devbox/build-manifest.json
r=$(jq -r ".skillset_snapshot_ref // empty" $j)
s=$(jq -r ".skillset_snapshot_tree_sha256 // empty" $j)
echo "ref=[$r] tree_sha256=[$s]" >&2
printf "%s" "$r" | grep -qxE "[0-9a-f]{40}" \
|| { echo "skillset_snapshot_ref is not a 40-hex commit" >&2; exit 1; }
printf "%s" "$s" | grep -qxE "[0-9a-f]{64}" \
|| { echo "skillset_snapshot_tree_sha256 is not a 64-hex digest" >&2; exit 1; }
'
# Recomputes over the whole DIRECTORY with the same tree_sha256() pipeline
# Dockerfile.variant used to measure it, not a plain `sha256sum SKILL.md` —
# a file-only compare here would pass even if the manifest recorded a
# fingerprint over a directory that has since grown a second file (this is
# not hypothetical: pi-extensions already ships two files for its skill).
run "manifest skill fingerprint matches the baked snapshot" '
j=/etc/pi-devbox/build-manifest.json
d=/usr/local/share/pi-devbox/skills/mempalace
m=$(jq -r ".skillset_snapshot_tree_sha256 // empty" $j)
a=$( (cd "$d" && find . -type f -print | LC_ALL=C sort | xargs -r sha256sum) | sha256sum | cut -d" " -f1)
echo "manifest=[$m] actual=[$a]" >&2
[ -n "$m" ] && [ "$m" = "$a" ]
'
# OCI labels live in the image config, not the container fs — inspect them
# from the host docker rather than via `docker run`.
LBL=$(docker inspect --format '{{ index .Config.Labels "se.jordbo.pi-devbox.pi-extensions-ref" }}' "$IMAGE" 2>/dev/null || true)
if [ -n "$LBL" ] && [ "$LBL" != "<no value>" ]; then
printf " ✅ OCI label se.jordbo.pi-devbox.pi-extensions-ref=%s\n" "$LBL"; PASS=$((PASS+1))
else
printf " ❌ OCI label se.jordbo.pi-devbox.pi-extensions-ref missing or empty\n"; FAIL=$((FAIL+1))
fi
# ── Runtime deployment (needs entrypoint to run) ──────────────────────
echo ""
echo "── Runtime deployment ──"
@@ -134,6 +530,9 @@ for i in $(seq 1 45); do
if docker exec "$CID" sh -c '
test -L /home/developer/.pi/agent/keybindings.json && \
test -L /home/developer/.pi/agent/extensions/mempalace.ts && \
test -L /home/developer/.agents/skills/pi-devbox-environment && \
test -L /home/developer/.agents/skills/pi-extensions && \
test -L /home/developer/.agents/skills/mempalace && \
count=$(ls -1 /home/developer/.pi/agent/extensions/*.ts 2>/dev/null | wc -l) && \
[ "$count" -ge 4 ]
' >/dev/null 2>&1; then
@@ -155,33 +554,168 @@ exec_test "keybindings.json (pi-toolkit)" 'test -L $HOME/.pi/agent/keybi
exec_test "extensions ≥ 4 (pi-extensions)" 'count=$(ls -1 $HOME/.pi/agent/extensions/*.ts 2>/dev/null | wc -l); [ $count -ge 4 ] && echo "$count extensions"'
exec_test "mempalace.ts bridge" 'test -L $HOME/.pi/agent/extensions/mempalace.ts && echo ok'
exec_test "settings.json bootstrapped" 'test -f $HOME/.pi/agent/settings.json && echo ok'
exec_test "pi-devbox-environment skill linked" 'test -L $HOME/.agents/skills/pi-devbox-environment && test -f $HOME/.agents/skills/pi-devbox-environment/SKILL.md && echo ok'
exec_test "pi-extensions skill linked (fallback)" 'test -L $HOME/.agents/skills/pi-extensions && test -f $HOME/.agents/skills/pi-extensions/SKILL.md && echo ok'
exec_test "mempalace skill linked (fallback)" 'test -L $HOME/.agents/skills/mempalace && test -f $HOME/.agents/skills/mempalace/SKILL.md && echo ok'
# The vendored mempalace snapshot is refreshed MANUALLY per release (see
# rootfs/usr/local/share/pi-devbox/skills/VENDORED.md). Through v1.8.4 it also
# silently SHADOWED the live skillset copy, so staleness was invisible — and the
# canary that was supposed to catch it could not: it grepped "Shared palace:
# multiple harnesses", a phrase present in BOTH the stale and the fresh copy.
# A snapshot canary must pin the NEWEST section, so update this string whenever
# the snapshot is refreshed — that is the point of it.
#
# v1.8.7: this fired for real, and on the release that changed the snapshot. The
# pinned phrase was "Attribute what you file yourself", the heading of the
# instruction telling agents to hand-stamp added_by — which that same release
# WITHDREW (RFC 001 §7.3.2 ranks agent-side stamping worst-possible; the bridge
# now does it). So the canary correctly reported "snapshot changed, expectation
# did not", and blocked publication of an otherwise-green build (81 passed, 1
# failed, twice). Two lessons kept in the assertion itself:
# * it is now BIDIRECTIONAL — the new phrase must be present AND the withdrawn
# one absent, so a re-vendored stale snapshot fails just as loudly as a
# forgotten bump. A one-way canary only catches half the drift.
# * a phrase canary can only ever detect "older than what I remembered to pin",
# never "older than skillset main".
#
# That structural limit is now addressed, but NOT by the "CI job diffing this
# file against the skillset repo" this comment used to point at (that pointer
# also dangled: it referenced an Unreleased changelog note that had become the
# v1.8.7 heading). A CI diff cannot be done without granting CI a credential
# for the PRIVATE skillset repo, and it would guard a file that on this fleet
# NO host reads — all four compose stacks mount a workspace containing the
# skillset, so devbox-skill-reconcile repoints this link at the live clone and
# the baked copy is a CI/no-mount fallback only. Instead the snapshot now
# carries its provenance (skillset_snapshot_ref + a measured
# skillset_snapshot_sha256 in build-manifest.json, written by
# scripts/vendor-mempalace-skill.sh), which moves the check to where the
# skillset actually IS: `scripts/vendor-mempalace-skill.sh --check` for a
# maintainer, and `pi-devbox-version` for an agent inside any container.
# This assertion is kept because it is orthogonal and free: it pins content,
# not provenance, so it still catches a re-vendored snapshot whose ref was
# bumped correctly but whose bytes came from the wrong place.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Provenance is stamped for you" "$f" && ! grep -q "Attribute what you file yourself" "$f" && echo ok'
# Link TARGETS, not just link existence: with no skillset mounted (as here) the
# baked tree must be what resolves, for all three vendored skills.
exec_test "vendored skills resolve to the baked tree (no skillset mounted)" \
'for s in mempalace pi-extensions pi-devbox-environment; do
case "$(readlink -f $HOME/.agents/skills/$s)" in
/usr/local/share/pi-devbox/skills/$s) ;;
*) echo "$s resolves to $(readlink -f $HOME/.agents/skills/$s)" >&2; exit 1 ;;
esac
done; echo ok'
# ... and that the tool REPORTS that resolution, which is the half that was
# missing: a stale baked snapshot and a current live clone were
# indistinguishable from inside the container. CI mounts no skillset, so every
# vendored skill must report "baked" here — which also makes this a real test of
# the fallback path rather than of the environment it happens to run in.
exec_test "pi-devbox-version reports skill sources (all baked, no skillset here)" \
'out=$(pi-devbox-version)
echo "$out" | grep -q "skills:" || { echo "no skills section" >&2; exit 1; }
for s in mempalace pi-extensions pi-devbox-environment; do
echo "$out" | grep -qE "^ $s +baked$" \
|| { echo "$s not reported as baked" >&2; exit 1; }
done; echo ok'
# The boot banner must NOT carry the section: entrypoint-user.sh prints the
# version FIRST, before the baked links exist and long before the skillset
# deploy + reconcile run last, so anything it said about skill sources would be
# a pre-reconcile state that is about to change.
# A bare negative (`! grep -q "skills:"`) passes if the tool crashes or
# prints nothing at all — it cannot tell "correctly omitted the section"
# apart from "the binary is broken". Anchor it positively: the command must
# still succeed and still print its normal release-tag line.
exec_test "pi-devbox-version --no-skills omits the skills section" \
'out=$(pi-devbox-version --no-skills) && echo "$out" | grep -q "^pi-devbox " && ! echo "$out" | grep -q "skills:"'
exec_test "entrypoint prints the version banner with --no-skills" \
'grep -q "pi-devbox-version --no-skills" /usr/local/bin/entrypoint-user.sh'
# The handover path itself. CI never mounts a skillset, so without this the
# v1.8.5 fix would ship untested: fabricate a skillset + a skills dir holding
# baked-style links, run the reconciler, and assert all three outcomes —
# owned skill repointed, unowned skill left baked, user override untouched.
exec_test "reconciler: owned skill handed to live clone, others untouched" \
'set -e; t=$(mktemp -d); mkdir -p $t/ss/skills/mempalace $t/ss/skills/pi-extensions $t/skills
echo LIVE > $t/ss/skills/mempalace/SKILL.md; echo LIVE > $t/ss/skills/pi-extensions/SKILL.md
ln -s /usr/local/share/pi-devbox/skills/mempalace $t/skills/mempalace
ln -s /usr/local/share/pi-devbox/skills/pi-extensions $t/skills/pi-extensions
mkdir -p $t/skills/mine; echo MINE > $t/skills/mine/SKILL.md
devbox-skill-reconcile $t/ss $t/skills >/dev/null
devbox-skill-reconcile $t/ss $t/skills >/dev/null # idempotent
[ "$(readlink $t/skills/mempalace)" = "$t/ss/skills/mempalace" ] || { echo "owned skill NOT repointed" >&2; exit 1; }
[ "$(readlink $t/skills/pi-extensions)" = /usr/local/share/pi-devbox/skills/pi-extensions ] || { echo "unowned skill was repointed" >&2; exit 1; }
[ "$(cat $t/skills/mine/SKILL.md)" = MINE ] || { echo "user override clobbered" >&2; exit 1; }
rm -rf $t; echo ok'
# The case above cannot fail if the reconciler stops checking WHERE a link
# points — a mutation test showed all three of its assertions still passing with
# that guard deleted, which is the same false-green shape as the old snapshot
# canary. This one discriminates: an OWNED name (so it is considered) whose link
# is a user override pointing outside the baked tree (so it must be left alone).
exec_test "reconciler: user override on an owned name is left alone" \
'set -e; t=$(mktemp -d); mkdir -p $t/ss/skills/mempalace $t/skills $t/mine-skill
echo LIVE > $t/ss/skills/mempalace/SKILL.md; echo USERLINK > $t/mine-skill/SKILL.md
ln -sfn $t/mine-skill $t/skills/mempalace
devbox-skill-reconcile $t/ss $t/skills >/dev/null
[ "$(cat $t/skills/mempalace/SKILL.md)" = USERLINK ] || { echo "user symlink override clobbered" >&2; exit 1; }
rm -rf $t; echo ok'
# mempalace-census gained a /usr/local/bin symlink in v1.8.3; its three siblings
# had one since they were added, so this asserts the set stays complete.
exec_test "mempalace-census on PATH" 'command -v mempalace-census >/dev/null && mempalace-census --help >/dev/null && echo ok'
# pi-fork + pi-observational-memory are registered by entrypoint-user.sh via
# `pi install /opt/<pkg>`, which runs slightly after the keybindings marker.
#
# Assert against the `packages` ARRAY, never a whole-file grep: the settings
# template ships a top-level "pi-fork" CONFIG block, so `grep -q pi-fork
# settings.json` passes even when `pi install /opt/pi-fork` never ran. That
# false green is exactly why the missing `fork` tool shipped unnoticed from
# v1.0.0 through v1.6.3.
pkg_registered_cmd() {
printf "jq -e --arg n %s '(.packages // []) | any((type == \"string\") and (. == \"npm:\" + \$n or endswith(\"/\" + \$n)))' \$HOME/.pi/agent/settings.json" "$1"
}
for i in $(seq 1 15); do
if docker exec "$CID" grep -q pi-observational-memory \
/home/developer/.pi/agent/settings.json 2>/dev/null; then
if docker exec -u developer "$CID" sh -c "$(pkg_registered_cmd pi-observational-memory)" \
>/dev/null 2>&1; then
break
fi
sleep 1
done
exec_test "pi-fork registered (fork tool)" 'grep -q pi-fork $HOME/.pi/agent/settings.json && echo ok'
exec_test "pi-observational-memory registered (recall tool)" 'grep -q pi-observational-memory $HOME/.pi/agent/settings.json && echo ok'
exec_test "pi-fork registered in packages[] (fork tool)" \
"$(pkg_registered_cmd pi-fork)"
exec_test "pi-observational-memory registered in packages[] (recall tool)" \
"$(pkg_registered_cmd pi-observational-memory)"
# pi-studio registration (studio variant only) — registered by the same
# entrypoint-user.sh local-path install loop as fork/obsmem.
if [ "${STUDIO_VARIANT:-0}" = "1" ]; then
for i in $(seq 1 15); do
if docker exec "$CID" grep -q pi-studio \
/home/developer/.pi/agent/settings.json 2>/dev/null; then
if docker exec -u developer "$CID" sh -c "$(pkg_registered_cmd pi-studio)" \
>/dev/null 2>&1; then
break
fi
sleep 1
done
exec_test "pi-studio registered (/studio command + studio_* tools)" \
'grep -q pi-studio $HOME/.pi/agent/settings.json && echo ok'
exec_test "pi-studio registered in packages[] (/studio command + studio_* tools)" \
"$(pkg_registered_cmd pi-studio)"
fi
# pi-atelier registration. It is LAST in the entrypoint's install loop, so a
# pass here also means that loop ran to completion rather than dying midway.
for i in $(seq 1 15); do
if docker exec -u developer "$CID" sh -c "$(pkg_registered_cmd pi-atelier)" \
>/dev/null 2>&1; then
break
fi
sleep 1
done
exec_test "pi-atelier registered in packages[] (TUI sidebar)" \
"$(pkg_registered_cmd pi-atelier)"
# ...and registered from the vendored /opt copy, NOT as `npm:pi-atelier`: an
# npm: entry resolves through ~/.pi/npm-global on the config VOLUME, which
# outlives image upgrades and would silently keep an old, unaudited atelier —
# exactly the shape that pairs a stale 0.6.x with a new pi and hangs at startup.
exec_test "pi-atelier registered from /opt, not npm: (volume-shadowing guard)" \
'jq -e "((.packages // []) | any((type == \"string\") and endswith(\"/pi-atelier\"))) and (((.packages // []) | any(. == \"npm:pi-atelier\")) | not)" $HOME/.pi/agent/settings.json'
# ── /tmp/sshcm directory created by entrypoint ────────────────────────
exec_test "/tmp/sshcm dir mode 700 (ssh ControlMaster)" \
'test -d /tmp/sshcm && [ "$(stat -c %a /tmp/sshcm)" = "700" ] && echo ok'
+293
View File
@@ -0,0 +1,293 @@
#!/usr/bin/env bash
# vendor-mempalace-skill.sh — refresh the vendored mempalace skill snapshot
# AND its recorded provenance, together, so the two cannot drift apart.
#
# WHY THIS EXISTS
# ---------------
# rootfs/usr/local/share/pi-devbox/skills/mempalace/SKILL.md is a snapshot of a
# file owned by the PRIVATE skillset repo (see VENDORED.md). Because the image
# cannot clone that repo, refreshing the snapshot was a manual `cp` — and the
# result was anonymous: nothing recorded WHICH skillset commit the bytes came
# from. The only staleness check available was a hand-maintained phrase canary
# in scripts/smoke-test.sh, which by construction detects "older than the phrase
# I remembered to pin", never "older than skillset main".
#
# Two facts now travel with the snapshot: the skillset commit it was taken from
# (ARG SKILLSET_SNAPSHOT_REF in Dockerfile.variant) and the sha256 of the bytes
# themselves (measured at build time into build-manifest.json). This script is
# the only thing that should ever write the first one, because a `cp` without a
# matching ARG bump produces a manifest that CONFIDENTLY LIES — worse than the
# anonymous snapshot it replaced.
#
# HARDENED after peer review (pi@emb-7kj4vr4g, logstream correlation
# skills-provenance-review, 2026-08-26) proved the original --check could print
# OK and exit 0 without actually verifying anything: `git show <ref>:<path>`
# emits NOTHING when the ref/path doesn't resolve, and `sha256sum` still hashes
# that empty stdin, so "ref not found" silently collided with "the file really
# is 0 bytes". Depending on which side of the comparison hit the collision this
# fell through as either a false MISMATCH (blaming provenance for what was
# really an incomplete clone) or, worse, a false OK. See EXIT STATUS below —
# "cannot determine" is now its own outcome, distinct from "confirmed wrong",
# which is the same distinction the phrase canary this script replaced lacked.
#
# USAGE
# scripts/vendor-mempalace-skill.sh [skillset-root] [--force]
# refresh: rewrite the snapshot and the ARG together.
# scripts/vendor-mempalace-skill.sh --check [skillset-root]
# verify only, writes nothing. The root path and any flag may appear in
# either order — a positional-only parser previously made `<root>
# --check` silently run a refresh instead of the verification asked for.
#
# skillset-root defaults to /workspace/skillset, then $HOME/skillset.
#
# --force (refresh mode only) proceed even when the recorded ref cannot be
# proven to be an ancestor of the skillset's current HEAD — i.e.
# skip the guard against silently REWINDING provenance, which a
# detached HEAD, an older checkout, or a shallow clone lacking the
# recorded commit can all trigger. Meant to be used deliberately,
# not habitually: each use is a human deciding a rewind is fine.
#
# --check answers "is the committed snapshot really skillset@<recorded ref>?"
# — the question CI cannot answer without a credential for the private repo,
# and which anyone with the skillset checked out can answer for free.
#
# EXIT STATUS (same three codes in both modes)
# 0 the operation succeeded, or (--check) the record is verified truthful.
# This INCLUDES a truthful record that is merely stale — upstream has
# moved on since the recorded ref, or the local working tree has since
# diverged. A NOTICE is printed to stderr, but the snapshot is not being
# accused of lying, so this is not a release-blocking failure. Skipping a
# refresh is a legitimate release-day choice (see AGENTS.md); this exit
# code is what makes that choice checkable rather than merely asserted.
# 1 refused: a CONFIRMED problem. Dirty upstream file; a refresh that would
# rewind past the recorded ref; or (--check) the vendored bytes provably
# do NOT match the file at the recorded ref — a lying record.
# 2 cannot determine: the recorded ref, or the path at that ref, is not
# resolvable in this clone. Commonly a shallow clone missing history, or
# a ref that was rewritten or never pushed. Deliberately NOT the same as
# 1 — "I can't tell" must never be reported as "it's wrong".
set -euo pipefail
cd "$(dirname "$0")/.."
DOCKERFILE="Dockerfile.variant"
VENDORED="rootfs/usr/local/share/pi-devbox/skills/mempalace/SKILL.md"
ARG_NAME="SKILLSET_SNAPSHOT_REF"
REL_PATH="skills/mempalace/SKILL.md"
die() { printf '%s: %s\n' "$(basename "$0")" "$1" >&2; exit 1; }
# Parse flags and the optional root path in either order, and reject anything
# unrecognised rather than silently absorbing it.
MODE="refresh"
FORCE=0
ROOT=""
for arg in "$@"; do
case "$arg" in
--check) MODE="check" ;;
--force) FORCE=1 ;;
# This is the one script whose argument ORDER was itself a landmine, so the
# path that documents the trap must not be the path that errors.
-h|--help)
awk 'NR>1 && /^#/ { sub(/^# ?/, ""); print; next } NR>1 { exit }' "$0"
exit 0
;;
--*) die "unknown option: $arg (try --help)" ;;
*)
[ -z "$ROOT" ] || die "unexpected extra argument: $arg (root already set to $ROOT)"
ROOT="$arg"
;;
esac
done
if [ "$MODE" = "check" ] && [ "$FORCE" = 1 ]; then
die "--force has no effect with --check (nothing is written); remove it"
fi
if [ -z "$ROOT" ]; then
for candidate in /workspace/skillset "$HOME/skillset"; do
if [ -d "$candidate/.git" ]; then
ROOT="$candidate"
break
fi
done
fi
[ -n "$ROOT" ] || die "no skillset clone found (pass one: $(basename "$0") /path/to/skillset)"
[ -d "$ROOT/.git" ] || die "not a git clone: $ROOT"
[ -f "$ROOT/$REL_PATH" ] || die "no $REL_PATH in $ROOT"
[ -f "$VENDORED" ] || die "vendored snapshot missing: $VENDORED"
# LOAD-BEARING, DO NOT DELETE AS "REDUNDANT WITH THE EXISTENCE PROBES": -f
# accepts an empty file, and sha256 of an empty file equals sha256 of a failed
# pipeline's empty stdin. Guarding it HERE, before mode dispatch, makes that
# collision unreachable by construction rather than by a probe further down --
# which also means no test below exercises the collision any more. Remove this
# line and the false "OK" for a nonexistent ref returns with nothing failing.
[ -s "$VENDORED" ] || die "vendored snapshot is empty: $VENDORED"
head_sha=$(git -C "$ROOT" rev-parse HEAD 2>/dev/null) || die "cannot read HEAD of $ROOT"
recorded=$(grep -oE "^ARG ${ARG_NAME}=[0-9a-f]{40}$" "$DOCKERFILE" | cut -d= -f2 || true)
[ -n "$recorded" ] || die "no 'ARG ${ARG_NAME}=<40-hex>' line in $DOCKERFILE"
sha_of() { sha256sum "$1" | cut -d' ' -f1; }
vendored_sha=$(sha_of "$VENDORED")
upstream_sha=$(sha_of "$ROOT/$REL_PATH")
# Does $REL_PATH exist at HEAD at all? Proven with `cat-file -e` BEFORE
# hashing anything. Piping a failed `git show` straight into sha256sum, as
# this script used to, hashes an EMPTY stream and produces sha256(""): a real,
# collidable value — not a representation of absence. That collapsed "doesn't
# exist" and "exists and happens to be empty" into the same signal, which is
# exactly the defect class the peer review found in --check's at_ref, below.
blob_sha=""
if git -C "$ROOT" cat-file -e "HEAD:$REL_PATH" 2>/dev/null; then
blob_sha=$(git -C "$ROOT" show "HEAD:$REL_PATH" | sha256sum | cut -d' ' -f1)
fi
upstream_dirty=""
if [ -z "$blob_sha" ]; then
upstream_dirty="not present at HEAD (untracked, or absent at this commit)"
elif [ "$blob_sha" != "$upstream_sha" ]; then
if ! git -C "$ROOT" diff --quiet -- "$REL_PATH" 2>/dev/null; then
upstream_dirty="modified but not committed"
elif ! git -C "$ROOT" diff --cached --quiet -- "$REL_PATH" 2>/dev/null; then
upstream_dirty="staged but not committed"
else
upstream_dirty="different at HEAD than in the working tree"
fi
fi
if [ "$MODE" = "check" ]; then
# Resolve the recorded ref the same careful way: existence is proven with
# `cat-file -e` before anything is hashed, and "the ref itself is missing"
# is reported distinctly from "the ref resolves but the path isn't there
# at it" — both used to be silently swallowed into a plausible sha256("").
ref_exists=0
path_at_ref_exists=0
at_ref=""
if git -C "$ROOT" cat-file -e "${recorded}^{commit}" 2>/dev/null; then
ref_exists=1
if git -C "$ROOT" cat-file -e "${recorded}:${REL_PATH}" 2>/dev/null; then
path_at_ref_exists=1
at_ref=$(git -C "$ROOT" show "${recorded}:${REL_PATH}" | sha256sum | cut -d' ' -f1)
fi
fi
printf 'recorded ref: %s\n' "$recorded"
printf 'vendored sha256: %s\n' "$vendored_sha"
if [ "$path_at_ref_exists" = 1 ]; then
printf 'sha256 at ref: %s\n' "$at_ref"
elif [ "$ref_exists" = 1 ]; then
printf 'sha256 at ref: <%s not present at %s>\n' "$REL_PATH" "${recorded:0:7}"
else
printf 'sha256 at ref: <%s not present in this clone>\n' "${recorded:0:7}"
fi
printf 'skillset HEAD: %s (%s)\n' "$head_sha" "$upstream_sha"
if [ -n "$upstream_dirty" ]; then
printf 'live working tree: %s\n' "$upstream_dirty"
fi
rc=0
if [ "$ref_exists" != 1 ]; then
printf 'CANNOT-DETERMINE: %s is not present in %s — fetch, or check against a complete clone\n' "$recorded" "$ROOT" >&2
rc=2
elif [ "$path_at_ref_exists" != 1 ]; then
printf 'MISMATCH: %s does not exist at %s in this clone — the recorded ref cannot be describing these bytes\n' "$REL_PATH" "$recorded" >&2
rc=1
elif [ "$at_ref" != "$vendored_sha" ]; then
printf 'MISMATCH: the vendored snapshot is NOT the file at the recorded ref\n' >&2
rc=1
else
printf 'OK: the vendored snapshot is exactly skillset@%s:%s\n' "${recorded:0:7}" "$REL_PATH"
fi
# Staleness is orthogonal to truthfulness: a record can correctly describe
# an old commit even after upstream has moved on, and a dirty local working
# tree in $ROOT doesn't rewrite git history either — it says nothing about
# whether the RECORDED, committed ref describes the RECORDED, committed
# bytes. Only worth reporting once we already know rc=0 (truthful) — a
# MISMATCH or CANNOT-DETERMINE is the dominant fact and a staleness note
# would only muddy it.
if [ "$rc" = 0 ] && [ "$vendored_sha" != "$upstream_sha" ]; then
# Name the ACTUAL cause. "working tree differs" is wrong when the tree is
# clean and the ref simply moved on — a message that names the wrong cause
# is the same defect class as a canary pinned to a deleted phrase.
if [ "$recorded" != "$head_sha" ] && [ "$blob_sha" = "$upstream_sha" ]; then
# Do not ASSERT which side is newer — test it. Asserting that HEAD is the
# newer side points the operator at a refresh (which costs a ~67-minute
# base rebuild) when the real remedy may be `git pull` in this clone. The
# refresh path below already uses this primitive; reuse it here.
if git -C "$ROOT" merge-base --is-ancestor "$recorded" "$head_sha" 2>/dev/null; then
printf 'NOTICE: %s has moved on to %s; the snapshot describes the older %s (stale, not untruthful — refresh to catch up)\n' \
"$ROOT" "${head_sha:0:7}" "${recorded:0:7}" >&2
elif git -C "$ROOT" merge-base --is-ancestor "$head_sha" "$recorded" 2>/dev/null; then
printf 'NOTICE: %s is BEHIND at %s; the snapshot describes the newer %s — pull this clone, do NOT refresh the snapshot\n' \
"$ROOT" "${head_sha:0:7}" "${recorded:0:7}" >&2
else
printf 'NOTICE: %s (HEAD %s) and the recorded %s have DIVERGED — neither is an ancestor of the other; reconcile the clone before refreshing\n' \
"$ROOT" "${head_sha:0:7}" "${recorded:0:7}" >&2
fi
else
printf 'NOTICE: the working tree of %s differs from the snapshot (HEAD %s)\n' \
"$ROOT" "${head_sha:0:7}" >&2
fi
fi
exit "$rc"
fi
[ -z "$upstream_dirty" ] || die "$ROOT/$REL_PATH is $upstream_dirty — commit it first, or the recorded ref would not describe these bytes"
if [ "$vendored_sha" = "$upstream_sha" ] && [ "$recorded" = "$head_sha" ]; then
printf 'already current: snapshot == skillset@%s\n' "${head_sha:0:7}"
exit 0
fi
# Refuse to silently REWIND provenance. `git checkout <tag>`, a detached HEAD,
# or an older checkout can all leave $ROOT's HEAD behind the already-recorded
# ref; without this guard a refresh there would happily rewrite both the ARG
# and the bytes backwards and report it as an ordinary update.
if [ "$recorded" != "$head_sha" ]; then
if git -C "$ROOT" cat-file -e "${recorded}^{commit}" 2>/dev/null; then
if ! git -C "$ROOT" merge-base --is-ancestor "$recorded" "$head_sha" 2>/dev/null; then
if [ "$FORCE" != 1 ]; then
die "refusing: $ROOT's HEAD ($head_sha) is not a descendant of the recorded ref ($recorded) — this looks like a rewind. Pass --force if this is intentional."
fi
printf 'WARNING: --force set; %s is not an ancestor of HEAD %s — proceeding anyway\n' "${recorded:0:7}" "${head_sha:0:7}" >&2
fi
else
if [ "$FORCE" != 1 ]; then
printf 'CANNOT-DETERMINE: %s is not present in %s (shallow clone?) — fetch full history to verify this refresh moves forward, or pass --force to proceed without that guarantee\n' "$recorded" "$ROOT" >&2
exit 2
fi
printf 'WARNING: --force set; %s could not be resolved in %s — proceeding without verifying forward motion\n' "${recorded:0:7}" "$ROOT" >&2
fi
fi
# Written FROM THE REF, not copied from the working tree, so the pair cannot
# be a lie by construction. Via a temp file so a failed write cannot leave a
# half-vendored snapshot behind.
snap_tmp=$(mktemp)
if ! git -C "$ROOT" show "HEAD:$REL_PATH" > "$snap_tmp" 2>/dev/null; then
rm -f -- "$snap_tmp"
die "cannot read HEAD:$REL_PATH from $ROOT"
fi
chmod 0644 -- "$snap_tmp"
mv -- "$snap_tmp" "$VENDORED"
[ "$(sha_of "$VENDORED")" = "$blob_sha" ] \
|| die "internal: written snapshot does not match HEAD:$REL_PATH"
# In-place, and only the exact pinned line: a broad sed on this Dockerfile
# could rewrite one of the other *_REF ARGs.
tmp=$(mktemp)
sed "s|^ARG ${ARG_NAME}=.*\$|ARG ${ARG_NAME}=${head_sha}|" "$DOCKERFILE" > "$tmp"
chmod 0644 -- "$tmp"
mv -- "$tmp" "$DOCKERFILE"
new_recorded=$(grep -oE "^ARG ${ARG_NAME}=[0-9a-f]{40}$" "$DOCKERFILE" | cut -d= -f2 || true)
[ "$new_recorded" = "$head_sha" ] || die "failed to rewrite ${ARG_NAME} in $DOCKERFILE"
printf 'snapshot: %s -> %s\n' "${vendored_sha:0:12}" "$(sha_of "$VENDORED" | cut -c1-12)"
printf 'ref: %s -> %s\n' "${recorded:0:7}" "${head_sha:0:7}"
printf '\nNOTE: %s is hashed into base_tag, so this costs a base rebuild\n' "$VENDORED"
printf 'on the next tag (~67 min). Also re-pin the phrase canary in\n'
printf 'scripts/smoke-test.sh if the section it names changed.\n'