Commit Graph

10 Commits

Author SHA1 Message Date
joakimp 2278b22ba7 fix(ssh): wire git core.sshCommand to the writable sidecar, and assert it
Lint / hadolint (push) Successful in 13s
Lint / skill-floor (push) Successful in 14s
Lint / doc-drift (push) Successful in 18s
Lint / actionlint (push) Successful in 29s
~/.ssh is commonly bind-mounted READ-ONLY from the host, so a per-host
`ControlPath ~/.ssh/cm/%r@%h:%p` — the standard CGNAT multiplexing recipe, and
correct on the host — resolves inside an unwritable dir in the container. Every
push dies `unix_listener: cannot bind to path ...: Read-only file system`,
behind git's misleading "make sure you have the correct access rights".

setup-lan-access.sh already writes the fix: ~/.ssh-local/config overrides
ControlPath into the writable ~/.ssh-local/cm BEFORE `Include ~/.ssh/config`,
so -F repairs the socket path and keeps every per-host User/Port/IdentityFile.
entrypoint-user.sh now points git at it, guarded on the sidecar existing —
setup-lan-access.sh writes none on native Linux Docker, where -F at a missing
file would break every git-over-ssh call instead of fixing one. An existing
core.sshCommand is left alone (first-wins, as for the three git settings above).

Why code and not another doc line: the remedy was already in the global
AGENTS.md, in pi-devbox-environment SKILL.md §3, in 24 MemPalace drawers from
three devices, and printed verbatim by recreate-sanity-check.sh — and an agent
that had run that script two hours earlier still hit the failure and reinvented
a /tmp/sshcm workaround. A fifth copy was not the missing piece.

Assertions, each where it can actually pass:
- smoke-test.sh: two STATIC greps (wiring line + its [ -r ] guard). `run` uses
  --entrypoint="", so asserting the runtime value there would repeat the v1.8.0
  mistake of an assertion that cannot pass, unvalidated until the next tag.
- smoke-test.sh runtime phase: a BICONDITIONAL — sidecar present => must route
  through it; absent => must be unset. The absent arm is the one CI exercises
  (native Linux runner), so "is set" would have failed CI for a correct image.
- recreate-sanity-check.sh: the runtime assertion, plus an explicit fail for the
  inverted state (set while the sidecar is missing). The permanent "default ssh
  precedence" warning keeps its severity but now states that it is structural and
  can never reach zero, and whether git is wired, unwired, or has no sidecar.

All five arms exercised against the real script before commit; that caught a
defect in the first draft, which reported "git IS wired ... unaffected" about a
state where the sidecar was gone and every git-over-ssh call failed.

Host ~/.ssh/config needs no change: the same line is right on the host and
unusable through a read-only mount, so the fix belongs in the container layer.
2026-09-22 23:08:59 +02:00
Joakim Persson f25efa074d fix(sanity): verify the ControlPath ssh RESOLVES, not the dir the image creates
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 13s
Lint / doc-drift (push) Successful in 12s
Lint / actionlint (push) Successful in 24s
recreate-sanity-check.sh asserted `/tmp/sshcm exists with mode 700` and printed
a green tick while every ssh in the container died rc=255. It was right about
what it checked: the breakage was in the directory a CONFIG named, not the one
the image creates, and the old check could not see the disagreement between them.

Found on the v1.9.2 first boot on emb-7kj4vr4g. A durable ~/.pi/ssh/config,
hand-written into the ~/.pi named volume by the previous session so it would
survive the recreate, declared `ControlPath /tmp/ssh-cm/%C` — with a hyphen.
Nothing here creates that path; the canonical spelling is /tmp/sshcm, identical
in Dockerfile.base, entrypoint-user.sh, recreate-sanity-check.sh and
smoke-test.sh. ControlMaster auto with an unusable ControlPath does not degrade
to an unmultiplexed connection — it fails hard:

    unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory

rc=255, remote command never runs.

Now: resolve the ControlPath ssh itself would use via `ssh -G` and require its
parent to exist and be writable. -G applies real config precedence (first-value-
wins, the system drop-in, Include, an -F override), so it answers "which rule
captured this host" instead of re-implementing the guess, and it never opens a
connection — 0.116s for 48 hosts.

This also puts a check under a caveat documented in prose in Dockerfile.base
("SSH client defaults") and verified nowhere: a per-host ControlPath under a
read-only bind-mounted ~/.ssh gives the identical failure, "Read-only file
system". On the machine this was built on that is 16 of 50 hosts — freeipa-1..6,
gitea.egl.lan, runner-1..3, tor-ms22 — never reported by anything before.

Severity split is deliberate. Default ssh precedence legitimately lands in the
read-only ~/.ssh for any host whose own config pins it there, and the supported
workaround (ssh -F ~/.ssh-local/config, generated every start by
setup-lan-access.sh) already exists, so that is a warn. Failing it would paint
the script red on every run of every device, and a check that fires benignly
every time is one you learn to ignore — the same reasoning that keeps
lint-shell.sh at -S error. The sidecar route is prescribed, so there it is a
hard fail. Host lists cap at six names plus a count: unreadable output is
ignored output.

Teeth proven both directions, each sabotage confirmed by diff BEFORE the result
was believed: ControlPath -> nonexistent dir => rc=1; -> existing-but-read-only
dir => rc=1; restored => rc=0 with the sidecar byte-identical.

No config was added to fix the original problem — the fix was to DELETE
~/.pi/ssh/config, a third hand-maintained copy of what setup-lan-access.sh
already generates from version control on every container start, with fewer
features and one typo. The host-owned ~/.ssh/config is left alone on purpose:
those ~/.ssh/cm paths are correct on the host, where ~/.ssh is writable.

Also lint-shell.sh: the SC2088 count in the severity-choice rationale said 19
and is now 20, with the reproduce command recorded. That command needs a `$ `
prefix — a comment whose first word is "shellcheck" is parsed as a directive,
and the malformed one tripped SC1072/SC1073 at severity error. The gate caught
it on the very commit that introduced it.
2026-09-15 22:59:43 +02:00
joakimp adcf56f829 release: audited bumps (pi 0.85.1, mempalace 3.9.0, atelier v0.10.1) + two guards
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 17s
Version audit for the next release. pi 0.84.4 -> 0.85.1, deliberately skipping
0.85.0 (it published internal experimental code and broke SDK imports,
upstream #9132). mempalace 3.8.0 -> 3.9.0. pi-atelier v0.10.0 -> v0.10.1.
PI_STUDIO_VERSION relabelled none -> v0.9.60-rc.0 so the floating main ref's
RC status is visible at docker-inspect time instead of discovered later.
PI_FORK_REF stays floating and adopts e69725c.

The pi bump was verified by running it under a pty in five combinations rather
than by reading the changelog, because this repo has already shipped a version
pair no changelog flagged (atelier < 0.7.1 hangs pi >= 0.84). CPU delta
0.00-0.01s over 5s against a ~5s sustained-CPU hang signature, two-sided via
the atelier sidebar painting identically to the 0.84.4 control.

NODE_VERSION stays 22 on purpose: node 24 is technically safe (pi's five
prebuilt addons are all NAPI, nothing declares a ceiling, agent-browser's
engines.node >=24 is vestigial for the shipped aarch64 ELF), but this release
already moves two minors and bakes an RC, and a node major would leave four
suspects if the image misbehaves. Own release, smoke suite as the gate.

Also corrects a stale claim at the mempalace ARG: synlig serves 3.8.0
server-side, not 3.7.1 (measured over ssh 2026-09-06).

agent-browser volume shadowing: the image has shipped 0.35.2, but every
session on mbp-m1-2020 ran 0.27.0 from a 2026-07-17 hand-install in
~/.pi/npm-global (a VOLUME, at PATH position 2 vs /usr/bin at 8). Third
package hit by this hazard after pi and pi-atelier, so the guard is now
generalised: entrypoint-user.sh retires the copy by moving it aside
(reversible, only when the image ships its own), recreate-sanity-check.sh
asserts resolution under /usr where the volume is real, smoke-test.sh carries
the build-time half and says in the source why it is weak. The real damage was
the stale BUNDLED SKILL (3 skillsets/17.6 KB vs 8/31.5 KB, ten subcommands
undocumented to the agent) - a stale tool errors, a stale skill quietly
teaches wrong commands.

pi-fork capability floor (extensions: []): forks were measured across four
dispatches ignoring their brief, answering in the user's voice, fabricating
self-referential measurements, and once filing a diary entry as agent_name=pi.
Cause is upstream by design - the child gets getHeader()+getBranch(), the
whole active session branch, with the brief as the final user message. Not a
model-capability problem: the same model as the fast profile obeyed the
identical brief perfectly with a fresh session and no inherited context.
extensions: [] runs children with --no-extensions, so the mempalace bridge is
absent and palace writes are impossible by construction (verified by asking a
child to enumerate its tools: read, bash, edit, write). Removes palace writes,
not filesystem writes.
2026-09-06 20:40:02 +02:00
joakimp aac4a1c323 release: v1.8.9 — the version flag that blamed the wrong component
Lint / hadolint (push) Successful in 15s
Lint / actionlint (push) Successful in 18s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 9s
Publish Docker Image / build-base (push) Successful in 41m49s
Publish Docker Image / smoke (push) Successful in 4m51s
Publish Docker Image / smoke-studio (push) Successful in 4m59s
Publish Docker Image / build-variant-studio (push) Successful in 16m58s
Publish Docker Image / build-variant (push) Successful in 28m35s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 17s
Two versions, two flags. `--expected-version` has only ever asserted
`pi --version`, but AGENTS.md step 4 spelled it `X.Y.Z` inside a checklist where
every other X.Y.Z is the pi-devbox tag. Run as documented for v1.8.8 the final
runtime gate of the release printed

    ✗ pi version mismatch: expected 1.8.8, got 0.84.3

and exited 1 — a red accusing the image of being the wrong version. Not one
reader's slip: the v1.8.8 release-readiness handoff from pi@emb-7kj4vr4g
propagated the same wrong spelling twice while correctly calling step 4 "not
ceremonial", so two independent readers converged on it. README.md had it right
all along, which means the two documents disagreed.

- new --expected-image-version asserts the pi-devbox release tag, read from
  release_tag in /etc/pi-devbox/build-manifest.json (no checkout, no network);
  leading `v` optional on either side
- both flags detect being handed the other one's value, and the test is exact
  rather than heuristic: the value is compared against the other quantity the
  image itself reports, so it can only fire on a real mix-up
- neither flag is required now. With none, live `pi --version` is asserted
  against the manifest's pi_version — not a tautology, since a stale pi in the
  ~/.pi/npm-global volume can shadow the baked one, exactly as a stale
  npm:pi-atelier can in packages[]
- the header note replaced was stale and load-bearing: it claimed pi is resolved
  from 'latest' and cannot be self-derived, while Dockerfile.variant pins
  ARG PI_VERSION=0.84.3 and docker-publish.yml reads that ARG as its source of
  truth. The same withdrawn claim also sat in cli_utils' pi-devbox-sanity --help
- argument parsing: a missing value, or a value that is another flag, is a usage
  error instead of silently consuming the next argument; --help works

All fourteen flag combinations exercised by execution, including the two
manifest-absent branches and the shadowing branch a healthy container cannot
reach — mutation-tested with a doctored manifest so each failure branch was
observed firing rather than assumed present.

CHANGELOG also names what no commit here causes: mempalace-toolkit main moved
e70bef2 -> 5b8d78f, so this tag ships the auto-delivered logstream mailbox
because base_tag folds the resolved toolkit SHA. It would have landed either
way; going unnamed is the 553d865 shape that already caused one cross-host
misattribution. Component audit found nothing else to bump — pi, mempalace,
pi-atelier all equal their upstream latest, and every other floating ref
resolves to the commit already baked.
2026-08-26 18:47:03 +02:00
joakimp 43cd6e22f2 v1.7.0: bundle pi-atelier at a pinned tag; pin pi to an audited 0.84.1
Publish Docker Image / resolve-versions (push) Successful in 9s
Lint / actionlint (push) Successful in 15s
Lint / hadolint (push) Successful in 13s
Publish Docker Image / base-decide (push) Successful in 8s
Publish Docker Image / build-base (push) Successful in 41m22s
Publish Docker Image / smoke-studio (push) Successful in 5m16s
Publish Docker Image / smoke (push) Successful in 7m31s
Publish Docker Image / build-variant-studio (push) Successful in 18m31s
Publish Docker Image / build-variant (push) Successful in 27m18s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / update-description (push) Successful in 12s
Two changes that belong together, because the first is what makes the second
dangerous to get wrong.

pi-atelier (TUI sidebar + status rail) is now vendored to /opt/pi-atelier at
PI_ATELIER_REF=v0.8.0 and registered by entrypoint-user.sh — the pi-fork /
pi-observational-memory / pi-studio pattern, deliberately NOT
`pi install npm:pi-atelier`, which writes into ~/.pi/npm-global on the config
volume where it shadows the image and pins nothing. Unlike its siblings it gets
no `npm install`: atelier declares zero runtime deps (peerDeps only, satisfied by
the baked pi) and has no build step, so pi loads its TypeScript straight from the
checkout via package.json `pi.extensions`.

pi is no longer resolved to npm `latest` at build time. The pin lives in
Dockerfile.variant and CI reads it from there, so a local `docker build` and a CI
release ship the same versions by construction. The pin is a CHECKPOINT, NOT A
FREEZE: bumping stays a one-line change; what stops is *unreviewed* adoption of
whatever shipped that morning, in the same build that then gets tagged and
published. CI fails when a pin is not concrete or not actually published on npm,
and warns — never adopts — when npm latest moves ahead, naming what to re-check.

Why this pairing needed care: pi-atelier 0.6.0/0.7.0 wrap pi's PRIVATE TUI
renderer, and under pi 0.84 that wrapper recurses — pi hangs at startup burning
CPU with no error. Upstream fixed the recursion in 0.7.1 and restored the
non-overlapping split in 0.7.2; 0.8.0 is additive on top. atelier's own
peerDependencies still say >=0.80.7, which does not express that floor, so
nothing in npm metadata could have warned us. The floor is therefore encoded as
an executable rule — pi >= 0.84 => pi-atelier >= 0.7.1 — asserted in both
smoke-test.sh (build time) and recreate-sanity-check.sh (after a real recreate),
verified against a 4x4 version matrix.

Existing volumes needed migration, not just vendoring: a hand-installed
`npm:pi-atelier` entry is counted as already-registered by the entrypoint guard,
so every existing volume would have kept its unpinned npm copy — and a 0.6.x copy
next to pi 0.84 is exactly the startup hang. The entrypoint now drops that one
exact string (settings.json.bak.atelier.<ts> backup, distinct prefix so it cannot
clobber the template merge's backup in the same second) and lets the pinned /opt
copy register. Tested against a real settings.json: only that entry removed,
other packages and all keys intact, idempotent, and unparseable JSON leaves the
file untouched. DEVBOX_ATELIER=0 opts out entirely — in the entrypoint rather
than via `pi uninstall`, because this component's failure mode is "pi will not
start", which cannot be repaired from inside pi.

0.84.1 was audited for this release, not merely adopted: theme/TUI changes are
additive, the session format is unchanged (CURRENT_SESSION_VERSION = 3 in both
0.83.0 and 0.84.1 with an identical migrateV1ToV2/migrateV2ToV3 ladder, so
existing transcripts are neither migrated nor at risk and pi-session-repair stays
valid), and the Node engine floor is unmoved at >=22.19.0. CI resolves the
atelier tag to its PEELED commit SHA — atelier uses annotated tags, so the
unpeeled ref is a tag object, not a commit; pi-studio's lightweight tags never
exposed that distinction.

Also: docs for overriding the read-only ~/.ssh/config from the container —
container-only keys in ~/.ssh-local, hardened authorized_keys, the fact that
`from=` must allow the HOST's addresses because container egress is NAT'd through
it, and the macOS-only-keyword trap (`UseKeychain` is fatal to Linux OpenSSH and
takes out dssh/pi --ssh while the host keeps working). Corrects two claims in
"Naming LAN peers": ssh-lan.conf is not ProxyJump-only, and first-time creation
does need one restart because the Include is emitted only when the file already
exists at start.
2026-08-07 21:34:41 +02:00
joakimp 8248688d58 fix(entrypoint,tests): register pi-fork — guard matched its own config block
Lint / actionlint (push) Successful in 32s
Lint / hadolint (push) Successful in 1m25s
The `pi install /opt/<pkg>` loop in entrypoint-user.sh guarded on a
whole-file substring grep of ~/.pi/agent/settings.json. settings.example.json
ships a top-level "pi-fork" CONFIG block (fork effort profiles, pi-toolkit
adb6907, 2026-06-17), so `grep -q pi-fork settings.json` matched the config
key itself and `pi install /opt/pi-fork` never ran — on fresh or preserved
volumes. The `fork` tool has therefore been absent since v1.0.0.

The non-destructive template merge runs earlier in the same startup than the
install loop, so the mechanism that delivers new template keys to an old
volume is what plants the string that defeats the guard. pi-observational-
memory and pi-studio escaped only by luck: the template key is
"observational-memory" (no pi- prefix) and there is no studio block.

Guard now inspects the `packages` array via jq, with a grep fallback on the
stored `.../opt/<name>"` path form, which a config key can never produce.
Existing volumes self-heal on the next container start.

Both test suites asserted the bug as green — smoke-test.sh:244 and
recreate-sanity-check.sh:204 used the same whole-file grep, so "pi-fork
registered (fork tool)" passed on every build and recreate while the tool was
missing. Both now assert against packages[] with the entrypoint's predicate,
labels say packages[], and the smoke readiness wait loop uses the array check
plus `docker exec -u developer` + $HOME instead of a hard-coded
/home/developer path.

Evidence: zero `fork` tool calls across all 19 sessions on this volume; the
v1.6.3 session that tuned pi-fork.deep to opus-5 was configuring a tool that
never loaded.
2026-07-29 19:21:49 +02:00
Joakim Persson 41c2c2b716 feat(entrypoint): non-destructively merge new template keys into settings.json
The settings.json bootstrap only fires when the file is ABSENT, so a
settings.json on a preserved named volume never picks up config added in a
later image (e.g. the observational-memory / pi-fork blocks, a newly-enabled
model). Users had to hand-merge after every upgrade.

On start, when settings.json already exists, deep-merge the template into it
with 'jq -s ".[0] * .[1]"' (template first, live second) so the user's values
always win and only MISSING keys are filled from the template. Arrays are
leaves (a model the user removed is not re-added). Rewrites only when the
merge changes something, backs up the original first, and skips safely (no
clobber) if either file is invalid JSON. Opt out with PI_SETTINGS_MERGE=0.

Add a recreate-sanity-check assertion that settings.json carries the
observational-memory + pi-fork blocks after recreate.
2026-06-17 20:49:41 +02:00
Joakim Persson 5c08bfc8a8 fix(shell): don't export DEVBOX_HIST_SET so nested shells flush history
The history-flush guard was exported, so it leaked into child processes.
Any nested shell -- crucially each tmux pane (which inherits the tmux
server's env) -- then saw the guard already set and skipped installing
'history -a' in PROMPT_COMMAND. Those shells only persisted history on a
clean exit, so abrupt termination (docker stop, tmux kill-server, SIGKILL)
silently lost their in-memory history. zoxide was less affected (its hook
is installed unguarded and writes immediately).

Make the guard shell-local (drop 'export') so every new interactive shell
re-installs its own per-prompt flush. Add a recreate-sanity-check assertion
that a nested login shell still wires up 'history -a'.

Storage was never the issue: ~/.cache/bash (devbox-shell-history) and
~/.local/share/zoxide (devbox-zoxide) are both persistent named volumes.
2026-06-17 17:22:30 +02:00
Joakim Persson 1371584634 sanity-check: verify global AGENTS.md symlink after recreate
pi-toolkit now symlinks pi-global-AGENTS.md -> ~/.pi/agent/AGENTS.md (pi's
global-instructions file, loaded at every start; directs the agent to read
the pi-extensions skill at session start). Add a recreate-sanity-check
assertion alongside the keybindings symlink check so a future image build
that bakes the new pi-toolkit verifies the wiring landed.
2026-06-17 16:58:18 +02:00
pi 4ed6764323 Add runtime post-recreate sanity check (peer of smoke-test.sh)
scripts/recreate-sanity-check.sh verifies what is actually live in a
recreated container — persisted volumes, pi runtime wiring (keybindings,
extensions, mempalace.ts bridge, settings.json, fork/obsmem/studio
registrations), /tmp/sshcm, skel defaults, /opt toolkits. smoke-test.sh
runs at build time with --entrypoint="" and cannot see any of this.

Variant (studio/plain) auto-detected via /opt/pi-studio. pi version is
asserted only with --expected-version (built from 'latest', no Dockerfile
pin to self-derive). Maintainer tooling, not baked into the image.

Documented in README and CHANGELOG.
2026-06-15 22:04:02 +02:00