Files
pi-devbox/AGENTS.md
T
Joakim Persson 852f900b53
Lint / doc-drift (push) Successful in 6s
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 9s
Lint / actionlint (push) Successful in 18m11s
test(smoke): make CI prove the natives still work — the runbook check didn't
The post-boot check v1.9.1 left for the next machine was

  node -e 'require("esbuild").transformSync("const x:number=1",{loader:"ts"})'

with "if this fails, the prune removed something needed on arm64 -> revert to
v1.8.14 and fix" attached. Run from /workspace on the freshly recreated v1.9.1
image it fails with MODULE_NOT_FOUND — on a perfectly healthy image. require()
resolves by walking up from the CURRENT DIRECTORY, esbuild is nested inside the
two pi trees, and a global install is not on node's require path (NODE_PATH is
unset). The version that "verified the prune" last night only passed because the
shell happened to sit inside the tree. So the check was CWD-dependent, its red
meant nothing, and the remedy it prescribed was reverting a good release.

Path-qualified, both sites compile TS today. Rather than leave a fragile command
in a note for a human to paste, CI now owns it:

  - esbuild must transformSync at EVERY install site found in the image
  - @mariozechner/clipboard must load with its native binding attached at every
    site — for the clipboard prune that IS the proof, since napi-rs resolves the
    platform package at require() time

Sites are discovered with find, so the studio variant's third site is covered
without naming it.

Verified as a four-way matrix, not just a green run: both assertions pass
against the real image FROM /workspace (the cwd that broke the old command), and
both fail (exit 1, "The package \"@esbuild/linux-arm64\" could not be found, and
is needed by esbuild") against copies of the same packages with the host
platform binary removed. An assertion never observed failing is not evidence.

Also documents both failure shapes in AGENTS.md next to the size gate: this one,
and the mode-700 /root one where `test ! -d` passes on a permission error.
2026-09-11 11:06:18 +02:00

375 lines
22 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# AGENTS.md — pi-devbox
Self-contained Docker image for the **pi coding-agent**. Decoupled from
opencode-devbox at v1.0.0 (2026-06-09); previously pi-devbox was a thin
re-brand of opencode-devbox's `pi-only` variant.
## Repository layout
- `Dockerfile.base` — multi-arch base layer with system packages,
GitHub-binary tools (fzf, eza, zoxide, neovim, bat, gosu, gitleaks,
git-lfs, uv, gitea-mcp, tealdeer), AWS CLI v2, mempalace + toolkit,
Node.js, Python toolchain, locales, ssh ControlMaster defaults, and
`/etc/tmux.conf` with 0-indexed sessions.
- `Dockerfile.variant` — `FROM base-<hash>`, adds pi + companions
(`pi-toolkit`, `pi-extensions`, `pi-fork`, `pi-observational-memory`)
and, when `INSTALL_STUDIO=true`, vendors `pi-studio` to `/opt/pi-studio`
(`-studio` variant). Also appends the pi-devbox managed block from
`pi-global-AGENTS.append.md` onto pi-toolkit's `pi-global-AGENTS.md` (the
single global instruction slot pi loads) so containers proactively load the
baked `pi-devbox-environment` skill. Idempotent via a marker grep. After the
pinned clones it also refreshes the vendored `pi-extensions` fallback skill
by copying `/opt/pi-extensions/skill/` over the committed `rootfs/` snapshot
(Option 1 over Option 2 — see `skills/VENDORED.md`).
- `entrypoint.sh` — UID/GID alignment as root, then drops to `developer`.
- `entrypoint-user.sh` — per-container start: prints the `pi-devbox-version`
banner first (which build/commit is running, from the manifest below),
then SSH ControlMaster socket dir, LAN-access setup, MemPalace init,
pi-toolkit + pi-extensions deploy, mempalace-bridge symlink, fork/recall +
pi-studio pi-install, optional `studio-expose` bridge (when
`STUDIO_EXPOSE=1`), image-baked skills symlink-in, skillset deploy.
- `rootfs/` — files baked into the image (bash aliases, inputrc,
setup-lan-access.sh, `studio-expose` helper, `pi-devbox-version` — wraps
`/etc/pi-devbox/build-manifest.json` into a human-readable summary + live
drift check, see README “Build provenance”). Also
`usr/local/share/pi-devbox/skills/<name>/SKILL.md` — image-baked agent
skills (the repo-authored `pi-devbox-environment`, plus vendored fallback
copies of `pi-extensions` and `mempalace` — see `skills/VENDORED.md`)
symlinked into `~/.agents/skills/` by the entrypoint, available with or
without a mounted skillset — plus
`usr/local/share/pi-devbox/pi-global-AGENTS.append.md` (the global-AGENTS
pointer concatenated in `Dockerfile.variant`).
- `scripts/smoke-test.sh` — sanity checks run by CI before pushing to Hub.
- `.gitea/workflows/docker-publish.yml` — two-phase CI (base-decide →
build-base → smoke → build-variant → promote-base-latest →
update-description). The `-studio` variant adds independent
`smoke-studio` + `build-variant-studio` jobs that gate only the
`-studio` tags (never the core `:latest` release).
## Versioning scheme
- Tags follow semver. **v1.0.0** is the first decoupled release; future
minor bumps add variants (`-studio`, `-studio-tex`) or significant base
additions (e.g. v1.2.0 image-baked agent skills); patch bumps follow
pi npm version updates and small fixes.
- Docker Hub tags: `joakimp/pi-devbox:vX.Y.Z` + `joakimp/pi-devbox:latest`
+ (since v1.1.0) `joakimp/pi-devbox:vX.Y.Z-studio` +
`joakimp/pi-devbox:latest-studio`.
Internal tags: `joakimp/pi-devbox:base-<hash>` (content-addressed) +
`joakimp/pi-devbox:base-latest` (alias of most recent base).
## Release-day checklist
1. Confirm `pi --version` resolves from npm to the expected version
(`curl -sf 'https://registry.npmjs.org/@earendil-works%2Fpi-coding-agent/latest' | jq -r .version`).
Check release notes at https://github.com/earendil-works/pi/releases for
the upstream changelog to include in `CHANGELOG.md`.
2. **Refresh the vendored mempalace skill snapshot if the skillset moved:**
`scripts/vendor-mempalace-skill.sh --check` (reads a real skillset clone,
writes nothing). Three exit codes, not two — a stale-but-truthful record is
**not** a release blocker, so don't treat any non-zero exit as "must
refresh" without reading which one it was:
- **0** — the record is truthful. This includes stale-but-truthful
(upstream has moved past the recorded ref, or the local clone has
uncommitted changes) — a `NOTICE` is printed, but nothing is lying.
**Skipping the refresh in this case is the legitimate, sanctioned
outcome** — every enrolled host reads its own live skillset clone, so
the baked copy is only a no-mount fallback. What is not legitimate is
skipping it *silently*: the drift is visible here, in
`pi-devbox-version`, and in the manifest, so decide rather than forget.
- **1** — a confirmed problem: the vendored bytes provably do NOT match
the file at the recorded ref (a lying record), or the recorded ref
doesn't even resolve to that path in this clone. Refresh.
- **2** — cannot determine (the recorded ref itself isn't resolvable in
this clone — commonly a shallow checkout missing history). Fetch full
history and re-check before deciding; don't refresh blind.
Refresh with `scripts/vendor-mempalace-skill.sh`, which rewrites the file
**and** the ARG together so they cannot drift apart, and refuses (exit 1)
rather than silently rewinding provenance if the skillset clone's HEAD is
behind the already-recorded ref (detached HEAD, older checkout) — pass
`--force` only if that rewind is genuinely intended.
Two consequences to accept deliberately on an actual refresh: the snapshot
is hashed into `base_tag`, so it costs a base rebuild (~67 min); and if the
section the phrase canary names has changed, re-pin it in
`scripts/smoke-test.sh`.
3. **Update the docs this release makes stale — BEFORE you tag.** Rename
`CHANGELOG.md`'s `## Unreleased` to `## vX.Y.Z — YYYY-MM-DD` (em dash, as
every prior release heading uses), then run the gate:
```bash
bash scripts/check-doc-drift.sh # 0 in sync / 1 drift / 2 cannot run
```
It compares README.md's version-pin table against the ARGs it names, and
DOCKER_HUB.md's Node claim against `ARG NODE_VERSION`, plus Hub's
25 000-char limit, unsubstituted `{{PLACEHOLDERS}}`, and stale `Unreleased`
pointers in user-facing docs.
**Why before and not after:** `docker-publish.yml` runs `actions/checkout@v4`
with no `ref:`, so every job reads `github.ref` — the **tag**. A doc fix
pushed to `main` after tagging does not reach the release, and for
`DOCKER_HUB.md` it does not reach the published Hub page either, because
`update-description` POSTs that file as Docker Hub's `full_description` from
the tag's tree. Getting it in afterwards means re-pointing the tag, which is
its own hazard (v1.8.14 went `601fc98` → `361babd` and broke deploy
verification until `git fetch --tags --force`).
The gate is deliberately narrow — it only checks claims verifiable from files
in this repo. Still eyeball, because these are NOT gated:
- counts and sizes (`~1.1 GB`, "N `mempalace_*` tools", "7 extensions") —
they need a running image; assert them in `scripts/smoke-test.sh` instead
- feature prose that quietly became false, e.g. a "Planned for an upcoming
release" section describing something that already shipped
- `Dockerfile.base`'s `# BASE_REBUILD_DATE:` marker. Ungated on purpose:
`base_tag` hashes Dockerfile.base's content, comments included, so
demanding it be current would force a ~60 min base rebuild on a release
that touched no base files. **Fix it when the base is already rebuilding —
then it is free.**
Measured cost of skipping this, 2026-09-10 (v1.9.0): five stale claims, one
of them published. README's pin table was wrong on all three rows, and
DOCKER_HUB.md — untouched for eight releases — still said Node v22 while the
image shipped Node 24.
4. Verify `docker compose up` works locally with the current `latest` image
if you're upgrading users from a previous version. Then run the
**post-recreate sanity check** inside the running container to confirm
persisted volumes survived and the pi runtime wiring re-deployed (not just
that the container booted):
`docker compose exec devbox bash scripts/recreate-sanity-check.sh --expected-image-version X.Y.Z`
(or just `pi-devbox-sanity --expected-image-version X.Y.Z` if
`cli_utils/bin` is on PATH). This is the runtime peer of the build-time
`smoke-test.sh` gate.
**`X.Y.Z` here is the pi-devbox release tag** you are shipping (e.g.
`1.8.9`), which is what the rest of this checklist means by `vX.Y.Z`.
`--expected-image-version` is the flag that asserts it. There is also an
`--expected-version`, and it means something else — the **pi coding agent**
version (e.g. `0.84.3`, the `ARG PI_VERSION` pin). Handing the release tag
to that one used to report *"pi version mismatch: expected 1.8.8, got
0.84.3"*, i.e. a red on the final gate of the release accusing the wrong
component; it now tells you to use `--expected-image-version` instead, and
the reverse mix-up is caught too. Both flags are optional — with neither,
the live pi version is asserted against the version recorded in the image's
own build manifest (which catches a stale `pi` in the `~/.pi/npm-global`
volume shadowing the baked one) and the image tag is reported
informationally.
5. Push tag: `git tag vX.Y.Z && git push origin vX.Y.Z`.
6. Watch CI: smoke job builds amd64 only and asserts size + extensions +
pi version + new-base-tooling presence. Variant build is multi-arch
(amd64 + arm64) only after smoke passes. A tag push fires **only**
`docker-publish.yml` — `lint.yml` is scoped to `branches: ['**']`, which
excludes tag refs on purpose (the tagged tree was already linted when the
commit hit `main`, and a fast lint run sorting above the slow publish run
made releases look finished before anything shipped). Verified on v1.8.4:
`refs/tags/v1.8.4` produced run 571 (publish) and nothing else. Still filter
discovery on `head_sha` **and** the workflow `path` — see *Gitea API access*
below — because that guard costs nothing and a future workflow added on `v*`
would silently reintroduce the ambiguity.
7. Verify the Hub tags appear (latest + vX.Y.Z, the `-studio` pair, plus
base-latest if the base was rebuilt this run).
8. **Revoke any short-lived Gitea PAT** used during the release at
`gitea.jordbo.se/user/settings/applications`. N/A if you used the
`GITEA_ACCESS_TOKEN` env var instead (see *Gitea API access* below) —
its lifecycle is managed host-side, nothing to revoke.
## Verifying this repo's reality from inside a container
Most work on this repo happens **inside** a pi-devbox container, inspecting a
host or a peer over SSH. That setup manufactures convincing false negatives, so
when you are about to report that something is **absent, unreachable, or not
running**, suspect your own command first. Recurring instances:
- **`docker` is not on the host's non-interactive SSH `PATH`.** `ssh mac 'docker
ps'` says *command not found* on a host that plainly runs Docker; use
`/usr/local/bin/docker` (or `command -v docker` first). Every step in the
*Release-day checklist* that inspects a running container hits this.
- **Don't `| head -N` a search whose answer you don't already know.** The host's
`~/.ssh/config` is ~500 lines; a `head -20` "proved" a peer absent that was
defined at line 454.
- **The deployment compose file is not this repo's.** `docker-compose.yml` here
is a template pinning `:latest`; a real host runs its own per-machine file
(find it with `docker inspect <container> --format '{{ index .Config.Labels
"com.docker.compose.project.config_files" }}'`). Recreating from the repo copy
can silently move a host off `:latest-studio` onto `:latest`.
- **A live SSH ControlMaster hides remote auth changes** — after editing a
peer's `authorized_keys`, prove access with `-o ControlPath=none -o
ControlMaster=no`, or the breakage surfaces in a later session instead.
Depth and further mechanisms: the repo-authored `pi-devbox-environment` skill
(`rootfs/usr/local/share/pi-devbox/skills/pi-devbox-environment/SKILL.md`) §2
and §3 — that file is the one an agent actually loads mid-session, whereas this
`AGENTS.md` is only auto-read when the cwd *is* this repo.
## Gitea API access (env token)
`GITEA_ACCESS_TOKEN` + `GITEA_HOST` are passed into the container from the
host `.env` via `docker-compose.yml` (`${GITEA_ACCESS_TOKEN:-}` /
`${GITEA_HOST:-}`), primarily to enable the `gitea-mcp` server. They are
**not** baked into the image. When configured, they are also available for
**any** direct Gitea API interaction from inside the container — inspecting
CI runs, checking published tags, listing commits — e.g.
`curl -H "Authorization: token $GITEA_ACCESS_TOKEN" "$GITEA_HOST/api/v1/repos/joakimp/pi-devbox/actions/runs?limit=20"`.
Prefer this over a short-lived PAT file when the env token is present (the
`ci-release-watcher` skill auto-detects it). Public-repo GET listings work
unauthenticated too, so the token matters mainly for private repos or
rate-limit headroom; its lifecycle is host-managed, so there is nothing to
revoke after use. Never echo the token value (including into logs).
**Gotcha — a tag push fires EVERY workflow whose triggers match the tag ref.**
`lint.yml` uses a bare `push:` trigger, so a release tag yields *both* a lint run
and the publish run. The listing is newest-first and lint sorts **above** the
publish run, so "take the first run whose `path` contains `refs/tags/<tag>`"
picks the wrong one **reliably, not occasionally**. Real listing for v1.6.4:
```
id=531 #104 lint.yml@refs/tags/v1.6.4 <- wrong; sorts first
id=530 #103 docker-publish.yml@refs/tags/v1.6.4 <- the release build
id=529 #102 lint.yml@refs/heads/main <- same commit, linted on push
```
Lint goes green in minutes while the image is still building, so watching it
makes a release look finished when nothing has been published yet.
**Gotcha — the jobs endpoint takes the internal `id`, NOT the `run_number` the
UI shows as `#104`.** The two diverge widely, and `GET
.../actions/runs/<run_number>/jobs` does **not** error — it silently returns a
*different* run's jobs. Always read `id` from the run listing:
```bash
# Which runs did this tag/commit trigger? Filter on head_sha; never trust
# ordering or run numbering. limit=20, not 5 — with two runs per push the
# publish run falls off a 5-item window fast.
curl -sS -H "Authorization: token $GITEA_ACCESS_TOKEN" \
"$GITEA_HOST/api/v1/repos/joakimp/pi-devbox/actions/runs?limit=20" \
| jq --arg sha "$(git rev-list -n1 vX.Y.Z)" \
'.workflow_runs[] | select(.head_sha==$sha) | {id, run_number, path, status, conclusion}'
# pick the id whose .path starts with docker-publish.yml, then:
curl -sS -H "Authorization: token $GITEA_ACCESS_TOKEN" \
"$GITEA_HOST/api/v1/repos/joakimp/pi-devbox/actions/runs/<id>/jobs" \
| jq '.jobs[] | {name, status, conclusion}'
```
**Watcher config for this repo** (`ci-release-watcher` skill, hub-only shape —
pi-devbox has no downstream host to deploy to):
- `EXPECT_WORKFLOW=docker-publish.yml` — the skill's `preflight_run()` aborts at
startup if the run id belongs to lint instead.
- `EXPECTED_FRESH_TAGS='vX.Y.Z latest vX.Y.Z-studio latest-studio'`
- `EXPECTED_EXISTS_TAGS='base-latest'` — existence only: it is content-addressed
and legitimately keeps its old timestamp when the base is a cache hit.
- `CRITICAL_JOBS='build-variant build-variant-studio'` — job names are matched
**exactly** (`critical.issubset(succeeded)`), so the studio variant must be
listed explicitly; the skill's default omits it. Leave `promote-base-latest`
out: it legitimately skips on a base cache hit, which would misclassify a good
run. `update-description` is the cosmetic post-publish job.
## Cache-hit footgun (must-know)
`PI_VERSION` defaults to `latest` in `Dockerfile.variant` but **CI must
resolve it to a concrete version string** before passing as a build-arg.
Otherwise the build-arg string is byte-identical across releases →
identical layer hash → registry buildcache silently reuses the old
layer. `resolve-versions` job in the workflow handles this.
Discovered in pi-devbox 2026-05-23 (every release v0.74.0..v0.75.5
shipped the same image bytes); preventatively fixed for `PI_VERSION` +
`PI_FORK_REF` + `PI_OBSMEM_REF`.
## Smoke-test gate
`scripts/smoke-test.sh` runs amd64-only against a freshly-built variant
image. Verifies binaries, repo clones, runtime deployment (waits for
keybindings + mempalace bridge + ≥4 extensions before sampling — fixes
the parallel-build-load race documented in opencode-devbox c6f9d11
2026-06-08), build-time leftovers (see below), and image size threshold
(3800 MB in `SIZE_THRESHOLD_MB`; revisit after a few releases as actuals
settle — this doc said 3500 until 2026-09-11, after the bar had already
moved twice).
If smoke fails on size threshold but build is otherwise fine: bump
`SIZE_THRESHOLD_MB` in scripts/smoke-test.sh in a follow-up commit and
re-run. The threshold exists to catch *runaway* growth (an accidental
texlive bake-in, a forgotten chrome dependency), not to block ordinary
upstream bumps.
**The size gate is not a substitute for naming the residue.** It carries
~225 MB of deliberate margin, so v1.9.1 shipped +131 MB of pure build
residue — 110 MB of it npm's own download cache under `/root/.npm`, the
rest foreign platform packages — and stayed green. Four named assertions
now cover that ground: no foreign npm-11 platform packages beyond the
host arch (`@esbuild/*`, `@mariozechner/clipboard-*`), no `/root/.npm` in
the image, and — because the prune's real risk is *removing something
needed*, not size — esbuild must compile TS and clipboard must load its
native binding at **every** install site.
Two failure shapes to copy from those, both of which bit here:
- `test ! -d /root/.npm` on mode-700 `/root` passes for a **permission**
error, so the cache assertion refuses to run as non-root. Watch for
this in any assertion about a path you may not be allowed to read.
- `node -e 'require("esbuild")'` resolves by walking up from the CURRENT
DIRECTORY, so it fails with `MODULE_NOT_FOUND` from `/workspace` on a
perfectly healthy image (esbuild is nested inside the pi trees;
`NODE_PATH` is unset). Always path-qualify: `require("<abs>/esbuild")`.
A runbook shipped the bare form with "if this fails, revert the
release" attached, and it duly went red for the wrong reason.
## Build pipeline notes
- **Two-phase**: base + variant. Base is rebuilt only when
`Dockerfile.base`, `rootfs/`, or `entrypoint*.sh` change (CI computes
a content hash and probes Hub for an existing `base-<hash>` tag).
- **`base-latest` alias** is promoted from `base-<hash>` via `crane copy`
(manifest copy, no rebuild) only when the base actually changed.
- **`docker buildx build --push` retry**: 3 attempts with backoff for
transient Hub blips. Deterministic failures fail all 3 and the job
fails as expected.
- **Registry buildcache disabled**: buildkit's cache-export hits HTTP 400
on Hub CDN since ~2026-05-23. Image push works fine; we pay the full
base build on Dockerfile.base change, but base tags are content-
addressed so unchanged bases short-circuit at the probe step.
## Decoupling history (briefly)
Pre-v1.0.0 pi-devbox was `FROM joakimp/pi-devbox:base-pi-only`, where
`base-pi-only` was a tag built by **opencode-devbox CI** (with
`INSTALL_OPENCODE=false` in their variant Dockerfile) and pushed under
the pi-devbox repo as an internal building-block tag. This setup
required rebuilding opencode-devbox before pi-devbox could be tagged
and meant pi-devbox docs needed cross-referencing into opencode-devbox.
v1.0.0 brings pi install logic into this repo, drops the cross-repo
dependency, and the `base-pi-only*` tags from opencode-devbox become
deprecated artifacts (to be removed in opencode-devbox v2.0.0).
## What we DON'T install (and why)
- **No texlive** (~600 MB–1 GB). PDF export from pandoc / pi-studio works
out of the box via **`typst`** (~30 MB static binary), which the base ships
as the pandoc PDF engine (`pandoc --pdf-engine=typst`) — small enough to live
in base rather than a dedicated `:latest-studio-tex` variant. We don't bake in
a full TeX Live: it's heavy and typst covers the common Markdown→PDF case.
Users needing LaTeX-exact output can install the higher-fidelity fallback on
demand: `sudo apt-get install texlive-xetex texlive-latex-recommended` (then
`pandoc --pdf-engine=xelatex`).
- **pi-studio** ships in the `:latest-studio` variant (since v1.1.0),
vendored to `/opt/pi-studio` and registered at container start via
`pi install /opt/pi-studio` (see Dockerfile.variant `INSTALL_STUDIO`).
The default `:latest` image stays studio-free. Note: pi-studio binds
`127.0.0.1` inside the container, so browser access needs host
networking or the bundled `studio-expose` bridge (socat; auto-starts
when `STUDIO_EXPOSE=1`) — see README "Using pi-studio".
- **No Julia/R/GHCi/Clojure runtimes**. Use `uv run --with X` for
Python REPLs; `apt install` other-language runtimes ad-hoc per
container if needed.
## Backward compatibility
- The host `~/.mempalace` bind-mount path is unchanged.
- Volume names (`devbox-pi-config`, `devbox-ssh-local`,
`devbox-shell-history`, `devbox-zoxide`, `devbox-nvim-data`,
`devbox-uv`; optional `devbox-palace`, `devbox-chroma-cache`) are
unchanged.
- `~/.pi/agent/` layout inside the container is unchanged; existing
named volumes work without recreation.
- The `:latest` and `vX.Y.Z` Hub tags continue to point at a "base + pi"
image. Same tag, same shape, just built differently.