entrypoint: put back the shell state a recreate eats
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 25s

cli_utils' install.sh reaches PATH by symlinking bin/ into ~/.local/bin.
That is persistent on a host and ephemeral in a container, so the same
installer produced opposite durability and every --force-recreate sent the
human back to typing /workspace/cli_utils/bin/git-status-all. Re-link at
start instead, and add a per-device boot hook so the next question of this
shape needs no image change at all.

Symlinks rather than a PATH edit in an rc file, deliberately: ~/.local/bin
is already ahead of /usr/local/bin in ENV PATH, so links resolve in
NON-interactive shells too (docker exec, agent tool shells, scripts). An
rc-file PATH edit cannot reach those because ~/.bashrc returns early when
not interactive — measured, that asymmetry is exactly why `command -v
git-status-all` failed in one shell and worked in another on the same box.
Shell FUNCTIONS remain the sourced file's job; a symlink cannot carry them.

Guards, because ~/.local/bin is shared: a real file is never clobbered, a
symlink pointing elsewhere is never stolen, ours are refreshed, and links
into a cli_utils/bin whose target vanished are pruned — a dangling link on
PATH reads as a broken container rather than a removed script.

The hook (~/.config/devbox-shell/init.sh) adds no trust boundary: that dir
is already sourced into every interactive shell by the baked bash_aliases,
so it is already arbitrary code from the same owner. Only WHEN it runs is
new. bash <file>, never sourced, exit status ignored, output to a log.

Caught before commit, and the reason the loops use `if` bodies instead of
`&&` chains: under `set -euo pipefail` a for-loop whose last command is a
false test exits non-zero, and with no match the /workspace/*/cli_utils
glob stays literal — so the first draft would have failed to START a
container on every machine that does not have this repo, rather than merely
skipping the links. Re-tested with set -e in place: no-cli_utils/empty-HOME
no-op, guards, idempotence, CLI_UTILS_LINK=0, a hook that exits 7, and a
hook that tries to mutate CLI_UTILS_BIN — all exit 0 with intact state.

Not covered by CI: docker-publish.yml runs only on tags and lint.yml lints
workflow run: steps, so neither executes this file. Validated by extracting
both sections and running them against fixtures, then for real in a live
v1.8.10 container. No smoke assertion added on purpose — the positive path
needs a /workspace mount smoke does not have, and asserting it there would
repeat the v1.8.0 mistake of a smoke check written against a stage that
does not exist at run time.

Moves the base hash (base-decide folds `cat entrypoint.sh
entrypoint-user.sh`), so this rides along with the next tag's ~40-minute
base rebuild rather than justifying a tag of its own.
This commit is contained in:
2026-08-27 21:52:51 +02:00
parent 6891dc32b8
commit 45850bc973
4 changed files with 211 additions and 0 deletions
+69
View File
@@ -11,6 +11,75 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
---
## Unreleased
**Shell state that the writable layer eats on every recreate now gets rebuilt at
start.** Two additions to `entrypoint-user.sh`, both idempotent, both silent
no-ops when the thing they wire up is absent.
**`cli_utils` commands are linked onto `PATH`.** If a `cli_utils` checkout is
mounted, every executable in its `bin/` is symlinked into `~/.local/bin` at
container start — `git-status-all`, `git-pull-all`, `devbox-sanity`,
`pi-devbox-sanity`, `pi-session-repair`, `docker-clean`, `vpn-status`. Detection:
`CLI_UTILS_CONTAINER_PATH` → `/workspace/cli_utils` → `$HOME/cli_utils` →
`/workspace/*/cli_utils`; `CLI_UTILS_LINK=0` disables it.
The reason this is an *image* concern and not the user's problem to re-solve: on a
host, `cli_utils/install.sh` puts those commands on `PATH` by symlinking them into
`~/.local/bin`, which is persistent there — and **ephemeral here**. Same installer,
same repo, opposite durability, so the fix died on every `--force-recreate` and the
next session was back to typing `/workspace/cli_utils/bin/git-status-all`. Running
`install.sh` *inside* a container is the trap rather than the fix: it re-creates
the same disposable state.
**Symlinks rather than a `PATH` edit in an rc file, deliberately.** `~/.local/bin`
is already ahead of `/usr/local/bin` in `ENV PATH`, so links resolve in
**non-interactive** shells too — `docker exec <c> git-status-all`, agent tool
shells, scripts. An rc-file `PATH` edit cannot reach those: `~/.bashrc` returns
early when the shell is not interactive. Measured on tor-ms22 2026-08-27,
`command -v git-status-all` failed in a non-interactive shell while succeeding in
an interactive one, from exactly that asymmetry. Guards, because `~/.local/bin` is
shared with other tooling: a real file is never clobbered, a symlink pointing
somewhere else is never stolen, our own links are refreshed, and links into a
`cli_utils/bin` whose target vanished are pruned — a dangling link on `PATH`
reports "No such file or directory" and reads as a broken container rather than a
removed script.
**A per-device boot hook: `~/.config/devbox-shell/init.sh`.** If the host provides
one, it runs once at start with output to `~/.pi/agent/devbox-init.log`. That
directory is the host-owned bind-mount already sourced into every interactive
shell by `/etc/skel-devbox/.bash_aliases`, so this is its boot-time twin — the
same ownership and the same persistence, but running *before any shell*, which is
what non-interactive fixups (symlinks, directories, one-off migrations) need. **It
introduces no new trust boundary**: that path is already arbitrary code from the
same owner; only *when* it runs is new. Invoked as `bash <file>`, never sourced,
and its exit status is ignored — a hook must not be able to mutate the
entrypoint's own shell state or stop a container from starting.
With the hook in place, the next "can this run on every recreate?" question needs
no image change at all — which is the point, given what the next paragraph costs.
**This moves the base hash.** `base-decide` folds `cat entrypoint.sh
entrypoint-user.sh` into it, so this change forces the ~40-minute base rebuild at
the next tag whether or not anything else in the base moved. It is a rider, not a
reason to tag.
**How it was validated, since CI cannot.** `docker-publish.yml` runs only on
`push: tags: v*`, and `lint.yml` runs `actionlint` over workflow `run:` steps —
neither one executes `entrypoint-user.sh`. So both sections were extracted and run
against fixtures in a throwaway `$HOME` before commit: real file not clobbered,
foreign symlink respected, stale link pruned, new command picked up, second run
byte-identical, `CLI_UTILS_LINK=0` honoured, and "no `cli_utils` anywhere" a silent
`exit 0`. Then run for real in a live v1.8.10 container, after which
`command -v git-status-all` resolved in a *non-interactive* shell. No
`smoke-test.sh` assertion was added on purpose: the positive path needs a
`/workspace` mount that smoke does not have, and asserting it there would repeat
the v1.8.0 mistake of a smoke assertion written against a stage that does not
exist at run time. `workflow_dispatch` with `smoke_only` remains the way to
exercise this against `HEAD` before a tag.
---
## v1.8.10 — 2026-08-27
**This tag exists to deploy a fix and a safety net that are currently running on