diff --git a/CHANGELOG.md b/CHANGELOG.md index 1c63de4..b344ae0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -13,6 +13,73 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`). ## Unreleased +**`scripts/recreate-sanity-check.sh` asserted that `/tmp/sshcm` exists while every +`ssh` in the container was dying `rc=255`, and it was right to — it was checking +the directory the *image* creates, and the breakage was in the directory a +*config* named.** Found on the v1.9.2 first boot on `emb-7kj4vr4g`. A durable +`~/.pi/ssh/config` (hand-written into the `~/.pi` named volume by the previous +session, so it would survive the recreate) declared `ControlPath +/tmp/ssh-cm/%C` — with a hyphen. Nothing in this repo creates that path; the +canonical directory is `/tmp/sshcm`, spelled the same way in four places +(`Dockerfile.base`, `entrypoint-user.sh`, `recreate-sanity-check.sh`, +`smoke-test.sh`). Result: + +``` +unix_listener: cannot bind to path /tmp/ssh-cm/: No such file or directory +``` + +`rc=255`, and the remote command never ran at all — `ControlMaster auto` with an +unusable `ControlPath` fails hard rather than falling back to an unmultiplexed +connection. The old check passed truthfully, about the wrong object. **Two +independent facts about the same subsystem can both be true while the subsystem +is dead; a check that asserts only one of them cannot see their disagreement.** + +**New in `recreate-sanity-check.sh`: resolve the ControlPath `ssh` itself would +use, via `ssh -G`, and require its parent to exist and be writable.** `-G` +applies real config precedence — first-obtained-value-wins, the system drop-in, +`Include`, an `-F` override — so it answers "which rule captured this host" +instead of re-implementing the guess. It never opens a connection: measured +0.116 s for 48 hosts. + +This also puts a check under a caveat that had been documented in prose in +`Dockerfile.base` ("SSH client defaults") and verified nowhere: a per-host +`ControlPath ~/.ssh/cm/%r@%h:%p` inherited from a **read-only** bind-mounted +`~/.ssh` produces the identical failure, `cannot bind … Read-only file system`. +On the machine where this was built that is **16 of 50 hosts** — `freeipa-1..6`, +`gitea.egl.lan`, `runner-1..3`, `tor-ms22` and more — none of which had ever been +reported by anything. + +**The two routes get different severities, deliberately.** Default `ssh` +precedence legitimately lands in the read-only `~/.ssh` on any host whose own +config pins it there, and the supported workaround (`ssh -F +~/.ssh-local/config`, generated every container start by `setup-lan-access.sh`) +already exists — so that is a `warn`. Making it a failure would paint the script +red on every run of every device, and **a check that fires benignly every time is +one you learn to ignore**, which is the same reasoning that keeps `lint-shell.sh` +at `-S error`. The sidecar route is the prescribed one, so there an unusable +directory is a hard `fail`. Host lists are capped at six names plus a count for +the same reason: unreadable output is ignored output. + +Teeth proven in both directions, with each sabotage confirmed by `diff` *before* +the result was believed — a vacuous sabotage that silently fails to apply +reports "gate passed" and is worse than no test: + +| Sabotage | Class | Result | +|---|---|---| +| `ControlPath` → nonexistent dir | the original hyphen bug | `✗` `rc=1` | +| `ControlPath` → existing but read-only dir | the `Dockerfile.base` caveat | `✗` `rc=1` | +| restored | — | `✓` `rc=0`, sidecar byte-identical | + +No config was added to this repo to fix the original problem, because the fix was +to **delete** the offending file: `~/.pi/ssh/config` was a third hand-maintained +copy of what `setup-lan-access.sh` already generates from version control on +every start, with fewer features (no `known_hosts` sidecar, no +`StrictHostKeyChecking accept-new`) and one typo. The host-owned `~/.ssh/config` +is also left alone on purpose — those `~/.ssh/cm` paths are correct *on the host*, +where `~/.ssh` is writable, and the container-side override is the right layer. + +--- + **The size numbers on the Docker Hub page were the only claim in these docs with nothing in the repo to check them against, and they had gone 20% wrong across eight releases.** Every other claim `scripts/check-doc-drift.sh` guards is diff --git a/scripts/lint-shell.sh b/scripts/lint-shell.sh index d9ab69f..2032cbe 100644 --- a/scripts/lint-shell.sh +++ b/scripts/lint-shell.sh @@ -32,10 +32,16 @@ # # SEVERITY CHOICE # -S error is 0 findings across this repo when clean, so it is free to add. -# -S warning is NOT free here (19x SC2088 tilde-in-quotes in +# -S warning is NOT free here (20x SC2088 tilde-in-quotes in # recreate-sanity-check.sh, plus assorted SC2016 — both intentional), and a # noisy gate trains people to ignore it. Error-only, matching the # SHELLCHECK_OPTS philosophy in lint.yml. +# Reproduce the count before editing it (the `$ ` prefix is load-bearing: a +# comment whose first word is "shellcheck" is parsed as a DIRECTIVE, and a +# malformed one is SC1072/SC1073 at severity error — this gate caught exactly +# that when the line was first written without it): +# $ shellcheck -S warning -f gcc scripts/*.sh rootfs/usr/local/bin/* \ +# entrypoint*.sh hooks/* | grep -c SC2088 # # Usage: bash scripts/lint-shell.sh [root] (default root: repo top level) set -uo pipefail diff --git a/scripts/recreate-sanity-check.sh b/scripts/recreate-sanity-check.sh index a4fcb73..7aa049d 100755 --- a/scripts/recreate-sanity-check.sh +++ b/scripts/recreate-sanity-check.sh @@ -11,7 +11,8 @@ # pi-observational-memory / (studio variant) pi-studio package # registrations in settings.json packages[] # - Shell defaults re-seeded from /etc/skel-devbox -# - /tmp/sshcm exists with mode 700 (ssh ControlMaster dir) +# - ssh ControlMaster works: /tmp/sshcm exists 700 AND the ControlPath that +# ssh actually resolves (ssh -G) is a writable directory # - /opt toolkits intact # - Known expected-absences don't regress # @@ -425,13 +426,115 @@ if [ -d /opt/pi-atelier ] && command -v jq >/dev/null 2>&1; then fi echo -echo "-- ssh ControlMaster dir --" +echo "-- ssh ControlMaster: socket dir + EFFECTIVE ControlPath --" +# TWO LAYERS, and the second is the one that has actually broken in the field. +# +# LAYER 1 (original check): /tmp/sshcm, the directory entrypoint-user.sh creates +# for the base image's system drop-in +# (/etc/ssh/ssh_config.d/00-devbox-controlmaster.conf). +# +# LAYER 2 (added 2026-09-15): the directory a config NAMES — which is not the +# same question, and asserting layer 1 is structurally blind to it. On +# emb-7kj4vr4g a durable ~/.pi/ssh/config pointed ControlPath at /tmp/ssh-cm +# (with a hyphen), a directory nothing in the image creates. EVERY ssh died +# unix_listener: cannot bind to path /tmp/ssh-cm/: No such file or directory +# rc=255 with the remote command never running — while this script printed a +# green tick for layer 1, truthfully, about the wrong object. +# +# The same rc=255 has a second, independent cause already documented in prose in +# Dockerfile.base ("SSH client defaults" CAVEAT) and never verified anywhere: a +# per-host `ControlPath ~/.ssh/cm/%r@%h:%p` inherited from a bind-mounted +# READ-ONLY ~/.ssh. Measured to be the identical failure class: +# unix_listener: cannot bind to path ~/.ssh/cm/...: Read-only file system +# So do not guess which config wins — ask ssh. `ssh -G` applies real config +# precedence (first-obtained-value-wins, system drop-in, Include, -F override) +# and prints the fully expanded ControlPath. Require its parent to exist and be +# writable. Cost measured at 0.116 s for 48 hosts; -G never opens a connection. if [ -d /tmp/sshcm ] && [ "$(stat -c %a /tmp/sshcm 2>/dev/null)" = "700" ]; then pass "/tmp/sshcm exists with mode 700" else fail "/tmp/sshcm missing or not mode 700" fi +# Probe one route. $1 = label, $2 = config to force with -F ("" = ssh's own +# default precedence), $3 = severity when a ControlPath dir is unusable. +# +# SEVERITY SPLIT IS DELIBERATE. The default route legitimately resolves into the +# read-only ~/.ssh on any host whose own config pins ControlPath there, and the +# supported workaround (`ssh -F ~/.ssh-local/config`) already exists — so that +# is a warn, not a fail. Failing it would paint this script red on every run of +# every device, and a check that fires benignly every time is one you learn to +# ignore. The sidecar route is the PRESCRIBED one, so there it is a hard fail. +_ssh_cm_probe() { + local label="$1" cfg="${2:-}" sev="${3:-fail}" + local h out cm cp dir n=0 shown + local bad=() + while IFS= read -r h; do + [ -n "$h" ] || continue + if [ -n "$cfg" ]; then + out=$(ssh -F "$cfg" -G "$h" 2>/dev/null) || continue + else + out=$(ssh -G "$h" 2>/dev/null) || continue + fi + cm=$(printf '%s\n' "$out" | awk '/^controlmaster /{print $2; exit}') + case "$cm" in '' | no | none | false) continue ;; esac + cp=$(printf '%s\n' "$out" | awk '/^controlpath /{print $2; exit}') + case "$cp" in '' | none) continue ;; esac + n=$((n + 1)) + dir=$(dirname "$cp") + if [ ! -d "$dir" ] || [ ! -w "$dir" ]; then + bad+=("$h") + fi + done <<< "$SSH_CM_HOSTS" + + # Cap the host list. A 50-name line is the noisy gate this repo already warns + # about in lint-shell.sh: unreadable output is ignored output. Six names plus + # a count is enough to identify the class and act. + if [ "${#bad[@]}" -gt 0 ]; then + shown="${bad[*]:0:6}" + if [ "${#bad[@]}" -gt 6 ]; then + shown="$shown (+$(( ${#bad[@]} - 6 )) more)" + fi + fi + + if [ "$n" -eq 0 ]; then + warn "$label: no host resolves to ControlMaster on — effective ControlPath not exercised" + elif [ "${#bad[@]}" -eq 0 ]; then + pass "$label: $n ControlMaster host(s), every ControlPath dir exists and is writable" + elif [ "$sev" = "warn" ]; then + warn "$label: ${#bad[@]}/$n host(s) resolve ControlPath to a missing or unwritable dir [$shown] — expected when ~/.ssh/config pins ControlPath inside the read-only ~/.ssh; use 'ssh -F ~/.ssh-local/config' (see Dockerfile.base CAVEAT)" + else + fail "$label: ${#bad[@]}/$n host(s) resolve ControlPath to a missing or unwritable dir [$shown] — ssh dies rc=255 'unix_listener: cannot bind to path' and the remote command never runs" + fi +} + +if command -v ssh >/dev/null 2>&1; then + # Concrete Host aliases only: patterns (*, ?) and negations (!) are not + # connectable targets, so `ssh -G` on them proves nothing. + _cm_cfgs=() + if [ -r "$HOME/.ssh/config" ]; then _cm_cfgs+=("$HOME/.ssh/config"); fi + if [ -r "$HOME/.ssh-local/config" ]; then _cm_cfgs+=("$HOME/.ssh-local/config"); fi + if [ "${#_cm_cfgs[@]}" -gt 0 ]; then + SSH_CM_HOSTS=$(awk 'tolower($1)=="host"{for(i=2;i<=NF;i++) if ($i !~ /[*?!]/) print $i}' \ + "${_cm_cfgs[@]}" 2>/dev/null | sort -u) + else + SSH_CM_HOSTS="" + fi + + if [ -z "$SSH_CM_HOSTS" ]; then + warn "no concrete Host aliases in ~/.ssh/config or ~/.ssh-local/config — effective ControlPath not verified" + else + _ssh_cm_probe "default ssh precedence" "" warn + if [ -r "$HOME/.ssh-local/config" ]; then + _ssh_cm_probe "ssh -F ~/.ssh-local/config" "$HOME/.ssh-local/config" fail + else + warn "~/.ssh-local/config absent — setup-lan-access.sh did not run; the prescribed multiplex route is unverified" + fi + fi +else + warn "ssh not on PATH — effective ControlPath not verified" +fi + echo echo "-- Shell defaults re-seeded from /etc/skel-devbox --" if [ -f "$HOME/.bash_aliases" ]; then