50153e65b75c4a266d8cc14a7f8a20b1375a9c60
2 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f25efa074d |
fix(sanity): verify the ControlPath ssh RESOLVES, not the dir the image creates
recreate-sanity-check.sh asserted `/tmp/sshcm exists with mode 700` and printed
a green tick while every ssh in the container died rc=255. It was right about
what it checked: the breakage was in the directory a CONFIG named, not the one
the image creates, and the old check could not see the disagreement between them.
Found on the v1.9.2 first boot on emb-7kj4vr4g. A durable ~/.pi/ssh/config,
hand-written into the ~/.pi named volume by the previous session so it would
survive the recreate, declared `ControlPath /tmp/ssh-cm/%C` — with a hyphen.
Nothing here creates that path; the canonical spelling is /tmp/sshcm, identical
in Dockerfile.base, entrypoint-user.sh, recreate-sanity-check.sh and
smoke-test.sh. ControlMaster auto with an unusable ControlPath does not degrade
to an unmultiplexed connection — it fails hard:
unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory
rc=255, remote command never runs.
Now: resolve the ControlPath ssh itself would use via `ssh -G` and require its
parent to exist and be writable. -G applies real config precedence (first-value-
wins, the system drop-in, Include, an -F override), so it answers "which rule
captured this host" instead of re-implementing the guess, and it never opens a
connection — 0.116s for 48 hosts.
This also puts a check under a caveat documented in prose in Dockerfile.base
("SSH client defaults") and verified nowhere: a per-host ControlPath under a
read-only bind-mounted ~/.ssh gives the identical failure, "Read-only file
system". On the machine this was built on that is 16 of 50 hosts — freeipa-1..6,
gitea.egl.lan, runner-1..3, tor-ms22 — never reported by anything before.
Severity split is deliberate. Default ssh precedence legitimately lands in the
read-only ~/.ssh for any host whose own config pins it there, and the supported
workaround (ssh -F ~/.ssh-local/config, generated every start by
setup-lan-access.sh) already exists, so that is a warn. Failing it would paint
the script red on every run of every device, and a check that fires benignly
every time is one you learn to ignore — the same reasoning that keeps
lint-shell.sh at -S error. The sidecar route is prescribed, so there it is a
hard fail. Host lists cap at six names plus a count: unreadable output is
ignored output.
Teeth proven both directions, each sabotage confirmed by diff BEFORE the result
was believed: ControlPath -> nonexistent dir => rc=1; -> existing-but-read-only
dir => rc=1; restored => rc=0 with the sidecar byte-identical.
No config was added to fix the original problem — the fix was to DELETE
~/.pi/ssh/config, a third hand-maintained copy of what setup-lan-access.sh
already generates from version control on every container start, with fewer
features and one typo. The host-owned ~/.ssh/config is left alone on purpose:
those ~/.ssh/cm paths are correct on the host, where ~/.ssh is writable.
Also lint-shell.sh: the SC2088 count in the severity-choice rationale said 19
and is now 20, with the reproduce command recorded. That command needs a `$ `
prefix — a comment whose first word is "shellcheck" is parsed as a directive,
and the malformed one tripped SC1072/SC1073 at severity error. The gate caught
it on the very commit that introduced it.
|
||
|
|
361babd4fd |
ci: gate the release on shell lint, from one shared script
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Publish Docker Image / smoke (push) Successful in 7m39s
Publish Docker Image / build-variant-studio (push) Successful in 17m15s
Publish Docker Image / build-variant (push) Successful in 17m58s
Publish Docker Image / update-description (push) Successful in 8s
Publish Docker Image / promote-base-latest (push) Successful in 12s
v1.8.14's first attempt spent ~46 minutes building a base image for a tree whose own lint had been failing for 24 hours. shellcheck had already flagged the defect (SC2289, severity error) on the push that introduced it; the lint workflow went red at run 186 and nobody read it. lint.yml deliberately skips tag pushes and its reasoning is sound -- the tagged tree was already linted on main, and a tag-ref lint run sorts above the publish run, making a release look finished before anything ships. The missing invariant was never "lint the tag". It was "do not RELEASE a tree whose lint failed", and only a job inside the publish workflow can enforce that. So: extract the shell-lint logic from lint.yml into scripts/lint-shell.sh and call it from both places, then add a lint-gate job that resolve-versions depends on. resolve-versions is the graph root, so gating it gates everything. Cost is ~40 s at the front of a release; the alternative already cost fifty minutes. Extracted rather than copied on purpose. A second copy of a check is the drift this repo keeps paying for -- the same evening produced a skillset mirror that had sat 9579 B behind its upstream through two consecutive edits. The script adds one behaviour the inline version lacked: if shellcheck is not installed it exits 2 rather than silently finding nothing, inheriting the existing "a gate that cannot run must not pass" rule from hooks/pre-commit in the skillset repo. Without that, reordering the install step away would turn the gate into a green tick over zero checks. Verified locally with a stubbed shellcheck (the real binary is not in the devbox), five cases, each with its expectation stated first: absent shellcheck -> rc=2; stub pass -> rc=0 and a non-zero file count; stub fail -> rc=1; a deliberately unterminated `if` planted in scripts/ -> rc=1 via the bash -n half, naming the file; removal -> rc=0 again. Discovery cross-checks against CI's own number: the inline version reported 12 files, the extracted one reports 13, the difference being lint-shell.sh itself. YAML re-parsed (10 jobs, was 9) with an assertion that resolve-versions needs lint-gate, and the repo's check-workflow-shell.sh guard still passes. |