ci: gate the release on shell lint, from one shared script
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Publish Docker Image / smoke (push) Successful in 7m39s
Publish Docker Image / build-variant-studio (push) Successful in 17m15s
Publish Docker Image / build-variant (push) Successful in 17m58s
Publish Docker Image / update-description (push) Successful in 8s
Publish Docker Image / promote-base-latest (push) Successful in 12s

v1.8.14's first attempt spent ~46 minutes building a base image for a tree whose
own lint had been failing for 24 hours. shellcheck had already flagged the
defect (SC2289, severity error) on the push that introduced it; the lint
workflow went red at run 186 and nobody read it.

lint.yml deliberately skips tag pushes and its reasoning is sound -- the tagged
tree was already linted on main, and a tag-ref lint run sorts above the publish
run, making a release look finished before anything ships. The missing invariant
was never "lint the tag". It was "do not RELEASE a tree whose lint failed", and
only a job inside the publish workflow can enforce that.

So: extract the shell-lint logic from lint.yml into scripts/lint-shell.sh and
call it from both places, then add a lint-gate job that resolve-versions depends
on. resolve-versions is the graph root, so gating it gates everything. Cost is
~40 s at the front of a release; the alternative already cost fifty minutes.

Extracted rather than copied on purpose. A second copy of a check is the drift
this repo keeps paying for -- the same evening produced a skillset mirror that
had sat 9579 B behind its upstream through two consecutive edits.

The script adds one behaviour the inline version lacked: if shellcheck is not
installed it exits 2 rather than silently finding nothing, inheriting the
existing "a gate that cannot run must not pass" rule from hooks/pre-commit in
the skillset repo. Without that, reordering the install step away would turn the
gate into a green tick over zero checks.

Verified locally with a stubbed shellcheck (the real binary is not in the
devbox), five cases, each with its expectation stated first: absent shellcheck
-> rc=2; stub pass -> rc=0 and a non-zero file count; stub fail -> rc=1; a
deliberately unterminated `if` planted in scripts/ -> rc=1 via the bash -n half,
naming the file; removal -> rc=0 again. Discovery cross-checks against CI's own
number: the inline version reported 12 files, the extracted one reports 13, the
difference being lint-shell.sh itself. YAML re-parsed (10 jobs, was 9) with an
assertion that resolve-versions needs lint-gate, and the repo's
check-workflow-shell.sh guard still passes.
This commit is contained in:
2026-09-08 23:41:44 +02:00
parent 70e675afee
commit 361babd4fd
4 changed files with 144 additions and 25 deletions
+14
View File
@@ -35,6 +35,20 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
> broken assumption was different, namely that a tree whose lint FAILED would not
> then be released. `docker-publish.yml` has no dependency on lint, so it built
> for 50 minutes on a tree known to be defective.
>
> **Fixed, then gated.** The prose moved above the `exec_test` call so an
> apostrophe cannot terminate anything, and the shell-lint logic moved out of
> `lint.yml` into **`scripts/lint-shell.sh`** — now called by both `lint.yml` and
> a new `lint-gate` job here that `resolve-versions` depends on. A release with a
> lint error refuses in ~40 s instead of failing after fifty minutes. One copy,
> not two: a duplicated check that drifts is the failure this repo keeps paying
> for. The script also refuses to pass when `shellcheck` is absent, inheriting
> the existing principle that a gate which cannot run must not pass.
>
> **`v1.8.14` was re-pointed** from `601fc98` to the fix commit. Nothing had
> consumed the original tag — no `v1.8.14` image was ever published, only the
> content-addressed `base-a365dd24de21`. `scripts/` does not feed the base hash,
> so the re-run reuses that base and skips the 46-minute rebuild.
**A test that was quietly checking nothing, and a version number that was wrong.**
Both found by delegating a read-only audit of this repo to a headless worker