feat(ci): gate the vendored pi-extensions skill floor, and bake python3-yaml
Followsecfd2fc, which refreshed the stale floor by hand. A one-off refresh fixes the symptom; this makes the drift impossible to reintroduce silently. scripts/check-skill-floor.sh compares the repo floor (rootfs/usr/local/share/pi-devbox/skills/pi-extensions/) against the package repo it is a snapshot of, wired in as a new `skill-floor` job in lint.yml. DIRECTORY hash, not `sha256sum SKILL.md`, using the same tree_sha256 pipeline Dockerfile.variant uses for skillset_snapshot_tree_sha256 and for the reason already documented there: a file-only compare answers "did this one file change", not "is this the same skill". Verified by NEGATIVE CONTROL rather than asserted -- with SKILL.md left byte-identical and only evaluate-extension-usage.py edited, the directory check fails (rc=1) where a file-only compare would have passed. Seven behaviour tests, each with its expected rc written down before running: in-sync via local dir (0), in-sync via anonymous remote clone (0), missing --package-dir (2), bad argument (2), content drift (1), the sibling-file case (1), and --warn-only over drift (0). Exit codes 0 in sync / 1 drift / 2 cannot-run, matching scripts/lint-shell.sh: a gate that cannot run must not pass, so an unreachable package repo is a red 2 and never a green tick. A ref with no skill/ is NOT drift -- that is the documented fallback -- but it emits ::warning:: because it is precisely the condition under which the floor ships. Gating on another repo is normally a smell. It is proportionate here because the check can only fire when skill/ itself changed, which is exactly when the floor has gone stale; pi-extensions commits that leave skill/ alone cannot turn this red. It also needs no secret: pi-extensions is anonymously clonable (verified with `git ls-remote` and no credentials), so it cannot start failing when a token expires. Also bakes python3-yaml (552 KB, zero extra deps) into Dockerfile.base. This is the shellcheck story repeating exactly: scripts/check-workflow-shell.sh -- the guard against the Gitea sh/dash footgun that broke resolve-versions (ed49b8d) and promote-base-latest (b7197e8) -- hard-exits with "python3 yaml module missing", so a gate this repo already owns could not be run locally by anyone. lint.yml installing it explicitly in CI was the evidence. Found while wiring the job above: the guard could not be run before pushing. CHANGELOG Unreleased updated for both this andecfd2fc, including an explicit note on what is NOT fixed -- the silent-fallback half still has no manifest flag recording which copy was served. Verified locally with every gate this repo owns, all green: lint-shell.sh (15 files clean), check-workflow-shell.sh, check-base-hash.sh, actionlint 1.7.7 (pinned, same version as CI), hadolint 2.14.0, and the new check itself.
This commit is contained in:
@@ -92,6 +92,66 @@ derivation's `mine` query at the newest end (`order: "desc"`); with the previous
|
||||
default `asc` + `limit: 100`, a device passing 100 authored events would have its
|
||||
recent replies fall out of the join window and see answered asks resurface.
|
||||
|
||||
**Four small packages, each chosen from a gap that was measured rather than
|
||||
imagined.** All four were picked by looking back at a real session — the
|
||||
`gitea.egl.lan`/FreeIPA debugging of 2026-09-09..10 — and asking which absences
|
||||
actually cost time, not which tools sound useful. `bind9-dnsutils` (~6.1 MB, 10
|
||||
packages): `dig`, `host` **and** `nslookup` were all absent, so the container
|
||||
could resolve names but had no way to interrogate a *specific* nameserver —
|
||||
`getent hosts` only follows the resolver's default path, so diagnosing "gateway
|
||||
`172.16.88.1` NXDOMAINs the `egl.lan` zone while `10.20.253.1` is authoritative
|
||||
for it" had to be hand-rolled in `python3`. Note the package name: plain
|
||||
`dnsutils` is transitional in trixie. `ldap-utils` (1244 KB, **zero** extra deps):
|
||||
the fleet authenticates against FreeIPA, yet every LDAP probe had to be run by
|
||||
SSHing to an already-enrolled host; this gives simple binds only, since GSSAPI
|
||||
would additionally need `krb5-user` + `libsasl2-modules-gssapi-mit`, which is a
|
||||
Kerberos-client decision rather than a tool. `xxd` (198 KB) is frank convenience
|
||||
— `od -c` already does the job. `python3-yaml` (552 KB, zero extra deps) is the
|
||||
shellcheck story repeating exactly: `scripts/check-workflow-shell.sh`, the guard
|
||||
against the Gitea `sh`/dash footgun that broke `resolve-versions` (`ed49b8d`) and
|
||||
`promote-base-latest` (`b7197e8`), hard-exits with "python3 yaml module missing"
|
||||
without it — and `lint.yml` installing it explicitly in CI was the evidence the
|
||||
image lacked it. **`netcat-openbsd` was proposed and deliberately rejected**:
|
||||
measured redundant, because `socat` is already baked and bash's `/dev/tcp` does
|
||||
reachability checks with zero packages. The reason is recorded in
|
||||
`Dockerfile.base` so the omission reads as a decision rather than an oversight.
|
||||
|
||||
**The vendored `pi-extensions` skill floor was 41 days stale, and is now gated so
|
||||
it cannot silently rot again.** `rootfs/usr/local/share/pi-devbox/skills/pi-extensions/`
|
||||
sat at 34284 B, untouched since `fa04d20` (2026-07-30), while the package copy
|
||||
was 38973 B — four copies of one skill existed across the fleet with three
|
||||
different sizes. `Dockerfile.variant` copies the freshly-cloned package copy over
|
||||
the **served** path but never writes back to the repo floor, so nothing in the
|
||||
repo ever noticed. That is worse than ordinary staleness because the floor is a
|
||||
**fallback**: the copy is guarded by `if [ -f /opt/pi-extensions/skill/SKILL.md ]`,
|
||||
so a build whose clone yields no `skill/` keeps the vendored snapshot and still
|
||||
goes **green**, with no manifest flag and no label recording which copy was
|
||||
served — the image would ship a July skill and nothing would say so. The floor is
|
||||
refreshed here from `pi-extensions@c64c122`, and the new `skill-floor` job in
|
||||
`lint.yml` runs `scripts/check-skill-floor.sh` to keep it that way.
|
||||
|
||||
The check compares the **directory** hash, using the same `tree_sha256` pipeline
|
||||
`Dockerfile.variant` uses for `skillset_snapshot_tree_sha256` and for the same
|
||||
documented reason: a `sha256sum SKILL.md` answers "did this one file change", not
|
||||
"is this the same skill", and `pi-extensions` ships two files. That is not
|
||||
hypothetical — it was **verified by negative control**: with `SKILL.md` left
|
||||
byte-identical and only `evaluate-extension-usage.py` edited, the directory check
|
||||
correctly fails while a file-only compare would have passed. Exit codes are `0`
|
||||
in sync / `1` drift / `2` cannot-run, matching `scripts/lint-shell.sh`, so an
|
||||
unreachable package repo is a red `2` rather than a green tick. Gating on another
|
||||
repo is normally a smell; it is proportionate here because the check can only
|
||||
fire when `skill/` itself changed — which is exactly when the floor has gone
|
||||
stale — and it needs no secret, since `pi-extensions` is anonymously clonable
|
||||
(verified with `git ls-remote` and no credentials).
|
||||
|
||||
> **What this does *not* fix, stated so nobody reads more into it than is there.**
|
||||
> The floor is now fresh and guarded, but the *silent-fallback* half remains:
|
||||
> if the build-time copy is ever absent, the build still succeeds with no
|
||||
> manifest flag or OCI label recording that the vendored snapshot was served
|
||||
> instead of the package copy. The durable fix for that is a manifest field
|
||||
> alongside the existing `skillset_snapshot_tree_sha256`, which this change does
|
||||
> not add.
|
||||
|
||||
---
|
||||
|
||||
## v1.8.14 — 2026-09-08
|
||||
|
||||
Reference in New Issue
Block a user