Commit Graph

10 Commits

Author SHA1 Message Date
Joakim Persson 3a44e81cad feat(ci): gate documentation drift, and make docs a pre-tag release step
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 17s
Lint / doc-drift (push) Successful in 7s
Lint / skill-floor (push) Successful in 10s
Five doc claims had rotted by v1.9.0, all the same shape: a value written once
by hand, in a file nothing verifies, about a number that lives elsewhere and
moved. README pin table wrong on all three rows; a "Planned" section describing
something already shipped; DOCKER_HUB.md claiming Node v22 against Node 24.

DOCKER_HUB.md is why this is a gate and not a resolution to be careful: it is
PUBLISHED (update-description POSTs it as Docker Hub full_description on every
tag), it had gone eight releases untouched, nothing generates it, and it is read
from the TAG -- so the stale page shipped with v1.9.0 regardless.

scripts/check-doc-drift.sh: seven checks, all repo-local (no network, token,
image, or sibling clone). Exit 0/1/2 matching lint-shell.sh; a renamed ARG is a
red 2, not a green tick. Wired as a fourth lint.yml job so "the docs lie" is its
own red name.

Verified with 15 controls, including two false-positive controls: the first
placeholder check flagged README.md:900, a Go template in a legitimate
`docker inspect --format` example. The gate was wrong, not the doc, so the
pattern is now anchored to the UPPER_SNAKE convention CI substitutes.

Not gated, deliberately: counts/sizes needing a running image (they belong in
smoke-test.sh -- a guessing gate is worse than none), and Dockerfile.base
BASE_REBUILD_DATE, because base_tag hashes that file content-wise and demanding
it be current would force a ~60 min rebuild on releases that touch no base
files. Free during a rebuild, expensive otherwise.

AGENTS.md step 3 rewritten around the mechanism: checkout@v4 with no ref: means
every job reads github.ref, the tag. Docs must be right BEFORE tagging.
2026-09-10 22:01:49 +02:00
Joakim Persson edc7659add chore(deps): node 22->24, actionlint 1.7.12, hadolint 2.15.1, skillset ref
Audited every component the image obtains OUTSIDE debian/apt. Of ~23, the 19
that resolve `latest` at build time were already current or refresh themselves
on the next rebuild, and the hard pins for pi (0.85.1), mempalace (3.9.0) and
pi-atelier (v0.10.1) were already newest. Four needed a human.

NODE_VERSION 22 -> 24 (LTS "Krypton"). This was a latent defect rather than
housekeeping: agent-browser publishes engines.node ">=24.0.0", so the image sat
BELOW a declared requirement -- v1.8.14 shipped node 22.23.2 with agent-browser
0.37.1, so every build installed it with an npm EBADENGINE warning and ran the
baked browser automation outside its supported range. pi (">=22.19.0") and
playwright (">=20") are satisfied either way. Verified before bumping, since a
missing NodeSource suite breaks every arch at once: setup_24.x returns HTTP 200
and node_24.x advertises `Architectures: amd64 arm64 armhf x86_64`, covering
the arm64 fleet and the amd64 CI runners. Nothing else pinned the node major.

actionlint 1.7.7 -> 1.7.12 and hadolint 2.14.0 -> 2.15.1, each RUN AGAINST THIS
TREE at the new version before being pinned -- both clean, no new findings. A
linter bump is the one dependency update that can turn CI red on unchanged
code, so it is verified locally rather than discovered on a round trip.

SKILLSET_SNAPSHOT_REF e9e09d9 -> 4d7c0ea via scripts/vendor-mempalace-skill.sh,
never by hand: that script is the only thing permitted to write the ARG,
because a cp without a matching bump yields a manifest that confidently lies.
This proved PROVENANCE-ONLY -- the ref was 6 commits behind, but
skills/mempalace/SKILL.md is byte-identical at both (3675bfab), so the snapshot
was already correct and only its recorded origin was stale. No rootfs/ bytes
changed, the smoke-test phrase canary stays valid, and this ARG alone would not
force a base rebuild (the node bump does).

Two measurement traps worth recording, since both would have produced a wrong
answer: GitHub's releases/latest reports pi-atelier v0.10.0 as newest because
v0.10.1 is a TAG WITH NO RELEASE OBJECT -- the pin was already current, and
`git ls-remote --tags` is the instrument that shows it. And gitea-mcp is hosted
on gitea.com, not GitHub, so querying api.github.com returned nothing at all
rather than an error.

Verified with every gate this repo owns, all green, using the NEW linter pins:
lint-shell.sh (15 files), check-workflow-shell.sh, check-base-hash.sh,
actionlint 1.7.12, hadolint 2.15.1, check-skill-floor.sh, and
vendor-mempalace-skill.sh --check.
2026-09-10 20:17:32 +02:00
Joakim Persson cac5e00a31 feat(ci): gate the vendored pi-extensions skill floor, and bake python3-yaml
Follows ecfd2fc, which refreshed the stale floor by hand. A one-off refresh
fixes the symptom; this makes the drift impossible to reintroduce silently.

scripts/check-skill-floor.sh compares the repo floor
(rootfs/usr/local/share/pi-devbox/skills/pi-extensions/) against the package
repo it is a snapshot of, wired in as a new `skill-floor` job in lint.yml.

DIRECTORY hash, not `sha256sum SKILL.md`, using the same tree_sha256 pipeline
Dockerfile.variant uses for skillset_snapshot_tree_sha256 and for the reason
already documented there: a file-only compare answers "did this one file
change", not "is this the same skill". Verified by NEGATIVE CONTROL rather
than asserted -- with SKILL.md left byte-identical and only
evaluate-extension-usage.py edited, the directory check fails (rc=1) where a
file-only compare would have passed. Seven behaviour tests, each with its
expected rc written down before running: in-sync via local dir (0), in-sync
via anonymous remote clone (0), missing --package-dir (2), bad argument (2),
content drift (1), the sibling-file case (1), and --warn-only over drift (0).

Exit codes 0 in sync / 1 drift / 2 cannot-run, matching scripts/lint-shell.sh:
a gate that cannot run must not pass, so an unreachable package repo is a red
2 and never a green tick. A ref with no skill/ is NOT drift -- that is the
documented fallback -- but it emits ::warning:: because it is precisely the
condition under which the floor ships.

Gating on another repo is normally a smell. It is proportionate here because
the check can only fire when skill/ itself changed, which is exactly when the
floor has gone stale; pi-extensions commits that leave skill/ alone cannot
turn this red. It also needs no secret: pi-extensions is anonymously clonable
(verified with `git ls-remote` and no credentials), so it cannot start failing
when a token expires.

Also bakes python3-yaml (552 KB, zero extra deps) into Dockerfile.base. This
is the shellcheck story repeating exactly: scripts/check-workflow-shell.sh --
the guard against the Gitea sh/dash footgun that broke resolve-versions
(ed49b8d) and promote-base-latest (b7197e8) -- hard-exits with "python3 yaml
module missing", so a gate this repo already owns could not be run locally by
anyone. lint.yml installing it explicitly in CI was the evidence. Found while
wiring the job above: the guard could not be run before pushing.

CHANGELOG Unreleased updated for both this and ecfd2fc, including an explicit
note on what is NOT fixed -- the silent-fallback half still has no manifest
flag recording which copy was served.

Verified locally with every gate this repo owns, all green: lint-shell.sh (15
files clean), check-workflow-shell.sh, check-base-hash.sh, actionlint 1.7.7
(pinned, same version as CI), hadolint 2.14.0, and the new check itself.
2026-09-10 19:14:01 +02:00
joakimp 361babd4fd ci: gate the release on shell lint, from one shared script
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Publish Docker Image / smoke (push) Successful in 7m39s
Publish Docker Image / build-variant-studio (push) Successful in 17m15s
Publish Docker Image / build-variant (push) Successful in 17m58s
Publish Docker Image / update-description (push) Successful in 8s
Publish Docker Image / promote-base-latest (push) Successful in 12s
v1.8.14's first attempt spent ~46 minutes building a base image for a tree whose
own lint had been failing for 24 hours. shellcheck had already flagged the
defect (SC2289, severity error) on the push that introduced it; the lint
workflow went red at run 186 and nobody read it.

lint.yml deliberately skips tag pushes and its reasoning is sound -- the tagged
tree was already linted on main, and a tag-ref lint run sorts above the publish
run, making a release look finished before anything ships. The missing invariant
was never "lint the tag". It was "do not RELEASE a tree whose lint failed", and
only a job inside the publish workflow can enforce that.

So: extract the shell-lint logic from lint.yml into scripts/lint-shell.sh and
call it from both places, then add a lint-gate job that resolve-versions depends
on. resolve-versions is the graph root, so gating it gates everything. Cost is
~40 s at the front of a release; the alternative already cost fifty minutes.

Extracted rather than copied on purpose. A second copy of a check is the drift
this repo keeps paying for -- the same evening produced a skillset mirror that
had sat 9579 B behind its upstream through two consecutive edits.

The script adds one behaviour the inline version lacked: if shellcheck is not
installed it exits 2 rather than silently finding nothing, inheriting the
existing "a gate that cannot run must not pass" rule from hooks/pre-commit in
the skillset repo. Without that, reordering the install step away would turn the
gate into a green tick over zero checks.

Verified locally with a stubbed shellcheck (the real binary is not in the
devbox), five cases, each with its expectation stated first: absent shellcheck
-> rc=2; stub pass -> rc=0 and a non-zero file count; stub fail -> rc=1; a
deliberately unterminated `if` planted in scripts/ -> rc=1 via the bash -n half,
naming the file; removal -> rc=0 again. Discovery cross-checks against CI's own
number: the inline version reported 12 files, the extracted one reports 13, the
difference being lint-shell.sh itself. YAML re-parsed (10 jobs, was 9) with an
assertion that resolve-versions needs lint-gate, and the repo's
check-workflow-shell.sh guard still passes.
2026-09-08 23:41:44 +02:00
joakimp 9e744d701f lint: shellcheck the repo's own shell scripts, not just workflow run: steps
Lint / actionlint (push) Successful in 15s
Lint / hadolint (push) Successful in 15s
lint.yml has shellchecked every workflow `run:` step since the dash-vs-bash
incidents, but nothing had ever pointed shellcheck at entrypoint.sh, scripts/*.sh
or the extensionless tools under rootfs/usr/local/bin/. That gap is not
hypothetical: the skillset repo's ci-release-watcher template shipped
`echo "$json" | python3 <<'EOF' ... json.load(sys.stdin)` for two months, where
the heredoc IS python's stdin (no script arg) so the load hit EOF and the
function silently returned nothing. shellcheck names exactly that at severity
ERROR — SC2259, "This redirection overrides piped input" — and could have named
it the whole time.

New step in the existing actionlint job, so no second container pull: shellcheck
-S error plus bash -n over every shell file, discovered as *.sh UNION a shebang
scan (the glob alone misses pi-devbox-version, devbox-skill-reconcile, dot-watch
and studio-expose; a shebang scan alone would miss a sourced fragment without
one). Fails loudly on a zero-file match, because a green tick over an empty set
is not a check.

Severity chosen by measurement, not taste: -S error is 0 findings across all 11
shell files today, so the gate is green on arrival with no cleanup, while
-S warning is NOT free (19x SC2088 tilde-in-quotes in recreate-sanity-check.sh
plus assorted SC2016, all intentional) and would train everyone to ignore the
job — the same reasoning as the SHELLCHECK_OPTS exclusions already on the
actionlint step.
2026-08-25 21:11:55 +02:00
pi 62a2a79b1c ci(lint): correct the rationale comment — runner contention was overstated
Lint / hadolint (push) Successful in 15s
Lint / actionlint (push) Successful in 4m29s
The previous commit justified excluding tag pushes partly on runner contention:
that the duplicate lint run stole one of two self-hosted runners from the release
build. Measured, that is false for THIS repo — pi-devbox lint runs take 0.3-0.9
min (ids 529/531/532/533) against a 77.6 min release build (id=530). I imported
the claim from opencode-devbox, where actionlint apt-installs shellcheck inside
the container and takes 6-15 min, so contention there is real.

The change stands on its actual merits: duplicate lint of an identical tree, and
release-run discovery ambiguity (the substantive one — it is what made the naive
"first run matching refs/tags/<tag>" rule pick lint over the publish run).

No functional change; comment only.
2026-08-04 18:08:28 +02:00
pi f20b2a7926 ci(lint): don't re-lint on tag pushes
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Has been cancelled
`on: push:` with no filter also fires on refs/tags/v*, which is duplicate work:
the tagged tree was already linted when that same commit was pushed to main
(v1.6.4 sha e86e5df linted as id=529 on main, then again as id=531 on the tag).

Two costs beyond the wasted run. It consumed one of the two self-hosted runners
while the release pipeline wanted both for its parallel multi-arch variant
builds; and it made release-run discovery ambiguous, since the newest-first runs
listing puts the tag-ref lint run above the publish run.

`branches: ['**']` keeps the documented intent exactly — lint fires early on
every branch push and PR, rather than only at tag time — while excluding tag
refs. docker-publish.yml is untouched and still tag-scoped.
2026-08-04 18:06:34 +02:00
pi 291ae5345e repo: add LICENSE, THIRD_PARTY.md, .dockerignore, hadolint lint, IDEAS backlog
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 16s
Repo/CI hygiene batch (none base-affecting; image contents unchanged):

- LICENSE: actual MIT file (repo previously declared MIT only in prose).
- THIRD_PARTY.md: notes bundled software + licenses (pi/pi-fork/pi-obsmem/
  pi-studio MIT, gosu Apache-2.0, Debian packages under their own terms).
- .dockerignore: trims build context to what the Dockerfiles COPY (rootfs/ +
  entrypoint*.sh); keeps .git/docs/scripts/compose out. Verified it excludes
  none of the required COPY sources.
- lint.yml: new hadolint job (pinned v2.14.0) lints both Dockerfiles;
  .hadolint.yaml grandfathers deliberate choices (DL3008/DL3016/DL4006/DL3003/
  SC2086, mirroring the shellcheck excludes), fails on anything new at warning+.
  Verified hadolint exit 0 and the repo shell-guard passes with the new job.
- IDEAS.md: parks deferred follow-ups (SHA-pin actions, trivy, buildx SBOM/
  provenance, Makefile, renovate).
- README/DOCKER_HUB License sections now link LICENSE + THIRD_PARTY.md.

No tag.
2026-07-13 18:20:44 +02:00
pi d1db595f17 ci(lint): pass explicit workflow paths to actionlint
Lint workflows / actionlint (push) Successful in 15s
actionlint's no-arg project auto-detection looks for .github/workflows
and hard-fails (exit 3, 'no project was found') on this .gitea/workflows
layout — observed on run 420. Glob the workflow files explicitly. The
Gitea shell guard step already passed in that run; only the actionlint
invocation needed the path fix.
2026-07-01 22:06:40 +02:00
pi 26384fe9f1 ci: eliminate the sh-vs-bash footgun class (defaults + lint guard)
Lint workflows / actionlint (push) Failing after 34s
Root cause of the recurring 'Illegal option -o pipefail' failures
(ed49b8d resolve-versions; b7197e8 promote-base-latest, run 418):
docker-publish.yml had no workflow-level default shell, so Gitea's
sh/dash default applied and every bash-syntax step had to individually
remember 'shell: bash'.

- docker-publish.yml: add 'defaults: run: shell: bash' — fixes the whole
  class; all pre-existing dash steps are POSIX so bash runs them unchanged.
- lint.yml: new workflow, runs on every push/PR (not just release tags):
    * scripts/check-workflow-shell.sh — Gitea-accurate guard: fails if any
      run: step doesn't resolve to bash. Catches the omit-shell+bash-syntax
      case that actionlint MISSES (actionlint models GitHub, where the
      default shell is bash, so a shell-less step is assumed bash).
    * actionlint + shellcheck — catches explicit 'shell: sh' + bash syntax
      (SC3040) and general workflow errors.
  Verified locally: guard + actionlint pass current workflows; guard fails
  a synthetic omit-shell+pipefail workflow; shellcheck clean.
2026-07-01 22:05:04 +02:00