Commit Graph

3 Commits

Author SHA1 Message Date
Joakim Persson 9d0b3dec0b release(v1.9.5): pi 0.87.1 with pi-obsmem pinned to the merged finishTurn fix
Lint / skill-floor (push) Successful in 7s
Lint / hadolint (push) Successful in 14s
Lint / actionlint (push) Successful in 19s
Lint / doc-drift (push) Successful in 12s
v1.9.4 held pi at 0.85.1 because 0.87.0 REMOVED `shouldStopAfterTurn`, which
pi-observational-memory 3.1.4 still used in all three workers. Its peerDeps are
`*`, so nothing refuses at install time and the breakage is silent at runtime --
turn caps ignored, workers losing their specialised prompts. PR #83 fixes that
and merged 2026-09-23.

Re-measured 2026-10-01 across src/agents/{observer,reflector,dropper}/agent.ts,
deliberately reusing the v1.9.4 audit's counting method so the numbers compare:

  e7d77dc (3.1.4, baked in v1.9.4): shouldStopAfterTurn 3, finishTurn 0, systemPrompt 3
  731c3d4 (pinned here)           : shouldStopAfterTurn 0, finishTurn 3, systemPrompt 0

All three workers migrated, and the 0.86.0 AgentContext.systemPrompt reads are
gone too. The coupling is asymmetric and that is why these move in ONE commit:
3.1.4 + 0.87.1 silently ignores turn caps, and 731c3d4 + 0.85.1 breaks the
workers outright, because finishTurn does not exist before 0.87.0.

Pinned to a SHA rather than waiting for a tag, departing from the v1.9.4
instruction to wait for a release: #83 is merged but the newest obsmem tag is
still 3.1.4, cut 2026-09-20, BEFORE the merge. Upstream tags slowly and moves
master often (e7d77dc -> 1529e14 -> 731c3d4 in nine days), so waiting means
holding pi indefinitely. A pinned SHA keeps the property `master` lacks:
rebuilding this tag later produces the same image.

The 40-char form is load-bearing, not pedantry. check-doc-drift.sh recognises a
literal SHA only through a 40-char match, so a 7-char pin would fall through to
its branch-or-tag lookup, fail to resolve, and downgrade pi-obsmem's drift check
to a silent SKIP -- a pin that reads correctly and is no longer verified. Full
gate run after this change: 23 OK, 0 DRIFT, 0 SKIP, 0 FAIL.

0.99.0/0.99.1/0.99.2 and 1.0.0 all exist upstream and are deliberately skipped:
obsmem's only compatibility work names Pi 0.87 (2b1dc1c) and a repo-wide
issue/PR search for 0.99 or 1.0 returns zero matches. 1.0.0 is its own round.

pi-atelier stays at v0.10.3 and is the residual risk. Its peerDeps declare pi
>=0.84.0 -- a floor, satisfied -- on both v0.10.3 and the current v0.12.1, so it
spans this bump. But atelier hooks pi TUI internals that a declared floor does
not protect, and an under-declared floor is exactly what failed to warn anyone
at pi 0.84. Acceptance must confirm the sidebar PAINTS, using the two-sided
check from 0.84.4/0.85.1 that distinguishes "loaded" from "silently absent".

Also here:
- check-doc-drift.sh gains a pi-obsmem pin check. The ref-move check covers the
  same component, but for a pinned SHA it can only answer "upstream did not
  move", never "the table still says what we bake".
- scripts/lint-shell.sh mode 100644 -> 100755. Pre-existing since f25efa0 and
  the only non-executable script in scripts/; it was latent because both CI
  steps call it as `bash scripts/lint-shell.sh`, but it failed rc=126 "bad
  interpreter" when invoked directly. Same dropped-exec-bit signature recorded
  on 2026-09-22, found the same way: by RUNNING it, not by reading a diff.
- Unreleased section renamed to `## v1.9.5 — 2026-10-02`, satisfying the
  release-gate rule that a tag's CHANGELOG must name its own version.
2026-10-02 00:23:28 +02:00
Joakim Persson f25efa074d fix(sanity): verify the ControlPath ssh RESOLVES, not the dir the image creates
Lint / hadolint (push) Successful in 9s
Lint / skill-floor (push) Successful in 13s
Lint / doc-drift (push) Successful in 12s
Lint / actionlint (push) Successful in 24s
recreate-sanity-check.sh asserted `/tmp/sshcm exists with mode 700` and printed
a green tick while every ssh in the container died rc=255. It was right about
what it checked: the breakage was in the directory a CONFIG named, not the one
the image creates, and the old check could not see the disagreement between them.

Found on the v1.9.2 first boot on emb-7kj4vr4g. A durable ~/.pi/ssh/config,
hand-written into the ~/.pi named volume by the previous session so it would
survive the recreate, declared `ControlPath /tmp/ssh-cm/%C` — with a hyphen.
Nothing here creates that path; the canonical spelling is /tmp/sshcm, identical
in Dockerfile.base, entrypoint-user.sh, recreate-sanity-check.sh and
smoke-test.sh. ControlMaster auto with an unusable ControlPath does not degrade
to an unmultiplexed connection — it fails hard:

    unix_listener: cannot bind to path /tmp/ssh-cm/<hash>: No such file or directory

rc=255, remote command never runs.

Now: resolve the ControlPath ssh itself would use via `ssh -G` and require its
parent to exist and be writable. -G applies real config precedence (first-value-
wins, the system drop-in, Include, an -F override), so it answers "which rule
captured this host" instead of re-implementing the guess, and it never opens a
connection — 0.116s for 48 hosts.

This also puts a check under a caveat documented in prose in Dockerfile.base
("SSH client defaults") and verified nowhere: a per-host ControlPath under a
read-only bind-mounted ~/.ssh gives the identical failure, "Read-only file
system". On the machine this was built on that is 16 of 50 hosts — freeipa-1..6,
gitea.egl.lan, runner-1..3, tor-ms22 — never reported by anything before.

Severity split is deliberate. Default ssh precedence legitimately lands in the
read-only ~/.ssh for any host whose own config pins it there, and the supported
workaround (ssh -F ~/.ssh-local/config, generated every start by
setup-lan-access.sh) already exists, so that is a warn. Failing it would paint
the script red on every run of every device, and a check that fires benignly
every time is one you learn to ignore — the same reasoning that keeps
lint-shell.sh at -S error. The sidecar route is prescribed, so there it is a
hard fail. Host lists cap at six names plus a count: unreadable output is
ignored output.

Teeth proven both directions, each sabotage confirmed by diff BEFORE the result
was believed: ControlPath -> nonexistent dir => rc=1; -> existing-but-read-only
dir => rc=1; restored => rc=0 with the sidecar byte-identical.

No config was added to fix the original problem — the fix was to DELETE
~/.pi/ssh/config, a third hand-maintained copy of what setup-lan-access.sh
already generates from version control on every container start, with fewer
features and one typo. The host-owned ~/.ssh/config is left alone on purpose:
those ~/.ssh/cm paths are correct on the host, where ~/.ssh is writable.

Also lint-shell.sh: the SC2088 count in the severity-choice rationale said 19
and is now 20, with the reproduce command recorded. That command needs a `$ `
prefix — a comment whose first word is "shellcheck" is parsed as a directive,
and the malformed one tripped SC1072/SC1073 at severity error. The gate caught
it on the very commit that introduced it.
2026-09-15 22:59:43 +02:00
joakimp 361babd4fd ci: gate the release on shell lint, from one shared script
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / lint-gate (push) Successful in 15s
Publish Docker Image / resolve-versions (push) Successful in 9s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Has been skipped
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Publish Docker Image / smoke (push) Successful in 7m39s
Publish Docker Image / build-variant-studio (push) Successful in 17m15s
Publish Docker Image / build-variant (push) Successful in 17m58s
Publish Docker Image / update-description (push) Successful in 8s
Publish Docker Image / promote-base-latest (push) Successful in 12s
v1.8.14's first attempt spent ~46 minutes building a base image for a tree whose
own lint had been failing for 24 hours. shellcheck had already flagged the
defect (SC2289, severity error) on the push that introduced it; the lint
workflow went red at run 186 and nobody read it.

lint.yml deliberately skips tag pushes and its reasoning is sound -- the tagged
tree was already linted on main, and a tag-ref lint run sorts above the publish
run, making a release look finished before anything ships. The missing invariant
was never "lint the tag". It was "do not RELEASE a tree whose lint failed", and
only a job inside the publish workflow can enforce that.

So: extract the shell-lint logic from lint.yml into scripts/lint-shell.sh and
call it from both places, then add a lint-gate job that resolve-versions depends
on. resolve-versions is the graph root, so gating it gates everything. Cost is
~40 s at the front of a release; the alternative already cost fifty minutes.

Extracted rather than copied on purpose. A second copy of a check is the drift
this repo keeps paying for -- the same evening produced a skillset mirror that
had sat 9579 B behind its upstream through two consecutive edits.

The script adds one behaviour the inline version lacked: if shellcheck is not
installed it exits 2 rather than silently finding nothing, inheriting the
existing "a gate that cannot run must not pass" rule from hooks/pre-commit in
the skillset repo. Without that, reordering the install step away would turn the
gate into a green tick over zero checks.

Verified locally with a stubbed shellcheck (the real binary is not in the
devbox), five cases, each with its expectation stated first: absent shellcheck
-> rc=2; stub pass -> rc=0 and a non-zero file count; stub fail -> rc=1; a
deliberately unterminated `if` planted in scripts/ -> rc=1 via the bash -n half,
naming the file; removal -> rc=0 again. Discovery cross-checks against CI's own
number: the inline version reported 12 files, the extracted one reports 13, the
difference being lint-shell.sh itself. YAML re-parsed (10 jobs, was 9) with an
assertion that resolve-versions needs lint-gate, and the repo's
check-workflow-shell.sh guard still passes.
2026-09-08 23:41:44 +02:00