b501120042
- task-pi-devbox-node22-pins.json: node-24 change-set. Result: Dockerfile.base:557 is the SOLE hard pin; the real find is that NO test asserts the node major (smoke-test.sh:94 is plain run(), not run_expect), so a bump passes silently. - task-ci-watcher-infra-vs-code.json: infra-vs-code CI failure. The delegate REJECTED the proposed zero-log-bytes heuristic on three false-positive grounds plus unreachability, and pointed at a better signal already fetched and unused (started_at -> RUN_START_TS, read only at watcher-hub-only.sh:266-267). Both specs assert only measured facts, including the retracted node-24/agent-browser claim, so the delegate could not resurrect it.
28 lines
3.2 KiB
JSON
28 lines
3.2 KiB
JSON
{
|
|
"id": "ci-watcher-infra-vs-code-failure",
|
|
"goal": "The ci-release-watcher templates poll a CI run and decide whether it succeeded. Determine how they currently distinguish an INFRASTRUCTURE failure (no runner ever picked the job up) from a CODE failure (the job ran and failed), then assess whether a proposed 'a completed run with zero job-log bytes means infrastructure, not code' check is sound and where exactly it belongs.",
|
|
"deliverable": "1) For EACH of the three templates, how it currently determines success/failure: file:line plus the exact API field, jq expression or exit code it reads. 2) A yes/no with pointer: today, can a run that was queued but NEVER started be distinguished from a run that started and failed? If the information is available in what the template already fetches but is unused, say so and point at it. 3) The exact insertion point (file:line) for the zero-log-bytes check and the precise condition to evaluate, expressed against the fields the template actually has in scope at that point. 4) AT LEAST THREE concrete scenarios where a completed run legitimately has zero or near-zero job-log bytes and the heuristic would therefore be a FALSE POSITIVE. Be adversarial here; do not simply endorse the proposal. 5) A recommendation on whether it should abort the watcher (hard failure) or only warn, justified by the false-positive rate you just described. Do NOT edit any file.",
|
|
"effort": "balanced",
|
|
"read_only": true,
|
|
"roots": ["/workspace/skillset"],
|
|
"context": {
|
|
"facts": [
|
|
"The skill lives at /workspace/skillset/skills/ci-release-watcher and contains exactly 6 files: SKILL.md, templates/watcher.sh, templates/watcher-hub-only.sh, templates/launch-release.sh, scripts/token-tmpfile.sh, scripts/ssh-control-master-setup.sh.",
|
|
"MEASURED: grep -rniE 'log.?byte|zero.?length|empty log|0 bytes|infrastructure' over that directory returns NO matches, so no form of this check exists today. You are specifying a new check, not finding an existing one.",
|
|
"The incident that motivates this (2026-09-06/07, Gitea): runner-2 was DISABLED in the Gitea admin UI. Jobs were queued but never fetched. The only symptom was repeated 'failed to fetch task' HTTP 500 responses in the runner's own journal on the runner host — invisible to anything polling the CI API. The runner host itself was verified healthy: 75 days uptime, load 0.00, Docker 29.5.0, correct registration (id=10, name=runner-2), and a 1.64 GB image pull completing in 32s.",
|
|
"This is a Gitea Actions deployment (Gitea's API is GitHub-Actions-shaped but NOT identical; do not assume a GitHub field exists without finding it used in these files).",
|
|
"/workspace/skillset is a git repository with a clean working tree."
|
|
],
|
|
"files": [
|
|
"/workspace/skillset/skills/ci-release-watcher/SKILL.md",
|
|
"/workspace/skillset/skills/ci-release-watcher/templates/watcher.sh",
|
|
"/workspace/skillset/skills/ci-release-watcher/templates/watcher-hub-only.sh",
|
|
"/workspace/skillset/skills/ci-release-watcher/templates/launch-release.sh"
|
|
],
|
|
"commands": [
|
|
"grep -n 'curl\\|jq\\|status\\|conclusion' /workspace/skillset/skills/ci-release-watcher/templates/watcher.sh"
|
|
]
|
|
},
|
|
"budget": {"wall_s": 700, "usd": 1.0}
|
|
}
|