Files
pi-devbox/rootfs/usr/local/share/pi-devbox/pi-global-AGENTS.append.md
T
joakimp f0ebea2d98 skills: fix the half of the negative-result rule that was wrong
pi-devbox-environment already warned that "a negative result is usually your
own filter" — baked, symlinked in at every container start, authored by an
earlier session. It survived every recreate, was available all of a later
session, and was violated five times anyway. So the gap was never persistence.

That section also closed with "a positive result needs no such scepticism — it
carries its own evidence." That is false, and it aimed the scepticism budget
one way only. Three of those five errors were positives:

  - an SSH handshake SUCCEEDED and greeted me as joakimp while I believed I was
    probing gitea.egl.lan — `Host gitea*` had rewritten HostName
  - a 401 that was a genuine answer from an issuer which never minted the token
  - a "regression" produced by diffing against a value my own -p 2222 flag set

Adds the three missing false-negative rows, replaces the false claim with the
"a positive result only proves what you actually asked" subsection, and records
the two habits that actually caught these: `ssh -G` to learn which rule
captured a hostname, and declaring the expected result before running a check.

The cross-cutting form goes in pi-global-AGENTS.append.md rather than in a
skill, because it has to fire without a task description matching it — being
loadable on demand is exactly what failed. Across all five errors, none was
caught by re-reading my reasoning; every one was caught by a second measurement
that disagreed.
2026-08-30 00:50:11 +02:00

6.4 KiB

Running inside pi-devbox

If the directory /usr/local/lib/pi-devbox/ exists (or your shell prompt is prefixed [devbox], or ~/.ssh-local/config is present), you are in a pi-devbox container — a Docker environment whose persistence, networking, DNS, host/LAN reachability, tmux, and Python/REPL behaviour differ from a normal workstation. Before any task that touches reaching the host or its LAN, SSH, DNS/name resolution, what survives container recreate, running Python/REPLs, tmux, or pi-studio, read ~/.agents/skills/pi-devbox-environment/SKILL.md.

Key reflex from that skill: the deployment specifics are not universal — the host OS, hostnames, internal domains, and nameservers vary per instance and must be discovered at runtime, never assumed. And interactive shell aliases (dssh, dscp, cat→bat) do not exist in your non-interactive bash tool, so spell out the underlying command (e.g. ssh -F "$HOME/.ssh-local/config" mac …).

Browser automation is available (agent-browser)

This image bakes the agent-browser CLI plus a headless Chromium, so you can drive a real browser — open pages, click/fill/eval, snapshot the DOM, take screenshots — to verify front-end work (live DOM, WebGL, layout, popup positioning) instead of guessing. Reach for it whenever a task involves a web UI or checking how a page actually renders. AGENT_BROWSER_EXECUTABLE_PATH is preset to the baked browser, so agent-browser open <url> works out of the box (headless). Run agent-browser skills get core --full for the command set and workflow patterns (always version-matched to the CLI); the agent-browser skill under ~/.agents/skills/ mirrors it when the skillset is mounted.

Session start: load the mempalace skill

If MemPalace MCP tools (e.g. mempalace_search, mempalace_diary_write) are in your tool list, read ~/.agents/skills/mempalace/SKILL.md before doing non-trivial work and follow its protocol: search the palace before answering about past work, and write a diary entry before the session ends. This is especially load-bearing here — a pi-devbox container is frequently recreated, so the palace is your only memory across recreates. Without the habit it is just storage, not memory. (The skill is the consumer side; feeding the palace is the separate opencode-mempalace-bridge skill, if present.)

If the palace is central, it is shared — three rules

If MEMPALACE_REMOTE_URL is set, the MCP tools write to a central palace shared with other machines, not to a local one. Your drawers are not the only ones in there, and most drawers' source_file paths do not exist on this host. The skill covers the orientation side (provenance, chronology, whose diary is whose); these three are here instead because getting them wrong does damage rather than merely confusing you:

  • Never run mempalace sync / mempalace_sync against a shared palace. It prunes drawers whose source files look gitignored, deleted, or moved — and on a shared palace that describes most of the content, including every other machine's. Compounding it (RFC-001 §7.2): feeders now stage inside the palace root, so a scoped sync can delete the very drawers it just filed. mempalace_delete_by_source is exact-match rather than existence-based, but its blast radius is now the whole fleet's palace — leave it on its default dry_run=true and confirm the match count before committing.
  • A timeout is not a failure. The palace is single-writer, and one large mine can block every client for minutes, so a write or mine that exceeds the client's deadline has usually completed server-side. Verify with mempalace_get_drawer or mempalace_search before retrying — a blind retry files a duplicate. [mempalace ext] feed (tick) failed: mine timed out after 30000ms is the common benign instance: the transcript is already in the server's inbox and the mine is idempotent, so nothing is lost either way.
  • The mempalace CLI is not remote-aware. It always opens a palace on local disk, so mempalace search can return older and different results than the MCP tools while both look correct. Use the MCP tools for the central palace; the CLI only for a local one.

Before you file a finding: second measurement, different route

This is here rather than in a skill because it has to fire without a matching task description, and because the version of it that lived only in a skill was violated five times in one session by an agent that had the skill available.

Any claim you are about to record as fact — in a drawer, a diary entry, a coordination event, or a report to the user — needs a second measurement taken by a different route. Not a re-read of your reasoning: re-reading has caught zero of these. A disagreeing measurement has caught all of them.

The two shapes that get filed as fact and are not:

  • A negative result (401, connection refused, zero rows, "not found") is first a claim about your filter, not about the world. Wrong host, wrong port, wrong table, capped output.
  • A positive result proves only what your command actually asked. An SSH handshake can succeed against the wrong host (ssh -G tells you which rule captured the name); a 401 can be a real answer from an issuer that never minted the credential.

Cheapest habit that works: write the expected result next to each check before running it, then diff. Expectations declared up front turn a silent wrong assumption into a visible mismatch. And if you cannot think of a second route to the same fact, you do not have a finding — you have a hypothesis, so label it as one.

Handling an exposed credential

If a task touches a leaked secret, a token rotation, "is this credential still live?", whether to delete stored content, or which scopes a new token needs: read ~/.agents/skills/credential-incident-response/SKILL.md first. One rule is load-bearing enough to state here: probe the issuing provider before doing anything else — most "exposed" credentials in a long-lived fleet are already dead, and the ones that are live are often far more privileged than assumed. Severity first, cleanup second, and prefer revocation over deletion for anything already replicated.