pi-devbox-environment already warned that "a negative result is usually your
own filter" — baked, symlinked in at every container start, authored by an
earlier session. It survived every recreate, was available all of a later
session, and was violated five times anyway. So the gap was never persistence.
That section also closed with "a positive result needs no such scepticism — it
carries its own evidence." That is false, and it aimed the scepticism budget
one way only. Three of those five errors were positives:
- an SSH handshake SUCCEEDED and greeted me as joakimp while I believed I was
probing gitea.egl.lan — `Host gitea*` had rewritten HostName
- a 401 that was a genuine answer from an issuer which never minted the token
- a "regression" produced by diffing against a value my own -p 2222 flag set
Adds the three missing false-negative rows, replaces the false claim with the
"a positive result only proves what you actually asked" subsection, and records
the two habits that actually caught these: `ssh -G` to learn which rule
captured a hostname, and declaring the expected result before running a check.
The cross-cutting form goes in pi-global-AGENTS.append.md rather than in a
skill, because it has to fire without a task description matching it — being
loadable on demand is exactly what failed. Across all five errors, none was
caught by re-reading my reasoning; every one was caught by a second measurement
that disagreed.
6.4 KiB
Running inside pi-devbox
If the directory /usr/local/lib/pi-devbox/ exists (or your shell prompt is
prefixed [devbox], or ~/.ssh-local/config is present), you are in a
pi-devbox container — a Docker environment whose persistence, networking,
DNS, host/LAN reachability, tmux, and Python/REPL behaviour differ from a normal
workstation. Before any task that touches reaching the host or its LAN, SSH,
DNS/name resolution, what survives container recreate, running Python/REPLs,
tmux, or pi-studio, read ~/.agents/skills/pi-devbox-environment/SKILL.md.
Key reflex from that skill: the deployment specifics are not universal — the
host OS, hostnames, internal domains, and nameservers vary per instance and must
be discovered at runtime, never assumed. And interactive shell aliases
(dssh, dscp, cat→bat) do not exist in your non-interactive bash
tool, so spell out the underlying command (e.g.
ssh -F "$HOME/.ssh-local/config" mac …).
Browser automation is available (agent-browser)
This image bakes the agent-browser CLI plus a headless Chromium, so you can
drive a real browser — open pages, click/fill/eval, snapshot the DOM, take
screenshots — to verify front-end work (live DOM, WebGL, layout, popup
positioning) instead of guessing. Reach for it whenever a task involves a web UI
or checking how a page actually renders. AGENT_BROWSER_EXECUTABLE_PATH is
preset to the baked browser, so agent-browser open <url> works out of the box
(headless). Run agent-browser skills get core --full for the command set and
workflow patterns (always version-matched to the CLI); the agent-browser skill
under ~/.agents/skills/ mirrors it when the skillset is mounted.
Session start: load the mempalace skill
If MemPalace MCP tools (e.g. mempalace_search, mempalace_diary_write) are in
your tool list, read ~/.agents/skills/mempalace/SKILL.md before doing
non-trivial work and follow its protocol: search the palace before answering
about past work, and write a diary entry before the session ends. This is
especially load-bearing here — a pi-devbox container is frequently recreated, so
the palace is your only memory across recreates. Without the habit it is just
storage, not memory. (The skill is the consumer side; feeding the palace is the
separate opencode-mempalace-bridge skill, if present.)
If the palace is central, it is shared — three rules
If MEMPALACE_REMOTE_URL is set, the MCP tools write to a central palace
shared with other machines, not to a local one. Your drawers are not the only
ones in there, and most drawers' source_file paths do not exist on this host.
The skill covers the orientation side (provenance, chronology, whose diary is
whose); these three are here instead because getting them wrong does damage
rather than merely confusing you:
- Never run
mempalace sync/mempalace_syncagainst a shared palace. It prunes drawers whose source files look gitignored, deleted, or moved — and on a shared palace that describes most of the content, including every other machine's. Compounding it (RFC-001 §7.2): feeders now stage inside the palace root, so a scoped sync can delete the very drawers it just filed.mempalace_delete_by_sourceis exact-match rather than existence-based, but its blast radius is now the whole fleet's palace — leave it on its defaultdry_run=trueand confirm the match count before committing. - A timeout is not a failure. The palace is single-writer, and one large
mine can block every client for minutes, so a write or mine that exceeds the
client's deadline has usually completed server-side. Verify with
mempalace_get_drawerormempalace_searchbefore retrying — a blind retry files a duplicate.[mempalace ext] feed (tick) failed: mine timed out after 30000msis the common benign instance: the transcript is already in the server's inbox and the mine is idempotent, so nothing is lost either way. - The
mempalaceCLI is not remote-aware. It always opens a palace on local disk, somempalace searchcan return older and different results than the MCP tools while both look correct. Use the MCP tools for the central palace; the CLI only for a local one.
Before you file a finding: second measurement, different route
This is here rather than in a skill because it has to fire without a matching task description, and because the version of it that lived only in a skill was violated five times in one session by an agent that had the skill available.
Any claim you are about to record as fact — in a drawer, a diary entry, a coordination event, or a report to the user — needs a second measurement taken by a different route. Not a re-read of your reasoning: re-reading has caught zero of these. A disagreeing measurement has caught all of them.
The two shapes that get filed as fact and are not:
- A negative result (
401, connection refused, zero rows, "not found") is first a claim about your filter, not about the world. Wrong host, wrong port, wrong table, capped output. - A positive result proves only what your command actually asked. An SSH
handshake can succeed against the wrong host (
ssh -Gtells you which rule captured the name); a401can be a real answer from an issuer that never minted the credential.
Cheapest habit that works: write the expected result next to each check before running it, then diff. Expectations declared up front turn a silent wrong assumption into a visible mismatch. And if you cannot think of a second route to the same fact, you do not have a finding — you have a hypothesis, so label it as one.
Handling an exposed credential
If a task touches a leaked secret, a token rotation, "is this credential still
live?", whether to delete stored content, or which scopes a new token needs:
read ~/.agents/skills/credential-incident-response/SKILL.md first. One rule
is load-bearing enough to state here: probe the issuing provider before doing
anything else — most "exposed" credentials in a long-lived fleet are already
dead, and the ones that are live are often far more privileged than assumed.
Severity first, cleanup second, and prefer revocation over deletion for
anything already replicated.