skills: fix the half of the negative-result rule that was wrong
pi-devbox-environment already warned that "a negative result is usually your
own filter" — baked, symlinked in at every container start, authored by an
earlier session. It survived every recreate, was available all of a later
session, and was violated five times anyway. So the gap was never persistence.
That section also closed with "a positive result needs no such scepticism — it
carries its own evidence." That is false, and it aimed the scepticism budget
one way only. Three of those five errors were positives:
- an SSH handshake SUCCEEDED and greeted me as joakimp while I believed I was
probing gitea.egl.lan — `Host gitea*` had rewritten HostName
- a 401 that was a genuine answer from an issuer which never minted the token
- a "regression" produced by diffing against a value my own -p 2222 flag set
Adds the three missing false-negative rows, replaces the false claim with the
"a positive result only proves what you actually asked" subsection, and records
the two habits that actually caught these: `ssh -G` to learn which rule
captured a hostname, and declaring the expected result before running a check.
The cross-cutting form goes in pi-global-AGENTS.append.md rather than in a
skill, because it has to fire without a task description matching it — being
loadable on demand is exactly what failed. Across all five errors, none was
caught by re-reading my reasoning; every one was caught by a second measurement
that disagreed.
This commit is contained in:
@@ -70,3 +70,41 @@ rather than merely confusing you:
|
|||||||
local disk, so `mempalace search` can return older and different results than
|
local disk, so `mempalace search` can return older and different results than
|
||||||
the MCP tools while both look correct. Use the MCP tools for the central
|
the MCP tools while both look correct. Use the MCP tools for the central
|
||||||
palace; the CLI only for a local one.
|
palace; the CLI only for a local one.
|
||||||
|
|
||||||
|
## Before you file a finding: second measurement, different route
|
||||||
|
|
||||||
|
This is here rather than in a skill because it has to fire *without* a matching
|
||||||
|
task description, and because the version of it that lived only in a skill was
|
||||||
|
violated five times in one session by an agent that had the skill available.
|
||||||
|
|
||||||
|
**Any claim you are about to record as fact — in a drawer, a diary entry, a
|
||||||
|
coordination event, or a report to the user — needs a second measurement taken
|
||||||
|
by a different route.** Not a re-read of your reasoning: re-reading has caught
|
||||||
|
zero of these. A disagreeing measurement has caught all of them.
|
||||||
|
|
||||||
|
The two shapes that get filed as fact and are not:
|
||||||
|
|
||||||
|
- **A negative result** (`401`, connection refused, zero rows, "not found") is
|
||||||
|
first a claim about *your filter*, not about the world. Wrong host, wrong port,
|
||||||
|
wrong table, capped output.
|
||||||
|
- **A positive result** proves only what your command *actually asked*. An SSH
|
||||||
|
handshake can succeed against the wrong host (`ssh -G` tells you which rule
|
||||||
|
captured the name); a `401` can be a real answer from an issuer that never
|
||||||
|
minted the credential.
|
||||||
|
|
||||||
|
Cheapest habit that works: **write the expected result next to each check before
|
||||||
|
running it**, then diff. Expectations declared up front turn a silent wrong
|
||||||
|
assumption into a visible mismatch. And if you cannot think of a second route to
|
||||||
|
the same fact, you do not have a finding — you have a hypothesis, so label it as
|
||||||
|
one.
|
||||||
|
|
||||||
|
## Handling an exposed credential
|
||||||
|
|
||||||
|
If a task touches a leaked secret, a token rotation, "is this credential still
|
||||||
|
live?", whether to delete stored content, or which scopes a new token needs:
|
||||||
|
**read `~/.agents/skills/credential-incident-response/SKILL.md` first.** One rule
|
||||||
|
is load-bearing enough to state here: **probe the issuing provider before doing
|
||||||
|
anything else** — most "exposed" credentials in a long-lived fleet are already
|
||||||
|
dead, and the ones that are live are often far more privileged than assumed.
|
||||||
|
Severity first, cleanup second, and prefer **revocation over deletion** for
|
||||||
|
anything already replicated.
|
||||||
|
|||||||
@@ -143,6 +143,9 @@ mine:
|
|||||||
| "`tor-ms22` is not in the SSH config" | `grep … \| head -20` — the entry was at **line 454**. `~/.ssh/config` here is ~500 lines. |
|
| "`tor-ms22` is not in the SSH config" | `grep … \| head -20` — the entry was at **line 454**. `~/.ssh/config` here is ~500 lines. |
|
||||||
| "the Docker host has no `docker`" | non-interactive SSH `PATH` lacks `/usr/local/bin` (§2, §3). It was at `/usr/local/bin/docker`. |
|
| "the Docker host has no `docker`" | non-interactive SSH `PATH` lacks `/usr/local/bin` (§2, §3). It was at `/usr/local/bin/docker`. |
|
||||||
| "no ControlMaster is running" | pattern `ssh ` (trailing space) cannot match a master: those processes **rename themselves** to `ssh: <controlpath> [mux]`. |
|
| "no ControlMaster is running" | pattern `ssh ` (trailing space) cannot match a master: those processes **rename themselves** to `ssh: <controlpath> [mux]`. |
|
||||||
|
| "the credential is not in the palace" | scanned `embedding_metadata.string_value` only. Drawer **text** lives in `embedding_fulltext_search_content.c0`; 554k metadata rows proved nothing. |
|
||||||
|
| "this token is dead — 401" | probed it against the **wrong issuer**. A 401 from an instance that never issued the credential is not evidence about the credential. |
|
||||||
|
| "that host is unreachable, can't test" | tried ports 443 and 80. It was on **3000**, and the env var I already held (`GITEA_EGL_HOST`) stated the scheme and port. |
|
||||||
|
|
||||||
Habits that would have caught all three:
|
Habits that would have caught all three:
|
||||||
|
|
||||||
@@ -157,8 +160,48 @@ ssh -F "$HOME/.ssh-local/config" mac 'command -v docker || ls /usr/local/bin/doc
|
|||||||
ps -eo pid,etime,args | grep -Ei 'mux|mosh|ssh'
|
ps -eo pid,etime,args | grep -Ei 'mux|mosh|ssh'
|
||||||
```
|
```
|
||||||
|
|
||||||
A positive result needs no such scepticism — it carries its own evidence. Only
|
Absence has to be *earned*, so spend the extra command there.
|
||||||
absence has to be *earned*, so spend the extra command there.
|
|
||||||
|
### …and a positive result only proves what you *actually asked*
|
||||||
|
|
||||||
|
An earlier version of this section claimed "a positive result needs no such
|
||||||
|
scepticism — it carries its own evidence." **That is false, and believing it
|
||||||
|
cost a later session three more wrong findings.** A positive result is evidence
|
||||||
|
about the question your command really posed, which may not be the question you
|
||||||
|
meant. The failure is invisible precisely *because* the command succeeded.
|
||||||
|
|
||||||
|
| Claim | The command succeeded — at answering something else |
|
||||||
|
|---|---|
|
||||||
|
| "EGL git over SSH works" | `ssh git@gitea.egl.lan` greeted me as `joakimp`. `~/.ssh/config` had `Host gitea*` → `HostName gitea.jordbo.se`, so I authenticated **to the wrong instance**. The real EGL account is `ecsjper`. |
|
||||||
|
| "the port config regressed" | compared `ssh -G` output against `2222` — a value produced by **my own earlier `-p 2222` flag**, not by the config. I reported the user's edit as a regression it never caused. |
|
||||||
|
| "the CI runners authenticate with this token" | pure fabrication, contradicted by my own scan output already on screen. The runners use per-runner `REGISTRATION_TOKEN`. |
|
||||||
|
|
||||||
|
Two habits that actually catch this class, both cheap:
|
||||||
|
|
||||||
|
```sh
|
||||||
|
# 1. ask which RULE captured your hostname before trusting any ssh result.
|
||||||
|
# ssh_config is first-obtained-value-wins PER KEYWORD, not per block: a
|
||||||
|
# specific block only wins the keywords it declares, so a later `Host gitea*`
|
||||||
|
# still supplies HostName unless the specific block restates it.
|
||||||
|
ssh -G git@thehost | grep -E '^(hostname|port|user|identityfile)'
|
||||||
|
|
||||||
|
# 2. state the expected result BEFORE running the check, and diff against it.
|
||||||
|
# This is the single technique that separated the one verification that went
|
||||||
|
# right (10/10, expectations declared per probe) from five that went wrong
|
||||||
|
# (results interpreted after the fact, each time in the direction I expected).
|
||||||
|
probe "/repos/.../actions/runs" 200 # must work
|
||||||
|
probe "/admin/users" 403 # must be denied
|
||||||
|
```
|
||||||
|
|
||||||
|
And the meta-observation, which is the reason this subsection exists: across all
|
||||||
|
five errors, **not one was caught by re-reading my own reasoning.** Every one was
|
||||||
|
caught by a second measurement that disagreed — the SSH lie surfaced only because
|
||||||
|
the greeting said `joakimp` while a token probe minutes earlier had said
|
||||||
|
`ecsjper`; the fabrication surfaced only because the user read my own output back
|
||||||
|
to me. So the operational rule is not "be careful". It is: **for a load-bearing
|
||||||
|
claim, produce a second measurement by a different route, and expect it to
|
||||||
|
disagree.** If you cannot think of a second route, you do not yet have a finding
|
||||||
|
— you have a hypothesis.
|
||||||
|
|
||||||
**`dscp`/`scp` with accented filenames on a macOS host.** macOS stores filenames
|
**`dscp`/`scp` with accented filenames on a macOS host.** macOS stores filenames
|
||||||
in Unicode **NFD** (decomposed — e.g. `ä` is `a` + combining U+0308), while the
|
in Unicode **NFD** (decomposed — e.g. `ä` is `a` + combining U+0308), while the
|
||||||
|
|||||||
Reference in New Issue
Block a user