skills: the fingerprint advice was missing its precondition, and the skill had no section on proving absence
Docs only; no image behaviour changes. WHY THIS AND NOT A PRIVATE NOTE. pi@emb-7kj4vr4g reported itself for printing sha256[:8] fingerprints of GIT_USER_EMAIL, GIT_USER_NAME and HOST_SSH_USER, and wrote a private rule forbidding it. It had not broken a rule. It followed §2 of this skill as written, and §2 is incomplete: it says a fingerprint lets you compare a credential "without ever materialising the secret" with no condition attached. When two agents independently make the same mistake, the artifact that taught them both is the bug. §2 NOW CARRIES THE PRECONDITION. A fingerprint is 32 bits over its INPUT SPACE, so publishing fp8(x) hands anyone a MEMBERSHIP ORACLE: they can test x == v for every candidate v they can generate. Safe for a 40-char random token; a wordlist for a hostname, username, e-mail, port, path, commit SHA or weak password. "High entropy" is the usual sufficient condition, NOT the test — a commit SHA is 160-bit and still fully enumerable from the repo. Operationally: if you can imagine writing the wordlist, you cannot publish the fingerprint. Also added: candidate fingerprints are working memory and never output (an extractor hashes hostnames and paths too, so the tempting "print what the scanner saw" debug step leaks low-entropy fingerprints wholesale), and a plain statement that a fingerprint register is a CONFIRMATION ORACLE for anyone already holding a candidate corpus — which is exactly how a retired token is identified in old transcripts, and works identically for someone else holding those same files. NEW §6, "Proving absence: instrument strength, and four ways a scan lies clean", placed next to §5 on purpose: §5 optimises against false POSITIVES, and every failure in §6 is a false NEGATIVE. Triage optimises precision, a gate optimises recall, and conflating them is what produced three clean reports over secrets that were really there. Contents: instrument ranking (exact-byte value search > class/structure pass > fingerprint census) with the instruction to state which one produced your zero; census vs class passes as different questions, both failure modes measured on this fleet; the tokenisation trap where quoting alone decided detectability; scan the index or pushed tree, never the working tree; git filters never run on symlinks while check-attr claims they do; two-sided self-tests that abort, incl. the fixture-interaction artifact; row-gone is not bytes-gone. Attribution kept per finding: the census/class split and the instrument ranking's provenance are pi@emb-7kj4vr4g's; exact-byte search over index blobs is pi@tor-ms22's. The credential sense of "census" originated in this skill, not with either agent. TRAP FOR THE NEXT EDITOR, also in the CHANGELOG: the frontmatter description is now 1022 of 1024 characters. Trim before adding, or the skill silently fails to load. Verified by parsing the frontmatter (1022 chars, name intact, every prior trigger phrase retained). Deployment: baked skill -> needs an image rebuild AND a container recreate to reach a running container.
This commit is contained in:
@@ -13,6 +13,65 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
|
||||
|
||||
## Unreleased
|
||||
|
||||
**`credential-incident-response` gained the section its own guidance had been
|
||||
missing, and §2 gained a precondition it should always have carried.** Docs only;
|
||||
no image behaviour moves. Both changes came out of a session where three separate
|
||||
detectors reported *clean* over secrets that were really there — the skill was
|
||||
the artifact that had taught two agents the pattern, so the fix belongs here
|
||||
rather than in either operator's private notes.
|
||||
|
||||
**§2 previously said an 8-hex fingerprint lets you compare a credential "without
|
||||
ever materialising the secret", with no condition attached.** That is true only
|
||||
when the *input space* is unreachable. A fingerprint is 32 bits over whatever it
|
||||
was computed from, so publishing `fp8(x)` hands anyone a **membership oracle**:
|
||||
they can test `x == v` for every candidate `v` they can generate. For a 40-char
|
||||
random token, fine. For a hostname, username, e-mail, port, path, commit SHA or
|
||||
weak password, that candidate set is a wordlist — and note that "high entropy" is
|
||||
the usual sufficient condition, not the test: a commit SHA is 160-bit and still
|
||||
fully enumerable from the repo. Two agents on this fleet published fingerprints of
|
||||
`GIT_USER_EMAIL`-class values while following this section as written; harmless in
|
||||
that instance, because those values sit in every commit trailer already, but the
|
||||
guidance licensed it. §2 now states the precondition, adds that candidate
|
||||
fingerprints are working memory and never output (a scanner hashes hostnames and
|
||||
paths too, so "print what it saw" leaks wholesale), and names what a fingerprint
|
||||
register *is* — a confirmation oracle for anyone already holding a candidate
|
||||
corpus, which is exactly how a retired token gets identified in old transcripts,
|
||||
and works the same way for someone else holding those files.
|
||||
|
||||
**New §6, "Proving absence: instrument strength, and four ways a scan lies
|
||||
clean".** Deliberately placed next to §5, because §5 optimises against false
|
||||
*positives* (name-anchoring, provenance — what stops a triage sweep drowning in
|
||||
session UUIDs) and every failure in §6 is a false *negative*. Triage optimises
|
||||
precision; a gate optimises recall, and conflating the two is what produced the
|
||||
clean reports. It carries: an instrument-strength ranking (exact-byte value search
|
||||
> class/structure pass > fingerprint census) with the standing instruction to say
|
||||
which one produced your zero; census and class passes answering different
|
||||
questions, with both failure modes measured here — a class-only pre-commit hook
|
||||
passed plaintext UUID API credentials to a shared repo twice because a UUID has no
|
||||
key header, while a census-only gate reported 0 hits with freshly-synced SSH
|
||||
private keys in the tree because no key is in the census; the tokenisation trap,
|
||||
where maximal-run extraction swallows an unquoted `VAR=<uuid>` so the value is
|
||||
never hashed alone while a *quoted* one is found, meaning quoting alone decided
|
||||
detectability; scan the index or the pushed tree, never the working tree, plus why
|
||||
a repo-only fix on an rsync-published mirror is temporary rather than weaker; git
|
||||
filters never running on symlinks, where `check-attr` answers `git-crypt` for a
|
||||
path it can never encrypt, so a coverage audit must join the attribute against the
|
||||
file mode and verify the blob magic; two-sided self-tests that abort, including
|
||||
the fixture-interaction artifact where a quoted and unquoted probe share one
|
||||
buffer and make the weak extractor look as strong as the union; and row-gone is
|
||||
not bytes-gone, since a correct sqlite DELETE leaves the payload in freelist pages
|
||||
until VACUUM.
|
||||
|
||||
Findings contributed by `pi@emb-7kj4vr4g` (the census/class split, and the
|
||||
instrument ranking's provenance) and `pi@tor-ms22` (exact-byte value search over
|
||||
index blobs). The description's trigger list grew accordingly and is 1022/1024
|
||||
characters — **it has almost no headroom, so trim before adding to it**, or the
|
||||
skill silently fails to load.
|
||||
|
||||
**Deployment:** the skill is baked at
|
||||
`/usr/local/share/pi-devbox/skills/credential-incident-response/`, so this needs
|
||||
an image rebuild **and** a container recreate to reach any running container.
|
||||
|
||||
**Two vendored skills changed, and one of the changes is a correction rather than
|
||||
an addition.** Nothing about the image's behaviour moves; this is entirely about
|
||||
what the next agent reads before it acts.
|
||||
|
||||
@@ -1,18 +1,18 @@
|
||||
---
|
||||
name: credential-incident-response
|
||||
description: >-
|
||||
Respond correctly when a live credential is found somewhere it should not be —
|
||||
Respond correctly when a live credential is found where it should not be —
|
||||
in a chat transcript, a MemPalace drawer, a log, a git-tracked config, or an
|
||||
agent-authored note. Load this whenever a task involves a leaked/exposed
|
||||
secret, a token rotation, a "is this credential still live?" question, deciding
|
||||
whether to delete or scrub stored content, or choosing scopes for a new API
|
||||
token. Covers the mandatory order of operations (probe the issuer FIRST —
|
||||
severity before cleanliness), leak-free credential identity via sha256[:8]
|
||||
fingerprints, why revocation beats deletion for anything already replicated,
|
||||
deriving least-privilege scopes from measured consumers instead of guessing,
|
||||
where this fleet's secrets actually live (age-encrypted .env.age in
|
||||
docker-compose-repo, plus gitignored plaintext .env drift), the three places a
|
||||
secret hides in a Chroma palace, and the exposures that rotation does NOT fix.
|
||||
whether to delete or scrub stored content, proving a corpus is clean, or
|
||||
choosing scopes for a new API token. Covers the mandatory order of operations
|
||||
(probe the issuer FIRST — severity before cleanliness), leak-free identity via
|
||||
sha256[:8] fingerprints and when publishing one is safe,
|
||||
why revocation beats deletion for anything already replicated, scopes derived
|
||||
from measured consumers, the three places a secret hides in a Chroma palace, how to prove ABSENCE rather than assume it (instrument strength,
|
||||
census vs class passes, the tokenisation trap where quoting decides detectability, why git filters never run on symlinks, self-tests that abort),
|
||||
where this fleet's secrets live, and what rotation does NOT fix.
|
||||
---
|
||||
|
||||
# Credential incident response
|
||||
@@ -54,6 +54,29 @@ shared credential; that is usually the important part. **Never** paste a live
|
||||
value into a search query, a palace drawer, an event body, or a chat message —
|
||||
in an agent context your own tool output is itself captured and re-filed.
|
||||
|
||||
**Precondition — only fingerprint what an adversary cannot enumerate.** An 8-hex
|
||||
fingerprint is 32 bits over its *input space*, so publishing `fp8(x)` hands
|
||||
anyone a **membership oracle**: they can test `x == v` for every candidate `v`
|
||||
they can generate. For a 40-char random token that space is unreachable. For a
|
||||
hostname, username, e-mail, port, path, commit SHA or weak password it is a
|
||||
wordlist. **If you can imagine writing the wordlist, you cannot publish the
|
||||
fingerprint** — reference those by name and location instead. "High entropy" is
|
||||
the usual *sufficient condition*, not the test: a commit SHA is 160-bit and still
|
||||
fully enumerable from the repo. `sha256("")` = `e3b0c442` is the degenerate case,
|
||||
recognisable on sight precisely because its input space has one member.
|
||||
|
||||
**Candidate fingerprints are working memory, never output.** A scanner that hashes
|
||||
every token in a file also hashes hostnames, paths and e-mails. Print only
|
||||
fingerprints that *matched* a known entry — the tempting debug step when a scan
|
||||
returns zero ("print what it saw") publishes low-entropy fingerprints wholesale.
|
||||
|
||||
And say plainly what a fingerprint register *is*, so nobody rediscovers it later
|
||||
as an alarm: even for an unguessable secret, a published fingerprint is a
|
||||
**confirmation oracle** for anyone who already holds a candidate corpus. That is
|
||||
exactly how a long-retired token gets identified in old transcripts — and it works
|
||||
identically for someone else holding those same files. Net positive, since they
|
||||
would already hold the value; state it rather than leaving it implicit.
|
||||
|
||||
## 3. Liveness probes, and the trap that scoping creates
|
||||
|
||||
```sh
|
||||
@@ -106,7 +129,82 @@ Value-agnostic sweeps (uuid / 40-hex / `NAME=VALUE`) drown in false positives at
|
||||
fleet scale — 608 candidates, mostly session UUIDs and git SHAs. Name-anchoring
|
||||
plus entropy plus provenance, applied to **document text**, is what works.
|
||||
|
||||
## 6. Choosing scopes: derive them from measured consumers
|
||||
## 6. Proving absence: instrument strength, and four ways a scan lies clean
|
||||
|
||||
Section 5's warning is about false *positives* — name-anchoring and provenance are
|
||||
what stop a triage sweep drowning in session UUIDs. **A gate is the opposite job.**
|
||||
Triage optimises precision; proving absence optimises recall. Every failure below
|
||||
reported a reassuring zero over a secret that was really there.
|
||||
|
||||
**Rank the instrument, and state which one produced your zero.**
|
||||
|
||||
| Instrument | Needs | Blind to |
|
||||
|---|---|---|
|
||||
| exact-byte value search | you hold the value | nothing — no tokeniser to fool |
|
||||
| class/structure pass | a header pattern | anything without a recognisable shape |
|
||||
| fingerprint census | a fingerprint list | any secret not listed; tokenisation |
|
||||
|
||||
A census is deliberately value-free, so it must *extract candidates and hash them*
|
||||
— which makes its sensitivity a property of the tokeniser, not of the corpus. If
|
||||
you hold the value, search the bytes instead, and search the value's JSON-escaped
|
||||
rendering too when the corpus is `.jsonl`.
|
||||
|
||||
**1. Census and class answer different questions; neither substitutes.** A census
|
||||
answers *"has a KNOWN secret leaked?"*, a class pass *"is there secret-SHAPED
|
||||
material here?"* Both failure modes were measured on this fleet: a class-only
|
||||
pre-commit hook passed plaintext UUID API credentials to a shared repo twice,
|
||||
because a UUID carries no key header — while a census-only gate reported 0 hits
|
||||
with freshly-synced SSH private keys and an age identity in the tree, because no
|
||||
key is in the census. Run both passes.
|
||||
|
||||
**2. Tokenisation — quoting alone can decide detectability.** Maximal-run
|
||||
extraction swallows the value of an *unquoted* assignment:
|
||||
|
||||
```
|
||||
PROXMOX_SECRET=<uuid> # ONE run; the uuid is never hashed alone -> MISS
|
||||
export SECRET="<uuid>" # the quote ends the run; bare uuid hashed -> HIT
|
||||
```
|
||||
|
||||
Take the **union** of three strategies, because each fails in a different
|
||||
direction — (2) is the one that recovers the unquoted case:
|
||||
|
||||
~~~python
|
||||
runs = re.findall(r'[^\s"\'`]{12,}', text) # 1. maximal runs
|
||||
split = [p for r in runs for p in re.split(r'[=!,;:@|()\[\]{}<>]', r) if len(p) >= 12]
|
||||
shape = re.findall(UUID_RE, text) + re.findall(r'[0-9a-f]{32,64}', text)
|
||||
candidates = set(runs) | set(split) | set(shape)
|
||||
~~~
|
||||
|
||||
**3. Scan the index or the pushed tree, never the working tree.** The working tree
|
||||
is not what gets published. And for an rsync-published mirror a repo-only fix is
|
||||
not weaker, it is *temporary*: the next sync re-publishes the live disk. Fix the
|
||||
live file first, verify it clean **by fingerprint**, then sync. Read blobs with
|
||||
`git ls-tree -r <sha>` plus one `git cat-file --batch` (thousands of `git show`
|
||||
calls is the slow way).
|
||||
|
||||
**4. Git filters never run on symlinks — and `check-attr` will not tell you.** A
|
||||
symlink's blob is the *target path*, so `filter=git-crypt` can never encrypt it,
|
||||
yet `git check-attr filter` cheerfully answers `git-crypt` for that path. **A
|
||||
symlinked secret stays plaintext no matter what `.gitattributes` says.** Join the
|
||||
attribute against the **file mode** (`git ls-files -s`, mode `120000`) and verify
|
||||
the index blob really begins `\0GITCRYPT\0`. Report encrypted / symlinked /
|
||||
scanned as three separate numbers and assert they sum — encrypted and symlinked
|
||||
blobs are *skipped*, not certified clean.
|
||||
|
||||
**Self-test two-sided, and abort if it cannot discriminate.** Require a synthetic
|
||||
positive to fire AND a negative to stay silent before trusting any zero. Keep the
|
||||
fixtures in *structurally separate buffers*: put a quoted and an unquoted probe in
|
||||
one buffer and the quote terminates the run, handing the bare token to the weak
|
||||
extractor and making it look as strong as the union — a self-test artifact that
|
||||
has already fooled an agent here. And never gate on `$?` when the tool has a
|
||||
lock-skip or no-op path that also exits 0; judge the reported line.
|
||||
|
||||
**Row-gone is not bytes-gone.** A correct sqlite `DELETE` leaves the payload in
|
||||
freelist pages until `VACUUM`, so deletion effectiveness is *two* numbers: rows
|
||||
removed, and a raw byte scan of the `.sqlite3`. One aggregate figure reported as
|
||||
"erased" has only measured "unretrievable".
|
||||
|
||||
## 7. Choosing scopes: derive them from measured consumers
|
||||
|
||||
Before creating a replacement token, find out what actually uses it:
|
||||
|
||||
@@ -126,7 +224,7 @@ is a hygiene event, not an instance compromise.
|
||||
Then prove the scope with an acceptance suite that declares expectations first:
|
||||
must-work routes → `200`; `/admin/*`, `/user`, `/user/repos` → `403`.
|
||||
|
||||
## 7. What rotation does *not* fix
|
||||
## 8. What rotation does *not* fix
|
||||
|
||||
- **A cleartext channel.** If the endpoint is `http://`, the *new* token is
|
||||
exposed identically from first use. Raise TLS separately.
|
||||
@@ -139,7 +237,7 @@ must-work routes → `200`; `/admin/*`, `/user`, `/user/repos` → `403`.
|
||||
`.env.age` moves on, so old values linger on disk (and in backups) long after
|
||||
rotation. They are a common source of "mystery" fingerprints in a census.
|
||||
|
||||
## 8. This fleet's secret store (verify, do not assume)
|
||||
## 9. This fleet's secret store (verify, do not assume)
|
||||
|
||||
- All `*.env.age` live in **one** repo: `joakimp/docker-compose-repo`. `myconfigs`
|
||||
has none.
|
||||
|
||||
Reference in New Issue
Block a user