\u00a75 said embedding_metadata.string_value holds "metadata fields only". False, measured directly: chroma also stores a copy of the document text there, under key chroma:document. Confirmed with a disposable sentinel drawer (pi@tor-ms22, 2026-08-30): one row in fts_content AND one row in embedding_metadata for the same drawer. This was a real mistake in shipped guidance, not a nitpick: this section's own scanning advice was written to guard against explaining a zero with a mechanism nobody verified from source, and the section itself did exactly that -- I downgraded a census to "a floor" on the strength of a metadata-blind claim I never checked against chroma's actual storage layout. The practical scan order is unchanged (fts_content is still the direct target, raw bytes are still the backstop); only the stated REASON for a metadata zero changes: it needs a different explanation now (key filter, query shape, escaping), not "structurally absent". \u00a76's row-gone/bytes-gone claim is upgraded from asserted to measured, same sentinel: delete_by_source took both fts_content and embedding_metadata 1->0, raw bytes stayed 4->4 (freed pages persist until VACUUM). Also records the method that unblocked the measurement: not a better instrument, a disposable sentinel drawer instead of risking real fleet data. No image behaviour changes.
15 KiB
name, description
| name | description |
|---|---|
| credential-incident-response | Respond correctly when a live credential is found where it should not be — in a chat transcript, a MemPalace drawer, a log, a git-tracked config, or an agent-authored note. Load this whenever a task involves a leaked/exposed secret, a token rotation, a "is this credential still live?" question, deciding whether to delete or scrub stored content, proving a corpus is clean, or choosing scopes for a new API token. Covers the mandatory order of operations (probe the issuer FIRST — severity before cleanliness), leak-free identity via sha256[:8] fingerprints and when publishing one is safe, why revocation beats deletion for anything already replicated, scopes derived from measured consumers, the three places a secret hides in a Chroma palace, how to prove ABSENCE rather than assume it (instrument strength, census vs class passes, the tokenisation trap where quoting decides detectability, why git filters never run on symlinks, self-tests that abort), where this fleet's secrets live, and what rotation does NOT fix. |
Credential incident response
A leaked credential is a severity question before it is a cleanliness question. Two days of scrubbing, redaction plumbing and deletion planning were once spent on a set of 13 credentials of which 11 were already dead at the provider — a fact that cost five HTTP requests to establish and was never checked. Meanwhile the two live ones turned out to be instance-owner admin tokens, which nobody had looked at either.
1. Order of operations — do not reorder this
- Is it still accepted? Probe the issuing provider. Dead credential → hygiene item, stop panicking. Live → incident, continue.
- What can it do? Read the identity back.
is_admin,id=1, scopes, which account. A read-only repo token and an instance-owner admin token are not the same finding. - What consumes it? Grep for real consumers before assuming breakage.
- Where does it live? Enumerate copies (store, palace, transcripts, git).
- Then rotate/revoke, and only then consider cleanup.
Doing 4→3→1 in reverse produces confident, wrong severity calls and wasted cleanup. If you only have time for one step, do step 1.
2. Leak-free identity: fingerprint, never the value
Publishing an 8-hex fingerprint lets you compare a credential across machines,
files, drawers and peers without ever materialising the secret. Same formula as
mempalace_redact.py:
printf '%s' "$SECRET" | sha256sum | cut -c1-8 # printf, NOT echo (no newline)
printf '%s' 'test' | sha256sum | cut -c1-8 # self-test -> 9f86d081
Report as (variable, fp, length). Equal fingerprints across hosts prove a
shared credential; that is usually the important part. Never paste a live
value into a search query, a palace drawer, an event body, or a chat message —
in an agent context your own tool output is itself captured and re-filed.
Precondition — only fingerprint what an adversary cannot enumerate. An 8-hex
fingerprint is 32 bits over its input space, so publishing fp8(x) hands
anyone a membership oracle: they can test x == v for every candidate v
they can generate. For a 40-char random token that space is unreachable. For a
hostname, username, e-mail, port, path, commit SHA or weak password it is a
wordlist. If you can imagine writing the wordlist, you cannot publish the
fingerprint — reference those by name and location instead. "High entropy" is
the usual sufficient condition, not the test: a commit SHA is 160-bit and still
fully enumerable from the repo. sha256("") = e3b0c442 is the degenerate case,
recognisable on sight precisely because its input space has one member.
Candidate fingerprints are working memory, never output. A scanner that hashes every token in a file also hashes hostnames, paths and e-mails. Print only fingerprints that matched a known entry — the tempting debug step when a scan returns zero ("print what it saw") publishes low-entropy fingerprints wholesale.
And say plainly what a fingerprint register is, so nobody rediscovers it later as an alarm: even for an unguessable secret, a published fingerprint is a confirmation oracle for anyone who already holds a candidate corpus. That is exactly how a long-retired token gets identified in old transcripts — and it works identically for someone else holding those same files. Net positive, since they would already hold the value; state it rather than leaving it implicit.
3. Liveness probes, and the trap that scoping creates
# Gitea
curl -sS -m 10 -o /dev/null -w '%{http_code}\n' -H "Authorization: token $T" \
"$GITEA_HOST/api/v1/repos/<owner>/<repo>/actions/runs?limit=1"
# GitHub
curl -sS -m 10 -o /dev/null -w '%{http_code}\n' -H "Authorization: token $T" \
https://api.github.com/user
200live ·401revoked/invalid ·403= wrong question, not a dead token- Probe the issuer that minted it. A 401 from an unrelated instance says
nothing. Resolve the host from config (
GITEA_EGL_HOSTetc.), do not assume. - Under scoped tokens,
/api/v1/userreturns 403 for a perfectly live token unlessuserscope was granted. So it cannot distinguish revoked from merely scoped. Use a repository route the token is authorised for. - Verify both directions after a rotation: old → 401, new → 200. The second check is what catches "deleted the wrong token".
- Port/scheme come from config, not habit: one instance here is
http://gitea.egl.lan:3000— plain HTTP, with 443 refused.
4. Revocation beats deletion — the load-bearing rule
Once revoked, stored copies are inert; you may leave them. Deleting them is
best-effort over an unbounded copy set: FTS shadow rows, feed inbox .jsonl
files on every host, sqlite free pages after the delete, mesh replicas that
already synced, and backups. Revocation invalidates every copy everywhere at
once, including copies nobody enumerated.
So: rotate + revoke first. Treat drawer deletion as optional hygiene, never as the remedy. Then record the retired fingerprints as known-dead so the next census recognises them instead of reopening the investigation.
Corollary: never reach for mempalace_sync or a bulk delete_by_source on a
shared palace as incident response. High blast radius, low actual benefit.
5. Finding a secret in a Chroma palace — three targets, in this order
embedding_fulltext_search_content.c0— document textembedding_metadata.string_value— metadata fields, and a second copy of the document text under keychroma:document- raw byte scan of every
*.sqlite3— backstop, covers FTS pages and free space
Correction, measured on chroma 1.5.9 with a sentinel drawer: one row in (1)
AND one row in (2) for the same drawer, so (2) is not structurally
content-blind — an earlier version of this section said it held "metadata
fields only", and that was wrong. Scan (1) and (3) regardless: (1) is the direct
target. But if a string_value query returns zero for a value you know is in a
drawer, the cause is a key filter, a query shape or escaping — not structural
absence, and the difference matters because the false explanation is what makes
the zero feel safe. See §6: do not explain a zero with a mechanism you have not
read from source.
Semantic search proves nothing about absence — it returns top-k. For
completeness, enumerate by filing window (list_drawers(since=T, before=T+1m)),
since one mine shares a minute.
Value-agnostic sweeps (uuid / 40-hex / NAME=VALUE) drown in false positives at
fleet scale — 608 candidates, mostly session UUIDs and git SHAs. Name-anchoring
plus entropy plus provenance, applied to document text, is what works.
6. Proving absence: instrument strength, and four ways a scan lies clean
Section 5's warning is about false positives — name-anchoring and provenance are what stop a triage sweep drowning in session UUIDs. A gate is the opposite job. Triage optimises precision; proving absence optimises recall. Every failure below reported a reassuring zero over a secret that was really there.
Rank the instrument, and state which one produced your zero.
| Instrument | Needs | Blind to |
|---|---|---|
| exact-byte value search | you hold the value | nothing — no tokeniser to fool |
| class/structure pass | a header pattern | anything without a recognisable shape |
| fingerprint census | a fingerprint list | any secret not listed; tokenisation |
A census is deliberately value-free, so it must extract candidates and hash them
— which makes its sensitivity a property of the tokeniser, not of the corpus. If
you hold the value, search the bytes instead, and search the value's JSON-escaped
rendering too when the corpus is .jsonl.
1. Census and class answer different questions; neither substitutes. A census answers "has a KNOWN secret leaked?", a class pass "is there secret-SHAPED material here?" Both failure modes were measured on this fleet: a class-only pre-commit hook passed plaintext UUID API credentials to a shared repo twice, because a UUID carries no key header — while a census-only gate reported 0 hits with freshly-synced SSH private keys and an age identity in the tree, because no key is in the census. Run both passes.
2. Tokenisation — quoting alone can decide detectability. Maximal-run extraction swallows the value of an unquoted assignment:
PROXMOX_SECRET=<uuid> # ONE run; the uuid is never hashed alone -> MISS
export SECRET="<uuid>" # the quote ends the run; bare uuid hashed -> HIT
Take the union of three strategies, because each fails in a different direction — (2) is the one that recovers the unquoted case:
runs = re.findall(r'[^\s"\'`]{12,}', text) # 1. maximal runs
split = [p for r in runs for p in re.split(r'[=!,;:@|()\[\]{}<>]', r) if len(p) >= 12]
shape = re.findall(UUID_RE, text) + re.findall(r'[0-9a-f]{32,64}', text)
candidates = set(runs) | set(split) | set(shape)
3. Scan the index or the pushed tree, never the working tree. The working tree
is not what gets published. And for an rsync-published mirror a repo-only fix is
not weaker, it is temporary: the next sync re-publishes the live disk. Fix the
live file first, verify it clean by fingerprint, then sync. Read blobs with
git ls-tree -r <sha> plus one git cat-file --batch (thousands of git show
calls is the slow way).
4. Git filters never run on symlinks — and check-attr will not tell you. A
symlink's blob is the target path, so filter=git-crypt can never encrypt it,
yet git check-attr filter cheerfully answers git-crypt for that path. A
symlinked secret stays plaintext no matter what .gitattributes says. Join the
attribute against the file mode (git ls-files -s, mode 120000) and verify
the index blob really begins \0GITCRYPT\0. Report encrypted / symlinked /
scanned as three separate numbers and assert they sum — encrypted and symlinked
blobs are skipped, not certified clean.
Self-test two-sided, and abort if it cannot discriminate. Require a synthetic
positive to fire AND a negative to stay silent before trusting any zero. Keep the
fixtures in structurally separate buffers: put a quoted and an unquoted probe in
one buffer and the quote terminates the run, handing the bare token to the weak
extractor and making it look as strong as the union — a self-test artifact that
has already fooled an agent here. And never gate on $? when the tool has a
lock-skip or no-op path that also exits 0; judge the reported line.
Row-gone is not bytes-gone. Measured, same sentinel drawer: after
delete_by_source the row count went 1 -> 0 in both the FTS content table and
embedding_metadata, while the raw byte count stayed 4 -> 4 — sqlite does not
zero freed pages, so the payload sits in free space until VACUUM. Deletion
effectiveness is therefore two numbers, and each direction has a trap: one
aggregate figure reported as "erased" has only measured "unretrievable", while a
raw byte scan used as the acceptance gate reads a CORRECT, complete deletion as a
failure. (Note how this was measured: the blocker was never a better instrument,
it was the subject — file your own disposable sentinel and delete that, instead
of testing deletion on real data.)
7. Choosing scopes: derive them from measured consumers
Before creating a replacement token, find out what actually uses it:
git -C <repo> remote get-url origin # ssh:// ? then git needs NO token
git config --global --list | grep -iE 'credential|insteadof' # and no helper?
grep -rhoE 'api/v1/[A-Za-z0-9/{}$_.-]+' <consumers> | sort -u # exact routes
grep -rhoE '\-X [A-Z]+' <consumers> # any writes?
Real outcome here: git used SSH keys throughout, and the token's only consumer
read three CI-run routes with GET. So repository: Read and nothing else
replaced two admin tokens. Scoping shrinks the blast radius of the next leak
far more than any redaction pipeline does — a read-only token in a transcript
is a hygiene event, not an instance compromise.
Then prove the scope with an acceptance suite that declares expectations first:
must-work routes → 200; /admin/*, /user, /user/repos → 403.
8. What rotation does not fix
- A cleartext channel. If the endpoint is
http://, the new token is exposed identically from first use. Raise TLS separately. - Git history. A secret committed and pushed cannot be fixed by any store or palace operation — it needs rotation and history surgery.
- Agent-authored content. Stage-write redactors see transcripts only, never
add_drawer/checkpoint/diary_writeoutput. Never type a secret into the palace yourself; nothing downstream will catch it. - Plaintext/encrypted drift. Gitignored plaintext
.envfiles go stale while.env.agemoves on, so old values linger on disk (and in backups) long after rotation. They are a common source of "mystery" fingerprints in a census.
9. This fleet's secret store (verify, do not assume)
- All
*.env.agelive in one repo:joakimp/docker-compose-repo.myconfigshas none. - Every
.agefile has one X25519 recipient — a single key tracked inmyconfigsunder git-crypt. Unlocking git-crypt therefore decrypts the entire fleet's secrets, including hosts you have no access to. The age layer adds no isolation beyond git-crypt. - Flow:
./fetch-secrets.sh <host>(decrypt →.env) → edit →./encrypt-secrets.sh <host>→ commit → push →docker compose up -d --force-recreate. - Always pass the host argument to
encrypt-secrets.sh. Bare, it walks the whole tree and re-encrypts every.envit finds, re-nonced, including stale ones — silently rolling back other hosts' secrets. - After any re-encrypt, check the header still shows exactly one X25519
recipient; a hand-rolled
age -rlocks the rest of the fleet out, and the failure only appears on another machine, later.