836e35b320
Review caught that the tier vocabulary was used in the report, the commit message and the docs without being defined anywhere the reader would land, and inspecting that turned up two real defects rather than just a wording gap. 1. STALE DOCSTRING. The module still described tier 3 as if it redacts, which stopped being true when the 403-hit measurement demoted it to report-only. It also credited tier 3 with resolving the 40-hex-PAT-vs-commit-sha collision — false by default, since a reporting rule resolves nothing. Corrected, with the consequence stated plainly: in the default configuration a sha-shaped PAT is caught if and only if it belongs to THIS machine, because only a known value (tier 1) or a naming key (tier 3, reporting) can separate it from a commit sha. That is an accepted gap; the alternative is redacting every sha in the palace. 2. THE VOCABULARY NEVER REACHED THE OUTPUT. The tool prints rule names (github-pat, env-value, url-credentials) and nothing printed a tier, so the docs' tier language was unconnected to what an operator actually sees. Added RULE_TIERS as the authoritative rule -> tier mapping, tier_of(), and Finding.tier; the feeder now prints "T2:github-pat=6" so what matched and how much to trust it are both visible on one line. A self-test asserts every rule that can appear in a Finding maps to a tier, so adding a rule without classifying it fails the tests instead of printing "T?". Tiers, for the record, are three kinds of EVIDENCE (not three severities): T1 known value from this process's env — near-certain, zero FP by construction; T2 known vendor shape — strong, the prefix is meaningful; T3 key name says secret — candidate only, measured FP-heavy, reported. T0 is reserved for suspicions(), which is a measured NON-detection. Docs gain worked one-line examples per tier and a "which tier fired?" section showing real output. 46 self-test cases pass.
186 lines
10 KiB
Markdown
186 lines
10 KiB
Markdown
# Secret hygiene: keeping credentials out of the palace
|
|
|
|
A palace is mined from **transcripts**, and transcripts contain whatever the
|
|
terminal printed. Print an environment, `cat` a `.env`, paste a `curl -H
|
|
"Authorization: …"`, and the secret becomes a drawer — searchable by every agent
|
|
on the fleet, on every machine, indefinitely. This document describes the
|
|
scrubbing that exists, what it deliberately does **not** try to do, and the two
|
|
layers it belongs in.
|
|
|
|
Measured on this fleet, 2026-08-27: a single leaked bearer token had reached
|
|
**3 drawers, 13 feeder inbox files across all three devices, and 10 local files
|
|
dating back 10 days**. Nobody was careless; an agent inspected an environment
|
|
variable while debugging. That frequency is the design premise — this is a
|
|
routine event to be handled by the pipeline, not an incident to be handled by
|
|
discipline.
|
|
|
|
## 1. Where the hook goes, and why there are two of them
|
|
|
|
| Layer | Implemented in | Sees | Catches |
|
|
|---|---|---|---|
|
|
| **Client, pre-staging** | `bin/mempalace_redact.py`, called from `bin/mempalace-pi-session` at the point the staged transcript is written | the machine's own environment **and** the transcript | the real-world case: a secret this machine holds, printed into a session, before it is uploaded anywhere |
|
|
| **Server, pre-persist** | *not yet implemented* — see §5 | only the text it is handed | secrets from clients that predate the scrubber, from non-pi clients, and secrets an agent types straight into `add_drawer` |
|
|
|
|
The client hook is the load-bearing one, because it is the only layer that can
|
|
compare text against **the actual secret values it holds** (see tier 1 below) and
|
|
because it stops the leak *before transmission*. The server hook is defence in
|
|
depth: pattern-only, but it covers write paths the feeder never sees.
|
|
|
|
Client placement matters for one specific reason: the pi feeder writes the staged
|
|
transcript from an in-memory structure, and the remote path then `rsync`s **that
|
|
same file** byte-for-byte. Scrubbing at the write point therefore covers local
|
|
mining and remote upload with a single hook — there is no second serialization to
|
|
forget.
|
|
|
|
## 2. Detection is anchored on meaning, never on entropy
|
|
|
|
The tempting design — "redact long random-looking strings" — is actively
|
|
destructive here, because a palace is *full* of high-entropy strings that are its
|
|
own primary keys:
|
|
|
|
```
|
|
drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b drawer id
|
|
evt_20260826T213924_397b1f710c2d event id
|
|
rep_d344e349ba276d6fc11997cd552f6937 replica id
|
|
1787780364792-000000-rep_d344e349ba276d6f… hybrid logical clock
|
|
ecc2a9c574f4156403cf4e90aea8b0d088c75700 git commit sha
|
|
```
|
|
|
|
An entropy detector fires on every one of those, and the resulting redaction is
|
|
silent, permanent, and destroys traceability. So detection uses three anchors
|
|
that carry meaning instead:
|
|
|
|
Each tier is a different *kind of evidence* that a string is a credential. The
|
|
tier is not a severity ranking of the secret — it is how much to trust the
|
|
detection. Rule names appear in the output; the tier tells you how to read them.
|
|
|
|
| Tier | Anchor | Default | False-positive risk |
|
|
|---|---|---|---|
|
|
| **1 — known values** | literal values from this process's env, for variables whose *name* says secret (`…TOKEN`, `…SECRET`, `…PASSWORD`, `…API_KEY`) | **redact** | none by construction: the value *is* the secret |
|
|
| **2 — known shapes** | vendor-prefixed credentials: `ghp_…`, `github_pat_…`, `glpat-…`, `xox[abprs]-…`, `sk-…`, `AKIA…`, `AIza…`, `hf_…`, JWTs, PEM private-key blocks, credentials inside URLs, `Authorization:` headers | **redact** | very low: the prefix is meaningful, not random |
|
|
| **3 — name=value** | an assignment whose *key* says secret | **report only** | measured **high** — see §3 |
|
|
|
|
### Worked examples, one line each
|
|
|
|
```
|
|
MEMPALACE_REMOTE_TOKEN=pfrDBfak… tier 1 — value matches this env's secret
|
|
the token is pfrDBfak… apparently tier 1 — same value, bare in prose, still caught
|
|
git clone https://joakim:hunter2@git/x tier 2 — credentials in a URL
|
|
Authorization: Bearer abcdefghijklmnop tier 2 — header shape
|
|
ghp_AbCdEf… / glpat-… / AKIA… / sk-ant-… tier 2 — vendor prefix
|
|
db_password = s3cr3t-p4ssw0rd-xyz tier 3 — only the KEY suggests it (reported)
|
|
GITEA_TOKEN=0123456789abcdef… (40 hex) tier 3 — indistinguishable from a commit sha
|
|
```
|
|
|
|
### Which tier fired? Read it off the output
|
|
|
|
The tool prints **rule names**, not tier numbers, because the rule says *what*
|
|
matched. The feeder prefixes them with the tier so both are visible:
|
|
|
|
```
|
|
[REDACTED] 2026-06-27T23-13-49.jsonl T1:env-value=1, T2:github-fine-grained=1 fp=ad78c7d4,fcd95ab5
|
|
[scrub] 29 redaction(s) applied, 141 name-anchored candidate(s) reported only
|
|
```
|
|
|
|
`RULE_TIERS` in `mempalace_redact.py` is the authoritative mapping, `tier_of()`
|
|
reads it, and a self-test fails if any rule is left unclassified — so the two
|
|
vocabularies cannot drift apart silently.
|
|
|
|
Tier 1 catches any presentation of a secret — env dump, JSON, error message,
|
|
URL, prose — because it matches the value itself. It is also, **in the default
|
|
configuration, the only tier that resolves a shape collision**: a 40-hex Gitea
|
|
PAT is byte-identical to a git commit sha, so only a known value (tier 1,
|
|
redacts) or a naming key (tier 3, reports) can tell them apart. A sha-shaped PAT
|
|
is therefore caught if and only if it belongs to this machine — an accepted gap,
|
|
since the alternative is redacting every commit sha in the palace.
|
|
|
|
## 3. The false-positive measurement, which changed the design
|
|
|
|
Tier 3 was originally going to redact. Measured against **52 MB of real fleet
|
|
transcripts (45 sessions)** it produced **403 hits**, and inspection of them (with
|
|
values masked) showed the overwhelming majority were not secrets:
|
|
|
|
- `GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}` — docker-compose interpolation
|
|
- `const tokens = countTokensForModel` — TypeScript source stored as memory
|
|
- `refreshToken(credentials: Credentials …)` — a *type annotation*
|
|
- `--setattr=krbPasswordExpiration=20260529…` — an IPA attribute holding a date
|
|
- `|1.FORGE.TOKENS:…` — AAAK diary shorthand
|
|
- `(root@host) Password:` followed by unrelated terminal output
|
|
|
|
Redacting those would corrupt code, configuration and documentation held as
|
|
memory, in order to catch secrets tier 1 already catches by value. After adding
|
|
guards for interpolation (`${VAR}`, `$VAR`, `%VAR%`, `{{tpl}}`), code context,
|
|
non-secret key suffixes (`…Expiration`, `…Count`, `…Type`) and all-digit values,
|
|
the enforced count on the same corpus fell from **403 to 29** — and those 29 are
|
|
tier-1 and tier-2 hits, i.e. real credential shapes.
|
|
|
|
So tier 3 **reports and does not rewrite** by default. Set
|
|
`MEMPALACE_REDACT_STRICT=1` to make it enforce, e.g. on a machine whose
|
|
transcripts are configuration-heavy rather than code-heavy.
|
|
|
|
## 4. Known false negatives — stated, not hidden
|
|
|
|
The scrubber will not catch a novel credential format pasted bare into prose
|
|
with no name nearby, a secret belonging to a machine whose environment this
|
|
process cannot see, a base64-of-a-secret, or a secret split across lines.
|
|
|
|
**"The scrubber ran" must never be read as "there are no secrets in here."** To
|
|
keep that measurable rather than assumed, two things are reported:
|
|
|
|
- every scrub prints a count — including `0 redaction(s)`, because printing
|
|
nothing is indistinguishable from a scrubber that never ran;
|
|
- `suspicions()` reports high-entropy strings it did **not** redact, as
|
|
`(length, fingerprint)` pairs rather than values, so the miss rate can be
|
|
tracked over time and a recurring fingerprint can be investigated by hand.
|
|
|
|
Findings never carry the secret. A `Finding` holds the rule, the label, the
|
|
length and `sha256(value)[:8]` — enough to recognise the same leak recurring,
|
|
not enough to recover it.
|
|
|
|
**Fail closed.** If the redactor cannot be imported, the feeder refuses to stage
|
|
rather than staging unscrubbed (`exit 3`). Override deliberately with
|
|
`MEMPALACE_FEED_ALLOW_UNSCRUBBED=1`.
|
|
|
|
## 5. The server-side layer (not yet implemented)
|
|
|
|
Three call sites, because the server has three near-duplicate validators rather
|
|
than one:
|
|
|
|
| Path | Function | Covers |
|
|
|---|---|---|
|
|
| drawers | `config.py` → `sanitize_content()` | `add_drawer`, `update_drawer`, `diary_write`, `checkpoint` — all four route through it |
|
|
| events | `logstream.py` → `_sanitize_body()` | `event_append`, and `event_ack` transitively |
|
|
| artifacts | `logstream.py` → `put_artifact()` | inlines its own checks; needs its own edit |
|
|
|
|
Server-side scrubbing is **tier 2 only** (plus optional tier-3 reporting): the
|
|
hub cannot see a client's environment, so tier 1 is structurally unavailable
|
|
there. That asymmetry is the reason the client hook is not redundant.
|
|
|
|
## 6. Cleaning up a leak that already landed
|
|
|
|
1. **Never re-echo the value while hunting it.** Pass it via stdin, never argv
|
|
(visible in `ps` on a shared host). Note `get_drawer` is a trap: reading a
|
|
drawer in order to redact it prints the secret back into the live transcript.
|
|
2. **Redact, don't delete.** Replacing the value in place keeps the mined
|
|
transcript's memory value; a stale embedding vector is a cheap price.
|
|
3. **Scrub the feeder inbox too**, or the next mine re-files it.
|
|
4. **A live session file needs an equal-length in-place overwrite** — the agent
|
|
holds it open, so temp-file-plus-rename loses everything appended afterwards.
|
|
5. **Expect residue.** SQLite keeps old page content in freed pages until
|
|
`VACUUM`, so a raw byte scan still matches after a successful `UPDATE`. Decide
|
|
explicitly whether that matters: if the credential store on the same host is
|
|
plaintext anyway, it usually does not.
|
|
|
|
## 7. Usage
|
|
|
|
```bash
|
|
# unit tests, including the must-not-redact corpus of palace id shapes
|
|
python3 bin/mempalace_redact.py --self-test
|
|
|
|
# scrub anything on stdin; report goes to stderr
|
|
some-command | python3 bin/mempalace_redact.py > clean.txt
|
|
|
|
# also list high-entropy strings that were NOT redacted, as fingerprints
|
|
python3 bin/mempalace_redact.py --suspicions < transcript.jsonl > clean.jsonl
|
|
```
|