Files
mempalace-toolkit/docs/secret-hygiene.md
T
joakimp 3d47937d06 feed: scrub secrets before staging a transcript, name-anchored not entropy-anchored
A palace is mined from transcripts, and transcripts contain whatever the terminal
printed. Measured on this fleet: one leaked bearer token had reached 3 drawers,
13 feeder inbox files across all three devices, and 10 local files spanning 10
days — from an agent inspecting an env var while debugging. That frequency is the
premise: this is a pipeline problem, not a discipline problem.

WHERE. bin/mempalace_redact.py, called from mempalace-pi-session at the point the
staged transcript is written. That single hook covers both transports, because
local mode mines the staged file and remote mode rsyncs that same file
byte-for-byte. Scrubbing operates on the parsed objects rather than the
serialized text, so string values are rewritten while keys, ids and structure are
untouched — scanning raw JSONL instead invents keys like "tapiKey" out of the \t
escape preceding a field name (observed, not theorised).

WHY NOT ENTROPY. The obvious "redact long random-looking strings" is actively
destructive here: drawer ids, chunk ids, event ids, replica ids, HLCs and commit
SHAs are all high-entropy, are the majority of random-looking text in a palace,
and redacting them is silent and permanent. Detection is anchored on meaning
instead: tier 1 literal values from this process's env whose NAME says secret
(zero false positives by construction, and the only tier that can tell a 40-hex
Gitea PAT from a git commit sha); tier 2 vendor-prefixed shapes (ghp_, glpat-,
xox*-, sk-, AKIA, AIza, hf_, JWT, PEM blocks, URL credentials, Authorization
headers); tier 3 name=value assignments.

THE MEASUREMENT THAT CHANGED THE DESIGN. Tier 3 was going to redact. Against
52 MB of real fleet transcripts it produced 403 hits, and inspection with values
masked showed most were ${VAR} interpolation in compose files, TypeScript
identifiers, a TYPE ANNOTATION (credentials: Credentials), an IPA attribute
holding a date (krbPasswordExpiration), AAAK diary shorthand, and terminal output
following a "Password:" prompt. Redacting those corrupts code and docs held as
memory to catch what tier 1 already catches by value. After adding guards for
interpolation, code context, non-secret key suffixes and all-digit values, the
enforced count fell 403 -> 29 on the same corpus. So tier 3 REPORTS and does not
rewrite unless MEMPALACE_REDACT_STRICT=1.

HONESTY ABOUT MISSES. Known false negatives are documented rather than papered
over: novel formats in bare prose, another machine's secrets, base64-of-a-secret,
line-split secrets. Every run prints a count including "0 redaction(s)", because
silence is indistinguishable from a scrubber that never ran, and suspicions()
reports high-entropy strings it did NOT redact as (length, fingerprint) so the
miss rate is measurable. Findings never carry the value — rule, label, length and
sha256[:8], enough to recognise a recurrence, not enough to recover the secret.

FAIL CLOSED: no scrubber, no staging (exit 3), overridable with
MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 for a machine older than this file.

Tested: 42-case corpus in --self-test, including every palace id shape as a
must-not-redact case, idempotency, and a compound case the corpus caught where a
vendor placeholder was then re-matched by the URL rule (nested placeholder — the
secret was hidden either way, so only an exact-output assertion catches it).
End-to-end on 45 real sessions: 29 enforced, fail-closed verified at rc=3,
override verified loud. Server-side layer specified in docs/secret-hygiene.md §5
but NOT implemented — it is tier 2 only there, since the hub cannot see a
client's environment.
2026-08-27 13:30:47 +02:00

152 lines
8.3 KiB
Markdown

# Secret hygiene: keeping credentials out of the palace
A palace is mined from **transcripts**, and transcripts contain whatever the
terminal printed. Print an environment, `cat` a `.env`, paste a `curl -H
"Authorization: …"`, and the secret becomes a drawer — searchable by every agent
on the fleet, on every machine, indefinitely. This document describes the
scrubbing that exists, what it deliberately does **not** try to do, and the two
layers it belongs in.
Measured on this fleet, 2026-08-27: a single leaked bearer token had reached
**3 drawers, 13 feeder inbox files across all three devices, and 10 local files
dating back 10 days**. Nobody was careless; an agent inspected an environment
variable while debugging. That frequency is the design premise — this is a
routine event to be handled by the pipeline, not an incident to be handled by
discipline.
## 1. Where the hook goes, and why there are two of them
| Layer | Implemented in | Sees | Catches |
|---|---|---|---|
| **Client, pre-staging** | `bin/mempalace_redact.py`, called from `bin/mempalace-pi-session` at the point the staged transcript is written | the machine's own environment **and** the transcript | the real-world case: a secret this machine holds, printed into a session, before it is uploaded anywhere |
| **Server, pre-persist** | *not yet implemented* — see §5 | only the text it is handed | secrets from clients that predate the scrubber, from non-pi clients, and secrets an agent types straight into `add_drawer` |
The client hook is the load-bearing one, because it is the only layer that can
compare text against **the actual secret values it holds** (see tier 1 below) and
because it stops the leak *before transmission*. The server hook is defence in
depth: pattern-only, but it covers write paths the feeder never sees.
Client placement matters for one specific reason: the pi feeder writes the staged
transcript from an in-memory structure, and the remote path then `rsync`s **that
same file** byte-for-byte. Scrubbing at the write point therefore covers local
mining and remote upload with a single hook — there is no second serialization to
forget.
## 2. Detection is anchored on meaning, never on entropy
The tempting design — "redact long random-looking strings" — is actively
destructive here, because a palace is *full* of high-entropy strings that are its
own primary keys:
```
drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b drawer id
evt_20260826T213924_397b1f710c2d event id
rep_d344e349ba276d6fc11997cd552f6937 replica id
1787780364792-000000-rep_d344e349ba276d6f… hybrid logical clock
ecc2a9c574f4156403cf4e90aea8b0d088c75700 git commit sha
```
An entropy detector fires on every one of those, and the resulting redaction is
silent, permanent, and destroys traceability. So detection uses three anchors
that carry meaning instead:
| Tier | Anchor | Default | False-positive risk |
|---|---|---|---|
| **1 — known values** | literal values from this process's env, for variables whose *name* says secret (`…TOKEN`, `…SECRET`, `…PASSWORD`, `…API_KEY`) | **redact** | none by construction: the value *is* the secret |
| **2 — known shapes** | vendor-prefixed credentials: `ghp_…`, `github_pat_…`, `glpat-…`, `xox[abprs]-…`, `sk-…`, `AKIA…`, `AIza…`, `hf_…`, JWTs, PEM private-key blocks, credentials inside URLs, `Authorization:` headers | **redact** | very low: the prefix is meaningful, not random |
| **3 — name=value** | an assignment whose *key* says secret | **report only** | measured **high** — see §3 |
Tier 1 catches any presentation of a secret — env dump, JSON, error message,
URL, prose — because it matches the value itself. It is also the only tier that
resolves this: a 40-hex Gitea PAT is byte-identical in shape to a git commit sha.
## 3. The false-positive measurement, which changed the design
Tier 3 was originally going to redact. Measured against **52 MB of real fleet
transcripts (45 sessions)** it produced **403 hits**, and inspection of them (with
values masked) showed the overwhelming majority were not secrets:
- `GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}` — docker-compose interpolation
- `const tokens = countTokensForModel` — TypeScript source stored as memory
- `refreshToken(credentials: Credentials …)` — a *type annotation*
- `--setattr=krbPasswordExpiration=20260529…` — an IPA attribute holding a date
- `|1.FORGE.TOKENS:…` — AAAK diary shorthand
- `(root@host) Password:` followed by unrelated terminal output
Redacting those would corrupt code, configuration and documentation held as
memory, in order to catch secrets tier 1 already catches by value. After adding
guards for interpolation (`${VAR}`, `$VAR`, `%VAR%`, `{{tpl}}`), code context,
non-secret key suffixes (`…Expiration`, `…Count`, `…Type`) and all-digit values,
the enforced count on the same corpus fell from **403 to 29** — and those 29 are
tier-1 and tier-2 hits, i.e. real credential shapes.
So tier 3 **reports and does not rewrite** by default. Set
`MEMPALACE_REDACT_STRICT=1` to make it enforce, e.g. on a machine whose
transcripts are configuration-heavy rather than code-heavy.
## 4. Known false negatives — stated, not hidden
The scrubber will not catch a novel credential format pasted bare into prose
with no name nearby, a secret belonging to a machine whose environment this
process cannot see, a base64-of-a-secret, or a secret split across lines.
**"The scrubber ran" must never be read as "there are no secrets in here."** To
keep that measurable rather than assumed, two things are reported:
- every scrub prints a count — including `0 redaction(s)`, because printing
nothing is indistinguishable from a scrubber that never ran;
- `suspicions()` reports high-entropy strings it did **not** redact, as
`(length, fingerprint)` pairs rather than values, so the miss rate can be
tracked over time and a recurring fingerprint can be investigated by hand.
Findings never carry the secret. A `Finding` holds the rule, the label, the
length and `sha256(value)[:8]` — enough to recognise the same leak recurring,
not enough to recover it.
**Fail closed.** If the redactor cannot be imported, the feeder refuses to stage
rather than staging unscrubbed (`exit 3`). Override deliberately with
`MEMPALACE_FEED_ALLOW_UNSCRUBBED=1`.
## 5. The server-side layer (not yet implemented)
Three call sites, because the server has three near-duplicate validators rather
than one:
| Path | Function | Covers |
|---|---|---|
| drawers | `config.py` → `sanitize_content()` | `add_drawer`, `update_drawer`, `diary_write`, `checkpoint` — all four route through it |
| events | `logstream.py` → `_sanitize_body()` | `event_append`, and `event_ack` transitively |
| artifacts | `logstream.py` → `put_artifact()` | inlines its own checks; needs its own edit |
Server-side scrubbing is **tier 2 only** (plus optional tier-3 reporting): the
hub cannot see a client's environment, so tier 1 is structurally unavailable
there. That asymmetry is the reason the client hook is not redundant.
## 6. Cleaning up a leak that already landed
1. **Never re-echo the value while hunting it.** Pass it via stdin, never argv
(visible in `ps` on a shared host). Note `get_drawer` is a trap: reading a
drawer in order to redact it prints the secret back into the live transcript.
2. **Redact, don't delete.** Replacing the value in place keeps the mined
transcript's memory value; a stale embedding vector is a cheap price.
3. **Scrub the feeder inbox too**, or the next mine re-files it.
4. **A live session file needs an equal-length in-place overwrite** — the agent
holds it open, so temp-file-plus-rename loses everything appended afterwards.
5. **Expect residue.** SQLite keeps old page content in freed pages until
`VACUUM`, so a raw byte scan still matches after a successful `UPDATE`. Decide
explicitly whether that matters: if the credential store on the same host is
plaintext anyway, it usually does not.
## 7. Usage
```bash
# unit tests, including the must-not-redact corpus of palace id shapes
python3 bin/mempalace_redact.py --self-test
# scrub anything on stdin; report goes to stderr
some-command | python3 bin/mempalace_redact.py > clean.txt
# also list high-entropy strings that were NOT redacted, as fingerprints
python3 bin/mempalace_redact.py --suspicions < transcript.jsonl > clean.jsonl
```