A palace is mined from transcripts, and transcripts contain whatever the terminal
printed. Measured on this fleet: one leaked bearer token had reached 3 drawers,
13 feeder inbox files across all three devices, and 10 local files spanning 10
days — from an agent inspecting an env var while debugging. That frequency is the
premise: this is a pipeline problem, not a discipline problem.
WHERE. bin/mempalace_redact.py, called from mempalace-pi-session at the point the
staged transcript is written. That single hook covers both transports, because
local mode mines the staged file and remote mode rsyncs that same file
byte-for-byte. Scrubbing operates on the parsed objects rather than the
serialized text, so string values are rewritten while keys, ids and structure are
untouched — scanning raw JSONL instead invents keys like "tapiKey" out of the \t
escape preceding a field name (observed, not theorised).
WHY NOT ENTROPY. The obvious "redact long random-looking strings" is actively
destructive here: drawer ids, chunk ids, event ids, replica ids, HLCs and commit
SHAs are all high-entropy, are the majority of random-looking text in a palace,
and redacting them is silent and permanent. Detection is anchored on meaning
instead: tier 1 literal values from this process's env whose NAME says secret
(zero false positives by construction, and the only tier that can tell a 40-hex
Gitea PAT from a git commit sha); tier 2 vendor-prefixed shapes (ghp_, glpat-,
xox*-, sk-, AKIA, AIza, hf_, JWT, PEM blocks, URL credentials, Authorization
headers); tier 3 name=value assignments.
THE MEASUREMENT THAT CHANGED THE DESIGN. Tier 3 was going to redact. Against
52 MB of real fleet transcripts it produced 403 hits, and inspection with values
masked showed most were ${VAR} interpolation in compose files, TypeScript
identifiers, a TYPE ANNOTATION (credentials: Credentials), an IPA attribute
holding a date (krbPasswordExpiration), AAAK diary shorthand, and terminal output
following a "Password:" prompt. Redacting those corrupts code and docs held as
memory to catch what tier 1 already catches by value. After adding guards for
interpolation, code context, non-secret key suffixes and all-digit values, the
enforced count fell 403 -> 29 on the same corpus. So tier 3 REPORTS and does not
rewrite unless MEMPALACE_REDACT_STRICT=1.
HONESTY ABOUT MISSES. Known false negatives are documented rather than papered
over: novel formats in bare prose, another machine's secrets, base64-of-a-secret,
line-split secrets. Every run prints a count including "0 redaction(s)", because
silence is indistinguishable from a scrubber that never ran, and suspicions()
reports high-entropy strings it did NOT redact as (length, fingerprint) so the
miss rate is measurable. Findings never carry the value — rule, label, length and
sha256[:8], enough to recognise a recurrence, not enough to recover the secret.
FAIL CLOSED: no scrubber, no staging (exit 3), overridable with
MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 for a machine older than this file.
Tested: 42-case corpus in --self-test, including every palace id shape as a
must-not-redact case, idempotency, and a compound case the corpus caught where a
vendor placeholder was then re-matched by the URL rule (nested placeholder — the
secret was hidden either way, so only an exact-output assertion catches it).
End-to-end on 45 real sessions: 29 enforced, fail-closed verified at rc=3,
override verified loud. Server-side layer specified in docs/secret-hygiene.md §5
but NOT implemented — it is tier 2 only there, since the hub cannot see a
client's environment.
8.3 KiB
Secret hygiene: keeping credentials out of the palace
A palace is mined from transcripts, and transcripts contain whatever the
terminal printed. Print an environment, cat a .env, paste a curl -H "Authorization: …", and the secret becomes a drawer — searchable by every agent
on the fleet, on every machine, indefinitely. This document describes the
scrubbing that exists, what it deliberately does not try to do, and the two
layers it belongs in.
Measured on this fleet, 2026-08-27: a single leaked bearer token had reached 3 drawers, 13 feeder inbox files across all three devices, and 10 local files dating back 10 days. Nobody was careless; an agent inspected an environment variable while debugging. That frequency is the design premise — this is a routine event to be handled by the pipeline, not an incident to be handled by discipline.
1. Where the hook goes, and why there are two of them
| Layer | Implemented in | Sees | Catches |
|---|---|---|---|
| Client, pre-staging | bin/mempalace_redact.py, called from bin/mempalace-pi-session at the point the staged transcript is written |
the machine's own environment and the transcript | the real-world case: a secret this machine holds, printed into a session, before it is uploaded anywhere |
| Server, pre-persist | not yet implemented — see §5 | only the text it is handed | secrets from clients that predate the scrubber, from non-pi clients, and secrets an agent types straight into add_drawer |
The client hook is the load-bearing one, because it is the only layer that can compare text against the actual secret values it holds (see tier 1 below) and because it stops the leak before transmission. The server hook is defence in depth: pattern-only, but it covers write paths the feeder never sees.
Client placement matters for one specific reason: the pi feeder writes the staged
transcript from an in-memory structure, and the remote path then rsyncs that
same file byte-for-byte. Scrubbing at the write point therefore covers local
mining and remote upload with a single hook — there is no second serialization to
forget.
2. Detection is anchored on meaning, never on entropy
The tempting design — "redact long random-looking strings" — is actively destructive here, because a palace is full of high-entropy strings that are its own primary keys:
drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b drawer id
evt_20260826T213924_397b1f710c2d event id
rep_d344e349ba276d6fc11997cd552f6937 replica id
1787780364792-000000-rep_d344e349ba276d6f… hybrid logical clock
ecc2a9c574f4156403cf4e90aea8b0d088c75700 git commit sha
An entropy detector fires on every one of those, and the resulting redaction is silent, permanent, and destroys traceability. So detection uses three anchors that carry meaning instead:
| Tier | Anchor | Default | False-positive risk |
|---|---|---|---|
| 1 — known values | literal values from this process's env, for variables whose name says secret (…TOKEN, …SECRET, …PASSWORD, …API_KEY) |
redact | none by construction: the value is the secret |
| 2 — known shapes | vendor-prefixed credentials: ghp_…, github_pat_…, glpat-…, xox[abprs]-…, sk-…, AKIA…, AIza…, hf_…, JWTs, PEM private-key blocks, credentials inside URLs, Authorization: headers |
redact | very low: the prefix is meaningful, not random |
| 3 — name=value | an assignment whose key says secret | report only | measured high — see §3 |
Tier 1 catches any presentation of a secret — env dump, JSON, error message, URL, prose — because it matches the value itself. It is also the only tier that resolves this: a 40-hex Gitea PAT is byte-identical in shape to a git commit sha.
3. The false-positive measurement, which changed the design
Tier 3 was originally going to redact. Measured against 52 MB of real fleet transcripts (45 sessions) it produced 403 hits, and inspection of them (with values masked) showed the overwhelming majority were not secrets:
GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}— docker-compose interpolationconst tokens = countTokensForModel— TypeScript source stored as memoryrefreshToken(credentials: Credentials …)— a type annotation--setattr=krbPasswordExpiration=20260529…— an IPA attribute holding a date|1.FORGE.TOKENS:…— AAAK diary shorthand(root@host) Password:followed by unrelated terminal output
Redacting those would corrupt code, configuration and documentation held as
memory, in order to catch secrets tier 1 already catches by value. After adding
guards for interpolation (${VAR}, $VAR, %VAR%, {{tpl}}), code context,
non-secret key suffixes (…Expiration, …Count, …Type) and all-digit values,
the enforced count on the same corpus fell from 403 to 29 — and those 29 are
tier-1 and tier-2 hits, i.e. real credential shapes.
So tier 3 reports and does not rewrite by default. Set
MEMPALACE_REDACT_STRICT=1 to make it enforce, e.g. on a machine whose
transcripts are configuration-heavy rather than code-heavy.
4. Known false negatives — stated, not hidden
The scrubber will not catch a novel credential format pasted bare into prose with no name nearby, a secret belonging to a machine whose environment this process cannot see, a base64-of-a-secret, or a secret split across lines.
"The scrubber ran" must never be read as "there are no secrets in here." To keep that measurable rather than assumed, two things are reported:
- every scrub prints a count — including
0 redaction(s), because printing nothing is indistinguishable from a scrubber that never ran; suspicions()reports high-entropy strings it did not redact, as(length, fingerprint)pairs rather than values, so the miss rate can be tracked over time and a recurring fingerprint can be investigated by hand.
Findings never carry the secret. A Finding holds the rule, the label, the
length and sha256(value)[:8] — enough to recognise the same leak recurring,
not enough to recover it.
Fail closed. If the redactor cannot be imported, the feeder refuses to stage
rather than staging unscrubbed (exit 3). Override deliberately with
MEMPALACE_FEED_ALLOW_UNSCRUBBED=1.
5. The server-side layer (not yet implemented)
Three call sites, because the server has three near-duplicate validators rather than one:
| Path | Function | Covers |
|---|---|---|
| drawers | config.py → sanitize_content() |
add_drawer, update_drawer, diary_write, checkpoint — all four route through it |
| events | logstream.py → _sanitize_body() |
event_append, and event_ack transitively |
| artifacts | logstream.py → put_artifact() |
inlines its own checks; needs its own edit |
Server-side scrubbing is tier 2 only (plus optional tier-3 reporting): the hub cannot see a client's environment, so tier 1 is structurally unavailable there. That asymmetry is the reason the client hook is not redundant.
6. Cleaning up a leak that already landed
- Never re-echo the value while hunting it. Pass it via stdin, never argv
(visible in
pson a shared host). Noteget_draweris a trap: reading a drawer in order to redact it prints the secret back into the live transcript. - Redact, don't delete. Replacing the value in place keeps the mined transcript's memory value; a stale embedding vector is a cheap price.
- Scrub the feeder inbox too, or the next mine re-files it.
- A live session file needs an equal-length in-place overwrite — the agent holds it open, so temp-file-plus-rename loses everything appended afterwards.
- Expect residue. SQLite keeps old page content in freed pages until
VACUUM, so a raw byte scan still matches after a successfulUPDATE. Decide explicitly whether that matters: if the credential store on the same host is plaintext anyway, it usually does not.
7. Usage
# unit tests, including the must-not-redact corpus of palace id shapes
python3 bin/mempalace_redact.py --self-test
# scrub anything on stdin; report goes to stderr
some-command | python3 bin/mempalace_redact.py > clean.txt
# also list high-entropy strings that were NOT redacted, as fingerprints
python3 bin/mempalace_redact.py --suspicions < transcript.jsonl > clean.jsonl