feed: scrub secrets before staging a transcript, name-anchored not entropy-anchored

A palace is mined from transcripts, and transcripts contain whatever the terminal
printed. Measured on this fleet: one leaked bearer token had reached 3 drawers,
13 feeder inbox files across all three devices, and 10 local files spanning 10
days — from an agent inspecting an env var while debugging. That frequency is the
premise: this is a pipeline problem, not a discipline problem.

WHERE. bin/mempalace_redact.py, called from mempalace-pi-session at the point the
staged transcript is written. That single hook covers both transports, because
local mode mines the staged file and remote mode rsyncs that same file
byte-for-byte. Scrubbing operates on the parsed objects rather than the
serialized text, so string values are rewritten while keys, ids and structure are
untouched — scanning raw JSONL instead invents keys like "tapiKey" out of the \t
escape preceding a field name (observed, not theorised).

WHY NOT ENTROPY. The obvious "redact long random-looking strings" is actively
destructive here: drawer ids, chunk ids, event ids, replica ids, HLCs and commit
SHAs are all high-entropy, are the majority of random-looking text in a palace,
and redacting them is silent and permanent. Detection is anchored on meaning
instead: tier 1 literal values from this process's env whose NAME says secret
(zero false positives by construction, and the only tier that can tell a 40-hex
Gitea PAT from a git commit sha); tier 2 vendor-prefixed shapes (ghp_, glpat-,
xox*-, sk-, AKIA, AIza, hf_, JWT, PEM blocks, URL credentials, Authorization
headers); tier 3 name=value assignments.

THE MEASUREMENT THAT CHANGED THE DESIGN. Tier 3 was going to redact. Against
52 MB of real fleet transcripts it produced 403 hits, and inspection with values
masked showed most were ${VAR} interpolation in compose files, TypeScript
identifiers, a TYPE ANNOTATION (credentials: Credentials), an IPA attribute
holding a date (krbPasswordExpiration), AAAK diary shorthand, and terminal output
following a "Password:" prompt. Redacting those corrupts code and docs held as
memory to catch what tier 1 already catches by value. After adding guards for
interpolation, code context, non-secret key suffixes and all-digit values, the
enforced count fell 403 -> 29 on the same corpus. So tier 3 REPORTS and does not
rewrite unless MEMPALACE_REDACT_STRICT=1.

HONESTY ABOUT MISSES. Known false negatives are documented rather than papered
over: novel formats in bare prose, another machine's secrets, base64-of-a-secret,
line-split secrets. Every run prints a count including "0 redaction(s)", because
silence is indistinguishable from a scrubber that never ran, and suspicions()
reports high-entropy strings it did NOT redact as (length, fingerprint) so the
miss rate is measurable. Findings never carry the value — rule, label, length and
sha256[:8], enough to recognise a recurrence, not enough to recover the secret.

FAIL CLOSED: no scrubber, no staging (exit 3), overridable with
MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 for a machine older than this file.

Tested: 42-case corpus in --self-test, including every palace id shape as a
must-not-redact case, idempotency, and a compound case the corpus caught where a
vendor placeholder was then re-matched by the URL rule (nested placeholder — the
secret was hidden either way, so only an exact-output assertion catches it).
End-to-end on 45 real sessions: 29 enforced, fail-closed verified at rc=3,
override verified loud. Server-side layer specified in docs/secret-hygiene.md §5
but NOT implemented — it is tier 2 only there, since the hub cannot see a
client's environment.
This commit is contained in:
2026-08-27 13:30:47 +02:00
parent ecc2a9c574
commit 3d47937d06
4 changed files with 748 additions and 0 deletions
+58
View File
@@ -534,11 +534,42 @@ fi
# Also classifies each export as NEW/ALREADY FILED (by source_file lookup)
# so --dry-run reports the real mine-set size. Classification is advisory;
# `mempalace mine --mode convos` is still the authoritative dedup.
# The redactor is a sibling module, imported by the heredoc below. Exported
# rather than passed as argv so the argv unpack stays stable.
MEMPALACE_REDACT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
export MEMPALACE_REDACT_DIR
export_count=$(python3 - "$PI_SESSIONS_DIR" "$STAGE" "$SESSION_ID" "$SINCE" "$MIN_MESSAGES" "$MIN_ASSISTANT_CHARS" "$MODE" <<'PY'
import json, os, sqlite3, sys
from datetime import datetime, timezone
from pathlib import Path
# ── Secret scrubbing before anything is staged ───────────────────────────────
# This is the last point at which a transcript is a plain in-memory value: the
# local mine reads the staged file, and the remote path rsyncs that same file
# byte-for-byte, so scrubbing here covers BOTH transports with one hook.
#
# FAIL CLOSED. If the redactor cannot be imported, staging is refused rather
# than done unscrubbed — a missing module means a broken install, and the whole
# point of this step is that a secret must not reach a shared palace. Override
# deliberately with MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 if you ever need to feed a
# machine whose toolkit is older than this file.
sys.path.insert(0, os.environ.get("MEMPALACE_REDACT_DIR", ""))
try:
import mempalace_redact as _redact
except Exception as _e: # noqa: BLE001 - any import failure is fatal by design
if os.environ.get("MEMPALACE_FEED_ALLOW_UNSCRUBBED", "").strip() in {"1", "true", "yes"}:
_redact = None
print(" [WARN] secret scrubber unavailable, staging UNSCRUBBED by request "
f"({_e})", file=sys.stderr)
else:
print(f" [FATAL] secret scrubber unavailable ({_e}); refusing to stage. "
"Set MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 to override.", file=sys.stderr)
raise SystemExit(3)
_known = _redact.env_secrets() if _redact else []
_redactions = 0
_reported = 0
sessions_dir, stage, session_filter, since, min_messages, min_assistant_chars, mode = sys.argv[1:8]
min_messages = int(min_messages)
min_assistant_chars = int(min_assistant_chars)
@@ -811,6 +842,26 @@ for path in paths:
)
continue
# Scrub the parsed objects, not the serialized text: string VALUES get
# rewritten while keys, ids and structure are left exactly as they are.
# (Scanning raw JSONL instead would also match escape artifacts like the
# "\\t" before a field name, inventing keys such as "tapiKey".)
if _redact is not None:
_found: list = []
out_lines = [_redact.scrub_obj(obj, _known, _found)[0] for obj in out_lines]
_hard = [x for x in _found if not x.rule.endswith("-reported")]
_soft = [x for x in _found if x.rule.endswith("-reported")]
_redactions += len(_hard)
_reported += len(_soft)
if _hard:
_by = {}
for x in _hard:
_by[x.rule] = _by.get(x.rule, 0) + 1
print(f" [REDACTED] {path.name} "
+ ", ".join(f"{k}={v}" for k, v in sorted(_by.items()))
+ " fp=" + ",".join(sorted({x.fingerprint for x in _hard})),
file=sys.stderr)
out_path = stage / f"pi_{session_uuid}.jsonl"
with out_path.open("w", encoding="utf-8") as f:
for obj in out_lines:
@@ -835,6 +886,13 @@ for path in paths:
print(f"EXPORTED {exported}")
print(f"ALREADY_FILED {-1 if already_filed is None else skipped_already_filed}")
# Report the scrub outcome even when it is zero: "0 redactions" is a measurement,
# whereas printing nothing is indistinguishable from a scrubber that never ran.
if _redact is not None:
print(f" [scrub] {_redactions} redaction(s) applied, "
f"{_reported} name-anchored candidate(s) reported only "
f"(set MEMPALACE_REDACT_STRICT=1 to redact those too)", file=sys.stderr)
if skipped_short:
print(f"SKIPPED_SHORT {skipped_short}", file=sys.stderr)
if skipped_quiet: