feed: scrub secrets before staging a transcript, name-anchored not entropy-anchored

A palace is mined from transcripts, and transcripts contain whatever the terminal
printed. Measured on this fleet: one leaked bearer token had reached 3 drawers,
13 feeder inbox files across all three devices, and 10 local files spanning 10
days — from an agent inspecting an env var while debugging. That frequency is the
premise: this is a pipeline problem, not a discipline problem.

WHERE. bin/mempalace_redact.py, called from mempalace-pi-session at the point the
staged transcript is written. That single hook covers both transports, because
local mode mines the staged file and remote mode rsyncs that same file
byte-for-byte. Scrubbing operates on the parsed objects rather than the
serialized text, so string values are rewritten while keys, ids and structure are
untouched — scanning raw JSONL instead invents keys like "tapiKey" out of the \t
escape preceding a field name (observed, not theorised).

WHY NOT ENTROPY. The obvious "redact long random-looking strings" is actively
destructive here: drawer ids, chunk ids, event ids, replica ids, HLCs and commit
SHAs are all high-entropy, are the majority of random-looking text in a palace,
and redacting them is silent and permanent. Detection is anchored on meaning
instead: tier 1 literal values from this process's env whose NAME says secret
(zero false positives by construction, and the only tier that can tell a 40-hex
Gitea PAT from a git commit sha); tier 2 vendor-prefixed shapes (ghp_, glpat-,
xox*-, sk-, AKIA, AIza, hf_, JWT, PEM blocks, URL credentials, Authorization
headers); tier 3 name=value assignments.

THE MEASUREMENT THAT CHANGED THE DESIGN. Tier 3 was going to redact. Against
52 MB of real fleet transcripts it produced 403 hits, and inspection with values
masked showed most were ${VAR} interpolation in compose files, TypeScript
identifiers, a TYPE ANNOTATION (credentials: Credentials), an IPA attribute
holding a date (krbPasswordExpiration), AAAK diary shorthand, and terminal output
following a "Password:" prompt. Redacting those corrupts code and docs held as
memory to catch what tier 1 already catches by value. After adding guards for
interpolation, code context, non-secret key suffixes and all-digit values, the
enforced count fell 403 -> 29 on the same corpus. So tier 3 REPORTS and does not
rewrite unless MEMPALACE_REDACT_STRICT=1.

HONESTY ABOUT MISSES. Known false negatives are documented rather than papered
over: novel formats in bare prose, another machine's secrets, base64-of-a-secret,
line-split secrets. Every run prints a count including "0 redaction(s)", because
silence is indistinguishable from a scrubber that never ran, and suspicions()
reports high-entropy strings it did NOT redact as (length, fingerprint) so the
miss rate is measurable. Findings never carry the value — rule, label, length and
sha256[:8], enough to recognise a recurrence, not enough to recover the secret.

FAIL CLOSED: no scrubber, no staging (exit 3), overridable with
MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 for a machine older than this file.

Tested: 42-case corpus in --self-test, including every palace id shape as a
must-not-redact case, idempotency, and a compound case the corpus caught where a
vendor placeholder was then re-matched by the URL rule (nested placeholder — the
secret was hidden either way, so only an exact-output assertion catches it).
End-to-end on 45 real sessions: 29 enforced, fail-closed verified at rc=3,
override verified loud. Server-side layer specified in docs/secret-hygiene.md §5
but NOT implemented — it is tier 2 only there, since the hub cannot see a
client's environment.
This commit is contained in:
2026-08-27 13:30:47 +02:00
parent ecc2a9c574
commit 3d47937d06
4 changed files with 748 additions and 0 deletions
+1
View File
@@ -66,6 +66,7 @@ Start here if you are deciding **what to put where**, or running MemPalace on mo
|---|---| |---|---|
| [`docs/fleet-memory.md`](docs/fleet-memory.md) | **Start here.** The five stores MemPalace keeps and what each is for; what a *central* palace buys a fleet of machines; when to file a memory versus when to message another device. Diagrams, worked examples. | | [`docs/fleet-memory.md`](docs/fleet-memory.md) | **Start here.** The five stores MemPalace keeps and what each is for; what a *central* palace buys a fleet of machines; when to file a memory versus when to message another device. Diagrams, worked examples. |
| [`docs/rfc-003-coordination-log.md`](docs/rfc-003-coordination-log.md) | The inter-agent / inter-device coordination log (`logstream`): storage, append and query semantics, delivery and latency, the trust model, and its landmines. | | [`docs/rfc-003-coordination-log.md`](docs/rfc-003-coordination-log.md) | The inter-agent / inter-device coordination log (`logstream`): storage, append and query semantics, delivery and latency, the trust model, and its landmines. |
| [`docs/secret-hygiene.md`](docs/secret-hygiene.md) | Keeping credentials out of the palace — what the feeder scrubs before staging, why detection is name-anchored and never entropy-anchored, and the measured false-positive rate that made one tier report-only |
| [`docs/rfc-001-global-palace.md`](docs/rfc-001-global-palace.md) | Why and how a palace is centralised, what should *not* be global, and the Phase-0 landmines (including: never run `mempalace sync` against a shared palace). | | [`docs/rfc-001-global-palace.md`](docs/rfc-001-global-palace.md) | Why and how a palace is centralised, what should *not* be global, and the Phase-0 landmines (including: never run `mempalace sync` against a shared palace). |
| [`docs/rfc-002-joiner.md`](docs/rfc-002-joiner.md) | Replaying a second palace into a shared primary, and what dedupes what. | | [`docs/rfc-002-joiner.md`](docs/rfc-002-joiner.md) | Replaying a second palace into a shared primary, and what dedupes what. |
| [`docs/phase-1-exposure-runbook.md`](docs/phase-1-exposure-runbook.md) | *(moved)* The exposure record for one deployment is now private operator data. The stub names the reusable mechanism it also carried — the `Host`/`Origin` pin above all — which is not yet published elsewhere. | | [`docs/phase-1-exposure-runbook.md`](docs/phase-1-exposure-runbook.md) | *(moved)* The exposure record for one deployment is now private operator data. The stub names the reusable mechanism it also carried — the `Host`/`Origin` pin above all — which is not yet published elsewhere. |
+58
View File
@@ -534,11 +534,42 @@ fi
# Also classifies each export as NEW/ALREADY FILED (by source_file lookup) # Also classifies each export as NEW/ALREADY FILED (by source_file lookup)
# so --dry-run reports the real mine-set size. Classification is advisory; # so --dry-run reports the real mine-set size. Classification is advisory;
# `mempalace mine --mode convos` is still the authoritative dedup. # `mempalace mine --mode convos` is still the authoritative dedup.
# The redactor is a sibling module, imported by the heredoc below. Exported
# rather than passed as argv so the argv unpack stays stable.
MEMPALACE_REDACT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
export MEMPALACE_REDACT_DIR
export_count=$(python3 - "$PI_SESSIONS_DIR" "$STAGE" "$SESSION_ID" "$SINCE" "$MIN_MESSAGES" "$MIN_ASSISTANT_CHARS" "$MODE" <<'PY' export_count=$(python3 - "$PI_SESSIONS_DIR" "$STAGE" "$SESSION_ID" "$SINCE" "$MIN_MESSAGES" "$MIN_ASSISTANT_CHARS" "$MODE" <<'PY'
import json, os, sqlite3, sys import json, os, sqlite3, sys
from datetime import datetime, timezone from datetime import datetime, timezone
from pathlib import Path from pathlib import Path
# ── Secret scrubbing before anything is staged ───────────────────────────────
# This is the last point at which a transcript is a plain in-memory value: the
# local mine reads the staged file, and the remote path rsyncs that same file
# byte-for-byte, so scrubbing here covers BOTH transports with one hook.
#
# FAIL CLOSED. If the redactor cannot be imported, staging is refused rather
# than done unscrubbed — a missing module means a broken install, and the whole
# point of this step is that a secret must not reach a shared palace. Override
# deliberately with MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 if you ever need to feed a
# machine whose toolkit is older than this file.
sys.path.insert(0, os.environ.get("MEMPALACE_REDACT_DIR", ""))
try:
import mempalace_redact as _redact
except Exception as _e: # noqa: BLE001 - any import failure is fatal by design
if os.environ.get("MEMPALACE_FEED_ALLOW_UNSCRUBBED", "").strip() in {"1", "true", "yes"}:
_redact = None
print(" [WARN] secret scrubber unavailable, staging UNSCRUBBED by request "
f"({_e})", file=sys.stderr)
else:
print(f" [FATAL] secret scrubber unavailable ({_e}); refusing to stage. "
"Set MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 to override.", file=sys.stderr)
raise SystemExit(3)
_known = _redact.env_secrets() if _redact else []
_redactions = 0
_reported = 0
sessions_dir, stage, session_filter, since, min_messages, min_assistant_chars, mode = sys.argv[1:8] sessions_dir, stage, session_filter, since, min_messages, min_assistant_chars, mode = sys.argv[1:8]
min_messages = int(min_messages) min_messages = int(min_messages)
min_assistant_chars = int(min_assistant_chars) min_assistant_chars = int(min_assistant_chars)
@@ -811,6 +842,26 @@ for path in paths:
) )
continue continue
# Scrub the parsed objects, not the serialized text: string VALUES get
# rewritten while keys, ids and structure are left exactly as they are.
# (Scanning raw JSONL instead would also match escape artifacts like the
# "\\t" before a field name, inventing keys such as "tapiKey".)
if _redact is not None:
_found: list = []
out_lines = [_redact.scrub_obj(obj, _known, _found)[0] for obj in out_lines]
_hard = [x for x in _found if not x.rule.endswith("-reported")]
_soft = [x for x in _found if x.rule.endswith("-reported")]
_redactions += len(_hard)
_reported += len(_soft)
if _hard:
_by = {}
for x in _hard:
_by[x.rule] = _by.get(x.rule, 0) + 1
print(f" [REDACTED] {path.name} "
+ ", ".join(f"{k}={v}" for k, v in sorted(_by.items()))
+ " fp=" + ",".join(sorted({x.fingerprint for x in _hard})),
file=sys.stderr)
out_path = stage / f"pi_{session_uuid}.jsonl" out_path = stage / f"pi_{session_uuid}.jsonl"
with out_path.open("w", encoding="utf-8") as f: with out_path.open("w", encoding="utf-8") as f:
for obj in out_lines: for obj in out_lines:
@@ -835,6 +886,13 @@ for path in paths:
print(f"EXPORTED {exported}") print(f"EXPORTED {exported}")
print(f"ALREADY_FILED {-1 if already_filed is None else skipped_already_filed}") print(f"ALREADY_FILED {-1 if already_filed is None else skipped_already_filed}")
# Report the scrub outcome even when it is zero: "0 redactions" is a measurement,
# whereas printing nothing is indistinguishable from a scrubber that never ran.
if _redact is not None:
print(f" [scrub] {_redactions} redaction(s) applied, "
f"{_reported} name-anchored candidate(s) reported only "
f"(set MEMPALACE_REDACT_STRICT=1 to redact those too)", file=sys.stderr)
if skipped_short: if skipped_short:
print(f"SKIPPED_SHORT {skipped_short}", file=sys.stderr) print(f"SKIPPED_SHORT {skipped_short}", file=sys.stderr)
if skipped_quiet: if skipped_quiet:
+538
View File
@@ -0,0 +1,538 @@
#!/usr/bin/env python3
"""Scrub secret-shaped values out of palace-bound text.
WHY THIS IS NAME-ANCHORED AND NOT ENTROPY-ANCHORED
==================================================
The obvious design — "redact long random-looking strings" — is actively wrong
for this corpus, and the reason is worth stating before anyone tries to
"improve" it. A palace is *full* of high-entropy strings that are its own
primary keys:
drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b drawer id
..._chunk_000000 chunk id
evt_20260826T213924_397b1f710c2d event id
rep_d344e349ba276d6fc11997cd552f6937 replica id
1787780364792-000000-rep_d344e349ba276d6f… hybrid logical clock
ecc2a9c574f4156403cf4e90aea8b0d088c75700 git commit sha
An entropy or "looks like base64/hex of length N" detector fires on every one
of those, i.e. on most identifiers in the corpus, and the redaction damages the
memory it was meant to protect. Worse, the damage is silent and unrecoverable:
a mined transcript whose ids have been replaced with <redacted> is no longer
traceable to anything.
So detection here is anchored on one of three things that carry *meaning*:
Tier 1 KNOWN VALUES — literal values taken from this process's own
environment, for variables whose NAME says secret.
Zero false positives by construction: the value
*is* the secret. Catches any presentation (env
dump, JSON, error message, URL, prose), which is
exactly the incident this exists for.
Tier 2 KNOWN SHAPES — vendor-prefixed credentials (ghp_…, glpat-…,
xox[abprs]-…, AKIA…, sk-ant-…, JWTs, PEM private
keys, credentials embedded in URLs, Authorization
headers). Very low false-positive rate because the
prefix is meaningful, not merely random.
Tier 3 NAME=VALUE — an assignment whose *key* says secret
(…TOKEN, …SECRET, …PASSWORD, …API_KEY, …).
This is the tier that resolves shape collisions
that no shape-based rule can: a 40-hex Gitea PAT
is byte-indistinguishable from a git commit sha,
and `GITEA_TOKEN=<40 hex>` is the only thing that
tells them apart.
KNOWN FALSE NEGATIVES — stated, not hidden
------------------------------------------
This will not catch: a novel credential format pasted bare into prose with no
name nearby; a secret from a machine whose env this process cannot see; a
base64-of-a-secret; a secret split across lines. Tier 1 covers the local
machine's own secrets completely, which is where the measured incidents came
from — but "the scrubber ran" must never be read as "there are no secrets in
here". `suspicions()` exists for exactly that reason: it reports high-entropy
candidates it did NOT redact, as fingerprints rather than values, so the false
negative rate can be *measured* over time instead of assumed to be zero.
Run `python3 mempalace_redact.py --self-test` for the test corpus, which
includes every benign id shape above as a must-not-redact case.
"""
from __future__ import annotations
import hashlib
import math
import os
import re
from typing import Any, Callable, Iterable
__all__ = ["scrub", "scrub_obj", "suspicions", "Finding"]
PLACEHOLDER = "<redacted:{label}>"
# ── Tier 1: names whose VALUES are secrets ────────────────────────────────────
# Anchored on the variable name, then applied as an exact literal match.
_SECRET_NAME = re.compile(
r"(TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|API_?KEY|APIKEY|ACCESS_?KEY"
r"|PRIVATE_?KEY|CREDENTIAL|BEARER|SESSION_?KEY|COOKIE|_PAT|^PAT$)",
re.IGNORECASE,
)
# …unless the name says it holds a location or a knob rather than a secret.
# SSH_KEY_PATH is the live example: it matches KEY, and its value is a path.
_NAME_NOT_SECRET = re.compile(
r"(_PATH|_FILE|_DIR|_NAME|_ID|_URL|_URI|_HOST|_PORT|_USER|_ENABLED"
r"|_TIMEOUT|_MS|_SECS|_SECONDS|_COUNT|_LIMIT|_MODE)$",
re.IGNORECASE,
)
# Keys that contain a secret word but describe a POLICY, COUNT or TYPE rather
# than holding a credential. Measured on real transcripts: krbPasswordExpiration
# (an IPA attribute whose value is a date), observationTokens (a token count),
# targetTokens, tokensOverTarget.
_KEY_NOT_A_SECRET = re.compile(
r"(expiration|expiry|_age|count|length|limit|usage|interval|policy|type|class"
r"|provider|manager|error|target|overtarget|file|path|dir|name|_id|url|uri"
r"|host|port|user|enabled|timeout|_ms|_secs)$",
re.IGNORECASE,
)
# Left of the match: a declaration keyword or a member access means this is CODE
# (`const token = x`, `obj.apiKey = y`), and the "value" is an expression.
_CODE_CONTEXT = re.compile(
r"(?:\b(?:const|let|var|def|func|fn|public|private|readonly|export|return|if|assert)\s+$)"
r"|[.\w]\s*$"
)
_TRIVIAL_VALUES = {
"", "0", "1", "true", "false", "none", "null", "nil", "unset", "changeme",
"change-me", "xxx", "***", "redacted", "your-token-here", "your_token_here",
"placeholder", "example", "dummy", "test", "secret",
}
# ── Tier 2: shapes that mean credential because the PREFIX means credential ──
_SHAPE_RULES: list[tuple[str, re.Pattern[str], Callable[[re.Match[str]], str]]] = [
# ORDER MATTERS, and the self-test is what proved it. The two structural
# rules (credentials inside a URL, Authorization header) run FIRST, before
# the vendor-prefix rules: a vendor rule rewriting `https://ghp_xxx@host` to
# `https://<redacted:github-token>@host` leaves a colon INSIDE the
# placeholder, which the URL rule then reads as user:password and redacts a
# second time, producing a nested placeholder. Ordering plus the (?!<redacted)
# guards below make the result independent of pass order.
("url-credentials",
re.compile(r"(?P<pre>[a-zA-Z][a-zA-Z0-9+.\-]*://(?!<redacted)[^/\s:@]+:)(?P<pw>(?!<redacted)[^/\s@]{4,})(?P<at>@)"),
lambda m: f"{m.group('pre')}{PLACEHOLDER.format(label='url-password')}{m.group('at')}"),
("authorization-header",
re.compile(r"(?P<pre>[Aa]uthorization:\s*(?:Bearer|Basic|token)\s+)(?P<val>(?!<redacted)[A-Za-z0-9._\-+/=]{8,})"),
lambda m: f"{m.group('pre')}{PLACEHOLDER.format(label='auth-header')}"),
("github-pat", re.compile(r"\bgh[pousr]_[A-Za-z0-9]{16,}\b"),
lambda m: PLACEHOLDER.format(label="github-token")),
("github-fine-grained", re.compile(r"\bgithub_pat_[A-Za-z0-9_]{22,}\b"),
lambda m: PLACEHOLDER.format(label="github-token")),
("gitlab-pat", re.compile(r"\bglpat-[A-Za-z0-9_\-]{16,}\b"),
lambda m: PLACEHOLDER.format(label="gitlab-token")),
("slack", re.compile(r"\bxox[abprs]-[A-Za-z0-9\-]{10,}\b"),
lambda m: PLACEHOLDER.format(label="slack-token")),
("openai-anthropic", re.compile(r"\bsk-(?:ant-)?[A-Za-z0-9_\-]{20,}\b"),
lambda m: PLACEHOLDER.format(label="api-key")),
("aws-access-key", re.compile(r"\b(?:AKIA|ASIA)[0-9A-Z]{16}\b"),
lambda m: PLACEHOLDER.format(label="aws-access-key-id")),
("google-api-key", re.compile(r"\bAIza[0-9A-Za-z_\-]{35}\b"),
lambda m: PLACEHOLDER.format(label="google-api-key")),
("huggingface", re.compile(r"\bhf_[A-Za-z0-9]{30,}\b"),
lambda m: PLACEHOLDER.format(label="hf-token")),
("npm", re.compile(r"\bnpm_[A-Za-z0-9]{30,}\b"),
lambda m: PLACEHOLDER.format(label="npm-token")),
("digitalocean", re.compile(r"\bdop_v1_[a-f0-9]{60,}\b"),
lambda m: PLACEHOLDER.format(label="do-token")),
# JWT: three base64url segments. Requires the "eyJ" header so a random
# dotted string cannot trip it.
("jwt", re.compile(r"\beyJ[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}\b"),
lambda m: PLACEHOLDER.format(label="jwt")),
# Whole PEM block, non-greedy so several in one file stay separate.
("pem-private-key",
re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----.*?-----END [A-Z ]*PRIVATE KEY-----", re.DOTALL),
lambda m: PLACEHOLDER.format(label="private-key-block")),
]
# ── Tier 3: assignment whose KEY says secret ─────────────────────────────────
# The key prefix is LAZY AND OPTIONAL, not "one char then anything": a mandatory
# leading character cannot match a key that BEGINS with the secret word, so
# `"api_key"` was missed while `"client_secret"` passed — the secret word sits
# mid-name in one and at position 0 in the other. Note plain `KEY` is absent from
# the alternation on purpose: it is too generic (SSH_KEY_PATH), and tier 1 covers
# the real ones by value anyway.
# The optional quote after the key is what makes JSON work: `"api_key": "v"`
# closes the key BEFORE the colon, so a pattern jumping straight from key to
# separator silently misses every JSON payload — i.e. most of a transcript.
_ASSIGNMENT = re.compile(
r"""(?P<key>\b[A-Za-z0-9_.\-]*?
(?:token|secret|password|passwd|passphrase|api[_-]?key|apikey
|access[_-]?key|private[_-]?key|credential)
[A-Za-z0-9_.\-]*)
(?P<kq>["']?)
(?P<sep>\s*(?:=>|[=:])\s*)
(?P<q>["']?)
(?P<val>[^\s"'<>,;)\]}]{12,})
(?P=q)""",
re.IGNORECASE | re.VERBOSE,
)
# A value that is obviously not a live secret: a reference, a path, a
# placeholder, or something we already redacted.
_NOT_A_SECRET_VALUE = re.compile(
r"""^(?:
\$\{[^}\s]*\}? # ${VAR} ${VAR:-default} ${VAR:?err}
# trailing } optional: the value
# charset stops before it, so the
# captured text is "${VAR" — which
# still means "a reference", and
# missing that is what made a
# compose file look like a secret
| \$[A-Za-z_][A-Za-z0-9_]* # $VAR
| %[A-Za-z_][A-Za-z0-9_]*%? # %VAR%
| \{\{[^}\s]*\}?\}? # {{ template }}
| <[^>]*> # <redacted:…> / <your-token>
| (?:/|~/|\./|\.\./)[^\s]* # a path
| [A-Za-z]:\\ # a windows path
| \*+ | x+ | \.+ # ***, xxx
| (?:true|false|none|null|nil|unset|changeme|placeholder|example)
)$""",
re.IGNORECASE | re.VERBOSE,
)
class Finding:
"""One redaction, recorded without ever carrying the secret itself."""
__slots__ = ("rule", "label", "length", "fingerprint")
def __init__(self, rule: str, label: str, value: str) -> None:
self.rule = rule
self.label = label
self.length = len(value)
# Stable across sessions, so a recurring leak is recognisable as the
# same one, while the value stays unrecoverable from the report.
self.fingerprint = hashlib.sha256(value.encode("utf-8")).hexdigest()[:8]
def __repr__(self) -> str: # pragma: no cover - diagnostics only
return f"<Finding {self.rule}:{self.label} len={self.length} fp={self.fingerprint}>"
def env_secrets(environ: dict[str, str] | None = None) -> list[tuple[str, str]]:
"""Tier 1 material: (name, value) pairs from the environment worth scrubbing.
Deliberately conservative: a secret-sounding NAME is not enough, the value
must also be long enough to be a credential and must not look like a path
or a knob. Longest values first so that if one secret contains another as a
substring, the longer one is replaced before the shorter can fragment it.
"""
env = os.environ if environ is None else environ
out: list[tuple[str, str]] = []
for name, value in env.items():
if not value or len(value) < 12:
continue
if not _SECRET_NAME.search(name) or _NAME_NOT_SECRET.search(name):
continue
if value.strip().lower() in _TRIVIAL_VALUES:
continue
if _NOT_A_SECRET_VALUE.match(value.strip()):
continue
if "/" in value and (value.startswith("/") or value.startswith("~")):
continue # a path that happened to be named …_KEY
if any(c.isspace() for c in value):
continue # credentials do not contain whitespace; prose does
out.append((name, value))
out.sort(key=lambda kv: len(kv[1]), reverse=True)
return out
def scrub(
text: str,
known: Iterable[tuple[str, str]] | None = None,
findings: list[Finding] | None = None,
strict: bool | None = None,
) -> tuple[str, list[Finding]]:
"""Redact secrets from one string. Returns (scrubbed, findings).
Idempotent: the replacement text matches none of the rules, so running this
twice changes nothing and cannot nest placeholders.
"""
found: list[Finding] = findings if findings is not None else []
if strict is None:
strict = os.environ.get("MEMPALACE_REDACT_STRICT", "").strip() in {"1", "true", "yes"}
if not text:
return text, found
# Tier 1 first: an exact known value should be labelled with its variable
# name (the most useful thing a reader can be told) rather than by shape.
for name, value in (env_secrets() if known is None else known):
if value and value in text:
count = text.count(value)
text = text.replace(value, PLACEHOLDER.format(label=name))
for _ in range(count):
found.append(Finding("env-value", name, value))
# Tier 2.
for rule, pattern, repl in _SHAPE_RULES:
def _sub(m: re.Match[str], _rule: str = rule) -> str:
found.append(Finding(_rule, _rule, m.group(0)))
return repl(m)
text = pattern.sub(_sub, text)
# Tier 3 — REPORT-ONLY unless strict. See STRICT note in the module docstring:
# measured on 52 MB of real fleet transcripts this rule produced 403 hits of
# which the overwhelming majority were ${VAR} interpolation, TypeScript
# identifiers, YAML env lists and documentation prose. Redacting those would
# corrupt code and docs stored as memory, to catch secrets that tier 1
# already catches by value. So by default a tier-3 hit is REPORTED (so the
# miss can be measured and the rule calibrated) and NOT rewritten.
def _sub_assignment(m: re.Match[str]) -> str:
val, key = m.group("val"), m.group("key")
pre = m.string[max(0, m.start() - 24):m.start()]
if (_NOT_A_SECRET_VALUE.match(val)
or val.strip().lower() in _TRIVIAL_VALUES
or _KEY_NOT_A_SECRET.search(key)
or _CODE_CONTEXT.search(pre)
or val.isdigit()):
return m.group(0)
found.append(Finding("named-assignment" if strict else "named-assignment-reported",
key, val))
if not strict:
return m.group(0)
return f"{m.group('key')}{m.group('kq')}{m.group('sep')}{m.group('q')}" \
f"{PLACEHOLDER.format(label=m.group('key'))}{m.group('q')}"
text = _ASSIGNMENT.sub(_sub_assignment, text)
return text, found
def scrub_obj(obj: Any, known: Iterable[tuple[str, str]] | None = None,
findings: list[Finding] | None = None,
strict: bool | None = None) -> tuple[Any, list[Finding]]:
"""Recursively scrub every string VALUE in a JSON-ish structure.
Keys are left alone: a key named "token" is metadata, not a credential, and
rewriting keys would corrupt the transcript schema.
"""
found: list[Finding] = findings if findings is not None else []
known = list(env_secrets() if known is None else known)
if isinstance(obj, str):
return scrub(obj, known, found, strict)[0], found
if isinstance(obj, dict):
return {k: scrub_obj(v, known, found, strict)[0] for k, v in obj.items()}, found
if isinstance(obj, list):
return [scrub_obj(v, known, found, strict)[0] for v in obj], found
return obj, found
# ── Measuring what we MISS, without leaking it ───────────────────────────────
_BENIGN_ID = re.compile(
r"""^(?:
drawer_[\w.\-]+ # drawer / chunk ids
| (?:evt|art)_\d{8}T\d{6}_[0-9a-f]{12} # event / artifact ids
| rep_(?:[0-9a-f]{12}|[0-9a-f]{32}) # replica ids
| \d{13}-[0-9a-f]{6}-rep_[0-9a-f]+ # hybrid logical clocks
| [0-9a-f]{7,12} # short git sha
| [0-9a-f]{40} # git sha / sha1
| [0-9a-f]{64} # sha256
| [0-9a-f]{8}(?:-[0-9a-f]{4}){3}-[0-9a-f]{12} # uuid
| sha256:[0-9a-f]{64}
| \d+
)$""",
re.VERBOSE | re.IGNORECASE,
)
_CANDIDATE = re.compile(r"[A-Za-z0-9+/_\-=]{24,}")
def _entropy(s: str) -> float:
if not s:
return 0.0
counts: dict[str, int] = {}
for ch in s:
counts[ch] = counts.get(ch, 0) + 1
n = len(s)
return -sum((c / n) * math.log2(c / n) for c in counts.values())
def suspicions(text: str, min_entropy: float = 3.6) -> list[Finding]:
"""High-entropy strings this module did NOT redact, as fingerprints.
This is the false-negative meter. It deliberately does not redact anything:
every one of these is far more likely to be one of the palace's own ids than
a credential, and redacting them would be the damage described at the top of
this file. Reporting them as (length, fingerprint) makes the miss rate
observable — and a fingerprint that keeps recurring is worth a human look.
"""
out: list[Finding] = []
for m in _CANDIDATE.finditer(text):
val = m.group(0)
if _BENIGN_ID.match(val) or val.startswith("<redacted:"):
continue
if _entropy(val) < min_entropy:
continue
out.append(Finding("suspicion", "high-entropy", val))
return out
# ── Self-test ────────────────────────────────────────────────────────────────
def _self_test() -> int:
fake_token = "pFrDBfakVKQB0SLN1hTzAoaRJeQ4J_go7xmJezojAwI" # shape-alike, not real
env = {
"MEMPALACE_REMOTE_TOKEN": fake_token,
"GITEA_TOKEN": "0123456789abcdef0123456789abcdef01234567", # 40 hex: sha-shaped
"SSH_KEY_PATH": "/Users/someone/.ssh", # name matches KEY, value is a path
"MEMPALACE_MAILBOX_POLL_MS": "300000", # knob, not a secret
"MEMPALACE_REMOTE_URL": "https://palace.example.com/mcp",
"PASSWORD": "changeme", # trivial value
"API_KEY": "${FROM_VAULT}", # a reference
}
known = env_secrets(env)
must_redact = [
("env dump line", f"MEMPALACE_REMOTE_TOKEN={fake_token}"),
("bare value in prose", f"the token is {fake_token} apparently"),
("value inside json", f'{{"env": {{"MEMPALACE_REMOTE_TOKEN": "{fake_token}"}}}}'),
("sha-shaped pat, name-anchored",
"GITEA_TOKEN=0123456789abcdef0123456789abcdef01234567"),
("github pat", "remote add origin https://ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345@git"),
("gitlab pat", "GLPAT is glpat-AbCdEfGhIjKlMnOpQrSt"),
("slack", "xoxb-1234567890-ABCdefGHIjklMNO"),
("anthropic-ish", "sk-ant-api03-AbCdEfGhIjKlMnOpQrStUvWx"),
("aws", "AKIAIOSFODNN7EXAMPLE"),
("jwt", "eyJhbGciOiJIUzI1NiJ9.eyJzdWIiOiIxMjM0NTY3ODkwIn0.dozjgNryP4J3jVmNHl0w5N_XgL0n3I9PlFUP0THsR8U"),
("pem", "-----BEGIN OPENSSH PRIVATE KEY-----\nb3BlbnNzaA\n-----END OPENSSH PRIVATE KEY-----"),
("url creds", "git clone https://joakim:hunter2hunter2@git.example.com/x.git"),
("auth header", "Authorization: Bearer abcdefghijklmnop"),
("json api_key", '"api_key": "AbCdEfGhIjKlMnOpQrSt"'),
("json nested quotes", '{"auth": {"client_secret": "AbCdEfGhIjKlMnOpQrStUv"}}'),
("yaml style", " gitea_token: 0123456789abcdefghij"),
("lowercase password kv", "db_password = s3cr3t-p4ssw0rd-xyz"),
]
must_not_redact = [
("drawer id", "drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b"),
("chunk id", "drawer_pi-devbox_landmines_ca8929fea82cf80e7770ea0d_chunk_000000"),
("event id", "evt_20260826T213924_397b1f710c2d"),
("replica id", "rep_d344e349ba276d6fc11997cd552f6937"),
("hlc", "1787780364792-000000-rep_d344e349ba276d6fc11997cd552f6937"),
("git sha", "ecc2a9c574f4156403cf4e90aea8b0d088c75700"),
("short sha", "ecc2a9c"),
("sha256", "a" * 64),
("uuid session file", "pi_01a03fd8-0dc6-73c2-a9c3-6b62ccf0318f.jsonl"),
("knob assignment", "MEMPALACE_MAILBOX_POLL_MS=300000"),
("path named key", "SSH_KEY_PATH=/Users/someone/.ssh"),
("empty token", "MEMPALACE_REMOTE_TOKEN="),
("already redacted", "MEMPALACE_REMOTE_TOKEN=<redacted:MEMPALACE_REMOTE_TOKEN>"),
("var reference", "MEMPALACE_REMOTE_TOKEN=${VAULT_TOKEN}"),
("placeholder", "MEMPALACE_REMOTE_TOKEN=<your-token-here>"),
("trivial", "PASSWORD=changeme"),
("url no creds", "MEMPALACE_REMOTE_URL=https://palace.example.com/mcp"),
("prose", "The mailbox is an obligation channel, not a news channel."),
("base64 in a diff line", "+ data = 'QUJDREVGR0hJSktMTU5PUFFSU1RVVldYWVo='"),
]
failures = 0
# Tier 3 is report-only by default, so the corpus below is checked in STRICT
# mode (where tier 3 rewrites) and the default mode is asserted separately.
print("MUST REDACT (strict: tier 3 enforcing)")
for name, sample in must_redact:
out, found = scrub(sample, known, strict=True)
ok = bool(found) and all(
secret not in out for secret in (fake_token, "hunter2hunter2",
"0123456789abcdef0123456789abcdef01234567")
)
failures += not ok
rules = ",".join(sorted({f.rule for f in found})) or "NONE"
print(f" {'pass' if ok else 'FAIL'} {name:32s} [{rules}]")
print("MUST NOT REDACT (strict: the hard cases)")
for name, sample in must_not_redact:
out, found = scrub(sample, known, strict=True)
ok = out == sample and not found
failures += not ok
detail = "" if ok else f" -> {out!r} {found}"
print(f" {'pass' if ok else 'FAIL'} {name:32s}{detail}")
print("DEFAULT MODE (tier 3 reports, does not rewrite)")
for name, sample, expect_rewrite in [
("env value still redacted", f"MEMPALACE_REMOTE_TOKEN={fake_token}", True),
("vendor shape still redacted", "ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345", True),
("name-anchored only reported", "db_password = s3cr3t-p4ssw0rd-xyz", False),
("compose interpolation untouched", "GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}", False),
("typescript identifier untouched", "const tokens = countTokensForModel", False),
("ipa attribute untouched", "--setattr=krbPasswordExpiration=20260529090000", False),
]:
out, found = scrub(sample, known)
rewritten = out != sample
ok = rewritten == expect_rewrite
if name == "name-anchored only reported":
ok = ok and any(f.rule == "named-assignment-reported" for f in found)
if name.endswith("untouched"):
ok = ok and not any(f.rule.startswith("named-assignment") for f in found)
failures += not ok
print(f" {'pass' if ok else 'FAIL'} {name:34s}"
+ ("" if ok else f" -> rewritten={rewritten} {found}"))
print("PROPERTIES")
once, _ = scrub(f"MEMPALACE_REMOTE_TOKEN={fake_token}", known, strict=True)
twice, second = scrub(once, known, strict=True)
ok = once == twice and not second
failures += not ok
print(f" {'pass' if ok else 'FAIL'} idempotent (no nested placeholders)")
# The compound case the corpus caught: assert on EXACT output, not merely
# that the secret is gone — a nested placeholder also hides the secret, and
# would have shipped unnoticed.
compound, _ = scrub("https://ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345@git.example.com/x", known, strict=True)
ok = compound == "https://<redacted:github-token>@git.example.com/x"
failures += not ok
print(f" {'pass' if ok else 'FAIL'} vendor placeholder not re-redacted by url rule"
+ ("" if ok else f" -> {compound!r}"))
already, found_a = scrub("https://joakim:<redacted:url-password>@git.example.com", known, strict=True)
ok = already == "https://joakim:<redacted:url-password>@git.example.com" and not found_a
failures += not ok
print(f" {'pass' if ok else 'FAIL'} already-redacted url left alone")
obj = {"role": "user", "content": [{"type": "text", "text": f"export TOK={fake_token}"}],
"id": "evt_20260826T213924_397b1f710c2d"}
scrubbed, found = scrub_obj(obj, known, strict=True)
ok = (fake_token not in repr(scrubbed)
and scrubbed["id"] == obj["id"]
and scrubbed["role"] == "user"
and len(found) >= 1)
failures += not ok
print(f" {'pass' if ok else 'FAIL'} nested structure scrubbed, ids and keys preserved")
ok = all(fake_token not in repr(f) for f in scrub(f"x {fake_token}", known, strict=True)[1])
failures += not ok
print(f" {'pass' if ok else 'FAIL'} findings never carry the secret value")
sus = suspicions("drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b "
"rep_d344e349ba276d6fc11997cd552f6937 "
"1787780364792-000000-rep_d344e349ba276d6fc11997cd552f6937 "
"ecc2a9c574f4156403cf4e90aea8b0d088c75700")
ok = not sus
failures += not ok
print(f" {'pass' if ok else 'FAIL'} suspicion meter ignores our own id shapes ({len(sus)} hits)")
sus2 = suspicions("blob=Zm9vYmFyYmF6cXV1eHF1dXhmb29iYXJiYXpxdXV4Zm9vYmFy")
ok = len(sus2) == 1 and all("Zm9v" not in repr(s) for s in sus2)
failures += not ok
print(f" {'pass' if ok else 'FAIL'} suspicion meter flags unknown blobs as fingerprints only")
print(f"\n{'ALL PASS' if not failures else f'{failures} FAILED'}")
return 1 if failures else 0
if __name__ == "__main__":
import sys
if "--self-test" in sys.argv:
raise SystemExit(_self_test())
# Filter mode: scrub stdin -> stdout, report to stderr. Useful for spot
# checks and for scrubbing a file by hand without writing throwaway code.
data = sys.stdin.read()
out, found = scrub(data)
sys.stdout.write(out)
if found:
by_rule: dict[str, int] = {}
for f in found:
by_rule[f.rule] = by_rule.get(f.rule, 0) + 1
print(f"[redact] {len(found)} redaction(s): "
+ ", ".join(f"{k}={v}" for k, v in sorted(by_rule.items())), file=sys.stderr)
if "--suspicions" in sys.argv:
for s in suspicions(out):
print(f"[suspicion] len={s.length} fp={s.fingerprint}", file=sys.stderr)
+151
View File
@@ -0,0 +1,151 @@
# Secret hygiene: keeping credentials out of the palace
A palace is mined from **transcripts**, and transcripts contain whatever the
terminal printed. Print an environment, `cat` a `.env`, paste a `curl -H
"Authorization: …"`, and the secret becomes a drawer — searchable by every agent
on the fleet, on every machine, indefinitely. This document describes the
scrubbing that exists, what it deliberately does **not** try to do, and the two
layers it belongs in.
Measured on this fleet, 2026-08-27: a single leaked bearer token had reached
**3 drawers, 13 feeder inbox files across all three devices, and 10 local files
dating back 10 days**. Nobody was careless; an agent inspected an environment
variable while debugging. That frequency is the design premise — this is a
routine event to be handled by the pipeline, not an incident to be handled by
discipline.
## 1. Where the hook goes, and why there are two of them
| Layer | Implemented in | Sees | Catches |
|---|---|---|---|
| **Client, pre-staging** | `bin/mempalace_redact.py`, called from `bin/mempalace-pi-session` at the point the staged transcript is written | the machine's own environment **and** the transcript | the real-world case: a secret this machine holds, printed into a session, before it is uploaded anywhere |
| **Server, pre-persist** | *not yet implemented* — see §5 | only the text it is handed | secrets from clients that predate the scrubber, from non-pi clients, and secrets an agent types straight into `add_drawer` |
The client hook is the load-bearing one, because it is the only layer that can
compare text against **the actual secret values it holds** (see tier 1 below) and
because it stops the leak *before transmission*. The server hook is defence in
depth: pattern-only, but it covers write paths the feeder never sees.
Client placement matters for one specific reason: the pi feeder writes the staged
transcript from an in-memory structure, and the remote path then `rsync`s **that
same file** byte-for-byte. Scrubbing at the write point therefore covers local
mining and remote upload with a single hook — there is no second serialization to
forget.
## 2. Detection is anchored on meaning, never on entropy
The tempting design — "redact long random-looking strings" — is actively
destructive here, because a palace is *full* of high-entropy strings that are its
own primary keys:
```
drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b drawer id
evt_20260826T213924_397b1f710c2d event id
rep_d344e349ba276d6fc11997cd552f6937 replica id
1787780364792-000000-rep_d344e349ba276d6f… hybrid logical clock
ecc2a9c574f4156403cf4e90aea8b0d088c75700 git commit sha
```
An entropy detector fires on every one of those, and the resulting redaction is
silent, permanent, and destroys traceability. So detection uses three anchors
that carry meaning instead:
| Tier | Anchor | Default | False-positive risk |
|---|---|---|---|
| **1 — known values** | literal values from this process's env, for variables whose *name* says secret (`…TOKEN`, `…SECRET`, `…PASSWORD`, `…API_KEY`) | **redact** | none by construction: the value *is* the secret |
| **2 — known shapes** | vendor-prefixed credentials: `ghp_…`, `github_pat_…`, `glpat-…`, `xox[abprs]-…`, `sk-…`, `AKIA…`, `AIza…`, `hf_…`, JWTs, PEM private-key blocks, credentials inside URLs, `Authorization:` headers | **redact** | very low: the prefix is meaningful, not random |
| **3 — name=value** | an assignment whose *key* says secret | **report only** | measured **high** — see §3 |
Tier 1 catches any presentation of a secret — env dump, JSON, error message,
URL, prose — because it matches the value itself. It is also the only tier that
resolves this: a 40-hex Gitea PAT is byte-identical in shape to a git commit sha.
## 3. The false-positive measurement, which changed the design
Tier 3 was originally going to redact. Measured against **52 MB of real fleet
transcripts (45 sessions)** it produced **403 hits**, and inspection of them (with
values masked) showed the overwhelming majority were not secrets:
- `GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}` — docker-compose interpolation
- `const tokens = countTokensForModel` — TypeScript source stored as memory
- `refreshToken(credentials: Credentials …)` — a *type annotation*
- `--setattr=krbPasswordExpiration=20260529…` — an IPA attribute holding a date
- `|1.FORGE.TOKENS:…` — AAAK diary shorthand
- `(root@host) Password:` followed by unrelated terminal output
Redacting those would corrupt code, configuration and documentation held as
memory, in order to catch secrets tier 1 already catches by value. After adding
guards for interpolation (`${VAR}`, `$VAR`, `%VAR%`, `{{tpl}}`), code context,
non-secret key suffixes (`…Expiration`, `…Count`, `…Type`) and all-digit values,
the enforced count on the same corpus fell from **403 to 29** — and those 29 are
tier-1 and tier-2 hits, i.e. real credential shapes.
So tier 3 **reports and does not rewrite** by default. Set
`MEMPALACE_REDACT_STRICT=1` to make it enforce, e.g. on a machine whose
transcripts are configuration-heavy rather than code-heavy.
## 4. Known false negatives — stated, not hidden
The scrubber will not catch a novel credential format pasted bare into prose
with no name nearby, a secret belonging to a machine whose environment this
process cannot see, a base64-of-a-secret, or a secret split across lines.
**"The scrubber ran" must never be read as "there are no secrets in here."** To
keep that measurable rather than assumed, two things are reported:
- every scrub prints a count — including `0 redaction(s)`, because printing
nothing is indistinguishable from a scrubber that never ran;
- `suspicions()` reports high-entropy strings it did **not** redact, as
`(length, fingerprint)` pairs rather than values, so the miss rate can be
tracked over time and a recurring fingerprint can be investigated by hand.
Findings never carry the secret. A `Finding` holds the rule, the label, the
length and `sha256(value)[:8]` — enough to recognise the same leak recurring,
not enough to recover it.
**Fail closed.** If the redactor cannot be imported, the feeder refuses to stage
rather than staging unscrubbed (`exit 3`). Override deliberately with
`MEMPALACE_FEED_ALLOW_UNSCRUBBED=1`.
## 5. The server-side layer (not yet implemented)
Three call sites, because the server has three near-duplicate validators rather
than one:
| Path | Function | Covers |
|---|---|---|
| drawers | `config.py` → `sanitize_content()` | `add_drawer`, `update_drawer`, `diary_write`, `checkpoint` — all four route through it |
| events | `logstream.py` → `_sanitize_body()` | `event_append`, and `event_ack` transitively |
| artifacts | `logstream.py` → `put_artifact()` | inlines its own checks; needs its own edit |
Server-side scrubbing is **tier 2 only** (plus optional tier-3 reporting): the
hub cannot see a client's environment, so tier 1 is structurally unavailable
there. That asymmetry is the reason the client hook is not redundant.
## 6. Cleaning up a leak that already landed
1. **Never re-echo the value while hunting it.** Pass it via stdin, never argv
(visible in `ps` on a shared host). Note `get_drawer` is a trap: reading a
drawer in order to redact it prints the secret back into the live transcript.
2. **Redact, don't delete.** Replacing the value in place keeps the mined
transcript's memory value; a stale embedding vector is a cheap price.
3. **Scrub the feeder inbox too**, or the next mine re-files it.
4. **A live session file needs an equal-length in-place overwrite** — the agent
holds it open, so temp-file-plus-rename loses everything appended afterwards.
5. **Expect residue.** SQLite keeps old page content in freed pages until
`VACUUM`, so a raw byte scan still matches after a successful `UPDATE`. Decide
explicitly whether that matters: if the credential store on the same host is
plaintext anyway, it usually does not.
## 7. Usage
```bash
# unit tests, including the must-not-redact corpus of palace id shapes
python3 bin/mempalace_redact.py --self-test
# scrub anything on stdin; report goes to stderr
some-command | python3 bin/mempalace_redact.py > clean.txt
# also list high-entropy strings that were NOT redacted, as fingerprints
python3 bin/mempalace_redact.py --suspicions < transcript.jsonl > clean.jsonl
```