836e35b320
Review caught that the tier vocabulary was used in the report, the commit message and the docs without being defined anywhere the reader would land, and inspecting that turned up two real defects rather than just a wording gap. 1. STALE DOCSTRING. The module still described tier 3 as if it redacts, which stopped being true when the 403-hit measurement demoted it to report-only. It also credited tier 3 with resolving the 40-hex-PAT-vs-commit-sha collision — false by default, since a reporting rule resolves nothing. Corrected, with the consequence stated plainly: in the default configuration a sha-shaped PAT is caught if and only if it belongs to THIS machine, because only a known value (tier 1) or a naming key (tier 3, reporting) can separate it from a commit sha. That is an accepted gap; the alternative is redacting every sha in the palace. 2. THE VOCABULARY NEVER REACHED THE OUTPUT. The tool prints rule names (github-pat, env-value, url-credentials) and nothing printed a tier, so the docs' tier language was unconnected to what an operator actually sees. Added RULE_TIERS as the authoritative rule -> tier mapping, tier_of(), and Finding.tier; the feeder now prints "T2:github-pat=6" so what matched and how much to trust it are both visible on one line. A self-test asserts every rule that can appear in a Finding maps to a tier, so adding a rule without classifying it fails the tests instead of printing "T?". Tiers, for the record, are three kinds of EVIDENCE (not three severities): T1 known value from this process's env — near-certain, zero FP by construction; T2 known vendor shape — strong, the prefix is meaningful; T3 key name says secret — candidate only, measured FP-heavy, reported. T0 is reserved for suspicions(), which is a measured NON-detection. Docs gain worked one-line examples per tier and a "which tier fired?" section showing real output. 46 self-test cases pass.
605 lines
30 KiB
Python
605 lines
30 KiB
Python
#!/usr/bin/env python3
|
|
"""Scrub secret-shaped values out of palace-bound text.
|
|
|
|
WHY THIS IS NAME-ANCHORED AND NOT ENTROPY-ANCHORED
|
|
==================================================
|
|
The obvious design — "redact long random-looking strings" — is actively wrong
|
|
for this corpus, and the reason is worth stating before anyone tries to
|
|
"improve" it. A palace is *full* of high-entropy strings that are its own
|
|
primary keys:
|
|
|
|
drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b drawer id
|
|
..._chunk_000000 chunk id
|
|
evt_20260826T213924_397b1f710c2d event id
|
|
rep_d344e349ba276d6fc11997cd552f6937 replica id
|
|
1787780364792-000000-rep_d344e349ba276d6f… hybrid logical clock
|
|
ecc2a9c574f4156403cf4e90aea8b0d088c75700 git commit sha
|
|
|
|
An entropy or "looks like base64/hex of length N" detector fires on every one
|
|
of those, i.e. on most identifiers in the corpus, and the redaction damages the
|
|
memory it was meant to protect. Worse, the damage is silent and unrecoverable:
|
|
a mined transcript whose ids have been replaced with <redacted> is no longer
|
|
traceable to anything.
|
|
|
|
So detection here is anchored on one of three things that carry *meaning*:
|
|
|
|
Tier 1 KNOWN VALUES — literal values taken from this process's own
|
|
environment, for variables whose NAME says secret.
|
|
Zero false positives by construction: the value
|
|
*is* the secret. Catches any presentation (env
|
|
dump, JSON, error message, URL, prose), which is
|
|
exactly the incident this exists for.
|
|
Tier 2 KNOWN SHAPES — vendor-prefixed credentials (ghp_…, glpat-…,
|
|
xox[abprs]-…, AKIA…, sk-ant-…, JWTs, PEM private
|
|
keys, credentials embedded in URLs, Authorization
|
|
headers). Very low false-positive rate because the
|
|
prefix is meaningful, not merely random.
|
|
Tier 3 NAME=VALUE — an assignment whose *key* says secret
|
|
(…TOKEN, …SECRET, …PASSWORD, …API_KEY, …).
|
|
REPORT-ONLY BY DEFAULT. Measured on 52 MB of real
|
|
fleet transcripts it fired 403 times, of which the
|
|
large majority were ${VAR} interpolation in compose
|
|
files, TypeScript identifiers, a type annotation
|
|
(`credentials: Credentials`), an IPA attribute
|
|
holding a date (krbPasswordExpiration) and terminal
|
|
output following a "Password:" prompt. Rewriting
|
|
those corrupts code and docs held as memory, to
|
|
catch what Tier 1 already catches by value.
|
|
Set MEMPALACE_REDACT_STRICT=1 to make it enforce.
|
|
|
|
SHAPE COLLISIONS, AND WHY TIER 1 IS THE LOAD-BEARING ONE
|
|
--------------------------------------------------------
|
|
A 40-hex Gitea PAT is byte-indistinguishable from a git commit sha, so no shape
|
|
rule can separate them. Only two things can: the value being known (Tier 1,
|
|
which redacts) or a key naming it (Tier 3, which by default only reports). So
|
|
in the default configuration a sha-shaped PAT is caught if and only if it
|
|
belongs to this machine. That is an accepted, documented gap — the alternative
|
|
is redacting every commit sha in the palace.
|
|
|
|
WHICH TIER DID THAT? — mapping printed rule names back to tiers
|
|
--------------------------------------------------------------
|
|
The output names the *rule* (`github-pat`, `env-value`), because that says what
|
|
matched. RULE_TIERS below maps rule -> tier, and `tier_of()` is what the feeder
|
|
uses to print `T2:github-pat=6`, because the tier is what tells a reader how
|
|
much to trust the hit. A self-test asserts every rule that can appear in a
|
|
Finding has a tier, so adding a rule without classifying it fails the tests.
|
|
|
|
KNOWN FALSE NEGATIVES — stated, not hidden
|
|
------------------------------------------
|
|
This will not catch: a novel credential format pasted bare into prose with no
|
|
name nearby; a secret from a machine whose env this process cannot see; a
|
|
base64-of-a-secret; a secret split across lines. Tier 1 covers the local
|
|
machine's own secrets completely, which is where the measured incidents came
|
|
from — but "the scrubber ran" must never be read as "there are no secrets in
|
|
here". `suspicions()` exists for exactly that reason: it reports high-entropy
|
|
candidates it did NOT redact, as fingerprints rather than values, so the false
|
|
negative rate can be *measured* over time instead of assumed to be zero.
|
|
|
|
Run `python3 mempalace_redact.py --self-test` for the test corpus, which
|
|
includes every benign id shape above as a must-not-redact case.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import hashlib
|
|
import math
|
|
import os
|
|
import re
|
|
from typing import Any, Callable, Iterable
|
|
|
|
__all__ = ["scrub", "scrub_obj", "suspicions", "Finding"]
|
|
|
|
PLACEHOLDER = "<redacted:{label}>"
|
|
|
|
# ── Tier 1: names whose VALUES are secrets ────────────────────────────────────
|
|
# Anchored on the variable name, then applied as an exact literal match.
|
|
_SECRET_NAME = re.compile(
|
|
r"(TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|API_?KEY|APIKEY|ACCESS_?KEY"
|
|
r"|PRIVATE_?KEY|CREDENTIAL|BEARER|SESSION_?KEY|COOKIE|_PAT|^PAT$)",
|
|
re.IGNORECASE,
|
|
)
|
|
# …unless the name says it holds a location or a knob rather than a secret.
|
|
# SSH_KEY_PATH is the live example: it matches KEY, and its value is a path.
|
|
_NAME_NOT_SECRET = re.compile(
|
|
r"(_PATH|_FILE|_DIR|_NAME|_ID|_URL|_URI|_HOST|_PORT|_USER|_ENABLED"
|
|
r"|_TIMEOUT|_MS|_SECS|_SECONDS|_COUNT|_LIMIT|_MODE)$",
|
|
re.IGNORECASE,
|
|
)
|
|
# Keys that contain a secret word but describe a POLICY, COUNT or TYPE rather
|
|
# than holding a credential. Measured on real transcripts: krbPasswordExpiration
|
|
# (an IPA attribute whose value is a date), observationTokens (a token count),
|
|
# targetTokens, tokensOverTarget.
|
|
_KEY_NOT_A_SECRET = re.compile(
|
|
r"(expiration|expiry|_age|count|length|limit|usage|interval|policy|type|class"
|
|
r"|provider|manager|error|target|overtarget|file|path|dir|name|_id|url|uri"
|
|
r"|host|port|user|enabled|timeout|_ms|_secs)$",
|
|
re.IGNORECASE,
|
|
)
|
|
# Left of the match: a declaration keyword or a member access means this is CODE
|
|
# (`const token = x`, `obj.apiKey = y`), and the "value" is an expression.
|
|
_CODE_CONTEXT = re.compile(
|
|
r"(?:\b(?:const|let|var|def|func|fn|public|private|readonly|export|return|if|assert)\s+$)"
|
|
r"|[.\w]\s*$"
|
|
)
|
|
_TRIVIAL_VALUES = {
|
|
"", "0", "1", "true", "false", "none", "null", "nil", "unset", "changeme",
|
|
"change-me", "xxx", "***", "redacted", "your-token-here", "your_token_here",
|
|
"placeholder", "example", "dummy", "test", "secret",
|
|
}
|
|
|
|
# ── Tier 2: shapes that mean credential because the PREFIX means credential ──
|
|
_SHAPE_RULES: list[tuple[str, re.Pattern[str], Callable[[re.Match[str]], str]]] = [
|
|
# ORDER MATTERS, and the self-test is what proved it. The two structural
|
|
# rules (credentials inside a URL, Authorization header) run FIRST, before
|
|
# the vendor-prefix rules: a vendor rule rewriting `https://ghp_xxx@host` to
|
|
# `https://<redacted:github-token>@host` leaves a colon INSIDE the
|
|
# placeholder, which the URL rule then reads as user:password and redacts a
|
|
# second time, producing a nested placeholder. Ordering plus the (?!<redacted)
|
|
# guards below make the result independent of pass order.
|
|
("url-credentials",
|
|
re.compile(r"(?P<pre>[a-zA-Z][a-zA-Z0-9+.\-]*://(?!<redacted)[^/\s:@]+:)(?P<pw>(?!<redacted)[^/\s@]{4,})(?P<at>@)"),
|
|
lambda m: f"{m.group('pre')}{PLACEHOLDER.format(label='url-password')}{m.group('at')}"),
|
|
("authorization-header",
|
|
re.compile(r"(?P<pre>[Aa]uthorization:\s*(?:Bearer|Basic|token)\s+)(?P<val>(?!<redacted)[A-Za-z0-9._\-+/=]{8,})"),
|
|
lambda m: f"{m.group('pre')}{PLACEHOLDER.format(label='auth-header')}"),
|
|
("github-pat", re.compile(r"\bgh[pousr]_[A-Za-z0-9]{16,}\b"),
|
|
lambda m: PLACEHOLDER.format(label="github-token")),
|
|
("github-fine-grained", re.compile(r"\bgithub_pat_[A-Za-z0-9_]{22,}\b"),
|
|
lambda m: PLACEHOLDER.format(label="github-token")),
|
|
("gitlab-pat", re.compile(r"\bglpat-[A-Za-z0-9_\-]{16,}\b"),
|
|
lambda m: PLACEHOLDER.format(label="gitlab-token")),
|
|
("slack", re.compile(r"\bxox[abprs]-[A-Za-z0-9\-]{10,}\b"),
|
|
lambda m: PLACEHOLDER.format(label="slack-token")),
|
|
("openai-anthropic", re.compile(r"\bsk-(?:ant-)?[A-Za-z0-9_\-]{20,}\b"),
|
|
lambda m: PLACEHOLDER.format(label="api-key")),
|
|
("aws-access-key", re.compile(r"\b(?:AKIA|ASIA)[0-9A-Z]{16}\b"),
|
|
lambda m: PLACEHOLDER.format(label="aws-access-key-id")),
|
|
("google-api-key", re.compile(r"\bAIza[0-9A-Za-z_\-]{35}\b"),
|
|
lambda m: PLACEHOLDER.format(label="google-api-key")),
|
|
("huggingface", re.compile(r"\bhf_[A-Za-z0-9]{30,}\b"),
|
|
lambda m: PLACEHOLDER.format(label="hf-token")),
|
|
("npm", re.compile(r"\bnpm_[A-Za-z0-9]{30,}\b"),
|
|
lambda m: PLACEHOLDER.format(label="npm-token")),
|
|
("digitalocean", re.compile(r"\bdop_v1_[a-f0-9]{60,}\b"),
|
|
lambda m: PLACEHOLDER.format(label="do-token")),
|
|
# JWT: three base64url segments. Requires the "eyJ" header so a random
|
|
# dotted string cannot trip it.
|
|
("jwt", re.compile(r"\beyJ[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}\b"),
|
|
lambda m: PLACEHOLDER.format(label="jwt")),
|
|
# Whole PEM block, non-greedy so several in one file stay separate.
|
|
("pem-private-key",
|
|
re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----.*?-----END [A-Z ]*PRIVATE KEY-----", re.DOTALL),
|
|
lambda m: PLACEHOLDER.format(label="private-key-block")),
|
|
]
|
|
|
|
# ── Tier 3: assignment whose KEY says secret ─────────────────────────────────
|
|
# The key prefix is LAZY AND OPTIONAL, not "one char then anything": a mandatory
|
|
# leading character cannot match a key that BEGINS with the secret word, so
|
|
# `"api_key"` was missed while `"client_secret"` passed — the secret word sits
|
|
# mid-name in one and at position 0 in the other. Note plain `KEY` is absent from
|
|
# the alternation on purpose: it is too generic (SSH_KEY_PATH), and tier 1 covers
|
|
# the real ones by value anyway.
|
|
# The optional quote after the key is what makes JSON work: `"api_key": "v"`
|
|
# closes the key BEFORE the colon, so a pattern jumping straight from key to
|
|
# separator silently misses every JSON payload — i.e. most of a transcript.
|
|
_ASSIGNMENT = re.compile(
|
|
r"""(?P<key>\b[A-Za-z0-9_.\-]*?
|
|
(?:token|secret|password|passwd|passphrase|api[_-]?key|apikey
|
|
|access[_-]?key|private[_-]?key|credential)
|
|
[A-Za-z0-9_.\-]*)
|
|
(?P<kq>["']?)
|
|
(?P<sep>\s*(?:=>|[=:])\s*)
|
|
(?P<q>["']?)
|
|
(?P<val>[^\s"'<>,;)\]}]{12,})
|
|
(?P=q)""",
|
|
re.IGNORECASE | re.VERBOSE,
|
|
)
|
|
# A value that is obviously not a live secret: a reference, a path, a
|
|
# placeholder, or something we already redacted.
|
|
_NOT_A_SECRET_VALUE = re.compile(
|
|
r"""^(?:
|
|
\$\{[^}\s]*\}? # ${VAR} ${VAR:-default} ${VAR:?err}
|
|
# trailing } optional: the value
|
|
# charset stops before it, so the
|
|
# captured text is "${VAR" — which
|
|
# still means "a reference", and
|
|
# missing that is what made a
|
|
# compose file look like a secret
|
|
| \$[A-Za-z_][A-Za-z0-9_]* # $VAR
|
|
| %[A-Za-z_][A-Za-z0-9_]*%? # %VAR%
|
|
| \{\{[^}\s]*\}?\}? # {{ template }}
|
|
| <[^>]*> # <redacted:…> / <your-token>
|
|
| (?:/|~/|\./|\.\./)[^\s]* # a path
|
|
| [A-Za-z]:\\ # a windows path
|
|
| \*+ | x+ | \.+ # ***, xxx
|
|
| (?:true|false|none|null|nil|unset|changeme|placeholder|example)
|
|
)$""",
|
|
re.IGNORECASE | re.VERBOSE,
|
|
)
|
|
|
|
|
|
class Finding:
|
|
"""One redaction, recorded without ever carrying the secret itself."""
|
|
|
|
__slots__ = ("rule", "label", "length", "fingerprint")
|
|
|
|
def __init__(self, rule: str, label: str, value: str) -> None:
|
|
self.rule = rule
|
|
self.label = label
|
|
self.length = len(value)
|
|
# Stable across sessions, so a recurring leak is recognisable as the
|
|
# same one, while the value stays unrecoverable from the report.
|
|
self.fingerprint = hashlib.sha256(value.encode("utf-8")).hexdigest()[:8]
|
|
|
|
@property
|
|
def tier(self) -> int:
|
|
"""Which detection tier produced this. See RULE_TIERS for the mapping."""
|
|
return tier_of(self.rule)
|
|
|
|
def __repr__(self) -> str: # pragma: no cover - diagnostics only
|
|
return (f"<Finding T{self.tier}:{self.rule}:{self.label} "
|
|
f"len={self.length} fp={self.fingerprint}>")
|
|
|
|
|
|
def env_secrets(environ: dict[str, str] | None = None) -> list[tuple[str, str]]:
|
|
"""Tier 1 material: (name, value) pairs from the environment worth scrubbing.
|
|
|
|
Deliberately conservative: a secret-sounding NAME is not enough, the value
|
|
must also be long enough to be a credential and must not look like a path
|
|
or a knob. Longest values first so that if one secret contains another as a
|
|
substring, the longer one is replaced before the shorter can fragment it.
|
|
"""
|
|
env = os.environ if environ is None else environ
|
|
out: list[tuple[str, str]] = []
|
|
for name, value in env.items():
|
|
if not value or len(value) < 12:
|
|
continue
|
|
if not _SECRET_NAME.search(name) or _NAME_NOT_SECRET.search(name):
|
|
continue
|
|
if value.strip().lower() in _TRIVIAL_VALUES:
|
|
continue
|
|
if _NOT_A_SECRET_VALUE.match(value.strip()):
|
|
continue
|
|
if "/" in value and (value.startswith("/") or value.startswith("~")):
|
|
continue # a path that happened to be named …_KEY
|
|
if any(c.isspace() for c in value):
|
|
continue # credentials do not contain whitespace; prose does
|
|
out.append((name, value))
|
|
out.sort(key=lambda kv: len(kv[1]), reverse=True)
|
|
return out
|
|
|
|
|
|
# Rule name -> tier. Printed output names the rule (WHAT matched); the tier says
|
|
# HOW MUCH TO TRUST IT: 1 = near-certain (the value is known to be a secret),
|
|
# 2 = strong (a vendor prefix means what it says), 3 = candidate, needs eyes.
|
|
RULE_TIERS: dict[str, int] = {
|
|
"env-value": 1,
|
|
**{name: 2 for name, _pat, _lbl in _SHAPE_RULES},
|
|
"named-assignment": 3,
|
|
"named-assignment-reported": 3,
|
|
# Tier 0 is not a detection: it is a measured NON-detection, reported so the
|
|
# false-negative rate is a number instead of an assumption.
|
|
"suspicion": 0,
|
|
}
|
|
|
|
|
|
def tier_of(rule: str) -> int:
|
|
"""Tier for a rule name, or -1 if the rule was never classified."""
|
|
return RULE_TIERS.get(rule, -1)
|
|
|
|
|
|
def scrub(
|
|
text: str,
|
|
known: Iterable[tuple[str, str]] | None = None,
|
|
findings: list[Finding] | None = None,
|
|
strict: bool | None = None,
|
|
) -> tuple[str, list[Finding]]:
|
|
"""Redact secrets from one string. Returns (scrubbed, findings).
|
|
|
|
Idempotent: the replacement text matches none of the rules, so running this
|
|
twice changes nothing and cannot nest placeholders.
|
|
"""
|
|
found: list[Finding] = findings if findings is not None else []
|
|
if strict is None:
|
|
strict = os.environ.get("MEMPALACE_REDACT_STRICT", "").strip() in {"1", "true", "yes"}
|
|
if not text:
|
|
return text, found
|
|
|
|
# Tier 1 first: an exact known value should be labelled with its variable
|
|
# name (the most useful thing a reader can be told) rather than by shape.
|
|
for name, value in (env_secrets() if known is None else known):
|
|
if value and value in text:
|
|
count = text.count(value)
|
|
text = text.replace(value, PLACEHOLDER.format(label=name))
|
|
for _ in range(count):
|
|
found.append(Finding("env-value", name, value))
|
|
|
|
# Tier 2.
|
|
for rule, pattern, repl in _SHAPE_RULES:
|
|
def _sub(m: re.Match[str], _rule: str = rule) -> str:
|
|
found.append(Finding(_rule, _rule, m.group(0)))
|
|
return repl(m)
|
|
text = pattern.sub(_sub, text)
|
|
|
|
# Tier 3 — REPORT-ONLY unless strict. See STRICT note in the module docstring:
|
|
# measured on 52 MB of real fleet transcripts this rule produced 403 hits of
|
|
# which the overwhelming majority were ${VAR} interpolation, TypeScript
|
|
# identifiers, YAML env lists and documentation prose. Redacting those would
|
|
# corrupt code and docs stored as memory, to catch secrets that tier 1
|
|
# already catches by value. So by default a tier-3 hit is REPORTED (so the
|
|
# miss can be measured and the rule calibrated) and NOT rewritten.
|
|
def _sub_assignment(m: re.Match[str]) -> str:
|
|
val, key = m.group("val"), m.group("key")
|
|
pre = m.string[max(0, m.start() - 24):m.start()]
|
|
if (_NOT_A_SECRET_VALUE.match(val)
|
|
or val.strip().lower() in _TRIVIAL_VALUES
|
|
or _KEY_NOT_A_SECRET.search(key)
|
|
or _CODE_CONTEXT.search(pre)
|
|
or val.isdigit()):
|
|
return m.group(0)
|
|
found.append(Finding("named-assignment" if strict else "named-assignment-reported",
|
|
key, val))
|
|
if not strict:
|
|
return m.group(0)
|
|
return f"{m.group('key')}{m.group('kq')}{m.group('sep')}{m.group('q')}" \
|
|
f"{PLACEHOLDER.format(label=m.group('key'))}{m.group('q')}"
|
|
|
|
text = _ASSIGNMENT.sub(_sub_assignment, text)
|
|
return text, found
|
|
|
|
|
|
def scrub_obj(obj: Any, known: Iterable[tuple[str, str]] | None = None,
|
|
findings: list[Finding] | None = None,
|
|
strict: bool | None = None) -> tuple[Any, list[Finding]]:
|
|
"""Recursively scrub every string VALUE in a JSON-ish structure.
|
|
|
|
Keys are left alone: a key named "token" is metadata, not a credential, and
|
|
rewriting keys would corrupt the transcript schema.
|
|
"""
|
|
found: list[Finding] = findings if findings is not None else []
|
|
known = list(env_secrets() if known is None else known)
|
|
if isinstance(obj, str):
|
|
return scrub(obj, known, found, strict)[0], found
|
|
if isinstance(obj, dict):
|
|
return {k: scrub_obj(v, known, found, strict)[0] for k, v in obj.items()}, found
|
|
if isinstance(obj, list):
|
|
return [scrub_obj(v, known, found, strict)[0] for v in obj], found
|
|
return obj, found
|
|
|
|
|
|
# ── Measuring what we MISS, without leaking it ───────────────────────────────
|
|
_BENIGN_ID = re.compile(
|
|
r"""^(?:
|
|
drawer_[\w.\-]+ # drawer / chunk ids
|
|
| (?:evt|art)_\d{8}T\d{6}_[0-9a-f]{12} # event / artifact ids
|
|
| rep_(?:[0-9a-f]{12}|[0-9a-f]{32}) # replica ids
|
|
| \d{13}-[0-9a-f]{6}-rep_[0-9a-f]+ # hybrid logical clocks
|
|
| [0-9a-f]{7,12} # short git sha
|
|
| [0-9a-f]{40} # git sha / sha1
|
|
| [0-9a-f]{64} # sha256
|
|
| [0-9a-f]{8}(?:-[0-9a-f]{4}){3}-[0-9a-f]{12} # uuid
|
|
| sha256:[0-9a-f]{64}
|
|
| \d+
|
|
)$""",
|
|
re.VERBOSE | re.IGNORECASE,
|
|
)
|
|
_CANDIDATE = re.compile(r"[A-Za-z0-9+/_\-=]{24,}")
|
|
|
|
|
|
def _entropy(s: str) -> float:
|
|
if not s:
|
|
return 0.0
|
|
counts: dict[str, int] = {}
|
|
for ch in s:
|
|
counts[ch] = counts.get(ch, 0) + 1
|
|
n = len(s)
|
|
return -sum((c / n) * math.log2(c / n) for c in counts.values())
|
|
|
|
|
|
def suspicions(text: str, min_entropy: float = 3.6) -> list[Finding]:
|
|
"""High-entropy strings this module did NOT redact, as fingerprints.
|
|
|
|
This is the false-negative meter. It deliberately does not redact anything:
|
|
every one of these is far more likely to be one of the palace's own ids than
|
|
a credential, and redacting them would be the damage described at the top of
|
|
this file. Reporting them as (length, fingerprint) makes the miss rate
|
|
observable — and a fingerprint that keeps recurring is worth a human look.
|
|
"""
|
|
out: list[Finding] = []
|
|
for m in _CANDIDATE.finditer(text):
|
|
val = m.group(0)
|
|
if _BENIGN_ID.match(val) or val.startswith("<redacted:"):
|
|
continue
|
|
if _entropy(val) < min_entropy:
|
|
continue
|
|
out.append(Finding("suspicion", "high-entropy", val))
|
|
return out
|
|
|
|
|
|
# ── Self-test ────────────────────────────────────────────────────────────────
|
|
def _self_test() -> int:
|
|
fake_token = "pFrDBfakVKQB0SLN1hTzAoaRJeQ4J_go7xmJezojAwI" # shape-alike, not real
|
|
env = {
|
|
"MEMPALACE_REMOTE_TOKEN": fake_token,
|
|
"GITEA_TOKEN": "0123456789abcdef0123456789abcdef01234567", # 40 hex: sha-shaped
|
|
"SSH_KEY_PATH": "/Users/someone/.ssh", # name matches KEY, value is a path
|
|
"MEMPALACE_MAILBOX_POLL_MS": "300000", # knob, not a secret
|
|
"MEMPALACE_REMOTE_URL": "https://palace.example.com/mcp",
|
|
"PASSWORD": "changeme", # trivial value
|
|
"API_KEY": "${FROM_VAULT}", # a reference
|
|
}
|
|
known = env_secrets(env)
|
|
|
|
must_redact = [
|
|
("env dump line", f"MEMPALACE_REMOTE_TOKEN={fake_token}"),
|
|
("bare value in prose", f"the token is {fake_token} apparently"),
|
|
("value inside json", f'{{"env": {{"MEMPALACE_REMOTE_TOKEN": "{fake_token}"}}}}'),
|
|
("sha-shaped pat, name-anchored",
|
|
"GITEA_TOKEN=0123456789abcdef0123456789abcdef01234567"),
|
|
("github pat", "remote add origin https://ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345@git"),
|
|
("gitlab pat", "GLPAT is glpat-AbCdEfGhIjKlMnOpQrSt"),
|
|
("slack", "xoxb-1234567890-ABCdefGHIjklMNO"),
|
|
("anthropic-ish", "sk-ant-api03-AbCdEfGhIjKlMnOpQrStUvWx"),
|
|
("aws", "AKIAIOSFODNN7EXAMPLE"),
|
|
("jwt", "eyJhbGciOiJIUzI1NiJ9.eyJzdWIiOiIxMjM0NTY3ODkwIn0.dozjgNryP4J3jVmNHl0w5N_XgL0n3I9PlFUP0THsR8U"),
|
|
("pem", "-----BEGIN OPENSSH PRIVATE KEY-----\nb3BlbnNzaA\n-----END OPENSSH PRIVATE KEY-----"),
|
|
("url creds", "git clone https://joakim:hunter2hunter2@git.example.com/x.git"),
|
|
("auth header", "Authorization: Bearer abcdefghijklmnop"),
|
|
("json api_key", '"api_key": "AbCdEfGhIjKlMnOpQrSt"'),
|
|
("json nested quotes", '{"auth": {"client_secret": "AbCdEfGhIjKlMnOpQrStUv"}}'),
|
|
("yaml style", " gitea_token: 0123456789abcdefghij"),
|
|
("lowercase password kv", "db_password = s3cr3t-p4ssw0rd-xyz"),
|
|
]
|
|
must_not_redact = [
|
|
("drawer id", "drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b"),
|
|
("chunk id", "drawer_pi-devbox_landmines_ca8929fea82cf80e7770ea0d_chunk_000000"),
|
|
("event id", "evt_20260826T213924_397b1f710c2d"),
|
|
("replica id", "rep_d344e349ba276d6fc11997cd552f6937"),
|
|
("hlc", "1787780364792-000000-rep_d344e349ba276d6fc11997cd552f6937"),
|
|
("git sha", "ecc2a9c574f4156403cf4e90aea8b0d088c75700"),
|
|
("short sha", "ecc2a9c"),
|
|
("sha256", "a" * 64),
|
|
("uuid session file", "pi_01a03fd8-0dc6-73c2-a9c3-6b62ccf0318f.jsonl"),
|
|
("knob assignment", "MEMPALACE_MAILBOX_POLL_MS=300000"),
|
|
("path named key", "SSH_KEY_PATH=/Users/someone/.ssh"),
|
|
("empty token", "MEMPALACE_REMOTE_TOKEN="),
|
|
("already redacted", "MEMPALACE_REMOTE_TOKEN=<redacted:MEMPALACE_REMOTE_TOKEN>"),
|
|
("var reference", "MEMPALACE_REMOTE_TOKEN=${VAULT_TOKEN}"),
|
|
("placeholder", "MEMPALACE_REMOTE_TOKEN=<your-token-here>"),
|
|
("trivial", "PASSWORD=changeme"),
|
|
("url no creds", "MEMPALACE_REMOTE_URL=https://palace.example.com/mcp"),
|
|
("prose", "The mailbox is an obligation channel, not a news channel."),
|
|
("base64 in a diff line", "+ data = 'QUJDREVGR0hJSktMTU5PUFFSU1RVVldYWVo='"),
|
|
]
|
|
|
|
failures = 0
|
|
# Tier 3 is report-only by default, so the corpus below is checked in STRICT
|
|
# mode (where tier 3 rewrites) and the default mode is asserted separately.
|
|
print("MUST REDACT (strict: tier 3 enforcing)")
|
|
for name, sample in must_redact:
|
|
out, found = scrub(sample, known, strict=True)
|
|
ok = bool(found) and all(
|
|
secret not in out for secret in (fake_token, "hunter2hunter2",
|
|
"0123456789abcdef0123456789abcdef01234567")
|
|
)
|
|
failures += not ok
|
|
rules = ",".join(sorted({f.rule for f in found})) or "NONE"
|
|
print(f" {'pass' if ok else 'FAIL'} {name:32s} [{rules}]")
|
|
print("MUST NOT REDACT (strict: the hard cases)")
|
|
for name, sample in must_not_redact:
|
|
out, found = scrub(sample, known, strict=True)
|
|
ok = out == sample and not found
|
|
failures += not ok
|
|
detail = "" if ok else f" -> {out!r} {found}"
|
|
print(f" {'pass' if ok else 'FAIL'} {name:32s}{detail}")
|
|
|
|
print("DEFAULT MODE (tier 3 reports, does not rewrite)")
|
|
for name, sample, expect_rewrite in [
|
|
("env value still redacted", f"MEMPALACE_REMOTE_TOKEN={fake_token}", True),
|
|
("vendor shape still redacted", "ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345", True),
|
|
("name-anchored only reported", "db_password = s3cr3t-p4ssw0rd-xyz", False),
|
|
("compose interpolation untouched", "GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}", False),
|
|
("typescript identifier untouched", "const tokens = countTokensForModel", False),
|
|
("ipa attribute untouched", "--setattr=krbPasswordExpiration=20260529090000", False),
|
|
]:
|
|
out, found = scrub(sample, known)
|
|
rewritten = out != sample
|
|
ok = rewritten == expect_rewrite
|
|
if name == "name-anchored only reported":
|
|
ok = ok and any(f.rule == "named-assignment-reported" for f in found)
|
|
if name.endswith("untouched"):
|
|
ok = ok and not any(f.rule.startswith("named-assignment") for f in found)
|
|
failures += not ok
|
|
print(f" {'pass' if ok else 'FAIL'} {name:34s}"
|
|
+ ("" if ok else f" -> rewritten={rewritten} {found}"))
|
|
|
|
print("TIER MAPPING")
|
|
_rules = ({"env-value", "named-assignment", "named-assignment-reported", "suspicion"}
|
|
| {n for n, _p, _l in _SHAPE_RULES})
|
|
_unclassified = sorted(r for r in _rules if tier_of(r) < 0)
|
|
failures += bool(_unclassified)
|
|
print(f" {'pass' if not _unclassified else 'FAIL'} every rule maps to a tier"
|
|
+ (f" -> unclassified: {_unclassified}" if _unclassified
|
|
else f" ({len(_rules)} rules)"))
|
|
for label, sample, want in [
|
|
("env value is tier 1", f"MEMPALACE_REMOTE_TOKEN={fake_token}", 1),
|
|
("vendor shape is tier 2", "ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345", 2),
|
|
("name-anchored is tier 3", "db_password = s3cr3t-p4ssw0rd-xyz", 3),
|
|
]:
|
|
got = scrub(sample, known)[1]
|
|
ok = bool(got) and all(f.tier == want for f in got)
|
|
failures += not ok
|
|
print(f" {'pass' if ok else 'FAIL'} {label:26s}"
|
|
+ ("" if ok else f" -> {[(f.rule, f.tier) for f in got]}"))
|
|
|
|
print("PROPERTIES")
|
|
once, _ = scrub(f"MEMPALACE_REMOTE_TOKEN={fake_token}", known, strict=True)
|
|
twice, second = scrub(once, known, strict=True)
|
|
ok = once == twice and not second
|
|
failures += not ok
|
|
print(f" {'pass' if ok else 'FAIL'} idempotent (no nested placeholders)")
|
|
|
|
# The compound case the corpus caught: assert on EXACT output, not merely
|
|
# that the secret is gone — a nested placeholder also hides the secret, and
|
|
# would have shipped unnoticed.
|
|
compound, _ = scrub("https://ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345@git.example.com/x", known, strict=True)
|
|
ok = compound == "https://<redacted:github-token>@git.example.com/x"
|
|
failures += not ok
|
|
print(f" {'pass' if ok else 'FAIL'} vendor placeholder not re-redacted by url rule"
|
|
+ ("" if ok else f" -> {compound!r}"))
|
|
|
|
already, found_a = scrub("https://joakim:<redacted:url-password>@git.example.com", known, strict=True)
|
|
ok = already == "https://joakim:<redacted:url-password>@git.example.com" and not found_a
|
|
failures += not ok
|
|
print(f" {'pass' if ok else 'FAIL'} already-redacted url left alone")
|
|
|
|
obj = {"role": "user", "content": [{"type": "text", "text": f"export TOK={fake_token}"}],
|
|
"id": "evt_20260826T213924_397b1f710c2d"}
|
|
scrubbed, found = scrub_obj(obj, known, strict=True)
|
|
ok = (fake_token not in repr(scrubbed)
|
|
and scrubbed["id"] == obj["id"]
|
|
and scrubbed["role"] == "user"
|
|
and len(found) >= 1)
|
|
failures += not ok
|
|
print(f" {'pass' if ok else 'FAIL'} nested structure scrubbed, ids and keys preserved")
|
|
|
|
ok = all(fake_token not in repr(f) for f in scrub(f"x {fake_token}", known, strict=True)[1])
|
|
failures += not ok
|
|
print(f" {'pass' if ok else 'FAIL'} findings never carry the secret value")
|
|
|
|
sus = suspicions("drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b "
|
|
"rep_d344e349ba276d6fc11997cd552f6937 "
|
|
"1787780364792-000000-rep_d344e349ba276d6fc11997cd552f6937 "
|
|
"ecc2a9c574f4156403cf4e90aea8b0d088c75700")
|
|
ok = not sus
|
|
failures += not ok
|
|
print(f" {'pass' if ok else 'FAIL'} suspicion meter ignores our own id shapes ({len(sus)} hits)")
|
|
|
|
sus2 = suspicions("blob=Zm9vYmFyYmF6cXV1eHF1dXhmb29iYXJiYXpxdXV4Zm9vYmFy")
|
|
ok = len(sus2) == 1 and all("Zm9v" not in repr(s) for s in sus2)
|
|
failures += not ok
|
|
print(f" {'pass' if ok else 'FAIL'} suspicion meter flags unknown blobs as fingerprints only")
|
|
|
|
print(f"\n{'ALL PASS' if not failures else f'{failures} FAILED'}")
|
|
return 1 if failures else 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
import sys
|
|
if "--self-test" in sys.argv:
|
|
raise SystemExit(_self_test())
|
|
# Filter mode: scrub stdin -> stdout, report to stderr. Useful for spot
|
|
# checks and for scrubbing a file by hand without writing throwaway code.
|
|
data = sys.stdin.read()
|
|
out, found = scrub(data)
|
|
sys.stdout.write(out)
|
|
if found:
|
|
by_rule: dict[str, int] = {}
|
|
for f in found:
|
|
by_rule[f.rule] = by_rule.get(f.rule, 0) + 1
|
|
print(f"[redact] {len(found)} redaction(s): "
|
|
+ ", ".join(f"{k}={v}" for k, v in sorted(by_rule.items())), file=sys.stderr)
|
|
if "--suspicions" in sys.argv:
|
|
for s in suspicions(out):
|
|
print(f"[suspicion] len={s.length} fp={s.fingerprint}", file=sys.stderr)
|