redact: name the tiers where they are used, and stop the docstring lying about tier 3
Review caught that the tier vocabulary was used in the report, the commit message and the docs without being defined anywhere the reader would land, and inspecting that turned up two real defects rather than just a wording gap. 1. STALE DOCSTRING. The module still described tier 3 as if it redacts, which stopped being true when the 403-hit measurement demoted it to report-only. It also credited tier 3 with resolving the 40-hex-PAT-vs-commit-sha collision — false by default, since a reporting rule resolves nothing. Corrected, with the consequence stated plainly: in the default configuration a sha-shaped PAT is caught if and only if it belongs to THIS machine, because only a known value (tier 1) or a naming key (tier 3, reporting) can separate it from a commit sha. That is an accepted gap; the alternative is redacting every sha in the palace. 2. THE VOCABULARY NEVER REACHED THE OUTPUT. The tool prints rule names (github-pat, env-value, url-credentials) and nothing printed a tier, so the docs' tier language was unconnected to what an operator actually sees. Added RULE_TIERS as the authoritative rule -> tier mapping, tier_of(), and Finding.tier; the feeder now prints "T2:github-pat=6" so what matched and how much to trust it are both visible on one line. A self-test asserts every rule that can appear in a Finding maps to a tier, so adding a rule without classifying it fails the tests instead of printing "T?". Tiers, for the record, are three kinds of EVIDENCE (not three severities): T1 known value from this process's env — near-certain, zero FP by construction; T2 known vendor shape — strong, the prefix is meaningful; T3 key name says secret — candidate only, measured FP-heavy, reported. T0 is reserved for suspicions(), which is a measured NON-detection. Docs gain worked one-line examples per tier and a "which tier fired?" section showing real output. 46 self-test cases pass.
This commit is contained in:
@@ -856,7 +856,10 @@ for path in paths:
|
|||||||
if _hard:
|
if _hard:
|
||||||
_by = {}
|
_by = {}
|
||||||
for x in _hard:
|
for x in _hard:
|
||||||
_by[x.rule] = _by.get(x.rule, 0) + 1
|
# "T2:github-pat" — rule says what matched, tier says how much to
|
||||||
|
# trust it, which is what the reader of this line actually needs.
|
||||||
|
_k = f"T{x.tier}:{x.rule}"
|
||||||
|
_by[_k] = _by.get(_k, 0) + 1
|
||||||
print(f" [REDACTED] {path.name} "
|
print(f" [REDACTED] {path.name} "
|
||||||
+ ", ".join(f"{k}={v}" for k, v in sorted(_by.items()))
|
+ ", ".join(f"{k}={v}" for k, v in sorted(_by.items()))
|
||||||
+ " fp=" + ",".join(sorted({x.fingerprint for x in _hard})),
|
+ " fp=" + ",".join(sorted({x.fingerprint for x in _hard})),
|
||||||
|
|||||||
+72
-6
@@ -36,11 +36,33 @@ So detection here is anchored on one of three things that carry *meaning*:
|
|||||||
prefix is meaningful, not merely random.
|
prefix is meaningful, not merely random.
|
||||||
Tier 3 NAME=VALUE — an assignment whose *key* says secret
|
Tier 3 NAME=VALUE — an assignment whose *key* says secret
|
||||||
(…TOKEN, …SECRET, …PASSWORD, …API_KEY, …).
|
(…TOKEN, …SECRET, …PASSWORD, …API_KEY, …).
|
||||||
This is the tier that resolves shape collisions
|
REPORT-ONLY BY DEFAULT. Measured on 52 MB of real
|
||||||
that no shape-based rule can: a 40-hex Gitea PAT
|
fleet transcripts it fired 403 times, of which the
|
||||||
is byte-indistinguishable from a git commit sha,
|
large majority were ${VAR} interpolation in compose
|
||||||
and `GITEA_TOKEN=<40 hex>` is the only thing that
|
files, TypeScript identifiers, a type annotation
|
||||||
tells them apart.
|
(`credentials: Credentials`), an IPA attribute
|
||||||
|
holding a date (krbPasswordExpiration) and terminal
|
||||||
|
output following a "Password:" prompt. Rewriting
|
||||||
|
those corrupts code and docs held as memory, to
|
||||||
|
catch what Tier 1 already catches by value.
|
||||||
|
Set MEMPALACE_REDACT_STRICT=1 to make it enforce.
|
||||||
|
|
||||||
|
SHAPE COLLISIONS, AND WHY TIER 1 IS THE LOAD-BEARING ONE
|
||||||
|
--------------------------------------------------------
|
||||||
|
A 40-hex Gitea PAT is byte-indistinguishable from a git commit sha, so no shape
|
||||||
|
rule can separate them. Only two things can: the value being known (Tier 1,
|
||||||
|
which redacts) or a key naming it (Tier 3, which by default only reports). So
|
||||||
|
in the default configuration a sha-shaped PAT is caught if and only if it
|
||||||
|
belongs to this machine. That is an accepted, documented gap — the alternative
|
||||||
|
is redacting every commit sha in the palace.
|
||||||
|
|
||||||
|
WHICH TIER DID THAT? — mapping printed rule names back to tiers
|
||||||
|
--------------------------------------------------------------
|
||||||
|
The output names the *rule* (`github-pat`, `env-value`), because that says what
|
||||||
|
matched. RULE_TIERS below maps rule -> tier, and `tier_of()` is what the feeder
|
||||||
|
uses to print `T2:github-pat=6`, because the tier is what tells a reader how
|
||||||
|
much to trust the hit. A self-test asserts every rule that can appear in a
|
||||||
|
Finding has a tier, so adding a rule without classifying it fails the tests.
|
||||||
|
|
||||||
KNOWN FALSE NEGATIVES — stated, not hidden
|
KNOWN FALSE NEGATIVES — stated, not hidden
|
||||||
------------------------------------------
|
------------------------------------------
|
||||||
@@ -209,8 +231,14 @@ class Finding:
|
|||||||
# same one, while the value stays unrecoverable from the report.
|
# same one, while the value stays unrecoverable from the report.
|
||||||
self.fingerprint = hashlib.sha256(value.encode("utf-8")).hexdigest()[:8]
|
self.fingerprint = hashlib.sha256(value.encode("utf-8")).hexdigest()[:8]
|
||||||
|
|
||||||
|
@property
|
||||||
|
def tier(self) -> int:
|
||||||
|
"""Which detection tier produced this. See RULE_TIERS for the mapping."""
|
||||||
|
return tier_of(self.rule)
|
||||||
|
|
||||||
def __repr__(self) -> str: # pragma: no cover - diagnostics only
|
def __repr__(self) -> str: # pragma: no cover - diagnostics only
|
||||||
return f"<Finding {self.rule}:{self.label} len={self.length} fp={self.fingerprint}>"
|
return (f"<Finding T{self.tier}:{self.rule}:{self.label} "
|
||||||
|
f"len={self.length} fp={self.fingerprint}>")
|
||||||
|
|
||||||
|
|
||||||
def env_secrets(environ: dict[str, str] | None = None) -> list[tuple[str, str]]:
|
def env_secrets(environ: dict[str, str] | None = None) -> list[tuple[str, str]]:
|
||||||
@@ -241,6 +269,25 @@ def env_secrets(environ: dict[str, str] | None = None) -> list[tuple[str, str]]:
|
|||||||
return out
|
return out
|
||||||
|
|
||||||
|
|
||||||
|
# Rule name -> tier. Printed output names the rule (WHAT matched); the tier says
|
||||||
|
# HOW MUCH TO TRUST IT: 1 = near-certain (the value is known to be a secret),
|
||||||
|
# 2 = strong (a vendor prefix means what it says), 3 = candidate, needs eyes.
|
||||||
|
RULE_TIERS: dict[str, int] = {
|
||||||
|
"env-value": 1,
|
||||||
|
**{name: 2 for name, _pat, _lbl in _SHAPE_RULES},
|
||||||
|
"named-assignment": 3,
|
||||||
|
"named-assignment-reported": 3,
|
||||||
|
# Tier 0 is not a detection: it is a measured NON-detection, reported so the
|
||||||
|
# false-negative rate is a number instead of an assumption.
|
||||||
|
"suspicion": 0,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def tier_of(rule: str) -> int:
|
||||||
|
"""Tier for a rule name, or -1 if the rule was never classified."""
|
||||||
|
return RULE_TIERS.get(rule, -1)
|
||||||
|
|
||||||
|
|
||||||
def scrub(
|
def scrub(
|
||||||
text: str,
|
text: str,
|
||||||
known: Iterable[tuple[str, str]] | None = None,
|
known: Iterable[tuple[str, str]] | None = None,
|
||||||
@@ -466,6 +513,25 @@ def _self_test() -> int:
|
|||||||
print(f" {'pass' if ok else 'FAIL'} {name:34s}"
|
print(f" {'pass' if ok else 'FAIL'} {name:34s}"
|
||||||
+ ("" if ok else f" -> rewritten={rewritten} {found}"))
|
+ ("" if ok else f" -> rewritten={rewritten} {found}"))
|
||||||
|
|
||||||
|
print("TIER MAPPING")
|
||||||
|
_rules = ({"env-value", "named-assignment", "named-assignment-reported", "suspicion"}
|
||||||
|
| {n for n, _p, _l in _SHAPE_RULES})
|
||||||
|
_unclassified = sorted(r for r in _rules if tier_of(r) < 0)
|
||||||
|
failures += bool(_unclassified)
|
||||||
|
print(f" {'pass' if not _unclassified else 'FAIL'} every rule maps to a tier"
|
||||||
|
+ (f" -> unclassified: {_unclassified}" if _unclassified
|
||||||
|
else f" ({len(_rules)} rules)"))
|
||||||
|
for label, sample, want in [
|
||||||
|
("env value is tier 1", f"MEMPALACE_REMOTE_TOKEN={fake_token}", 1),
|
||||||
|
("vendor shape is tier 2", "ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345", 2),
|
||||||
|
("name-anchored is tier 3", "db_password = s3cr3t-p4ssw0rd-xyz", 3),
|
||||||
|
]:
|
||||||
|
got = scrub(sample, known)[1]
|
||||||
|
ok = bool(got) and all(f.tier == want for f in got)
|
||||||
|
failures += not ok
|
||||||
|
print(f" {'pass' if ok else 'FAIL'} {label:26s}"
|
||||||
|
+ ("" if ok else f" -> {[(f.rule, f.tier) for f in got]}"))
|
||||||
|
|
||||||
print("PROPERTIES")
|
print("PROPERTIES")
|
||||||
once, _ = scrub(f"MEMPALACE_REMOTE_TOKEN={fake_token}", known, strict=True)
|
once, _ = scrub(f"MEMPALACE_REMOTE_TOKEN={fake_token}", known, strict=True)
|
||||||
twice, second = scrub(once, known, strict=True)
|
twice, second = scrub(once, known, strict=True)
|
||||||
|
|||||||
+36
-2
@@ -50,15 +50,49 @@ An entropy detector fires on every one of those, and the resulting redaction is
|
|||||||
silent, permanent, and destroys traceability. So detection uses three anchors
|
silent, permanent, and destroys traceability. So detection uses three anchors
|
||||||
that carry meaning instead:
|
that carry meaning instead:
|
||||||
|
|
||||||
|
Each tier is a different *kind of evidence* that a string is a credential. The
|
||||||
|
tier is not a severity ranking of the secret — it is how much to trust the
|
||||||
|
detection. Rule names appear in the output; the tier tells you how to read them.
|
||||||
|
|
||||||
| Tier | Anchor | Default | False-positive risk |
|
| Tier | Anchor | Default | False-positive risk |
|
||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| **1 — known values** | literal values from this process's env, for variables whose *name* says secret (`…TOKEN`, `…SECRET`, `…PASSWORD`, `…API_KEY`) | **redact** | none by construction: the value *is* the secret |
|
| **1 — known values** | literal values from this process's env, for variables whose *name* says secret (`…TOKEN`, `…SECRET`, `…PASSWORD`, `…API_KEY`) | **redact** | none by construction: the value *is* the secret |
|
||||||
| **2 — known shapes** | vendor-prefixed credentials: `ghp_…`, `github_pat_…`, `glpat-…`, `xox[abprs]-…`, `sk-…`, `AKIA…`, `AIza…`, `hf_…`, JWTs, PEM private-key blocks, credentials inside URLs, `Authorization:` headers | **redact** | very low: the prefix is meaningful, not random |
|
| **2 — known shapes** | vendor-prefixed credentials: `ghp_…`, `github_pat_…`, `glpat-…`, `xox[abprs]-…`, `sk-…`, `AKIA…`, `AIza…`, `hf_…`, JWTs, PEM private-key blocks, credentials inside URLs, `Authorization:` headers | **redact** | very low: the prefix is meaningful, not random |
|
||||||
| **3 — name=value** | an assignment whose *key* says secret | **report only** | measured **high** — see §3 |
|
| **3 — name=value** | an assignment whose *key* says secret | **report only** | measured **high** — see §3 |
|
||||||
|
|
||||||
|
### Worked examples, one line each
|
||||||
|
|
||||||
|
```
|
||||||
|
MEMPALACE_REMOTE_TOKEN=pfrDBfak… tier 1 — value matches this env's secret
|
||||||
|
the token is pfrDBfak… apparently tier 1 — same value, bare in prose, still caught
|
||||||
|
git clone https://joakim:hunter2@git/x tier 2 — credentials in a URL
|
||||||
|
Authorization: Bearer abcdefghijklmnop tier 2 — header shape
|
||||||
|
ghp_AbCdEf… / glpat-… / AKIA… / sk-ant-… tier 2 — vendor prefix
|
||||||
|
db_password = s3cr3t-p4ssw0rd-xyz tier 3 — only the KEY suggests it (reported)
|
||||||
|
GITEA_TOKEN=0123456789abcdef… (40 hex) tier 3 — indistinguishable from a commit sha
|
||||||
|
```
|
||||||
|
|
||||||
|
### Which tier fired? Read it off the output
|
||||||
|
|
||||||
|
The tool prints **rule names**, not tier numbers, because the rule says *what*
|
||||||
|
matched. The feeder prefixes them with the tier so both are visible:
|
||||||
|
|
||||||
|
```
|
||||||
|
[REDACTED] 2026-06-27T23-13-49.jsonl T1:env-value=1, T2:github-fine-grained=1 fp=ad78c7d4,fcd95ab5
|
||||||
|
[scrub] 29 redaction(s) applied, 141 name-anchored candidate(s) reported only
|
||||||
|
```
|
||||||
|
|
||||||
|
`RULE_TIERS` in `mempalace_redact.py` is the authoritative mapping, `tier_of()`
|
||||||
|
reads it, and a self-test fails if any rule is left unclassified — so the two
|
||||||
|
vocabularies cannot drift apart silently.
|
||||||
|
|
||||||
Tier 1 catches any presentation of a secret — env dump, JSON, error message,
|
Tier 1 catches any presentation of a secret — env dump, JSON, error message,
|
||||||
URL, prose — because it matches the value itself. It is also the only tier that
|
URL, prose — because it matches the value itself. It is also, **in the default
|
||||||
resolves this: a 40-hex Gitea PAT is byte-identical in shape to a git commit sha.
|
configuration, the only tier that resolves a shape collision**: a 40-hex Gitea
|
||||||
|
PAT is byte-identical to a git commit sha, so only a known value (tier 1,
|
||||||
|
redacts) or a naming key (tier 3, reports) can tell them apart. A sha-shaped PAT
|
||||||
|
is therefore caught if and only if it belongs to this machine — an accepted gap,
|
||||||
|
since the alternative is redacting every commit sha in the palace.
|
||||||
|
|
||||||
## 3. The false-positive measurement, which changed the design
|
## 3. The false-positive measurement, which changed the design
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user