redact: name the tiers where they are used, and stop the docstring lying about tier 3
Review caught that the tier vocabulary was used in the report, the commit message and the docs without being defined anywhere the reader would land, and inspecting that turned up two real defects rather than just a wording gap. 1. STALE DOCSTRING. The module still described tier 3 as if it redacts, which stopped being true when the 403-hit measurement demoted it to report-only. It also credited tier 3 with resolving the 40-hex-PAT-vs-commit-sha collision — false by default, since a reporting rule resolves nothing. Corrected, with the consequence stated plainly: in the default configuration a sha-shaped PAT is caught if and only if it belongs to THIS machine, because only a known value (tier 1) or a naming key (tier 3, reporting) can separate it from a commit sha. That is an accepted gap; the alternative is redacting every sha in the palace. 2. THE VOCABULARY NEVER REACHED THE OUTPUT. The tool prints rule names (github-pat, env-value, url-credentials) and nothing printed a tier, so the docs' tier language was unconnected to what an operator actually sees. Added RULE_TIERS as the authoritative rule -> tier mapping, tier_of(), and Finding.tier; the feeder now prints "T2:github-pat=6" so what matched and how much to trust it are both visible on one line. A self-test asserts every rule that can appear in a Finding maps to a tier, so adding a rule without classifying it fails the tests instead of printing "T?". Tiers, for the record, are three kinds of EVIDENCE (not three severities): T1 known value from this process's env — near-certain, zero FP by construction; T2 known vendor shape — strong, the prefix is meaningful; T3 key name says secret — candidate only, measured FP-heavy, reported. T0 is reserved for suspicions(), which is a measured NON-detection. Docs gain worked one-line examples per tier and a "which tier fired?" section showing real output. 46 self-test cases pass.
This commit is contained in:
+36
-2
@@ -50,15 +50,49 @@ An entropy detector fires on every one of those, and the resulting redaction is
|
||||
silent, permanent, and destroys traceability. So detection uses three anchors
|
||||
that carry meaning instead:
|
||||
|
||||
Each tier is a different *kind of evidence* that a string is a credential. The
|
||||
tier is not a severity ranking of the secret — it is how much to trust the
|
||||
detection. Rule names appear in the output; the tier tells you how to read them.
|
||||
|
||||
| Tier | Anchor | Default | False-positive risk |
|
||||
|---|---|---|---|
|
||||
| **1 — known values** | literal values from this process's env, for variables whose *name* says secret (`…TOKEN`, `…SECRET`, `…PASSWORD`, `…API_KEY`) | **redact** | none by construction: the value *is* the secret |
|
||||
| **2 — known shapes** | vendor-prefixed credentials: `ghp_…`, `github_pat_…`, `glpat-…`, `xox[abprs]-…`, `sk-…`, `AKIA…`, `AIza…`, `hf_…`, JWTs, PEM private-key blocks, credentials inside URLs, `Authorization:` headers | **redact** | very low: the prefix is meaningful, not random |
|
||||
| **3 — name=value** | an assignment whose *key* says secret | **report only** | measured **high** — see §3 |
|
||||
|
||||
### Worked examples, one line each
|
||||
|
||||
```
|
||||
MEMPALACE_REMOTE_TOKEN=pfrDBfak… tier 1 — value matches this env's secret
|
||||
the token is pfrDBfak… apparently tier 1 — same value, bare in prose, still caught
|
||||
git clone https://joakim:hunter2@git/x tier 2 — credentials in a URL
|
||||
Authorization: Bearer abcdefghijklmnop tier 2 — header shape
|
||||
ghp_AbCdEf… / glpat-… / AKIA… / sk-ant-… tier 2 — vendor prefix
|
||||
db_password = s3cr3t-p4ssw0rd-xyz tier 3 — only the KEY suggests it (reported)
|
||||
GITEA_TOKEN=0123456789abcdef… (40 hex) tier 3 — indistinguishable from a commit sha
|
||||
```
|
||||
|
||||
### Which tier fired? Read it off the output
|
||||
|
||||
The tool prints **rule names**, not tier numbers, because the rule says *what*
|
||||
matched. The feeder prefixes them with the tier so both are visible:
|
||||
|
||||
```
|
||||
[REDACTED] 2026-06-27T23-13-49.jsonl T1:env-value=1, T2:github-fine-grained=1 fp=ad78c7d4,fcd95ab5
|
||||
[scrub] 29 redaction(s) applied, 141 name-anchored candidate(s) reported only
|
||||
```
|
||||
|
||||
`RULE_TIERS` in `mempalace_redact.py` is the authoritative mapping, `tier_of()`
|
||||
reads it, and a self-test fails if any rule is left unclassified — so the two
|
||||
vocabularies cannot drift apart silently.
|
||||
|
||||
Tier 1 catches any presentation of a secret — env dump, JSON, error message,
|
||||
URL, prose — because it matches the value itself. It is also the only tier that
|
||||
resolves this: a 40-hex Gitea PAT is byte-identical in shape to a git commit sha.
|
||||
URL, prose — because it matches the value itself. It is also, **in the default
|
||||
configuration, the only tier that resolves a shape collision**: a 40-hex Gitea
|
||||
PAT is byte-identical to a git commit sha, so only a known value (tier 1,
|
||||
redacts) or a naming key (tier 3, reports) can tell them apart. A sha-shaped PAT
|
||||
is therefore caught if and only if it belongs to this machine — an accepted gap,
|
||||
since the alternative is redacting every commit sha in the palace.
|
||||
|
||||
## 3. The false-positive measurement, which changed the design
|
||||
|
||||
|
||||
Reference in New Issue
Block a user