b2b50afcc115d07646f0e4b6000478ab803cc7fa
10 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f0bffd1b93 |
feed: chase the symlink, or fail-closed takes the whole fleet's feed down
The image installs /usr/local/bin/mempalace-pi-session as a symlink into
/opt/mempalace-toolkit/bin, and ${BASH_SOURCE[0]} reports the path the script was
INVOKED as, not the resolved target. So the sibling-module lookup pointed at
/usr/local/bin, the redactor was not there, and the fail-closed import did exactly
what it was told: refused to stage.
MEASURED, not theorised: /usr/local/bin/mempalace-pi-session --dry-run exited
"[FATAL] secret scrubber unavailable ... refusing to stage" while
/opt/mempalace-toolkit/bin/mempalace-pi-session --dry-run scrubbed 40 findings on
the same input. The symlink is how every device invokes it, so at the next image
bake every feeder tick on every machine would have stopped staging — a silent,
fleet-wide memory outage, which is a worse outcome than the leak the scrubber
exists to prevent. My own tests missed it by calling bin/... directly from the
checkout, i.e. the one invocation path the fleet never uses.
FIX: chase the symlink chain in portable shell and offer fallbacks instead of
betting on a single answer. MEMPALACE_REDACT_DIR is now colon-separated —
resolved dir, invoked dir, then the image's known install path /opt/... — and the
Python side inserts the first existing candidate. readlink -f is deliberately NOT
used: it is GNU/newer-BSD only and this script also runs directly on macOS hosts,
so the chase is a plain while [ -L ] loop handling relative link targets.
Verified on all three invocation shapes: via the /usr/local/bin symlink, via the
direct /opt path, and via a second-hop symlink in an unrelated directory. All
three now report the same 40 redactions.
LESSON worth keeping: fail-closed is correct for a secret scrubber, but it
converts "module not found" into an outage, so the module lookup becomes
load-bearing infrastructure and must be tested through the real invocation path,
not the convenient one.
|
||
|
|
836e35b320 |
redact: name the tiers where they are used, and stop the docstring lying about tier 3
Review caught that the tier vocabulary was used in the report, the commit message and the docs without being defined anywhere the reader would land, and inspecting that turned up two real defects rather than just a wording gap. 1. STALE DOCSTRING. The module still described tier 3 as if it redacts, which stopped being true when the 403-hit measurement demoted it to report-only. It also credited tier 3 with resolving the 40-hex-PAT-vs-commit-sha collision — false by default, since a reporting rule resolves nothing. Corrected, with the consequence stated plainly: in the default configuration a sha-shaped PAT is caught if and only if it belongs to THIS machine, because only a known value (tier 1) or a naming key (tier 3, reporting) can separate it from a commit sha. That is an accepted gap; the alternative is redacting every sha in the palace. 2. THE VOCABULARY NEVER REACHED THE OUTPUT. The tool prints rule names (github-pat, env-value, url-credentials) and nothing printed a tier, so the docs' tier language was unconnected to what an operator actually sees. Added RULE_TIERS as the authoritative rule -> tier mapping, tier_of(), and Finding.tier; the feeder now prints "T2:github-pat=6" so what matched and how much to trust it are both visible on one line. A self-test asserts every rule that can appear in a Finding maps to a tier, so adding a rule without classifying it fails the tests instead of printing "T?". Tiers, for the record, are three kinds of EVIDENCE (not three severities): T1 known value from this process's env — near-certain, zero FP by construction; T2 known vendor shape — strong, the prefix is meaningful; T3 key name says secret — candidate only, measured FP-heavy, reported. T0 is reserved for suspicions(), which is a measured NON-detection. Docs gain worked one-line examples per tier and a "which tier fired?" section showing real output. 46 self-test cases pass. |
||
|
|
3d47937d06 |
feed: scrub secrets before staging a transcript, name-anchored not entropy-anchored
A palace is mined from transcripts, and transcripts contain whatever the terminal
printed. Measured on this fleet: one leaked bearer token had reached 3 drawers,
13 feeder inbox files across all three devices, and 10 local files spanning 10
days — from an agent inspecting an env var while debugging. That frequency is the
premise: this is a pipeline problem, not a discipline problem.
WHERE. bin/mempalace_redact.py, called from mempalace-pi-session at the point the
staged transcript is written. That single hook covers both transports, because
local mode mines the staged file and remote mode rsyncs that same file
byte-for-byte. Scrubbing operates on the parsed objects rather than the
serialized text, so string values are rewritten while keys, ids and structure are
untouched — scanning raw JSONL instead invents keys like "tapiKey" out of the \t
escape preceding a field name (observed, not theorised).
WHY NOT ENTROPY. The obvious "redact long random-looking strings" is actively
destructive here: drawer ids, chunk ids, event ids, replica ids, HLCs and commit
SHAs are all high-entropy, are the majority of random-looking text in a palace,
and redacting them is silent and permanent. Detection is anchored on meaning
instead: tier 1 literal values from this process's env whose NAME says secret
(zero false positives by construction, and the only tier that can tell a 40-hex
Gitea PAT from a git commit sha); tier 2 vendor-prefixed shapes (ghp_, glpat-,
xox*-, sk-, AKIA, AIza, hf_, JWT, PEM blocks, URL credentials, Authorization
headers); tier 3 name=value assignments.
THE MEASUREMENT THAT CHANGED THE DESIGN. Tier 3 was going to redact. Against
52 MB of real fleet transcripts it produced 403 hits, and inspection with values
masked showed most were ${VAR} interpolation in compose files, TypeScript
identifiers, a TYPE ANNOTATION (credentials: Credentials), an IPA attribute
holding a date (krbPasswordExpiration), AAAK diary shorthand, and terminal output
following a "Password:" prompt. Redacting those corrupts code and docs held as
memory to catch what tier 1 already catches by value. After adding guards for
interpolation, code context, non-secret key suffixes and all-digit values, the
enforced count fell 403 -> 29 on the same corpus. So tier 3 REPORTS and does not
rewrite unless MEMPALACE_REDACT_STRICT=1.
HONESTY ABOUT MISSES. Known false negatives are documented rather than papered
over: novel formats in bare prose, another machine's secrets, base64-of-a-secret,
line-split secrets. Every run prints a count including "0 redaction(s)", because
silence is indistinguishable from a scrubber that never ran, and suspicions()
reports high-entropy strings it did NOT redact as (length, fingerprint) so the
miss rate is measurable. Findings never carry the value — rule, label, length and
sha256[:8], enough to recognise a recurrence, not enough to recover the secret.
FAIL CLOSED: no scrubber, no staging (exit 3), overridable with
MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 for a machine older than this file.
Tested: 42-case corpus in --self-test, including every palace id shape as a
must-not-redact case, idempotency, and a compound case the corpus caught where a
vendor placeholder was then re-matched by the URL rule (nested placeholder — the
secret was hidden either way, so only an exact-output assertion catches it).
End-to-end on 45 real sessions: 29 enforced, fail-closed verified at rc=3,
override verified loud. Server-side layer specified in docs/secret-hygiene.md §5
but NOT implemented — it is tier 2 only there, since the hub cannot see a
client's environment.
|
||
|
|
0fe64c480e |
docs: MEMPALACE_PI_DEVICE now also attributes drawers
Follow-up to
|
||
|
|
c64ffa1d93 |
feeder: default agent to pi@<device> so palace writes carry provenance
mempalace core 3.7.1 records neither which machine nor which harness produced a drawer, and the single shared bearer token means the server cannot distinguish clients. Today both facts survive only incidentally -- device in the per-device inbox path, harness in the pi_*.jsonl filename -- so attribution for anything filed outside the feeder has to be inferred after the fact. Defaulting --agent to pi@$MEMPALACE_PI_DEVICE records both explicitly in a field that already flows through to drawer metadata (added_by), for every enrolled device, with no core change. Falls back to $USER, which is what pre-existing drawers carry (added_by=joakim on fed transcripts, hence the ambiguity). Not pushed: lands on machines at the next image build. |
||
|
|
b609cf5a69 |
docs(synlig): the transcript inbox is the primary's third moving part; and survive an unset HOME
Runbook gaps found while fixing the 2026-08-15 feed failure: - §2.5 still titled "written but not installed" and still asserting "Not installed, not enabled" — false since 2026-08-12. A reader landing there got a flat contradiction of §4 item 3. Retitled, with the verified-2026-08-16 process line, and it now states the fact §2.6 depends on: the server is a NATIVE process (no mempalace container on synlig), so it sees host paths. - New §2.6 documents ~/mempalace-feed/<device>/: why transcripts cannot travel over the HTTPS leg at all, the three client variables, and why MEMPALACE_PI_REMOTE_PATH is the trap (its /data/feed default assumes a containerized server; here it must equal the ssh-target path, and a mismatch fails with rsync succeeding and only the mine failing). Plus the operational notes that cost time: dedup keys on the absolute path so the inbox path is load-bearing, grown sessions are purged+refiled by mtime, and a client-side MCP timeout is NOT a failed mine. - Header status: counts refreshed with an explicit "treat counts as timestamps". Also, bin/mempalace-pi-session: default HOME from the passwd database when it is unset. `docker run --entrypoint="" <image>` inherits no HOME when the image config declares none, and every default is HOME-anchored under `set -u`, so the script — including the palace-free --self-test — died with "HOME: unbound variable" in exactly the environment pi-devbox's smoke suite uses. pi-devbox v1.8.0 lost a release to the same assumption from the other side. |
||
|
|
6e1f4f30fc |
fix(pi-session): a failed remote mine reported success — MCP escapes the payload
The remote-mine leg decided success with `'"error"' in body`. MCP answers a
hard tool failure with HTTP 200 and a JSON-RPC *result* whose content[].text
carries the tool's own JSON as an ESCAPED string, so those bytes are
\"error\" and the substring never matches. On 2026-08-15 (EMB-7KJ4VR4G, first
boot of the fresh pi-devbox image) a mine that failed with
{"success": false, "error": "source directory not found: '/data/feed/emb-7kj4vr4g'"}
printed "Done. Wing 'wing_conversations' updated." directly under that error
and exited 0. Transcripts had been rsynced for the whole session and filed
nowhere; the only artifact anyone would check said it worked.
- classify(): parse the envelope instead of grepping it. Catches JSON-RPC
errors, MCP isError, and inner success=false/error, and distinguishes
"verified ok" from "unverified: no JSON tool payload" rather than assuming.
- --self-test: six recorded MCP responses (fixture 1 is the real 2026-08-15
body) plus a regression guard asserting the old substring check is blind to
it. Needs no palace, no network, no sessions dir.
- Preflight warning when the rsync destination path and
MEMPALACE_PI_REMOTE_PATH disagree. The /data/feed default assumes a
CONTAINERIZED palace server; a native one (systemd unit / uv tool) sees host
paths, and then the two must match. Warned in preflight so --dry-run and
--prepare surface it too.
- Remote mode no longer previews NEW/SKIP from the LOCAL palace: dedup happens
on the palace host keyed on the remote inbox path, so this machine cannot
answer it. Tags become [?] and the summary says who decides. It had been
reporting "6 already filed" about a palace it was not feeding.
|
||
|
|
29e660e18f |
feeders: stage beside the palace, not in ~/.cache; document Phase 1 exposure
Staging default moves out of ~/.cache to <palace-root>/pi-stage (pi) and <palace-root>/opencode-stage (opencode), resolved with mempalace's own palace-path precedence ($MEMPALACE_PALACE_PATH -> $MEMPAL_PALACE_PATH -> ~/.mempalace/config.json -> ~/.mempalace/palace), then dirname. Why: the convos miner keys dedup on the *staged* path, so a wiped stage plus a sync scoped to include it prunes the drawers mined from those sources -- deleting memories, not a cache. Under ~/.cache that state was reachable by anything treating a cache as disposable. Staging inside the palace makes the coupling structural: the stage cannot be wiped without touching the palace itself. Overrides ($MEMPALACE_PI_STAGE / $MEMPALACE_SESSION_STAGE, --stage) are unchanged. Note the old default had never been created on any host, so this closed a latent hazard, not a live one. Measured, and the docs now claim only this much: sync prunes only within the scope it is given -- wing-only, 1299 scanned / 1299 out of scope / 0 removed; scoped at the palace root, 651 kept / 648 out of scope. The previous blanket "sync prunes every drawer" wording overstated it, which is a liability: the next reader disproves the overstatement and discards the real constraint with it. Also in this change: - cron log dir ~/.cache/mempalace-session -> ~/.cache/mempalace-logs. The stage left that namespace, so the old name now read as "the stage". - AGENTS.md: the convos miner *does* check mtime (verified against upstream convo_miner.py); the previous "no mtime check" claim was wrong. - smoke-test assertions use `mktemp -d` for --sessions-dir. One pointed at /tmp, which still held earlier synthetic transcripts, so a --dry-run exported a fake session into the real stage: --dry-run skips the mine, not the export. docs/phase-1-exposure-runbook.md -- the newt/DNS/auth step that RFC 001 and the synlig runbook leave open (runbook section 4, items 2 and 5). Port 8765 at /mcp, newt targets 172.17.0.1, and the authentication is the single shared bearer token (RFC 6.2, decided 2026-08-09) rather than per-device proxy users. The latter cannot work today: mempalace validates exactly one token, and Pangolin's SSO/PIN/password are browser-shaped while every client here is a headless JSON-RPC POST -- enabling that protection breaks the clients it protects. The per-device axis that *does* exist is the feeder's SSH key + per-device inbox. New finding recorded there: a loopback bind does not merely 403 behind a tunnel (already known, runbook 2.4) -- it also silently starts the server with no token at all, because auto-minting is gated on the bind being non-loopback. extensions/pi/README.md: the HTTP transport IS authenticated as of mempalace 3.6.0; the "sessionless and unauthenticated" note dated from the v1.3.0 era. Closes the RFC section 8 Phase-0 hygiene item. |
||
|
|
6352373a1f |
fix(feeders): make post-mine repair opt-in, not default
The three feeder wrappers (mempalace-docs, mempalace-pi-session,
mempalace-session) unconditionally ran 'mempalace repair --yes' after
mining, controllable only via --no-repair opt-out. The contrib launchd
and systemd templates did not pass --no-repair, so every scheduled tick
invoked the destructive in-place HNSW rebuild.
This has bitten us twice:
- 2026-05-04 09:08: a kickstart triggered repair while an MCP
subprocess held the DB open; the live collection was wiped (0
drawers) and had to be restored from the palace.backup snapshot.
- 2026-05-05 10:00: post-mine repair crashed mid-rebuild with
'NotFoundError: Collection [<uuid>] does not exist' - chromadb's
rebuild recreated the collection under a new UUID while the code
still held the old handle. Live DB survived only by luck (crash
hit before the swap).
Fix: flip the default.
- New flag: --repair (opt-in). Prints a warning and sleeps 3s before
invoking 'mempalace repair --yes'.
- --no-repair is retained as a deprecated no-op alias for backward
compatibility with any scripts/units still passing it.
- Default behavior: no repair. Routine ChromaDB add() keeps HNSW
consistent; repair is a recovery op, not a maintenance tick.
Docs updated to match: README, SKILL, ARCHITECTURE, AGENTS,
contrib/README. Scheduling guidance now explicitly warns against
enabling --repair on cron/launchd/systemd-timer runs.
|
||
|
|
9450a45194 |
feat(pi-session): add mempalace-pi-session feeder for pi coding-agent sessions
Parallel to mempalace-session, this wrapper walks ~/.pi/agent/sessions/
JSONL files and mines qualifying sessions into wing_conversations via
'mempalace mine --mode convos'.
Design choices mirror mempalace-session:
- Export-stage-mine idiom with deterministic per-session staging paths
under ~/.cache/mempalace-pi-session/<wing>/, so 'mempalace mine' dedup
on source_file makes re-runs idempotent.
- --dry-run classifies each export as [NEW] or [SKIP] by matching staging
path against the palace's already-filed source_files.
- --min-messages filter skips throwaway single-prompt sessions.
Pi-specific parsing:
- Pi JSONL is a typed tree (id/parentId) per docs/session-format.md;
this walks in file order, which is correct for the overwhelmingly
linear case and harmlessly duplicative on branched sessions (palace
semantic dedup handles it).
- Roles mapped to Claude Code JSONL shape:
user -> {type:user, content:text}
assistant -> {type:assistant, content:[text, tool_use]}
toolResult-> {type:human, content:[tool_result]} (folded back by normalizer)
bashExecution/custom(display)/branchSummary/compactionSummary
-> rendered as text annotations
- thinking blocks and image blocks dropped (noise / palace is text-only).
Source labelling:
- Staging filenames prefixed 'pi_<uuid>.jsonl' so every drawer's
source_file metadata (visible in search results) unambiguously
identifies the harness. Opencode's convention ('<slug>_<id>.jsonl')
is preserved to keep the existing 19k+ drawers deduped.
- Inline synthetic header on first chunk:
[session: <title> | <cwd> | <date> | source: pi]
as a secondary signal.
|