Compare commits
19 Commits
bfe9c5cd4f
...
main
| Author | SHA1 | Date | |
|---|---|---|---|
| 975ab92943 | |||
| 2167a1b033 | |||
| 817b3a82b7 | |||
| dab989b068 | |||
| e68ee2071c | |||
| 309980b62c | |||
| e2b060a940 | |||
| e45f6b4301 | |||
| 21023e7aa0 | |||
| a361b71c40 | |||
| b2b50afcc1 | |||
| de59571966 | |||
| f0bffd1b93 | |||
| 836e35b320 | |||
| 3d47937d06 | |||
| ecc2a9c574 | |||
| e91766286e | |||
| a92c75d070 | |||
| 982b001001 |
@@ -66,6 +66,7 @@ Start here if you are deciding **what to put where**, or running MemPalace on mo
|
||||
|---|---|
|
||||
| [`docs/fleet-memory.md`](docs/fleet-memory.md) | **Start here.** The five stores MemPalace keeps and what each is for; what a *central* palace buys a fleet of machines; when to file a memory versus when to message another device. Diagrams, worked examples. |
|
||||
| [`docs/rfc-003-coordination-log.md`](docs/rfc-003-coordination-log.md) | The inter-agent / inter-device coordination log (`logstream`): storage, append and query semantics, delivery and latency, the trust model, and its landmines. |
|
||||
| [`docs/secret-hygiene.md`](docs/secret-hygiene.md) | Keeping credentials out of the palace — what the feeder scrubs before staging, why detection is name-anchored and never entropy-anchored, and the measured false-positive rate that made one tier report-only |
|
||||
| [`docs/rfc-001-global-palace.md`](docs/rfc-001-global-palace.md) | Why and how a palace is centralised, what should *not* be global, and the Phase-0 landmines (including: never run `mempalace sync` against a shared palace). |
|
||||
| [`docs/rfc-002-joiner.md`](docs/rfc-002-joiner.md) | Replaying a second palace into a shared primary, and what dedupes what. |
|
||||
| [`docs/phase-1-exposure-runbook.md`](docs/phase-1-exposure-runbook.md) | *(moved)* The exposure record for one deployment is now private operator data. The stub names the reusable mechanism it also carried — the `Host`/`Origin` pin above all — which is not yet published elsewhere. |
|
||||
|
||||
+102
-2
@@ -534,11 +534,70 @@ fi
|
||||
# Also classifies each export as NEW/ALREADY FILED (by source_file lookup)
|
||||
# so --dry-run reports the real mine-set size. Classification is advisory;
|
||||
# `mempalace mine --mode convos` is still the authoritative dedup.
|
||||
# The redactor is a sibling module, imported by the heredoc below. Exported
|
||||
# rather than passed as argv so the argv unpack stays stable.
|
||||
#
|
||||
# ${BASH_SOURCE[0]} reports the path the script was INVOKED as, and the image
|
||||
# installs /usr/local/bin/mempalace-pi-session as a symlink into
|
||||
# /opt/mempalace-toolkit/bin. A naive dirname therefore yields /usr/local/bin,
|
||||
# where the redactor does not exist — and because the import is fail-closed, that
|
||||
# turns every feeder tick on every device into "refusing to stage". Measured: the
|
||||
# symlinked invocation exited FATAL while the direct one worked, i.e. it would
|
||||
# have stopped the whole fleet's memory feed at the next image bake. So chase the
|
||||
# symlink chain, and offer fallbacks rather than betting on one answer.
|
||||
#
|
||||
# readlink -f is avoided deliberately: it is GNU/newer-BSD only, and this script
|
||||
# also runs directly on macOS hosts.
|
||||
_mp_resolve() {
|
||||
local p="$1" target
|
||||
while [ -L "$p" ]; do
|
||||
target="$(readlink "$p")" || break
|
||||
case "$target" in
|
||||
/*) p="$target" ;;
|
||||
*) p="$(dirname "$p")/$target" ;;
|
||||
esac
|
||||
done
|
||||
printf '%s' "$p"
|
||||
}
|
||||
_mp_self="$(_mp_resolve "${BASH_SOURCE[0]}")"
|
||||
MEMPALACE_REDACT_DIR="$(cd "$(dirname "$_mp_self")" && pwd):$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd):/opt/mempalace-toolkit/bin"
|
||||
export MEMPALACE_REDACT_DIR
|
||||
export_count=$(python3 - "$PI_SESSIONS_DIR" "$STAGE" "$SESSION_ID" "$SINCE" "$MIN_MESSAGES" "$MIN_ASSISTANT_CHARS" "$MODE" <<'PY'
|
||||
import json, os, sqlite3, sys
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
|
||||
# ── Secret scrubbing before anything is staged ───────────────────────────────
|
||||
# This is the last point at which a transcript is a plain in-memory value: the
|
||||
# local mine reads the staged file, and the remote path rsyncs that same file
|
||||
# byte-for-byte, so scrubbing here covers BOTH transports with one hook.
|
||||
#
|
||||
# FAIL CLOSED. If the redactor cannot be imported, staging is refused rather
|
||||
# than done unscrubbed — a missing module means a broken install, and the whole
|
||||
# point of this step is that a secret must not reach a shared palace. Override
|
||||
# deliberately with MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 if you ever need to feed a
|
||||
# machine whose toolkit is older than this file.
|
||||
# Colon-separated candidates: symlink-resolved dir, invoked dir, then the image's
|
||||
# known install path. First one that has the module wins.
|
||||
for _cand in os.environ.get("MEMPALACE_REDACT_DIR", "").split(":"):
|
||||
if _cand and os.path.isdir(_cand):
|
||||
sys.path.insert(0, _cand)
|
||||
try:
|
||||
import mempalace_redact as _redact
|
||||
except Exception as _e: # noqa: BLE001 - any import failure is fatal by design
|
||||
if os.environ.get("MEMPALACE_FEED_ALLOW_UNSCRUBBED", "").strip() in {"1", "true", "yes"}:
|
||||
_redact = None
|
||||
print(" [WARN] secret scrubber unavailable, staging UNSCRUBBED by request "
|
||||
f"({_e})", file=sys.stderr)
|
||||
else:
|
||||
print(f" [FATAL] secret scrubber unavailable ({_e}); refusing to stage. "
|
||||
"Set MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 to override.", file=sys.stderr)
|
||||
raise SystemExit(3)
|
||||
|
||||
_known = _redact.env_secrets() if _redact else []
|
||||
_redactions = 0
|
||||
_reported = 0
|
||||
|
||||
sessions_dir, stage, session_filter, since, min_messages, min_assistant_chars, mode = sys.argv[1:8]
|
||||
min_messages = int(min_messages)
|
||||
min_assistant_chars = int(min_assistant_chars)
|
||||
@@ -811,6 +870,29 @@ for path in paths:
|
||||
)
|
||||
continue
|
||||
|
||||
# Scrub the parsed objects, not the serialized text: string VALUES get
|
||||
# rewritten while keys, ids and structure are left exactly as they are.
|
||||
# (Scanning raw JSONL instead would also match escape artifacts like the
|
||||
# "\\t" before a field name, inventing keys such as "tapiKey".)
|
||||
if _redact is not None:
|
||||
_found: list = []
|
||||
out_lines = [_redact.scrub_obj(obj, _known, _found)[0] for obj in out_lines]
|
||||
_hard = [x for x in _found if not x.rule.endswith("-reported")]
|
||||
_soft = [x for x in _found if x.rule.endswith("-reported")]
|
||||
_redactions += len(_hard)
|
||||
_reported += len(_soft)
|
||||
if _hard:
|
||||
_by = {}
|
||||
for x in _hard:
|
||||
# "T2:github-pat" — rule says what matched, tier says how much to
|
||||
# trust it, which is what the reader of this line actually needs.
|
||||
_k = f"T{x.tier}:{x.rule}"
|
||||
_by[_k] = _by.get(_k, 0) + 1
|
||||
print(f" [REDACTED] {path.name} "
|
||||
+ ", ".join(f"{k}={v}" for k, v in sorted(_by.items()))
|
||||
+ " fp=" + ",".join(sorted({x.fingerprint for x in _hard})),
|
||||
file=sys.stderr)
|
||||
|
||||
out_path = stage / f"pi_{session_uuid}.jsonl"
|
||||
with out_path.open("w", encoding="utf-8") as f:
|
||||
for obj in out_lines:
|
||||
@@ -835,6 +917,13 @@ for path in paths:
|
||||
|
||||
print(f"EXPORTED {exported}")
|
||||
print(f"ALREADY_FILED {-1 if already_filed is None else skipped_already_filed}")
|
||||
# Report the scrub outcome even when it is zero: "0 redactions" is a measurement,
|
||||
# whereas printing nothing is indistinguishable from a scrubber that never ran.
|
||||
if _redact is not None:
|
||||
print(f" [scrub] {_redactions} redaction(s) applied, "
|
||||
f"{_reported} name-anchored candidate(s) reported only "
|
||||
f"(set MEMPALACE_REDACT_STRICT=1 to redact those too)", file=sys.stderr)
|
||||
|
||||
if skipped_short:
|
||||
print(f"SKIPPED_SHORT {skipped_short}", file=sys.stderr)
|
||||
if skipped_quiet:
|
||||
@@ -885,15 +974,26 @@ fi
|
||||
|
||||
# ── Ship to the palace host (remote mode only) ───────────────────────
|
||||
# mempalace_mine expands its source path in the SERVER process, so in remote
|
||||
# mode the exports have to physically exist over there. rsync --update is the
|
||||
# mode the exports have to physically exist over there. --checksum is the
|
||||
# idempotent half; the mine is the other half.
|
||||
#
|
||||
# NOT --update: the stage file's mtime is deliberately the SOURCE transcript's
|
||||
# mtime (see the os.utime() in the exporter, "preserve session mtime for dedup
|
||||
# stability"), so a re-export of a session that has not been appended to since
|
||||
# the last ship carries an mtime that is NOT newer than the receiver copy. With
|
||||
# --update rsync then SKIPS it silently -- which is exactly the case that must
|
||||
# ship after a redactor change, because the content differs while the mtime does
|
||||
# not. Measured on mbp-m1-2020 2026-08-27: a scrubbed re-export of a dormant
|
||||
# session was skipped and unscrubbed bytes stayed in the palace host's inbox,
|
||||
# while every local signal reported a clean stage. --checksum compares content
|
||||
# and keeps the ship idempotent without trusting timestamps.
|
||||
MINE_SOURCE="$STAGE"
|
||||
if [[ "$MODE" == "remote" ]]; then
|
||||
ssh_cmd="ssh"
|
||||
[[ -n "$SSH_CONFIG" ]] && ssh_cmd="ssh -F $SSH_CONFIG"
|
||||
echo ""
|
||||
echo "Shipping stage to ${SSH_TARGET%/}/$DEVICE/ ..."
|
||||
if ! rsync -a --update --no-owner --no-group \
|
||||
if ! rsync -a --checksum --no-owner --no-group \
|
||||
-e "$ssh_cmd" \
|
||||
--include='*.jsonl' --exclude='*' \
|
||||
"$STAGE/" "${SSH_TARGET%/}/$DEVICE/"; then
|
||||
|
||||
@@ -0,0 +1,604 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Scrub secret-shaped values out of palace-bound text.
|
||||
|
||||
WHY THIS IS NAME-ANCHORED AND NOT ENTROPY-ANCHORED
|
||||
==================================================
|
||||
The obvious design — "redact long random-looking strings" — is actively wrong
|
||||
for this corpus, and the reason is worth stating before anyone tries to
|
||||
"improve" it. A palace is *full* of high-entropy strings that are its own
|
||||
primary keys:
|
||||
|
||||
drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b drawer id
|
||||
..._chunk_000000 chunk id
|
||||
evt_20260826T213924_397b1f710c2d event id
|
||||
rep_d344e349ba276d6fc11997cd552f6937 replica id
|
||||
1787780364792-000000-rep_d344e349ba276d6f… hybrid logical clock
|
||||
ecc2a9c574f4156403cf4e90aea8b0d088c75700 git commit sha
|
||||
|
||||
An entropy or "looks like base64/hex of length N" detector fires on every one
|
||||
of those, i.e. on most identifiers in the corpus, and the redaction damages the
|
||||
memory it was meant to protect. Worse, the damage is silent and unrecoverable:
|
||||
a mined transcript whose ids have been replaced with <redacted> is no longer
|
||||
traceable to anything.
|
||||
|
||||
So detection here is anchored on one of three things that carry *meaning*:
|
||||
|
||||
Tier 1 KNOWN VALUES — literal values taken from this process's own
|
||||
environment, for variables whose NAME says secret.
|
||||
Zero false positives by construction: the value
|
||||
*is* the secret. Catches any presentation (env
|
||||
dump, JSON, error message, URL, prose), which is
|
||||
exactly the incident this exists for.
|
||||
Tier 2 KNOWN SHAPES — vendor-prefixed credentials (ghp_…, glpat-…,
|
||||
xox[abprs]-…, AKIA…, sk-ant-…, JWTs, PEM private
|
||||
keys, credentials embedded in URLs, Authorization
|
||||
headers). Very low false-positive rate because the
|
||||
prefix is meaningful, not merely random.
|
||||
Tier 3 NAME=VALUE — an assignment whose *key* says secret
|
||||
(…TOKEN, …SECRET, …PASSWORD, …API_KEY, …).
|
||||
REPORT-ONLY BY DEFAULT. Measured on 52 MB of real
|
||||
fleet transcripts it fired 403 times, of which the
|
||||
large majority were ${VAR} interpolation in compose
|
||||
files, TypeScript identifiers, a type annotation
|
||||
(`credentials: Credentials`), an IPA attribute
|
||||
holding a date (krbPasswordExpiration) and terminal
|
||||
output following a "Password:" prompt. Rewriting
|
||||
those corrupts code and docs held as memory, to
|
||||
catch what Tier 1 already catches by value.
|
||||
Set MEMPALACE_REDACT_STRICT=1 to make it enforce.
|
||||
|
||||
SHAPE COLLISIONS, AND WHY TIER 1 IS THE LOAD-BEARING ONE
|
||||
--------------------------------------------------------
|
||||
A 40-hex Gitea PAT is byte-indistinguishable from a git commit sha, so no shape
|
||||
rule can separate them. Only two things can: the value being known (Tier 1,
|
||||
which redacts) or a key naming it (Tier 3, which by default only reports). So
|
||||
in the default configuration a sha-shaped PAT is caught if and only if it
|
||||
belongs to this machine. That is an accepted, documented gap — the alternative
|
||||
is redacting every commit sha in the palace.
|
||||
|
||||
WHICH TIER DID THAT? — mapping printed rule names back to tiers
|
||||
--------------------------------------------------------------
|
||||
The output names the *rule* (`github-pat`, `env-value`), because that says what
|
||||
matched. RULE_TIERS below maps rule -> tier, and `tier_of()` is what the feeder
|
||||
uses to print `T2:github-pat=6`, because the tier is what tells a reader how
|
||||
much to trust the hit. A self-test asserts every rule that can appear in a
|
||||
Finding has a tier, so adding a rule without classifying it fails the tests.
|
||||
|
||||
KNOWN FALSE NEGATIVES — stated, not hidden
|
||||
------------------------------------------
|
||||
This will not catch: a novel credential format pasted bare into prose with no
|
||||
name nearby; a secret from a machine whose env this process cannot see; a
|
||||
base64-of-a-secret; a secret split across lines. Tier 1 covers the local
|
||||
machine's own secrets completely, which is where the measured incidents came
|
||||
from — but "the scrubber ran" must never be read as "there are no secrets in
|
||||
here". `suspicions()` exists for exactly that reason: it reports high-entropy
|
||||
candidates it did NOT redact, as fingerprints rather than values, so the false
|
||||
negative rate can be *measured* over time instead of assumed to be zero.
|
||||
|
||||
Run `python3 mempalace_redact.py --self-test` for the test corpus, which
|
||||
includes every benign id shape above as a must-not-redact case.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import math
|
||||
import os
|
||||
import re
|
||||
from typing import Any, Callable, Iterable
|
||||
|
||||
__all__ = ["scrub", "scrub_obj", "suspicions", "Finding"]
|
||||
|
||||
PLACEHOLDER = "<redacted:{label}>"
|
||||
|
||||
# ── Tier 1: names whose VALUES are secrets ────────────────────────────────────
|
||||
# Anchored on the variable name, then applied as an exact literal match.
|
||||
_SECRET_NAME = re.compile(
|
||||
r"(TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|API_?KEY|APIKEY|ACCESS_?KEY"
|
||||
r"|PRIVATE_?KEY|CREDENTIAL|BEARER|SESSION_?KEY|COOKIE|_PAT|^PAT$)",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
# …unless the name says it holds a location or a knob rather than a secret.
|
||||
# SSH_KEY_PATH is the live example: it matches KEY, and its value is a path.
|
||||
_NAME_NOT_SECRET = re.compile(
|
||||
r"(_PATH|_FILE|_DIR|_NAME|_ID|_URL|_URI|_HOST|_PORT|_USER|_ENABLED"
|
||||
r"|_TIMEOUT|_MS|_SECS|_SECONDS|_COUNT|_LIMIT|_MODE)$",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
# Keys that contain a secret word but describe a POLICY, COUNT or TYPE rather
|
||||
# than holding a credential. Measured on real transcripts: krbPasswordExpiration
|
||||
# (an IPA attribute whose value is a date), observationTokens (a token count),
|
||||
# targetTokens, tokensOverTarget.
|
||||
_KEY_NOT_A_SECRET = re.compile(
|
||||
r"(expiration|expiry|_age|count|length|limit|usage|interval|policy|type|class"
|
||||
r"|provider|manager|error|target|overtarget|file|path|dir|name|_id|url|uri"
|
||||
r"|host|port|user|enabled|timeout|_ms|_secs)$",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
# Left of the match: a declaration keyword or a member access means this is CODE
|
||||
# (`const token = x`, `obj.apiKey = y`), and the "value" is an expression.
|
||||
_CODE_CONTEXT = re.compile(
|
||||
r"(?:\b(?:const|let|var|def|func|fn|public|private|readonly|export|return|if|assert)\s+$)"
|
||||
r"|[.\w]\s*$"
|
||||
)
|
||||
_TRIVIAL_VALUES = {
|
||||
"", "0", "1", "true", "false", "none", "null", "nil", "unset", "changeme",
|
||||
"change-me", "xxx", "***", "redacted", "your-token-here", "your_token_here",
|
||||
"placeholder", "example", "dummy", "test", "secret",
|
||||
}
|
||||
|
||||
# ── Tier 2: shapes that mean credential because the PREFIX means credential ──
|
||||
_SHAPE_RULES: list[tuple[str, re.Pattern[str], Callable[[re.Match[str]], str]]] = [
|
||||
# ORDER MATTERS, and the self-test is what proved it. The two structural
|
||||
# rules (credentials inside a URL, Authorization header) run FIRST, before
|
||||
# the vendor-prefix rules: a vendor rule rewriting `https://ghp_xxx@host` to
|
||||
# `https://<redacted:github-token>@host` leaves a colon INSIDE the
|
||||
# placeholder, which the URL rule then reads as user:password and redacts a
|
||||
# second time, producing a nested placeholder. Ordering plus the (?!<redacted)
|
||||
# guards below make the result independent of pass order.
|
||||
("url-credentials",
|
||||
re.compile(r"(?P<pre>[a-zA-Z][a-zA-Z0-9+.\-]*://(?!<redacted)[^/\s:@]+:)(?P<pw>(?!<redacted)[^/\s@]{4,})(?P<at>@)"),
|
||||
lambda m: f"{m.group('pre')}{PLACEHOLDER.format(label='url-password')}{m.group('at')}"),
|
||||
("authorization-header",
|
||||
re.compile(r"(?P<pre>[Aa]uthorization:\s*(?:Bearer|Basic|token)\s+)(?P<val>(?!<redacted)[A-Za-z0-9._\-+/=]{8,})"),
|
||||
lambda m: f"{m.group('pre')}{PLACEHOLDER.format(label='auth-header')}"),
|
||||
("github-pat", re.compile(r"\bgh[pousr]_[A-Za-z0-9]{16,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="github-token")),
|
||||
("github-fine-grained", re.compile(r"\bgithub_pat_[A-Za-z0-9_]{22,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="github-token")),
|
||||
("gitlab-pat", re.compile(r"\bglpat-[A-Za-z0-9_\-]{16,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="gitlab-token")),
|
||||
("slack", re.compile(r"\bxox[abprs]-[A-Za-z0-9\-]{10,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="slack-token")),
|
||||
("openai-anthropic", re.compile(r"\bsk-(?:ant-)?[A-Za-z0-9_\-]{20,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="api-key")),
|
||||
("aws-access-key", re.compile(r"\b(?:AKIA|ASIA)[0-9A-Z]{16}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="aws-access-key-id")),
|
||||
("google-api-key", re.compile(r"\bAIza[0-9A-Za-z_\-]{35}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="google-api-key")),
|
||||
("huggingface", re.compile(r"\bhf_[A-Za-z0-9]{30,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="hf-token")),
|
||||
("npm", re.compile(r"\bnpm_[A-Za-z0-9]{30,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="npm-token")),
|
||||
("digitalocean", re.compile(r"\bdop_v1_[a-f0-9]{60,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="do-token")),
|
||||
# JWT: three base64url segments. Requires the "eyJ" header so a random
|
||||
# dotted string cannot trip it.
|
||||
("jwt", re.compile(r"\beyJ[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="jwt")),
|
||||
# Whole PEM block, non-greedy so several in one file stay separate.
|
||||
("pem-private-key",
|
||||
re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----.*?-----END [A-Z ]*PRIVATE KEY-----", re.DOTALL),
|
||||
lambda m: PLACEHOLDER.format(label="private-key-block")),
|
||||
]
|
||||
|
||||
# ── Tier 3: assignment whose KEY says secret ─────────────────────────────────
|
||||
# The key prefix is LAZY AND OPTIONAL, not "one char then anything": a mandatory
|
||||
# leading character cannot match a key that BEGINS with the secret word, so
|
||||
# `"api_key"` was missed while `"client_secret"` passed — the secret word sits
|
||||
# mid-name in one and at position 0 in the other. Note plain `KEY` is absent from
|
||||
# the alternation on purpose: it is too generic (SSH_KEY_PATH), and tier 1 covers
|
||||
# the real ones by value anyway.
|
||||
# The optional quote after the key is what makes JSON work: `"api_key": "v"`
|
||||
# closes the key BEFORE the colon, so a pattern jumping straight from key to
|
||||
# separator silently misses every JSON payload — i.e. most of a transcript.
|
||||
_ASSIGNMENT = re.compile(
|
||||
r"""(?P<key>\b[A-Za-z0-9_.\-]*?
|
||||
(?:token|secret|password|passwd|passphrase|api[_-]?key|apikey
|
||||
|access[_-]?key|private[_-]?key|credential)
|
||||
[A-Za-z0-9_.\-]*)
|
||||
(?P<kq>["']?)
|
||||
(?P<sep>\s*(?:=>|[=:])\s*)
|
||||
(?P<q>["']?)
|
||||
(?P<val>[^\s"'<>,;)\]}]{12,})
|
||||
(?P=q)""",
|
||||
re.IGNORECASE | re.VERBOSE,
|
||||
)
|
||||
# A value that is obviously not a live secret: a reference, a path, a
|
||||
# placeholder, or something we already redacted.
|
||||
_NOT_A_SECRET_VALUE = re.compile(
|
||||
r"""^(?:
|
||||
\$\{[^}\s]*\}? # ${VAR} ${VAR:-default} ${VAR:?err}
|
||||
# trailing } optional: the value
|
||||
# charset stops before it, so the
|
||||
# captured text is "${VAR" — which
|
||||
# still means "a reference", and
|
||||
# missing that is what made a
|
||||
# compose file look like a secret
|
||||
| \$[A-Za-z_][A-Za-z0-9_]* # $VAR
|
||||
| %[A-Za-z_][A-Za-z0-9_]*%? # %VAR%
|
||||
| \{\{[^}\s]*\}?\}? # {{ template }}
|
||||
| <[^>]*> # <redacted:…> / <your-token>
|
||||
| (?:/|~/|\./|\.\./)[^\s]* # a path
|
||||
| [A-Za-z]:\\ # a windows path
|
||||
| \*+ | x+ | \.+ # ***, xxx
|
||||
| (?:true|false|none|null|nil|unset|changeme|placeholder|example)
|
||||
)$""",
|
||||
re.IGNORECASE | re.VERBOSE,
|
||||
)
|
||||
|
||||
|
||||
class Finding:
|
||||
"""One redaction, recorded without ever carrying the secret itself."""
|
||||
|
||||
__slots__ = ("rule", "label", "length", "fingerprint")
|
||||
|
||||
def __init__(self, rule: str, label: str, value: str) -> None:
|
||||
self.rule = rule
|
||||
self.label = label
|
||||
self.length = len(value)
|
||||
# Stable across sessions, so a recurring leak is recognisable as the
|
||||
# same one, while the value stays unrecoverable from the report.
|
||||
self.fingerprint = hashlib.sha256(value.encode("utf-8")).hexdigest()[:8]
|
||||
|
||||
@property
|
||||
def tier(self) -> int:
|
||||
"""Which detection tier produced this. See RULE_TIERS for the mapping."""
|
||||
return tier_of(self.rule)
|
||||
|
||||
def __repr__(self) -> str: # pragma: no cover - diagnostics only
|
||||
return (f"<Finding T{self.tier}:{self.rule}:{self.label} "
|
||||
f"len={self.length} fp={self.fingerprint}>")
|
||||
|
||||
|
||||
def env_secrets(environ: dict[str, str] | None = None) -> list[tuple[str, str]]:
|
||||
"""Tier 1 material: (name, value) pairs from the environment worth scrubbing.
|
||||
|
||||
Deliberately conservative: a secret-sounding NAME is not enough, the value
|
||||
must also be long enough to be a credential and must not look like a path
|
||||
or a knob. Longest values first so that if one secret contains another as a
|
||||
substring, the longer one is replaced before the shorter can fragment it.
|
||||
"""
|
||||
env = os.environ if environ is None else environ
|
||||
out: list[tuple[str, str]] = []
|
||||
for name, value in env.items():
|
||||
if not value or len(value) < 12:
|
||||
continue
|
||||
if not _SECRET_NAME.search(name) or _NAME_NOT_SECRET.search(name):
|
||||
continue
|
||||
if value.strip().lower() in _TRIVIAL_VALUES:
|
||||
continue
|
||||
if _NOT_A_SECRET_VALUE.match(value.strip()):
|
||||
continue
|
||||
if "/" in value and (value.startswith("/") or value.startswith("~")):
|
||||
continue # a path that happened to be named …_KEY
|
||||
if any(c.isspace() for c in value):
|
||||
continue # credentials do not contain whitespace; prose does
|
||||
out.append((name, value))
|
||||
out.sort(key=lambda kv: len(kv[1]), reverse=True)
|
||||
return out
|
||||
|
||||
|
||||
# Rule name -> tier. Printed output names the rule (WHAT matched); the tier says
|
||||
# HOW MUCH TO TRUST IT: 1 = near-certain (the value is known to be a secret),
|
||||
# 2 = strong (a vendor prefix means what it says), 3 = candidate, needs eyes.
|
||||
RULE_TIERS: dict[str, int] = {
|
||||
"env-value": 1,
|
||||
**{name: 2 for name, _pat, _lbl in _SHAPE_RULES},
|
||||
"named-assignment": 3,
|
||||
"named-assignment-reported": 3,
|
||||
# Tier 0 is not a detection: it is a measured NON-detection, reported so the
|
||||
# false-negative rate is a number instead of an assumption.
|
||||
"suspicion": 0,
|
||||
}
|
||||
|
||||
|
||||
def tier_of(rule: str) -> int:
|
||||
"""Tier for a rule name, or -1 if the rule was never classified."""
|
||||
return RULE_TIERS.get(rule, -1)
|
||||
|
||||
|
||||
def scrub(
|
||||
text: str,
|
||||
known: Iterable[tuple[str, str]] | None = None,
|
||||
findings: list[Finding] | None = None,
|
||||
strict: bool | None = None,
|
||||
) -> tuple[str, list[Finding]]:
|
||||
"""Redact secrets from one string. Returns (scrubbed, findings).
|
||||
|
||||
Idempotent: the replacement text matches none of the rules, so running this
|
||||
twice changes nothing and cannot nest placeholders.
|
||||
"""
|
||||
found: list[Finding] = findings if findings is not None else []
|
||||
if strict is None:
|
||||
strict = os.environ.get("MEMPALACE_REDACT_STRICT", "").strip() in {"1", "true", "yes"}
|
||||
if not text:
|
||||
return text, found
|
||||
|
||||
# Tier 1 first: an exact known value should be labelled with its variable
|
||||
# name (the most useful thing a reader can be told) rather than by shape.
|
||||
for name, value in (env_secrets() if known is None else known):
|
||||
if value and value in text:
|
||||
count = text.count(value)
|
||||
text = text.replace(value, PLACEHOLDER.format(label=name))
|
||||
for _ in range(count):
|
||||
found.append(Finding("env-value", name, value))
|
||||
|
||||
# Tier 2.
|
||||
for rule, pattern, repl in _SHAPE_RULES:
|
||||
def _sub(m: re.Match[str], _rule: str = rule) -> str:
|
||||
found.append(Finding(_rule, _rule, m.group(0)))
|
||||
return repl(m)
|
||||
text = pattern.sub(_sub, text)
|
||||
|
||||
# Tier 3 — REPORT-ONLY unless strict. See STRICT note in the module docstring:
|
||||
# measured on 52 MB of real fleet transcripts this rule produced 403 hits of
|
||||
# which the overwhelming majority were ${VAR} interpolation, TypeScript
|
||||
# identifiers, YAML env lists and documentation prose. Redacting those would
|
||||
# corrupt code and docs stored as memory, to catch secrets that tier 1
|
||||
# already catches by value. So by default a tier-3 hit is REPORTED (so the
|
||||
# miss can be measured and the rule calibrated) and NOT rewritten.
|
||||
def _sub_assignment(m: re.Match[str]) -> str:
|
||||
val, key = m.group("val"), m.group("key")
|
||||
pre = m.string[max(0, m.start() - 24):m.start()]
|
||||
if (_NOT_A_SECRET_VALUE.match(val)
|
||||
or val.strip().lower() in _TRIVIAL_VALUES
|
||||
or _KEY_NOT_A_SECRET.search(key)
|
||||
or _CODE_CONTEXT.search(pre)
|
||||
or val.isdigit()):
|
||||
return m.group(0)
|
||||
found.append(Finding("named-assignment" if strict else "named-assignment-reported",
|
||||
key, val))
|
||||
if not strict:
|
||||
return m.group(0)
|
||||
return f"{m.group('key')}{m.group('kq')}{m.group('sep')}{m.group('q')}" \
|
||||
f"{PLACEHOLDER.format(label=m.group('key'))}{m.group('q')}"
|
||||
|
||||
text = _ASSIGNMENT.sub(_sub_assignment, text)
|
||||
return text, found
|
||||
|
||||
|
||||
def scrub_obj(obj: Any, known: Iterable[tuple[str, str]] | None = None,
|
||||
findings: list[Finding] | None = None,
|
||||
strict: bool | None = None) -> tuple[Any, list[Finding]]:
|
||||
"""Recursively scrub every string VALUE in a JSON-ish structure.
|
||||
|
||||
Keys are left alone: a key named "token" is metadata, not a credential, and
|
||||
rewriting keys would corrupt the transcript schema.
|
||||
"""
|
||||
found: list[Finding] = findings if findings is not None else []
|
||||
known = list(env_secrets() if known is None else known)
|
||||
if isinstance(obj, str):
|
||||
return scrub(obj, known, found, strict)[0], found
|
||||
if isinstance(obj, dict):
|
||||
return {k: scrub_obj(v, known, found, strict)[0] for k, v in obj.items()}, found
|
||||
if isinstance(obj, list):
|
||||
return [scrub_obj(v, known, found, strict)[0] for v in obj], found
|
||||
return obj, found
|
||||
|
||||
|
||||
# ── Measuring what we MISS, without leaking it ───────────────────────────────
|
||||
_BENIGN_ID = re.compile(
|
||||
r"""^(?:
|
||||
drawer_[\w.\-]+ # drawer / chunk ids
|
||||
| (?:evt|art)_\d{8}T\d{6}_[0-9a-f]{12} # event / artifact ids
|
||||
| rep_(?:[0-9a-f]{12}|[0-9a-f]{32}) # replica ids
|
||||
| \d{13}-[0-9a-f]{6}-rep_[0-9a-f]+ # hybrid logical clocks
|
||||
| [0-9a-f]{7,12} # short git sha
|
||||
| [0-9a-f]{40} # git sha / sha1
|
||||
| [0-9a-f]{64} # sha256
|
||||
| [0-9a-f]{8}(?:-[0-9a-f]{4}){3}-[0-9a-f]{12} # uuid
|
||||
| sha256:[0-9a-f]{64}
|
||||
| \d+
|
||||
)$""",
|
||||
re.VERBOSE | re.IGNORECASE,
|
||||
)
|
||||
_CANDIDATE = re.compile(r"[A-Za-z0-9+/_\-=]{24,}")
|
||||
|
||||
|
||||
def _entropy(s: str) -> float:
|
||||
if not s:
|
||||
return 0.0
|
||||
counts: dict[str, int] = {}
|
||||
for ch in s:
|
||||
counts[ch] = counts.get(ch, 0) + 1
|
||||
n = len(s)
|
||||
return -sum((c / n) * math.log2(c / n) for c in counts.values())
|
||||
|
||||
|
||||
def suspicions(text: str, min_entropy: float = 3.6) -> list[Finding]:
|
||||
"""High-entropy strings this module did NOT redact, as fingerprints.
|
||||
|
||||
This is the false-negative meter. It deliberately does not redact anything:
|
||||
every one of these is far more likely to be one of the palace's own ids than
|
||||
a credential, and redacting them would be the damage described at the top of
|
||||
this file. Reporting them as (length, fingerprint) makes the miss rate
|
||||
observable — and a fingerprint that keeps recurring is worth a human look.
|
||||
"""
|
||||
out: list[Finding] = []
|
||||
for m in _CANDIDATE.finditer(text):
|
||||
val = m.group(0)
|
||||
if _BENIGN_ID.match(val) or val.startswith("<redacted:"):
|
||||
continue
|
||||
if _entropy(val) < min_entropy:
|
||||
continue
|
||||
out.append(Finding("suspicion", "high-entropy", val))
|
||||
return out
|
||||
|
||||
|
||||
# ── Self-test ────────────────────────────────────────────────────────────────
|
||||
def _self_test() -> int:
|
||||
fake_token = "pFrDBfakVKQB0SLN1hTzAoaRJeQ4J_go7xmJezojAwI" # shape-alike, not real
|
||||
env = {
|
||||
"MEMPALACE_REMOTE_TOKEN": fake_token,
|
||||
"GITEA_TOKEN": "0123456789abcdef0123456789abcdef01234567", # 40 hex: sha-shaped
|
||||
"SSH_KEY_PATH": "/Users/someone/.ssh", # name matches KEY, value is a path
|
||||
"MEMPALACE_MAILBOX_POLL_MS": "300000", # knob, not a secret
|
||||
"MEMPALACE_REMOTE_URL": "https://palace.example.com/mcp",
|
||||
"PASSWORD": "changeme", # trivial value
|
||||
"API_KEY": "${FROM_VAULT}", # a reference
|
||||
}
|
||||
known = env_secrets(env)
|
||||
|
||||
must_redact = [
|
||||
("env dump line", f"MEMPALACE_REMOTE_TOKEN={fake_token}"),
|
||||
("bare value in prose", f"the token is {fake_token} apparently"),
|
||||
("value inside json", f'{{"env": {{"MEMPALACE_REMOTE_TOKEN": "{fake_token}"}}}}'),
|
||||
("sha-shaped pat, name-anchored",
|
||||
"GITEA_TOKEN=0123456789abcdef0123456789abcdef01234567"),
|
||||
("github pat", "remote add origin https://ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345@git"),
|
||||
("gitlab pat", "GLPAT is glpat-AbCdEfGhIjKlMnOpQrSt"),
|
||||
("slack", "xoxb-1234567890-ABCdefGHIjklMNO"),
|
||||
("anthropic-ish", "sk-ant-api03-AbCdEfGhIjKlMnOpQrStUvWx"),
|
||||
("aws", "AKIAIOSFODNN7EXAMPLE"),
|
||||
("jwt", "eyJhbGciOiJIUzI1NiJ9.eyJzdWIiOiIxMjM0NTY3ODkwIn0.dozjgNryP4J3jVmNHl0w5N_XgL0n3I9PlFUP0THsR8U"),
|
||||
("pem", "-----BEGIN OPENSSH PRIVATE KEY-----\nb3BlbnNzaA\n-----END OPENSSH PRIVATE KEY-----"),
|
||||
("url creds", "git clone https://joakim:hunter2hunter2@git.example.com/x.git"),
|
||||
("auth header", "Authorization: Bearer abcdefghijklmnop"),
|
||||
("json api_key", '"api_key": "AbCdEfGhIjKlMnOpQrSt"'),
|
||||
("json nested quotes", '{"auth": {"client_secret": "AbCdEfGhIjKlMnOpQrStUv"}}'),
|
||||
("yaml style", " gitea_token: 0123456789abcdefghij"),
|
||||
("lowercase password kv", "db_password = s3cr3t-p4ssw0rd-xyz"),
|
||||
]
|
||||
must_not_redact = [
|
||||
("drawer id", "drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b"),
|
||||
("chunk id", "drawer_pi-devbox_landmines_ca8929fea82cf80e7770ea0d_chunk_000000"),
|
||||
("event id", "evt_20260826T213924_397b1f710c2d"),
|
||||
("replica id", "rep_d344e349ba276d6fc11997cd552f6937"),
|
||||
("hlc", "1787780364792-000000-rep_d344e349ba276d6fc11997cd552f6937"),
|
||||
("git sha", "ecc2a9c574f4156403cf4e90aea8b0d088c75700"),
|
||||
("short sha", "ecc2a9c"),
|
||||
("sha256", "a" * 64),
|
||||
("uuid session file", "pi_01a03fd8-0dc6-73c2-a9c3-6b62ccf0318f.jsonl"),
|
||||
("knob assignment", "MEMPALACE_MAILBOX_POLL_MS=300000"),
|
||||
("path named key", "SSH_KEY_PATH=/Users/someone/.ssh"),
|
||||
("empty token", "MEMPALACE_REMOTE_TOKEN="),
|
||||
("already redacted", "MEMPALACE_REMOTE_TOKEN=<redacted:MEMPALACE_REMOTE_TOKEN>"),
|
||||
("var reference", "MEMPALACE_REMOTE_TOKEN=${VAULT_TOKEN}"),
|
||||
("placeholder", "MEMPALACE_REMOTE_TOKEN=<your-token-here>"),
|
||||
("trivial", "PASSWORD=changeme"),
|
||||
("url no creds", "MEMPALACE_REMOTE_URL=https://palace.example.com/mcp"),
|
||||
("prose", "The mailbox is an obligation channel, not a news channel."),
|
||||
("base64 in a diff line", "+ data = 'QUJDREVGR0hJSktMTU5PUFFSU1RVVldYWVo='"),
|
||||
]
|
||||
|
||||
failures = 0
|
||||
# Tier 3 is report-only by default, so the corpus below is checked in STRICT
|
||||
# mode (where tier 3 rewrites) and the default mode is asserted separately.
|
||||
print("MUST REDACT (strict: tier 3 enforcing)")
|
||||
for name, sample in must_redact:
|
||||
out, found = scrub(sample, known, strict=True)
|
||||
ok = bool(found) and all(
|
||||
secret not in out for secret in (fake_token, "hunter2hunter2",
|
||||
"0123456789abcdef0123456789abcdef01234567")
|
||||
)
|
||||
failures += not ok
|
||||
rules = ",".join(sorted({f.rule for f in found})) or "NONE"
|
||||
print(f" {'pass' if ok else 'FAIL'} {name:32s} [{rules}]")
|
||||
print("MUST NOT REDACT (strict: the hard cases)")
|
||||
for name, sample in must_not_redact:
|
||||
out, found = scrub(sample, known, strict=True)
|
||||
ok = out == sample and not found
|
||||
failures += not ok
|
||||
detail = "" if ok else f" -> {out!r} {found}"
|
||||
print(f" {'pass' if ok else 'FAIL'} {name:32s}{detail}")
|
||||
|
||||
print("DEFAULT MODE (tier 3 reports, does not rewrite)")
|
||||
for name, sample, expect_rewrite in [
|
||||
("env value still redacted", f"MEMPALACE_REMOTE_TOKEN={fake_token}", True),
|
||||
("vendor shape still redacted", "ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345", True),
|
||||
("name-anchored only reported", "db_password = s3cr3t-p4ssw0rd-xyz", False),
|
||||
("compose interpolation untouched", "GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}", False),
|
||||
("typescript identifier untouched", "const tokens = countTokensForModel", False),
|
||||
("ipa attribute untouched", "--setattr=krbPasswordExpiration=20260529090000", False),
|
||||
]:
|
||||
out, found = scrub(sample, known)
|
||||
rewritten = out != sample
|
||||
ok = rewritten == expect_rewrite
|
||||
if name == "name-anchored only reported":
|
||||
ok = ok and any(f.rule == "named-assignment-reported" for f in found)
|
||||
if name.endswith("untouched"):
|
||||
ok = ok and not any(f.rule.startswith("named-assignment") for f in found)
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} {name:34s}"
|
||||
+ ("" if ok else f" -> rewritten={rewritten} {found}"))
|
||||
|
||||
print("TIER MAPPING")
|
||||
_rules = ({"env-value", "named-assignment", "named-assignment-reported", "suspicion"}
|
||||
| {n for n, _p, _l in _SHAPE_RULES})
|
||||
_unclassified = sorted(r for r in _rules if tier_of(r) < 0)
|
||||
failures += bool(_unclassified)
|
||||
print(f" {'pass' if not _unclassified else 'FAIL'} every rule maps to a tier"
|
||||
+ (f" -> unclassified: {_unclassified}" if _unclassified
|
||||
else f" ({len(_rules)} rules)"))
|
||||
for label, sample, want in [
|
||||
("env value is tier 1", f"MEMPALACE_REMOTE_TOKEN={fake_token}", 1),
|
||||
("vendor shape is tier 2", "ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345", 2),
|
||||
("name-anchored is tier 3", "db_password = s3cr3t-p4ssw0rd-xyz", 3),
|
||||
]:
|
||||
got = scrub(sample, known)[1]
|
||||
ok = bool(got) and all(f.tier == want for f in got)
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} {label:26s}"
|
||||
+ ("" if ok else f" -> {[(f.rule, f.tier) for f in got]}"))
|
||||
|
||||
print("PROPERTIES")
|
||||
once, _ = scrub(f"MEMPALACE_REMOTE_TOKEN={fake_token}", known, strict=True)
|
||||
twice, second = scrub(once, known, strict=True)
|
||||
ok = once == twice and not second
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} idempotent (no nested placeholders)")
|
||||
|
||||
# The compound case the corpus caught: assert on EXACT output, not merely
|
||||
# that the secret is gone — a nested placeholder also hides the secret, and
|
||||
# would have shipped unnoticed.
|
||||
compound, _ = scrub("https://ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345@git.example.com/x", known, strict=True)
|
||||
ok = compound == "https://<redacted:github-token>@git.example.com/x"
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} vendor placeholder not re-redacted by url rule"
|
||||
+ ("" if ok else f" -> {compound!r}"))
|
||||
|
||||
already, found_a = scrub("https://joakim:<redacted:url-password>@git.example.com", known, strict=True)
|
||||
ok = already == "https://joakim:<redacted:url-password>@git.example.com" and not found_a
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} already-redacted url left alone")
|
||||
|
||||
obj = {"role": "user", "content": [{"type": "text", "text": f"export TOK={fake_token}"}],
|
||||
"id": "evt_20260826T213924_397b1f710c2d"}
|
||||
scrubbed, found = scrub_obj(obj, known, strict=True)
|
||||
ok = (fake_token not in repr(scrubbed)
|
||||
and scrubbed["id"] == obj["id"]
|
||||
and scrubbed["role"] == "user"
|
||||
and len(found) >= 1)
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} nested structure scrubbed, ids and keys preserved")
|
||||
|
||||
ok = all(fake_token not in repr(f) for f in scrub(f"x {fake_token}", known, strict=True)[1])
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} findings never carry the secret value")
|
||||
|
||||
sus = suspicions("drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b "
|
||||
"rep_d344e349ba276d6fc11997cd552f6937 "
|
||||
"1787780364792-000000-rep_d344e349ba276d6fc11997cd552f6937 "
|
||||
"ecc2a9c574f4156403cf4e90aea8b0d088c75700")
|
||||
ok = not sus
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} suspicion meter ignores our own id shapes ({len(sus)} hits)")
|
||||
|
||||
sus2 = suspicions("blob=Zm9vYmFyYmF6cXV1eHF1dXhmb29iYXJiYXpxdXV4Zm9vYmFy")
|
||||
ok = len(sus2) == 1 and all("Zm9v" not in repr(s) for s in sus2)
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} suspicion meter flags unknown blobs as fingerprints only")
|
||||
|
||||
print(f"\n{'ALL PASS' if not failures else f'{failures} FAILED'}")
|
||||
return 1 if failures else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
import sys
|
||||
if "--self-test" in sys.argv:
|
||||
raise SystemExit(_self_test())
|
||||
# Filter mode: scrub stdin -> stdout, report to stderr. Useful for spot
|
||||
# checks and for scrubbing a file by hand without writing throwaway code.
|
||||
data = sys.stdin.read()
|
||||
out, found = scrub(data)
|
||||
sys.stdout.write(out)
|
||||
if found:
|
||||
by_rule: dict[str, int] = {}
|
||||
for f in found:
|
||||
by_rule[f.rule] = by_rule.get(f.rule, 0) + 1
|
||||
print(f"[redact] {len(found)} redaction(s): "
|
||||
+ ", ".join(f"{k}={v}" for k, v in sorted(by_rule.items())), file=sys.stderr)
|
||||
if "--suspicions" in sys.argv:
|
||||
for s in suspicions(out):
|
||||
print(f"[suspicion] len={s.length} fp={s.fingerprint}", file=sys.stderr)
|
||||
+11
-11
@@ -13,14 +13,14 @@ This document covers: what the stores are, what a central palace changes when se
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Q["ask by MEANING<br/>'what do we know about X?'"] --> D["<b>Drawers</b><br/>chroma.sqlite3<br/>verbatim text, embedded"]
|
||||
R["ask by ENTITY + TIME<br/>'what was true in June?'"] --> G["<b>Knowledge graph</b><br/>knowledge_graph.sqlite3<br/>typed facts, time-bounded"]
|
||||
S["ask by ADDRESS + ORDER<br/>'what is waiting for me?'"] --> L["<b>Coordination log</b><br/>logstream.sqlite3<br/>addressed events + artifacts"]
|
||||
T["ask by ASSOCIATION<br/>'what else touches this?'"] --> P["<b>Palace graph</b><br/>tunnels.json + hallways.json<br/>links between rooms"]
|
||||
D --> DI["<b>Diaries</b> live here too:<br/>drawers with room=diary,<br/>read by recency, not similarity"]
|
||||
Q["ask by MEANING<br/>'what do we know about X?'"] --> D["<b>Drawers</b><br/>chroma.sqlite3"]
|
||||
R["ask by ENTITY + TIME<br/>'what was true in June?'"] --> G["<b>Knowledge graph</b><br/>knowledge_graph.sqlite3"]
|
||||
S["ask by ADDRESS<br/>'what is waiting for me?'"] --> L["<b>Coordination log</b><br/>logstream.sqlite3"]
|
||||
T["ask by ASSOCIATION<br/>'what else touches this?'"] --> P["<b>Palace graph</b><br/>tunnels.json + hallways.json"]
|
||||
D --> DI["<b>Diaries</b> live here too:<br/>drawers with room=diary"]
|
||||
```
|
||||
|
||||
All four stores are files inside **one palace directory**, so "the palace" is a directory you can back up in one go.
|
||||
All four stores are files inside **one palace directory**, so "the palace" is a directory you can back up in one go. Each holds a different *shape* of thing: drawers hold verbatim text, embedded for similarity; the knowledge graph holds typed facts that are time-bounded; the coordination log holds addressed events and their artifacts, read in append order; the palace graph holds links between rooms. Diaries are drawers, but they are read by recency rather than by similarity.
|
||||
|
||||
| Store | You get things back by | Typical use | Wrong use |
|
||||
|---|---|---|---|
|
||||
@@ -110,19 +110,19 @@ sequenceDiagram
|
||||
|
||||
## 3. Deciding where something goes
|
||||
|
||||
The question that matters is not "is this important?" but **"how will I want this back, and does anyone need to act?"**
|
||||
The question that matters is not "is this important?" but **"how will I want this back, and does a specific machine or agent need to act?"**
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Start["I have something worth keeping"] --> Act{"Must a specific<br/>machine or agent<br/>DO something?"}
|
||||
Start["I have something worth keeping"] --> Act{"Must someone specific<br/>DO something?"}
|
||||
Act -->|no| Change{"Is it a fact that<br/>changes over time?"}
|
||||
Act -->|yes| Know{"Do they also need<br/>to KNOW it later?"}
|
||||
Change -->|yes| KG["Knowledge graph<br/>kg_add / kg_supersede"]
|
||||
Change -->|no| Mine{"Is it about MY session<br/>— what I tried, felt, learned?"}
|
||||
Change -->|no| Mine{"Is it about MY session<br/>— tried, felt, learned?"}
|
||||
Mine -->|yes| Diary["Diary entry"]
|
||||
Mine -->|no| Drawer["Drawer<br/>wing + room, verbatim"]
|
||||
Know -->|yes| Both["BOTH:<br/>drawer for the knowledge,<br/>event pointing at it"]
|
||||
Know -->|no| Event["Coordination event<br/>to_agent = the specific agent"]
|
||||
Know -->|yes| Both["BOTH:<br/>drawer + event pointing at it"]
|
||||
Know -->|no| Event["Coordination event<br/>to_agent = that agent"]
|
||||
```
|
||||
|
||||
Worked examples:
|
||||
|
||||
@@ -171,10 +171,76 @@ An event is **still owed** when all of these hold:
|
||||
|
||||
1. it is directed at you (`to_agent = <you>`, and `to_agent = '*'` is excluded — §7.4);
|
||||
2. it was not written by you (you cannot owe yourself);
|
||||
3. **no** event of yours has *all* of: a strictly higher `seq`, a join to it (`metadata.ack_of = <id>` or a shared `correlation_id`), **and** a terminal status (`applied` / `superseded` / `failed` / `blocked`).
|
||||
3. **no** event of yours has *all* of: a strictly higher `seq`, a join to it (`metadata.ack_of = <id>` or a shared `correlation_id`), **and** a terminal status (`applied` / `superseded` / `failed` / `blocked`);
|
||||
4. **the requester has not explicitly withdrawn it** — see §3.3.1.
|
||||
|
||||
The strictly-higher-`seq` test is load-bearing: without it, one terminal reply on a `correlation_id` suppresses every *later* ask on that same correlation, permanently. And because `seq` is local arrival order, this test is only sound on a single replica — §7.3.
|
||||
|
||||
#### 3.3.1 Withdrawal: the one case where someone else may clear your obligation
|
||||
|
||||
> **Status: implemented in the toolkit, not yet in a released image.** Clause 4 is
|
||||
> live in `extensions/pi/mempalace.ts` (`isWithdrawn`) with rule tests in
|
||||
> `scripts/test-owed-withdrawal.sh`. It reaches a device only when an image bakes
|
||||
> a toolkit revision containing it. Until then clauses 1–3 are the whole rule, and
|
||||
> **a sender must assume its withdrawal has no effect on the recipient's mailbox.**
|
||||
|
||||
Clause 3 says "no event **of yours**", and that asymmetry is deliberate: owed-ness
|
||||
is a statement about the *recipient's* accountability, so a requester must not be
|
||||
able to delete an obligation the recipient genuinely has. It is wrong in exactly
|
||||
one case — the requester retracting its **own** ask.
|
||||
|
||||
**Measured, 2026-09-08/09.** `pi@mbp-m1-2020` withdrew a v1.8.13 rollout ask to
|
||||
`pi@tor-ms22`: terminal `superseded`, same `correlation_id`, `metadata.closes`
|
||||
naming the thread, body "DO NOT SPEND A MINUTE ON v1.8.13". It recorded the
|
||||
withdrawal as done. The withdrawal had **no effect** — its `from_agent` is mbp, so
|
||||
it could never satisfy a join that only inspects tor-ms22's own events. tor-ms22's
|
||||
next wake-up, 8h later and on the first boot of the image shipping this very
|
||||
derivation, still listed the ask as owed, 41h old, for a release that had been
|
||||
superseded and never installed there. Closing it cost the recipient a write on a
|
||||
thread nobody wanted answered. **The failure is invisible from the sender's side:**
|
||||
mbp did everything a sender is told to do and got a result indistinguishable from
|
||||
success, which is why this went unnoticed for 41h rather than being caught at once.
|
||||
|
||||
A withdrawal is honoured only when **all** of these hold. Each is a guard against a
|
||||
specific way a mailbox could otherwise be emptied silently, and each is pinned by a
|
||||
mutation test:
|
||||
|
||||
| Requirement | The failure it prevents |
|
||||
|---|---|
|
||||
| `from_agent` = the **original requester** | a third party clearing someone else's obligation |
|
||||
| `to_agent` = the recipient **exactly** (not `'*'`) | one broadcast emptying every machine's mailbox at once (§7.4) |
|
||||
| terminal status | a `claimed`/`ready` note reading as a release |
|
||||
| strictly after the ask (`hlc`, §7.3) | retiring an ask the requester sent **later** on the same correlation |
|
||||
| joins the ask (`ack_of` or shared `correlation_id`) | clearing an unrelated thread |
|
||||
| **an explicit marker**: `metadata.withdraws` **or** `metadata.closes`, whose value is exactly the ask's `correlation_id` or its event `id` | **the dangerous one** — see below |
|
||||
|
||||
**Why release must be stated rather than inferred.** The tempting rule is "any
|
||||
terminal event from the requester clears it". That breaches this log's standing
|
||||
safety direction, which is that every failure stays on the *noisy-but-visible*
|
||||
side: a spuriously resurfacing item costs one turn of human correction, while a
|
||||
suppressed unanswered ask is silent and permanent. Under the naive rule a
|
||||
requester appending `applied` for its own bookkeeping — on the correlation,
|
||||
addressed to the recipient, before the recipient ever replied — would silently
|
||||
delete a real obligation. `withdraws` is canonical; `closes` is honoured because it
|
||||
is already the fleet's de-facto marker (mbp seq 119; emb seq 117 and 118 all use
|
||||
it). Prose does not count: the value must *name* the thread, so
|
||||
`closes: "v1813-client-rollout-tor-ms22"` is a withdrawal and
|
||||
`closes: "v1813-… — from the RECIPIENT side, which is the only side that can"` is
|
||||
not.
|
||||
|
||||
**This does not break the fixed point in §9.2, and a reader will reach for §9.2 to
|
||||
object.** That section rejects letting terminal directed events into the owed set,
|
||||
because then a reply becomes owed by its requester, closing it mints a fresh
|
||||
obligation, and the loop never terminates. That argument is about **candidates**.
|
||||
Clause 4 adds a **clearer**. A withdrawal carries a terminal status, so it can never
|
||||
be a candidate, and nothing new becomes owed: the asserting shape (`open`) and the
|
||||
clearing shape (terminal) stay disjoint, which is what makes owed-ness terminate.
|
||||
|
||||
**Recipients keep the last word.** A withdrawal suppresses; it does not rewrite
|
||||
history. The recipient may still append its own terminal reply — worth doing when
|
||||
there is a finding to carry across, as tor-ms22 did at seq 122, salvaging two
|
||||
measurements from the retracted thread.
|
||||
|
||||
### 3.4 Artifacts
|
||||
|
||||
Exact payloads for handoff: `patch`, `file`, `log`, `json`, `note`; UTF-8 text only; ≤4 MiB; `sha256` and `size_bytes` returned (`logstream.py:744-810`).
|
||||
@@ -279,6 +345,10 @@ Measured on real data: the positive-control pair carries `seq` 26/27 and `hlc` `
|
||||
|
||||
Two consequences: **broadcasting an ask reaches no owed set at all** (the "don't broadcast an ask" anti-pattern is mechanically enforced, not merely advised), and to reach a whole fleet with something actionable you must write **one directed event per device**, sharing a `correlation_id` so the thread stays joinable. Separately, an event written with **no** `to_agent` matches no `to_agent=` query ever and is addressed to nobody.
|
||||
|
||||
**✅ Measured 2026-08-26, on two devices independently.** This was inferred from source until a two-arm control settled it. A `to_agent='*'` event with `status='open'` — the anti-pattern, planted deliberately — survives the raw candidate filter and so reaches the exclusion branch, which every previously written broadcast (all non-open) never did. Result: the raw query returned it, the derived owed set did not, and the delivery named only the directed arm of the pair. Confirmed on a second machine that had merely *received* the broadcast rather than planted it. Two arms differing only in `to_agent`, same poll and same fake sender, is what makes it evidence rather than an absence: "delivered exactly one of two" cannot be explained by a mailbox that delivers everything or nothing.
|
||||
|
||||
> ⚠️ **Correction to an earlier record.** A note filed the same day claimed this branch *could not* be exercised, reasoning that every broadcast in the log carries `status='ready'` and the protocol forbids writing an open broadcast. The reasoning was wrong in a specific and instructive way: the protocol forbids it *as production traffic*, which does not forbid planting one as a labelled, self-closed control. Where a rule blocks a measurement, check whether it blocks the *use* or the *test* before recording the branch as untestable.
|
||||
|
||||
### 7.5 `since_event_id` raises on an unknown id
|
||||
|
||||
An unresolvable cursor raises `ValueError` rather than returning empty (`logstream.py:966-970`) — stronger than the tool description promises, and benign on one replica. On a second replica, a resuming watcher whose cursor has not yet replicated will **hard-fail instead of waiting**.
|
||||
@@ -309,6 +379,49 @@ The complete GET route table is `/healthz`, `/statusz`, `/logstream/stream`, `/s
|
||||
|
||||
WAL. **Measured 2026-08-26** on the primary: `logstream.sqlite3` mtime three days old while `logstream.sqlite3-wal` was 865 KiB and seconds old. An operator checking whether coordination is live must look at the `-wal` file, or query.
|
||||
|
||||
### 7.11 Delivered is not read: the human is the trigger
|
||||
|
||||
The mid-session poll delivers with `deliverAs: "steer"` and **deliberately no `triggerTurn`** — waking a model on inbound fleet traffic is a much larger behavioural change than auto-delivery, and was not approved. The poll fires on `agent_settled`, which means the agent is *idle*: nothing is running to react. So the delivered text is queued and read at the top of the **next turn**, whenever a human happens to start one.
|
||||
|
||||
**Measured 2026-08-26:** a delivery landed in a session window at ~20:50Z and sat there, visibly unreacted-to, until the operator asked "do I have to nudge you for you to read it?" — which is precisely the question this design produces. The delay is not the agent choosing to ignore its mailbox; between the poll and the next turn there is no inference at all.
|
||||
|
||||
The consequence is a **UX gap, not a defect**: from the outside, an inbound ask plus an idle agent reads as neglect. The delivered text explained how to *close* an ask but said nothing about *when* it would be seen — so the one reader who needed that fact, the human watching the window, was the one not told.
|
||||
|
||||
**✅ Addressed 2026-08-26 (toolkit `extensions/pi/mempalace.ts`), without touching the no-`triggerTurn` decision.** Two changes, in the order their value is realised:
|
||||
|
||||
1. **A delivery note in the message itself** — it says this is a queued message, that nothing woke the agent, and that any message starts the turn that will handle it. Costs nothing, changes no behaviour, and converts "why is it ignoring me?" into "right, I nudge it."
|
||||
2. **A notification at poll time**, because the note only reaches someone already looking, and the case that actually loses an ask is nobody looking. `MEMPALACE_MAILBOX_NOTIFY` unset → in-TUI `ctx.ui.notify` (the surface `session_start` already uses); `=desktop` → additionally a terminal-native notification (Kitty `OSC 99`, else `OSC 777`), which is how a ping escapes a container with no `notify-send`, DBus or host access — the escape sequence is interpreted by the terminal emulator on the human's machine; `=0` → silent. Title and body are stripped of `;` and control bytes so a payload cannot forge an OSC field or terminate the sequence early. It fires only when something is due, because a ping on an empty poll trains its reader to ignore it.
|
||||
|
||||
**⚠️ Unverified through a multiplexer — measured 2026-08-26, tor-ms22.** The terminal-native path is written and syntax-checked but has **not** been observed to fire. On this client the layering is `kitty → tmux (on the host) → docker exec → pi (in the container)`, with two clients attached to the same tmux session (one local, one remote over SSH). Test sequences written directly to pi's tty produced **no notification on the remote client**; the local client is unobserved so far. The likely mechanism is that **tmux drops OSC sequences it does not itself implement** — reaching the outer terminal requires wrapping them in tmux's DCS passthrough (`ESC P tmux; … ESC \`, with inner `ESC` bytes doubled) *and* `allow-passthrough on`, which is not the default.
|
||||
|
||||
This compounds §7.11's detection problem rather than repeating it: **the container cannot see either layer**. `KITTY_WINDOW_ID` is absent because `docker exec` does not forward it, and `TMUX` is absent because tmux is running one level further out, on the host — so the client cannot detect the terminal it is speaking to *or* the multiplexer it must speak through. Explicit configuration is the only reliable route; autodetection is structurally blind here.
|
||||
|
||||
Open, and deliberately not settled by guessing: whether the local client shows it (pending); whether to add DCS-wrapping modes; and with two clients attached, **which** client should be notified — a pane's output reaches every attached client, so a work laptop and a home machine would both ping.
|
||||
|
||||
Still deliberately **not** done: `MEMPALACE_MAILBOX_TRIGGER=1` (opt-in, default off) to start a turn on arrival. Worth revisiting once notification is proven in daily use — but a machine that auto-turns on inbound events writes to a shared log with nobody watching, and §7.1 (no idempotency guard) is the reason to be slow about it.
|
||||
|
||||
### 7.12 The mailbox is an obligation channel, so a report addressed to you is never delivered
|
||||
|
||||
Candidates are drawn with `status="open"` (§3.3), so an event carrying a **terminal** status is not a mailbox candidate at all — no matter who it is addressed to. A `task.reply` sent to a named device to share a finding is therefore delivered to nobody, ever, and neither is any `event.ack`. It sits in the log until someone reads the log.
|
||||
|
||||
Read the filter precisely: candidacy requires **exactly `open`**, not merely "non-terminal". So a `claimed` announcement is equally undelivered — announcing that you have started work reaches the requester's mailbox no more than your finished report does. Claiming is still worth doing on long or handoff-prone work, because it puts *pickup* in the log for whoever later asks "did anyone start this before that container died?", but its audience is a log reader, not the requester. Note also that claiming does not quiet **your own** mailbox: the original ask stays owed until a terminal event of yours joins it (§3.3), which is exactly what the state machine means by `open --> claimed` *does NOT clear*.
|
||||
|
||||
**Measured 2026-08-26:** a device wrote a detailed report addressed to `pi@<peer>` with `status="applied"`, on a thread whose ask belonged to a third party. The recipient never saw it and only read it when the operator pointed at it by id — with the mailbox working exactly as specified throughout.
|
||||
|
||||
This is the shape of the channel, and it is worth stating because the natural inter-device message — *"here is something you should know"* — is exactly the shape that gets no delivery. Two ways to make news reach a peer: address it as a **directed `status="open"` ask** so it enters the owed set and gets closed when acted on (correct when a response is genuinely wanted), or accept that it is **pull-only** and pair it with a palace drawer, which the peer's search *will* surface later. What does not work is a terminal-status report and an expectation of attention.
|
||||
|
||||
### 7.13 An event addressed to an identity no session runs as is write-only
|
||||
|
||||
The owed set (§3.3) is the log's only push channel — it is what a client injects at wake-up. It keys on `to_agent`. A `from_agent` is whatever string the writer puts there (§6: "authenticates the fleet, not the agent"), so nothing stops a writer from authoring under an identity that no live session has ever run as. When it later receives a reply, that reply is addressed right back to the string the writer chose — and if nothing runs as that string, the reply is stored, queryable, and delivered to no one. Unlike §7.12, this is not about status; a **directed, `status="open"`** reply is undeliverable here, because the *address* rather than the *shape* is the defect.
|
||||
|
||||
**Measured 2026-08-27.** A device planted two directed `task.request`s under a synthetic `from_agent` (a release-provenance label, not an identity any session runs as), intending only to record where the ask came from. Both recipients replied correctly, addressed to that synthetic sender as the protocol requires — one an `event.ack` (`status="claimed"`), one a `task.reply` (`status="applied"`). Neither reply was ever delivered to anyone. One contained a live-credential exposure finding. It sat unread for roughly two and a half hours and was found only because a human asked whether mail had arrived.
|
||||
|
||||
**The general rule, not the instance:** before addressing an event, the writer is responsible for the address being an identity a live session actually runs as; and a writer that authors under an identity other than its own session identity has made every reply to that event undeliverable to itself. This fleet has already produced the failure in three unrelated shapes — a synthetic sender (above); an event addressed to a device since decommissioned; and a device-identity case mismatch (an uppercase form written where the fleet's convention is lowercase) — which is why this is recorded as a property of the channel rather than a mistake to remember not to repeat.
|
||||
|
||||
**The permitted exception, and its obligation.** Authoring under a synthetic identity is legitimate for controlled experiments — this fleet has run positive/negative mailbox controls this way deliberately — and this RFC does not forbid it. But a synthetic sender then **must name the real identity to reply to** in the body; the address line is not a place to record provenance if a reply is ever wanted.
|
||||
|
||||
**Proposed, not implemented:** on `event_append`, if `from_agent` differs from the writer's own stamped session identity, warn that replies to this event will be addressed to that other identity, and say whether any event in the log has ever been authored under it — a plain `DISTINCT from_agent` lookup. State the heuristic's limit honestly: an identity that has never authored is certainly unread by anyone; one that *has* authored may still belong to a device that no longer exists, which this check cannot see.
|
||||
|
||||
---
|
||||
|
||||
## 8. Phasing
|
||||
@@ -338,6 +451,9 @@ Replication exists in code — `version_vector()`, `list_ops()`, `apply_remote_e
|
||||
5. **Should `GET /logstream/events` exist?** A read-only HTTP tail would let non-MCP consumers (dashboards, CI) follow a stream without an MCP client. Today they must hold SSE or speak MCP.
|
||||
6. **Undocumented limits.** The 64 KiB `metadata` cap and the lowercase-`type` regex are enforced server-side but absent from the tool schemas (§3.1). Document, or relax.
|
||||
7. **Agent-name registry.** Nothing prevents two devices sharing one `from_agent` (§7.7). A warning at append time would be cheap.
|
||||
8. ~~**Does an arriving ask deserve a notification?**~~ **Done 2026-08-26** — delivery note plus `MEMPALACE_MAILBOX_NOTIFY` (§7.11). What stays open is the narrower question: should `desktop` be the *default* rather than opt-in, and does an arriving ask ever deserve a turn (`MEMPALACE_MAILBOX_TRIGGER`)?
|
||||
9. **Is there a delivery path for news rather than obligations?** (§7.12) Today the only delivered shape is a directed open ask. Options: leave it pull-only and rely on the paired drawer, or give informational events a distinct low-priority surface at wake-up ("3 reports since your last session") separate from the owed set. **Proposed direction in §9.2** — including why the obvious fix (widening the owed set) is not available.
|
||||
10. **Should a write warn when `from_agent` is not the writer's own identity?** (§7.13) Related to decision 7 but distinct: decision 7 is about two devices *sharing* one identity by accident; this is about a writer deliberately authoring under an identity nobody runs, which makes every reply to that event undeliverable to itself. A `DISTINCT from_agent` lookup at append time is cheap and catches the unread-for-certain case; it cannot catch a decommissioned device that once wrote, which is the harder half of the same problem.
|
||||
|
||||
### 9.1 Retention — decided direction (2026-08-26)
|
||||
|
||||
@@ -356,6 +472,27 @@ Three constraints any implementation has to respect:
|
||||
|
||||
Open sub-questions: the window length (a quarter is the obvious first guess); whether cold storage is a sibling `logstream-archive-<period>.sqlite3` or plain files on disk; and whether archives are queryable through the same tools behind an explicit opt-in flag, or simply left as files for a human to open when a question reaches back that far.
|
||||
|
||||
### 9.2 A news surface, without new state — proposed direction (2026-08-27)
|
||||
|
||||
Decision 9 above is **proposed in direction**: news should get its own low-priority surface at wake-up, and the owed set should not be touched.
|
||||
|
||||
**Why the obvious fix is not available.** The tempting change is to let terminal directed events into the owed set, so a report addressed to a device reaches it. That breaks the derivation's fixed point. Owed-ness is defined (§3.3) as *directed at me, not written by me, and not joined by a later terminal event of mine* — so the asserting shape (`open`) and the clearing shape (terminal) must be disjoint. Make terminal events owed, and: a reply becomes owed by its requester; the requester clears it by writing a terminal event joined to it; that event is directed, terminal, and therefore owed by the original author; and every closure mints a fresh obligation. The loop never terminates. You could exempt `task.reply` and `event.ack` by `type`, but that is the same filter re-entered through a different door, with more surface to get wrong. **`status="open"` is not an arbitrary editorial choice about tone; it is what makes owed-ness terminate.**
|
||||
|
||||
**The proposal.** A second, clearly-separated section at wake-up — "N reports since your last session" — listing directed non-`open` events newer than the reader's own last activity, count-capped, bodies not inlined.
|
||||
|
||||
The cost that sinks the naive version is state: "since your last session" implies a per-device read cursor, and a cursor in container-local storage is precisely what a `--force-recreate` erases (the palace exists because container state does not survive). No new state is needed, because **the device's own last authored event is already a cursor**: surface directed events whose `hlc` is greater than the newest `hlc` among events written by this device. That anchor lives in the shared log, survives recreate, and is comparable across replicas for the same reason the owed-set join now uses `hlc` (§7.3).
|
||||
|
||||
Properties worth stating before anyone implements it:
|
||||
|
||||
- **Per stream, not global.** Take the anchor as the newest `hlc` of this device's events *in that stream*. A global anchor lets a chatty stream advance past unread news in a quiet one — the failure is silent, so this is not a refinement to leave for later.
|
||||
- **A device that has never written sees everything.** Bounded by the count cap, and arguably correct on first boot; it is the same shape as a new joiner reading recent history.
|
||||
- **A busy device gets a narrow window.** Correct by construction: you saw the log when you last wrote to it.
|
||||
- **There is no read receipt, so repeats are possible** until the anchor advances. Acceptable only because this surface is a summary, not an interruption: count and one-line subjects, never bodies.
|
||||
- **It must not resurface.** Owed items re-announce (`MEMPALACE_MAILBOX_RESURFACE_MS`); news must not, or it becomes nagging without obligation.
|
||||
- **Opt-in first**, following the notification precedent (§7.11): inert unless explicitly enabled, and never counted in the owed total.
|
||||
|
||||
**What this does not fix.** It is still not push (§7.12 and the 404 on `/logstream/stream` for this deployment), still queued into the next turn rather than waking anyone (§7.11), and still useless to a device that never starts a session. For anything that must survive indefinitely, the paired **drawer** remains the durable channel; this surface only shortens the delay before a peer notices something already written.
|
||||
|
||||
---
|
||||
|
||||
## 10. Evidence index
|
||||
|
||||
@@ -0,0 +1,185 @@
|
||||
# Secret hygiene: keeping credentials out of the palace
|
||||
|
||||
A palace is mined from **transcripts**, and transcripts contain whatever the
|
||||
terminal printed. Print an environment, `cat` a `.env`, paste a `curl -H
|
||||
"Authorization: …"`, and the secret becomes a drawer — searchable by every agent
|
||||
on the fleet, on every machine, indefinitely. This document describes the
|
||||
scrubbing that exists, what it deliberately does **not** try to do, and the two
|
||||
layers it belongs in.
|
||||
|
||||
Measured on this fleet, 2026-08-27: a single leaked bearer token had reached
|
||||
**3 drawers, 13 feeder inbox files across all three devices, and 10 local files
|
||||
dating back 10 days**. Nobody was careless; an agent inspected an environment
|
||||
variable while debugging. That frequency is the design premise — this is a
|
||||
routine event to be handled by the pipeline, not an incident to be handled by
|
||||
discipline.
|
||||
|
||||
## 1. Where the hook goes, and why there are two of them
|
||||
|
||||
| Layer | Implemented in | Sees | Catches |
|
||||
|---|---|---|---|
|
||||
| **Client, pre-staging** | `bin/mempalace_redact.py`, called from `bin/mempalace-pi-session` at the point the staged transcript is written | the machine's own environment **and** the transcript | the real-world case: a secret this machine holds, printed into a session, before it is uploaded anywhere |
|
||||
| **Server, pre-persist** | *not yet implemented* — see §5 | only the text it is handed | secrets from clients that predate the scrubber, from non-pi clients, and secrets an agent types straight into `add_drawer` |
|
||||
|
||||
The client hook is the load-bearing one, because it is the only layer that can
|
||||
compare text against **the actual secret values it holds** (see tier 1 below) and
|
||||
because it stops the leak *before transmission*. The server hook is defence in
|
||||
depth: pattern-only, but it covers write paths the feeder never sees.
|
||||
|
||||
Client placement matters for one specific reason: the pi feeder writes the staged
|
||||
transcript from an in-memory structure, and the remote path then `rsync`s **that
|
||||
same file** byte-for-byte. Scrubbing at the write point therefore covers local
|
||||
mining and remote upload with a single hook — there is no second serialization to
|
||||
forget.
|
||||
|
||||
## 2. Detection is anchored on meaning, never on entropy
|
||||
|
||||
The tempting design — "redact long random-looking strings" — is actively
|
||||
destructive here, because a palace is *full* of high-entropy strings that are its
|
||||
own primary keys:
|
||||
|
||||
```
|
||||
drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b drawer id
|
||||
evt_20260826T213924_397b1f710c2d event id
|
||||
rep_d344e349ba276d6fc11997cd552f6937 replica id
|
||||
1787780364792-000000-rep_d344e349ba276d6f… hybrid logical clock
|
||||
ecc2a9c574f4156403cf4e90aea8b0d088c75700 git commit sha
|
||||
```
|
||||
|
||||
An entropy detector fires on every one of those, and the resulting redaction is
|
||||
silent, permanent, and destroys traceability. So detection uses three anchors
|
||||
that carry meaning instead:
|
||||
|
||||
Each tier is a different *kind of evidence* that a string is a credential. The
|
||||
tier is not a severity ranking of the secret — it is how much to trust the
|
||||
detection. Rule names appear in the output; the tier tells you how to read them.
|
||||
|
||||
| Tier | Anchor | Default | False-positive risk |
|
||||
|---|---|---|---|
|
||||
| **1 — known values** | literal values from this process's env, for variables whose *name* says secret (`…TOKEN`, `…SECRET`, `…PASSWORD`, `…API_KEY`) | **redact** | none by construction: the value *is* the secret |
|
||||
| **2 — known shapes** | vendor-prefixed credentials: `ghp_…`, `github_pat_…`, `glpat-…`, `xox[abprs]-…`, `sk-…`, `AKIA…`, `AIza…`, `hf_…`, JWTs, PEM private-key blocks, credentials inside URLs, `Authorization:` headers | **redact** | very low: the prefix is meaningful, not random |
|
||||
| **3 — name=value** | an assignment whose *key* says secret | **report only** | measured **high** — see §3 |
|
||||
|
||||
### Worked examples, one line each
|
||||
|
||||
```
|
||||
MEMPALACE_REMOTE_TOKEN=pfrDBfak… tier 1 — value matches this env's secret
|
||||
the token is pfrDBfak… apparently tier 1 — same value, bare in prose, still caught
|
||||
git clone https://joakim:hunter2@git/x tier 2 — credentials in a URL
|
||||
Authorization: Bearer abcdefghijklmnop tier 2 — header shape
|
||||
ghp_AbCdEf… / glpat-… / AKIA… / sk-ant-… tier 2 — vendor prefix
|
||||
db_password = s3cr3t-p4ssw0rd-xyz tier 3 — only the KEY suggests it (reported)
|
||||
GITEA_TOKEN=0123456789abcdef… (40 hex) tier 3 — indistinguishable from a commit sha
|
||||
```
|
||||
|
||||
### Which tier fired? Read it off the output
|
||||
|
||||
The tool prints **rule names**, not tier numbers, because the rule says *what*
|
||||
matched. The feeder prefixes them with the tier so both are visible:
|
||||
|
||||
```
|
||||
[REDACTED] 2026-06-27T23-13-49.jsonl T1:env-value=1, T2:github-fine-grained=1 fp=ad78c7d4,fcd95ab5
|
||||
[scrub] 29 redaction(s) applied, 141 name-anchored candidate(s) reported only
|
||||
```
|
||||
|
||||
`RULE_TIERS` in `mempalace_redact.py` is the authoritative mapping, `tier_of()`
|
||||
reads it, and a self-test fails if any rule is left unclassified — so the two
|
||||
vocabularies cannot drift apart silently.
|
||||
|
||||
Tier 1 catches any presentation of a secret — env dump, JSON, error message,
|
||||
URL, prose — because it matches the value itself. It is also, **in the default
|
||||
configuration, the only tier that resolves a shape collision**: a 40-hex Gitea
|
||||
PAT is byte-identical to a git commit sha, so only a known value (tier 1,
|
||||
redacts) or a naming key (tier 3, reports) can tell them apart. A sha-shaped PAT
|
||||
is therefore caught if and only if it belongs to this machine — an accepted gap,
|
||||
since the alternative is redacting every commit sha in the palace.
|
||||
|
||||
## 3. The false-positive measurement, which changed the design
|
||||
|
||||
Tier 3 was originally going to redact. Measured against **52 MB of real fleet
|
||||
transcripts (45 sessions)** it produced **403 hits**, and inspection of them (with
|
||||
values masked) showed the overwhelming majority were not secrets:
|
||||
|
||||
- `GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}` — docker-compose interpolation
|
||||
- `const tokens = countTokensForModel` — TypeScript source stored as memory
|
||||
- `refreshToken(credentials: Credentials …)` — a *type annotation*
|
||||
- `--setattr=krbPasswordExpiration=20260529…` — an IPA attribute holding a date
|
||||
- `|1.FORGE.TOKENS:…` — AAAK diary shorthand
|
||||
- `(root@host) Password:` followed by unrelated terminal output
|
||||
|
||||
Redacting those would corrupt code, configuration and documentation held as
|
||||
memory, in order to catch secrets tier 1 already catches by value. After adding
|
||||
guards for interpolation (`${VAR}`, `$VAR`, `%VAR%`, `{{tpl}}`), code context,
|
||||
non-secret key suffixes (`…Expiration`, `…Count`, `…Type`) and all-digit values,
|
||||
the enforced count on the same corpus fell from **403 to 29** — and those 29 are
|
||||
tier-1 and tier-2 hits, i.e. real credential shapes.
|
||||
|
||||
So tier 3 **reports and does not rewrite** by default. Set
|
||||
`MEMPALACE_REDACT_STRICT=1` to make it enforce, e.g. on a machine whose
|
||||
transcripts are configuration-heavy rather than code-heavy.
|
||||
|
||||
## 4. Known false negatives — stated, not hidden
|
||||
|
||||
The scrubber will not catch a novel credential format pasted bare into prose
|
||||
with no name nearby, a secret belonging to a machine whose environment this
|
||||
process cannot see, a base64-of-a-secret, or a secret split across lines.
|
||||
|
||||
**"The scrubber ran" must never be read as "there are no secrets in here."** To
|
||||
keep that measurable rather than assumed, two things are reported:
|
||||
|
||||
- every scrub prints a count — including `0 redaction(s)`, because printing
|
||||
nothing is indistinguishable from a scrubber that never ran;
|
||||
- `suspicions()` reports high-entropy strings it did **not** redact, as
|
||||
`(length, fingerprint)` pairs rather than values, so the miss rate can be
|
||||
tracked over time and a recurring fingerprint can be investigated by hand.
|
||||
|
||||
Findings never carry the secret. A `Finding` holds the rule, the label, the
|
||||
length and `sha256(value)[:8]` — enough to recognise the same leak recurring,
|
||||
not enough to recover it.
|
||||
|
||||
**Fail closed.** If the redactor cannot be imported, the feeder refuses to stage
|
||||
rather than staging unscrubbed (`exit 3`). Override deliberately with
|
||||
`MEMPALACE_FEED_ALLOW_UNSCRUBBED=1`.
|
||||
|
||||
## 5. The server-side layer (not yet implemented)
|
||||
|
||||
Three call sites, because the server has three near-duplicate validators rather
|
||||
than one:
|
||||
|
||||
| Path | Function | Covers |
|
||||
|---|---|---|
|
||||
| drawers | `config.py` → `sanitize_content()` | `add_drawer`, `update_drawer`, `diary_write`, `checkpoint` — all four route through it |
|
||||
| events | `logstream.py` → `_sanitize_body()` | `event_append`, and `event_ack` transitively |
|
||||
| artifacts | `logstream.py` → `put_artifact()` | inlines its own checks; needs its own edit |
|
||||
|
||||
Server-side scrubbing is **tier 2 only** (plus optional tier-3 reporting): the
|
||||
hub cannot see a client's environment, so tier 1 is structurally unavailable
|
||||
there. That asymmetry is the reason the client hook is not redundant.
|
||||
|
||||
## 6. Cleaning up a leak that already landed
|
||||
|
||||
1. **Never re-echo the value while hunting it.** Pass it via stdin, never argv
|
||||
(visible in `ps` on a shared host). Note `get_drawer` is a trap: reading a
|
||||
drawer in order to redact it prints the secret back into the live transcript.
|
||||
2. **Redact, don't delete.** Replacing the value in place keeps the mined
|
||||
transcript's memory value; a stale embedding vector is a cheap price.
|
||||
3. **Scrub the feeder inbox too**, or the next mine re-files it.
|
||||
4. **A live session file needs an equal-length in-place overwrite** — the agent
|
||||
holds it open, so temp-file-plus-rename loses everything appended afterwards.
|
||||
5. **Expect residue.** SQLite keeps old page content in freed pages until
|
||||
`VACUUM`, so a raw byte scan still matches after a successful `UPDATE`. Decide
|
||||
explicitly whether that matters: if the credential store on the same host is
|
||||
plaintext anyway, it usually does not.
|
||||
|
||||
## 7. Usage
|
||||
|
||||
```bash
|
||||
# unit tests, including the must-not-redact corpus of palace id shapes
|
||||
python3 bin/mempalace_redact.py --self-test
|
||||
|
||||
# scrub anything on stdin; report goes to stderr
|
||||
some-command | python3 bin/mempalace_redact.py > clean.txt
|
||||
|
||||
# also list high-entropy strings that were NOT redacted, as fingerprints
|
||||
python3 bin/mempalace_redact.py --suspicions < transcript.jsonl > clean.jsonl
|
||||
```
|
||||
+141
-2
@@ -117,7 +117,7 @@ side of the wiring.
|
||||
| `MEMPALACE_FEED_WING` | `wing_conversations` | Target wing — passed to both the exporter and the `mempalace_mine` call. |
|
||||
| `MEMPALACE_FEED_DEBOUNCE_MS` | `600000` (10 min) | Minimum gap between mid-session (`agent_settled`) feeds. Bounds crash loss to one window instead of a whole session. |
|
||||
| `MEMPALACE_FEED_PREPARE_TIMEOUT_MS` | `120000` | Kills a wedged `--prepare` subprocess. |
|
||||
| `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `30000` | Caps the `mempalace_mine` call so a stalled palace can't hang session exit. |
|
||||
| `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `300000` (5 min) | Bounds how long the extension *waits* for `mempalace_mine`, so a stalled palace can't hang session exit. It does **not** cancel the mine — see [Debugging](#debugging). Raised from `30000` in 2026-09: the mine is the slowest call this extension makes (30–60s in normal operation), so the old deadline fired routinely and reported healthy behaviour as an error. |
|
||||
|
||||
**Remote palace:** if `$MEMPALACE_REMOTE_URL` is set (see
|
||||
[Transport](#transport-local-vs-external)), `mempalace_mine`'s source path is
|
||||
@@ -316,6 +316,46 @@ waits for the agent to think of asking. Two delivery points, both fail-silent:
|
||||
`triggerTurn` — waking the model on inbound fleet traffic is a much larger
|
||||
behavioural change than auto-delivery.
|
||||
|
||||
**Which means a human is the trigger, and the mailbox now says so.** Measured
|
||||
2026-08-26 on two devices: a delivery lands, the agent is idle, nothing happens,
|
||||
and the operator asks *"do I have to nudge you for you to read this?"*. Yes —
|
||||
because between the poll and the next turn no inference is running. The old text
|
||||
explained how to *close* an ask and never said when it would be *seen*, so the
|
||||
only reader who needed that fact was the one not told. Two additions, neither of
|
||||
which touches the no-`triggerTurn` decision:
|
||||
|
||||
- **A delivery note in the message itself** — states that this is a queued
|
||||
message, that nothing woke the agent, and that any message starts the turn
|
||||
that handles it. Free, and aimed at the human reading the window.
|
||||
- **A notification at poll time**, because the note only helps someone who is
|
||||
already looking, and the case that loses an ask is nobody looking:
|
||||
|
||||
| `MEMPALACE_MAILBOX_NOTIFY` | Behaviour |
|
||||
|---|---|
|
||||
| *unset* (default) | in-TUI `ctx.ui.notify`, the same surface `session_start` already uses |
|
||||
| `desktop` | additionally a terminal-native notification — Kitty `OSC 99`, else `OSC 777` (iTerm2, WezTerm, Ghostty, rxvt-unicode) |
|
||||
| `0` / `off` | silent; mailbox still delivers |
|
||||
|
||||
The `desktop` path is how a notification escapes a container without
|
||||
`notify-send`, DBus or any host access: the escape sequence is written to stdout
|
||||
and interpreted by the terminal emulator on the human's own machine. It is
|
||||
opt-in because writing raw escapes is a behaviour change on a shared machine,
|
||||
not because it is unreliable. Title and body are stripped of `;` and control
|
||||
bytes, so a payload can neither forge an OSC field nor end the sequence early.
|
||||
|
||||
**⚠️ Not yet observed firing through tmux (2026-08-26).** If pi runs inside a
|
||||
multiplexer — e.g. `kitty → tmux on the host → docker exec → pi` — tmux drops
|
||||
OSC sequences it does not implement, so the ping can vanish silently between
|
||||
the container and the human. Reaching the outer terminal needs tmux's DCS
|
||||
passthrough plus `allow-passthrough on`, which is not implemented here yet.
|
||||
Note the client can detect *neither* layer: `KITTY_WINDOW_ID` is not forwarded
|
||||
by `docker exec`, and `TMUX` is unset because tmux runs one level further out.
|
||||
Verify with a hand-written sequence in your own stack before trusting it.
|
||||
|
||||
It fires **only when something is due** — the same condition as the delivery
|
||||
itself. A ping on an empty poll would train its reader to ignore it, which is
|
||||
the failure this whole feature exists to reverse.
|
||||
|
||||
Owed-ness is **derived, never read off a field**, because `event_ack` appends and
|
||||
`status` is written once: a directed `open` event matches the mailbox query
|
||||
*forever*, answered or not. Two calls (`to_agent=<me> status=open`, and
|
||||
@@ -338,9 +378,61 @@ the one failure this derivation exists to prevent.
|
||||
protocol says a broadcast owes nobody a reply — which also means broadcasting an
|
||||
ask demonstrably reaches no owed set, giving that documented anti-pattern teeth.
|
||||
|
||||
**Dormant asks — open, owed, and deliberately not announced.** Some asks are
|
||||
planted *unanswerable on purpose*: a first-boot acceptance describes work that
|
||||
becomes possible only when the device is next recreated, and it must stay owed
|
||||
until then, because closing it early to tidy the mailbox is exactly how the work
|
||||
gets lost. Derivation alone cannot tell that apart from neglected work, so such
|
||||
an ask was announced at every session start and every poll for as long as it was
|
||||
correctly waiting — measured at three days running on `emb-7kj4vr4g`
|
||||
(2026-09-28 → 2026-10-01), and across three consecutive releases before that.
|
||||
The ask was right; announcing it was wrong, and the cost landed on the human
|
||||
reading the window, the one reader who cannot filter it.
|
||||
|
||||
An ask may therefore declare, in its own `metadata`, the condition under which it
|
||||
is merely waiting:
|
||||
|
||||
```json
|
||||
"dormant_unless": [
|
||||
{ "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json",
|
||||
"field": "release_tag", "baseline": "v1.9.4" },
|
||||
{ "kind": "file_mtime", "path": "/etc/hostname",
|
||||
"baseline": "2026-09-22T18:12:49Z" }
|
||||
]
|
||||
```
|
||||
|
||||
It is dormant while **every** condition still matches its baseline, and goes live
|
||||
the moment **any** of them differs — the trigger such asks already stated in
|
||||
prose ("act when EITHER differs"), now in a form the bridge can check. Two kinds,
|
||||
both local-file-only: `file_mtime` (compared at whole seconds in UTC, because a
|
||||
filesystem mtime carries sub-second residue that a reported ISO baseline never
|
||||
will — comparing raw milliseconds would mark every such predicate permanently
|
||||
"changed" and silently disable the feature) and `json_field` (dotted paths
|
||||
allowed, compared as strings so a manifest holding `3` matches a baseline of
|
||||
`"3"`). There is no expression language, no shell and no network: a general
|
||||
evaluator in the path that decides whether work is *visible* is a far worse trade
|
||||
than a clumsy schema.
|
||||
|
||||
Dormancy is **withheld from the announcement, never from the mailbox**: the
|
||||
wake-up injection lists dormant asks once per session with their ids, and a
|
||||
mid-session poll mentions only a count, and only when the window is already open
|
||||
for something else. They are not added to the resurface map, so one becomes
|
||||
announceable the instant its baseline moves.
|
||||
|
||||
**Dormancy must be proven, never assumed** — every unevaluable predicate shows
|
||||
the ask as owed. A missing file, an unreadable one, a baseline that will not
|
||||
parse, an unknown `kind`, a vanished field, a relative path, more than eight
|
||||
conditions: each announces. The dangerous failure here is not a spurious nag but
|
||||
work that disappears because a predicate could not be evaluated, which would be
|
||||
indistinguishable from the ask being lost and would not surface until a release
|
||||
needed it. An ask with no `dormant_unless` behaves exactly as it did before the
|
||||
feature existed. `scripts/test-dormancy.sh` exercises all of it, positive and
|
||||
negative arms both.
|
||||
|
||||
Gated on `MEMPALACE_PI_DEVICE` **and** `MEMPALACE_REMOTE_URL` (the same pair as
|
||||
the stamper, since an unstamped client has no address to be reached at), and
|
||||
disabled outright with `MEMPALACE_MAILBOX=0`. Inert on a solitary palace: no
|
||||
disabled outright with `MEMPALACE_MAILBOX=0` (notifications alone with
|
||||
`MEMPALACE_MAILBOX_NOTIFY=0`). Inert on a solitary palace: no
|
||||
calls, no injection. Delivery is the mechanism; the *norms* — what a reply owes,
|
||||
and that only a terminal event closes a thread — remain normative in the
|
||||
*consumer* skill (`~/.agents/skills/mempalace/SKILL.md`). **This file documents
|
||||
@@ -382,6 +474,9 @@ one — the long-lived server is only killed when a request genuinely stalls.
|
||||
|
||||
- `MEMPALACE_MCP_TIMEOUT_MS` — tool-call/request timeout. Default `60000`.
|
||||
Kept short on purpose: a *query* taking this long is genuinely wedged.
|
||||
The one call that is not a query — the feed's `mempalace_mine` — passes its
|
||||
own deadline (`MEMPALACE_FEED_MINE_TIMEOUT_MS`) down to the transport per
|
||||
call, so this default does not apply to it (since 2026-09-18; see below).
|
||||
- `MEMPALACE_MCP_INIT_TIMEOUT_MS` — `initialize` + `tools/list` handshake
|
||||
timeout. Default `300000`. Deliberately generous: a genuine first
|
||||
cold-open over virtiofs can legitimately take minutes, and killing a
|
||||
@@ -424,6 +519,50 @@ the next tool call transparently respawns `mempalace-mcp` and retries.
|
||||
`mempalace-mcp` manually with raw JSON-RPC on stdin to read the
|
||||
server-side error — much faster than guessing.
|
||||
|
||||
### `feed (tick) failed: mine timed out after …ms`
|
||||
|
||||
**Nothing has been lost.** The deadline bounds only how long the extension
|
||||
*waits*; it cannot cancel the mine, which continues on the server. The
|
||||
transcript is already staged before the mine is invoked, and
|
||||
`mempalace mine --mode convos` dedups by `source_file` and is idempotent, so the
|
||||
work either completed after the deadline or is redone by the next tick.
|
||||
|
||||
**Do not "fix" it by retrying harder from the client.** The palace is a single
|
||||
writer; a blind retry is what turns one slow mine into a queue of them.
|
||||
|
||||
Before 2026-09 this message appeared many times per session, which made it look
|
||||
like a persistent failure. That was a real defect, now fixed: `lastFeedAt` was
|
||||
recorded only after a *successful* wait, so a timeout left the debounce clock
|
||||
stale and every following settled turn started another mine on top of the one
|
||||
still running. Two changes — recording the attempt before the wait, and raising
|
||||
the deadline to sit far above the slowest honest completion — mean a healthy
|
||||
fleet should now never see it.
|
||||
|
||||
If you *do* still see it, it is now informative rather than noise: a mine
|
||||
exceeded five minutes. Check palace size and whether another writer (a
|
||||
host-side feeder, a scheduled `mine`) is holding the write lock, rather than
|
||||
raising the timeout again.
|
||||
|
||||
### `feed (tick) failed: mempalace remote request 'tools/call' failed: timed out after 60000ms`
|
||||
|
||||
Same event, different deadline — and the same reassurance: **nothing has been
|
||||
lost**, the mine continues server-side.
|
||||
|
||||
This is what the previous message turned into after the 2026-09 change, and
|
||||
it exposed that the change was incomplete. The feed's five-minute deadline was
|
||||
only *raced* against the call; `callTool()` had no way to carry it, so the
|
||||
transport's generic per-request timeout (`MEMPALACE_MCP_TIMEOUT_MS`, 60 s)
|
||||
fired first on every honest 60 s+ mine, and the 300 s was unreachable. On the
|
||||
stdio transport this was worse than noise: a per-request timeout there kills the
|
||||
server child, so the mine really was aborted at 60 s.
|
||||
|
||||
Fixed 2026-09-18: `callTool(name, args, { timeoutMs })` passes a per-call
|
||||
deadline to both transports and the feed uses it for the mine.
|
||||
`scripts/test-mcp-call-timeout.sh` pins the contract (a plain call still honours
|
||||
the short default; the override is honoured and is itself a deadline). Seeing
|
||||
this message on a fixed build means a mine exceeded *five* minutes — treat it as
|
||||
the previous section says.
|
||||
|
||||
## The `Type.Unsafe` gotcha
|
||||
|
||||
Earlier versions of this extension registered every MCP tool with
|
||||
|
||||
+587
-40
@@ -67,7 +67,9 @@
|
||||
* child, so pi gets an error instead of hanging and later calls fail fast.
|
||||
* This is a per-REQUEST timeout, not a process-lifetime one — the
|
||||
* long-lived server is only killed when a request genuinely stalls.
|
||||
* - MEMPALACE_MCP_TIMEOUT_MS tool-call/request timeout (default 60000)
|
||||
* - MEMPALACE_MCP_TIMEOUT_MS tool-call/request timeout (default 60000);
|
||||
* the feed's mine carries its own, longer
|
||||
* deadline (MEMPALACE_FEED_MINE_TIMEOUT_MS)
|
||||
* - MEMPALACE_MCP_INIT_TIMEOUT_MS initialize+tools/list timeout (default 300000)
|
||||
* Set either to 0 to disable (legacy unbounded behavior).
|
||||
*
|
||||
@@ -90,6 +92,7 @@
|
||||
*/
|
||||
|
||||
import { type ChildProcessWithoutNullStreams, spawn } from "node:child_process";
|
||||
import { readFileSync, statSync } from "node:fs";
|
||||
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
|
||||
import { Type } from "typebox";
|
||||
|
||||
@@ -116,7 +119,13 @@ interface IMcpClient {
|
||||
readonly alive: boolean;
|
||||
onExit: (() => void) | null;
|
||||
start(): Promise<void>;
|
||||
callTool(name: string, args: Record<string, unknown>): Promise<any>;
|
||||
/**
|
||||
* `opts.timeoutMs` overrides the transport's generic per-request deadline
|
||||
* for THIS call only. Callers that knowingly invoke a long server-side job
|
||||
* (the feed's `mempalace_mine`) pass their own deadline here; everything
|
||||
* else keeps the short default, which is the wedged-query guard.
|
||||
*/
|
||||
callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any>;
|
||||
ensureAlive(): Promise<boolean>;
|
||||
stop(): void | Promise<void>;
|
||||
}
|
||||
@@ -149,6 +158,165 @@ const sleep = (ms: number): Promise<void> =>
|
||||
if (typeof t.unref === "function") t.unref();
|
||||
});
|
||||
|
||||
// ── Dormancy ──────────────────────────────────────────────────────────────────
|
||||
/**
|
||||
* Is a standing ask DORMANT — correct to leave open, wrong to announce?
|
||||
*
|
||||
* The problem this solves, measured rather than imagined. A first-boot
|
||||
* acceptance ask is planted deliberately unanswerable: it describes work that
|
||||
* becomes possible only when the device is next recreated, and it must STAY
|
||||
* owed until then, because closing it early to tidy the mailbox is exactly how
|
||||
* that work gets lost. But deriveOwed could only see "directed, open, not
|
||||
* answered, not withdrawn", so such an ask is announced at every session start
|
||||
* and every poll for as long as it is correctly waiting — three days running on
|
||||
* emb-7kj4vr4g (2026-09-28 → 2026-10-01), and three consecutive releases before
|
||||
* that. The ask was right; announcing it was wrong. The cost lands on the human
|
||||
* reading the window, who is the one reader that cannot filter it.
|
||||
*
|
||||
* So an ask may now declare, in its own metadata, the condition under which it
|
||||
* is merely waiting:
|
||||
*
|
||||
* "dormant_unless": [
|
||||
* { "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json",
|
||||
* "field": "release_tag", "baseline": "v1.9.4" },
|
||||
* { "kind": "file_mtime", "path": "/etc/hostname",
|
||||
* "baseline": "2026-09-22T18:12:49Z" }
|
||||
* ]
|
||||
*
|
||||
* It is dormant while EVERY condition still matches its baseline, and goes live
|
||||
* the moment ANY of them differs — which is the trigger those asks already
|
||||
* state in prose ("act when EITHER differs"), now in a form a machine can
|
||||
* check. Dormant asks are withheld from the ANNOUNCED owed set, never from the
|
||||
* mailbox: the wake-up injection still lists them once per session, so they
|
||||
* stay discoverable and cannot be silently dropped.
|
||||
*
|
||||
* THE SAFETY RULE, and the reason every branch below reports `unchanged: false`
|
||||
* on doubt: dormancy must be PROVEN, never assumed. A missing file, an
|
||||
* unreadable one, a baseline that will not parse, an unknown `kind`, a field
|
||||
* that has vanished — each is a reason to SHOW the ask, not to hide it. The
|
||||
* dangerous failure here is not a spurious nag; it is work that vanishes
|
||||
* because a predicate could not be evaluated, which would be indistinguishable
|
||||
* from the ask being lost and would not surface until a release needed it.
|
||||
*
|
||||
* No eval, no shell, no network: the predicate language is deliberately two
|
||||
* fixed checks against local files. A general expression evaluator sitting in
|
||||
* the path that decides whether work is VISIBLE is a far worse trade than a
|
||||
* slightly clumsy schema.
|
||||
*/
|
||||
export type DormancyVerdict = { dormant: boolean; reason: string };
|
||||
|
||||
/** Split of the derived owed-set: what to announce, and what is merely waiting. */
|
||||
export type OwedSplit = { owed: LogEvent[]; dormant: LogEvent[] };
|
||||
|
||||
// Bounds, so a malformed or hostile event cannot turn a mailbox read into real
|
||||
// work. Both are deliberately small: these predicates describe boot state, and
|
||||
// anything needing more than a handful of local checks is not a dormancy rule.
|
||||
const MAX_DORMANCY_CONDITIONS = 8;
|
||||
const MAX_DORMANCY_JSON_BYTES = 256 * 1024;
|
||||
|
||||
const dormancyCondition = (c: unknown): { unchanged: boolean; reason: string } => {
|
||||
if (typeof c !== "object" || c === null || Array.isArray(c))
|
||||
return { unchanged: false, reason: "condition is not an object" };
|
||||
const { kind, path, baseline, field } = c as Record<string, unknown>;
|
||||
// Absolute paths only: a relative one would resolve against whatever cwd the
|
||||
// pi process happens to hold, which is not a property of the device whose
|
||||
// state the baseline describes.
|
||||
if (typeof path !== "string" || !path.startsWith("/") || path.includes("\0"))
|
||||
return {
|
||||
unchanged: false,
|
||||
reason: `path must be an absolute string, got ${JSON.stringify(path)}`,
|
||||
};
|
||||
if (typeof baseline !== "string" && typeof baseline !== "number")
|
||||
return { unchanged: false, reason: "baseline must be a string or a number" };
|
||||
|
||||
if (kind === "file_mtime") {
|
||||
// Compared at WHOLE SECONDS in UTC, because the two routes that produce
|
||||
// these baselines disagree below that: a filesystem mtime carries
|
||||
// sub-second residue (measured: /etc/hostname at .773761009) that a
|
||||
// hand-written or reported ISO baseline never will. Comparing raw
|
||||
// milliseconds would make every such predicate permanently "changed",
|
||||
// i.e. would silently disable the feature while appearing to work.
|
||||
const expected = Date.parse(String(baseline));
|
||||
if (!Number.isFinite(expected))
|
||||
return {
|
||||
unchanged: false,
|
||||
reason: `baseline is not a parseable date: ${JSON.stringify(baseline)}`,
|
||||
};
|
||||
let mtimeMs: number;
|
||||
try {
|
||||
mtimeMs = statSync(path).mtimeMs;
|
||||
} catch {
|
||||
return { unchanged: false, reason: `cannot stat ${path}` };
|
||||
}
|
||||
return Math.floor(mtimeMs / 1000) === Math.floor(expected / 1000)
|
||||
? { unchanged: true, reason: `${path} mtime still at baseline` }
|
||||
: { unchanged: false, reason: `${path} mtime moved from baseline` };
|
||||
}
|
||||
|
||||
if (kind === "json_field") {
|
||||
if (typeof field !== "string" || field.length === 0)
|
||||
return { unchanged: false, reason: "json_field needs a non-empty field name" };
|
||||
let text: string;
|
||||
try {
|
||||
const buf = readFileSync(path);
|
||||
if (buf.byteLength > MAX_DORMANCY_JSON_BYTES)
|
||||
return {
|
||||
unchanged: false,
|
||||
reason: `${path} exceeds the ${MAX_DORMANCY_JSON_BYTES}-byte cap`,
|
||||
};
|
||||
text = buf.toString("utf8");
|
||||
} catch {
|
||||
return { unchanged: false, reason: `cannot read ${path}` };
|
||||
}
|
||||
let doc: unknown;
|
||||
try {
|
||||
doc = JSON.parse(text);
|
||||
} catch {
|
||||
return { unchanged: false, reason: `${path} is not valid JSON` };
|
||||
}
|
||||
let cur: unknown = doc;
|
||||
for (const seg of field.split(".")) {
|
||||
if (
|
||||
typeof cur !== "object" ||
|
||||
cur === null ||
|
||||
!Object.prototype.hasOwnProperty.call(cur, seg)
|
||||
)
|
||||
return { unchanged: false, reason: `${path} has no field ${field}` };
|
||||
cur = (cur as Record<string, unknown>)[seg];
|
||||
}
|
||||
if (cur !== null && typeof cur === "object")
|
||||
return { unchanged: false, reason: `field ${field} is not a scalar` };
|
||||
// Compared as strings on purpose: the baseline arrives as JSON metadata, so
|
||||
// a manifest holding 3 and a baseline saying "3" are the same fact.
|
||||
return String(cur) === String(baseline)
|
||||
? { unchanged: true, reason: `${field} still at baseline` }
|
||||
: { unchanged: false, reason: `${field} moved from baseline` };
|
||||
}
|
||||
|
||||
return { unchanged: false, reason: `unknown dormancy kind ${JSON.stringify(kind)}` };
|
||||
};
|
||||
|
||||
export function evaluateDormancy(
|
||||
metadata: Record<string, unknown> | null | undefined,
|
||||
): DormancyVerdict {
|
||||
const raw = metadata?.dormant_unless;
|
||||
// Every one of these returns NOT dormant, i.e. "announce it". An ask without
|
||||
// a predicate behaves exactly as it did before this feature existed.
|
||||
if (raw === undefined || raw === null) return { dormant: false, reason: "no dormant_unless" };
|
||||
if (!Array.isArray(raw)) return { dormant: false, reason: "dormant_unless is not an array" };
|
||||
if (raw.length === 0) return { dormant: false, reason: "dormant_unless is empty" };
|
||||
if (raw.length > MAX_DORMANCY_CONDITIONS)
|
||||
return {
|
||||
dormant: false,
|
||||
reason: `too many conditions (${raw.length} > ${MAX_DORMANCY_CONDITIONS})`,
|
||||
};
|
||||
for (const c of raw) {
|
||||
const v = dormancyCondition(c);
|
||||
if (!v.unchanged) return { dormant: false, reason: v.reason };
|
||||
}
|
||||
return { dormant: true, reason: `all ${raw.length} condition(s) still at baseline` };
|
||||
}
|
||||
|
||||
class StdioMcpClient implements IMcpClient {
|
||||
private proc: ChildProcessWithoutNullStreams | null = null;
|
||||
private nextId = 1;
|
||||
@@ -372,8 +540,12 @@ class StdioMcpClient implements IMcpClient {
|
||||
});
|
||||
}
|
||||
|
||||
async callTool(name: string, args: Record<string, unknown>): Promise<any> {
|
||||
return this.request("tools/call", { name, arguments: args });
|
||||
async callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any> {
|
||||
// The per-call override matters MORE here than for the HTTP client: on
|
||||
// timeout this transport kills the server child, so a generic deadline
|
||||
// that undercuts a long mine does not merely abandon the wait — it
|
||||
// aborts the mine.
|
||||
return this.request("tools/call", { name, arguments: args }, opts?.timeoutMs ?? this.requestTimeoutMs);
|
||||
}
|
||||
|
||||
/** SIGTERM then SIGKILL grace, for stall recovery. */
|
||||
@@ -421,6 +593,10 @@ class StdioMcpClient implements IMcpClient {
|
||||
// • per-request AbortController timeout honouring MEMPALACE_MCP_TIMEOUT_MS /
|
||||
// MEMPALACE_MCP_INIT_TIMEOUT_MS, mirroring StdioMcpClient's timeout ethos.
|
||||
// • alive / ensureAlive / onExit to satisfy IMcpClient.
|
||||
// • callTool() takes an optional per-call `{ timeoutMs }` (IMcpClient
|
||||
// contract, see there) so the feed's long-running mine is not cut off by
|
||||
// the generic per-request deadline. Not a protocol change; sync token
|
||||
// unchanged.
|
||||
//
|
||||
// NOTE: mempalace-mcp --transport http is a SESSIONLESS, stateless JSON-RPC
|
||||
// server (no Mcp-Session-Id, always application/json, Connection: close), so
|
||||
@@ -466,8 +642,8 @@ class RemoteMcpClient implements IMcpClient {
|
||||
this.healthy = true;
|
||||
}
|
||||
|
||||
async callTool(name: string, args: Record<string, unknown>): Promise<any> {
|
||||
return this.request("tools/call", { name, arguments: args });
|
||||
async callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any> {
|
||||
return this.request("tools/call", { name, arguments: args }, { timeoutMs: opts?.timeoutMs });
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -855,7 +1031,22 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
const feedWing = process.env.MEMPALACE_FEED_WING ?? "wing_conversations";
|
||||
const feedDebounceMs = num(process.env.MEMPALACE_FEED_DEBOUNCE_MS, 600_000);
|
||||
const feedPrepareTimeoutMs = num(process.env.MEMPALACE_FEED_PREPARE_TIMEOUT_MS, 120_000);
|
||||
const feedMineTimeoutMs = num(process.env.MEMPALACE_FEED_MINE_TIMEOUT_MS, 30_000);
|
||||
// 300_000, raised from 30_000 on 2026-09-10. The mine is the SLOWEST thing
|
||||
// this extension does — it offers every qualifying session transcript to a
|
||||
// SINGLE-WRITER palace, measured at 30–60s in normal operation and growing
|
||||
// with the corpus — yet it carried by far the TIGHTEST deadline: 4x tighter
|
||||
// than the prepare step that precedes it (120_000) and 10x tighter than the
|
||||
// init handshake (300_000), which is a fast call. All three were introduced
|
||||
// together in 29e660e (2026-08-12) and this one was never revisited, so the
|
||||
// deadline fired during entirely normal operation and the resulting message
|
||||
// read as an error when nothing had gone wrong.
|
||||
//
|
||||
// Matched to the init timeout because liveness is the ONLY legitimate job
|
||||
// left for this deadline: the Promise.race below abandons our WAIT, it cannot
|
||||
// cancel the server's work, so the deadline buys nothing except an escape
|
||||
// from a permanently hung call. It must therefore sit far above the slowest
|
||||
// honest completion, not near it.
|
||||
const feedMineTimeoutMs = num(process.env.MEMPALACE_FEED_MINE_TIMEOUT_MS, 300_000);
|
||||
let lastFeedAt = 0; // 0 => the first settled turn also acts as a catch-up
|
||||
let feedInFlight: Promise<void> | null = null;
|
||||
|
||||
@@ -909,17 +1100,46 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
try {
|
||||
const source = await prepareFeed(reason);
|
||||
if (!source) return;
|
||||
// Record the attempt HERE, before awaiting — not after a successful
|
||||
// wait. The race below abandons only our WAIT; the mine keeps running
|
||||
// server-side, and `mine --mode convos` dedups by source_file and is
|
||||
// idempotent, so a timeout is emphatically not a "did not happen".
|
||||
//
|
||||
// Leaving lastFeedAt stale on the timeout path defeated the debounce
|
||||
// guard in the agent_settled handler below (`Date.now() - lastFeedAt <
|
||||
// feedDebounceMs`): with lastFeedAt unchanged that guard passed on EVERY
|
||||
// settled turn, and because `run` had already settled, feedInFlight was
|
||||
// null too — so BOTH guards stood open. Each settled turn then launched
|
||||
// another mine while the previous one was still running: overlapping
|
||||
// writers queueing on a single-writer palace, each making the next one
|
||||
// slower and the next timeout likelier. That positive feedback loop,
|
||||
// not the tight deadline by itself, is why operators saw "mine timed out
|
||||
// after 30000ms" many times per session rather than at most once per
|
||||
// debounce window.
|
||||
lastFeedAt = Date.now();
|
||||
// The deadline is passed DOWN to the transport as well as raced
|
||||
// here. Until 2026-09-18 it was only raced: callTool() had no way
|
||||
// to carry it, so the transport's generic per-request timeout
|
||||
// (MEMPALACE_MCP_TIMEOUT_MS, 60 000) fired first on every honest
|
||||
// 60 s+ mine — "remote request 'tools/call' failed: timed out
|
||||
// after 60000ms" — and the 300 000 below was unreachable. The
|
||||
// race stays as the liveness guard for a transport whose timeout
|
||||
// is disabled (0).
|
||||
await Promise.race([
|
||||
client.callTool("mempalace_mine", {
|
||||
source,
|
||||
mode: "convos",
|
||||
wing: feedWing,
|
||||
// Internal call: it does not pass through the registered tool's
|
||||
// execute(), so it stamps itself. These ARE this harness's own
|
||||
// transcripts from this device, so the harness segment is the
|
||||
// agent (not `miner`) even though the tool is `mine`.
|
||||
agent: stampProvenance ? `${agentName}@${device}` : agentName,
|
||||
}),
|
||||
client.callTool(
|
||||
"mempalace_mine",
|
||||
{
|
||||
source,
|
||||
mode: "convos",
|
||||
wing: feedWing,
|
||||
// Internal call: it does not pass through the registered tool's
|
||||
// execute(), so it stamps itself. These ARE this harness's own
|
||||
// transcripts from this device, so the harness segment is the
|
||||
// agent (not `miner`) even though the tool is `mine`.
|
||||
agent: stampProvenance ? `${agentName}@${device}` : agentName,
|
||||
},
|
||||
{ timeoutMs: feedMineTimeoutMs },
|
||||
),
|
||||
new Promise((_resolve, reject) =>
|
||||
setTimeout(
|
||||
() => reject(new Error(`mine timed out after ${feedMineTimeoutMs}ms`)),
|
||||
@@ -927,7 +1147,6 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
),
|
||||
),
|
||||
]);
|
||||
lastFeedAt = Date.now();
|
||||
} catch (err) {
|
||||
process.stderr.write(
|
||||
`[mempalace ext] feed (${reason}) failed: ${(err as Error).message}\n`,
|
||||
@@ -1068,17 +1287,120 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
return Boolean(m.correlation_id && m.correlation_id === candidate.correlation_id);
|
||||
});
|
||||
|
||||
/** Directed asks addressed to this device with no terminal reply from it. */
|
||||
async function deriveOwed(): Promise<LogEvent[]> {
|
||||
if (!mailboxEnabled || !available) return [];
|
||||
/**
|
||||
* Has the ORIGINAL REQUESTER explicitly withdrawn this candidate?
|
||||
*
|
||||
* RFC 003 §3.3 clause 3 clears an ask only on "no event **of yours**", and the
|
||||
* skill states the consequence outright: "there is nothing anyone can do about
|
||||
* it from the other end". That asymmetry is deliberate and mostly right — owed
|
||||
* ness is a statement about the RECIPIENT's accountability, and a requester
|
||||
* must not be able to delete an obligation the recipient genuinely has.
|
||||
*
|
||||
* It is wrong in exactly one case: the requester withdrawing its OWN ask.
|
||||
*
|
||||
* MEASURED COST, 2026-09-08/09. pi@mbp-m1-2020 withdrew a v1.8.13 rollout ask
|
||||
* to pi@tor-ms22 (evt seq 119: terminal `superseded`, same correlation_id,
|
||||
* `metadata.closes` naming it, body "DO NOT SPEND A MINUTE ON v1.8.13") and
|
||||
* recorded the withdrawal as done. It had no effect: seq 119's `from_agent` is
|
||||
* mbp, so it could never satisfy a join that only looks at tor-ms22's own
|
||||
* events. tor-ms22's next wake-up — 8h later, on the first boot of the image
|
||||
* shipping this very derivation — still listed the ask as owed, 41h old, for a
|
||||
* release that had been superseded and never installed there. It had to spend a
|
||||
* write closing an ask nobody wanted answered. The asymmetry is also INVISIBLE
|
||||
* from the sender's side: mbp did everything a sender is told to do and got a
|
||||
* result indistinguishable from success.
|
||||
*
|
||||
* WHY AN EXPLICIT MARKER AND NOT "ANY TERMINAL EVENT FROM THE REQUESTER".
|
||||
* This file's standing rule is that every failure mode stays on the
|
||||
* noisy-but-visible side, because a spuriously resurfacing item is one turn of
|
||||
* human correction whereas a suppressed unanswered ask is silent and permanent.
|
||||
* Inferring release from any terminal event would breach that: a requester
|
||||
* appending `applied` for its own bookkeeping — on the correlation, addressed
|
||||
* to me, before I ever replied — would silently delete a real obligation.
|
||||
* So release must be STATED, not inferred. `withdraws` is the canonical key;
|
||||
* `closes` is honoured because it is already this fleet's de-facto marker (mbp
|
||||
* seq 119, emb seq 117/118 all use it), and either must name this exact ask —
|
||||
* its `correlation_id` or its event `id`. Prose does not count.
|
||||
*
|
||||
* WHY THIS DOES NOT BREAK THE FIXED POINT IN RFC 003 §9.2. §9.2 rejects letting
|
||||
* terminal directed events into the owed set, because then every closure mints a
|
||||
* fresh obligation and the loop never terminates. That argument is about
|
||||
* CANDIDATES. This adds a CLEARER. A withdrawal carries a terminal status, so it
|
||||
* can never be a candidate, and nothing new becomes owed. The asserting shape
|
||||
* (`open`) and the clearing shape (terminal) stay disjoint. A reader who reaches
|
||||
* for §9.2 to object is answering a different question.
|
||||
*/
|
||||
const isWithdrawn = (candidate: LogEvent, inbound: LogEvent[]): boolean => {
|
||||
const requester = candidate.from_agent;
|
||||
if (!requester) return false;
|
||||
// Names THIS ask specifically: its correlation or its id. A marker naming
|
||||
// something else, or carrying prose, is not a withdrawal of this ask.
|
||||
const names = (v: unknown): boolean =>
|
||||
typeof v === "string" &&
|
||||
v.length > 0 &&
|
||||
((Boolean(candidate.correlation_id) && v === candidate.correlation_id) ||
|
||||
(Boolean(candidate.id) && v === candidate.id));
|
||||
return inbound.some((e) => {
|
||||
// Only the party that ASKED may retract. A third party writing a terminal
|
||||
// event on a shared correlation must never clear my obligation.
|
||||
if (e.from_agent !== requester) return false;
|
||||
// Addressed to the party being released. `to_agent=<me>` also matches '*'
|
||||
// at the SQL level (RFC 003 §7.4), and a broadcast must not be able to
|
||||
// quietly empty every machine's mailbox at once.
|
||||
if (e.to_agent !== mailboxAddress) return false;
|
||||
if (!TERMINAL_STATUS.has((e.status ?? "").toLowerCase())) return false;
|
||||
// Same ordering guard, same reason as isAnswered: a withdrawal must not
|
||||
// retire an ask the requester sent LATER on the same correlation.
|
||||
if (!isStrictlyAfter(e, candidate)) return false;
|
||||
const ackOf = e.metadata?.ack_of;
|
||||
const joins =
|
||||
(typeof ackOf === "string" && Boolean(candidate.id) && ackOf === candidate.id) ||
|
||||
Boolean(e.correlation_id && e.correlation_id === candidate.correlation_id);
|
||||
if (!joins) return false;
|
||||
return names(e.metadata?.withdraws) || names(e.metadata?.closes);
|
||||
});
|
||||
};
|
||||
|
||||
/**
|
||||
* Directed asks addressed to this device with no terminal reply from it, and
|
||||
* not explicitly withdrawn by whoever sent them (see isWithdrawn).
|
||||
*/
|
||||
async function deriveOwed(): Promise<OwedSplit> {
|
||||
if (!mailboxEnabled || !available) return { owed: [], dormant: [] };
|
||||
try {
|
||||
const [candidatesRaw, mineRaw] = await Promise.all([
|
||||
const [candidatesRaw, mineRaw, inboundRaw] = await Promise.all([
|
||||
// Every cursor-less event_list call in this file passes `order` explicitly.
|
||||
// The server's default is not ours to lean on: mempalace <= 3.9.0 defaults
|
||||
// to `asc`, 3.10.0 flips a cursor-less listing to newest-first. Both
|
||||
// derive* joins want the NEWEST window (see the comment below), so say so
|
||||
// and the verdict stops depending on which server version answers.
|
||||
client.callTool("mempalace_event_list", {
|
||||
to_agent: mailboxAddress,
|
||||
status: "open",
|
||||
limit: 50,
|
||||
order: "desc",
|
||||
}),
|
||||
// `order: "desc"` is a fix, not a flourish. event_list defaults to `asc`
|
||||
// (append order), so this asked for the OLDEST 100 events this device
|
||||
// ever wrote — meaning that once a device passes 100 authored events its
|
||||
// most RECENT replies fall out of the join window and every ask it just
|
||||
// answered resurfaces as owed. Latent rather than theoretical: tor-ms22
|
||||
// was at ~20 authored events when this was written. The window has to be
|
||||
// anchored at the newest end for the same reason the join uses `hlc`.
|
||||
client.callTool("mempalace_event_list", {
|
||||
from_agent: mailboxAddress,
|
||||
limit: 100,
|
||||
order: "desc",
|
||||
}),
|
||||
// Inbound with NO status filter, deliberately: a withdrawal carries a
|
||||
// TERMINAL status and is therefore structurally invisible to the
|
||||
// `status: "open"` candidate query above — the same omission deriveClosed
|
||||
// is built on. No amount of polling the open set could ever see one.
|
||||
client.callTool("mempalace_event_list", {
|
||||
to_agent: mailboxAddress,
|
||||
limit: 100,
|
||||
order: "desc",
|
||||
}),
|
||||
client.callTool("mempalace_event_list", { from_agent: mailboxAddress, limit: 100 }),
|
||||
]);
|
||||
const candidates = parseEvents(candidatesRaw).filter((c) => {
|
||||
// `to_agent: <me>` ALSO matches `*` broadcasts, per the tool contract.
|
||||
@@ -1092,11 +1414,80 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
// You cannot owe yourself.
|
||||
return c.from_agent !== mailboxAddress;
|
||||
});
|
||||
if (candidates.length === 0) return [];
|
||||
if (candidates.length === 0) return { owed: [], dormant: [] };
|
||||
const mine = parseEvents(mineRaw);
|
||||
return candidates.filter((c) => !isAnswered(c, mine));
|
||||
const inbound = parseEvents(inboundRaw);
|
||||
const live = candidates.filter((c) => !isAnswered(c, mine) && !isWithdrawn(c, inbound));
|
||||
// PARTITION, not filter. A dormant ask is still owed in the protocol
|
||||
// sense — it has no terminal reply and must not be closed — so it stays in
|
||||
// the return value where the mailbox can still account for it, and only
|
||||
// the ANNOUNCING paths below treat the two halves differently.
|
||||
const owed: LogEvent[] = [];
|
||||
const dormant: LogEvent[] = [];
|
||||
for (const c of live) (evaluateDormancy(c.metadata).dormant ? dormant : owed).push(c);
|
||||
return { owed, dormant };
|
||||
} catch {
|
||||
// Fail silent and open: a mailbox read must never break a session.
|
||||
return { owed: [], dormant: [] };
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Terminal replies to asks THIS device sent. NOT work — news.
|
||||
*
|
||||
* Why this is a second query rather than a widened deriveOwed(): deriveOwed()
|
||||
* asks the log for `status: "open"`, and a reply that CLOSES an ask is by
|
||||
* definition not open, so it is structurally invisible to that filter — no
|
||||
* amount of polling or waiting could ever have surfaced it.
|
||||
*
|
||||
* Measured 2026-09-07: emb-7kj4vr4g closed a v1.8.13 rollout ask with a
|
||||
* task.reply at status=applied, and the operator reasonably expected to be
|
||||
* told. The mailbox stayed silent and was CORRECT to — "owed" means "you must
|
||||
* reply", and nothing was owed. But the single most useful thing a fleet can
|
||||
* tell a human is "the thing you asked for is done" (here: done, AND your
|
||||
* premise was wrong), and that was the one category the feature could not
|
||||
* report. Silence was right by its own definition and wrong by the user's.
|
||||
*/
|
||||
async function deriveClosed(): Promise<LogEvent[]> {
|
||||
if (!mailboxEnabled || !available) return [];
|
||||
try {
|
||||
// `order: "desc"` for the same reason as deriveOwed: without it a 3.9.0
|
||||
// server hands back the OLDEST window, so past 100 authored events this
|
||||
// device's newest asks (and past 50 inbound, the replies that close them)
|
||||
// fall outside the join. The selection below is order-independent
|
||||
// (isStrictlyAfter), so this only fixes WHICH events are in the window.
|
||||
const [mineRaw, inboundRaw] = await Promise.all([
|
||||
client.callTool("mempalace_event_list", { from_agent: mailboxAddress, limit: 100, order: "desc" }),
|
||||
// NO status filter, deliberately: that omission is the entire fix.
|
||||
client.callTool("mempalace_event_list", { to_agent: mailboxAddress, limit: 50, order: "desc" }),
|
||||
]);
|
||||
// Correlations this device opened as a DIRECTED ask. A broadcast owes
|
||||
// nobody a reply, so it cannot be closed by one either.
|
||||
const myAsks = new Map<string, LogEvent>();
|
||||
for (const e of parseEvents(mineRaw)) {
|
||||
if (!e.correlation_id || !e.to_agent) continue;
|
||||
if (e.to_agent === "*" || e.to_agent === mailboxAddress) continue;
|
||||
if (e.type !== "task.request") continue;
|
||||
myAsks.set(e.correlation_id, e);
|
||||
}
|
||||
if (myAsks.size === 0) return [];
|
||||
// Newest terminal reply per correlation only. A peer that appends
|
||||
// applied-then-superseded should cost one line, not a wall of them.
|
||||
const best = new Map<string, LogEvent>();
|
||||
for (const e of parseEvents(inboundRaw)) {
|
||||
if (e.from_agent === mailboxAddress) continue; // cannot inform myself
|
||||
const cid = e.correlation_id;
|
||||
if (!cid || !myAsks.has(cid)) continue;
|
||||
// NOTE the deliberate asymmetry with deriveOwed: a broadcast is excluded
|
||||
// there (it owes nobody) but allowed here, because a peer answering on MY
|
||||
// correlation_id is news to me regardless of how widely it was addressed.
|
||||
if (!TERMINAL_STATUS.has((e.status ?? "").toLowerCase())) continue;
|
||||
const prev = best.get(cid);
|
||||
if (!prev || isStrictlyAfter(e, prev)) best.set(cid, e);
|
||||
}
|
||||
return [...best.values()];
|
||||
} catch {
|
||||
// Same contract as deriveOwed: a mailbox read never breaks a session.
|
||||
return [];
|
||||
}
|
||||
}
|
||||
@@ -1133,11 +1524,41 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
})
|
||||
.join("\n");
|
||||
|
||||
const DORMANT_NOTE =
|
||||
"Each of these declares a `dormant_unless` predicate in its metadata that STILL " +
|
||||
"matches this device's current state, so the work it describes is not yet " +
|
||||
"possible. They are NOT closed and NOT answered — leave them open. They return to " +
|
||||
"the announced owed-set by themselves the moment a baseline stops matching. If a " +
|
||||
"predicate cannot be evaluated at all (missing file, unparseable baseline, unknown " +
|
||||
"kind) the ask is announced as owed instead, deliberately: dormancy has to be " +
|
||||
"proven, never assumed.";
|
||||
|
||||
const OWED_HOWTO =
|
||||
"To close one, append an event on the SAME correlation_id with a terminal status " +
|
||||
"(applied/superseded/failed/blocked). An ack alone does NOT clear it, and neither " +
|
||||
"does claimed/ready — those deliberately keep it visible as taken-but-unfinished.";
|
||||
|
||||
// Written for the ONE reader the rest of this text ignores: the human watching
|
||||
// the window. Measured 2026-08-26 on two devices — a delivery lands, the agent
|
||||
// is idle, nothing happens, and the operator asks "do I have to nudge you for
|
||||
// you to read this?". Yes, and it is by design (see the deliverAs comment
|
||||
// below): at agent_settled no inference is running, so the text is queued for
|
||||
// the next turn. The message explained how to CLOSE an ask but never when it
|
||||
// would be SEEN, so the person who needed that fact was the only one not told.
|
||||
// Cheapest possible fix, and it changes no behaviour: say so in the text.
|
||||
const QUEUED_NOTE =
|
||||
"DELIVERY NOTE — this is a QUEUED message, not an action: it arrived while this " +
|
||||
"agent was idle, so no model call was made and nothing woke it. It is read at the " +
|
||||
"start of the next turn. If you are the human watching and want it handled now, " +
|
||||
"send any message to start a turn; the agent is not ignoring the ask, it is not running.";
|
||||
|
||||
const CLOSED_NOTE =
|
||||
"NO ACTION IS OWED on these — they are replies to asks THIS device sent, shown " +
|
||||
"once each because 'the thing you asked for is done' is news you wanted and the " +
|
||||
"owed-set could never carry it. Read the reply before assuming your original ask " +
|
||||
"was right: a peer that did the work is the most likely party to have found your " +
|
||||
"premise wrong.";
|
||||
|
||||
// Mid-session mailbox. Tier 1 (the wake-up injection below) owns the first
|
||||
// look; this exists because arrivals are bursty and correlate with our own
|
||||
// activity — measured 2026-08-26, 11 of 22 events on this log landed inside a
|
||||
@@ -1149,7 +1570,68 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
// the owed set holds something not already surfaced. Re-announcing the same
|
||||
// ask every poll is precisely how a channel teaches its reader to ignore it,
|
||||
// which is the failure this whole feature exists to reverse.
|
||||
pi.on("agent_settled", async () => {
|
||||
// --- A: tell the human at poll time (RFC 003 §7.11) -----------------------
|
||||
//
|
||||
// B (the QUEUED_NOTE above) fixes the confusion of someone who IS looking at
|
||||
// the window. It does nothing for the case that actually loses an ask: nobody
|
||||
// is looking. The delivered text is queued for a turn that only a human starts,
|
||||
// so an ask can wait indefinitely on an idle session while its recipient makes
|
||||
// coffee. A notification at poll time is the only part of this feature that
|
||||
// reaches a person who is not watching.
|
||||
//
|
||||
// Two surfaces, deliberately split by risk:
|
||||
// default in-TUI ctx.ui.notify, exactly as session_start already does.
|
||||
// Zero new output channels; visible only if you are looking.
|
||||
// =desktop additionally emit a terminal-native notification, reusing
|
||||
// the detection the fleet's own notify.ts already proves in
|
||||
// this harness (Kitty OSC 99 / Windows toast / OSC 777). This
|
||||
// escapes the container without notify-send, DBus or any host
|
||||
// access: the escape sequence is interpreted by the terminal
|
||||
// emulator on the human's machine. Opt-in because writing raw
|
||||
// escapes to stdout is a behaviour change on shared machines,
|
||||
// not because it is unreliable.
|
||||
// =kitty =osc777 force one protocol, skipping detection entirely.
|
||||
// =0 silent, for anyone who wants the mailbox without pings.
|
||||
//
|
||||
// WHY FORCING EXISTS — measured on tor-ms22 2026-08-26, and it invalidates
|
||||
// detection in exactly the deployment this ships in. Terminal identity lives in
|
||||
// env vars set by the emulator (KITTY_WINDOW_ID, TERM_PROGRAM) and `docker exec`
|
||||
// does NOT forward them: inside the container pi sees only TERM=xterm-256color
|
||||
// no matter what is rendering it. So detection can never see Kitty from in here
|
||||
// and always falls through to OSC 777, which Kitty does not implement — on a
|
||||
// containerised client the notification would silently do nothing, the worst
|
||||
// possible failure for a feature whose whole job is to break a silence. Naming
|
||||
// the protocol ends the guessing.
|
||||
//
|
||||
// Fired only when something is actually due, i.e. the same condition as the
|
||||
// delivery itself: a notification that fires on an empty poll would teach its
|
||||
// reader to ignore it, which is the failure this whole feature exists to undo.
|
||||
const mailboxNotifyMode = (process.env.MEMPALACE_MAILBOX_NOTIFY ?? "").trim().toLowerCase();
|
||||
const mailboxNotifyEnabled = mailboxNotifyMode !== "0" && mailboxNotifyMode !== "off";
|
||||
const mailboxNotifyTerminal =
|
||||
mailboxNotifyMode === "desktop" ||
|
||||
mailboxNotifyMode === "kitty" ||
|
||||
mailboxNotifyMode === "osc777";
|
||||
|
||||
/** Terminal-native notification. Best-effort; never throws into a handler. */
|
||||
const notifyTerminal = (title: string, body: string): void => {
|
||||
try {
|
||||
const esc = (s: string): string => s.replace(/[\x00-\x1f\x7f;]/g, " ");
|
||||
if (
|
||||
mailboxNotifyMode === "kitty" ||
|
||||
(mailboxNotifyMode === "desktop" && !!process.env.KITTY_WINDOW_ID)
|
||||
) {
|
||||
process.stdout.write(`\x1b]99;i=1:d=0;${esc(title)}\x1b\\`);
|
||||
process.stdout.write(`\x1b]99;i=1:p=body;${esc(body)}\x1b\\`);
|
||||
return;
|
||||
}
|
||||
process.stdout.write(`\x1b]777;notify;${esc(title)};${esc(body)}\x07`);
|
||||
} catch {
|
||||
/* a notification is never worth an exception */
|
||||
}
|
||||
};
|
||||
|
||||
pi.on("agent_settled", async (_event, ctx) => {
|
||||
if (!mailboxEnabled || !available) return;
|
||||
if (!wokeUp) return; // the wake-up injection has not run yet; it does look #1
|
||||
if (Date.now() - lastMailboxPollAt < mailboxPollMs) return;
|
||||
@@ -1158,22 +1640,48 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
lastMailboxPollAt = Date.now();
|
||||
void (async () => {
|
||||
try {
|
||||
const owed = await deriveOwed();
|
||||
const [owedSplit, closed] = await Promise.all([deriveOwed(), deriveClosed()]);
|
||||
const { owed, dormant } = owedSplit;
|
||||
const now = Date.now();
|
||||
const due = owed.filter((e) => {
|
||||
if (!e.id) return false;
|
||||
const last = surfaced.get(e.id);
|
||||
return last === undefined || now - last >= mailboxResurfaceMs;
|
||||
});
|
||||
if (due.length === 0) return; // silence is the correct output here
|
||||
for (const e of due) if (e.id) surfaced.set(e.id, now);
|
||||
const unseen = (list: LogEvent[]): LogEvent[] =>
|
||||
list.filter((e) => {
|
||||
if (!e.id) return false;
|
||||
const last = surfaced.get(e.id);
|
||||
return last === undefined || now - last >= mailboxResurfaceMs;
|
||||
});
|
||||
const due = unseen(owed);
|
||||
// Closing replies are announced ONCE and never resurface: an ask you
|
||||
// already know is finished is not a nag, and re-announcing it hourly is
|
||||
// the train-the-reader-to-ignore-it failure this window exists to stop.
|
||||
const newsRaw = closed.filter((e) => e.id && !surfaced.has(e.id));
|
||||
if (due.length === 0 && newsRaw.length === 0) return; // silence is correct
|
||||
for (const e of [...due, ...newsRaw]) if (e.id) surfaced.set(e.id, now);
|
||||
const blocks: string[] = [];
|
||||
if (due.length > 0) {
|
||||
blocks.push(
|
||||
`${due.length} directed ask(s) addressed to "${mailboxAddress}" with no ` +
|
||||
`terminal reply from this device. ${OWED_HOWTO}\n\n${formatOwed(due)}`,
|
||||
);
|
||||
}
|
||||
if (newsRaw.length > 0) {
|
||||
blocks.push(
|
||||
`${newsRaw.length} repl(y|ies) CLOSING an ask this device sent. ` +
|
||||
`${CLOSED_NOTE}\n\n${formatOwed(newsRaw)}`,
|
||||
);
|
||||
}
|
||||
// A dormant ask NEVER opens this window on its own — the early return
|
||||
// above is what actually stops the nag. But once the window is open for
|
||||
// something else, one line of count is nearly free and keeps a waiting
|
||||
// ask from feeling lost between session starts.
|
||||
if (dormant.length > 0) {
|
||||
blocks.push(`${dormant.length} dormant ask(s), not listed here. ${DORMANT_NOTE}`);
|
||||
}
|
||||
pi.sendMessage(
|
||||
{
|
||||
customType: "mempalace-mailbox",
|
||||
content:
|
||||
`MemPalace logstream mailbox: ${due.length} directed ask(s) addressed to ` +
|
||||
`"${mailboxAddress}" with no terminal reply from this device. ${OWED_HOWTO}\n\n` +
|
||||
formatOwed(due),
|
||||
`MemPalace logstream mailbox.\n\n${QUEUED_NOTE}\n\n` +
|
||||
blocks.join("\n\n---\n\n"),
|
||||
display: true,
|
||||
},
|
||||
// "steer" is queue-on-arrival, NOT an interrupt: at agent_settled the
|
||||
@@ -1185,6 +1693,23 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
// than auto-delivery and is not what was approved.
|
||||
{ deliverAs: "steer" },
|
||||
);
|
||||
// AFTER the queueing call, so a notification can never be the only
|
||||
// thing that happened: the message is in the thread first, then we
|
||||
// point at it. Wording names the nudge explicitly, because "you have
|
||||
// mail" without "press a key" reproduces the exact confusion B fixes.
|
||||
if (mailboxNotifyEnabled) {
|
||||
const parts = [
|
||||
due.length > 0 ? `${due.length} ask(s) owed` : "",
|
||||
newsRaw.length > 0 ? `${newsRaw.length} closed` : "",
|
||||
].filter(Boolean);
|
||||
const summary = `${parts.join(" + ")} for ${mailboxAddress} — send any message to handle`;
|
||||
try {
|
||||
if (ctx?.hasUI) ctx.ui.notify(`MemPalace mailbox: ${summary}`, "info");
|
||||
} catch {
|
||||
/* ignore: UI may be gone by the time the poll resolves */
|
||||
}
|
||||
if (mailboxNotifyTerminal) notifyTerminal("MemPalace mailbox", summary);
|
||||
}
|
||||
} catch {
|
||||
/* best-effort delivery: never break a turn over a mailbox read */
|
||||
}
|
||||
@@ -1215,10 +1740,11 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
sections.push(`## mempalace_diary_read\n\n(error: ${(err as Error).message})`);
|
||||
}
|
||||
// Tier 1 mailbox: what this device owes a reply to, derived rather than read
|
||||
// off a filter. deriveOwed() swallows its own failures and returns [], so a
|
||||
// broken palace costs a missing section, never a broken wake-up.
|
||||
// off a filter. deriveOwed() swallows its own failures and returns an empty
|
||||
// split, so a broken palace costs a missing section, never a broken wake-up.
|
||||
if (mailboxEnabled) {
|
||||
const owed = await deriveOwed();
|
||||
const [owedSplit, closed] = await Promise.all([deriveOwed(), deriveClosed()]);
|
||||
const { owed, dormant } = owedSplit;
|
||||
if (owed.length > 0) {
|
||||
// Record what this injection showed, so the first mid-session poll does
|
||||
// not re-announce the identical list minutes later. Without this the
|
||||
@@ -1232,6 +1758,27 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
`this device. ${OWED_HOWTO}\n\n${formatOwed(owed)}`,
|
||||
);
|
||||
}
|
||||
// Once per session, WITH ids. The wake-up injection is the one place a
|
||||
// dormant ask should be visible: "what is this device still carrying?" is a
|
||||
// session-start question, and answering it here is what keeps the mid-session
|
||||
// silence honest rather than concealing. Deliberately NOT added to
|
||||
// `surfaced`: that map exists to suppress repeats of ANNOUNCED work, and a
|
||||
// dormant ask must stay announceable the instant its baseline moves.
|
||||
if (dormant.length > 0) {
|
||||
sections.push(
|
||||
`## logstream mailbox (${dormant.length} dormant — waiting, nothing owed yet)\n\n` +
|
||||
`${DORMANT_NOTE}\n\n${formatOwed(dormant)}`,
|
||||
);
|
||||
}
|
||||
const news = closed.filter((e) => e.id && !surfaced.has(e.id));
|
||||
if (news.length > 0) {
|
||||
const now = Date.now();
|
||||
for (const e of news) if (e.id) surfaced.set(e.id, now);
|
||||
sections.push(
|
||||
`## logstream mailbox (${news.length} closed — no action owed)\n\n` +
|
||||
`${CLOSED_NOTE}\n\n${formatOwed(news)}`,
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
if (sections.length === 0) return;
|
||||
|
||||
Executable
+156
@@ -0,0 +1,156 @@
|
||||
#!/usr/bin/env bash
|
||||
# test-dormancy.sh — exercise the mailbox dormancy predicate (dormant_unless).
|
||||
#
|
||||
# Why this exists as a script and not a note: `evaluateDormancy` decides whether
|
||||
# an ask is SHOWN to the agent, so a silent regression there does not look like a
|
||||
# bug — it looks like an empty mailbox. The positive arms matter as much as the
|
||||
# negative ones: a predicate that never fires makes the feature a no-op, and a
|
||||
# predicate that fires too eagerly hides real work. Both arms run here.
|
||||
#
|
||||
# It cannot import extensions/pi/mempalace.ts in place, because that file imports
|
||||
# `typebox`, which pi provides at RUNTIME and this repo has no node_modules for.
|
||||
# So it copies the file into a temp tree with the resolved deps symlinked beside
|
||||
# it. The copy is made BY this script on every run, so it cannot drift from the
|
||||
# source the way a vendored duplicate would.
|
||||
#
|
||||
# Fixtures are created here rather than read from the host: the first draft used
|
||||
# /etc/hostname and /etc/pi-devbox/build-manifest.json, which made the positive
|
||||
# arms pass only on a pi-devbox container and silently flip to "not dormant"
|
||||
# anywhere else — a device-dependent test that reports success by doing nothing.
|
||||
#
|
||||
# Exit codes, deliberately distinct:
|
||||
# 0 every case behaved as specified
|
||||
# 1 at least one case FAILED (a real regression)
|
||||
# 2 INCONCLUSIVE — deps could not be resolved, so nothing was proven
|
||||
# (never 0: a test that skips quietly is the failure mode it should catch)
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
# ── Args ──────────────────────────────────────────────────────────────────────
|
||||
if [[ "${1:-}" == "--help" || "${1:-}" == "-h" ]]; then
|
||||
sed -n '2,25p' "$0" | sed 's/^# \?//'
|
||||
exit 0
|
||||
fi
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
SRC="$REPO_ROOT/extensions/pi/mempalace.ts"
|
||||
[[ -r "$SRC" ]] || {
|
||||
echo "INCONCLUSIVE: cannot read $SRC" >&2
|
||||
exit 2
|
||||
}
|
||||
|
||||
# ── Resolve pi's runtime deps (discovered, never hardcoded) ───────────────────
|
||||
GLOBAL_ROOT="$(npm root -g 2>/dev/null || true)"
|
||||
PI_PKG=""
|
||||
for cand in \
|
||||
"${GLOBAL_ROOT:+$GLOBAL_ROOT/@earendil-works/pi-coding-agent}" \
|
||||
"/usr/lib/node_modules/@earendil-works/pi-coding-agent" \
|
||||
"/usr/local/lib/node_modules/@earendil-works/pi-coding-agent"; do
|
||||
[[ -n "$cand" && -d "$cand" ]] && {
|
||||
PI_PKG="$cand"
|
||||
break
|
||||
}
|
||||
done
|
||||
[[ -n "$PI_PKG" && -d "$PI_PKG/node_modules/typebox" ]] || {
|
||||
echo "INCONCLUSIVE: pi-coding-agent / typebox not resolvable (looked under 'npm root -g')." >&2
|
||||
echo " Nothing was proven. Install pi, or run this where pi is installed." >&2
|
||||
exit 2
|
||||
}
|
||||
|
||||
# Node must be able to strip types from a .ts entry point (Node >= 22.6).
|
||||
node -e 'process.exit(0)' 2>/dev/null || {
|
||||
echo "INCONCLUSIVE: no usable node" >&2
|
||||
exit 2
|
||||
}
|
||||
|
||||
# ── Build the throwaway tree ──────────────────────────────────────────────────
|
||||
WORK="$(mktemp -d)"
|
||||
trap 'rm -rf "$WORK"' EXIT
|
||||
mkdir -p "$WORK/extensions/pi" "$WORK/node_modules/@earendil-works"
|
||||
cp "$SRC" "$WORK/extensions/pi/mempalace.ts"
|
||||
ln -s "$PI_PKG/node_modules/typebox" "$WORK/node_modules/typebox"
|
||||
ln -s "$PI_PKG" "$WORK/node_modules/@earendil-works/pi-coding-agent"
|
||||
printf '{"type":"module"}\n' > "$WORK/package.json"
|
||||
|
||||
# Prove the copy is the source, so a PASS cannot be about a stale file.
|
||||
if ! cmp -s "$SRC" "$WORK/extensions/pi/mempalace.ts"; then
|
||||
echo "INCONCLUSIVE: the copy differs from the source" >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
# ── Fixtures, owned by this test ──────────────────────────────────────────────
|
||||
printf '{"release_tag":"v1.9.4","nested":{"deep":{"leaf":"found"}},"count":3}\n' > "$WORK/manifest.json"
|
||||
printf 'not json at all\n' > "$WORK/notjson.txt"
|
||||
: > "$WORK/stamp"
|
||||
|
||||
# ── The truth table ───────────────────────────────────────────────────────────
|
||||
cat > "$WORK/run.mjs" <<'EOF'
|
||||
import { statSync } from "node:fs";
|
||||
import { evaluateDormancy } from "./extensions/pi/mempalace.ts";
|
||||
|
||||
const W = process.env.WORK;
|
||||
const STAMP = `${W}/stamp`;
|
||||
const MAN = `${W}/manifest.json`;
|
||||
// Floor to whole seconds: that is the contract, and the residue below proves
|
||||
// why the implementation must do the same.
|
||||
const atBaseline = new Date(Math.floor(statSync(STAMP).mtimeMs / 1000) * 1000).toISOString();
|
||||
|
||||
const mt = (baseline, path = STAMP) => ({ kind: "file_mtime", path, baseline });
|
||||
const jf = (field, baseline, path = MAN) => ({ kind: "json_field", path, field, baseline });
|
||||
|
||||
const cases = [
|
||||
// No predicate => behave exactly as before the feature existed.
|
||||
["no metadata", undefined, false],
|
||||
["metadata without the key", { foo: "bar" }, false],
|
||||
["dormant_unless is a string", { dormant_unless: "release != x" }, false],
|
||||
["dormant_unless is empty", { dormant_unless: [] }, false],
|
||||
["more than 8 conditions", { dormant_unless: Array(9).fill(mt(atBaseline)) }, false],
|
||||
|
||||
// POSITIVE arms. If these ever read false the feature is inert.
|
||||
["file_mtime at baseline", { dormant_unless: [mt(atBaseline)] }, true],
|
||||
["json_field at baseline", { dormant_unless: [jf("release_tag", "v1.9.4")] }, true],
|
||||
["both at baseline", { dormant_unless: [jf("release_tag", "v1.9.4"), mt(atBaseline)] }, true],
|
||||
["dotted field path", { dormant_unless: [jf("nested.deep.leaf", "found")] }, true],
|
||||
["number vs string baseline", { dormant_unless: [jf("count", "3")] }, true],
|
||||
|
||||
// The trigger firing: ANY condition differing wakes the ask.
|
||||
["file_mtime moved", { dormant_unless: [mt("2020-01-01T00:00:00Z")] }, false],
|
||||
["json_field moved", { dormant_unless: [jf("release_tag", "v1.9.5")] }, false],
|
||||
["one same, one moved", { dormant_unless: [mt(atBaseline), jf("release_tag", "v1.9.5")] }, false],
|
||||
|
||||
// FAIL-VISIBLE arms: dormancy unproven => announce.
|
||||
["missing file", { dormant_unless: [mt(atBaseline, "/nonexistent/path")] }, false],
|
||||
["unparseable baseline", { dormant_unless: [mt("not-a-date")] }, false],
|
||||
["relative path", { dormant_unless: [{ kind: "file_mtime", path: "etc/hostname", baseline: atBaseline }] }, false],
|
||||
["unknown kind", { dormant_unless: [{ kind: "uptime_lt", path: "/etc/hostname", baseline: "1d" }] }, false],
|
||||
["missing json field", { dormant_unless: [jf("no_such_field", "x")] }, false],
|
||||
["file is not JSON", { dormant_unless: [jf("a", "b", `${W}/notjson.txt`)] }, false],
|
||||
["field is not scalar", { dormant_unless: [jf("nested", "x")] }, false],
|
||||
["condition is not an object", { dormant_unless: ["release_tag"] }, false],
|
||||
["baseline is an object", { dormant_unless: [{ kind: "file_mtime", path: STAMP, baseline: {} }] }, false],
|
||||
];
|
||||
|
||||
let pass = 0;
|
||||
const failed = [];
|
||||
for (const [name, meta, want] of cases) {
|
||||
const v = evaluateDormancy(meta);
|
||||
const ok = v.dormant === want;
|
||||
if (ok) pass++;
|
||||
else failed.push(name);
|
||||
console.log(
|
||||
` ${ok ? "PASS" : "FAIL"} dormant=${String(v.dormant).padEnd(5)} want=${String(want).padEnd(5)} ` +
|
||||
`${name.padEnd(28)} :: ${v.reason}`,
|
||||
);
|
||||
}
|
||||
console.log(`\n ${pass}/${cases.length} passed`);
|
||||
if (failed.length) console.log(` FAILED: ${failed.join(", ")}`);
|
||||
// Recorded because it is the reason file_mtime floors to seconds: a real mtime
|
||||
// carries sub-second residue that an ISO baseline does not.
|
||||
console.log(` (stamp mtimeMs=${statSync(STAMP).mtimeMs}, baseline=${atBaseline})`);
|
||||
process.exit(failed.length === 0 ? 0 : 1);
|
||||
EOF
|
||||
|
||||
cd "$WORK" || exit 2
|
||||
export WORK
|
||||
node "$WORK/run.mjs"
|
||||
exit $?
|
||||
Executable
+134
@@ -0,0 +1,134 @@
|
||||
#!/usr/bin/env bash
|
||||
# test-mcp-call-timeout.sh — the per-call deadline of RemoteMcpClient in
|
||||
# extensions/pi/mempalace.ts, and the override that the feed's mine relies on.
|
||||
#
|
||||
# WHY THIS EXISTS. The feed tick calls `mempalace_mine` through
|
||||
# `client.callTool()`. Its own deadline (MEMPALACE_FEED_MINE_TIMEOUT_MS,
|
||||
# 300 000) was raised in 2026-09 because the mine legitimately takes 30-60 s on
|
||||
# a shared single-writer palace — but callTool() carried no way to pass that
|
||||
# deadline down, so the transport's generic per-request timeout
|
||||
# (MEMPALACE_MCP_TIMEOUT_MS, 60 000) fired first, and operators saw
|
||||
# feed (tick) failed: mempalace remote request 'tools/call' failed: timed out after 60000ms
|
||||
# instead of the message the 2026-09 change had aimed at. The outer race was
|
||||
# unreachable in practice. This harness pins both halves of the contract:
|
||||
# 1. a plain callTool() still honours the short per-request timeout
|
||||
# (a query taking that long really is wedged — keep it short);
|
||||
# 2. callTool(name, args, { timeoutMs }) honours the override, so a long
|
||||
# server-side job can be given its own, longer, deadline.
|
||||
# Before the fix, (2) fails: the third argument was silently ignored.
|
||||
#
|
||||
# Like test-owed-withdrawal.sh, it runs the SHIPPED text: the class is cut out
|
||||
# of mempalace.ts by brace matching, type-stripped with node's own stripper,
|
||||
# and driven against a local JSON-RPC server that delays tools/call.
|
||||
set -euo pipefail
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
SRC="${1:-$REPO_ROOT/extensions/pi/mempalace.ts}"
|
||||
WORK="$(mktemp -d)"
|
||||
trap 'rm -rf "$WORK"' EXIT
|
||||
[ -r "$SRC" ] || { echo "FAIL: cannot read $SRC" >&2; exit 2; }
|
||||
|
||||
# ---------------------------------------------------------------- extractor ---
|
||||
cat >"$WORK/extract.mjs" <<'EXTRACT'
|
||||
import { readFileSync, writeFileSync } from "node:fs";
|
||||
import { stripTypeScriptTypes } from "node:module";
|
||||
const src = readFileSync(process.argv[2], "utf8");
|
||||
|
||||
/** Cut a top-level `const NAME = ...;` out of the source, verbatim. */
|
||||
function decl(name) {
|
||||
const start = src.indexOf(`const ${name} =`);
|
||||
if (start < 0) throw new Error(`declaration not found: ${name}`);
|
||||
return src.slice(start, span(start));
|
||||
}
|
||||
/** Cut `class NAME ... { ... }` out of the source, verbatim. */
|
||||
function klass(name) {
|
||||
const start = src.indexOf(`class ${name} `);
|
||||
if (start < 0) throw new Error(`class not found: ${name}`);
|
||||
return src.slice(start, span(start, true));
|
||||
}
|
||||
/** End offset of the statement starting at `start` (brace-matched). */
|
||||
function span(start, braceOnly = false) {
|
||||
let depth = 0, inStr = null, sawBrace = false;
|
||||
for (let i = start; i < src.length; i++) {
|
||||
const c = src[i], prev = src[i - 1];
|
||||
if (inStr) { if (c === inStr && prev !== "\\") inStr = null; continue; }
|
||||
if (c === '"' || c === "'" || c === "`") { inStr = c; continue; }
|
||||
if (c === "/" && src[i + 1] === "/") { i = src.indexOf("\n", i); if (i < 0) break; continue; }
|
||||
if (c === "/" && src[i + 1] === "*") { i = src.indexOf("*/", i) + 1; continue; }
|
||||
if (c === "{" || c === "(" || c === "[") { depth++; if (c === "{") sawBrace = true; }
|
||||
else if (c === "}" || c === ")" || c === "]") { depth--; if (braceOnly && sawBrace && depth === 0) return i + 1; }
|
||||
else if (!braceOnly && c === ";" && depth === 0) return i + 1;
|
||||
}
|
||||
throw new Error(`unterminated statement at ${start}`);
|
||||
}
|
||||
|
||||
const parts = [decl("num"), decl("REMOTE_PROTOCOL_VERSION"), decl("REMOTE_CLIENT_INFO"), klass("RemoteMcpClient")];
|
||||
for (const [i, p] of parts.entries()) if (p.length < 30) throw new Error(`extraction ${i} implausibly short: ${p}`);
|
||||
if (!parts[3].includes("tools/call")) throw new Error("RemoteMcpClient does not mention tools/call");
|
||||
const ts = parts.join("\n\n") + "\nexport { RemoteMcpClient };\n";
|
||||
writeFileSync(process.argv[3], stripTypeScriptTypes(ts, { mode: "strip" }));
|
||||
EXTRACT
|
||||
node --no-warnings "$WORK/extract.mjs" "$SRC" "$WORK/client.mjs" || exit 2
|
||||
node --check "$WORK/client.mjs" || { echo "FAIL: extracted client does not parse" >&2; exit 2; }
|
||||
echo "[extract] pulled RemoteMcpClient from $(basename "$SRC") ($(wc -c <"$WORK/client.mjs") bytes)"
|
||||
|
||||
# ------------------------------------------------------------------ harness ---
|
||||
cat >"$WORK/run.mjs" <<'RUN'
|
||||
import { createServer } from "node:http";
|
||||
import { RemoteMcpClient } from "./client.mjs";
|
||||
|
||||
// A sessionless JSON-RPC server like `mempalace-mcp --transport http`:
|
||||
// initialize / tools/list answer at once; tools/call sleeps SLOW_MS first.
|
||||
const SLOW_MS = 1500;
|
||||
const server = createServer((req, res) => {
|
||||
let body = "";
|
||||
req.on("data", (c) => (body += c));
|
||||
req.on("end", () => {
|
||||
const msg = JSON.parse(body);
|
||||
if (msg.id === undefined) { res.writeHead(202); res.end(); return; } // notification
|
||||
const reply = (result) => {
|
||||
res.writeHead(200, { "content-type": "application/json" });
|
||||
res.end(JSON.stringify({ jsonrpc: "2.0", id: msg.id, result }));
|
||||
};
|
||||
if (msg.method === "initialize") return reply({ protocolVersion: "2024-11-05", capabilities: {}, serverInfo: { name: "fake", version: "0" } });
|
||||
if (msg.method === "tools/list") return reply({ tools: [{ name: "slow", description: "", inputSchema: { type: "object" } }] });
|
||||
if (msg.method === "tools/call") return void setTimeout(() => reply({ content: [{ type: "text", text: "done" }] }), SLOW_MS);
|
||||
reply({});
|
||||
});
|
||||
});
|
||||
await new Promise((r) => server.listen(0, "127.0.0.1", r));
|
||||
const url = `http://127.0.0.1:${server.address().port}/mcp`;
|
||||
|
||||
let failures = 0;
|
||||
const check = (ok, label) => { console.log(`${ok ? "ok " : "FAIL"} ${label}`); if (!ok) failures++; };
|
||||
|
||||
// Short per-request deadline, deliberately below SLOW_MS.
|
||||
process.env.MEMPALACE_MCP_TIMEOUT_MS = "400";
|
||||
const client = new RemoteMcpClient(url);
|
||||
await client.start();
|
||||
check(client.alive === true, "start(): initialize + tools/list complete, client alive");
|
||||
|
||||
// 1. A plain call honours the short deadline — this is the wedged-query guard.
|
||||
let err = null;
|
||||
const t0 = Date.now();
|
||||
try { await client.callTool("slow", {}); } catch (e) { err = e; }
|
||||
check(err !== null && /timed out after 400ms/.test(err.message), `plain callTool() rejects at the per-request deadline (${err && err.message})`);
|
||||
check(Date.now() - t0 < SLOW_MS, "…and rejects BEFORE the server would have answered");
|
||||
check(client.alive === false, "a timeout marks the client unhealthy (documented side effect; ensureAlive() revives)");
|
||||
|
||||
// 2. A call with its own deadline outlives the generic one — the feed's mine.
|
||||
await client.ensureAlive();
|
||||
err = null;
|
||||
let result = null;
|
||||
try { result = await client.callTool("slow", {}, { timeoutMs: 5000 }); } catch (e) { err = e; }
|
||||
check(err === null && result && result.content?.[0]?.text === "done", `callTool(name, args, { timeoutMs: 5000 }) waits past the generic deadline and gets the result (${err ? err.message : "ok"})`);
|
||||
|
||||
// 3. The override is itself a deadline, not "no deadline".
|
||||
err = null;
|
||||
try { await client.callTool("slow", {}, { timeoutMs: 200 }); } catch (e) { err = e; }
|
||||
check(err !== null && /timed out after 200ms/.test(err.message), `callTool(..., { timeoutMs: 200 }) rejects at ITS deadline (${err && err.message})`);
|
||||
|
||||
server.close();
|
||||
console.log(failures === 0 ? "PASS: all assertions held" : `FAIL: ${failures} assertion(s) failed`);
|
||||
process.exit(failures === 0 ? 0 : 1);
|
||||
RUN
|
||||
cd "$WORK" && node --no-warnings run.mjs
|
||||
Executable
+312
@@ -0,0 +1,312 @@
|
||||
#!/usr/bin/env bash
|
||||
# test-owed-withdrawal.sh — rule tests for the owed-set derivation in
|
||||
# extensions/pi/mempalace.ts: isStrictlyAfter, isAnswered, isWithdrawn.
|
||||
#
|
||||
# WHY THIS EXISTS
|
||||
# isWithdrawn changes who is allowed to clear an obligation, which is the one
|
||||
# thing in this extension that can go wrong SILENTLY. A wrong rule here does
|
||||
# not throw and does not show up in a log — it makes a real unanswered ask
|
||||
# vanish from a mailbox forever. So the rules get pinned down before they ship.
|
||||
#
|
||||
# WHY IT EXTRACTS THE PREDICATES FROM THE SOURCE INSTEAD OF RESTATING THEM
|
||||
# These are closures inside createExtension(), so they cannot be imported. The
|
||||
# tempting shortcut is to paste a copy of the logic into the test — which tests
|
||||
# the copy. This repo has already paid for a divergent second copy (the
|
||||
# pi-extensions skill mirror, 9579 B behind for weeks; the shell-lint logic
|
||||
# extracted to one file for the same reason). So the harness cuts the actual
|
||||
# declaration text out of mempalace.ts by brace matching and runs THAT. Change
|
||||
# the source and this test follows; paraphrase the source and it cannot.
|
||||
#
|
||||
# FIXTURES ARE VERBATIM REAL EVENTS from the fleet log (project/pi-devbox), not
|
||||
# invented shapes — including the exact incident that motivated isWithdrawn:
|
||||
# mbp-m1-2020 withdrew a v1.8.13 rollout ask to tor-ms22 at seq 119 and the
|
||||
# withdrawal had no effect, so tor-ms22 was still being told it owed a reply 41h
|
||||
# later for a release it never installed.
|
||||
#
|
||||
# Usage: scripts/test-owed-withdrawal.sh [source.ts]
|
||||
# exit 0 = all rules behave 1 = a rule broke
|
||||
# exit 2 = source unreadable or does not parse 3 = the gate itself cannot run
|
||||
#
|
||||
# The optional argument exists so the suite can be pointed at a deliberately
|
||||
# MUTATED copy of the source to prove it is sensitive — a suite that has never
|
||||
# been observed to fail is not evidence. See the mutation check in the CHANGELOG
|
||||
# entry that introduced isWithdrawn.
|
||||
set -euo pipefail
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
SRC="${1:-$REPO_ROOT/extensions/pi/mempalace.ts}"
|
||||
WORK="$(mktemp -d)"
|
||||
trap 'rm -rf "$WORK"' EXIT
|
||||
|
||||
[ -r "$SRC" ] || { echo "FAIL: cannot read $SRC" >&2; exit 2; }
|
||||
|
||||
# A gate that cannot run must not pass — the standing rule in this repo.
|
||||
#
|
||||
# NOT `node --check`: it does not type-strip, so it rejects ANY TypeScript —
|
||||
# `const x: number = 1` included, not just the inline type-import at the top of
|
||||
# mempalace.ts. It passed on node 22.x and stopped passing on node 24.x
|
||||
# (measured on pi-devbox v1.9.1, node v24.21.0: this gate exited 2 on an
|
||||
# unmodified mempalace.ts, so all 17 assertions below refused to run). Strip
|
||||
# first, then syntax-check the emitted JS. `mode: "strip"` blanks type syntax
|
||||
# without moving anything, so byte offsets and line numbers survive and a
|
||||
# reported error line still points at the right line of the ORIGINAL .ts.
|
||||
#
|
||||
# The two failure modes are reported separately and on purpose. "Cannot strip"
|
||||
# is a fact about the toolchain; "does not parse" is a fact about the source.
|
||||
# Collapsing them is what made this very defect present itself as
|
||||
# "mempalace.ts does not parse" when mempalace.ts was fine.
|
||||
#
|
||||
# SC2016 is disabled deliberately: the single quotes are the point. What follows
|
||||
# is JavaScript, and `${process.version}` must reach node, not be expanded by the
|
||||
# shell first.
|
||||
# shellcheck disable=SC2016
|
||||
node --no-warnings -e '
|
||||
const { readFileSync, writeFileSync } = require("node:fs");
|
||||
const { stripTypeScriptTypes } = require("node:module");
|
||||
if (typeof stripTypeScriptTypes !== "function") {
|
||||
console.error(`FAIL: node ${process.version} cannot strip TypeScript ` +
|
||||
`(module.stripTypeScriptTypes needs >= 22.13) — the gate is unavailable, ` +
|
||||
`which is NOT a statement about the source`);
|
||||
process.exit(3);
|
||||
}
|
||||
let js;
|
||||
try {
|
||||
js = stripTypeScriptTypes(readFileSync(process.argv[1], "utf8"), { mode: "strip" });
|
||||
} catch (err) {
|
||||
// The stripper is itself a parser, so a syntax error lands HERE, not in the
|
||||
// --check below. This is a statement about the source: exit 2, not 3.
|
||||
console.error(`FAIL: ${process.argv[1]} does not parse: ${err.message}`);
|
||||
process.exit(2);
|
||||
}
|
||||
writeFileSync(process.argv[2], js);
|
||||
' "$SRC" "$WORK/stripped.mjs" || exit $?
|
||||
node --check "$WORK/stripped.mjs" \
|
||||
|| { echo "FAIL: $SRC does not parse (stripped output rejected)" >&2; exit 2; }
|
||||
|
||||
# ---------------------------------------------------------------- extractor ---
|
||||
cat >"$WORK/extract.mjs" <<'EXTRACT'
|
||||
import { readFileSync, writeFileSync } from "node:fs";
|
||||
|
||||
const src = readFileSync(process.argv[2], "utf8");
|
||||
|
||||
/**
|
||||
* Cut `const <name> = ...;` out of the source by matching delimiters, so the
|
||||
* test runs the shipped text. Returns the declaration verbatim.
|
||||
*/
|
||||
function decl(name) {
|
||||
const start = src.indexOf(`const ${name} =`);
|
||||
if (start < 0) throw new Error(`declaration not found in source: ${name}`);
|
||||
let depth = 0;
|
||||
let inStr = null;
|
||||
for (let i = start; i < src.length; i++) {
|
||||
const c = src[i];
|
||||
const prev = src[i - 1];
|
||||
if (inStr) {
|
||||
if (c === inStr && prev !== "\\") inStr = null;
|
||||
continue;
|
||||
}
|
||||
if (c === '"' || c === "'" || c === "`") { inStr = c; continue; }
|
||||
if (c === "/" && src[i + 1] === "/") { i = src.indexOf("\n", i); if (i < 0) break; continue; }
|
||||
if (c === "{" || c === "(" || c === "[") depth++;
|
||||
else if (c === "}" || c === ")" || c === "]") depth--;
|
||||
else if (c === ";" && depth === 0) return src.slice(start, i + 1);
|
||||
}
|
||||
throw new Error(`unterminated declaration: ${name}`);
|
||||
}
|
||||
|
||||
const parts = ["TERMINAL_STATUS", "isStrictlyAfter", "isAnswered", "isWithdrawn"].map(decl);
|
||||
|
||||
// Sanity: the extractor must have found real bodies, not empty matches. A
|
||||
// silently-empty extraction would make every assertion below pass vacuously.
|
||||
for (const [i, p] of parts.entries()) {
|
||||
if (p.length < 40) throw new Error(`extracted declaration ${i} is implausibly short: ${p}`);
|
||||
}
|
||||
if (!parts[3].includes("withdraws")) throw new Error("isWithdrawn does not mention its marker key");
|
||||
|
||||
writeFileSync(process.argv[3], parts.join("\n\n"));
|
||||
EXTRACT
|
||||
|
||||
node "$WORK/extract.mjs" "$SRC" "$WORK/extracted.ts"
|
||||
echo "[extract] pulled 4 declarations from $(basename "$SRC") ($(wc -c <"$WORK/extracted.ts") bytes)"
|
||||
|
||||
# ------------------------------------------------------------------ harness ---
|
||||
{
|
||||
cat <<'HEAD'
|
||||
type LogEvent = {
|
||||
id?: string;
|
||||
seq?: number;
|
||||
hlc?: string;
|
||||
type?: string;
|
||||
status?: string;
|
||||
from_agent?: string;
|
||||
to_agent?: string;
|
||||
correlation_id?: string | null;
|
||||
created_at?: string;
|
||||
body?: string;
|
||||
metadata?: Record<string, unknown> | null;
|
||||
};
|
||||
|
||||
const mailboxAddress = "pi@tor-ms22";
|
||||
|
||||
HEAD
|
||||
cat "$WORK/extracted.ts"
|
||||
cat <<'TAIL'
|
||||
|
||||
// ---- VERBATIM fixtures from the real fleet log, stream project/pi-devbox ----
|
||||
const REP = "rep_d344e349ba276d6fc11997cd552f6937";
|
||||
|
||||
/** seq 112 — mbp's v1.8.13 rollout ask to tor-ms22. The obligation. */
|
||||
const ask112: LogEvent = {
|
||||
id: "evt_20260907T125110_2bef6f9b07a8",
|
||||
seq: 112,
|
||||
hlc: `1788785470242-000000-${REP}`,
|
||||
type: "task.request",
|
||||
status: "open",
|
||||
from_agent: "pi@mbp-m1-2020",
|
||||
to_agent: "pi@tor-ms22",
|
||||
correlation_id: "v1813-client-rollout-tor-ms22",
|
||||
};
|
||||
|
||||
/** seq 119 — mbp's withdrawal of its OWN ask. Terminal, directed, marker present. */
|
||||
const withdraw119: LogEvent = {
|
||||
id: "evt_20260908T223949_8f5b9bec7c92",
|
||||
seq: 119,
|
||||
hlc: `1788907189286-000000-${REP}`,
|
||||
type: "task.reply",
|
||||
status: "superseded",
|
||||
from_agent: "pi@mbp-m1-2020",
|
||||
to_agent: "pi@tor-ms22",
|
||||
correlation_id: "v1813-client-rollout-tor-ms22",
|
||||
metadata: {
|
||||
closes: "v1813-client-rollout-tor-ms22",
|
||||
nothing_owed: "no reply required to this event",
|
||||
replacement: "v1814-client-rollout-tor-ms22",
|
||||
},
|
||||
};
|
||||
|
||||
/** seq 120 — the LIVE v1.8.14 ask. Must survive the v1813 withdrawal. */
|
||||
const ask120: LogEvent = {
|
||||
id: "evt_20260908T224118_808de43d6982",
|
||||
seq: 120,
|
||||
hlc: `1788907278075-000000-${REP}`,
|
||||
type: "task.request",
|
||||
status: "open",
|
||||
from_agent: "pi@mbp-m1-2020",
|
||||
to_agent: "pi@tor-ms22",
|
||||
correlation_id: "v1814-client-rollout-tor-ms22",
|
||||
};
|
||||
|
||||
/** seq 122 — tor-ms22's OWN terminal reply, which is what actually cleared 112. */
|
||||
const myReply122: LogEvent = {
|
||||
id: "evt_20260909T061851_cf9bd18800db",
|
||||
seq: 122,
|
||||
hlc: `1788934731146-000000-${REP}`,
|
||||
type: "task.reply",
|
||||
status: "superseded",
|
||||
from_agent: "pi@tor-ms22",
|
||||
to_agent: "pi@mbp-m1-2020",
|
||||
correlation_id: "v1813-client-rollout-tor-ms22",
|
||||
metadata: { closes: "v1813-client-rollout-tor-ms22 — from the RECIPIENT side, which is the only side that can" },
|
||||
};
|
||||
|
||||
const clone = (e: LogEvent, over: Partial<LogEvent>): LogEvent => ({ ...e, ...over });
|
||||
|
||||
let failed = 0;
|
||||
const check = (name: string, expected: boolean, actual: boolean): void => {
|
||||
const ok = expected === actual;
|
||||
if (!ok) failed++;
|
||||
console.log(`${ok ? " ok " : " FAIL"} ${name} (expected ${expected}, got ${actual})`);
|
||||
};
|
||||
|
||||
console.log("\nisWithdrawn — the requester retracting its own ask");
|
||||
// THE INCIDENT. Before this rule existed the answer was false and tor-ms22 was
|
||||
// told it owed a reply for a release that no longer existed.
|
||||
check("real seq 119 withdraws real seq 112", true, isWithdrawn(ask112, [withdraw119]));
|
||||
check("withdrawal does NOT touch the live v1814 ask", false, isWithdrawn(ask120, [withdraw119]));
|
||||
|
||||
console.log("\nisWithdrawn — the ways a mailbox must NOT be silently emptied");
|
||||
check(
|
||||
"no marker: a bare terminal event from the requester is not a withdrawal",
|
||||
false,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { metadata: { nothing_owed: "no reply required" } })]),
|
||||
);
|
||||
check(
|
||||
"marker naming a DIFFERENT thread does not clear this one",
|
||||
false,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { metadata: { closes: "v1814-client-rollout-tor-ms22" } })]),
|
||||
);
|
||||
check(
|
||||
"marker carrying PROSE does not count (tor-ms22's own seq 122 shape)",
|
||||
false,
|
||||
isWithdrawn(ask112, [
|
||||
clone(withdraw119, {
|
||||
metadata: { closes: "v1813-client-rollout-tor-ms22 — from the RECIPIENT side, which is the only side that can" },
|
||||
}),
|
||||
]),
|
||||
);
|
||||
check(
|
||||
"a THIRD PARTY cannot retract someone else's ask",
|
||||
false,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { from_agent: "pi@emb-7kj4vr4g" })]),
|
||||
);
|
||||
check(
|
||||
"a BROADCAST cannot empty every machine's mailbox at once",
|
||||
false,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { to_agent: "*" })]),
|
||||
);
|
||||
check(
|
||||
"a NON-TERMINAL status is not a withdrawal",
|
||||
false,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { status: "open" })]),
|
||||
);
|
||||
check(
|
||||
"an event addressed to a DIFFERENT device does not clear my ask",
|
||||
false,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { to_agent: "pi@emb-7kj4vr4g" })]),
|
||||
);
|
||||
check(
|
||||
"ORDERING: a withdrawal cannot retire an ask the requester sent LATER",
|
||||
false,
|
||||
isWithdrawn(clone(ask112, { seq: 121, hlc: `1788907300000-000000-${REP}` }), [withdraw119]),
|
||||
);
|
||||
|
||||
console.log("\nisWithdrawn — accepted marker spellings");
|
||||
check(
|
||||
"canonical `withdraws` naming the correlation",
|
||||
true,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { metadata: { withdraws: "v1813-client-rollout-tor-ms22" } })]),
|
||||
);
|
||||
check(
|
||||
"marker naming the ask's EVENT ID instead of its correlation",
|
||||
true,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { metadata: { withdraws: ask112.id } })]),
|
||||
);
|
||||
|
||||
console.log("\nisAnswered — regression guard, unchanged behaviour");
|
||||
check("my own terminal reply still closes my own ask", true, isAnswered(ask112, [myReply122]));
|
||||
check("my reply on one thread does not close another", false, isAnswered(ask120, [myReply122]));
|
||||
check(
|
||||
"ORDERING still load-bearing: my reply cannot pre-close a LATER ask",
|
||||
false,
|
||||
isAnswered(clone(ask112, { seq: 130, hlc: `1788999999999-000000-${REP}` }), [myReply122]),
|
||||
);
|
||||
|
||||
console.log("\nEND TO END — the derived owed set for tor-ms22 on 2026-09-09");
|
||||
// The behaviour change, stated as the derivation states it. `mine` deliberately
|
||||
// EXCLUDES seq 122 so this measures the new rule rather than the reply that
|
||||
// happened to be written first.
|
||||
const candidates = [ask112, ask120];
|
||||
const mine: LogEvent[] = [];
|
||||
const inbound = [withdraw119];
|
||||
const owed = candidates.filter((c) => !isAnswered(c, mine) && !isWithdrawn(c, inbound));
|
||||
const owedIds = owed.map((e) => e.correlation_id).join(", ");
|
||||
check("exactly one ask remains owed", true, owed.length === 1);
|
||||
check("and it is the v1814 one, not the withdrawn v1813", true, owedIds === "v1814-client-rollout-tor-ms22");
|
||||
console.log(` derived owed set: [${owedIds}]`);
|
||||
|
||||
console.log(failed === 0 ? "\n=== PASSED ===" : `\n=== FAILED: ${failed} ===`);
|
||||
process.exit(failed === 0 ? 0 : 1);
|
||||
TAIL
|
||||
} >"$WORK/harness.ts"
|
||||
|
||||
node --experimental-strip-types "$WORK/harness.ts"
|
||||
Executable
+88
@@ -0,0 +1,88 @@
|
||||
#!/usr/bin/env bash
|
||||
# test-rsync-ship-idempotency.sh — regression test for the ship-step bug fixed
|
||||
# 2026-08-27 (bin/mempalace-pi-session: rsync --update -> --checksum).
|
||||
#
|
||||
# THE BUG: the stage file's mtime is deliberately the SOURCE transcript's mtime
|
||||
# (os.utime() at :903, "preserve session mtime for dedup stability"), so a
|
||||
# re-export of a session that has not been appended to since the last ship
|
||||
# carries a mtime that is NOT newer than the receiver copy. rsync --update
|
||||
# skips a file whose mtime is not strictly newer than the destination's, so a
|
||||
# scrubbed re-export of a DORMANT session was silently dropped: content
|
||||
# differed, mtime did not, "success" was reported, and unscrubbed bytes stayed
|
||||
# in the palace host's inbox indefinitely. Measured on mbp-m1-2020 2026-08-27.
|
||||
#
|
||||
# THE ASSERTION mirrors that exactly and needs no ssh, no palace, no secret:
|
||||
# stage a file, ship it, rewrite the content while RESTORING the original
|
||||
# mtime (the same thing os.utime() does), ship again, assert the receiver's
|
||||
# sha256 changed. On the pre-fix flag (--update) this fails; with --checksum
|
||||
# (content comparison, mtime-independent) it passes.
|
||||
#
|
||||
# Deliberately does NOT invoke bin/mempalace-pi-session itself: that needs a
|
||||
# live ssh target, a palace, MEMPALACE_* env — disproportionate scaffolding for
|
||||
# what is, at its core, one rsync flag. Pulls the flag list out of the script
|
||||
# instead of hand-copying it, so a future change to the ship command either
|
||||
# updates this test's expectation or fails it loudly rather than drifting
|
||||
# silently out of sync with what actually ships.
|
||||
#
|
||||
# Usage: scripts/test-rsync-ship-idempotency.sh (exit 0 = fix still holds)
|
||||
set -euo pipefail
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
SHIP_SCRIPT="$REPO_ROOT/bin/mempalace-pi-session"
|
||||
SRC="$(mktemp -d)"; DST="$(mktemp -d)"
|
||||
trap 'rm -rf "$SRC" "$DST"' EXIT
|
||||
|
||||
EXPECT='rsync -a --checksum --no-owner --no-group'
|
||||
if ! grep -qF "$EXPECT" "$SHIP_SCRIPT"; then
|
||||
echo "FAIL: $SHIP_SCRIPT no longer ships with '$EXPECT'." >&2
|
||||
echo " Either the fix regressed, or the flags changed -- update this" >&2
|
||||
echo " test's EXPECT string to match, don't just delete the test." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
ship() {
|
||||
rsync -a --checksum --no-owner --no-group \
|
||||
--include='*.jsonl' --exclude='*' \
|
||||
"$SRC/" "$DST/"
|
||||
}
|
||||
|
||||
# Same BYTE LENGTH on purpose, mirroring mbp-m1-2020's own reasoning for why
|
||||
# dropping --update outright would not be enough: rsync's default quick check
|
||||
# (no --update, no --checksum) already transfers when SIZE differs, so a test
|
||||
# with mismatched lengths would pass by accident and prove nothing about
|
||||
# --checksum specifically. A real redaction whose placeholder happens to match
|
||||
# the secret's length is exactly the case that stays silently undetected unless
|
||||
# content itself, not size or mtime, is compared.
|
||||
UNSCRUBBED='live: MEMPALACE_REMOTE_TOKEN=AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA'
|
||||
SCRUBBED='live: MEMPALACE_REMOTE_TOKEN=<redacted:MEMPALACE_REMOTE_TOKEN---------->'
|
||||
if [[ ${#UNSCRUBBED} -ne ${#SCRUBBED} ]]; then
|
||||
echo "FAIL: test fixture bug -- the two fixture strings must be equal length (${#UNSCRUBBED} vs ${#SCRUBBED}); this test would not exercise --checksum at all otherwise." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# 1. Baseline: ship a session, receiver now matches sender.
|
||||
printf '%s\n' "$UNSCRUBBED" > "$SRC/session.jsonl"
|
||||
touch -d '2026-08-23T23:53:18' "$SRC/session.jsonl"
|
||||
ship
|
||||
before="$(sha256sum "$DST/session.jsonl" | cut -d' ' -f1)"
|
||||
|
||||
# 2. The bug's exact shape: rewrite content (same length!) as a scrubbed
|
||||
# re-export would, then restore the ORIGINAL mtime, as the exporter's
|
||||
# os.utime() does.
|
||||
printf '%s\n' "$SCRUBBED" > "$SRC/session.jsonl"
|
||||
touch -d '2026-08-23T23:53:18' "$SRC/session.jsonl"
|
||||
|
||||
ship
|
||||
after="$(sha256sum "$DST/session.jsonl" | cut -d' ' -f1)"
|
||||
|
||||
if [[ "$before" == "$after" ]]; then
|
||||
echo "FAIL: receiver sha256 unchanged after a content-only re-export at a held-constant mtime." >&2
|
||||
echo " This is the exact defect --checksum was added to fix." >&2
|
||||
exit 1
|
||||
fi
|
||||
if ! grep -q '<redacted:MEMPALACE_REMOTE_TOKEN' "$DST/session.jsonl"; then
|
||||
echo "FAIL: receiver does not contain the scrubbed content." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "PASS: ship step transfers changed content even when mtime is deliberately held constant (before=$before after=$after)"
|
||||
Reference in New Issue
Block a user