Compare commits
23 Commits
5b8d78f946
...
main
| Author | SHA1 | Date | |
|---|---|---|---|
| 975ab92943 | |||
| 2167a1b033 | |||
| 817b3a82b7 | |||
| dab989b068 | |||
| e68ee2071c | |||
| 309980b62c | |||
| e2b060a940 | |||
| e45f6b4301 | |||
| 21023e7aa0 | |||
| a361b71c40 | |||
| b2b50afcc1 | |||
| de59571966 | |||
| f0bffd1b93 | |||
| 836e35b320 | |||
| 3d47937d06 | |||
| ecc2a9c574 | |||
| e91766286e | |||
| a92c75d070 | |||
| 982b001001 | |||
| bfe9c5cd4f | |||
| d2764bf78e | |||
| d4d8bb6109 | |||
| e1cc7592a1 |
@@ -58,6 +58,28 @@ Who owns what:
|
||||
|
||||
---
|
||||
|
||||
## Documentation
|
||||
|
||||
Start here if you are deciding **what to put where**, or running MemPalace on more than one machine:
|
||||
|
||||
| Document | What it answers |
|
||||
|---|---|
|
||||
| [`docs/fleet-memory.md`](docs/fleet-memory.md) | **Start here.** The five stores MemPalace keeps and what each is for; what a *central* palace buys a fleet of machines; when to file a memory versus when to message another device. Diagrams, worked examples. |
|
||||
| [`docs/rfc-003-coordination-log.md`](docs/rfc-003-coordination-log.md) | The inter-agent / inter-device coordination log (`logstream`): storage, append and query semantics, delivery and latency, the trust model, and its landmines. |
|
||||
| [`docs/secret-hygiene.md`](docs/secret-hygiene.md) | Keeping credentials out of the palace — what the feeder scrubs before staging, why detection is name-anchored and never entropy-anchored, and the measured false-positive rate that made one tier report-only |
|
||||
| [`docs/rfc-001-global-palace.md`](docs/rfc-001-global-palace.md) | Why and how a palace is centralised, what should *not* be global, and the Phase-0 landmines (including: never run `mempalace sync` against a shared palace). |
|
||||
| [`docs/rfc-002-joiner.md`](docs/rfc-002-joiner.md) | Replaying a second palace into a shared primary, and what dedupes what. |
|
||||
| [`docs/phase-1-exposure-runbook.md`](docs/phase-1-exposure-runbook.md) | *(moved)* The exposure record for one deployment is now private operator data. The stub names the reusable mechanism it also carried — the `Host`/`Origin` pin above all — which is not yet published elsewhere. |
|
||||
| [`docs/backup-and-recovery.md`](docs/backup-and-recovery.md) | Why a palace needs its own backup procedure, and how to restore one. |
|
||||
| [`extensions/pi/README.md`](extensions/pi/README.md) | The pi-side client: provenance stamping at the edge, and the auto-delivered mailbox. |
|
||||
|
||||
Behaviour for the *agents* is normative in the `mempalace` skill (`SKILL.md` here,
|
||||
installed to `~/.agents/skills/mempalace/`), not in these documents — the split is
|
||||
deliberate: the skill says what an agent should do, the docs say what the machinery
|
||||
actually does.
|
||||
|
||||
---
|
||||
|
||||
## Why this exists
|
||||
|
||||
MemPalace is the agent memory layer. Its stock CLI has two gaps that bite on a machine running opencode with a docs-first palace policy:
|
||||
@@ -575,9 +597,12 @@ the palace host and asks the server to mine its own local copy. Requires
|
||||
`MEMPALACE_PI_REMOTE_PATH`, and `MEMPALACE_PI_DEVICE` in `--help` for the
|
||||
rest. `MEMPALACE_PI_DEVICE` also defaults `--agent` to `pi@<device>` so each
|
||||
drawer's `added_by` records which harness and which machine produced it —
|
||||
mempalace itself stores neither. Deploying that primary — newt, DNS, and why the
|
||||
auth is a shared bearer token rather than per-device proxy users — is
|
||||
[`docs/phase-1-exposure-runbook.md`](docs/phase-1-exposure-runbook.md).
|
||||
mempalace itself stores neither. Deploying a primary — tunnel, DNS, token
|
||||
custody — is deployment-specific operator data and lives in a private fleet
|
||||
repository; [`docs/phase-1-exposure-runbook.md`](docs/phase-1-exposure-runbook.md)
|
||||
records what moved and why. The *consequence* of choosing one shared fleet token
|
||||
over per-device proxy users is public, in
|
||||
[`docs/rfc-003-coordination-log.md`](docs/rfc-003-coordination-log.md) §6.
|
||||
|
||||
---
|
||||
|
||||
|
||||
+102
-2
@@ -534,11 +534,70 @@ fi
|
||||
# Also classifies each export as NEW/ALREADY FILED (by source_file lookup)
|
||||
# so --dry-run reports the real mine-set size. Classification is advisory;
|
||||
# `mempalace mine --mode convos` is still the authoritative dedup.
|
||||
# The redactor is a sibling module, imported by the heredoc below. Exported
|
||||
# rather than passed as argv so the argv unpack stays stable.
|
||||
#
|
||||
# ${BASH_SOURCE[0]} reports the path the script was INVOKED as, and the image
|
||||
# installs /usr/local/bin/mempalace-pi-session as a symlink into
|
||||
# /opt/mempalace-toolkit/bin. A naive dirname therefore yields /usr/local/bin,
|
||||
# where the redactor does not exist — and because the import is fail-closed, that
|
||||
# turns every feeder tick on every device into "refusing to stage". Measured: the
|
||||
# symlinked invocation exited FATAL while the direct one worked, i.e. it would
|
||||
# have stopped the whole fleet's memory feed at the next image bake. So chase the
|
||||
# symlink chain, and offer fallbacks rather than betting on one answer.
|
||||
#
|
||||
# readlink -f is avoided deliberately: it is GNU/newer-BSD only, and this script
|
||||
# also runs directly on macOS hosts.
|
||||
_mp_resolve() {
|
||||
local p="$1" target
|
||||
while [ -L "$p" ]; do
|
||||
target="$(readlink "$p")" || break
|
||||
case "$target" in
|
||||
/*) p="$target" ;;
|
||||
*) p="$(dirname "$p")/$target" ;;
|
||||
esac
|
||||
done
|
||||
printf '%s' "$p"
|
||||
}
|
||||
_mp_self="$(_mp_resolve "${BASH_SOURCE[0]}")"
|
||||
MEMPALACE_REDACT_DIR="$(cd "$(dirname "$_mp_self")" && pwd):$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd):/opt/mempalace-toolkit/bin"
|
||||
export MEMPALACE_REDACT_DIR
|
||||
export_count=$(python3 - "$PI_SESSIONS_DIR" "$STAGE" "$SESSION_ID" "$SINCE" "$MIN_MESSAGES" "$MIN_ASSISTANT_CHARS" "$MODE" <<'PY'
|
||||
import json, os, sqlite3, sys
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
|
||||
# ── Secret scrubbing before anything is staged ───────────────────────────────
|
||||
# This is the last point at which a transcript is a plain in-memory value: the
|
||||
# local mine reads the staged file, and the remote path rsyncs that same file
|
||||
# byte-for-byte, so scrubbing here covers BOTH transports with one hook.
|
||||
#
|
||||
# FAIL CLOSED. If the redactor cannot be imported, staging is refused rather
|
||||
# than done unscrubbed — a missing module means a broken install, and the whole
|
||||
# point of this step is that a secret must not reach a shared palace. Override
|
||||
# deliberately with MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 if you ever need to feed a
|
||||
# machine whose toolkit is older than this file.
|
||||
# Colon-separated candidates: symlink-resolved dir, invoked dir, then the image's
|
||||
# known install path. First one that has the module wins.
|
||||
for _cand in os.environ.get("MEMPALACE_REDACT_DIR", "").split(":"):
|
||||
if _cand and os.path.isdir(_cand):
|
||||
sys.path.insert(0, _cand)
|
||||
try:
|
||||
import mempalace_redact as _redact
|
||||
except Exception as _e: # noqa: BLE001 - any import failure is fatal by design
|
||||
if os.environ.get("MEMPALACE_FEED_ALLOW_UNSCRUBBED", "").strip() in {"1", "true", "yes"}:
|
||||
_redact = None
|
||||
print(" [WARN] secret scrubber unavailable, staging UNSCRUBBED by request "
|
||||
f"({_e})", file=sys.stderr)
|
||||
else:
|
||||
print(f" [FATAL] secret scrubber unavailable ({_e}); refusing to stage. "
|
||||
"Set MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 to override.", file=sys.stderr)
|
||||
raise SystemExit(3)
|
||||
|
||||
_known = _redact.env_secrets() if _redact else []
|
||||
_redactions = 0
|
||||
_reported = 0
|
||||
|
||||
sessions_dir, stage, session_filter, since, min_messages, min_assistant_chars, mode = sys.argv[1:8]
|
||||
min_messages = int(min_messages)
|
||||
min_assistant_chars = int(min_assistant_chars)
|
||||
@@ -811,6 +870,29 @@ for path in paths:
|
||||
)
|
||||
continue
|
||||
|
||||
# Scrub the parsed objects, not the serialized text: string VALUES get
|
||||
# rewritten while keys, ids and structure are left exactly as they are.
|
||||
# (Scanning raw JSONL instead would also match escape artifacts like the
|
||||
# "\\t" before a field name, inventing keys such as "tapiKey".)
|
||||
if _redact is not None:
|
||||
_found: list = []
|
||||
out_lines = [_redact.scrub_obj(obj, _known, _found)[0] for obj in out_lines]
|
||||
_hard = [x for x in _found if not x.rule.endswith("-reported")]
|
||||
_soft = [x for x in _found if x.rule.endswith("-reported")]
|
||||
_redactions += len(_hard)
|
||||
_reported += len(_soft)
|
||||
if _hard:
|
||||
_by = {}
|
||||
for x in _hard:
|
||||
# "T2:github-pat" — rule says what matched, tier says how much to
|
||||
# trust it, which is what the reader of this line actually needs.
|
||||
_k = f"T{x.tier}:{x.rule}"
|
||||
_by[_k] = _by.get(_k, 0) + 1
|
||||
print(f" [REDACTED] {path.name} "
|
||||
+ ", ".join(f"{k}={v}" for k, v in sorted(_by.items()))
|
||||
+ " fp=" + ",".join(sorted({x.fingerprint for x in _hard})),
|
||||
file=sys.stderr)
|
||||
|
||||
out_path = stage / f"pi_{session_uuid}.jsonl"
|
||||
with out_path.open("w", encoding="utf-8") as f:
|
||||
for obj in out_lines:
|
||||
@@ -835,6 +917,13 @@ for path in paths:
|
||||
|
||||
print(f"EXPORTED {exported}")
|
||||
print(f"ALREADY_FILED {-1 if already_filed is None else skipped_already_filed}")
|
||||
# Report the scrub outcome even when it is zero: "0 redactions" is a measurement,
|
||||
# whereas printing nothing is indistinguishable from a scrubber that never ran.
|
||||
if _redact is not None:
|
||||
print(f" [scrub] {_redactions} redaction(s) applied, "
|
||||
f"{_reported} name-anchored candidate(s) reported only "
|
||||
f"(set MEMPALACE_REDACT_STRICT=1 to redact those too)", file=sys.stderr)
|
||||
|
||||
if skipped_short:
|
||||
print(f"SKIPPED_SHORT {skipped_short}", file=sys.stderr)
|
||||
if skipped_quiet:
|
||||
@@ -885,15 +974,26 @@ fi
|
||||
|
||||
# ── Ship to the palace host (remote mode only) ───────────────────────
|
||||
# mempalace_mine expands its source path in the SERVER process, so in remote
|
||||
# mode the exports have to physically exist over there. rsync --update is the
|
||||
# mode the exports have to physically exist over there. --checksum is the
|
||||
# idempotent half; the mine is the other half.
|
||||
#
|
||||
# NOT --update: the stage file's mtime is deliberately the SOURCE transcript's
|
||||
# mtime (see the os.utime() in the exporter, "preserve session mtime for dedup
|
||||
# stability"), so a re-export of a session that has not been appended to since
|
||||
# the last ship carries an mtime that is NOT newer than the receiver copy. With
|
||||
# --update rsync then SKIPS it silently -- which is exactly the case that must
|
||||
# ship after a redactor change, because the content differs while the mtime does
|
||||
# not. Measured on mbp-m1-2020 2026-08-27: a scrubbed re-export of a dormant
|
||||
# session was skipped and unscrubbed bytes stayed in the palace host's inbox,
|
||||
# while every local signal reported a clean stage. --checksum compares content
|
||||
# and keeps the ship idempotent without trusting timestamps.
|
||||
MINE_SOURCE="$STAGE"
|
||||
if [[ "$MODE" == "remote" ]]; then
|
||||
ssh_cmd="ssh"
|
||||
[[ -n "$SSH_CONFIG" ]] && ssh_cmd="ssh -F $SSH_CONFIG"
|
||||
echo ""
|
||||
echo "Shipping stage to ${SSH_TARGET%/}/$DEVICE/ ..."
|
||||
if ! rsync -a --update --no-owner --no-group \
|
||||
if ! rsync -a --checksum --no-owner --no-group \
|
||||
-e "$ssh_cmd" \
|
||||
--include='*.jsonl' --exclude='*' \
|
||||
"$STAGE/" "${SSH_TARGET%/}/$DEVICE/"; then
|
||||
|
||||
@@ -0,0 +1,604 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Scrub secret-shaped values out of palace-bound text.
|
||||
|
||||
WHY THIS IS NAME-ANCHORED AND NOT ENTROPY-ANCHORED
|
||||
==================================================
|
||||
The obvious design — "redact long random-looking strings" — is actively wrong
|
||||
for this corpus, and the reason is worth stating before anyone tries to
|
||||
"improve" it. A palace is *full* of high-entropy strings that are its own
|
||||
primary keys:
|
||||
|
||||
drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b drawer id
|
||||
..._chunk_000000 chunk id
|
||||
evt_20260826T213924_397b1f710c2d event id
|
||||
rep_d344e349ba276d6fc11997cd552f6937 replica id
|
||||
1787780364792-000000-rep_d344e349ba276d6f… hybrid logical clock
|
||||
ecc2a9c574f4156403cf4e90aea8b0d088c75700 git commit sha
|
||||
|
||||
An entropy or "looks like base64/hex of length N" detector fires on every one
|
||||
of those, i.e. on most identifiers in the corpus, and the redaction damages the
|
||||
memory it was meant to protect. Worse, the damage is silent and unrecoverable:
|
||||
a mined transcript whose ids have been replaced with <redacted> is no longer
|
||||
traceable to anything.
|
||||
|
||||
So detection here is anchored on one of three things that carry *meaning*:
|
||||
|
||||
Tier 1 KNOWN VALUES — literal values taken from this process's own
|
||||
environment, for variables whose NAME says secret.
|
||||
Zero false positives by construction: the value
|
||||
*is* the secret. Catches any presentation (env
|
||||
dump, JSON, error message, URL, prose), which is
|
||||
exactly the incident this exists for.
|
||||
Tier 2 KNOWN SHAPES — vendor-prefixed credentials (ghp_…, glpat-…,
|
||||
xox[abprs]-…, AKIA…, sk-ant-…, JWTs, PEM private
|
||||
keys, credentials embedded in URLs, Authorization
|
||||
headers). Very low false-positive rate because the
|
||||
prefix is meaningful, not merely random.
|
||||
Tier 3 NAME=VALUE — an assignment whose *key* says secret
|
||||
(…TOKEN, …SECRET, …PASSWORD, …API_KEY, …).
|
||||
REPORT-ONLY BY DEFAULT. Measured on 52 MB of real
|
||||
fleet transcripts it fired 403 times, of which the
|
||||
large majority were ${VAR} interpolation in compose
|
||||
files, TypeScript identifiers, a type annotation
|
||||
(`credentials: Credentials`), an IPA attribute
|
||||
holding a date (krbPasswordExpiration) and terminal
|
||||
output following a "Password:" prompt. Rewriting
|
||||
those corrupts code and docs held as memory, to
|
||||
catch what Tier 1 already catches by value.
|
||||
Set MEMPALACE_REDACT_STRICT=1 to make it enforce.
|
||||
|
||||
SHAPE COLLISIONS, AND WHY TIER 1 IS THE LOAD-BEARING ONE
|
||||
--------------------------------------------------------
|
||||
A 40-hex Gitea PAT is byte-indistinguishable from a git commit sha, so no shape
|
||||
rule can separate them. Only two things can: the value being known (Tier 1,
|
||||
which redacts) or a key naming it (Tier 3, which by default only reports). So
|
||||
in the default configuration a sha-shaped PAT is caught if and only if it
|
||||
belongs to this machine. That is an accepted, documented gap — the alternative
|
||||
is redacting every commit sha in the palace.
|
||||
|
||||
WHICH TIER DID THAT? — mapping printed rule names back to tiers
|
||||
--------------------------------------------------------------
|
||||
The output names the *rule* (`github-pat`, `env-value`), because that says what
|
||||
matched. RULE_TIERS below maps rule -> tier, and `tier_of()` is what the feeder
|
||||
uses to print `T2:github-pat=6`, because the tier is what tells a reader how
|
||||
much to trust the hit. A self-test asserts every rule that can appear in a
|
||||
Finding has a tier, so adding a rule without classifying it fails the tests.
|
||||
|
||||
KNOWN FALSE NEGATIVES — stated, not hidden
|
||||
------------------------------------------
|
||||
This will not catch: a novel credential format pasted bare into prose with no
|
||||
name nearby; a secret from a machine whose env this process cannot see; a
|
||||
base64-of-a-secret; a secret split across lines. Tier 1 covers the local
|
||||
machine's own secrets completely, which is where the measured incidents came
|
||||
from — but "the scrubber ran" must never be read as "there are no secrets in
|
||||
here". `suspicions()` exists for exactly that reason: it reports high-entropy
|
||||
candidates it did NOT redact, as fingerprints rather than values, so the false
|
||||
negative rate can be *measured* over time instead of assumed to be zero.
|
||||
|
||||
Run `python3 mempalace_redact.py --self-test` for the test corpus, which
|
||||
includes every benign id shape above as a must-not-redact case.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import math
|
||||
import os
|
||||
import re
|
||||
from typing import Any, Callable, Iterable
|
||||
|
||||
__all__ = ["scrub", "scrub_obj", "suspicions", "Finding"]
|
||||
|
||||
PLACEHOLDER = "<redacted:{label}>"
|
||||
|
||||
# ── Tier 1: names whose VALUES are secrets ────────────────────────────────────
|
||||
# Anchored on the variable name, then applied as an exact literal match.
|
||||
_SECRET_NAME = re.compile(
|
||||
r"(TOKEN|SECRET|PASSWORD|PASSWD|PASSPHRASE|API_?KEY|APIKEY|ACCESS_?KEY"
|
||||
r"|PRIVATE_?KEY|CREDENTIAL|BEARER|SESSION_?KEY|COOKIE|_PAT|^PAT$)",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
# …unless the name says it holds a location or a knob rather than a secret.
|
||||
# SSH_KEY_PATH is the live example: it matches KEY, and its value is a path.
|
||||
_NAME_NOT_SECRET = re.compile(
|
||||
r"(_PATH|_FILE|_DIR|_NAME|_ID|_URL|_URI|_HOST|_PORT|_USER|_ENABLED"
|
||||
r"|_TIMEOUT|_MS|_SECS|_SECONDS|_COUNT|_LIMIT|_MODE)$",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
# Keys that contain a secret word but describe a POLICY, COUNT or TYPE rather
|
||||
# than holding a credential. Measured on real transcripts: krbPasswordExpiration
|
||||
# (an IPA attribute whose value is a date), observationTokens (a token count),
|
||||
# targetTokens, tokensOverTarget.
|
||||
_KEY_NOT_A_SECRET = re.compile(
|
||||
r"(expiration|expiry|_age|count|length|limit|usage|interval|policy|type|class"
|
||||
r"|provider|manager|error|target|overtarget|file|path|dir|name|_id|url|uri"
|
||||
r"|host|port|user|enabled|timeout|_ms|_secs)$",
|
||||
re.IGNORECASE,
|
||||
)
|
||||
# Left of the match: a declaration keyword or a member access means this is CODE
|
||||
# (`const token = x`, `obj.apiKey = y`), and the "value" is an expression.
|
||||
_CODE_CONTEXT = re.compile(
|
||||
r"(?:\b(?:const|let|var|def|func|fn|public|private|readonly|export|return|if|assert)\s+$)"
|
||||
r"|[.\w]\s*$"
|
||||
)
|
||||
_TRIVIAL_VALUES = {
|
||||
"", "0", "1", "true", "false", "none", "null", "nil", "unset", "changeme",
|
||||
"change-me", "xxx", "***", "redacted", "your-token-here", "your_token_here",
|
||||
"placeholder", "example", "dummy", "test", "secret",
|
||||
}
|
||||
|
||||
# ── Tier 2: shapes that mean credential because the PREFIX means credential ──
|
||||
_SHAPE_RULES: list[tuple[str, re.Pattern[str], Callable[[re.Match[str]], str]]] = [
|
||||
# ORDER MATTERS, and the self-test is what proved it. The two structural
|
||||
# rules (credentials inside a URL, Authorization header) run FIRST, before
|
||||
# the vendor-prefix rules: a vendor rule rewriting `https://ghp_xxx@host` to
|
||||
# `https://<redacted:github-token>@host` leaves a colon INSIDE the
|
||||
# placeholder, which the URL rule then reads as user:password and redacts a
|
||||
# second time, producing a nested placeholder. Ordering plus the (?!<redacted)
|
||||
# guards below make the result independent of pass order.
|
||||
("url-credentials",
|
||||
re.compile(r"(?P<pre>[a-zA-Z][a-zA-Z0-9+.\-]*://(?!<redacted)[^/\s:@]+:)(?P<pw>(?!<redacted)[^/\s@]{4,})(?P<at>@)"),
|
||||
lambda m: f"{m.group('pre')}{PLACEHOLDER.format(label='url-password')}{m.group('at')}"),
|
||||
("authorization-header",
|
||||
re.compile(r"(?P<pre>[Aa]uthorization:\s*(?:Bearer|Basic|token)\s+)(?P<val>(?!<redacted)[A-Za-z0-9._\-+/=]{8,})"),
|
||||
lambda m: f"{m.group('pre')}{PLACEHOLDER.format(label='auth-header')}"),
|
||||
("github-pat", re.compile(r"\bgh[pousr]_[A-Za-z0-9]{16,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="github-token")),
|
||||
("github-fine-grained", re.compile(r"\bgithub_pat_[A-Za-z0-9_]{22,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="github-token")),
|
||||
("gitlab-pat", re.compile(r"\bglpat-[A-Za-z0-9_\-]{16,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="gitlab-token")),
|
||||
("slack", re.compile(r"\bxox[abprs]-[A-Za-z0-9\-]{10,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="slack-token")),
|
||||
("openai-anthropic", re.compile(r"\bsk-(?:ant-)?[A-Za-z0-9_\-]{20,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="api-key")),
|
||||
("aws-access-key", re.compile(r"\b(?:AKIA|ASIA)[0-9A-Z]{16}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="aws-access-key-id")),
|
||||
("google-api-key", re.compile(r"\bAIza[0-9A-Za-z_\-]{35}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="google-api-key")),
|
||||
("huggingface", re.compile(r"\bhf_[A-Za-z0-9]{30,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="hf-token")),
|
||||
("npm", re.compile(r"\bnpm_[A-Za-z0-9]{30,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="npm-token")),
|
||||
("digitalocean", re.compile(r"\bdop_v1_[a-f0-9]{60,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="do-token")),
|
||||
# JWT: three base64url segments. Requires the "eyJ" header so a random
|
||||
# dotted string cannot trip it.
|
||||
("jwt", re.compile(r"\beyJ[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}\.[A-Za-z0-9_\-]{8,}\b"),
|
||||
lambda m: PLACEHOLDER.format(label="jwt")),
|
||||
# Whole PEM block, non-greedy so several in one file stay separate.
|
||||
("pem-private-key",
|
||||
re.compile(r"-----BEGIN [A-Z ]*PRIVATE KEY-----.*?-----END [A-Z ]*PRIVATE KEY-----", re.DOTALL),
|
||||
lambda m: PLACEHOLDER.format(label="private-key-block")),
|
||||
]
|
||||
|
||||
# ── Tier 3: assignment whose KEY says secret ─────────────────────────────────
|
||||
# The key prefix is LAZY AND OPTIONAL, not "one char then anything": a mandatory
|
||||
# leading character cannot match a key that BEGINS with the secret word, so
|
||||
# `"api_key"` was missed while `"client_secret"` passed — the secret word sits
|
||||
# mid-name in one and at position 0 in the other. Note plain `KEY` is absent from
|
||||
# the alternation on purpose: it is too generic (SSH_KEY_PATH), and tier 1 covers
|
||||
# the real ones by value anyway.
|
||||
# The optional quote after the key is what makes JSON work: `"api_key": "v"`
|
||||
# closes the key BEFORE the colon, so a pattern jumping straight from key to
|
||||
# separator silently misses every JSON payload — i.e. most of a transcript.
|
||||
_ASSIGNMENT = re.compile(
|
||||
r"""(?P<key>\b[A-Za-z0-9_.\-]*?
|
||||
(?:token|secret|password|passwd|passphrase|api[_-]?key|apikey
|
||||
|access[_-]?key|private[_-]?key|credential)
|
||||
[A-Za-z0-9_.\-]*)
|
||||
(?P<kq>["']?)
|
||||
(?P<sep>\s*(?:=>|[=:])\s*)
|
||||
(?P<q>["']?)
|
||||
(?P<val>[^\s"'<>,;)\]}]{12,})
|
||||
(?P=q)""",
|
||||
re.IGNORECASE | re.VERBOSE,
|
||||
)
|
||||
# A value that is obviously not a live secret: a reference, a path, a
|
||||
# placeholder, or something we already redacted.
|
||||
_NOT_A_SECRET_VALUE = re.compile(
|
||||
r"""^(?:
|
||||
\$\{[^}\s]*\}? # ${VAR} ${VAR:-default} ${VAR:?err}
|
||||
# trailing } optional: the value
|
||||
# charset stops before it, so the
|
||||
# captured text is "${VAR" — which
|
||||
# still means "a reference", and
|
||||
# missing that is what made a
|
||||
# compose file look like a secret
|
||||
| \$[A-Za-z_][A-Za-z0-9_]* # $VAR
|
||||
| %[A-Za-z_][A-Za-z0-9_]*%? # %VAR%
|
||||
| \{\{[^}\s]*\}?\}? # {{ template }}
|
||||
| <[^>]*> # <redacted:…> / <your-token>
|
||||
| (?:/|~/|\./|\.\./)[^\s]* # a path
|
||||
| [A-Za-z]:\\ # a windows path
|
||||
| \*+ | x+ | \.+ # ***, xxx
|
||||
| (?:true|false|none|null|nil|unset|changeme|placeholder|example)
|
||||
)$""",
|
||||
re.IGNORECASE | re.VERBOSE,
|
||||
)
|
||||
|
||||
|
||||
class Finding:
|
||||
"""One redaction, recorded without ever carrying the secret itself."""
|
||||
|
||||
__slots__ = ("rule", "label", "length", "fingerprint")
|
||||
|
||||
def __init__(self, rule: str, label: str, value: str) -> None:
|
||||
self.rule = rule
|
||||
self.label = label
|
||||
self.length = len(value)
|
||||
# Stable across sessions, so a recurring leak is recognisable as the
|
||||
# same one, while the value stays unrecoverable from the report.
|
||||
self.fingerprint = hashlib.sha256(value.encode("utf-8")).hexdigest()[:8]
|
||||
|
||||
@property
|
||||
def tier(self) -> int:
|
||||
"""Which detection tier produced this. See RULE_TIERS for the mapping."""
|
||||
return tier_of(self.rule)
|
||||
|
||||
def __repr__(self) -> str: # pragma: no cover - diagnostics only
|
||||
return (f"<Finding T{self.tier}:{self.rule}:{self.label} "
|
||||
f"len={self.length} fp={self.fingerprint}>")
|
||||
|
||||
|
||||
def env_secrets(environ: dict[str, str] | None = None) -> list[tuple[str, str]]:
|
||||
"""Tier 1 material: (name, value) pairs from the environment worth scrubbing.
|
||||
|
||||
Deliberately conservative: a secret-sounding NAME is not enough, the value
|
||||
must also be long enough to be a credential and must not look like a path
|
||||
or a knob. Longest values first so that if one secret contains another as a
|
||||
substring, the longer one is replaced before the shorter can fragment it.
|
||||
"""
|
||||
env = os.environ if environ is None else environ
|
||||
out: list[tuple[str, str]] = []
|
||||
for name, value in env.items():
|
||||
if not value or len(value) < 12:
|
||||
continue
|
||||
if not _SECRET_NAME.search(name) or _NAME_NOT_SECRET.search(name):
|
||||
continue
|
||||
if value.strip().lower() in _TRIVIAL_VALUES:
|
||||
continue
|
||||
if _NOT_A_SECRET_VALUE.match(value.strip()):
|
||||
continue
|
||||
if "/" in value and (value.startswith("/") or value.startswith("~")):
|
||||
continue # a path that happened to be named …_KEY
|
||||
if any(c.isspace() for c in value):
|
||||
continue # credentials do not contain whitespace; prose does
|
||||
out.append((name, value))
|
||||
out.sort(key=lambda kv: len(kv[1]), reverse=True)
|
||||
return out
|
||||
|
||||
|
||||
# Rule name -> tier. Printed output names the rule (WHAT matched); the tier says
|
||||
# HOW MUCH TO TRUST IT: 1 = near-certain (the value is known to be a secret),
|
||||
# 2 = strong (a vendor prefix means what it says), 3 = candidate, needs eyes.
|
||||
RULE_TIERS: dict[str, int] = {
|
||||
"env-value": 1,
|
||||
**{name: 2 for name, _pat, _lbl in _SHAPE_RULES},
|
||||
"named-assignment": 3,
|
||||
"named-assignment-reported": 3,
|
||||
# Tier 0 is not a detection: it is a measured NON-detection, reported so the
|
||||
# false-negative rate is a number instead of an assumption.
|
||||
"suspicion": 0,
|
||||
}
|
||||
|
||||
|
||||
def tier_of(rule: str) -> int:
|
||||
"""Tier for a rule name, or -1 if the rule was never classified."""
|
||||
return RULE_TIERS.get(rule, -1)
|
||||
|
||||
|
||||
def scrub(
|
||||
text: str,
|
||||
known: Iterable[tuple[str, str]] | None = None,
|
||||
findings: list[Finding] | None = None,
|
||||
strict: bool | None = None,
|
||||
) -> tuple[str, list[Finding]]:
|
||||
"""Redact secrets from one string. Returns (scrubbed, findings).
|
||||
|
||||
Idempotent: the replacement text matches none of the rules, so running this
|
||||
twice changes nothing and cannot nest placeholders.
|
||||
"""
|
||||
found: list[Finding] = findings if findings is not None else []
|
||||
if strict is None:
|
||||
strict = os.environ.get("MEMPALACE_REDACT_STRICT", "").strip() in {"1", "true", "yes"}
|
||||
if not text:
|
||||
return text, found
|
||||
|
||||
# Tier 1 first: an exact known value should be labelled with its variable
|
||||
# name (the most useful thing a reader can be told) rather than by shape.
|
||||
for name, value in (env_secrets() if known is None else known):
|
||||
if value and value in text:
|
||||
count = text.count(value)
|
||||
text = text.replace(value, PLACEHOLDER.format(label=name))
|
||||
for _ in range(count):
|
||||
found.append(Finding("env-value", name, value))
|
||||
|
||||
# Tier 2.
|
||||
for rule, pattern, repl in _SHAPE_RULES:
|
||||
def _sub(m: re.Match[str], _rule: str = rule) -> str:
|
||||
found.append(Finding(_rule, _rule, m.group(0)))
|
||||
return repl(m)
|
||||
text = pattern.sub(_sub, text)
|
||||
|
||||
# Tier 3 — REPORT-ONLY unless strict. See STRICT note in the module docstring:
|
||||
# measured on 52 MB of real fleet transcripts this rule produced 403 hits of
|
||||
# which the overwhelming majority were ${VAR} interpolation, TypeScript
|
||||
# identifiers, YAML env lists and documentation prose. Redacting those would
|
||||
# corrupt code and docs stored as memory, to catch secrets that tier 1
|
||||
# already catches by value. So by default a tier-3 hit is REPORTED (so the
|
||||
# miss can be measured and the rule calibrated) and NOT rewritten.
|
||||
def _sub_assignment(m: re.Match[str]) -> str:
|
||||
val, key = m.group("val"), m.group("key")
|
||||
pre = m.string[max(0, m.start() - 24):m.start()]
|
||||
if (_NOT_A_SECRET_VALUE.match(val)
|
||||
or val.strip().lower() in _TRIVIAL_VALUES
|
||||
or _KEY_NOT_A_SECRET.search(key)
|
||||
or _CODE_CONTEXT.search(pre)
|
||||
or val.isdigit()):
|
||||
return m.group(0)
|
||||
found.append(Finding("named-assignment" if strict else "named-assignment-reported",
|
||||
key, val))
|
||||
if not strict:
|
||||
return m.group(0)
|
||||
return f"{m.group('key')}{m.group('kq')}{m.group('sep')}{m.group('q')}" \
|
||||
f"{PLACEHOLDER.format(label=m.group('key'))}{m.group('q')}"
|
||||
|
||||
text = _ASSIGNMENT.sub(_sub_assignment, text)
|
||||
return text, found
|
||||
|
||||
|
||||
def scrub_obj(obj: Any, known: Iterable[tuple[str, str]] | None = None,
|
||||
findings: list[Finding] | None = None,
|
||||
strict: bool | None = None) -> tuple[Any, list[Finding]]:
|
||||
"""Recursively scrub every string VALUE in a JSON-ish structure.
|
||||
|
||||
Keys are left alone: a key named "token" is metadata, not a credential, and
|
||||
rewriting keys would corrupt the transcript schema.
|
||||
"""
|
||||
found: list[Finding] = findings if findings is not None else []
|
||||
known = list(env_secrets() if known is None else known)
|
||||
if isinstance(obj, str):
|
||||
return scrub(obj, known, found, strict)[0], found
|
||||
if isinstance(obj, dict):
|
||||
return {k: scrub_obj(v, known, found, strict)[0] for k, v in obj.items()}, found
|
||||
if isinstance(obj, list):
|
||||
return [scrub_obj(v, known, found, strict)[0] for v in obj], found
|
||||
return obj, found
|
||||
|
||||
|
||||
# ── Measuring what we MISS, without leaking it ───────────────────────────────
|
||||
_BENIGN_ID = re.compile(
|
||||
r"""^(?:
|
||||
drawer_[\w.\-]+ # drawer / chunk ids
|
||||
| (?:evt|art)_\d{8}T\d{6}_[0-9a-f]{12} # event / artifact ids
|
||||
| rep_(?:[0-9a-f]{12}|[0-9a-f]{32}) # replica ids
|
||||
| \d{13}-[0-9a-f]{6}-rep_[0-9a-f]+ # hybrid logical clocks
|
||||
| [0-9a-f]{7,12} # short git sha
|
||||
| [0-9a-f]{40} # git sha / sha1
|
||||
| [0-9a-f]{64} # sha256
|
||||
| [0-9a-f]{8}(?:-[0-9a-f]{4}){3}-[0-9a-f]{12} # uuid
|
||||
| sha256:[0-9a-f]{64}
|
||||
| \d+
|
||||
)$""",
|
||||
re.VERBOSE | re.IGNORECASE,
|
||||
)
|
||||
_CANDIDATE = re.compile(r"[A-Za-z0-9+/_\-=]{24,}")
|
||||
|
||||
|
||||
def _entropy(s: str) -> float:
|
||||
if not s:
|
||||
return 0.0
|
||||
counts: dict[str, int] = {}
|
||||
for ch in s:
|
||||
counts[ch] = counts.get(ch, 0) + 1
|
||||
n = len(s)
|
||||
return -sum((c / n) * math.log2(c / n) for c in counts.values())
|
||||
|
||||
|
||||
def suspicions(text: str, min_entropy: float = 3.6) -> list[Finding]:
|
||||
"""High-entropy strings this module did NOT redact, as fingerprints.
|
||||
|
||||
This is the false-negative meter. It deliberately does not redact anything:
|
||||
every one of these is far more likely to be one of the palace's own ids than
|
||||
a credential, and redacting them would be the damage described at the top of
|
||||
this file. Reporting them as (length, fingerprint) makes the miss rate
|
||||
observable — and a fingerprint that keeps recurring is worth a human look.
|
||||
"""
|
||||
out: list[Finding] = []
|
||||
for m in _CANDIDATE.finditer(text):
|
||||
val = m.group(0)
|
||||
if _BENIGN_ID.match(val) or val.startswith("<redacted:"):
|
||||
continue
|
||||
if _entropy(val) < min_entropy:
|
||||
continue
|
||||
out.append(Finding("suspicion", "high-entropy", val))
|
||||
return out
|
||||
|
||||
|
||||
# ── Self-test ────────────────────────────────────────────────────────────────
|
||||
def _self_test() -> int:
|
||||
fake_token = "pFrDBfakVKQB0SLN1hTzAoaRJeQ4J_go7xmJezojAwI" # shape-alike, not real
|
||||
env = {
|
||||
"MEMPALACE_REMOTE_TOKEN": fake_token,
|
||||
"GITEA_TOKEN": "0123456789abcdef0123456789abcdef01234567", # 40 hex: sha-shaped
|
||||
"SSH_KEY_PATH": "/Users/someone/.ssh", # name matches KEY, value is a path
|
||||
"MEMPALACE_MAILBOX_POLL_MS": "300000", # knob, not a secret
|
||||
"MEMPALACE_REMOTE_URL": "https://palace.example.com/mcp",
|
||||
"PASSWORD": "changeme", # trivial value
|
||||
"API_KEY": "${FROM_VAULT}", # a reference
|
||||
}
|
||||
known = env_secrets(env)
|
||||
|
||||
must_redact = [
|
||||
("env dump line", f"MEMPALACE_REMOTE_TOKEN={fake_token}"),
|
||||
("bare value in prose", f"the token is {fake_token} apparently"),
|
||||
("value inside json", f'{{"env": {{"MEMPALACE_REMOTE_TOKEN": "{fake_token}"}}}}'),
|
||||
("sha-shaped pat, name-anchored",
|
||||
"GITEA_TOKEN=0123456789abcdef0123456789abcdef01234567"),
|
||||
("github pat", "remote add origin https://ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345@git"),
|
||||
("gitlab pat", "GLPAT is glpat-AbCdEfGhIjKlMnOpQrSt"),
|
||||
("slack", "xoxb-1234567890-ABCdefGHIjklMNO"),
|
||||
("anthropic-ish", "sk-ant-api03-AbCdEfGhIjKlMnOpQrStUvWx"),
|
||||
("aws", "AKIAIOSFODNN7EXAMPLE"),
|
||||
("jwt", "eyJhbGciOiJIUzI1NiJ9.eyJzdWIiOiIxMjM0NTY3ODkwIn0.dozjgNryP4J3jVmNHl0w5N_XgL0n3I9PlFUP0THsR8U"),
|
||||
("pem", "-----BEGIN OPENSSH PRIVATE KEY-----\nb3BlbnNzaA\n-----END OPENSSH PRIVATE KEY-----"),
|
||||
("url creds", "git clone https://joakim:hunter2hunter2@git.example.com/x.git"),
|
||||
("auth header", "Authorization: Bearer abcdefghijklmnop"),
|
||||
("json api_key", '"api_key": "AbCdEfGhIjKlMnOpQrSt"'),
|
||||
("json nested quotes", '{"auth": {"client_secret": "AbCdEfGhIjKlMnOpQrStUv"}}'),
|
||||
("yaml style", " gitea_token: 0123456789abcdefghij"),
|
||||
("lowercase password kv", "db_password = s3cr3t-p4ssw0rd-xyz"),
|
||||
]
|
||||
must_not_redact = [
|
||||
("drawer id", "drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b"),
|
||||
("chunk id", "drawer_pi-devbox_landmines_ca8929fea82cf80e7770ea0d_chunk_000000"),
|
||||
("event id", "evt_20260826T213924_397b1f710c2d"),
|
||||
("replica id", "rep_d344e349ba276d6fc11997cd552f6937"),
|
||||
("hlc", "1787780364792-000000-rep_d344e349ba276d6fc11997cd552f6937"),
|
||||
("git sha", "ecc2a9c574f4156403cf4e90aea8b0d088c75700"),
|
||||
("short sha", "ecc2a9c"),
|
||||
("sha256", "a" * 64),
|
||||
("uuid session file", "pi_01a03fd8-0dc6-73c2-a9c3-6b62ccf0318f.jsonl"),
|
||||
("knob assignment", "MEMPALACE_MAILBOX_POLL_MS=300000"),
|
||||
("path named key", "SSH_KEY_PATH=/Users/someone/.ssh"),
|
||||
("empty token", "MEMPALACE_REMOTE_TOKEN="),
|
||||
("already redacted", "MEMPALACE_REMOTE_TOKEN=<redacted:MEMPALACE_REMOTE_TOKEN>"),
|
||||
("var reference", "MEMPALACE_REMOTE_TOKEN=${VAULT_TOKEN}"),
|
||||
("placeholder", "MEMPALACE_REMOTE_TOKEN=<your-token-here>"),
|
||||
("trivial", "PASSWORD=changeme"),
|
||||
("url no creds", "MEMPALACE_REMOTE_URL=https://palace.example.com/mcp"),
|
||||
("prose", "The mailbox is an obligation channel, not a news channel."),
|
||||
("base64 in a diff line", "+ data = 'QUJDREVGR0hJSktMTU5PUFFSU1RVVldYWVo='"),
|
||||
]
|
||||
|
||||
failures = 0
|
||||
# Tier 3 is report-only by default, so the corpus below is checked in STRICT
|
||||
# mode (where tier 3 rewrites) and the default mode is asserted separately.
|
||||
print("MUST REDACT (strict: tier 3 enforcing)")
|
||||
for name, sample in must_redact:
|
||||
out, found = scrub(sample, known, strict=True)
|
||||
ok = bool(found) and all(
|
||||
secret not in out for secret in (fake_token, "hunter2hunter2",
|
||||
"0123456789abcdef0123456789abcdef01234567")
|
||||
)
|
||||
failures += not ok
|
||||
rules = ",".join(sorted({f.rule for f in found})) or "NONE"
|
||||
print(f" {'pass' if ok else 'FAIL'} {name:32s} [{rules}]")
|
||||
print("MUST NOT REDACT (strict: the hard cases)")
|
||||
for name, sample in must_not_redact:
|
||||
out, found = scrub(sample, known, strict=True)
|
||||
ok = out == sample and not found
|
||||
failures += not ok
|
||||
detail = "" if ok else f" -> {out!r} {found}"
|
||||
print(f" {'pass' if ok else 'FAIL'} {name:32s}{detail}")
|
||||
|
||||
print("DEFAULT MODE (tier 3 reports, does not rewrite)")
|
||||
for name, sample, expect_rewrite in [
|
||||
("env value still redacted", f"MEMPALACE_REMOTE_TOKEN={fake_token}", True),
|
||||
("vendor shape still redacted", "ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345", True),
|
||||
("name-anchored only reported", "db_password = s3cr3t-p4ssw0rd-xyz", False),
|
||||
("compose interpolation untouched", "GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}", False),
|
||||
("typescript identifier untouched", "const tokens = countTokensForModel", False),
|
||||
("ipa attribute untouched", "--setattr=krbPasswordExpiration=20260529090000", False),
|
||||
]:
|
||||
out, found = scrub(sample, known)
|
||||
rewritten = out != sample
|
||||
ok = rewritten == expect_rewrite
|
||||
if name == "name-anchored only reported":
|
||||
ok = ok and any(f.rule == "named-assignment-reported" for f in found)
|
||||
if name.endswith("untouched"):
|
||||
ok = ok and not any(f.rule.startswith("named-assignment") for f in found)
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} {name:34s}"
|
||||
+ ("" if ok else f" -> rewritten={rewritten} {found}"))
|
||||
|
||||
print("TIER MAPPING")
|
||||
_rules = ({"env-value", "named-assignment", "named-assignment-reported", "suspicion"}
|
||||
| {n for n, _p, _l in _SHAPE_RULES})
|
||||
_unclassified = sorted(r for r in _rules if tier_of(r) < 0)
|
||||
failures += bool(_unclassified)
|
||||
print(f" {'pass' if not _unclassified else 'FAIL'} every rule maps to a tier"
|
||||
+ (f" -> unclassified: {_unclassified}" if _unclassified
|
||||
else f" ({len(_rules)} rules)"))
|
||||
for label, sample, want in [
|
||||
("env value is tier 1", f"MEMPALACE_REMOTE_TOKEN={fake_token}", 1),
|
||||
("vendor shape is tier 2", "ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345", 2),
|
||||
("name-anchored is tier 3", "db_password = s3cr3t-p4ssw0rd-xyz", 3),
|
||||
]:
|
||||
got = scrub(sample, known)[1]
|
||||
ok = bool(got) and all(f.tier == want for f in got)
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} {label:26s}"
|
||||
+ ("" if ok else f" -> {[(f.rule, f.tier) for f in got]}"))
|
||||
|
||||
print("PROPERTIES")
|
||||
once, _ = scrub(f"MEMPALACE_REMOTE_TOKEN={fake_token}", known, strict=True)
|
||||
twice, second = scrub(once, known, strict=True)
|
||||
ok = once == twice and not second
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} idempotent (no nested placeholders)")
|
||||
|
||||
# The compound case the corpus caught: assert on EXACT output, not merely
|
||||
# that the secret is gone — a nested placeholder also hides the secret, and
|
||||
# would have shipped unnoticed.
|
||||
compound, _ = scrub("https://ghp_AbCdEfGhIjKlMnOpQrStUvWxYz012345@git.example.com/x", known, strict=True)
|
||||
ok = compound == "https://<redacted:github-token>@git.example.com/x"
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} vendor placeholder not re-redacted by url rule"
|
||||
+ ("" if ok else f" -> {compound!r}"))
|
||||
|
||||
already, found_a = scrub("https://joakim:<redacted:url-password>@git.example.com", known, strict=True)
|
||||
ok = already == "https://joakim:<redacted:url-password>@git.example.com" and not found_a
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} already-redacted url left alone")
|
||||
|
||||
obj = {"role": "user", "content": [{"type": "text", "text": f"export TOK={fake_token}"}],
|
||||
"id": "evt_20260826T213924_397b1f710c2d"}
|
||||
scrubbed, found = scrub_obj(obj, known, strict=True)
|
||||
ok = (fake_token not in repr(scrubbed)
|
||||
and scrubbed["id"] == obj["id"]
|
||||
and scrubbed["role"] == "user"
|
||||
and len(found) >= 1)
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} nested structure scrubbed, ids and keys preserved")
|
||||
|
||||
ok = all(fake_token not in repr(f) for f in scrub(f"x {fake_token}", known, strict=True)[1])
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} findings never carry the secret value")
|
||||
|
||||
sus = suspicions("drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b "
|
||||
"rep_d344e349ba276d6fc11997cd552f6937 "
|
||||
"1787780364792-000000-rep_d344e349ba276d6fc11997cd552f6937 "
|
||||
"ecc2a9c574f4156403cf4e90aea8b0d088c75700")
|
||||
ok = not sus
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} suspicion meter ignores our own id shapes ({len(sus)} hits)")
|
||||
|
||||
sus2 = suspicions("blob=Zm9vYmFyYmF6cXV1eHF1dXhmb29iYXJiYXpxdXV4Zm9vYmFy")
|
||||
ok = len(sus2) == 1 and all("Zm9v" not in repr(s) for s in sus2)
|
||||
failures += not ok
|
||||
print(f" {'pass' if ok else 'FAIL'} suspicion meter flags unknown blobs as fingerprints only")
|
||||
|
||||
print(f"\n{'ALL PASS' if not failures else f'{failures} FAILED'}")
|
||||
return 1 if failures else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
import sys
|
||||
if "--self-test" in sys.argv:
|
||||
raise SystemExit(_self_test())
|
||||
# Filter mode: scrub stdin -> stdout, report to stderr. Useful for spot
|
||||
# checks and for scrubbing a file by hand without writing throwaway code.
|
||||
data = sys.stdin.read()
|
||||
out, found = scrub(data)
|
||||
sys.stdout.write(out)
|
||||
if found:
|
||||
by_rule: dict[str, int] = {}
|
||||
for f in found:
|
||||
by_rule[f.rule] = by_rule.get(f.rule, 0) + 1
|
||||
print(f"[redact] {len(found)} redaction(s): "
|
||||
+ ", ".join(f"{k}={v}" for k, v in sorted(by_rule.items())), file=sys.stderr)
|
||||
if "--suspicions" in sys.argv:
|
||||
for s in suspicions(out):
|
||||
print(f"[suspicion] len={s.length} fp={s.fingerprint}", file=sys.stderr)
|
||||
+2
-3
@@ -48,9 +48,8 @@ The pi variants are drop-in copies of the opencode variants with script name and
|
||||
|
||||
The odd one out in this directory: every other template *feeds* a palace on a schedule, this one
|
||||
**serves** a palace over HTTP so several machines can share it (RFC-001). **This unit currently runs
|
||||
the fleet primary on `synlig`** — serving since 2026-08-12, seeded 2026-08-14, and the palace behind
|
||||
it is the only copy. Read [`docs/rfc-001-global-palace.md`](../docs/rfc-001-global-palace.md) and
|
||||
[`docs/phase-1-exposure-runbook.md`](../docs/phase-1-exposure-runbook.md) before installing a second one.
|
||||
a fleet primary** — whose palace may be the only copy. Read [`docs/rfc-001-global-palace.md`](../docs/rfc-001-global-palace.md)
|
||||
before installing a second one.
|
||||
|
||||
It is a **user** unit (`systemctl --user`), so it dies with your login session unless lingering is
|
||||
enabled — that is the one `sudo` this recipe needs:
|
||||
|
||||
@@ -0,0 +1,218 @@
|
||||
# Fleet memory — what MemPalace stores, and what to put where
|
||||
|
||||
**Audience:** the person running MemPalace on more than one machine, or thinking about it.
|
||||
**Companion documents:** `docs/rfc-003-coordination-log.md` for the coordination log's mechanism, `docs/rfc-001-global-palace.md` for the centralisation design, and `~/.agents/skills/mempalace/SKILL.md` for what the *agents* are told to do.
|
||||
|
||||
MemPalace is usually described as "memory for agents", which is true and not very actionable. It is really **five stores with different retrieval models**, and most of the value — especially across a fleet — comes from putting each kind of thing in the store whose retrieval model matches how you will want it back.
|
||||
|
||||
This document covers: what the stores are, what a central palace changes when several machines share one, and how to decide between filing a memory and sending a message.
|
||||
|
||||
---
|
||||
|
||||
## 1. The stores at a glance
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
Q["ask by MEANING<br/>'what do we know about X?'"] --> D["<b>Drawers</b><br/>chroma.sqlite3"]
|
||||
R["ask by ENTITY + TIME<br/>'what was true in June?'"] --> G["<b>Knowledge graph</b><br/>knowledge_graph.sqlite3"]
|
||||
S["ask by ADDRESS<br/>'what is waiting for me?'"] --> L["<b>Coordination log</b><br/>logstream.sqlite3"]
|
||||
T["ask by ASSOCIATION<br/>'what else touches this?'"] --> P["<b>Palace graph</b><br/>tunnels.json + hallways.json"]
|
||||
D --> DI["<b>Diaries</b> live here too:<br/>drawers with room=diary"]
|
||||
```
|
||||
|
||||
All four stores are files inside **one palace directory**, so "the palace" is a directory you can back up in one go. Each holds a different *shape* of thing: drawers hold verbatim text, embedded for similarity; the knowledge graph holds typed facts that are time-bounded; the coordination log holds addressed events and their artifacts, read in append order; the palace graph holds links between rooms. Diaries are drawers, but they are read by recency rather than by similarity.
|
||||
|
||||
| Store | You get things back by | Typical use | Wrong use |
|
||||
|---|---|---|---|
|
||||
| **Drawers** (wings → rooms) | semantic similarity | a verbatim finding, a decision and its reasoning, a runbook, a transcript excerpt | anything a specific machine must *act* on; anything whose value is its exact byte content |
|
||||
| **Diaries** (drawers with `room="diary"`) | agent + recency | "what did I do last session, and what did it feel like" — the continuity thread across sessions | facts other agents need to find by searching; a diary is read by *its author*, chronologically |
|
||||
| **Knowledge graph** | entity, relationship, point in time | facts that *change*: versions, employers, who owns what, an injury that heals | prose, reasoning, anything you'd want to read rather than query |
|
||||
| **Coordination log** | address, correlation, append order | "device B must review this patch"; "this claim is retracted, stop building on it" | durable knowledge — an event is invisible to semantic search |
|
||||
| **Palace graph** | traversal from a room | discovering that an API design in one project touches a schema in another | primary storage — it links drawers, it does not hold content |
|
||||
|
||||
Two smaller files exist and are implementation detail, not user surface: `sqlite_exact.sqlite3` (an exact-match index over drawer metadata) and, if the daemon runs, `queue.sqlite3` (its job queue).
|
||||
|
||||
### 1.1 Drawers: wings and rooms
|
||||
|
||||
A **wing** is a project or domain; a **room** is an aspect within it. `wing="pi-devbox", room="landmines"` is a good pair; `wing="misc", room="stuff"` is how a palace becomes a landfill. Content is stored **verbatim and chunked** — never summarised — and retrieved by embedding similarity, so a drawer is found by someone who *doesn't already know it exists*. That is the property to optimise for: write the drawer that the next person's search will match.
|
||||
|
||||
The one counter-intuitive consequence: **fresh drawers rank worst.** A drawer filed an hour ago has no advantage in a similarity search, and a well-worn older drawer will outrank it. For anything recent, enumerate by date (`list_drawers(since=…)`) instead of searching.
|
||||
|
||||
### 1.2 Diaries
|
||||
|
||||
A diary entry is a drawer with `room="diary"`, filed by default into `wing_<agent>`, tagged with the writing agent. It is the *first-person* record: what I did, what surprised me, what I would do differently. Agents are told to write one before a session ends, and to read the last few at session start.
|
||||
|
||||
In a fleet this is the highest-signal store per byte, for a reason that is easy to miss: a diary entry is the only place that records **what did not work**. A drawer tends to record the conclusion; the diary records the three hours that produced it.
|
||||
|
||||
### 1.3 Knowledge graph
|
||||
|
||||
Triples — subject, predicate, object — with `valid_from` / `valid_to`, so a fact can *stop* being true without being deleted. `supersede` replaces a single-valued fact at a shared boundary, so a point-in-time query at that instant returns exactly one value.
|
||||
|
||||
Use it for anything you will later want to ask "what was true at time T?" about: which version was released when, who owned a service, what model an assistant was using. Do not use it for prose — a triple whose object is a paragraph is a drawer wearing a costume.
|
||||
|
||||
### 1.4 The coordination log
|
||||
|
||||
Addressed, ordered, exact. `mempalace_event_*` carries messages between agents; `mempalace_artifact_*` carries byte-exact payloads (patches, logs, files) that events can reference. This is the only store where one machine can *reach* another.
|
||||
|
||||
Its full mechanism, limits and landmines are in `docs/rfc-003-coordination-log.md`. §4 below covers what an operator needs to decide.
|
||||
|
||||
---
|
||||
|
||||
## 2. What changes when a fleet shares one palace
|
||||
|
||||
A single machine's palace is a notebook. A shared palace is something different in kind: **the fleet stops being a set of independent agents that each learn the same lessons separately.**
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
A["laptop<br/>agent session"] -->|MCP over HTTPS| H(("central palace<br/>one server"))
|
||||
B["workstation<br/>agent session"] -->|MCP over HTTPS| H
|
||||
C["build box<br/>agent session"] -->|MCP over HTTPS| H
|
||||
H --> D["drawers + diaries"]
|
||||
H --> G["knowledge graph"]
|
||||
H --> L["coordination log"]
|
||||
```
|
||||
|
||||
Note the topology: in the common deployment the machines are **thin clients of one server**, not peer replicas. Everything one machine writes is immediately visible to the others — there is no sync delay to reason about, and equally no local copy to fall back on when the server is unreachable. (`docs/rfc-001-global-palace.md` §4 designs an edge proxy with a local palace and a durable outbox for deployments that need to keep working offline; the plain thin-client shape above does not.)
|
||||
|
||||
### 2.1 The three things this actually buys
|
||||
|
||||
**Awareness.** "What has anyone been doing?" becomes answerable. Each machine's diary is readable by every other machine, so an agent starting work on a shared project can see that another machine spent yesterday on it, and how far it got.
|
||||
|
||||
**Non-repetition of expensive work.** This is the biggest measurable win. Anything that cost real time to obtain — a scraped API surface, a spec read end to end, a bisect, a benchmark, a long investigation into why a build fails on one platform — is filed once and searchable everywhere. The second machine's cost drops from hours to one search.
|
||||
|
||||
**Mistakes and retractions travel.** The subtle one, and the reason a shared palace is worth more than a shared wiki. When a machine discovers that a belief was *wrong*, it can file the retraction where every other machine will hit it. Without that, each machine independently rediscovers the same dead end — and worse, a machine can spend a day rebuilding something another machine already proved doesn't work.
|
||||
|
||||
```mermaid
|
||||
sequenceDiagram
|
||||
participant W as workstation
|
||||
participant P as central palace
|
||||
participant L as laptop, asleep 9 days
|
||||
W->>W: spends 3h finding why the build breaks
|
||||
W->>P: drawer — the finding, verbatim, with evidence
|
||||
W->>P: KG fact — "component X requires flag Y"
|
||||
W->>P: diary — what was tried and failed
|
||||
Note over L: ...9 days pass, laptop asleep...
|
||||
L->>P: wake-up: read diaries + search before starting
|
||||
P-->>L: the finding, the failed attempts, the fact
|
||||
Note over L: cost: one search instead of 3 hours
|
||||
```
|
||||
|
||||
### 2.2 The two costs, stated plainly
|
||||
|
||||
**Everything is visible to everyone.** One shared token, no per-agent read scoping. Anything filed into a shared palace should be considered readable by every machine and every agent on it. Do not put secrets in drawers.
|
||||
|
||||
**Provenance stops being obvious.** On a single machine, every drawer is yours and every path exists. On a shared palace, most drawers came from other machines and most `source_file` paths **do not exist locally**. Two consequences worth internalising:
|
||||
|
||||
- A file path in a drawer is evidence about *some* machine, not necessarily this one.
|
||||
- ⚠️ **Never run `mempalace sync` against a shared palace.** It prunes drawers whose source files look gitignored, deleted or moved — which on a shared palace describes most of the content, including every other machine's. See `docs/rfc-001-global-palace.md` §7.2. The coordination log is *not* affected by this (RFC 003 §2), but drawers very much are.
|
||||
|
||||
---
|
||||
|
||||
## 3. Deciding where something goes
|
||||
|
||||
The question that matters is not "is this important?" but **"how will I want this back, and does a specific machine or agent need to act?"**
|
||||
|
||||
```mermaid
|
||||
flowchart TD
|
||||
Start["I have something worth keeping"] --> Act{"Must someone specific<br/>DO something?"}
|
||||
Act -->|no| Change{"Is it a fact that<br/>changes over time?"}
|
||||
Act -->|yes| Know{"Do they also need<br/>to KNOW it later?"}
|
||||
Change -->|yes| KG["Knowledge graph<br/>kg_add / kg_supersede"]
|
||||
Change -->|no| Mine{"Is it about MY session<br/>— tried, felt, learned?"}
|
||||
Mine -->|yes| Diary["Diary entry"]
|
||||
Mine -->|no| Drawer["Drawer<br/>wing + room, verbatim"]
|
||||
Know -->|yes| Both["BOTH:<br/>drawer + event pointing at it"]
|
||||
Know -->|no| Event["Coordination event<br/>to_agent = that agent"]
|
||||
```
|
||||
|
||||
Worked examples:
|
||||
|
||||
| Situation | Where | Why |
|
||||
|---|---|---|
|
||||
| "The release takes 76 min, and 40 of those are the base image build." | Drawer | durable, nobody must act, next person finds it by searching "release timing" |
|
||||
| "v1.8.9 is the released version, as of this timestamp." | KG (`supersede`) | it will change; you will want "what was released in August?" |
|
||||
| "I spent two hours chasing a watcher that was already dead." | Diary | first-person, chronological, tells the next session what *not* to retry |
|
||||
| "Build box: this patch is ready, please review and apply." | Event (directed) | a named machine must act; the patch itself goes in as an artifact |
|
||||
| "The claim in that drawer is wrong — I measured the opposite." | Both | file the corrected finding as a drawer, then an event so the machine building on it stops |
|
||||
| "Everyone should know the new toolkit is live." | Drawer + broadcast event | the drawer is what anyone will *find*; the broadcast is a notice, not an ask (§4.2) |
|
||||
|
||||
The failure mode in each direction is worth naming, because both are common:
|
||||
|
||||
- **A finding filed only as an event** is invisible to semantic search. Nobody will ever find it again, and the next agent will re-derive it.
|
||||
- **An ask filed only as a drawer** is addressed to nobody. It will be found, if ever, by accident — long after it mattered.
|
||||
|
||||
---
|
||||
|
||||
## 4. What the coordination log can and cannot do
|
||||
|
||||
This is the part most likely to be mis-set expectations, so it is worth being blunt: **it is a durable log, not a chat.** Nothing is listening. Events are appended and persist; there is no delivery window; nothing is lost by being offline when one is written. A message waits indefinitely, and your reply waits just as patiently for a sender who has since gone away.
|
||||
|
||||
That sounds like a limitation and is actually the correct design for a fleet where few machines are awake at once and any given machine may sleep for weeks.
|
||||
|
||||
### 4.1 Latency, honestly
|
||||
|
||||
| Recipient state | When they see it | Notes |
|
||||
|---|---|---|
|
||||
| In a live session, mailbox-enabled client | **≈2–5 minutes** | measured ≈2–3 min on first live delivery; a poll floor of 5 min applies between checks |
|
||||
| Holding an SSE connection (`GET /logstream/stream`) | sub-second | for daemons/dashboards, not interactive agents |
|
||||
| Asleep, next session tomorrow | tomorrow | delivered in the session-start wake-up |
|
||||
| Asleep for three weeks | in three weeks | nothing expires; the log is permanent |
|
||||
| Never runs again | never | there is no re-routing and no dead-letter path |
|
||||
|
||||
So: appropriate for "handle this when you next wake", "here is a patch", "stop building on that claim". Not appropriate for anything with a deadline inside the hour, unless you know the recipient is awake.
|
||||
|
||||
One pleasant property, worth knowing because it is counter-intuitive: coordination traffic is **exempt from the palace's write lock**, so you can message another machine and it can reply *while* a long mine is running on the server (RFC 003 §4). Coordination stays alive when memory writes are blocked.
|
||||
|
||||
### 4.2 The one thing that does not work: broadcasting an ask
|
||||
|
||||
You can write a broadcast (`to_agent="*"`) and every machine that *lists* events will see it. But a broadcast **never enters any machine's mailbox** and is never auto-delivered — by design, because "everyone owes this answer" degenerates into either N duplicate replies or nobody acting.
|
||||
|
||||
- **Broadcast** = a notice on a wall. Fine for "v1.8.9 is out".
|
||||
- **Directed event** = a message in a named mailbox. Required for anything that must be done.
|
||||
|
||||
To reach a whole fleet with something actionable, **fan out**: one directed event per device, sharing one `correlation_id` so the thread stays joinable. Each machine then owes its own reply.
|
||||
|
||||
```mermaid
|
||||
flowchart LR
|
||||
You["you"] -->|"to_agent=pi@laptop"| A["laptop owes a reply"]
|
||||
You -->|"to_agent=pi@workstation"| B["workstation owes a reply"]
|
||||
You -->|"to_agent=pi@build-box"| C["build box owes a reply"]
|
||||
A --> Corr["one shared correlation_id<br/>joins the three threads"]
|
||||
B --> Corr
|
||||
C --> Corr
|
||||
```
|
||||
|
||||
### 4.3 Addressing, and why the format matters
|
||||
|
||||
Addresses are `<harness>@<device>` — `pi@laptop`, `opencode@build-box`. Two rules follow:
|
||||
|
||||
1. **An unstamped client is unreachable.** If a machine's events say `from_agent: pi` with no device, nobody can address it, because "pi" is every machine.
|
||||
2. **Two machines must never share one address.** Nothing prevents it, nothing warns, and the result is that each silently discards the other's asks as its own (RFC 003 §7.7).
|
||||
|
||||
### 4.4 A reply is owed until it is *terminally* closed
|
||||
|
||||
An acknowledgement does not close a thread. Neither does "claimed" or "ready". Only a terminal status — `applied`, `superseded`, `failed`, `blocked` — clears an ask from the recipient's mailbox. Until then, a mailbox-enabled client will keep resurfacing it, which is the intended behaviour: an unanswered ask should nag.
|
||||
|
||||
---
|
||||
|
||||
## 5. Habits that make a shared palace work
|
||||
|
||||
Small, and the whole value rests on them:
|
||||
|
||||
1. **Search before you answer, and enumerate before you conclude.** One empty search is not proof of silence — fresh drawers rank worst, so for anything from the last couple of days list by date and read the other machines' diaries.
|
||||
2. **Write the diary entry before the session ends.** It is the store other machines learn from fastest, and the only one that records failed attempts.
|
||||
3. **File retractions as loudly as findings.** "I was wrong about X, here is the measurement" is worth more than a new finding, because it stops N machines repeating a dead end.
|
||||
4. **Check the mailbox at wake-up, even when you expect nothing.** An empty result costs one call. Silence is only informative once you know delivery works.
|
||||
5. **Say which machine you are talking about.** On a shared palace, "the container" and "the host" are ambiguous and a path is not self-identifying.
|
||||
6. **After a write times out, verify — do not blindly retry.** The palace is single-writer for memory writes, so a timeout usually means the write *completed*. For coordination events this matters twice over: there is no idempotency guard, so a retried event forks the thread into two (RFC 003 §7.1).
|
||||
|
||||
---
|
||||
|
||||
## 6. See also
|
||||
|
||||
- `docs/rfc-003-coordination-log.md` — the coordination log: storage, semantics, security model, landmines.
|
||||
- `docs/rfc-001-global-palace.md` — how and why a palace is centralised; §5 what should not be global; §7 the landmines, including the `sync` hazard.
|
||||
- `docs/phase-1-exposure-runbook.md` — (moved) the site-specific exposure record is now private; the stub names the mechanism still awaiting extraction.
|
||||
- `docs/backup-and-recovery.md` — why a palace needs its own backup procedure.
|
||||
- `extensions/pi/README.md` — the pi-side client: provenance stamping and the auto-delivered mailbox.
|
||||
- `~/.agents/skills/mempalace/SKILL.md` — the protocol the agents themselves follow.
|
||||
@@ -1,463 +1,47 @@
|
||||
# Phase 1 exposure — newt on synlig, DNS, and the client auth model
|
||||
|
||||
Companion to [`rfc-001-global-palace.md`](./rfc-001-global-palace.md) (design + decisions) and
|
||||
[`synlig-primary-runbook.md`](./synlig-primary-runbook.md) (what is already installed on the primary).
|
||||
This doc covers only the step the other two leave open: **making the primary reachable** — runbook §4
|
||||
items 2 and 5.
|
||||
|
||||
> **Status 2026-08-14 — DONE. Exposed, seeded, and one client flipped.**
|
||||
> `https://mempalace.jordbo.se/mcp` has been serving since 2026-08-12 (§3.3–§3.6 are all ✅ below).
|
||||
> The palace was seeded 2026-08-14 15:07 from EMB-7KJ4VR4G (14,777 → 14,803 drawers) — whose palace was
|
||||
> itself carried over from the previous work computer **EMB-X1JY06WJ** around 2026-07-06, so the
|
||||
> primary's lineage is EMB-X1JY06WJ → EMB-7KJ4VR4G → synlig by two file-level copies (rfc-001 §4.4
|
||||
> Deviation) — and that machine's pi-devbox container is flipped and **verified end-to-end — see §3.8**, which is the
|
||||
> verification procedure that did not exist when the first flip was performed.
|
||||
> Still outstanding: the transcript feeder is **inert on every deployed image** (§3.7), and §7.6 is
|
||||
> still a hard blocker for the *second* machine to join (§4).
|
||||
|
||||
Original status, kept for the record — it contradicted this file's own ✅ section markers for two days,
|
||||
which is the failure mode a file-top status block invites:
|
||||
|
||||
**Status 2026-08-12 — Pangolin updated on nyvaken (done, yours). newt not yet installed on synlig.
|
||||
Nothing exposed. No client `.env` flipped.**
|
||||
|
||||
Read this before touching Pangolin: three of the four questions this step raises were **already decided**
|
||||
in RFC §6.2 on 2026-08-09, and re-deciding them differently is how the fleet ends up in two states.
|
||||
|
||||
---
|
||||
|
||||
## 1. The four questions, answered
|
||||
|
||||
| Question | Answer | Where it was decided |
|
||||
| --- | --- | --- |
|
||||
| Which port? | **8765**, path **`/mcp`** (liveness: `/healthz`) | `cli.py:2141` default; runbook §2.4 |
|
||||
| What does newt target? | **`172.17.0.1:8765`** (docker0), **never** `127.0.0.1` | RFC §6.2 Transport; runbook §2.4 |
|
||||
| Open, or authenticated? | **Authenticated. The primary is never an open public resource.** | RFC §6.2 Network posture |
|
||||
| Per-device credentials? | **No — Phase 1 ships the single shared bearer token.** Per-device tokens are Phase 4. | RFC §6.2 Authentication |
|
||||
|
||||
### 1.1 Why not per-device users at the proxy
|
||||
|
||||
The instinct — "create a Pangolin user per container, put the credentials in each `.env`, keep the
|
||||
usernames distinct" — is the right *goal* (revocation, attribution) reached through the wrong *layer*,
|
||||
twice over:
|
||||
|
||||
1. **mempalace validates exactly one token.** `hmac.compare_digest(provided, f"Bearer {srv.auth_token}")`
|
||||
(`mcp_server.py:5292-5295`) — there is no user table and no second credential. Per-device HTTP identity
|
||||
is not a configuration you can express today; it is Phase 4 work (a server-side
|
||||
`token → {device_id, scopes}` registry). RFC §6.2 chose the shared token for Phase 1 deliberately:
|
||||
*"iterate more feature rich but more complex solutions over time."*
|
||||
|
||||
2. **Pangolin's HTTP auth is browser-shaped; the clients are not.** SSO login, resource PIN and resource
|
||||
password all assume something that can follow a redirect, render a form and hold a session cookie.
|
||||
Every MemPalace client here is a headless JSON-RPC `POST` with an `Authorization` header — pi's
|
||||
extension, opencode's `type:remote` MCP entry, and `mempalace-pi-session --mode remote`'s
|
||||
`urllib.request.urlopen`. Point those at a user-authenticated resource and they receive a login page
|
||||
where JSON should be. Enabling that protection breaks precisely the clients it is meant to protect.
|
||||
|
||||
So: **Pangolin terminates TLS and nothing more** (RFC §6.2 Transport, decided 2026-08-09). The bearer token
|
||||
is the authentication. This is not "unprotected" — an unauthenticated request to `/mcp` gets a 401 from
|
||||
mempalace itself, verified A4/A5 in runbook §2.4.
|
||||
|
||||
**Consequence to accept consciously** (RFC §6.2, §7.3.2): until Phase 4 the primary **cannot tell devices
|
||||
apart**. `origin_device` is client-asserted and advisory — nothing load-bearing may depend on it, and
|
||||
revoking one laptop means rotating the token everywhere.
|
||||
|
||||
### 1.2 The one place per-device identity *does* exist today
|
||||
|
||||
Remote mode is not only HTTP. `mempalace_mine` expands its source path in the **server** process, so a
|
||||
client's staged transcripts must physically exist on the primary. The feeder therefore ships them over
|
||||
SSH into a **per-device inbox** before asking the server to mine its own copy:
|
||||
|
||||
```sh
|
||||
rsync -a --update -e "$ssh_cmd" "$STAGE/" "${SSH_TARGET%/}/$DEVICE/" # bin/mempalace-pi-session:677-680
|
||||
```
|
||||
|
||||
That SSH key **is** per-device identity, and it is individually revocable (one line out of
|
||||
`authorized_keys`) years before Phase 4 lands. It costs nothing extra, because the mining path needs SSH
|
||||
regardless.
|
||||
|
||||
Two implications people miss:
|
||||
|
||||
- **`mempalace.jordbo.se` alone does not enable mining.** HTTPS covers the read/write tool surface
|
||||
(`search`, `add_drawer`, `diary_write`, `kg_*`) — genuinely useful on its own, and the reason to do this
|
||||
at all. But `--mode remote` also needs `MEMPALACE_PI_SSH_TARGET` reachable. Budget for both paths.
|
||||
- **`DEVICE` defaults to `$(hostname)`** (`bin/mempalace-pi-session:163`). In a container that is the
|
||||
container hostname: either random per recreate (inboxes proliferate; each recreate re-mines into a fresh
|
||||
empty inbox) or identical across sibling devboxes (two containers writing one inbox). **Set
|
||||
`MEMPALACE_PI_DEVICE` explicitly per container.** It is a label, not a secret, so put it somewhere
|
||||
reviewable — a committed compose file — where duplicates are visible. That, not username hygiene in
|
||||
`.env`, is the discipline this design actually asks of you.
|
||||
|
||||
### 1.3 "Then why Pangolin at all, if the feeder uses SSH?"
|
||||
|
||||
Because they are not alternatives — they carry different traffic, and neither substitutes for the other.
|
||||
|
||||
| | Pangolin/newt (HTTPS) | SSH + rsync |
|
||||
| --- | --- | --- |
|
||||
| Carries | the **MCP tool surface**: `search`, `add_drawer`, `diary_write`, `kg_*` — every live tool call | **transcript files only**, once per session or cron run |
|
||||
| Used by | the pi extension, opencode `type:remote`, any MCP client | the feeder, internally (`bin/mempalace-pi-session:677-680`) |
|
||||
| Needed because | clients need one stable URL, reachable from wherever they are | `mempalace_mine` expands its source path **server-side**, so the server can only mine files on its own disk |
|
||||
|
||||
HTTPS alone is a palace you can query but cannot feed. SSH alone is files shipped with no live query API.
|
||||
The rsync is not a transport preference; it is a workaround for *where `mine` resolves paths*.
|
||||
|
||||
**Could SSH replace Pangolin?** Partly, and it is worth being honest about it:
|
||||
`ssh -L 8765:172.17.0.1:8765 synlig` yields a working local MCP endpoint with no public HTTPS at all.
|
||||
Three reasons this runbook does not do that:
|
||||
|
||||
1. **Direction.** synlig dials *out* through newt. That we reached for a dial-out tunnel rather than a
|
||||
port-forward is itself the evidence that inbound was not available — a corporate host does not accept
|
||||
connections from a phone on a foreign network.
|
||||
2. **MCP clients want a durable URL**, not a per-session forwarded port. opencode `type:remote` takes a
|
||||
URL; a forward that drops takes the tools down mid-session.
|
||||
3. The forward must be up on **every device before every session**. Pangolin is up once.
|
||||
|
||||
**The weak point, stated plainly.** The rsync runs *client → synlig*, so it needs synlig's SSH reachable
|
||||
**from the client**. Were that already true everywhere, no tunnel would be needed for MCP either. So the
|
||||
honest expectation after Phase 1 is: **query and write from anywhere, mine only from devices that can
|
||||
reach synlig's SSH** (corporate network / VPN / LAN). See §4 for the change that would remove that limit.
|
||||
|
||||
---
|
||||
|
||||
## 2. The bind trap, in full
|
||||
|
||||
RFC §6.2 and runbook §2.4 already say **do not bind loopback behind the tunnel**, because
|
||||
`enforce_host_pin = _http_is_loopback(host)` (`mcp_server.py:5367`) makes a loopback bind reject the
|
||||
proxy's forwarded `Host:` with a **403** that reads exactly like a Pangolin misconfiguration.
|
||||
|
||||
**Additional finding, 2026-08-12 — the same reflex also silently removes authentication.** Token
|
||||
resolution in `cmd_serve` (`cli.py:1447-1450`) is:
|
||||
|
||||
```python
|
||||
loopback = _server_is_loopback(host)
|
||||
if not token and not loopback and not args.allow_insecure:
|
||||
token, token_created = _load_or_create_server_token(palace_path)
|
||||
```
|
||||
|
||||
Auto-minting is gated on the bind being **non-loopback**. A loopback bind therefore starts with **no token
|
||||
at all** — no error, no warning, `--allow-insecure` not required — because the server has concluded it is
|
||||
only reachable locally, while the tunnel is serving it to the internet. Bind loopback behind newt and you
|
||||
get a 403 wall *and*, the moment anything relaxes the Host pin, an unauthenticated palace.
|
||||
|
||||
Both failure modes have the same cure, already implemented in
|
||||
`contrib/systemd/mempalace-serve.service`: **bind `172.17.0.1`**. Non-loopback, so the Host pin relaxes and
|
||||
the token is mandatory; docker0-only, so newt reaches it and the LAN does not.
|
||||
|
||||
> Belt and braces: set `MEMPALACE_MCP_HTTP_TOKEN` explicitly in the unit rather than relying on
|
||||
> auto-minting. Then no future bind change can quietly drop authentication.
|
||||
|
||||
---
|
||||
|
||||
## 3. Steps
|
||||
|
||||
Ordered so nothing is reachable before it is authenticated.
|
||||
|
||||
### 3.1 Start the primary (runbook §4.3 — one `sudo`, unit already staged)
|
||||
|
||||
```sh
|
||||
sudo loginctl enable-linger ecsjper
|
||||
cd ~/.config/systemd/user && mv mempalace-serve.service.staged mempalace-serve.service
|
||||
systemctl --user daemon-reload && systemctl --user enable --now mempalace-serve
|
||||
|
||||
curl -s 172.17.0.1:8765/healthz # expect ok
|
||||
curl -s 127.0.0.1:8765/healthz # expect NOTHING — connection refused, exit 7 (see below)
|
||||
ss -ltnp | grep 8765 # expect 172.17.0.1:8765 only
|
||||
```
|
||||
|
||||
⚠ **The loopback probe returns empty, not 403** — corrected 2026-08-12 against the real run. Nothing is
|
||||
listening on `127.0.0.1`, so the connection is refused at TCP level and `curl -s` prints nothing; check it
|
||||
with `-w '%{http_code}'` → `000` and `$?` → `7`. The 403 belongs to a *different* configuration: server
|
||||
bound **to loopback**, receiving a proxy-forwarded foreign `Host:` (§2, verified 2026-08-10). With a
|
||||
docker0-only bind you cannot get 403 from loopback, because you never get far enough to send a header.
|
||||
Refusal is the stronger signal of the two: it proves the loopback and LAN surface is not listening at all.
|
||||
If it *hangs* instead, or `ss` shows `0.0.0.0:8765`, stop — that is not this configuration.
|
||||
|
||||
### 3.2 Collect the shared token
|
||||
|
||||
```sh
|
||||
cat ~/.mempalace/server/f5d849287f6d73f0141b29d7/token
|
||||
```
|
||||
|
||||
Directory name is `sha256(realpath(palace))[:24]` — it changes if the palace path ever moves. Store via the
|
||||
`.env.age` flow, 0600 (RFC §6.2).
|
||||
|
||||
### 3.3 newt on synlig — ✅ done 2026-08-12 (installed, connected to Pangolin)
|
||||
|
||||
synlig runs Docker (Gitea Actions runner + digikam) but **no tunnel client** — runbook §4.2. Pangolin on
|
||||
nyvaken cannot dial in; synlig must dial out. Add a `newt` container with the credentials Pangolin issues
|
||||
for a new site.
|
||||
|
||||
Because newt runs in Docker on this box, the docker0 bind is already correct for it: from inside the
|
||||
container the primary is `172.17.0.1:8765`. Verify from *inside* newt's network namespace, not from the
|
||||
host, before touching DNS.
|
||||
|
||||
> synlig has 7.8 GiB shared with a CI runner (runbook §1). newt is small, but do not colocate anything
|
||||
> else here casually.
|
||||
|
||||
**Confirm next, now that newt is up.** "Connected to Pangolin" proves newt reached *nyvaken* — a different
|
||||
claim from newt reaching *the palace*, and the two fail independently:
|
||||
|
||||
```sh
|
||||
# from inside newt's namespace, not from the host
|
||||
docker exec <newt-container> wget -qO- http://172.17.0.1:8765/healthz # expect ok
|
||||
```
|
||||
|
||||
If that hangs or refuses while the Pangolin dashboard shows the site online, the tunnel is fine and the
|
||||
*target* is wrong — look at the resource's upstream address (§3.5), not at newt. Note this check needs
|
||||
§3.1 done first: if `mempalace-serve` is not running yet, it fails for that reason alone.
|
||||
|
||||
### 3.4 DNS at the web hotel — ✅ done 2026-08-12
|
||||
|
||||
One CNAME: `mempalace` → **the same target your existing Pangolin resources use** (nyvaken's public
|
||||
hostname). RFC §6.2 costed this as *"one DNS record per service on the web hotel is the whole setup cost."*
|
||||
Done: `mempalace.jordbo.se` resolves and terminates TLS at Pangolin on nyvaken — the §6.2 estimate held.
|
||||
|
||||
⚠️ Not verified from here: nyvaken's public FQDN, and whether your web hotel permits a CNAME at that label
|
||||
(some require an A record, or forbid CNAME where other records exist). Confirm before assuming a 5-minute job.
|
||||
|
||||
### 3.5 Pangolin resource
|
||||
|
||||
- Target: newt site → **`http://172.17.0.1:8765`** — path `/mcp` (plus `/healthz` for the external probe).
|
||||
- ⚠️ **The scheme is `http`, not `https`.** The primary runs `serve --host 172.17.0.1 --port 8765` with no
|
||||
cert (`contrib/systemd/mempalace-serve.service`): TLS terminates **at Pangolin**, which is the entire
|
||||
point of the §6.2 decision. Point the resource at `https://172.17.0.1:8765` and Pangolin attempts a TLS
|
||||
handshake against a plaintext listener — you get a 502/Bad Gateway from outside while the server itself
|
||||
looks perfectly healthy on `curl 172.17.0.1:8765/healthz`. (Got this wrong on the first attempt
|
||||
2026-08-12, because this line used to omit the scheme.)
|
||||
- **Auth: none at the Pangolin layer** (§1.1). TLS termination only.
|
||||
- ⚠️ **If resource auth is left on, the signature is a `302`, not a 401 or 403** — hit for real 2026-08-12:
|
||||
```
|
||||
HTTP/2 302
|
||||
location: https://pangolin.jordbo.se/auth/resource/<uuid>?redirect=https%3A%2F%2Fmempalace.jordbo.se%2Fhealthz
|
||||
content-length: 0
|
||||
```
|
||||
This is §1.1's "browser-shaped auth" arriving as a concrete symptom: Pangolin sends the login redirect
|
||||
**before** proxying, so the palace never sees the request and its journal stays silent. `curl -s` shows
|
||||
an empty body and an MCP client sees non-JSON. Diagnose with `-D-` or
|
||||
`-w '%{http_code} %{redirect_url}'` — a `location:` pointing at `/auth/resource/…` means the fix is in
|
||||
the Pangolin UI (switch the site's authentication off), not in the palace, the unit, or newt.
|
||||
**A 302 is unambiguously good news:** DNS, TLS and routing all worked — only the auth layer intervened.
|
||||
- Do **not** attach an `Origin`-injecting proxy or browser client: a *present* non-loopback `Origin` is a
|
||||
hard 403 with no override (runbook §2.4 B3).
|
||||
|
||||
**The server's entire HTTP surface is two exact paths**, so path-scoped rules cover it completely
|
||||
(`mcp_server.py:5299-5318`, read 2026-08-12):
|
||||
|
||||
| Method + path | Auth | Notes |
|
||||
| --- | --- | --- |
|
||||
| `GET /healthz` | none (Host/Origin gated only) | the liveness probe; works with no creds by design |
|
||||
| `POST /mcp` | `Authorization: Bearer <token>`, `hmac.compare_digest` on the exact string | the whole tool surface |
|
||||
| anything else | — | `send_error(404)` from the palace itself |
|
||||
|
||||
Three consequences worth having in writing:
|
||||
|
||||
- **Path-scoped Pangolin rules are not a compromise here, they are tighter than a host-wide proxy** and
|
||||
lose nothing — there is no third endpoint to forget.
|
||||
- **`/mcp` is matched exactly** (`if path != "/mcp"`), so a client URL with a trailing slash gets a 404
|
||||
from the palace. Configure clients as `https://mempalace.jordbo.se/mcp` — no trailing slash.
|
||||
- **There is no `GET /mcp`, no SSE, no session id, no `DELETE`.** This is plain JSON-RPC over POST, not MCP
|
||||
streamable-HTTP. A strict client that opens with a `GET` handshake will see 404; pi's extension,
|
||||
opencode `type:remote` and the feeder all POST directly and are fine.
|
||||
|
||||
### 3.6 Verify end-to-end before flipping any client — ✅ passed 2026-08-12
|
||||
|
||||
```sh
|
||||
curl -s https://mempalace.jordbo.se/healthz # ok ✓ from devbox AND synlig
|
||||
curl -s -o /dev/null -w '%{http_code}\n' -X POST \
|
||||
https://mempalace.jordbo.se/mcp # 401 ✓ the token is doing its job
|
||||
TOKEN=$(cat ~/.mempalace/server/f5d849287f6d73f0141b29d7/token) # on synlig
|
||||
curl -s -X POST https://mempalace.jordbo.se/mcp \
|
||||
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
|
||||
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | head -c 300 # 36 tools
|
||||
```
|
||||
|
||||
No `initialize` and no `Accept: text/event-stream` needed — see the surface table above; a bare
|
||||
`tools/list` POST is a complete request. Run the 200 and the 401 from **two different networks** (a client
|
||||
box and the primary itself): passing from only one leaves split-horizon DNS untested.
|
||||
|
||||
The 401 check matters as much as the 200: it is the only evidence that the thing you just published to the
|
||||
internet is not open. Then, and only then, Phase 1 client flip — **one machine first** (RFC §8), and
|
||||
remember opencode containers need the §4.1 sidecar merge before their `.env` takes effect.
|
||||
|
||||
### 3.7 Flipping a client: set three variables, or none
|
||||
|
||||
⚠ **`MEMPALACE_REMOTE_URL` on its own does not degrade to local feeding — it stops feeding.** `auto` mode
|
||||
switches to `remote` the moment the URL is set, and remote mode then refuses to run without an SSH target:
|
||||
|
||||
```sh
|
||||
auto) if [[ -n "$REMOTE_URL" ]]; then MODE="remote"; else MODE="local"; fi ;; # :286
|
||||
...
|
||||
command -v rsync >/dev/null 2>&1 || { echo "error: rsync not found ..."; exit 3; } # :297
|
||||
if [[ -z "$SSH_TARGET" ]]; then
|
||||
echo "error: MEMPALACE_PI_SSH_TARGET unset (needed for --mode remote)" >&2; exit 1 # :298-300
|
||||
fi
|
||||
```
|
||||
|
||||
That exit happens **before anything is staged or filed**, and a cron-driven feeder will simply start
|
||||
failing — the loudest symptom is silence, which is the hardest kind to notice. Two safe orders:
|
||||
|
||||
- **Both paths at once:** set `MEMPALACE_REMOTE_URL`, `MEMPALACE_REMOTE_TOKEN` **and**
|
||||
`MEMPALACE_PI_SSH_TARGET` (plus `MEMPALACE_PI_DEVICE`, §1.2) in the same edit.
|
||||
- **HTTPS first, mining later:** set the URL and token, and pin the feeder to `--mode local` until the SSH
|
||||
target exists. Tools then read/write the shared palace while transcripts keep landing in the local one.
|
||||
|
||||
`MEMPALACE_REMOTE_URL` is the **full endpoint including `/mcp`, with no trailing slash** — the feeder POSTs
|
||||
to it verbatim (`urllib.request.Request(url, data=payload, …)`, `:718`), and §3.5 shows the server matches
|
||||
`/mcp` exactly. Matches the existing examples (`docker-compose.mempalace.yml:7`,
|
||||
`MEMPALACE_REMOTE_URL=http://<reachable-host>:8765/mcp`). For this fleet:
|
||||
`MEMPALACE_REMOTE_URL=https://mempalace.jordbo.se/mcp`
|
||||
|
||||
Either way, run the feeder once by hand and read its exit code before trusting the timer. This is
|
||||
precisely the failure "one machine first" is meant to contain.
|
||||
|
||||
**The same misconfiguration has two different symptoms depending on who invokes the feeder** — checked in
|
||||
the code 2026-08-12, and the quieter one is the trap:
|
||||
|
||||
| Caller | Behaviour with `REMOTE_URL` set and `SSH_TARGET` unset |
|
||||
| --- | --- |
|
||||
| direct run, session-end hook, cron | `exit 1` with `error: MEMPALACE_PI_SSH_TARGET unset` (`:298-300`) |
|
||||
| **pi-devbox container start, images ≤ v1.7.0** | **silently skips — no error, no log file** (`entrypoint-user.sh:134`: `: # remote palace but no inbox configured — nothing we can ship to; skip quietly`) |
|
||||
| pi-devbox container start, images after `cbd7cf5` | prints `MemPalace catch-up skipped: remote palace with no transcript inbox` to the start output *and* to `mempalace-catchup.log`, naming both variables |
|
||||
|
||||
The silent skip is fixed in pi-devbox (`cbd7cf5`), but the fix is in
|
||||
`entrypoint-user.sh`, which is `COPY`d in `Dockerfile.base` — so **every container running an image built
|
||||
before that base rebuild still skips silently.** That is the whole fleet today. Until the rebuild lands,
|
||||
assume silence and check by hand.
|
||||
|
||||
The old skip happened *before* the subshell that writes `~/.pi/agent/mempalace-catchup.log`, so there was
|
||||
not even an empty log to notice. Someone asking "why is nothing from this container in the palace?" found
|
||||
no artifact at all. The skip itself is correct — there is genuinely nothing to ship to — but it was
|
||||
indistinguishable from a healthy run that had nothing to do, which is the worst property a memory system
|
||||
can have: **the failure looks exactly like success.**
|
||||
|
||||
→ On a pi-devbox container, confirm the feeder is actually alive after a flip rather than assuming:
|
||||
|
||||
```sh
|
||||
mempalace-pi-session --reason manual-check; echo "exit=$?" # exit=0 and a filed count, not silence
|
||||
cat ~/.pi/agent/mempalace-catchup.log # missing file = the entrypoint skipped
|
||||
```
|
||||
|
||||
### 3.8 Verify the flip actually took — ✅ done 2026-08-14 on EMB-7KJ4VR4G
|
||||
|
||||
§3.6 verifies the *endpoint* before you flip. §3.7 tells you *how* to flip. Neither verifies the thing
|
||||
you actually care about afterwards: **that the agent's palace tools are now talking to synlig.** That
|
||||
gap is why the first flip was "done" for an hour before anyone could say whether it had worked.
|
||||
|
||||
There are **three separate claims** here and they fail independently. Check them in order; each one is
|
||||
cheap and rules out a different fault.
|
||||
|
||||
**(a) Did the container receive the variable?** In a shell *inside* the container:
|
||||
|
||||
```sh
|
||||
env | grep MEMPALACE_REMOTE_URL # -> https://mempalace.jordbo.se/mcp
|
||||
```
|
||||
|
||||
⚠ Do **not** run a bare `env | grep MEMPALACE` — that prints the bearer token into your scrollback.
|
||||
If the variable is absent, the cause is almost always §3.7's edit not having been applied: `env_file`
|
||||
is read at container **create** time and baked into the container config, so **`docker compose up -d`
|
||||
is required; `docker compose restart` silently reuses the old config.** Evidence from 2026-08-14 —
|
||||
running container config-hash `3a55e09ac19e0118` vs compose-computed `8a7e0cd4a677a9c1`; a differing
|
||||
hash is what makes `up -d` recreate. `--force-recreate` is not needed.
|
||||
**Never `docker compose down -v`** to "pick up" a change: on pi-devbox that destroys seven named
|
||||
volumes, including `devbox-pi-config` (pi's config **and every session transcript**), `devbox-uv`, and
|
||||
`devbox-chroma-cache` (a large embedding-model re-download).
|
||||
|
||||
**(b) Is the server reachable and the token accepted?** Still inside the container — this proves the
|
||||
network path and the credential, independently of any agent:
|
||||
|
||||
```sh
|
||||
curl -s -X POST "$MEMPALACE_REMOTE_URL" \
|
||||
-H "Authorization: Bearer $MEMPALACE_REMOTE_TOKEN" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | head -c 200
|
||||
```
|
||||
|
||||
`401` = token problem (compare `md5sum` of the client value against synlig's token file — trailing
|
||||
whitespace from an editor is the classic cause). Connection failure = DNS/tunnel, not auth.
|
||||
|
||||
**(c) Did the agent's bridge actually switch?** This is the claim that matters and the one that cannot
|
||||
be checked from bash — there is no log file. Ask the agent running in the container for
|
||||
`mempalace_status` and read the palace path it reports:
|
||||
|
||||
```
|
||||
sqlite_integrity.palace = /home/ecsjper/.mempalace/palace # ← proves remote
|
||||
```
|
||||
|
||||
**This is the cheapest and strongest discriminator: that path cannot exist inside the container**,
|
||||
whose user is `developer` with `HOME=/home/developer` and whose local palace is
|
||||
`/home/developer/.mempalace/palace`. One call, no token in scrollback, no writes. If instead it
|
||||
reports the local path, the bridge fell back: `createClient()` reads `MEMPALACE_REMOTE_URL` **once at
|
||||
load**, so re-check (a) and restart the agent, not just the container.
|
||||
|
||||
#### What does *not* verify the flip — read this before inventing your own check
|
||||
|
||||
- **❌ A drawer count.** Central was seeded *from* the client's own palace, so both report ~14,777.
|
||||
Counts cannot tell the two apart. Worse, `mempalace status` counts **chunk rows, not logical
|
||||
drawers** (three drawers plus a 2-chunk diary presented as +9), so a delta does not even mean what
|
||||
it looks like. Never reason about palace identity or contents from a count.
|
||||
- **❌ Writing a drawer through the palace tools and reading it back through the same tools.** This
|
||||
succeeds *identically whether or not the flip worked* — both palaces are healthy and were seeded
|
||||
from the same source, so a write-then-read round-trips either way. It only becomes evidence if the
|
||||
drawer is read back over a **different transport** (the `curl` in (b), by drawer id) or is proven
|
||||
**absent** from the local sqlite. This is the same false-positive class as the CLI below, one layer
|
||||
up, and it is an easy trap to fall into precisely because it feels like an end-to-end test.
|
||||
- **❌ The `mempalace` CLI, in any form.** The CLI has **no remote support whatsoever** — its only
|
||||
selector is `--palace <path>` — so it reads and writes the LOCAL on-disk archive via
|
||||
`config.json`. After a flip that archive is dead, yet `mempalace status` / `mempalace search` will
|
||||
cheerfully report ~14,777 drawers and look exactly like success. It is a false-positive machine.
|
||||
**Corollary that bites later:** memories filed with the CLI after a flip land in the dead archive,
|
||||
not in central, and a transcript backfill must therefore be mined **on synlig**, where the CLI's
|
||||
local palace *is* the central one.
|
||||
|
||||
#### Expected failure behaviour, so you can recognise it
|
||||
|
||||
The bridge is **fail-closed, not fail-local.** If synlig is unreachable, DNS fails, or the token is
|
||||
rejected, the extension retries a bounded number of times, prints `mempalace-mcp unavailable after
|
||||
retries; continuing without palace tools`, and **does not register the palace tools at all**. It does
|
||||
not silently write to the local palace; there is no dual-write and no local mirror, and in remote mode
|
||||
no local `mempalace-mcp` process is spawned (verified: zero mempalace processes in the flipped
|
||||
container). So **"the agent has no `mempalace_*` tools" is the expected symptom of a server, token, or
|
||||
DNS fault** — not of a broken container. Diagnose with (b).
|
||||
|
||||
#### Rollback — ~30 seconds, loses nothing
|
||||
|
||||
Comment out `MEMPALACE_REMOTE_URL` in `.env`, then `docker compose up -d`. The bridge falls back to the
|
||||
local stdio palace, which is intact. Note the local archive is **frozen, not empty**: it stops at the
|
||||
moment of the flip, so anything the agent filed into central since then will not be there.
|
||||
|
||||
---
|
||||
|
||||
## 4. Still open
|
||||
|
||||
- **Feeding without SSH — the upstream ask that would close §1.3's gap.** Have the feeder send *content*
|
||||
over MCP (`add_drawer` / `diary_write`, which it already calls) instead of asking the server to mine a
|
||||
path it must first rsync there. HTTPS would then be genuinely sufficient and mining would work from any
|
||||
network. Until then, mining is limited to devices that can reach synlig's SSH.
|
||||
- **Per-device tokens** — Phase 4. Until then `origin_device` is advisory (§1.1).
|
||||
- **§7.6 diary dedup** must be settled *before* the **next** §4.4 join; replay duplicates every entry.
|
||||
⚠ **2026-08-14 — the first join did not resolve this, it SIDESTEPPED it.** The seed was a file-level
|
||||
copy of one palace, which replays no diaries and therefore cannot duplicate them. The success of
|
||||
that join is *not* evidence the replay path is safe — it never exercised it. §7.6 remains a hard
|
||||
blocker for the second machine, which is the one that will actually need merge semantics.
|
||||
- **The transcript feeder is inert on every deployed image**, so nothing is being mined automatically
|
||||
anywhere in the fleet (§3.7). Regaining it needs a **new tagged pi-devbox release** — its CI
|
||||
publishes only on `v*` tags and the latest tag *is* the currently deployed image — built on
|
||||
pi-devbox ≥ `7c00dd6` **and** mempalace-toolkit ≥ `29e660e`. Both are required: the first restores
|
||||
the container-start catch-up, the second is where extension-side feeding was implemented at all.
|
||||
- **Whether *native* pi on a flipped machine was also flipped** is a per-machine question. Answered for
|
||||
**EMB-7KJ4VR4G (2026-08-14): native pi is not installed there *yet*** — no `pi`/`mempalace` on the host
|
||||
PATH, no `~/.config/pi/`, no `~/.pi/agent/extensions/`; it is a replacement machine that has not had
|
||||
native pi set up. So nothing further needs flipping there **today** — but this is a deferred hazard,
|
||||
not a closed one: a native install resolves its palace to `~/.mempalace`, which on that host is the
|
||||
bind-mounted **frozen archive**, so it would start writing there and split the machine's memory from
|
||||
central silently. **Flip native pi at install time, not after.** **Still open for every other host**
|
||||
running native pi beside a flipped container. Verify with the §3.8(c) path check, per machine.
|
||||
Note native pi has **no `.env` to edit** — pi loads no dotenv file and has no `env` block in
|
||||
`settings.json`, so the variables must come from the shell that launches it; recipe in
|
||||
[`extensions/pi/README.md`](../extensions/pi/README.md#transport-local-vs-external) § Transport. On
|
||||
EMB-7KJ4VR4G specifically, that tree's `config.json` also points at
|
||||
`palace_path=/home/developer/.mempalace/palace` — a *container* path that does not exist on macOS — so
|
||||
a native install must not inherit it unexamined.
|
||||
- **§7.2**: never run `mempalace sync` against the shared palace. Doubly true now that the pi/opencode
|
||||
feeders stage *inside* the palace root, which puts staged sources in scope for a sync of the palace dir.
|
||||
- **nyvaken's public FQDN and the web hotel's CNAME rules** — unverified (§3.4).
|
||||
# (moved) Phase 1 exposure runbook
|
||||
|
||||
This file used to contain the Phase 1 exposure record for one specific
|
||||
deployment — the primary host and tunnel host by name, the DNS registrar step,
|
||||
the Pangolin/newt resource wiring, the shared-token handling, per-machine flip
|
||||
dates, and a palace lineage naming three work machines.
|
||||
|
||||
**That content now lives in a private repository**, for the same reason
|
||||
[`synlig-primary-runbook.md`](synlig-primary-runbook.md) does: a host inventory
|
||||
is operator data for one deployment, not part of the toolkit. This repository is
|
||||
public and keeps only host-agnostic *mechanism*.
|
||||
|
||||
What lives where:
|
||||
|
||||
| Content | Home |
|
||||
|---|---|
|
||||
| Why a palace needs a special backup, and how to restore one | [`backup-and-recovery.md`](backup-and-recovery.md) (here) |
|
||||
| The coordination log: semantics, delivery, trust model, landmines | [`rfc-003-coordination-log.md`](rfc-003-coordination-log.md) (here) |
|
||||
| What the stores are for, and how a fleet shares one palace | [`fleet-memory.md`](fleet-memory.md) (here) |
|
||||
| Unit/timer/plist templates | [`contrib/`](../contrib/) (here) |
|
||||
| Which host is primary and which fronts the tunnel, at which addresses, as which user | private fleet repository |
|
||||
| Registrar/DNS records, tunnel resource config, token custody | private fleet repository |
|
||||
| Per-machine flip dates, seeding history, palace lineage | private fleet repository |
|
||||
|
||||
## The mechanism this file also carried — not yet extracted
|
||||
|
||||
Unlike the primary-host runbook, this file was **mixed**: several sections were
|
||||
reusable mechanism that would be true of anyone's palace, and those are not
|
||||
published anywhere else yet. Named here so the debt is visible rather than lost:
|
||||
|
||||
- **The bind trap, in full.** Why binding a palace to a public interface is not
|
||||
the same as exposing it, and the `Host`/`Origin` pin that makes an MCP
|
||||
endpoint refuse requests that arrive with the wrong hostname — the single
|
||||
most surprising failure in the whole exposure path.
|
||||
- **Why not per-device users at the proxy.** The reasoning behind one shared
|
||||
fleet token instead of per-device proxy credentials, and the consequence
|
||||
documented in [`rfc-003-coordination-log.md`](rfc-003-coordination-log.md) §6:
|
||||
the deployment authenticates the *fleet*, not the *agent*.
|
||||
- **The one place per-device identity does exist**, and why that is the feeder
|
||||
path rather than the HTTP path.
|
||||
- **Flipping a client: three variables, or none.** The all-or-nothing shape of
|
||||
pointing a machine at a remote palace, and how a half-flipped client fails.
|
||||
|
||||
Until that extraction happens, the mechanism is readable only in the private
|
||||
record. If you are standing up your own palace over HTTP, the two pieces you
|
||||
must not skip are the `Host`/`Origin` pin and the fact that a client is flipped
|
||||
by environment variables that travel as a set.
|
||||
|
||||
@@ -183,12 +183,12 @@ correctness, not for this join.
|
||||
`filed_at` spread over the same parent drawers — 12 in May, 52 in June, 13,619 in July, 882 in August — is
|
||||
the concrete case for regime A: an MCP replay would restamp all of it to the join date.
|
||||
|
||||
## 5. Open decisions for ALC
|
||||
## 5. Open decisions
|
||||
|
||||
1. **Chronology: keep it or flatten it?** Regime A keeps it and costs a service stop plus an rsync of the
|
||||
source palace onto synlig. Regime B is simpler and loses it. This is the fork in the road.
|
||||
source palace onto the primary host. Regime B is simpler and loses it. This is the fork in the road.
|
||||
2. **Hallways for joined content:** re-mine to rebuild, or accept degraded `traverse`?
|
||||
3. **Do MBP-M1-2020 and tor-ms22 keep their palaces on persistent storage?** If either is a Docker named
|
||||
3. **Do the secondary machines keep their palaces on persistent storage?** If a palace lives in a Docker named
|
||||
volume rather than a bind mount, its un-migrated content dies on the next container recreate — so the
|
||||
census (Phase A) is time-sensitive there, and flipping before censusing is risky.
|
||||
4. **Does §7.6 get fixed client-side (in the joiner) or upstream (probe-and-skip in `diary_write`)?** The
|
||||
|
||||
@@ -0,0 +1,552 @@
|
||||
# RFC 003 — The coordination log (`logstream`)
|
||||
|
||||
**Status:** implemented and in production use since 3.7.x. This document is a *retrospective* specification, written after the fact.
|
||||
**Author:** pi (agent), 2026-08-26.
|
||||
**Applies to:** mempalace 3.8.0 (`logstream.py`, 1261 lines), mempalace-toolkit `5b8d78f` (the pi-side mailbox).
|
||||
**Context:** every `mempalace_event_*` and `mempalace_artifact_*` tool description cites "RFC 003". `logstream.py`'s module docstring is headed *"Agent coordination event log for MemPalace (RFC 003)"* and enumerates five "Design constraints (RFC 003)". Inline comments cite "RFC 003 phase 5", "RFC 003 suggested defaults" and "the first RFC 003 dogfood". **No such document has ever existed** — verified 2026-08-26 by searching both the toolkit repository and the primary host. This RFC transcribes the spec the implementation already believes in, and — more usefully — records what it does *not* do.
|
||||
|
||||
Read §7 before you build anything on this log; it is the part that is not obvious from the tool descriptions. Every mechanical claim below cites `file.py:LINE` in mempalace 3.8.0, and §10 indexes them so any claim can be re-verified without re-reading 1261 lines. Claims marked **measured** were executed against the live fleet on the date given; claims without that marker are reads of the source. Where the code and the shipped tool descriptions disagree, §10.1 says so explicitly rather than quietly siding with one.
|
||||
|
||||
One scoping note up front. This RFC covers the **log**: its storage, its append and query semantics, delivery, and the trust model. It does not specify **replication**, which the source already attributes to a different document (`# ── Replication (RFC 004 step 0: logstream multi-master) ──`, `logstream.py:1096`). RFC 004 does not exist either; §8.2 states what it owes.
|
||||
|
||||
---
|
||||
|
||||
## 1. The problem
|
||||
|
||||
The palace stores what an agent *knows*: drawers are semantic, retrieved by meaning, and deliberately have no addressee. That shape is wrong for four things a fleet of agents actually needs:
|
||||
|
||||
1. **Addressing.** "This is for the machine that owns the release" cannot be expressed as a drawer. A drawer is found by whoever happens to search for the right words.
|
||||
2. **A reply that closes something.** Semantic memory has no notion of an outstanding question. Nothing in a drawer can be *owed*.
|
||||
3. **Exact payloads.** A unified diff must survive byte-for-byte. Drawers are chunked and embedded; that is a feature for prose and a defect for patches.
|
||||
4. **Order.** "What happened after this?" needs an append cursor, not a similarity score.
|
||||
|
||||
The coordination log adds exactly those four properties and nothing else. It is a second store beside the palace, not a new kind of drawer.
|
||||
|
||||
### 1.1 What it is not
|
||||
|
||||
Stated first, because every misuse of this log so far has come from assuming one of these:
|
||||
|
||||
- **Not a bus.** Nothing subscribes by default; nothing is delivered "live" unless a client is holding an SSE connection or polling. There is no delivery window and nothing is lost by being offline when an event is written.
|
||||
- **Not a chat channel.** Latency is bounded by *when the recipient next runs*, which in a fleet of workstations is hours, days or weeks. §3.5 gives the measured numbers for the case where the recipient is awake.
|
||||
- **Not authenticated per agent.** `from_agent` is a routing label with the trust properties of an e-mail `From:` header (§6).
|
||||
- **Not a queue.** Nothing is consumed, acknowledged-and-removed, or retried. Events are permanent (§9.1) and "handled" is a *derived* property (§3.3).
|
||||
|
||||
---
|
||||
|
||||
## 2. Verified starting point (2026-08-26, mempalace 3.8.0)
|
||||
|
||||
| Piece | State | Evidence |
|
||||
|---|---|---|
|
||||
| Separate SQLite store, inside the palace dir | ✅ `logstream.sqlite3` beside `chroma.sqlite3`; WAL; dir `chmod 0700` best-effort | `logstream.py:51`, `mcp_server.py:1108-1117`, `logstream.py:430-433` |
|
||||
| `events`, `artifacts`, `event_artifacts` tables | ✅ Created idempotently at open | `logstream.py:447-497` |
|
||||
| Hybrid logical clock on every event | ✅ `hlc` populated on every local append | `logstream.py:668`, `hlc.py:1-21` |
|
||||
| Seven MCP tools (append/list/wait/ack, artifact put/get, patch_submit) | ✅ | `mcp_server.py:4546-4771` |
|
||||
| Server-sent events push | ✅ `GET /logstream/stream`, auth required, 15 s heartbeat, ≤8 clients | `mcp_server.py:7570-7671`, `7189-7194` |
|
||||
| `GET /logstream/events` | ❌ **Never implemented.** Only `/logstream/stream` exists | see §7.8 |
|
||||
| Auto-delivered mailbox (owed-set derivation + injection) | ✅ **Client-side, not server-side** — lives in the toolkit's pi extension | `extensions/pi/mempalace.ts` (toolkit `5b8d78f`) |
|
||||
| Idempotency guard on append | ❌ **Absent.** See §7.1 | `logstream.py:625-716` |
|
||||
| Retention / TTL / compaction | ❌ Absent by design-so-far. See §9.1 | no `DELETE FROM events` in the package |
|
||||
| Interaction with `mempalace sync` | ✅ **None.** `sync` prunes drawers only | `cli.py:1057`, `sync.py` (no logstream references) |
|
||||
|
||||
Two things that sound like this feature and are not:
|
||||
|
||||
- **`mempalace_mesh_peers` is not a fleet roster.** It reports peer *replicas*. A hub-and-spoke deployment — many thin MCP clients of one server — correctly reports `peers: []` while every machine in the fleet is actively writing to the same log. Deciding "coordination does not apply to me" from an empty peer list is a measured failure mode, not a hypothetical one.
|
||||
- **`origin_replica` is not the writing machine.** It is a property of the palace *directory* (§3.2).
|
||||
|
||||
---
|
||||
|
||||
## 3. Design
|
||||
|
||||
The five constraints in the module docstring (`logstream.py:1-18`) are the design, quoted verbatim because they are already normative in the implementation:
|
||||
|
||||
> - No Chroma dependency, no vector index open — plain SQLite only.
|
||||
> - Append-only: events are immutable; corrections are new events that reference prior events.
|
||||
> - Exact payloads: event bodies and artifact content are stored verbatim.
|
||||
> - Safe under concurrent HTTP requests (WAL + per-instance lock, same pattern as `knowledge_graph.py`).
|
||||
> - Explicit size limits with clear errors, never silent truncation.
|
||||
|
||||
The shape, in the same idiom as RFC 001 §4.1:
|
||||
|
||||
```
|
||||
agent on device A ─┐ palace directory
|
||||
agent on device B ─┼── MCP /mcp ──► server ─┬─ chroma.sqlite3 (drawers: what you know)
|
||||
agent on device C ─┘ │ ├─ knowledge_graph.sqlite3 (facts, temporal)
|
||||
│ └─ logstream.sqlite3 (events + artifacts)
|
||||
GET /logstream/stream events ── append-only, permanent
|
||||
(SSE, auth, ≤8 clients) artifacts ── verbatim, ≤4 MiB, sha256
|
||||
event_artifacts ── join, checked at append
|
||||
```
|
||||
|
||||
### 3.1 Data model
|
||||
|
||||
```sql
|
||||
CREATE TABLE IF NOT EXISTS events (
|
||||
id TEXT PRIMARY KEY,
|
||||
type TEXT NOT NULL,
|
||||
stream TEXT NOT NULL,
|
||||
room TEXT NOT NULL,
|
||||
from_agent TEXT NOT NULL,
|
||||
to_agent TEXT,
|
||||
correlation_id TEXT,
|
||||
branch TEXT,
|
||||
base_commit TEXT,
|
||||
status TEXT,
|
||||
body TEXT NOT NULL DEFAULT '',
|
||||
created_at TEXT NOT NULL,
|
||||
metadata_json TEXT NOT NULL DEFAULT '{}',
|
||||
origin_replica TEXT,
|
||||
origin_seq INTEGER,
|
||||
hlc TEXT
|
||||
);
|
||||
CREATE UNIQUE INDEX IF NOT EXISTS events_origin_seq_idx ON events(origin_replica, origin_seq);
|
||||
```
|
||||
(`logstream.py:447-497`, unique index at `552-556`. `artifacts` and `event_artifacts` in the same script; `artifacts_sha256_idx` is **not** unique — see §3.4.)
|
||||
|
||||
Validation, all server-side, all raising `ValueError` naming the allowed set:
|
||||
|
||||
| Field | Required | Rule | Where |
|
||||
|---|---|---|---|
|
||||
| `type` | yes | `^[a-z0-9][a-z0-9_.-]{0,63}$` — **lowercase only, ≤64** | `logstream.py:76, 124-134` |
|
||||
| `stream`, `room`, `from_agent` | yes | non-empty, ≤256, no control chars | `logstream.py:78, 103-122` |
|
||||
| `status` | no | one of `open claimed ready applied blocked failed superseded` | `logstream.py:67-69, 136-143` |
|
||||
| `body` | no (empty allowed) | ≤256 KiB, no NUL | `logstream.py:55, 145-158` |
|
||||
| `metadata` | no | canonical JSON, **≤64 KiB** | `logstream.py:57, 160-177` |
|
||||
| `artifact_ids` | no | each must already exist, else `ValueError` | `logstream.py:687-694` |
|
||||
|
||||
Three consequences worth stating because they are invisible from the tool descriptions: an event body **may be empty** while artifact content **may not** (`logstream.py:766-767`); `to_agent` is **nullable**, and an event with `to_agent = NULL` is addressed to nobody and can never match a `to_agent=` query; and the 64 KiB `metadata` cap and the lowercase-`type` regex are **enforced but undocumented** in the shipped tool schemas.
|
||||
|
||||
### 3.2 Three orderings, and one identity that is not what it looks like
|
||||
|
||||
The log carries three distinct notions of order, and conflating them is the single most likely design error in any consumer:
|
||||
|
||||
| Field | Meaning | Scope | Where |
|
||||
|---|---|---|---|
|
||||
| `seq` | **local arrival cursor** — the SQLite `rowid`, surfaced at read time | this replica only | `logstream.py:580-586` |
|
||||
| `origin_seq` | the author replica's own gap-free counter, assigned by `UPDATE … SET origin_seq = rowid` in the insert transaction | per author | `logstream.py:700-706` |
|
||||
| `hlc` | hybrid logical clock, `<unix_ms>-<counter_hex>-<replica_id>`, fixed-width so TEXT comparison *is* causal comparison | fleet-wide | `logstream.py:668`, `hlc.py:1-21` |
|
||||
|
||||
The rule, from `hlc.py:19-21`: *"Cursor semantics stay LOCAL (rowid arrival order) … a tail consumer must see late-arriving remote ops even though their HLC is older. HLC is the display/merge order; arrival is the delivery order."*
|
||||
|
||||
**`origin_replica` identifies the palace, not the writer.** It comes from `get_replica_id(db_parent)` — the `replica.json` in the palace directory (`logstream.py:422-436`, `replica.py:31-49`). In a hub-and-spoke deployment every client writes through the server's single `Logstream`, so **every event from every machine carries the same `origin_replica`** and it cannot distinguish two devices. This is mechanically forced by the design, not a misconfiguration.
|
||||
|
||||
Therefore: **`from_agent` / `to_agent` carry the entire distinction between machines.** The convention that makes them able to is `<harness>@<device>` — e.g. `pi@laptop`, `opencode@build-box` — stamped at the client edge, which RFC 001 §7.3.5 specifies. An unstamped client addressed as bare `pi` is unreachable in a fleet, because nobody can name it.
|
||||
|
||||
### 3.3 The ack contract: owed-ness is derived, never read
|
||||
|
||||
`status` is written **once**, into an append-only row. Nothing updates it. So `status="open"` means *"the sender declared this an ask at the moment of writing"* — it does **not** mean unanswered, and a directed `open` event keeps matching a mailbox query forever, answered or not.
|
||||
|
||||
`ack_event` (`logstream.py:812-845`) never touches the target row. It *appends*:
|
||||
|
||||
```python
|
||||
return self.append_event(
|
||||
type=ACK_EVENT_TYPE, # "event.ack"
|
||||
stream=target["stream"], room=target["room"],
|
||||
from_agent=from_agent,
|
||||
to_agent=target["from_agent"], # routes back to the sender
|
||||
correlation_id=target["correlation_id"] or target["id"], # falls back to the target id
|
||||
status=status, body=body,
|
||||
metadata={"ack_of": event_id}, # the join key lives in metadata
|
||||
)
|
||||
```
|
||||
|
||||
So a consumer that wants "what do I still owe?" must derive it. The derivation the toolkit implements, and which this RFC adopts as the reference semantics:
|
||||
|
||||
```mermaid
|
||||
stateDiagram-v2
|
||||
[*] --> open : task.request written, addressed to you
|
||||
open --> claimed : you announce you are working (does NOT clear)
|
||||
open --> ready : work exists, thread still open (does NOT clear)
|
||||
claimed --> applied : terminal, clears
|
||||
ready --> applied : terminal, clears
|
||||
open --> blocked : terminal, clears (with a reason)
|
||||
open --> failed : terminal, clears (with a reason)
|
||||
open --> superseded : terminal, clears (replaced by another thread)
|
||||
applied --> [*]
|
||||
blocked --> [*]
|
||||
failed --> [*]
|
||||
superseded --> [*]
|
||||
```
|
||||
|
||||
An event is **still owed** when all of these hold:
|
||||
|
||||
1. it is directed at you (`to_agent = <you>`, and `to_agent = '*'` is excluded — §7.4);
|
||||
2. it was not written by you (you cannot owe yourself);
|
||||
3. **no** event of yours has *all* of: a strictly higher `seq`, a join to it (`metadata.ack_of = <id>` or a shared `correlation_id`), **and** a terminal status (`applied` / `superseded` / `failed` / `blocked`);
|
||||
4. **the requester has not explicitly withdrawn it** — see §3.3.1.
|
||||
|
||||
The strictly-higher-`seq` test is load-bearing: without it, one terminal reply on a `correlation_id` suppresses every *later* ask on that same correlation, permanently. And because `seq` is local arrival order, this test is only sound on a single replica — §7.3.
|
||||
|
||||
#### 3.3.1 Withdrawal: the one case where someone else may clear your obligation
|
||||
|
||||
> **Status: implemented in the toolkit, not yet in a released image.** Clause 4 is
|
||||
> live in `extensions/pi/mempalace.ts` (`isWithdrawn`) with rule tests in
|
||||
> `scripts/test-owed-withdrawal.sh`. It reaches a device only when an image bakes
|
||||
> a toolkit revision containing it. Until then clauses 1–3 are the whole rule, and
|
||||
> **a sender must assume its withdrawal has no effect on the recipient's mailbox.**
|
||||
|
||||
Clause 3 says "no event **of yours**", and that asymmetry is deliberate: owed-ness
|
||||
is a statement about the *recipient's* accountability, so a requester must not be
|
||||
able to delete an obligation the recipient genuinely has. It is wrong in exactly
|
||||
one case — the requester retracting its **own** ask.
|
||||
|
||||
**Measured, 2026-09-08/09.** `pi@mbp-m1-2020` withdrew a v1.8.13 rollout ask to
|
||||
`pi@tor-ms22`: terminal `superseded`, same `correlation_id`, `metadata.closes`
|
||||
naming the thread, body "DO NOT SPEND A MINUTE ON v1.8.13". It recorded the
|
||||
withdrawal as done. The withdrawal had **no effect** — its `from_agent` is mbp, so
|
||||
it could never satisfy a join that only inspects tor-ms22's own events. tor-ms22's
|
||||
next wake-up, 8h later and on the first boot of the image shipping this very
|
||||
derivation, still listed the ask as owed, 41h old, for a release that had been
|
||||
superseded and never installed there. Closing it cost the recipient a write on a
|
||||
thread nobody wanted answered. **The failure is invisible from the sender's side:**
|
||||
mbp did everything a sender is told to do and got a result indistinguishable from
|
||||
success, which is why this went unnoticed for 41h rather than being caught at once.
|
||||
|
||||
A withdrawal is honoured only when **all** of these hold. Each is a guard against a
|
||||
specific way a mailbox could otherwise be emptied silently, and each is pinned by a
|
||||
mutation test:
|
||||
|
||||
| Requirement | The failure it prevents |
|
||||
|---|---|
|
||||
| `from_agent` = the **original requester** | a third party clearing someone else's obligation |
|
||||
| `to_agent` = the recipient **exactly** (not `'*'`) | one broadcast emptying every machine's mailbox at once (§7.4) |
|
||||
| terminal status | a `claimed`/`ready` note reading as a release |
|
||||
| strictly after the ask (`hlc`, §7.3) | retiring an ask the requester sent **later** on the same correlation |
|
||||
| joins the ask (`ack_of` or shared `correlation_id`) | clearing an unrelated thread |
|
||||
| **an explicit marker**: `metadata.withdraws` **or** `metadata.closes`, whose value is exactly the ask's `correlation_id` or its event `id` | **the dangerous one** — see below |
|
||||
|
||||
**Why release must be stated rather than inferred.** The tempting rule is "any
|
||||
terminal event from the requester clears it". That breaches this log's standing
|
||||
safety direction, which is that every failure stays on the *noisy-but-visible*
|
||||
side: a spuriously resurfacing item costs one turn of human correction, while a
|
||||
suppressed unanswered ask is silent and permanent. Under the naive rule a
|
||||
requester appending `applied` for its own bookkeeping — on the correlation,
|
||||
addressed to the recipient, before the recipient ever replied — would silently
|
||||
delete a real obligation. `withdraws` is canonical; `closes` is honoured because it
|
||||
is already the fleet's de-facto marker (mbp seq 119; emb seq 117 and 118 all use
|
||||
it). Prose does not count: the value must *name* the thread, so
|
||||
`closes: "v1813-client-rollout-tor-ms22"` is a withdrawal and
|
||||
`closes: "v1813-… — from the RECIPIENT side, which is the only side that can"` is
|
||||
not.
|
||||
|
||||
**This does not break the fixed point in §9.2, and a reader will reach for §9.2 to
|
||||
object.** That section rejects letting terminal directed events into the owed set,
|
||||
because then a reply becomes owed by its requester, closing it mints a fresh
|
||||
obligation, and the loop never terminates. That argument is about **candidates**.
|
||||
Clause 4 adds a **clearer**. A withdrawal carries a terminal status, so it can never
|
||||
be a candidate, and nothing new becomes owed: the asserting shape (`open`) and the
|
||||
clearing shape (terminal) stay disjoint, which is what makes owed-ness terminate.
|
||||
|
||||
**Recipients keep the last word.** A withdrawal suppresses; it does not rewrite
|
||||
history. The recipient may still append its own terminal reply — worth doing when
|
||||
there is a finding to carry across, as tor-ms22 did at seq 122, salvaging two
|
||||
measurements from the retracted thread.
|
||||
|
||||
### 3.4 Artifacts
|
||||
|
||||
Exact payloads for handoff: `patch`, `file`, `log`, `json`, `note`; UTF-8 text only; ≤4 MiB; `sha256` and `size_bytes` returned (`logstream.py:744-810`).
|
||||
|
||||
⚠️ **The sha256 is for verification, not deduplication.** `artifacts_sha256_idx` is a plain index and `put_artifact` always INSERTs, so storing the same patch twice stores the content twice. Referential integrity in the other direction *is* enforced: `artifact_ids` naming an unknown artifact raises at append time, in both the local and the remote path (`logstream.py:687-694`, `1181-1189`), *"so readers never see a dangling reference"*.
|
||||
|
||||
`kind="patch"` gets advisory warnings — never mutation — for a missing trailing newline or CRLF endings (`logstream.py:718-742`). The comment records why: *"the first RFC 003 dogfood: a patch stored without its final trailing newline truncates the last hunk line."*
|
||||
|
||||
### 3.5 Delivery: three mechanisms, one client feature
|
||||
|
||||
| Mechanism | Shape | Latency | Where |
|
||||
|---|---|---|---|
|
||||
| `event_list` | pull, filtered, `ORDER BY rowid ASC`, default 50 / max 500 (silently clamped) | whenever you ask | `logstream.py:914-990` |
|
||||
| `event_wait` | in-request long poll, jittered backoff 0.25 s→1 s, default 60 s, hard cap 300 s, returns `{timed_out: true, events: []}` rather than raising | seconds, while you hold the call | `logstream.py:1004-1039` |
|
||||
| `GET /logstream/stream` | SSE; resume by `?since_event_id=` or `Last-Event-ID`; `: ping` every 15 s; ≤8 concurrent clients then `503` + `Retry-After` | sub-second, while connected | `mcp_server.py:7570-7671` |
|
||||
|
||||
None of these push anything into an agent that is not already running. That gap is closed **client-side** by the toolkit's *mailbox*: it derives the owed set per §3.3 and injects it at two points — once in the session-start wake-up, and mid-session on a settled-agent poll floored at 5 minutes, delivered as a queued steer rather than an interrupt.
|
||||
|
||||
**Measured 2026-08-26** (first live delivery on a published image, positive control with a synthetic sender): planted → delivered in **≈2–3 minutes** to a session that was already running, as exactly one item; three already-answered asks were correctly excluded and a `*` broadcast was correctly ignored. The wake-up path and the resurface path were **not** exercised by that test and remain unmeasured.
|
||||
|
||||
This is why the mailbox is a client concern and belongs in the toolkit rather than here: the server has no idea which agents exist, and no way to reach one that is not calling it.
|
||||
|
||||
---
|
||||
|
||||
## 4. Availability: the logstream is deliberately exempt from both palace locks
|
||||
|
||||
The rest of the palace is a single-writer service — `_HTTP_REQUEST_LOCK` wraps every dispatch (RFC 001 §7.5), and a Chroma peer-writer lease protects against a concurrent CLI mine. **Coordination traffic is exempt from both**, in two separate places, for two different reasons:
|
||||
|
||||
- `_HTTP_LOCK_FREE_TOOLS` (`mcp_server.py:6887-6902`) — all seven event/artifact tools dispatch outside the request lock: *"Dispatching them outside `_HTTP_REQUEST_LOCK` keeps a five-minute `mempalace_event_wait` long-poll from stalling every other agent on a shared hub, and lets the SSE stream coexist with normal tool traffic."*
|
||||
- `_PEER_WRITER_EXEMPT_TOOLS` (`mcp_server.py:434-449`) — the four mutating ones bypass the writer lease: *"Exempting them keeps agent coordination alive while a CLI mine or a peer stdio writer holds the palace lock."*
|
||||
|
||||
They remain in `_MUTATING_TOOLS`, so `--read-only` still refuses them.
|
||||
|
||||
✅ **Consequence, and it corrects a widely-held belief in this fleet:** "one large mine blocks every client for minutes" is true of drawer writes and **false of coordination writes**. You can message another device, and it can reply, while a mine is running.
|
||||
|
||||
Concurrency within the log is a per-instance `threading.Lock` over a WAL database with a 10 s busy timeout, one cached `Logstream` per palace path per process (`logstream.py:437`, `mcp_server.py:333-334, 1120-1141`).
|
||||
|
||||
---
|
||||
|
||||
## 5. Not everything belongs in the log
|
||||
|
||||
RFC 001 §5 draws this line for wings; the same discipline applies between the two stores. The short form, expanded for operators in `docs/fleet-memory.md`:
|
||||
|
||||
| Put it in | When |
|
||||
|---|---|
|
||||
| A **drawer** (palace) | It will still be true, and worth finding, next month. Nobody in particular needs to act. Retrieval is by meaning. |
|
||||
| An **event** (log) | A named agent must act, reply, or be stopped. Retrieval is by address and order. |
|
||||
| **Both** | The durable finding goes in a drawer; the event says "there is a new finding, here is the drawer id". This is the recommended pattern for anything a peer must *know* rather than *do*. |
|
||||
|
||||
The failure mode in each direction: a finding filed only as an event is invisible to semantic search and will be re-derived by the next agent; an ask filed only as a drawer is addressed to nobody and will be found, if ever, by accident.
|
||||
|
||||
---
|
||||
|
||||
## 6. Security model
|
||||
|
||||
**The log authenticates the fleet, not the agent.** Precisely:
|
||||
|
||||
| Control | State | Evidence |
|
||||
|---|---|---|
|
||||
| Transport auth | One shared bearer token, `hmac.compare_digest`, 401 otherwise; required for non-loopback binds | `mcp_server.py:7137-7141, 7685` |
|
||||
| `from_agent` authenticity | ❌ **Not checked against anything.** Shape-validated only; any client may append an event claiming to be any agent | `logstream.py:103-122` is the only check |
|
||||
| Read authorization | ❌ **None.** Any token holder may list every event addressed to anyone, and fetch any artifact by id — including patch contents | no scoping in the query path |
|
||||
| Stream auth | ✅ SSE follows the same bearer policy: *"Events and artifacts expose work metadata and patch contents"* | `mcp_server.py:7189-7194` |
|
||||
| Transport security | TLS via `--tls-cert`/`--tls-key` (both or neither); Host pin against DNS rebinding; Origin check; 16 MiB request cap | `mcp_server.py:6920-6947` |
|
||||
|
||||
For a single-operator fleet behind one token this is adequate, and it must be stated rather than implied, because two useful consequences follow directly from it:
|
||||
|
||||
1. **Impersonation is trivial** — do not treat `from_agent` as evidence of origin in any security decision.
|
||||
2. **That same property is the only way to test the mailbox.** A positive control requires writing an event *from* an agent you are not, so that your own client does not self-filter it. This is a supported technique precisely because the field is unauthenticated (**measured 2026-08-26**).
|
||||
|
||||
If per-agent identity is ever needed, the correct home is the same place RFC 001 §7.3 puts provenance: the authenticated credential at the boundary, not a self-asserted field.
|
||||
|
||||
---
|
||||
|
||||
## 7. Landmines
|
||||
|
||||
### 7.1 There is no idempotency guard on append — a retried write duplicates
|
||||
|
||||
`append_event` mints a fresh id and INSERTs, with no dedup lookup of any kind (`logstream.py:625-716`). `put_artifact` likewise. Two byte-identical calls produce two events with different ids and different `seq`.
|
||||
|
||||
The asymmetry is stark: the *replication* path is rigorously idempotent (`logstream.py:1174-1180` checks `id` **or** `(origin_replica, origin_seq)` before applying, and returns early for its own echoed ops). Peer replay is safe; **client retry is not.**
|
||||
|
||||
**Action:** this makes the palace's general rule — *a timeout usually means the write completed; verify, don't retry* — load-bearing rather than advisory. For a drawer, a blind retry costs a dedup-detectable duplicate. For an event it silently forks a coordination thread into two ids, and a mailbox will then show two owed items that must each be closed. After any `event_append` / `patch_submit` timeout, verify with `event_list(correlation_id=…)` before re-issuing.
|
||||
|
||||
### 7.2 `status="open"` is not owed-ness
|
||||
|
||||
Covered in §3.3 and repeated here because it is the mistake most likely to be made by someone reading only the tool descriptions: the mailbox query `to_agent=<me> status=open` **never shrinks as you work**. Treating its length as a to-do count means re-answering answered asks forever.
|
||||
|
||||
**Action:** derive per §3.3, or use a client that does.
|
||||
|
||||
### 7.3 `seq` is local arrival order and is meaningless across replicas
|
||||
|
||||
`seq` is the local `rowid`. The moment a second replica exists, a remote event that was *authored* earlier can arrive *later* and receive a higher `seq`. An owed-set derivation that compares `seq` is therefore sound only on a single replica: on a mesh, a reply can land before the ask it answers, the join concludes "no later reply exists", and an already-answered ask reappears as owed — permanently, on that machine.
|
||||
|
||||
**✅ Resolved 2026-08-26 (toolkit `extensions/pi/mempalace.ts`).** The derivation now compares `hlc` when both events carry one, falling back to `seq` only when either lacks it (a server predating the field, or an un-backfilled row). `hlc` is rendered fixed-width — `<unix_ms:13 digits>-<counter:6 hex>-<replica_id>` (`hlc.py:1-21`) — so a plain string comparison *is* the causal comparison, with the replica id as final tiebreak. The field is present on the wire: `event_list` returns it per event (`logstream.py:593`), verified against the live hub the same day.
|
||||
|
||||
Measured on real data: the positive-control pair carries `seq` 26/27 and `hlc` `1787773071857-…`/`1787773574085-…`, so on today's single replica the two orderings agree and the switch is a **no-op now and correct later** — which is the whole reason to make it before a second replica exists rather than after. `created_at` remains the wrong key in both worlds: server-generated at second precision, so ties are routine and a tie can suppress an *unanswered* ask outright.
|
||||
|
||||
### 7.4 A broadcast reaches no mailbox, and `to_agent=NULL` reaches nobody at all
|
||||
|
||||
`to_agent=<x>` matches `x` **or** `'*'` at the SQL level (`logstream.py:958-960`), so a broadcast *is* visible to a listing agent. But the reference owed-set derivation excludes `to_agent='*'` deliberately — a broadcast owes nobody a reply, and if it entered every mailbox, every machine would think it personally owed the same answer.
|
||||
|
||||
Two consequences: **broadcasting an ask reaches no owed set at all** (the "don't broadcast an ask" anti-pattern is mechanically enforced, not merely advised), and to reach a whole fleet with something actionable you must write **one directed event per device**, sharing a `correlation_id` so the thread stays joinable. Separately, an event written with **no** `to_agent` matches no `to_agent=` query ever and is addressed to nobody.
|
||||
|
||||
**✅ Measured 2026-08-26, on two devices independently.** This was inferred from source until a two-arm control settled it. A `to_agent='*'` event with `status='open'` — the anti-pattern, planted deliberately — survives the raw candidate filter and so reaches the exclusion branch, which every previously written broadcast (all non-open) never did. Result: the raw query returned it, the derived owed set did not, and the delivery named only the directed arm of the pair. Confirmed on a second machine that had merely *received* the broadcast rather than planted it. Two arms differing only in `to_agent`, same poll and same fake sender, is what makes it evidence rather than an absence: "delivered exactly one of two" cannot be explained by a mailbox that delivers everything or nothing.
|
||||
|
||||
> ⚠️ **Correction to an earlier record.** A note filed the same day claimed this branch *could not* be exercised, reasoning that every broadcast in the log carries `status='ready'` and the protocol forbids writing an open broadcast. The reasoning was wrong in a specific and instructive way: the protocol forbids it *as production traffic*, which does not forbid planting one as a labelled, self-closed control. Where a rule blocks a measurement, check whether it blocks the *use* or the *test* before recording the branch as untestable.
|
||||
|
||||
### 7.5 `since_event_id` raises on an unknown id
|
||||
|
||||
An unresolvable cursor raises `ValueError` rather than returning empty (`logstream.py:966-970`) — stronger than the tool description promises, and benign on one replica. On a second replica, a resuming watcher whose cursor has not yet replicated will **hard-fail instead of waiting**.
|
||||
|
||||
### 7.6 A rejected append returns HTTP 200
|
||||
|
||||
Failures come back as `{"success": false, "error": …}` in a 200 response, not a JSON-RPC error (`mcp_server.py:4581-4584`).
|
||||
|
||||
**Action:** a client that checks only transport status silently loses the event. Check the payload.
|
||||
|
||||
### 7.7 The same agent name on two devices is indistinguishable
|
||||
|
||||
No uniqueness, no registry, no warning. Both machines' events carry one `from_agent`, and each machine's "you cannot owe yourself" filter will discard the other's asks. This is the failure the `<harness>@<device>` convention exists to prevent (§3.2).
|
||||
|
||||
### 7.8 `GET /logstream/events` does not exist
|
||||
|
||||
The complete GET route table is `/healthz`, `/statusz`, `/logstream/stream`, `/sync/{version_vector,ops,artifact,peers}`; POST serves `/mcp` only. A recursive grep for `/logstream` in the package yields three hits, all `/logstream/stream`.
|
||||
|
||||
⚠️ **Correction.** An earlier measurement observed `/logstream/events`, `/logstream/stream` and `/sync/peers` all returning 404 against a deployment and attributed all three to a reverse proxy exposing only `/mcp`. That inference was right for `/sync/*` and `/logstream/stream` and **wrong for `/logstream/events`**, which would 404 on a directly-reachable server too. Distinguish "route absent" from "route blocked" before blaming infrastructure.
|
||||
|
||||
### 7.9 Dogfood scars, preserved because each cost someone a session
|
||||
|
||||
- A `patch` artifact stored without its trailing newline truncates the last hunk line — hence the advisory warnings (`logstream.py:718-742`).
|
||||
- `--type Task.Request` was rejected while `--type Task.Request --type patch.ready` was *silently accepted and matched nothing*, leaving a watcher waiting forever: single-valued filters were validated by pushdown, multi-valued ones compared raw (`sanitize_watch_spec`, `logstream.py:245-278`). `type` is lowercase-only for this reason.
|
||||
- `event_wait` rejected a `limit` that `event_list` accepted — *"reported by windows-codex during dogfood"* (`mcp_server.py:4666-4670`).
|
||||
|
||||
### 7.10 The main database file's mtime is not a liveness signal
|
||||
|
||||
WAL. **Measured 2026-08-26** on the primary: `logstream.sqlite3` mtime three days old while `logstream.sqlite3-wal` was 865 KiB and seconds old. An operator checking whether coordination is live must look at the `-wal` file, or query.
|
||||
|
||||
### 7.11 Delivered is not read: the human is the trigger
|
||||
|
||||
The mid-session poll delivers with `deliverAs: "steer"` and **deliberately no `triggerTurn`** — waking a model on inbound fleet traffic is a much larger behavioural change than auto-delivery, and was not approved. The poll fires on `agent_settled`, which means the agent is *idle*: nothing is running to react. So the delivered text is queued and read at the top of the **next turn**, whenever a human happens to start one.
|
||||
|
||||
**Measured 2026-08-26:** a delivery landed in a session window at ~20:50Z and sat there, visibly unreacted-to, until the operator asked "do I have to nudge you for you to read it?" — which is precisely the question this design produces. The delay is not the agent choosing to ignore its mailbox; between the poll and the next turn there is no inference at all.
|
||||
|
||||
The consequence is a **UX gap, not a defect**: from the outside, an inbound ask plus an idle agent reads as neglect. The delivered text explained how to *close* an ask but said nothing about *when* it would be seen — so the one reader who needed that fact, the human watching the window, was the one not told.
|
||||
|
||||
**✅ Addressed 2026-08-26 (toolkit `extensions/pi/mempalace.ts`), without touching the no-`triggerTurn` decision.** Two changes, in the order their value is realised:
|
||||
|
||||
1. **A delivery note in the message itself** — it says this is a queued message, that nothing woke the agent, and that any message starts the turn that will handle it. Costs nothing, changes no behaviour, and converts "why is it ignoring me?" into "right, I nudge it."
|
||||
2. **A notification at poll time**, because the note only reaches someone already looking, and the case that actually loses an ask is nobody looking. `MEMPALACE_MAILBOX_NOTIFY` unset → in-TUI `ctx.ui.notify` (the surface `session_start` already uses); `=desktop` → additionally a terminal-native notification (Kitty `OSC 99`, else `OSC 777`), which is how a ping escapes a container with no `notify-send`, DBus or host access — the escape sequence is interpreted by the terminal emulator on the human's machine; `=0` → silent. Title and body are stripped of `;` and control bytes so a payload cannot forge an OSC field or terminate the sequence early. It fires only when something is due, because a ping on an empty poll trains its reader to ignore it.
|
||||
|
||||
**⚠️ Unverified through a multiplexer — measured 2026-08-26, tor-ms22.** The terminal-native path is written and syntax-checked but has **not** been observed to fire. On this client the layering is `kitty → tmux (on the host) → docker exec → pi (in the container)`, with two clients attached to the same tmux session (one local, one remote over SSH). Test sequences written directly to pi's tty produced **no notification on the remote client**; the local client is unobserved so far. The likely mechanism is that **tmux drops OSC sequences it does not itself implement** — reaching the outer terminal requires wrapping them in tmux's DCS passthrough (`ESC P tmux; … ESC \`, with inner `ESC` bytes doubled) *and* `allow-passthrough on`, which is not the default.
|
||||
|
||||
This compounds §7.11's detection problem rather than repeating it: **the container cannot see either layer**. `KITTY_WINDOW_ID` is absent because `docker exec` does not forward it, and `TMUX` is absent because tmux is running one level further out, on the host — so the client cannot detect the terminal it is speaking to *or* the multiplexer it must speak through. Explicit configuration is the only reliable route; autodetection is structurally blind here.
|
||||
|
||||
Open, and deliberately not settled by guessing: whether the local client shows it (pending); whether to add DCS-wrapping modes; and with two clients attached, **which** client should be notified — a pane's output reaches every attached client, so a work laptop and a home machine would both ping.
|
||||
|
||||
Still deliberately **not** done: `MEMPALACE_MAILBOX_TRIGGER=1` (opt-in, default off) to start a turn on arrival. Worth revisiting once notification is proven in daily use — but a machine that auto-turns on inbound events writes to a shared log with nobody watching, and §7.1 (no idempotency guard) is the reason to be slow about it.
|
||||
|
||||
### 7.12 The mailbox is an obligation channel, so a report addressed to you is never delivered
|
||||
|
||||
Candidates are drawn with `status="open"` (§3.3), so an event carrying a **terminal** status is not a mailbox candidate at all — no matter who it is addressed to. A `task.reply` sent to a named device to share a finding is therefore delivered to nobody, ever, and neither is any `event.ack`. It sits in the log until someone reads the log.
|
||||
|
||||
Read the filter precisely: candidacy requires **exactly `open`**, not merely "non-terminal". So a `claimed` announcement is equally undelivered — announcing that you have started work reaches the requester's mailbox no more than your finished report does. Claiming is still worth doing on long or handoff-prone work, because it puts *pickup* in the log for whoever later asks "did anyone start this before that container died?", but its audience is a log reader, not the requester. Note also that claiming does not quiet **your own** mailbox: the original ask stays owed until a terminal event of yours joins it (§3.3), which is exactly what the state machine means by `open --> claimed` *does NOT clear*.
|
||||
|
||||
**Measured 2026-08-26:** a device wrote a detailed report addressed to `pi@<peer>` with `status="applied"`, on a thread whose ask belonged to a third party. The recipient never saw it and only read it when the operator pointed at it by id — with the mailbox working exactly as specified throughout.
|
||||
|
||||
This is the shape of the channel, and it is worth stating because the natural inter-device message — *"here is something you should know"* — is exactly the shape that gets no delivery. Two ways to make news reach a peer: address it as a **directed `status="open"` ask** so it enters the owed set and gets closed when acted on (correct when a response is genuinely wanted), or accept that it is **pull-only** and pair it with a palace drawer, which the peer's search *will* surface later. What does not work is a terminal-status report and an expectation of attention.
|
||||
|
||||
### 7.13 An event addressed to an identity no session runs as is write-only
|
||||
|
||||
The owed set (§3.3) is the log's only push channel — it is what a client injects at wake-up. It keys on `to_agent`. A `from_agent` is whatever string the writer puts there (§6: "authenticates the fleet, not the agent"), so nothing stops a writer from authoring under an identity that no live session has ever run as. When it later receives a reply, that reply is addressed right back to the string the writer chose — and if nothing runs as that string, the reply is stored, queryable, and delivered to no one. Unlike §7.12, this is not about status; a **directed, `status="open"`** reply is undeliverable here, because the *address* rather than the *shape* is the defect.
|
||||
|
||||
**Measured 2026-08-27.** A device planted two directed `task.request`s under a synthetic `from_agent` (a release-provenance label, not an identity any session runs as), intending only to record where the ask came from. Both recipients replied correctly, addressed to that synthetic sender as the protocol requires — one an `event.ack` (`status="claimed"`), one a `task.reply` (`status="applied"`). Neither reply was ever delivered to anyone. One contained a live-credential exposure finding. It sat unread for roughly two and a half hours and was found only because a human asked whether mail had arrived.
|
||||
|
||||
**The general rule, not the instance:** before addressing an event, the writer is responsible for the address being an identity a live session actually runs as; and a writer that authors under an identity other than its own session identity has made every reply to that event undeliverable to itself. This fleet has already produced the failure in three unrelated shapes — a synthetic sender (above); an event addressed to a device since decommissioned; and a device-identity case mismatch (an uppercase form written where the fleet's convention is lowercase) — which is why this is recorded as a property of the channel rather than a mistake to remember not to repeat.
|
||||
|
||||
**The permitted exception, and its obligation.** Authoring under a synthetic identity is legitimate for controlled experiments — this fleet has run positive/negative mailbox controls this way deliberately — and this RFC does not forbid it. But a synthetic sender then **must name the real identity to reply to** in the body; the address line is not a place to record provenance if a reply is ever wanted.
|
||||
|
||||
**Proposed, not implemented:** on `event_append`, if `from_agent` differs from the writer's own stamped session identity, warn that replies to this event will be addressed to that other identity, and say whether any event in the log has ever been authored under it — a plain `DISTINCT from_agent` lookup. State the heuristic's limit honestly: an identity that has never authored is certainly unread by anyone; one that *has* authored may still belong to a device that no longer exists, which this check cannot see.
|
||||
|
||||
---
|
||||
|
||||
## 8. Phasing
|
||||
|
||||
### 8.1 Implemented
|
||||
|
||||
| Phase | Deliverable | State |
|
||||
|---|---|---|
|
||||
| A | `events`/`artifacts`/`event_artifacts` + append/list/ack + size limits | ✅ 3.7.x |
|
||||
| B | Long poll (`event_wait`) and SSE push | ✅ |
|
||||
| C | `patch_submit` convenience + patch advisories | ✅ |
|
||||
| D | HLC on every event; `(origin_replica, origin_seq)` unique index | ✅ 3.8.0 |
|
||||
| E | Client-side auto-delivered mailbox | ✅ toolkit `5b8d78f`; first live delivery measured 2026-08-26 |
|
||||
|
||||
### 8.2 Deferred, and what RFC 004 owes
|
||||
|
||||
Replication exists in code — `version_vector()`, `list_ops()`, `apply_remote_event()`, `apply_remote_artifact()` and the four `GET /sync/*` routes (`logstream.py:1096-1261`, `mcp_server.py:7206-7256`) — and is attributed in comments to "RFC 004 step 0". **That document does not exist.** Until it does, this RFC records the two properties a reader most needs: the apply path *is* idempotent (§7.1), and switching the owed-set derivation from `seq` to `hlc` is a precondition for a second replica, not a follow-up (§7.3).
|
||||
|
||||
---
|
||||
|
||||
## 9. Open decisions
|
||||
|
||||
1. **Retention.** No TTL, compaction or pruning exists, and `mempalace sync` does not touch the log (§2). For a fleet log this is mostly a feature — nothing is lost by being offline for weeks — but every 4 MiB artifact is permanent. Decide a policy before the log outgrows a comfortable backup, or decide explicitly that permanence is the policy.
|
||||
2. **Idempotency key.** Should `event_append` accept an optional client-supplied dedupe key so a retried call is a no-op? This is the one change that would make §7.1 disappear.
|
||||
3. ~~**`hlc` as the mailbox join key**, replacing `seq`.~~ **Done 2026-08-26** — see §7.3. What remains open is the *reverse* direction: `mempalace_event_list`'s `since_event_id` cursor is deliberately local-arrival-ordered (`hlc.py:19-21` — "a tail consumer must see late-arriving remote ops even though their HLC is older"), so cursor semantics and join semantics use different orderings on purpose. That is correct, and worth stating loudly before someone "fixes" the cursor to match the join.
|
||||
4. **Per-agent identity.** Do we ever want `from_agent` to be authenticated (§6), or is "authenticates the fleet, not the agent" the permanent contract?
|
||||
5. **Should `GET /logstream/events` exist?** A read-only HTTP tail would let non-MCP consumers (dashboards, CI) follow a stream without an MCP client. Today they must hold SSE or speak MCP.
|
||||
6. **Undocumented limits.** The 64 KiB `metadata` cap and the lowercase-`type` regex are enforced server-side but absent from the tool schemas (§3.1). Document, or relax.
|
||||
7. **Agent-name registry.** Nothing prevents two devices sharing one `from_agent` (§7.7). A warning at append time would be cheap.
|
||||
8. ~~**Does an arriving ask deserve a notification?**~~ **Done 2026-08-26** — delivery note plus `MEMPALACE_MAILBOX_NOTIFY` (§7.11). What stays open is the narrower question: should `desktop` be the *default* rather than opt-in, and does an arriving ask ever deserve a turn (`MEMPALACE_MAILBOX_TRIGGER`)?
|
||||
9. **Is there a delivery path for news rather than obligations?** (§7.12) Today the only delivered shape is a directed open ask. Options: leave it pull-only and rely on the paired drawer, or give informational events a distinct low-priority surface at wake-up ("3 reports since your last session") separate from the owed set. **Proposed direction in §9.2** — including why the obvious fix (widening the owed set) is not available.
|
||||
10. **Should a write warn when `from_agent` is not the writer's own identity?** (§7.13) Related to decision 7 but distinct: decision 7 is about two devices *sharing* one identity by accident; this is about a writer deliberately authoring under an identity nobody runs, which makes every reply to that event undeliverable to itself. A `DISTINCT from_agent` lookup at append time is cheap and catches the unread-for-certain case; it cannot catch a decommissioned device that once wrote, which is the harder half of the same problem.
|
||||
|
||||
### 9.1 Retention — decided direction (2026-08-26)
|
||||
|
||||
Decision 1 above is **settled in direction**: old traffic should eventually move out of the way, `logrotate`-style. Moved aside, not necessarily destroyed — the audit trail of who asked whom for what is worth more than the disk it occupies.
|
||||
|
||||
The economics point at a two-tier design rather than one sweep. Events are tiny (a body cap of 256 KiB, and in practice a few KB); artifacts are capped at **4 MiB each** and are stored as content in-row. So:
|
||||
|
||||
- **Tier the artifacts first.** Keep every `events` row indefinitely — they are the index and the audit trail — and move artifact *content* to cold storage past the window, keeping the `kind`, `sha256`, `size_bytes` and `created_by` stub in place. A stub still answers "what was handed over, by whom, verified how"; only the bytes go cold. This reclaims nearly all of the space with none of the join risk below.
|
||||
- **Rotate events by thread, never by row.** If rotation is wanted for events too, the unit must be the **`correlation_id` thread**, and only a thread whose latest event is terminal (§3.3) and older than the window. Archiving an *ask* while leaving its *reply* behind — or the reverse — breaks the owed-set join, and the two failure modes are both bad: an ask that can never be cleared resurfaces as owed forever, or a reply is orphaned from what it answered.
|
||||
|
||||
Three constraints any implementation has to respect:
|
||||
|
||||
1. **Never rotate an event that is still owed.** The owed set is derived at read time (§3.3), so an unanswered ask has no marker distinguishing it from a stale one except that derivation. A time-based sweep alone would silently discard live obligations from a machine that has simply been offline for a month — exactly the case this log exists to serve.
|
||||
2. **Rotation invalidates held cursors.** `since_event_id` raises on an unknown id rather than returning empty (§7.5), so archiving an event that a watcher still holds as its resume point turns that watcher's next poll into an error. Either rotation is announced far enough ahead of any live cursor, or the anchor lookup learns to fall back to `created_at`/`hlc` when the id is gone.
|
||||
3. **Archive before delete, and verify.** Same discipline as any palace destructive op: write the cold copy, verify row counts and `sha256` for artifacts, and only then `DELETE` + `VACUUM` the live database. A `--dry-run` that reports what would move, in thread units, is the minimum interface.
|
||||
|
||||
Open sub-questions: the window length (a quarter is the obvious first guess); whether cold storage is a sibling `logstream-archive-<period>.sqlite3` or plain files on disk; and whether archives are queryable through the same tools behind an explicit opt-in flag, or simply left as files for a human to open when a question reaches back that far.
|
||||
|
||||
### 9.2 A news surface, without new state — proposed direction (2026-08-27)
|
||||
|
||||
Decision 9 above is **proposed in direction**: news should get its own low-priority surface at wake-up, and the owed set should not be touched.
|
||||
|
||||
**Why the obvious fix is not available.** The tempting change is to let terminal directed events into the owed set, so a report addressed to a device reaches it. That breaks the derivation's fixed point. Owed-ness is defined (§3.3) as *directed at me, not written by me, and not joined by a later terminal event of mine* — so the asserting shape (`open`) and the clearing shape (terminal) must be disjoint. Make terminal events owed, and: a reply becomes owed by its requester; the requester clears it by writing a terminal event joined to it; that event is directed, terminal, and therefore owed by the original author; and every closure mints a fresh obligation. The loop never terminates. You could exempt `task.reply` and `event.ack` by `type`, but that is the same filter re-entered through a different door, with more surface to get wrong. **`status="open"` is not an arbitrary editorial choice about tone; it is what makes owed-ness terminate.**
|
||||
|
||||
**The proposal.** A second, clearly-separated section at wake-up — "N reports since your last session" — listing directed non-`open` events newer than the reader's own last activity, count-capped, bodies not inlined.
|
||||
|
||||
The cost that sinks the naive version is state: "since your last session" implies a per-device read cursor, and a cursor in container-local storage is precisely what a `--force-recreate` erases (the palace exists because container state does not survive). No new state is needed, because **the device's own last authored event is already a cursor**: surface directed events whose `hlc` is greater than the newest `hlc` among events written by this device. That anchor lives in the shared log, survives recreate, and is comparable across replicas for the same reason the owed-set join now uses `hlc` (§7.3).
|
||||
|
||||
Properties worth stating before anyone implements it:
|
||||
|
||||
- **Per stream, not global.** Take the anchor as the newest `hlc` of this device's events *in that stream*. A global anchor lets a chatty stream advance past unread news in a quiet one — the failure is silent, so this is not a refinement to leave for later.
|
||||
- **A device that has never written sees everything.** Bounded by the count cap, and arguably correct on first boot; it is the same shape as a new joiner reading recent history.
|
||||
- **A busy device gets a narrow window.** Correct by construction: you saw the log when you last wrote to it.
|
||||
- **There is no read receipt, so repeats are possible** until the anchor advances. Acceptable only because this surface is a summary, not an interruption: count and one-line subjects, never bodies.
|
||||
- **It must not resurface.** Owed items re-announce (`MEMPALACE_MAILBOX_RESURFACE_MS`); news must not, or it becomes nagging without obligation.
|
||||
- **Opt-in first**, following the notification precedent (§7.11): inert unless explicitly enabled, and never counted in the owed total.
|
||||
|
||||
**What this does not fix.** It is still not push (§7.12 and the 404 on `/logstream/stream` for this deployment), still queued into the next turn rather than waking anyone (§7.11), and still useless to a device that never starts a session. For anything that must survive indefinitely, the paired **drawer** remains the durable channel; this surface only shortens the delay before a peer notices something already written.
|
||||
|
||||
---
|
||||
|
||||
## 10. Evidence index
|
||||
|
||||
| Claim | Where |
|
||||
|---|---|
|
||||
| Store is `logstream.sqlite3` inside the palace dir | `logstream.py:51`; `mcp_server.py:1108-1117` |
|
||||
| Full DDL, indexes, WAL | `logstream.py:447-497`, `552-556` |
|
||||
| Design constraints (quoted in §3) | `logstream.py:1-18` |
|
||||
| Event id format, ordering not carried by id | `logstream.py:92-102` |
|
||||
| `seq` = rowid, surfaced at read | `logstream.py:580-586` |
|
||||
| `origin_seq` assigned in-transaction | `logstream.py:700-706` |
|
||||
| `hlc` populated per append; format and sort property | `logstream.py:668`; `hlc.py:1-21` |
|
||||
| Cursor stays local, HLC is merge order | `hlc.py:19-21` |
|
||||
| `origin_replica` from palace dir | `logstream.py:422-436`; `replica.py:31-49` |
|
||||
| Server-side `created_at`, second precision | `logstream.py:667` |
|
||||
| No dedup on append; remote apply *is* idempotent | `logstream.py:625-716` vs `1174-1180` |
|
||||
| Artifact sha256 non-unique index | `logstream.py:447-497`, `744-810` |
|
||||
| `artifact_ids` referential check | `logstream.py:687-694`, `1181-1189` |
|
||||
| Status set; artifact kinds; size limits | `logstream.py:67-69`, `55-57`, `765-778` |
|
||||
| `type` regex, lowercase ≤64 | `logstream.py:76, 124-134` |
|
||||
| Broadcast matching in SQL | `logstream.py:958-960` |
|
||||
| `since_event_id` anchor + raise | `logstream.py:963-978` |
|
||||
| Order by rowid; limit clamp 500 | `logstream.py:947` |
|
||||
| `preview` truncates body to 200 chars | `mcp_server.py:4587-4607` |
|
||||
| `event_wait` backoff, cap, timeout result | `logstream.py:1004-1039` |
|
||||
| `ack_event` appends, never mutates; `ack_of` in metadata | `logstream.py:812-845` |
|
||||
| Patch advisories | `logstream.py:718-742` |
|
||||
| Lock exemptions, both, with rationale | `mcp_server.py:6887-6902`, `434-449` |
|
||||
| Bearer auth; SSE auth required | `mcp_server.py:7137-7141`, `7189-7194` |
|
||||
| SSE implementation, heartbeat, client cap | `mcp_server.py:7570-7671` |
|
||||
| Rejected append returns 200 + `success:false` | `mcp_server.py:4581-4584` |
|
||||
| `sync` does not touch the log | `cli.py:1057`; `sync.py` (no logstream refs) |
|
||||
| No retention/pruning anywhere | no `DELETE FROM events`/`artifacts` in the package |
|
||||
| Replication surface, attributed to RFC 004 | `logstream.py:1096-1261`; `mcp_server.py:7206-7256` |
|
||||
|
||||
### 10.1 Where the shipped tool descriptions and the code disagree
|
||||
|
||||
| Tool-description claim | Verdict |
|
||||
|---|---|
|
||||
| "append-only; corrections are new events" | True of content. Precisely: no event's *content* is ever mutated; `append_event` does issue one `UPDATE … SET origin_seq = rowid` on the row it just inserted, and the 3.8.0 migration backfills `origin_replica`/`origin_seq`/`hlc` on pre-existing rows (`logstream.py:504-550`). |
|
||||
| "`since_event_id` cannot skip anything" | True, and stronger than stated — it raises on an unknown id (§7.5). |
|
||||
| "body max 256 KiB", "artifact max 4 MiB, UTF-8 only", "timeout default 60 s max 5 min" | All true; the timeout clamps silently rather than erroring. |
|
||||
| "`to_agent=<you>` also matches `*` broadcasts" | True, SQL-level — but see §7.4 for why a broadcast still reaches no mailbox. |
|
||||
| "prefer the push stream at `GET /logstream/stream`" | True; route exists and requires auth. |
|
||||
| `GET /logstream/events` | ❌ Does not exist (§7.8). |
|
||||
| `metadata` size limit; `type` charset | ❌ Enforced but undocumented (§3.1). |
|
||||
|
||||
---
|
||||
|
||||
## 11. See also
|
||||
|
||||
- `docs/fleet-memory.md` — the operator-facing companion: what the palace's stores are *for*, and when to use a drawer versus an event.
|
||||
- `docs/rfc-001-global-palace.md` — §5 what should not be global, §7.2 the `sync` hazard, §7.3 provenance at the boundary, §7.5 single-writer expectations.
|
||||
- `docs/rfc-002-joiner.md` — replaying a second palace into a shared primary.
|
||||
- `extensions/pi/README.md` §2 — the client-side mailbox: delivery points, gating, and why it is a client concern.
|
||||
- `~/.agents/skills/mempalace/SKILL.md` §"Cross-Machine Coordination" — **normative for agent behaviour**; this RFC is normative for mechanism.
|
||||
@@ -0,0 +1,185 @@
|
||||
# Secret hygiene: keeping credentials out of the palace
|
||||
|
||||
A palace is mined from **transcripts**, and transcripts contain whatever the
|
||||
terminal printed. Print an environment, `cat` a `.env`, paste a `curl -H
|
||||
"Authorization: …"`, and the secret becomes a drawer — searchable by every agent
|
||||
on the fleet, on every machine, indefinitely. This document describes the
|
||||
scrubbing that exists, what it deliberately does **not** try to do, and the two
|
||||
layers it belongs in.
|
||||
|
||||
Measured on this fleet, 2026-08-27: a single leaked bearer token had reached
|
||||
**3 drawers, 13 feeder inbox files across all three devices, and 10 local files
|
||||
dating back 10 days**. Nobody was careless; an agent inspected an environment
|
||||
variable while debugging. That frequency is the design premise — this is a
|
||||
routine event to be handled by the pipeline, not an incident to be handled by
|
||||
discipline.
|
||||
|
||||
## 1. Where the hook goes, and why there are two of them
|
||||
|
||||
| Layer | Implemented in | Sees | Catches |
|
||||
|---|---|---|---|
|
||||
| **Client, pre-staging** | `bin/mempalace_redact.py`, called from `bin/mempalace-pi-session` at the point the staged transcript is written | the machine's own environment **and** the transcript | the real-world case: a secret this machine holds, printed into a session, before it is uploaded anywhere |
|
||||
| **Server, pre-persist** | *not yet implemented* — see §5 | only the text it is handed | secrets from clients that predate the scrubber, from non-pi clients, and secrets an agent types straight into `add_drawer` |
|
||||
|
||||
The client hook is the load-bearing one, because it is the only layer that can
|
||||
compare text against **the actual secret values it holds** (see tier 1 below) and
|
||||
because it stops the leak *before transmission*. The server hook is defence in
|
||||
depth: pattern-only, but it covers write paths the feeder never sees.
|
||||
|
||||
Client placement matters for one specific reason: the pi feeder writes the staged
|
||||
transcript from an in-memory structure, and the remote path then `rsync`s **that
|
||||
same file** byte-for-byte. Scrubbing at the write point therefore covers local
|
||||
mining and remote upload with a single hook — there is no second serialization to
|
||||
forget.
|
||||
|
||||
## 2. Detection is anchored on meaning, never on entropy
|
||||
|
||||
The tempting design — "redact long random-looking strings" — is actively
|
||||
destructive here, because a palace is *full* of high-entropy strings that are its
|
||||
own primary keys:
|
||||
|
||||
```
|
||||
drawer_pi-devbox_gotchas_f7b1c8d4c7b590196351ab9b drawer id
|
||||
evt_20260826T213924_397b1f710c2d event id
|
||||
rep_d344e349ba276d6fc11997cd552f6937 replica id
|
||||
1787780364792-000000-rep_d344e349ba276d6f… hybrid logical clock
|
||||
ecc2a9c574f4156403cf4e90aea8b0d088c75700 git commit sha
|
||||
```
|
||||
|
||||
An entropy detector fires on every one of those, and the resulting redaction is
|
||||
silent, permanent, and destroys traceability. So detection uses three anchors
|
||||
that carry meaning instead:
|
||||
|
||||
Each tier is a different *kind of evidence* that a string is a credential. The
|
||||
tier is not a severity ranking of the secret — it is how much to trust the
|
||||
detection. Rule names appear in the output; the tier tells you how to read them.
|
||||
|
||||
| Tier | Anchor | Default | False-positive risk |
|
||||
|---|---|---|---|
|
||||
| **1 — known values** | literal values from this process's env, for variables whose *name* says secret (`…TOKEN`, `…SECRET`, `…PASSWORD`, `…API_KEY`) | **redact** | none by construction: the value *is* the secret |
|
||||
| **2 — known shapes** | vendor-prefixed credentials: `ghp_…`, `github_pat_…`, `glpat-…`, `xox[abprs]-…`, `sk-…`, `AKIA…`, `AIza…`, `hf_…`, JWTs, PEM private-key blocks, credentials inside URLs, `Authorization:` headers | **redact** | very low: the prefix is meaningful, not random |
|
||||
| **3 — name=value** | an assignment whose *key* says secret | **report only** | measured **high** — see §3 |
|
||||
|
||||
### Worked examples, one line each
|
||||
|
||||
```
|
||||
MEMPALACE_REMOTE_TOKEN=pfrDBfak… tier 1 — value matches this env's secret
|
||||
the token is pfrDBfak… apparently tier 1 — same value, bare in prose, still caught
|
||||
git clone https://joakim:hunter2@git/x tier 2 — credentials in a URL
|
||||
Authorization: Bearer abcdefghijklmnop tier 2 — header shape
|
||||
ghp_AbCdEf… / glpat-… / AKIA… / sk-ant-… tier 2 — vendor prefix
|
||||
db_password = s3cr3t-p4ssw0rd-xyz tier 3 — only the KEY suggests it (reported)
|
||||
GITEA_TOKEN=0123456789abcdef… (40 hex) tier 3 — indistinguishable from a commit sha
|
||||
```
|
||||
|
||||
### Which tier fired? Read it off the output
|
||||
|
||||
The tool prints **rule names**, not tier numbers, because the rule says *what*
|
||||
matched. The feeder prefixes them with the tier so both are visible:
|
||||
|
||||
```
|
||||
[REDACTED] 2026-06-27T23-13-49.jsonl T1:env-value=1, T2:github-fine-grained=1 fp=ad78c7d4,fcd95ab5
|
||||
[scrub] 29 redaction(s) applied, 141 name-anchored candidate(s) reported only
|
||||
```
|
||||
|
||||
`RULE_TIERS` in `mempalace_redact.py` is the authoritative mapping, `tier_of()`
|
||||
reads it, and a self-test fails if any rule is left unclassified — so the two
|
||||
vocabularies cannot drift apart silently.
|
||||
|
||||
Tier 1 catches any presentation of a secret — env dump, JSON, error message,
|
||||
URL, prose — because it matches the value itself. It is also, **in the default
|
||||
configuration, the only tier that resolves a shape collision**: a 40-hex Gitea
|
||||
PAT is byte-identical to a git commit sha, so only a known value (tier 1,
|
||||
redacts) or a naming key (tier 3, reports) can tell them apart. A sha-shaped PAT
|
||||
is therefore caught if and only if it belongs to this machine — an accepted gap,
|
||||
since the alternative is redacting every commit sha in the palace.
|
||||
|
||||
## 3. The false-positive measurement, which changed the design
|
||||
|
||||
Tier 3 was originally going to redact. Measured against **52 MB of real fleet
|
||||
transcripts (45 sessions)** it produced **403 hits**, and inspection of them (with
|
||||
values masked) showed the overwhelming majority were not secrets:
|
||||
|
||||
- `GITEA_ACCESS_TOKEN=${GITEA_ACCESS_TOKEN:-}` — docker-compose interpolation
|
||||
- `const tokens = countTokensForModel` — TypeScript source stored as memory
|
||||
- `refreshToken(credentials: Credentials …)` — a *type annotation*
|
||||
- `--setattr=krbPasswordExpiration=20260529…` — an IPA attribute holding a date
|
||||
- `|1.FORGE.TOKENS:…` — AAAK diary shorthand
|
||||
- `(root@host) Password:` followed by unrelated terminal output
|
||||
|
||||
Redacting those would corrupt code, configuration and documentation held as
|
||||
memory, in order to catch secrets tier 1 already catches by value. After adding
|
||||
guards for interpolation (`${VAR}`, `$VAR`, `%VAR%`, `{{tpl}}`), code context,
|
||||
non-secret key suffixes (`…Expiration`, `…Count`, `…Type`) and all-digit values,
|
||||
the enforced count on the same corpus fell from **403 to 29** — and those 29 are
|
||||
tier-1 and tier-2 hits, i.e. real credential shapes.
|
||||
|
||||
So tier 3 **reports and does not rewrite** by default. Set
|
||||
`MEMPALACE_REDACT_STRICT=1` to make it enforce, e.g. on a machine whose
|
||||
transcripts are configuration-heavy rather than code-heavy.
|
||||
|
||||
## 4. Known false negatives — stated, not hidden
|
||||
|
||||
The scrubber will not catch a novel credential format pasted bare into prose
|
||||
with no name nearby, a secret belonging to a machine whose environment this
|
||||
process cannot see, a base64-of-a-secret, or a secret split across lines.
|
||||
|
||||
**"The scrubber ran" must never be read as "there are no secrets in here."** To
|
||||
keep that measurable rather than assumed, two things are reported:
|
||||
|
||||
- every scrub prints a count — including `0 redaction(s)`, because printing
|
||||
nothing is indistinguishable from a scrubber that never ran;
|
||||
- `suspicions()` reports high-entropy strings it did **not** redact, as
|
||||
`(length, fingerprint)` pairs rather than values, so the miss rate can be
|
||||
tracked over time and a recurring fingerprint can be investigated by hand.
|
||||
|
||||
Findings never carry the secret. A `Finding` holds the rule, the label, the
|
||||
length and `sha256(value)[:8]` — enough to recognise the same leak recurring,
|
||||
not enough to recover it.
|
||||
|
||||
**Fail closed.** If the redactor cannot be imported, the feeder refuses to stage
|
||||
rather than staging unscrubbed (`exit 3`). Override deliberately with
|
||||
`MEMPALACE_FEED_ALLOW_UNSCRUBBED=1`.
|
||||
|
||||
## 5. The server-side layer (not yet implemented)
|
||||
|
||||
Three call sites, because the server has three near-duplicate validators rather
|
||||
than one:
|
||||
|
||||
| Path | Function | Covers |
|
||||
|---|---|---|
|
||||
| drawers | `config.py` → `sanitize_content()` | `add_drawer`, `update_drawer`, `diary_write`, `checkpoint` — all four route through it |
|
||||
| events | `logstream.py` → `_sanitize_body()` | `event_append`, and `event_ack` transitively |
|
||||
| artifacts | `logstream.py` → `put_artifact()` | inlines its own checks; needs its own edit |
|
||||
|
||||
Server-side scrubbing is **tier 2 only** (plus optional tier-3 reporting): the
|
||||
hub cannot see a client's environment, so tier 1 is structurally unavailable
|
||||
there. That asymmetry is the reason the client hook is not redundant.
|
||||
|
||||
## 6. Cleaning up a leak that already landed
|
||||
|
||||
1. **Never re-echo the value while hunting it.** Pass it via stdin, never argv
|
||||
(visible in `ps` on a shared host). Note `get_drawer` is a trap: reading a
|
||||
drawer in order to redact it prints the secret back into the live transcript.
|
||||
2. **Redact, don't delete.** Replacing the value in place keeps the mined
|
||||
transcript's memory value; a stale embedding vector is a cheap price.
|
||||
3. **Scrub the feeder inbox too**, or the next mine re-files it.
|
||||
4. **A live session file needs an equal-length in-place overwrite** — the agent
|
||||
holds it open, so temp-file-plus-rename loses everything appended afterwards.
|
||||
5. **Expect residue.** SQLite keeps old page content in freed pages until
|
||||
`VACUUM`, so a raw byte scan still matches after a successful `UPDATE`. Decide
|
||||
explicitly whether that matters: if the credential store on the same host is
|
||||
plaintext anyway, it usually does not.
|
||||
|
||||
## 7. Usage
|
||||
|
||||
```bash
|
||||
# unit tests, including the must-not-redact corpus of palace id shapes
|
||||
python3 bin/mempalace_redact.py --self-test
|
||||
|
||||
# scrub anything on stdin; report goes to stderr
|
||||
some-command | python3 bin/mempalace_redact.py > clean.txt
|
||||
|
||||
# also list high-entropy strings that were NOT redacted, as fingerprints
|
||||
python3 bin/mempalace_redact.py --suspicions < transcript.jsonl > clean.jsonl
|
||||
```
|
||||
@@ -11,7 +11,7 @@ What lives where:
|
||||
|
||||
| Content | Home |
|
||||
|---|---|
|
||||
| How to expose a palace over HTTP, and the Host/Origin pin | `docs/phase-1-exposure-runbook.md` (here) |
|
||||
| How this deployment was exposed over HTTP (tunnel, DNS, token custody) | private fleet repository — see [`phase-1-exposure-runbook.md`](phase-1-exposure-runbook.md), also moved |
|
||||
| Why a palace needs a special backup, and how to restore one | `docs/backup-and-recovery.md` (here) |
|
||||
| Unit/timer/plist templates | `contrib/` (here) |
|
||||
| Which machine is primary, its addresses, users, tunnels, offsite target | private fleet repository |
|
||||
|
||||
+171
-15
@@ -117,7 +117,7 @@ side of the wiring.
|
||||
| `MEMPALACE_FEED_WING` | `wing_conversations` | Target wing — passed to both the exporter and the `mempalace_mine` call. |
|
||||
| `MEMPALACE_FEED_DEBOUNCE_MS` | `600000` (10 min) | Minimum gap between mid-session (`agent_settled`) feeds. Bounds crash loss to one window instead of a whole session. |
|
||||
| `MEMPALACE_FEED_PREPARE_TIMEOUT_MS` | `120000` | Kills a wedged `--prepare` subprocess. |
|
||||
| `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `30000` | Caps the `mempalace_mine` call so a stalled palace can't hang session exit. |
|
||||
| `MEMPALACE_FEED_MINE_TIMEOUT_MS` | `300000` (5 min) | Bounds how long the extension *waits* for `mempalace_mine`, so a stalled palace can't hang session exit. It does **not** cancel the mine — see [Debugging](#debugging). Raised from `30000` in 2026-09: the mine is the slowest call this extension makes (30–60s in normal operation), so the old deadline fired routinely and reported healthy behaviour as an error. |
|
||||
|
||||
**Remote palace:** if `$MEMPALACE_REMOTE_URL` is set (see
|
||||
[Transport](#transport-local-vs-external)), `mempalace_mine`'s source path is
|
||||
@@ -187,8 +187,8 @@ chosen at load time:
|
||||
transport is chosen once at extension load. Confirm the result the same way
|
||||
as a container flip: ask the agent for `mempalace_status` and check the
|
||||
reported palace path is the **remote** host's, not your own
|
||||
`$HOME/.mempalace/palace` — see
|
||||
[`docs/phase-1-exposure-runbook.md`](../../docs/phase-1-exposure-runbook.md) §3.8.
|
||||
`$HOME/.mempalace/palace`. That one check is the whole verification: a
|
||||
half-flipped client reports a local path while looking healthy.
|
||||
|
||||
Serve such an endpoint with `mempalace serve --host 172.17.0.1 --port 8765`
|
||||
(the `pi-devbox` / `opencode-devbox` repos ship a
|
||||
@@ -207,8 +207,9 @@ chosen at load time:
|
||||
token auto-minting is gated on the bind being non-loopback, so it starts with
|
||||
**no authentication at all**, no warning. Bind the docker0 gateway
|
||||
(`172.17.0.1`): reachable from the host and its containers, not from the LAN.
|
||||
See
|
||||
[`docs/phase-1-exposure-runbook.md`](../../docs/phase-1-exposure-runbook.md).
|
||||
Binding an interface is not the same as exposing a palace — an MCP endpoint
|
||||
also pins `Host`/`Origin`, so a request arriving under the wrong hostname is
|
||||
refused even when the port is open.
|
||||
|
||||
Implementation note: the HTTP client (`RemoteMcpClient`) is **vendored** from
|
||||
[`pi-extensions`](https://gitea.jordbo.se/joakimp/pi-extensions)'
|
||||
@@ -234,9 +235,9 @@ local palace, so a remote outage can never scatter memories into a local copy
|
||||
nobody will look at again. The practical corollary, worth knowing before you
|
||||
debug the wrong layer: **"the agent has no `mempalace_*` tools" is the
|
||||
expected symptom of a server, token, or DNS fault**, not of a broken install.
|
||||
Diagnose it with a direct `curl` to `MEMPALACE_REMOTE_URL` — see
|
||||
[`docs/phase-1-exposure-runbook.md`](../../docs/phase-1-exposure-runbook.md)
|
||||
§3.8. The design rationale for de-registering rather than degrading is in
|
||||
Diagnose it with a direct `curl` to `MEMPALACE_REMOTE_URL`, and confirm the flip
|
||||
with `mempalace_status` — the reported palace path must be the remote host's.
|
||||
The design rationale for de-registering rather than degrading is in
|
||||
[`docs/rfc-001-global-palace.md`](../../docs/rfc-001-global-palace.md) §2 and §4.1.
|
||||
|
||||
## Identity
|
||||
@@ -315,24 +316,123 @@ waits for the agent to think of asking. Two delivery points, both fail-silent:
|
||||
`triggerTurn` — waking the model on inbound fleet traffic is a much larger
|
||||
behavioural change than auto-delivery.
|
||||
|
||||
**Which means a human is the trigger, and the mailbox now says so.** Measured
|
||||
2026-08-26 on two devices: a delivery lands, the agent is idle, nothing happens,
|
||||
and the operator asks *"do I have to nudge you for you to read this?"*. Yes —
|
||||
because between the poll and the next turn no inference is running. The old text
|
||||
explained how to *close* an ask and never said when it would be *seen*, so the
|
||||
only reader who needed that fact was the one not told. Two additions, neither of
|
||||
which touches the no-`triggerTurn` decision:
|
||||
|
||||
- **A delivery note in the message itself** — states that this is a queued
|
||||
message, that nothing woke the agent, and that any message starts the turn
|
||||
that handles it. Free, and aimed at the human reading the window.
|
||||
- **A notification at poll time**, because the note only helps someone who is
|
||||
already looking, and the case that loses an ask is nobody looking:
|
||||
|
||||
| `MEMPALACE_MAILBOX_NOTIFY` | Behaviour |
|
||||
|---|---|
|
||||
| *unset* (default) | in-TUI `ctx.ui.notify`, the same surface `session_start` already uses |
|
||||
| `desktop` | additionally a terminal-native notification — Kitty `OSC 99`, else `OSC 777` (iTerm2, WezTerm, Ghostty, rxvt-unicode) |
|
||||
| `0` / `off` | silent; mailbox still delivers |
|
||||
|
||||
The `desktop` path is how a notification escapes a container without
|
||||
`notify-send`, DBus or any host access: the escape sequence is written to stdout
|
||||
and interpreted by the terminal emulator on the human's own machine. It is
|
||||
opt-in because writing raw escapes is a behaviour change on a shared machine,
|
||||
not because it is unreliable. Title and body are stripped of `;` and control
|
||||
bytes, so a payload can neither forge an OSC field nor end the sequence early.
|
||||
|
||||
**⚠️ Not yet observed firing through tmux (2026-08-26).** If pi runs inside a
|
||||
multiplexer — e.g. `kitty → tmux on the host → docker exec → pi` — tmux drops
|
||||
OSC sequences it does not implement, so the ping can vanish silently between
|
||||
the container and the human. Reaching the outer terminal needs tmux's DCS
|
||||
passthrough plus `allow-passthrough on`, which is not implemented here yet.
|
||||
Note the client can detect *neither* layer: `KITTY_WINDOW_ID` is not forwarded
|
||||
by `docker exec`, and `TMUX` is unset because tmux runs one level further out.
|
||||
Verify with a hand-written sequence in your own stack before trusting it.
|
||||
|
||||
It fires **only when something is due** — the same condition as the delivery
|
||||
itself. A ping on an empty poll would train its reader to ignore it, which is
|
||||
the failure this whole feature exists to reverse.
|
||||
|
||||
Owed-ness is **derived, never read off a field**, because `event_ack` appends and
|
||||
`status` is written once: a directed `open` event matches the mailbox query
|
||||
*forever*, answered or not. Two calls (`to_agent=<me> status=open`, and
|
||||
`from_agent=<me>`), then a candidate counts as answered only when one of this
|
||||
device's own events has a **strictly higher `seq`**, joins via
|
||||
device's own events is **strictly later**, joins via
|
||||
`metadata.ack_of` or a shared `correlation_id`, *and* carries a terminal status
|
||||
(`applied`/`superseded`/`failed`/`blocked`). The `seq` test is load-bearing:
|
||||
(`applied`/`superseded`/`failed`/`blocked`). The ordering test is load-bearing:
|
||||
without it one terminal reply suppresses every later ask on that correlation
|
||||
forever. `seq` is replica-local, so `hlc` is the correct key once `mesh_peers`
|
||||
reports actual peers.
|
||||
forever.
|
||||
|
||||
"Strictly later" means `hlc` when both events carry one — a hybrid logical clock
|
||||
rendered fixed-width, so a string comparison is a causal comparison across
|
||||
replicas — falling back to `seq` only when either side lacks an `hlc`. `seq` is
|
||||
this database's arrival `rowid`, so on a mesh the same event has a different
|
||||
`seq` per replica and a reply can arrive before its ask. Never `created_at`: it
|
||||
is second-precision, and a tie there can suppress an *unanswered* ask, which is
|
||||
the one failure this derivation exists to prevent.
|
||||
|
||||
`*` broadcasts are excluded even though `to_agent=<me>` matches them, because the
|
||||
protocol says a broadcast owes nobody a reply — which also means broadcasting an
|
||||
ask demonstrably reaches no owed set, giving that documented anti-pattern teeth.
|
||||
|
||||
**Dormant asks — open, owed, and deliberately not announced.** Some asks are
|
||||
planted *unanswerable on purpose*: a first-boot acceptance describes work that
|
||||
becomes possible only when the device is next recreated, and it must stay owed
|
||||
until then, because closing it early to tidy the mailbox is exactly how the work
|
||||
gets lost. Derivation alone cannot tell that apart from neglected work, so such
|
||||
an ask was announced at every session start and every poll for as long as it was
|
||||
correctly waiting — measured at three days running on `emb-7kj4vr4g`
|
||||
(2026-09-28 → 2026-10-01), and across three consecutive releases before that.
|
||||
The ask was right; announcing it was wrong, and the cost landed on the human
|
||||
reading the window, the one reader who cannot filter it.
|
||||
|
||||
An ask may therefore declare, in its own `metadata`, the condition under which it
|
||||
is merely waiting:
|
||||
|
||||
```json
|
||||
"dormant_unless": [
|
||||
{ "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json",
|
||||
"field": "release_tag", "baseline": "v1.9.4" },
|
||||
{ "kind": "file_mtime", "path": "/etc/hostname",
|
||||
"baseline": "2026-09-22T18:12:49Z" }
|
||||
]
|
||||
```
|
||||
|
||||
It is dormant while **every** condition still matches its baseline, and goes live
|
||||
the moment **any** of them differs — the trigger such asks already stated in
|
||||
prose ("act when EITHER differs"), now in a form the bridge can check. Two kinds,
|
||||
both local-file-only: `file_mtime` (compared at whole seconds in UTC, because a
|
||||
filesystem mtime carries sub-second residue that a reported ISO baseline never
|
||||
will — comparing raw milliseconds would mark every such predicate permanently
|
||||
"changed" and silently disable the feature) and `json_field` (dotted paths
|
||||
allowed, compared as strings so a manifest holding `3` matches a baseline of
|
||||
`"3"`). There is no expression language, no shell and no network: a general
|
||||
evaluator in the path that decides whether work is *visible* is a far worse trade
|
||||
than a clumsy schema.
|
||||
|
||||
Dormancy is **withheld from the announcement, never from the mailbox**: the
|
||||
wake-up injection lists dormant asks once per session with their ids, and a
|
||||
mid-session poll mentions only a count, and only when the window is already open
|
||||
for something else. They are not added to the resurface map, so one becomes
|
||||
announceable the instant its baseline moves.
|
||||
|
||||
**Dormancy must be proven, never assumed** — every unevaluable predicate shows
|
||||
the ask as owed. A missing file, an unreadable one, a baseline that will not
|
||||
parse, an unknown `kind`, a vanished field, a relative path, more than eight
|
||||
conditions: each announces. The dangerous failure here is not a spurious nag but
|
||||
work that disappears because a predicate could not be evaluated, which would be
|
||||
indistinguishable from the ask being lost and would not surface until a release
|
||||
needed it. An ask with no `dormant_unless` behaves exactly as it did before the
|
||||
feature existed. `scripts/test-dormancy.sh` exercises all of it, positive and
|
||||
negative arms both.
|
||||
|
||||
Gated on `MEMPALACE_PI_DEVICE` **and** `MEMPALACE_REMOTE_URL` (the same pair as
|
||||
the stamper, since an unstamped client has no address to be reached at), and
|
||||
disabled outright with `MEMPALACE_MAILBOX=0`. Inert on a solitary palace: no
|
||||
disabled outright with `MEMPALACE_MAILBOX=0` (notifications alone with
|
||||
`MEMPALACE_MAILBOX_NOTIFY=0`). Inert on a solitary palace: no
|
||||
calls, no injection. Delivery is the mechanism; the *norms* — what a reply owes,
|
||||
and that only a terminal event closes a thread — remain normative in the
|
||||
*consumer* skill (`~/.agents/skills/mempalace/SKILL.md`). **This file documents
|
||||
@@ -342,12 +442,21 @@ the mechanism; the skill is normative for behaviour.**
|
||||
an SSE endpoint (`GET /logstream/stream`, `text/event-stream` in
|
||||
`mempalace/mcp_server.py`), but a deployment may expose only the MCP endpoint
|
||||
through its reverse proxy — verified 2026-08-26 against
|
||||
`https://mempalace.jordbo.se`, where `/logstream/events`, `/logstream/stream`
|
||||
and `/sync/peers` all return 404 while `/mcp` serves normally. Where that is the
|
||||
`https://mempalace.jordbo.se`, where `/logstream/stream` and `/sync/peers`
|
||||
return 404 while `/mcp` serves normally. Where that is the
|
||||
case, polling through the existing MCP client is the only available path — which
|
||||
is what the mailbox in §2 does — and enabling SSE means a proxy route plus an
|
||||
auth decision, not an extension change.
|
||||
|
||||
> ⚠️ **Corrected 2026-08-26.** An earlier revision of this paragraph listed
|
||||
> `/logstream/events` alongside those two as proxy-blocked. That route **does not
|
||||
> exist in the server at all** — the complete GET table in mempalace 3.8.0 is
|
||||
> `/healthz`, `/statusz`, `/logstream/stream` and `/sync/{version_vector,ops,artifact,peers}`,
|
||||
> so `/logstream/events` would 404 against a directly-reachable server too. The
|
||||
> proxy inference was right for the other two and wrong for that one; see
|
||||
> [RFC 003 §7.8](../../docs/rfc-003-coordination-log.md). Distinguish *route absent*
|
||||
> from *route blocked* before blaming infrastructure.
|
||||
|
||||
As with stamping, all of this is inert unless `MEMPALACE_REMOTE_URL` points at a
|
||||
shared palace. On a solitary palace the event tools work fine and the log
|
||||
contains only this machine's own events.
|
||||
@@ -365,6 +474,9 @@ one — the long-lived server is only killed when a request genuinely stalls.
|
||||
|
||||
- `MEMPALACE_MCP_TIMEOUT_MS` — tool-call/request timeout. Default `60000`.
|
||||
Kept short on purpose: a *query* taking this long is genuinely wedged.
|
||||
The one call that is not a query — the feed's `mempalace_mine` — passes its
|
||||
own deadline (`MEMPALACE_FEED_MINE_TIMEOUT_MS`) down to the transport per
|
||||
call, so this default does not apply to it (since 2026-09-18; see below).
|
||||
- `MEMPALACE_MCP_INIT_TIMEOUT_MS` — `initialize` + `tools/list` handshake
|
||||
timeout. Default `300000`. Deliberately generous: a genuine first
|
||||
cold-open over virtiofs can legitimately take minutes, and killing a
|
||||
@@ -407,6 +519,50 @@ the next tool call transparently respawns `mempalace-mcp` and retries.
|
||||
`mempalace-mcp` manually with raw JSON-RPC on stdin to read the
|
||||
server-side error — much faster than guessing.
|
||||
|
||||
### `feed (tick) failed: mine timed out after …ms`
|
||||
|
||||
**Nothing has been lost.** The deadline bounds only how long the extension
|
||||
*waits*; it cannot cancel the mine, which continues on the server. The
|
||||
transcript is already staged before the mine is invoked, and
|
||||
`mempalace mine --mode convos` dedups by `source_file` and is idempotent, so the
|
||||
work either completed after the deadline or is redone by the next tick.
|
||||
|
||||
**Do not "fix" it by retrying harder from the client.** The palace is a single
|
||||
writer; a blind retry is what turns one slow mine into a queue of them.
|
||||
|
||||
Before 2026-09 this message appeared many times per session, which made it look
|
||||
like a persistent failure. That was a real defect, now fixed: `lastFeedAt` was
|
||||
recorded only after a *successful* wait, so a timeout left the debounce clock
|
||||
stale and every following settled turn started another mine on top of the one
|
||||
still running. Two changes — recording the attempt before the wait, and raising
|
||||
the deadline to sit far above the slowest honest completion — mean a healthy
|
||||
fleet should now never see it.
|
||||
|
||||
If you *do* still see it, it is now informative rather than noise: a mine
|
||||
exceeded five minutes. Check palace size and whether another writer (a
|
||||
host-side feeder, a scheduled `mine`) is holding the write lock, rather than
|
||||
raising the timeout again.
|
||||
|
||||
### `feed (tick) failed: mempalace remote request 'tools/call' failed: timed out after 60000ms`
|
||||
|
||||
Same event, different deadline — and the same reassurance: **nothing has been
|
||||
lost**, the mine continues server-side.
|
||||
|
||||
This is what the previous message turned into after the 2026-09 change, and
|
||||
it exposed that the change was incomplete. The feed's five-minute deadline was
|
||||
only *raced* against the call; `callTool()` had no way to carry it, so the
|
||||
transport's generic per-request timeout (`MEMPALACE_MCP_TIMEOUT_MS`, 60 s)
|
||||
fired first on every honest 60 s+ mine, and the 300 s was unreachable. On the
|
||||
stdio transport this was worse than noise: a per-request timeout there kills the
|
||||
server child, so the mine really was aborted at 60 s.
|
||||
|
||||
Fixed 2026-09-18: `callTool(name, args, { timeoutMs })` passes a per-call
|
||||
deadline to both transports and the feed uses it for the mine.
|
||||
`scripts/test-mcp-call-timeout.sh` pins the contract (a plain call still honours
|
||||
the short default; the override is honoured and is itself a deadline). Seeing
|
||||
this message on a fixed build means a mine exceeded *five* minutes — treat it as
|
||||
the previous section says.
|
||||
|
||||
## The `Type.Unsafe` gotcha
|
||||
|
||||
Earlier versions of this extension registered every MCP tool with
|
||||
|
||||
+624
-49
@@ -67,7 +67,9 @@
|
||||
* child, so pi gets an error instead of hanging and later calls fail fast.
|
||||
* This is a per-REQUEST timeout, not a process-lifetime one — the
|
||||
* long-lived server is only killed when a request genuinely stalls.
|
||||
* - MEMPALACE_MCP_TIMEOUT_MS tool-call/request timeout (default 60000)
|
||||
* - MEMPALACE_MCP_TIMEOUT_MS tool-call/request timeout (default 60000);
|
||||
* the feed's mine carries its own, longer
|
||||
* deadline (MEMPALACE_FEED_MINE_TIMEOUT_MS)
|
||||
* - MEMPALACE_MCP_INIT_TIMEOUT_MS initialize+tools/list timeout (default 300000)
|
||||
* Set either to 0 to disable (legacy unbounded behavior).
|
||||
*
|
||||
@@ -90,6 +92,7 @@
|
||||
*/
|
||||
|
||||
import { type ChildProcessWithoutNullStreams, spawn } from "node:child_process";
|
||||
import { readFileSync, statSync } from "node:fs";
|
||||
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
|
||||
import { Type } from "typebox";
|
||||
|
||||
@@ -116,7 +119,13 @@ interface IMcpClient {
|
||||
readonly alive: boolean;
|
||||
onExit: (() => void) | null;
|
||||
start(): Promise<void>;
|
||||
callTool(name: string, args: Record<string, unknown>): Promise<any>;
|
||||
/**
|
||||
* `opts.timeoutMs` overrides the transport's generic per-request deadline
|
||||
* for THIS call only. Callers that knowingly invoke a long server-side job
|
||||
* (the feed's `mempalace_mine`) pass their own deadline here; everything
|
||||
* else keeps the short default, which is the wedged-query guard.
|
||||
*/
|
||||
callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any>;
|
||||
ensureAlive(): Promise<boolean>;
|
||||
stop(): void | Promise<void>;
|
||||
}
|
||||
@@ -132,6 +141,7 @@ const num = (envVal: string | undefined, fallback: number): number => {
|
||||
type LogEvent = {
|
||||
id?: string;
|
||||
seq?: number;
|
||||
hlc?: string;
|
||||
type?: string;
|
||||
status?: string;
|
||||
from_agent?: string;
|
||||
@@ -148,6 +158,165 @@ const sleep = (ms: number): Promise<void> =>
|
||||
if (typeof t.unref === "function") t.unref();
|
||||
});
|
||||
|
||||
// ── Dormancy ──────────────────────────────────────────────────────────────────
|
||||
/**
|
||||
* Is a standing ask DORMANT — correct to leave open, wrong to announce?
|
||||
*
|
||||
* The problem this solves, measured rather than imagined. A first-boot
|
||||
* acceptance ask is planted deliberately unanswerable: it describes work that
|
||||
* becomes possible only when the device is next recreated, and it must STAY
|
||||
* owed until then, because closing it early to tidy the mailbox is exactly how
|
||||
* that work gets lost. But deriveOwed could only see "directed, open, not
|
||||
* answered, not withdrawn", so such an ask is announced at every session start
|
||||
* and every poll for as long as it is correctly waiting — three days running on
|
||||
* emb-7kj4vr4g (2026-09-28 → 2026-10-01), and three consecutive releases before
|
||||
* that. The ask was right; announcing it was wrong. The cost lands on the human
|
||||
* reading the window, who is the one reader that cannot filter it.
|
||||
*
|
||||
* So an ask may now declare, in its own metadata, the condition under which it
|
||||
* is merely waiting:
|
||||
*
|
||||
* "dormant_unless": [
|
||||
* { "kind": "json_field", "path": "/etc/pi-devbox/build-manifest.json",
|
||||
* "field": "release_tag", "baseline": "v1.9.4" },
|
||||
* { "kind": "file_mtime", "path": "/etc/hostname",
|
||||
* "baseline": "2026-09-22T18:12:49Z" }
|
||||
* ]
|
||||
*
|
||||
* It is dormant while EVERY condition still matches its baseline, and goes live
|
||||
* the moment ANY of them differs — which is the trigger those asks already
|
||||
* state in prose ("act when EITHER differs"), now in a form a machine can
|
||||
* check. Dormant asks are withheld from the ANNOUNCED owed set, never from the
|
||||
* mailbox: the wake-up injection still lists them once per session, so they
|
||||
* stay discoverable and cannot be silently dropped.
|
||||
*
|
||||
* THE SAFETY RULE, and the reason every branch below reports `unchanged: false`
|
||||
* on doubt: dormancy must be PROVEN, never assumed. A missing file, an
|
||||
* unreadable one, a baseline that will not parse, an unknown `kind`, a field
|
||||
* that has vanished — each is a reason to SHOW the ask, not to hide it. The
|
||||
* dangerous failure here is not a spurious nag; it is work that vanishes
|
||||
* because a predicate could not be evaluated, which would be indistinguishable
|
||||
* from the ask being lost and would not surface until a release needed it.
|
||||
*
|
||||
* No eval, no shell, no network: the predicate language is deliberately two
|
||||
* fixed checks against local files. A general expression evaluator sitting in
|
||||
* the path that decides whether work is VISIBLE is a far worse trade than a
|
||||
* slightly clumsy schema.
|
||||
*/
|
||||
export type DormancyVerdict = { dormant: boolean; reason: string };
|
||||
|
||||
/** Split of the derived owed-set: what to announce, and what is merely waiting. */
|
||||
export type OwedSplit = { owed: LogEvent[]; dormant: LogEvent[] };
|
||||
|
||||
// Bounds, so a malformed or hostile event cannot turn a mailbox read into real
|
||||
// work. Both are deliberately small: these predicates describe boot state, and
|
||||
// anything needing more than a handful of local checks is not a dormancy rule.
|
||||
const MAX_DORMANCY_CONDITIONS = 8;
|
||||
const MAX_DORMANCY_JSON_BYTES = 256 * 1024;
|
||||
|
||||
const dormancyCondition = (c: unknown): { unchanged: boolean; reason: string } => {
|
||||
if (typeof c !== "object" || c === null || Array.isArray(c))
|
||||
return { unchanged: false, reason: "condition is not an object" };
|
||||
const { kind, path, baseline, field } = c as Record<string, unknown>;
|
||||
// Absolute paths only: a relative one would resolve against whatever cwd the
|
||||
// pi process happens to hold, which is not a property of the device whose
|
||||
// state the baseline describes.
|
||||
if (typeof path !== "string" || !path.startsWith("/") || path.includes("\0"))
|
||||
return {
|
||||
unchanged: false,
|
||||
reason: `path must be an absolute string, got ${JSON.stringify(path)}`,
|
||||
};
|
||||
if (typeof baseline !== "string" && typeof baseline !== "number")
|
||||
return { unchanged: false, reason: "baseline must be a string or a number" };
|
||||
|
||||
if (kind === "file_mtime") {
|
||||
// Compared at WHOLE SECONDS in UTC, because the two routes that produce
|
||||
// these baselines disagree below that: a filesystem mtime carries
|
||||
// sub-second residue (measured: /etc/hostname at .773761009) that a
|
||||
// hand-written or reported ISO baseline never will. Comparing raw
|
||||
// milliseconds would make every such predicate permanently "changed",
|
||||
// i.e. would silently disable the feature while appearing to work.
|
||||
const expected = Date.parse(String(baseline));
|
||||
if (!Number.isFinite(expected))
|
||||
return {
|
||||
unchanged: false,
|
||||
reason: `baseline is not a parseable date: ${JSON.stringify(baseline)}`,
|
||||
};
|
||||
let mtimeMs: number;
|
||||
try {
|
||||
mtimeMs = statSync(path).mtimeMs;
|
||||
} catch {
|
||||
return { unchanged: false, reason: `cannot stat ${path}` };
|
||||
}
|
||||
return Math.floor(mtimeMs / 1000) === Math.floor(expected / 1000)
|
||||
? { unchanged: true, reason: `${path} mtime still at baseline` }
|
||||
: { unchanged: false, reason: `${path} mtime moved from baseline` };
|
||||
}
|
||||
|
||||
if (kind === "json_field") {
|
||||
if (typeof field !== "string" || field.length === 0)
|
||||
return { unchanged: false, reason: "json_field needs a non-empty field name" };
|
||||
let text: string;
|
||||
try {
|
||||
const buf = readFileSync(path);
|
||||
if (buf.byteLength > MAX_DORMANCY_JSON_BYTES)
|
||||
return {
|
||||
unchanged: false,
|
||||
reason: `${path} exceeds the ${MAX_DORMANCY_JSON_BYTES}-byte cap`,
|
||||
};
|
||||
text = buf.toString("utf8");
|
||||
} catch {
|
||||
return { unchanged: false, reason: `cannot read ${path}` };
|
||||
}
|
||||
let doc: unknown;
|
||||
try {
|
||||
doc = JSON.parse(text);
|
||||
} catch {
|
||||
return { unchanged: false, reason: `${path} is not valid JSON` };
|
||||
}
|
||||
let cur: unknown = doc;
|
||||
for (const seg of field.split(".")) {
|
||||
if (
|
||||
typeof cur !== "object" ||
|
||||
cur === null ||
|
||||
!Object.prototype.hasOwnProperty.call(cur, seg)
|
||||
)
|
||||
return { unchanged: false, reason: `${path} has no field ${field}` };
|
||||
cur = (cur as Record<string, unknown>)[seg];
|
||||
}
|
||||
if (cur !== null && typeof cur === "object")
|
||||
return { unchanged: false, reason: `field ${field} is not a scalar` };
|
||||
// Compared as strings on purpose: the baseline arrives as JSON metadata, so
|
||||
// a manifest holding 3 and a baseline saying "3" are the same fact.
|
||||
return String(cur) === String(baseline)
|
||||
? { unchanged: true, reason: `${field} still at baseline` }
|
||||
: { unchanged: false, reason: `${field} moved from baseline` };
|
||||
}
|
||||
|
||||
return { unchanged: false, reason: `unknown dormancy kind ${JSON.stringify(kind)}` };
|
||||
};
|
||||
|
||||
export function evaluateDormancy(
|
||||
metadata: Record<string, unknown> | null | undefined,
|
||||
): DormancyVerdict {
|
||||
const raw = metadata?.dormant_unless;
|
||||
// Every one of these returns NOT dormant, i.e. "announce it". An ask without
|
||||
// a predicate behaves exactly as it did before this feature existed.
|
||||
if (raw === undefined || raw === null) return { dormant: false, reason: "no dormant_unless" };
|
||||
if (!Array.isArray(raw)) return { dormant: false, reason: "dormant_unless is not an array" };
|
||||
if (raw.length === 0) return { dormant: false, reason: "dormant_unless is empty" };
|
||||
if (raw.length > MAX_DORMANCY_CONDITIONS)
|
||||
return {
|
||||
dormant: false,
|
||||
reason: `too many conditions (${raw.length} > ${MAX_DORMANCY_CONDITIONS})`,
|
||||
};
|
||||
for (const c of raw) {
|
||||
const v = dormancyCondition(c);
|
||||
if (!v.unchanged) return { dormant: false, reason: v.reason };
|
||||
}
|
||||
return { dormant: true, reason: `all ${raw.length} condition(s) still at baseline` };
|
||||
}
|
||||
|
||||
class StdioMcpClient implements IMcpClient {
|
||||
private proc: ChildProcessWithoutNullStreams | null = null;
|
||||
private nextId = 1;
|
||||
@@ -371,8 +540,12 @@ class StdioMcpClient implements IMcpClient {
|
||||
});
|
||||
}
|
||||
|
||||
async callTool(name: string, args: Record<string, unknown>): Promise<any> {
|
||||
return this.request("tools/call", { name, arguments: args });
|
||||
async callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any> {
|
||||
// The per-call override matters MORE here than for the HTTP client: on
|
||||
// timeout this transport kills the server child, so a generic deadline
|
||||
// that undercuts a long mine does not merely abandon the wait — it
|
||||
// aborts the mine.
|
||||
return this.request("tools/call", { name, arguments: args }, opts?.timeoutMs ?? this.requestTimeoutMs);
|
||||
}
|
||||
|
||||
/** SIGTERM then SIGKILL grace, for stall recovery. */
|
||||
@@ -420,6 +593,10 @@ class StdioMcpClient implements IMcpClient {
|
||||
// • per-request AbortController timeout honouring MEMPALACE_MCP_TIMEOUT_MS /
|
||||
// MEMPALACE_MCP_INIT_TIMEOUT_MS, mirroring StdioMcpClient's timeout ethos.
|
||||
// • alive / ensureAlive / onExit to satisfy IMcpClient.
|
||||
// • callTool() takes an optional per-call `{ timeoutMs }` (IMcpClient
|
||||
// contract, see there) so the feed's long-running mine is not cut off by
|
||||
// the generic per-request deadline. Not a protocol change; sync token
|
||||
// unchanged.
|
||||
//
|
||||
// NOTE: mempalace-mcp --transport http is a SESSIONLESS, stateless JSON-RPC
|
||||
// server (no Mcp-Session-Id, always application/json, Connection: close), so
|
||||
@@ -465,8 +642,8 @@ class RemoteMcpClient implements IMcpClient {
|
||||
this.healthy = true;
|
||||
}
|
||||
|
||||
async callTool(name: string, args: Record<string, unknown>): Promise<any> {
|
||||
return this.request("tools/call", { name, arguments: args });
|
||||
async callTool(name: string, args: Record<string, unknown>, opts?: { timeoutMs?: number }): Promise<any> {
|
||||
return this.request("tools/call", { name, arguments: args }, { timeoutMs: opts?.timeoutMs });
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -854,7 +1031,22 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
const feedWing = process.env.MEMPALACE_FEED_WING ?? "wing_conversations";
|
||||
const feedDebounceMs = num(process.env.MEMPALACE_FEED_DEBOUNCE_MS, 600_000);
|
||||
const feedPrepareTimeoutMs = num(process.env.MEMPALACE_FEED_PREPARE_TIMEOUT_MS, 120_000);
|
||||
const feedMineTimeoutMs = num(process.env.MEMPALACE_FEED_MINE_TIMEOUT_MS, 30_000);
|
||||
// 300_000, raised from 30_000 on 2026-09-10. The mine is the SLOWEST thing
|
||||
// this extension does — it offers every qualifying session transcript to a
|
||||
// SINGLE-WRITER palace, measured at 30–60s in normal operation and growing
|
||||
// with the corpus — yet it carried by far the TIGHTEST deadline: 4x tighter
|
||||
// than the prepare step that precedes it (120_000) and 10x tighter than the
|
||||
// init handshake (300_000), which is a fast call. All three were introduced
|
||||
// together in 29e660e (2026-08-12) and this one was never revisited, so the
|
||||
// deadline fired during entirely normal operation and the resulting message
|
||||
// read as an error when nothing had gone wrong.
|
||||
//
|
||||
// Matched to the init timeout because liveness is the ONLY legitimate job
|
||||
// left for this deadline: the Promise.race below abandons our WAIT, it cannot
|
||||
// cancel the server's work, so the deadline buys nothing except an escape
|
||||
// from a permanently hung call. It must therefore sit far above the slowest
|
||||
// honest completion, not near it.
|
||||
const feedMineTimeoutMs = num(process.env.MEMPALACE_FEED_MINE_TIMEOUT_MS, 300_000);
|
||||
let lastFeedAt = 0; // 0 => the first settled turn also acts as a catch-up
|
||||
let feedInFlight: Promise<void> | null = null;
|
||||
|
||||
@@ -908,17 +1100,46 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
try {
|
||||
const source = await prepareFeed(reason);
|
||||
if (!source) return;
|
||||
// Record the attempt HERE, before awaiting — not after a successful
|
||||
// wait. The race below abandons only our WAIT; the mine keeps running
|
||||
// server-side, and `mine --mode convos` dedups by source_file and is
|
||||
// idempotent, so a timeout is emphatically not a "did not happen".
|
||||
//
|
||||
// Leaving lastFeedAt stale on the timeout path defeated the debounce
|
||||
// guard in the agent_settled handler below (`Date.now() - lastFeedAt <
|
||||
// feedDebounceMs`): with lastFeedAt unchanged that guard passed on EVERY
|
||||
// settled turn, and because `run` had already settled, feedInFlight was
|
||||
// null too — so BOTH guards stood open. Each settled turn then launched
|
||||
// another mine while the previous one was still running: overlapping
|
||||
// writers queueing on a single-writer palace, each making the next one
|
||||
// slower and the next timeout likelier. That positive feedback loop,
|
||||
// not the tight deadline by itself, is why operators saw "mine timed out
|
||||
// after 30000ms" many times per session rather than at most once per
|
||||
// debounce window.
|
||||
lastFeedAt = Date.now();
|
||||
// The deadline is passed DOWN to the transport as well as raced
|
||||
// here. Until 2026-09-18 it was only raced: callTool() had no way
|
||||
// to carry it, so the transport's generic per-request timeout
|
||||
// (MEMPALACE_MCP_TIMEOUT_MS, 60 000) fired first on every honest
|
||||
// 60 s+ mine — "remote request 'tools/call' failed: timed out
|
||||
// after 60000ms" — and the 300 000 below was unreachable. The
|
||||
// race stays as the liveness guard for a transport whose timeout
|
||||
// is disabled (0).
|
||||
await Promise.race([
|
||||
client.callTool("mempalace_mine", {
|
||||
source,
|
||||
mode: "convos",
|
||||
wing: feedWing,
|
||||
// Internal call: it does not pass through the registered tool's
|
||||
// execute(), so it stamps itself. These ARE this harness's own
|
||||
// transcripts from this device, so the harness segment is the
|
||||
// agent (not `miner`) even though the tool is `mine`.
|
||||
agent: stampProvenance ? `${agentName}@${device}` : agentName,
|
||||
}),
|
||||
client.callTool(
|
||||
"mempalace_mine",
|
||||
{
|
||||
source,
|
||||
mode: "convos",
|
||||
wing: feedWing,
|
||||
// Internal call: it does not pass through the registered tool's
|
||||
// execute(), so it stamps itself. These ARE this harness's own
|
||||
// transcripts from this device, so the harness segment is the
|
||||
// agent (not `miner`) even though the tool is `mine`.
|
||||
agent: stampProvenance ? `${agentName}@${device}` : agentName,
|
||||
},
|
||||
{ timeoutMs: feedMineTimeoutMs },
|
||||
),
|
||||
new Promise((_resolve, reject) =>
|
||||
setTimeout(
|
||||
() => reject(new Error(`mine timed out after ${feedMineTimeoutMs}ms`)),
|
||||
@@ -926,7 +1147,6 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
),
|
||||
),
|
||||
]);
|
||||
lastFeedAt = Date.now();
|
||||
} catch (err) {
|
||||
process.stderr.write(
|
||||
`[mempalace ext] feed (${reason}) failed: ${(err as Error).message}\n`,
|
||||
@@ -1007,6 +1227,40 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
}
|
||||
};
|
||||
|
||||
/**
|
||||
* Is `later` strictly after `earlier` in the fleet's ordering?
|
||||
*
|
||||
* Prefer `hlc` — a hybrid logical clock rendered fixed-width
|
||||
* (`<unix_ms:13 digits>-<counter:6 hex>-<replica_id>`), so a plain string
|
||||
* comparison IS the causal comparison, and it stays correct once a second
|
||||
* replica exists. `seq` is the arrival rowid of THIS database, so on a mesh
|
||||
* the same event carries a different `seq` per replica and a reply can land
|
||||
* before the ask it answers.
|
||||
*
|
||||
* Fall back to `seq` only when either side lacks an `hlc` (a server predating
|
||||
* the field, or a row whose backfill did not run). On a single replica the two
|
||||
* agree, so the fallback is not a downgrade today — it is the single-replica
|
||||
* case still being handled after the mesh case became primary.
|
||||
*
|
||||
* NEVER compare `created_at`: it is server-generated at second precision, so
|
||||
* ties are routine, and a tie or a skew there can suppress an UNANSWERED ask
|
||||
* outright. Every failure mode here is deliberately kept on the
|
||||
* noisy-but-visible side — an answered item resurfacing is annoying, an
|
||||
* unanswered ask going silent defeats the mailbox.
|
||||
*/
|
||||
const isStrictlyAfter = (later: LogEvent, earlier: LogEvent): boolean => {
|
||||
if (
|
||||
typeof later.hlc === "string" &&
|
||||
typeof earlier.hlc === "string" &&
|
||||
later.hlc !== "" &&
|
||||
earlier.hlc !== ""
|
||||
) {
|
||||
return later.hlc > earlier.hlc;
|
||||
}
|
||||
if (typeof later.seq !== "number" || typeof earlier.seq !== "number") return false;
|
||||
return later.seq > earlier.seq;
|
||||
};
|
||||
|
||||
/**
|
||||
* Has one of MY events closed this candidate?
|
||||
*
|
||||
@@ -1023,15 +1277,8 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
// it one terminal reply suppresses every LATER ask on the same
|
||||
// correlation_id forever — silently, permanently, and worst on exactly
|
||||
// the long-running threads the correlation join is for.
|
||||
//
|
||||
// Compare `seq`, NEVER `created_at`. `seq` is replica-local (it equals
|
||||
// origin_seq only while a single replica authors for every machine); the
|
||||
// durable key once `mempalace_mesh_peers` reports real peers is `hlc`,
|
||||
// which is total and causally consistent. Local-seq skew can only make an
|
||||
// ANSWERED item resurface (noise, visible), whereas a timestamp
|
||||
// comparison can suppress an UNANSWERED ask outright.
|
||||
if (typeof m.seq !== "number" || typeof candidate.seq !== "number") return false;
|
||||
if (m.seq <= candidate.seq) return false;
|
||||
// See isStrictlyAfter for why the key is `hlc` and not `seq`/`created_at`.
|
||||
if (!isStrictlyAfter(m, candidate)) return false;
|
||||
// (b) join on the exact ack (written for us by event_ack), else on a
|
||||
// shared correlation_id — which is why correlation_id is mandatory on a
|
||||
// directed open: without it there is no key to join a reply back to.
|
||||
@@ -1040,17 +1287,120 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
return Boolean(m.correlation_id && m.correlation_id === candidate.correlation_id);
|
||||
});
|
||||
|
||||
/** Directed asks addressed to this device with no terminal reply from it. */
|
||||
async function deriveOwed(): Promise<LogEvent[]> {
|
||||
if (!mailboxEnabled || !available) return [];
|
||||
/**
|
||||
* Has the ORIGINAL REQUESTER explicitly withdrawn this candidate?
|
||||
*
|
||||
* RFC 003 §3.3 clause 3 clears an ask only on "no event **of yours**", and the
|
||||
* skill states the consequence outright: "there is nothing anyone can do about
|
||||
* it from the other end". That asymmetry is deliberate and mostly right — owed
|
||||
* ness is a statement about the RECIPIENT's accountability, and a requester
|
||||
* must not be able to delete an obligation the recipient genuinely has.
|
||||
*
|
||||
* It is wrong in exactly one case: the requester withdrawing its OWN ask.
|
||||
*
|
||||
* MEASURED COST, 2026-09-08/09. pi@mbp-m1-2020 withdrew a v1.8.13 rollout ask
|
||||
* to pi@tor-ms22 (evt seq 119: terminal `superseded`, same correlation_id,
|
||||
* `metadata.closes` naming it, body "DO NOT SPEND A MINUTE ON v1.8.13") and
|
||||
* recorded the withdrawal as done. It had no effect: seq 119's `from_agent` is
|
||||
* mbp, so it could never satisfy a join that only looks at tor-ms22's own
|
||||
* events. tor-ms22's next wake-up — 8h later, on the first boot of the image
|
||||
* shipping this very derivation — still listed the ask as owed, 41h old, for a
|
||||
* release that had been superseded and never installed there. It had to spend a
|
||||
* write closing an ask nobody wanted answered. The asymmetry is also INVISIBLE
|
||||
* from the sender's side: mbp did everything a sender is told to do and got a
|
||||
* result indistinguishable from success.
|
||||
*
|
||||
* WHY AN EXPLICIT MARKER AND NOT "ANY TERMINAL EVENT FROM THE REQUESTER".
|
||||
* This file's standing rule is that every failure mode stays on the
|
||||
* noisy-but-visible side, because a spuriously resurfacing item is one turn of
|
||||
* human correction whereas a suppressed unanswered ask is silent and permanent.
|
||||
* Inferring release from any terminal event would breach that: a requester
|
||||
* appending `applied` for its own bookkeeping — on the correlation, addressed
|
||||
* to me, before I ever replied — would silently delete a real obligation.
|
||||
* So release must be STATED, not inferred. `withdraws` is the canonical key;
|
||||
* `closes` is honoured because it is already this fleet's de-facto marker (mbp
|
||||
* seq 119, emb seq 117/118 all use it), and either must name this exact ask —
|
||||
* its `correlation_id` or its event `id`. Prose does not count.
|
||||
*
|
||||
* WHY THIS DOES NOT BREAK THE FIXED POINT IN RFC 003 §9.2. §9.2 rejects letting
|
||||
* terminal directed events into the owed set, because then every closure mints a
|
||||
* fresh obligation and the loop never terminates. That argument is about
|
||||
* CANDIDATES. This adds a CLEARER. A withdrawal carries a terminal status, so it
|
||||
* can never be a candidate, and nothing new becomes owed. The asserting shape
|
||||
* (`open`) and the clearing shape (terminal) stay disjoint. A reader who reaches
|
||||
* for §9.2 to object is answering a different question.
|
||||
*/
|
||||
const isWithdrawn = (candidate: LogEvent, inbound: LogEvent[]): boolean => {
|
||||
const requester = candidate.from_agent;
|
||||
if (!requester) return false;
|
||||
// Names THIS ask specifically: its correlation or its id. A marker naming
|
||||
// something else, or carrying prose, is not a withdrawal of this ask.
|
||||
const names = (v: unknown): boolean =>
|
||||
typeof v === "string" &&
|
||||
v.length > 0 &&
|
||||
((Boolean(candidate.correlation_id) && v === candidate.correlation_id) ||
|
||||
(Boolean(candidate.id) && v === candidate.id));
|
||||
return inbound.some((e) => {
|
||||
// Only the party that ASKED may retract. A third party writing a terminal
|
||||
// event on a shared correlation must never clear my obligation.
|
||||
if (e.from_agent !== requester) return false;
|
||||
// Addressed to the party being released. `to_agent=<me>` also matches '*'
|
||||
// at the SQL level (RFC 003 §7.4), and a broadcast must not be able to
|
||||
// quietly empty every machine's mailbox at once.
|
||||
if (e.to_agent !== mailboxAddress) return false;
|
||||
if (!TERMINAL_STATUS.has((e.status ?? "").toLowerCase())) return false;
|
||||
// Same ordering guard, same reason as isAnswered: a withdrawal must not
|
||||
// retire an ask the requester sent LATER on the same correlation.
|
||||
if (!isStrictlyAfter(e, candidate)) return false;
|
||||
const ackOf = e.metadata?.ack_of;
|
||||
const joins =
|
||||
(typeof ackOf === "string" && Boolean(candidate.id) && ackOf === candidate.id) ||
|
||||
Boolean(e.correlation_id && e.correlation_id === candidate.correlation_id);
|
||||
if (!joins) return false;
|
||||
return names(e.metadata?.withdraws) || names(e.metadata?.closes);
|
||||
});
|
||||
};
|
||||
|
||||
/**
|
||||
* Directed asks addressed to this device with no terminal reply from it, and
|
||||
* not explicitly withdrawn by whoever sent them (see isWithdrawn).
|
||||
*/
|
||||
async function deriveOwed(): Promise<OwedSplit> {
|
||||
if (!mailboxEnabled || !available) return { owed: [], dormant: [] };
|
||||
try {
|
||||
const [candidatesRaw, mineRaw] = await Promise.all([
|
||||
const [candidatesRaw, mineRaw, inboundRaw] = await Promise.all([
|
||||
// Every cursor-less event_list call in this file passes `order` explicitly.
|
||||
// The server's default is not ours to lean on: mempalace <= 3.9.0 defaults
|
||||
// to `asc`, 3.10.0 flips a cursor-less listing to newest-first. Both
|
||||
// derive* joins want the NEWEST window (see the comment below), so say so
|
||||
// and the verdict stops depending on which server version answers.
|
||||
client.callTool("mempalace_event_list", {
|
||||
to_agent: mailboxAddress,
|
||||
status: "open",
|
||||
limit: 50,
|
||||
order: "desc",
|
||||
}),
|
||||
// `order: "desc"` is a fix, not a flourish. event_list defaults to `asc`
|
||||
// (append order), so this asked for the OLDEST 100 events this device
|
||||
// ever wrote — meaning that once a device passes 100 authored events its
|
||||
// most RECENT replies fall out of the join window and every ask it just
|
||||
// answered resurfaces as owed. Latent rather than theoretical: tor-ms22
|
||||
// was at ~20 authored events when this was written. The window has to be
|
||||
// anchored at the newest end for the same reason the join uses `hlc`.
|
||||
client.callTool("mempalace_event_list", {
|
||||
from_agent: mailboxAddress,
|
||||
limit: 100,
|
||||
order: "desc",
|
||||
}),
|
||||
// Inbound with NO status filter, deliberately: a withdrawal carries a
|
||||
// TERMINAL status and is therefore structurally invisible to the
|
||||
// `status: "open"` candidate query above — the same omission deriveClosed
|
||||
// is built on. No amount of polling the open set could ever see one.
|
||||
client.callTool("mempalace_event_list", {
|
||||
to_agent: mailboxAddress,
|
||||
limit: 100,
|
||||
order: "desc",
|
||||
}),
|
||||
client.callTool("mempalace_event_list", { from_agent: mailboxAddress, limit: 100 }),
|
||||
]);
|
||||
const candidates = parseEvents(candidatesRaw).filter((c) => {
|
||||
// `to_agent: <me>` ALSO matches `*` broadcasts, per the tool contract.
|
||||
@@ -1064,11 +1414,80 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
// You cannot owe yourself.
|
||||
return c.from_agent !== mailboxAddress;
|
||||
});
|
||||
if (candidates.length === 0) return [];
|
||||
if (candidates.length === 0) return { owed: [], dormant: [] };
|
||||
const mine = parseEvents(mineRaw);
|
||||
return candidates.filter((c) => !isAnswered(c, mine));
|
||||
const inbound = parseEvents(inboundRaw);
|
||||
const live = candidates.filter((c) => !isAnswered(c, mine) && !isWithdrawn(c, inbound));
|
||||
// PARTITION, not filter. A dormant ask is still owed in the protocol
|
||||
// sense — it has no terminal reply and must not be closed — so it stays in
|
||||
// the return value where the mailbox can still account for it, and only
|
||||
// the ANNOUNCING paths below treat the two halves differently.
|
||||
const owed: LogEvent[] = [];
|
||||
const dormant: LogEvent[] = [];
|
||||
for (const c of live) (evaluateDormancy(c.metadata).dormant ? dormant : owed).push(c);
|
||||
return { owed, dormant };
|
||||
} catch {
|
||||
// Fail silent and open: a mailbox read must never break a session.
|
||||
return { owed: [], dormant: [] };
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Terminal replies to asks THIS device sent. NOT work — news.
|
||||
*
|
||||
* Why this is a second query rather than a widened deriveOwed(): deriveOwed()
|
||||
* asks the log for `status: "open"`, and a reply that CLOSES an ask is by
|
||||
* definition not open, so it is structurally invisible to that filter — no
|
||||
* amount of polling or waiting could ever have surfaced it.
|
||||
*
|
||||
* Measured 2026-09-07: emb-7kj4vr4g closed a v1.8.13 rollout ask with a
|
||||
* task.reply at status=applied, and the operator reasonably expected to be
|
||||
* told. The mailbox stayed silent and was CORRECT to — "owed" means "you must
|
||||
* reply", and nothing was owed. But the single most useful thing a fleet can
|
||||
* tell a human is "the thing you asked for is done" (here: done, AND your
|
||||
* premise was wrong), and that was the one category the feature could not
|
||||
* report. Silence was right by its own definition and wrong by the user's.
|
||||
*/
|
||||
async function deriveClosed(): Promise<LogEvent[]> {
|
||||
if (!mailboxEnabled || !available) return [];
|
||||
try {
|
||||
// `order: "desc"` for the same reason as deriveOwed: without it a 3.9.0
|
||||
// server hands back the OLDEST window, so past 100 authored events this
|
||||
// device's newest asks (and past 50 inbound, the replies that close them)
|
||||
// fall outside the join. The selection below is order-independent
|
||||
// (isStrictlyAfter), so this only fixes WHICH events are in the window.
|
||||
const [mineRaw, inboundRaw] = await Promise.all([
|
||||
client.callTool("mempalace_event_list", { from_agent: mailboxAddress, limit: 100, order: "desc" }),
|
||||
// NO status filter, deliberately: that omission is the entire fix.
|
||||
client.callTool("mempalace_event_list", { to_agent: mailboxAddress, limit: 50, order: "desc" }),
|
||||
]);
|
||||
// Correlations this device opened as a DIRECTED ask. A broadcast owes
|
||||
// nobody a reply, so it cannot be closed by one either.
|
||||
const myAsks = new Map<string, LogEvent>();
|
||||
for (const e of parseEvents(mineRaw)) {
|
||||
if (!e.correlation_id || !e.to_agent) continue;
|
||||
if (e.to_agent === "*" || e.to_agent === mailboxAddress) continue;
|
||||
if (e.type !== "task.request") continue;
|
||||
myAsks.set(e.correlation_id, e);
|
||||
}
|
||||
if (myAsks.size === 0) return [];
|
||||
// Newest terminal reply per correlation only. A peer that appends
|
||||
// applied-then-superseded should cost one line, not a wall of them.
|
||||
const best = new Map<string, LogEvent>();
|
||||
for (const e of parseEvents(inboundRaw)) {
|
||||
if (e.from_agent === mailboxAddress) continue; // cannot inform myself
|
||||
const cid = e.correlation_id;
|
||||
if (!cid || !myAsks.has(cid)) continue;
|
||||
// NOTE the deliberate asymmetry with deriveOwed: a broadcast is excluded
|
||||
// there (it owes nobody) but allowed here, because a peer answering on MY
|
||||
// correlation_id is news to me regardless of how widely it was addressed.
|
||||
if (!TERMINAL_STATUS.has((e.status ?? "").toLowerCase())) continue;
|
||||
const prev = best.get(cid);
|
||||
if (!prev || isStrictlyAfter(e, prev)) best.set(cid, e);
|
||||
}
|
||||
return [...best.values()];
|
||||
} catch {
|
||||
// Same contract as deriveOwed: a mailbox read never breaks a session.
|
||||
return [];
|
||||
}
|
||||
}
|
||||
@@ -1105,11 +1524,41 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
})
|
||||
.join("\n");
|
||||
|
||||
const DORMANT_NOTE =
|
||||
"Each of these declares a `dormant_unless` predicate in its metadata that STILL " +
|
||||
"matches this device's current state, so the work it describes is not yet " +
|
||||
"possible. They are NOT closed and NOT answered — leave them open. They return to " +
|
||||
"the announced owed-set by themselves the moment a baseline stops matching. If a " +
|
||||
"predicate cannot be evaluated at all (missing file, unparseable baseline, unknown " +
|
||||
"kind) the ask is announced as owed instead, deliberately: dormancy has to be " +
|
||||
"proven, never assumed.";
|
||||
|
||||
const OWED_HOWTO =
|
||||
"To close one, append an event on the SAME correlation_id with a terminal status " +
|
||||
"(applied/superseded/failed/blocked). An ack alone does NOT clear it, and neither " +
|
||||
"does claimed/ready — those deliberately keep it visible as taken-but-unfinished.";
|
||||
|
||||
// Written for the ONE reader the rest of this text ignores: the human watching
|
||||
// the window. Measured 2026-08-26 on two devices — a delivery lands, the agent
|
||||
// is idle, nothing happens, and the operator asks "do I have to nudge you for
|
||||
// you to read this?". Yes, and it is by design (see the deliverAs comment
|
||||
// below): at agent_settled no inference is running, so the text is queued for
|
||||
// the next turn. The message explained how to CLOSE an ask but never when it
|
||||
// would be SEEN, so the person who needed that fact was the only one not told.
|
||||
// Cheapest possible fix, and it changes no behaviour: say so in the text.
|
||||
const QUEUED_NOTE =
|
||||
"DELIVERY NOTE — this is a QUEUED message, not an action: it arrived while this " +
|
||||
"agent was idle, so no model call was made and nothing woke it. It is read at the " +
|
||||
"start of the next turn. If you are the human watching and want it handled now, " +
|
||||
"send any message to start a turn; the agent is not ignoring the ask, it is not running.";
|
||||
|
||||
const CLOSED_NOTE =
|
||||
"NO ACTION IS OWED on these — they are replies to asks THIS device sent, shown " +
|
||||
"once each because 'the thing you asked for is done' is news you wanted and the " +
|
||||
"owed-set could never carry it. Read the reply before assuming your original ask " +
|
||||
"was right: a peer that did the work is the most likely party to have found your " +
|
||||
"premise wrong.";
|
||||
|
||||
// Mid-session mailbox. Tier 1 (the wake-up injection below) owns the first
|
||||
// look; this exists because arrivals are bursty and correlate with our own
|
||||
// activity — measured 2026-08-26, 11 of 22 events on this log landed inside a
|
||||
@@ -1121,7 +1570,68 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
// the owed set holds something not already surfaced. Re-announcing the same
|
||||
// ask every poll is precisely how a channel teaches its reader to ignore it,
|
||||
// which is the failure this whole feature exists to reverse.
|
||||
pi.on("agent_settled", async () => {
|
||||
// --- A: tell the human at poll time (RFC 003 §7.11) -----------------------
|
||||
//
|
||||
// B (the QUEUED_NOTE above) fixes the confusion of someone who IS looking at
|
||||
// the window. It does nothing for the case that actually loses an ask: nobody
|
||||
// is looking. The delivered text is queued for a turn that only a human starts,
|
||||
// so an ask can wait indefinitely on an idle session while its recipient makes
|
||||
// coffee. A notification at poll time is the only part of this feature that
|
||||
// reaches a person who is not watching.
|
||||
//
|
||||
// Two surfaces, deliberately split by risk:
|
||||
// default in-TUI ctx.ui.notify, exactly as session_start already does.
|
||||
// Zero new output channels; visible only if you are looking.
|
||||
// =desktop additionally emit a terminal-native notification, reusing
|
||||
// the detection the fleet's own notify.ts already proves in
|
||||
// this harness (Kitty OSC 99 / Windows toast / OSC 777). This
|
||||
// escapes the container without notify-send, DBus or any host
|
||||
// access: the escape sequence is interpreted by the terminal
|
||||
// emulator on the human's machine. Opt-in because writing raw
|
||||
// escapes to stdout is a behaviour change on shared machines,
|
||||
// not because it is unreliable.
|
||||
// =kitty =osc777 force one protocol, skipping detection entirely.
|
||||
// =0 silent, for anyone who wants the mailbox without pings.
|
||||
//
|
||||
// WHY FORCING EXISTS — measured on tor-ms22 2026-08-26, and it invalidates
|
||||
// detection in exactly the deployment this ships in. Terminal identity lives in
|
||||
// env vars set by the emulator (KITTY_WINDOW_ID, TERM_PROGRAM) and `docker exec`
|
||||
// does NOT forward them: inside the container pi sees only TERM=xterm-256color
|
||||
// no matter what is rendering it. So detection can never see Kitty from in here
|
||||
// and always falls through to OSC 777, which Kitty does not implement — on a
|
||||
// containerised client the notification would silently do nothing, the worst
|
||||
// possible failure for a feature whose whole job is to break a silence. Naming
|
||||
// the protocol ends the guessing.
|
||||
//
|
||||
// Fired only when something is actually due, i.e. the same condition as the
|
||||
// delivery itself: a notification that fires on an empty poll would teach its
|
||||
// reader to ignore it, which is the failure this whole feature exists to undo.
|
||||
const mailboxNotifyMode = (process.env.MEMPALACE_MAILBOX_NOTIFY ?? "").trim().toLowerCase();
|
||||
const mailboxNotifyEnabled = mailboxNotifyMode !== "0" && mailboxNotifyMode !== "off";
|
||||
const mailboxNotifyTerminal =
|
||||
mailboxNotifyMode === "desktop" ||
|
||||
mailboxNotifyMode === "kitty" ||
|
||||
mailboxNotifyMode === "osc777";
|
||||
|
||||
/** Terminal-native notification. Best-effort; never throws into a handler. */
|
||||
const notifyTerminal = (title: string, body: string): void => {
|
||||
try {
|
||||
const esc = (s: string): string => s.replace(/[\x00-\x1f\x7f;]/g, " ");
|
||||
if (
|
||||
mailboxNotifyMode === "kitty" ||
|
||||
(mailboxNotifyMode === "desktop" && !!process.env.KITTY_WINDOW_ID)
|
||||
) {
|
||||
process.stdout.write(`\x1b]99;i=1:d=0;${esc(title)}\x1b\\`);
|
||||
process.stdout.write(`\x1b]99;i=1:p=body;${esc(body)}\x1b\\`);
|
||||
return;
|
||||
}
|
||||
process.stdout.write(`\x1b]777;notify;${esc(title)};${esc(body)}\x07`);
|
||||
} catch {
|
||||
/* a notification is never worth an exception */
|
||||
}
|
||||
};
|
||||
|
||||
pi.on("agent_settled", async (_event, ctx) => {
|
||||
if (!mailboxEnabled || !available) return;
|
||||
if (!wokeUp) return; // the wake-up injection has not run yet; it does look #1
|
||||
if (Date.now() - lastMailboxPollAt < mailboxPollMs) return;
|
||||
@@ -1130,22 +1640,48 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
lastMailboxPollAt = Date.now();
|
||||
void (async () => {
|
||||
try {
|
||||
const owed = await deriveOwed();
|
||||
const [owedSplit, closed] = await Promise.all([deriveOwed(), deriveClosed()]);
|
||||
const { owed, dormant } = owedSplit;
|
||||
const now = Date.now();
|
||||
const due = owed.filter((e) => {
|
||||
if (!e.id) return false;
|
||||
const last = surfaced.get(e.id);
|
||||
return last === undefined || now - last >= mailboxResurfaceMs;
|
||||
});
|
||||
if (due.length === 0) return; // silence is the correct output here
|
||||
for (const e of due) if (e.id) surfaced.set(e.id, now);
|
||||
const unseen = (list: LogEvent[]): LogEvent[] =>
|
||||
list.filter((e) => {
|
||||
if (!e.id) return false;
|
||||
const last = surfaced.get(e.id);
|
||||
return last === undefined || now - last >= mailboxResurfaceMs;
|
||||
});
|
||||
const due = unseen(owed);
|
||||
// Closing replies are announced ONCE and never resurface: an ask you
|
||||
// already know is finished is not a nag, and re-announcing it hourly is
|
||||
// the train-the-reader-to-ignore-it failure this window exists to stop.
|
||||
const newsRaw = closed.filter((e) => e.id && !surfaced.has(e.id));
|
||||
if (due.length === 0 && newsRaw.length === 0) return; // silence is correct
|
||||
for (const e of [...due, ...newsRaw]) if (e.id) surfaced.set(e.id, now);
|
||||
const blocks: string[] = [];
|
||||
if (due.length > 0) {
|
||||
blocks.push(
|
||||
`${due.length} directed ask(s) addressed to "${mailboxAddress}" with no ` +
|
||||
`terminal reply from this device. ${OWED_HOWTO}\n\n${formatOwed(due)}`,
|
||||
);
|
||||
}
|
||||
if (newsRaw.length > 0) {
|
||||
blocks.push(
|
||||
`${newsRaw.length} repl(y|ies) CLOSING an ask this device sent. ` +
|
||||
`${CLOSED_NOTE}\n\n${formatOwed(newsRaw)}`,
|
||||
);
|
||||
}
|
||||
// A dormant ask NEVER opens this window on its own — the early return
|
||||
// above is what actually stops the nag. But once the window is open for
|
||||
// something else, one line of count is nearly free and keeps a waiting
|
||||
// ask from feeling lost between session starts.
|
||||
if (dormant.length > 0) {
|
||||
blocks.push(`${dormant.length} dormant ask(s), not listed here. ${DORMANT_NOTE}`);
|
||||
}
|
||||
pi.sendMessage(
|
||||
{
|
||||
customType: "mempalace-mailbox",
|
||||
content:
|
||||
`MemPalace logstream mailbox: ${due.length} directed ask(s) addressed to ` +
|
||||
`"${mailboxAddress}" with no terminal reply from this device. ${OWED_HOWTO}\n\n` +
|
||||
formatOwed(due),
|
||||
`MemPalace logstream mailbox.\n\n${QUEUED_NOTE}\n\n` +
|
||||
blocks.join("\n\n---\n\n"),
|
||||
display: true,
|
||||
},
|
||||
// "steer" is queue-on-arrival, NOT an interrupt: at agent_settled the
|
||||
@@ -1157,6 +1693,23 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
// than auto-delivery and is not what was approved.
|
||||
{ deliverAs: "steer" },
|
||||
);
|
||||
// AFTER the queueing call, so a notification can never be the only
|
||||
// thing that happened: the message is in the thread first, then we
|
||||
// point at it. Wording names the nudge explicitly, because "you have
|
||||
// mail" without "press a key" reproduces the exact confusion B fixes.
|
||||
if (mailboxNotifyEnabled) {
|
||||
const parts = [
|
||||
due.length > 0 ? `${due.length} ask(s) owed` : "",
|
||||
newsRaw.length > 0 ? `${newsRaw.length} closed` : "",
|
||||
].filter(Boolean);
|
||||
const summary = `${parts.join(" + ")} for ${mailboxAddress} — send any message to handle`;
|
||||
try {
|
||||
if (ctx?.hasUI) ctx.ui.notify(`MemPalace mailbox: ${summary}`, "info");
|
||||
} catch {
|
||||
/* ignore: UI may be gone by the time the poll resolves */
|
||||
}
|
||||
if (mailboxNotifyTerminal) notifyTerminal("MemPalace mailbox", summary);
|
||||
}
|
||||
} catch {
|
||||
/* best-effort delivery: never break a turn over a mailbox read */
|
||||
}
|
||||
@@ -1187,10 +1740,11 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
sections.push(`## mempalace_diary_read\n\n(error: ${(err as Error).message})`);
|
||||
}
|
||||
// Tier 1 mailbox: what this device owes a reply to, derived rather than read
|
||||
// off a filter. deriveOwed() swallows its own failures and returns [], so a
|
||||
// broken palace costs a missing section, never a broken wake-up.
|
||||
// off a filter. deriveOwed() swallows its own failures and returns an empty
|
||||
// split, so a broken palace costs a missing section, never a broken wake-up.
|
||||
if (mailboxEnabled) {
|
||||
const owed = await deriveOwed();
|
||||
const [owedSplit, closed] = await Promise.all([deriveOwed(), deriveClosed()]);
|
||||
const { owed, dormant } = owedSplit;
|
||||
if (owed.length > 0) {
|
||||
// Record what this injection showed, so the first mid-session poll does
|
||||
// not re-announce the identical list minutes later. Without this the
|
||||
@@ -1204,6 +1758,27 @@ export default async function mempalaceExtension(pi: ExtensionAPI) {
|
||||
`this device. ${OWED_HOWTO}\n\n${formatOwed(owed)}`,
|
||||
);
|
||||
}
|
||||
// Once per session, WITH ids. The wake-up injection is the one place a
|
||||
// dormant ask should be visible: "what is this device still carrying?" is a
|
||||
// session-start question, and answering it here is what keeps the mid-session
|
||||
// silence honest rather than concealing. Deliberately NOT added to
|
||||
// `surfaced`: that map exists to suppress repeats of ANNOUNCED work, and a
|
||||
// dormant ask must stay announceable the instant its baseline moves.
|
||||
if (dormant.length > 0) {
|
||||
sections.push(
|
||||
`## logstream mailbox (${dormant.length} dormant — waiting, nothing owed yet)\n\n` +
|
||||
`${DORMANT_NOTE}\n\n${formatOwed(dormant)}`,
|
||||
);
|
||||
}
|
||||
const news = closed.filter((e) => e.id && !surfaced.has(e.id));
|
||||
if (news.length > 0) {
|
||||
const now = Date.now();
|
||||
for (const e of news) if (e.id) surfaced.set(e.id, now);
|
||||
sections.push(
|
||||
`## logstream mailbox (${news.length} closed — no action owed)\n\n` +
|
||||
`${CLOSED_NOTE}\n\n${formatOwed(news)}`,
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
if (sections.length === 0) return;
|
||||
|
||||
Executable
+156
@@ -0,0 +1,156 @@
|
||||
#!/usr/bin/env bash
|
||||
# test-dormancy.sh — exercise the mailbox dormancy predicate (dormant_unless).
|
||||
#
|
||||
# Why this exists as a script and not a note: `evaluateDormancy` decides whether
|
||||
# an ask is SHOWN to the agent, so a silent regression there does not look like a
|
||||
# bug — it looks like an empty mailbox. The positive arms matter as much as the
|
||||
# negative ones: a predicate that never fires makes the feature a no-op, and a
|
||||
# predicate that fires too eagerly hides real work. Both arms run here.
|
||||
#
|
||||
# It cannot import extensions/pi/mempalace.ts in place, because that file imports
|
||||
# `typebox`, which pi provides at RUNTIME and this repo has no node_modules for.
|
||||
# So it copies the file into a temp tree with the resolved deps symlinked beside
|
||||
# it. The copy is made BY this script on every run, so it cannot drift from the
|
||||
# source the way a vendored duplicate would.
|
||||
#
|
||||
# Fixtures are created here rather than read from the host: the first draft used
|
||||
# /etc/hostname and /etc/pi-devbox/build-manifest.json, which made the positive
|
||||
# arms pass only on a pi-devbox container and silently flip to "not dormant"
|
||||
# anywhere else — a device-dependent test that reports success by doing nothing.
|
||||
#
|
||||
# Exit codes, deliberately distinct:
|
||||
# 0 every case behaved as specified
|
||||
# 1 at least one case FAILED (a real regression)
|
||||
# 2 INCONCLUSIVE — deps could not be resolved, so nothing was proven
|
||||
# (never 0: a test that skips quietly is the failure mode it should catch)
|
||||
|
||||
set -uo pipefail
|
||||
|
||||
# ── Args ──────────────────────────────────────────────────────────────────────
|
||||
if [[ "${1:-}" == "--help" || "${1:-}" == "-h" ]]; then
|
||||
sed -n '2,25p' "$0" | sed 's/^# \?//'
|
||||
exit 0
|
||||
fi
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
SRC="$REPO_ROOT/extensions/pi/mempalace.ts"
|
||||
[[ -r "$SRC" ]] || {
|
||||
echo "INCONCLUSIVE: cannot read $SRC" >&2
|
||||
exit 2
|
||||
}
|
||||
|
||||
# ── Resolve pi's runtime deps (discovered, never hardcoded) ───────────────────
|
||||
GLOBAL_ROOT="$(npm root -g 2>/dev/null || true)"
|
||||
PI_PKG=""
|
||||
for cand in \
|
||||
"${GLOBAL_ROOT:+$GLOBAL_ROOT/@earendil-works/pi-coding-agent}" \
|
||||
"/usr/lib/node_modules/@earendil-works/pi-coding-agent" \
|
||||
"/usr/local/lib/node_modules/@earendil-works/pi-coding-agent"; do
|
||||
[[ -n "$cand" && -d "$cand" ]] && {
|
||||
PI_PKG="$cand"
|
||||
break
|
||||
}
|
||||
done
|
||||
[[ -n "$PI_PKG" && -d "$PI_PKG/node_modules/typebox" ]] || {
|
||||
echo "INCONCLUSIVE: pi-coding-agent / typebox not resolvable (looked under 'npm root -g')." >&2
|
||||
echo " Nothing was proven. Install pi, or run this where pi is installed." >&2
|
||||
exit 2
|
||||
}
|
||||
|
||||
# Node must be able to strip types from a .ts entry point (Node >= 22.6).
|
||||
node -e 'process.exit(0)' 2>/dev/null || {
|
||||
echo "INCONCLUSIVE: no usable node" >&2
|
||||
exit 2
|
||||
}
|
||||
|
||||
# ── Build the throwaway tree ──────────────────────────────────────────────────
|
||||
WORK="$(mktemp -d)"
|
||||
trap 'rm -rf "$WORK"' EXIT
|
||||
mkdir -p "$WORK/extensions/pi" "$WORK/node_modules/@earendil-works"
|
||||
cp "$SRC" "$WORK/extensions/pi/mempalace.ts"
|
||||
ln -s "$PI_PKG/node_modules/typebox" "$WORK/node_modules/typebox"
|
||||
ln -s "$PI_PKG" "$WORK/node_modules/@earendil-works/pi-coding-agent"
|
||||
printf '{"type":"module"}\n' > "$WORK/package.json"
|
||||
|
||||
# Prove the copy is the source, so a PASS cannot be about a stale file.
|
||||
if ! cmp -s "$SRC" "$WORK/extensions/pi/mempalace.ts"; then
|
||||
echo "INCONCLUSIVE: the copy differs from the source" >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
# ── Fixtures, owned by this test ──────────────────────────────────────────────
|
||||
printf '{"release_tag":"v1.9.4","nested":{"deep":{"leaf":"found"}},"count":3}\n' > "$WORK/manifest.json"
|
||||
printf 'not json at all\n' > "$WORK/notjson.txt"
|
||||
: > "$WORK/stamp"
|
||||
|
||||
# ── The truth table ───────────────────────────────────────────────────────────
|
||||
cat > "$WORK/run.mjs" <<'EOF'
|
||||
import { statSync } from "node:fs";
|
||||
import { evaluateDormancy } from "./extensions/pi/mempalace.ts";
|
||||
|
||||
const W = process.env.WORK;
|
||||
const STAMP = `${W}/stamp`;
|
||||
const MAN = `${W}/manifest.json`;
|
||||
// Floor to whole seconds: that is the contract, and the residue below proves
|
||||
// why the implementation must do the same.
|
||||
const atBaseline = new Date(Math.floor(statSync(STAMP).mtimeMs / 1000) * 1000).toISOString();
|
||||
|
||||
const mt = (baseline, path = STAMP) => ({ kind: "file_mtime", path, baseline });
|
||||
const jf = (field, baseline, path = MAN) => ({ kind: "json_field", path, field, baseline });
|
||||
|
||||
const cases = [
|
||||
// No predicate => behave exactly as before the feature existed.
|
||||
["no metadata", undefined, false],
|
||||
["metadata without the key", { foo: "bar" }, false],
|
||||
["dormant_unless is a string", { dormant_unless: "release != x" }, false],
|
||||
["dormant_unless is empty", { dormant_unless: [] }, false],
|
||||
["more than 8 conditions", { dormant_unless: Array(9).fill(mt(atBaseline)) }, false],
|
||||
|
||||
// POSITIVE arms. If these ever read false the feature is inert.
|
||||
["file_mtime at baseline", { dormant_unless: [mt(atBaseline)] }, true],
|
||||
["json_field at baseline", { dormant_unless: [jf("release_tag", "v1.9.4")] }, true],
|
||||
["both at baseline", { dormant_unless: [jf("release_tag", "v1.9.4"), mt(atBaseline)] }, true],
|
||||
["dotted field path", { dormant_unless: [jf("nested.deep.leaf", "found")] }, true],
|
||||
["number vs string baseline", { dormant_unless: [jf("count", "3")] }, true],
|
||||
|
||||
// The trigger firing: ANY condition differing wakes the ask.
|
||||
["file_mtime moved", { dormant_unless: [mt("2020-01-01T00:00:00Z")] }, false],
|
||||
["json_field moved", { dormant_unless: [jf("release_tag", "v1.9.5")] }, false],
|
||||
["one same, one moved", { dormant_unless: [mt(atBaseline), jf("release_tag", "v1.9.5")] }, false],
|
||||
|
||||
// FAIL-VISIBLE arms: dormancy unproven => announce.
|
||||
["missing file", { dormant_unless: [mt(atBaseline, "/nonexistent/path")] }, false],
|
||||
["unparseable baseline", { dormant_unless: [mt("not-a-date")] }, false],
|
||||
["relative path", { dormant_unless: [{ kind: "file_mtime", path: "etc/hostname", baseline: atBaseline }] }, false],
|
||||
["unknown kind", { dormant_unless: [{ kind: "uptime_lt", path: "/etc/hostname", baseline: "1d" }] }, false],
|
||||
["missing json field", { dormant_unless: [jf("no_such_field", "x")] }, false],
|
||||
["file is not JSON", { dormant_unless: [jf("a", "b", `${W}/notjson.txt`)] }, false],
|
||||
["field is not scalar", { dormant_unless: [jf("nested", "x")] }, false],
|
||||
["condition is not an object", { dormant_unless: ["release_tag"] }, false],
|
||||
["baseline is an object", { dormant_unless: [{ kind: "file_mtime", path: STAMP, baseline: {} }] }, false],
|
||||
];
|
||||
|
||||
let pass = 0;
|
||||
const failed = [];
|
||||
for (const [name, meta, want] of cases) {
|
||||
const v = evaluateDormancy(meta);
|
||||
const ok = v.dormant === want;
|
||||
if (ok) pass++;
|
||||
else failed.push(name);
|
||||
console.log(
|
||||
` ${ok ? "PASS" : "FAIL"} dormant=${String(v.dormant).padEnd(5)} want=${String(want).padEnd(5)} ` +
|
||||
`${name.padEnd(28)} :: ${v.reason}`,
|
||||
);
|
||||
}
|
||||
console.log(`\n ${pass}/${cases.length} passed`);
|
||||
if (failed.length) console.log(` FAILED: ${failed.join(", ")}`);
|
||||
// Recorded because it is the reason file_mtime floors to seconds: a real mtime
|
||||
// carries sub-second residue that an ISO baseline does not.
|
||||
console.log(` (stamp mtimeMs=${statSync(STAMP).mtimeMs}, baseline=${atBaseline})`);
|
||||
process.exit(failed.length === 0 ? 0 : 1);
|
||||
EOF
|
||||
|
||||
cd "$WORK" || exit 2
|
||||
export WORK
|
||||
node "$WORK/run.mjs"
|
||||
exit $?
|
||||
Executable
+134
@@ -0,0 +1,134 @@
|
||||
#!/usr/bin/env bash
|
||||
# test-mcp-call-timeout.sh — the per-call deadline of RemoteMcpClient in
|
||||
# extensions/pi/mempalace.ts, and the override that the feed's mine relies on.
|
||||
#
|
||||
# WHY THIS EXISTS. The feed tick calls `mempalace_mine` through
|
||||
# `client.callTool()`. Its own deadline (MEMPALACE_FEED_MINE_TIMEOUT_MS,
|
||||
# 300 000) was raised in 2026-09 because the mine legitimately takes 30-60 s on
|
||||
# a shared single-writer palace — but callTool() carried no way to pass that
|
||||
# deadline down, so the transport's generic per-request timeout
|
||||
# (MEMPALACE_MCP_TIMEOUT_MS, 60 000) fired first, and operators saw
|
||||
# feed (tick) failed: mempalace remote request 'tools/call' failed: timed out after 60000ms
|
||||
# instead of the message the 2026-09 change had aimed at. The outer race was
|
||||
# unreachable in practice. This harness pins both halves of the contract:
|
||||
# 1. a plain callTool() still honours the short per-request timeout
|
||||
# (a query taking that long really is wedged — keep it short);
|
||||
# 2. callTool(name, args, { timeoutMs }) honours the override, so a long
|
||||
# server-side job can be given its own, longer, deadline.
|
||||
# Before the fix, (2) fails: the third argument was silently ignored.
|
||||
#
|
||||
# Like test-owed-withdrawal.sh, it runs the SHIPPED text: the class is cut out
|
||||
# of mempalace.ts by brace matching, type-stripped with node's own stripper,
|
||||
# and driven against a local JSON-RPC server that delays tools/call.
|
||||
set -euo pipefail
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
SRC="${1:-$REPO_ROOT/extensions/pi/mempalace.ts}"
|
||||
WORK="$(mktemp -d)"
|
||||
trap 'rm -rf "$WORK"' EXIT
|
||||
[ -r "$SRC" ] || { echo "FAIL: cannot read $SRC" >&2; exit 2; }
|
||||
|
||||
# ---------------------------------------------------------------- extractor ---
|
||||
cat >"$WORK/extract.mjs" <<'EXTRACT'
|
||||
import { readFileSync, writeFileSync } from "node:fs";
|
||||
import { stripTypeScriptTypes } from "node:module";
|
||||
const src = readFileSync(process.argv[2], "utf8");
|
||||
|
||||
/** Cut a top-level `const NAME = ...;` out of the source, verbatim. */
|
||||
function decl(name) {
|
||||
const start = src.indexOf(`const ${name} =`);
|
||||
if (start < 0) throw new Error(`declaration not found: ${name}`);
|
||||
return src.slice(start, span(start));
|
||||
}
|
||||
/** Cut `class NAME ... { ... }` out of the source, verbatim. */
|
||||
function klass(name) {
|
||||
const start = src.indexOf(`class ${name} `);
|
||||
if (start < 0) throw new Error(`class not found: ${name}`);
|
||||
return src.slice(start, span(start, true));
|
||||
}
|
||||
/** End offset of the statement starting at `start` (brace-matched). */
|
||||
function span(start, braceOnly = false) {
|
||||
let depth = 0, inStr = null, sawBrace = false;
|
||||
for (let i = start; i < src.length; i++) {
|
||||
const c = src[i], prev = src[i - 1];
|
||||
if (inStr) { if (c === inStr && prev !== "\\") inStr = null; continue; }
|
||||
if (c === '"' || c === "'" || c === "`") { inStr = c; continue; }
|
||||
if (c === "/" && src[i + 1] === "/") { i = src.indexOf("\n", i); if (i < 0) break; continue; }
|
||||
if (c === "/" && src[i + 1] === "*") { i = src.indexOf("*/", i) + 1; continue; }
|
||||
if (c === "{" || c === "(" || c === "[") { depth++; if (c === "{") sawBrace = true; }
|
||||
else if (c === "}" || c === ")" || c === "]") { depth--; if (braceOnly && sawBrace && depth === 0) return i + 1; }
|
||||
else if (!braceOnly && c === ";" && depth === 0) return i + 1;
|
||||
}
|
||||
throw new Error(`unterminated statement at ${start}`);
|
||||
}
|
||||
|
||||
const parts = [decl("num"), decl("REMOTE_PROTOCOL_VERSION"), decl("REMOTE_CLIENT_INFO"), klass("RemoteMcpClient")];
|
||||
for (const [i, p] of parts.entries()) if (p.length < 30) throw new Error(`extraction ${i} implausibly short: ${p}`);
|
||||
if (!parts[3].includes("tools/call")) throw new Error("RemoteMcpClient does not mention tools/call");
|
||||
const ts = parts.join("\n\n") + "\nexport { RemoteMcpClient };\n";
|
||||
writeFileSync(process.argv[3], stripTypeScriptTypes(ts, { mode: "strip" }));
|
||||
EXTRACT
|
||||
node --no-warnings "$WORK/extract.mjs" "$SRC" "$WORK/client.mjs" || exit 2
|
||||
node --check "$WORK/client.mjs" || { echo "FAIL: extracted client does not parse" >&2; exit 2; }
|
||||
echo "[extract] pulled RemoteMcpClient from $(basename "$SRC") ($(wc -c <"$WORK/client.mjs") bytes)"
|
||||
|
||||
# ------------------------------------------------------------------ harness ---
|
||||
cat >"$WORK/run.mjs" <<'RUN'
|
||||
import { createServer } from "node:http";
|
||||
import { RemoteMcpClient } from "./client.mjs";
|
||||
|
||||
// A sessionless JSON-RPC server like `mempalace-mcp --transport http`:
|
||||
// initialize / tools/list answer at once; tools/call sleeps SLOW_MS first.
|
||||
const SLOW_MS = 1500;
|
||||
const server = createServer((req, res) => {
|
||||
let body = "";
|
||||
req.on("data", (c) => (body += c));
|
||||
req.on("end", () => {
|
||||
const msg = JSON.parse(body);
|
||||
if (msg.id === undefined) { res.writeHead(202); res.end(); return; } // notification
|
||||
const reply = (result) => {
|
||||
res.writeHead(200, { "content-type": "application/json" });
|
||||
res.end(JSON.stringify({ jsonrpc: "2.0", id: msg.id, result }));
|
||||
};
|
||||
if (msg.method === "initialize") return reply({ protocolVersion: "2024-11-05", capabilities: {}, serverInfo: { name: "fake", version: "0" } });
|
||||
if (msg.method === "tools/list") return reply({ tools: [{ name: "slow", description: "", inputSchema: { type: "object" } }] });
|
||||
if (msg.method === "tools/call") return void setTimeout(() => reply({ content: [{ type: "text", text: "done" }] }), SLOW_MS);
|
||||
reply({});
|
||||
});
|
||||
});
|
||||
await new Promise((r) => server.listen(0, "127.0.0.1", r));
|
||||
const url = `http://127.0.0.1:${server.address().port}/mcp`;
|
||||
|
||||
let failures = 0;
|
||||
const check = (ok, label) => { console.log(`${ok ? "ok " : "FAIL"} ${label}`); if (!ok) failures++; };
|
||||
|
||||
// Short per-request deadline, deliberately below SLOW_MS.
|
||||
process.env.MEMPALACE_MCP_TIMEOUT_MS = "400";
|
||||
const client = new RemoteMcpClient(url);
|
||||
await client.start();
|
||||
check(client.alive === true, "start(): initialize + tools/list complete, client alive");
|
||||
|
||||
// 1. A plain call honours the short deadline — this is the wedged-query guard.
|
||||
let err = null;
|
||||
const t0 = Date.now();
|
||||
try { await client.callTool("slow", {}); } catch (e) { err = e; }
|
||||
check(err !== null && /timed out after 400ms/.test(err.message), `plain callTool() rejects at the per-request deadline (${err && err.message})`);
|
||||
check(Date.now() - t0 < SLOW_MS, "…and rejects BEFORE the server would have answered");
|
||||
check(client.alive === false, "a timeout marks the client unhealthy (documented side effect; ensureAlive() revives)");
|
||||
|
||||
// 2. A call with its own deadline outlives the generic one — the feed's mine.
|
||||
await client.ensureAlive();
|
||||
err = null;
|
||||
let result = null;
|
||||
try { result = await client.callTool("slow", {}, { timeoutMs: 5000 }); } catch (e) { err = e; }
|
||||
check(err === null && result && result.content?.[0]?.text === "done", `callTool(name, args, { timeoutMs: 5000 }) waits past the generic deadline and gets the result (${err ? err.message : "ok"})`);
|
||||
|
||||
// 3. The override is itself a deadline, not "no deadline".
|
||||
err = null;
|
||||
try { await client.callTool("slow", {}, { timeoutMs: 200 }); } catch (e) { err = e; }
|
||||
check(err !== null && /timed out after 200ms/.test(err.message), `callTool(..., { timeoutMs: 200 }) rejects at ITS deadline (${err && err.message})`);
|
||||
|
||||
server.close();
|
||||
console.log(failures === 0 ? "PASS: all assertions held" : `FAIL: ${failures} assertion(s) failed`);
|
||||
process.exit(failures === 0 ? 0 : 1);
|
||||
RUN
|
||||
cd "$WORK" && node --no-warnings run.mjs
|
||||
Executable
+312
@@ -0,0 +1,312 @@
|
||||
#!/usr/bin/env bash
|
||||
# test-owed-withdrawal.sh — rule tests for the owed-set derivation in
|
||||
# extensions/pi/mempalace.ts: isStrictlyAfter, isAnswered, isWithdrawn.
|
||||
#
|
||||
# WHY THIS EXISTS
|
||||
# isWithdrawn changes who is allowed to clear an obligation, which is the one
|
||||
# thing in this extension that can go wrong SILENTLY. A wrong rule here does
|
||||
# not throw and does not show up in a log — it makes a real unanswered ask
|
||||
# vanish from a mailbox forever. So the rules get pinned down before they ship.
|
||||
#
|
||||
# WHY IT EXTRACTS THE PREDICATES FROM THE SOURCE INSTEAD OF RESTATING THEM
|
||||
# These are closures inside createExtension(), so they cannot be imported. The
|
||||
# tempting shortcut is to paste a copy of the logic into the test — which tests
|
||||
# the copy. This repo has already paid for a divergent second copy (the
|
||||
# pi-extensions skill mirror, 9579 B behind for weeks; the shell-lint logic
|
||||
# extracted to one file for the same reason). So the harness cuts the actual
|
||||
# declaration text out of mempalace.ts by brace matching and runs THAT. Change
|
||||
# the source and this test follows; paraphrase the source and it cannot.
|
||||
#
|
||||
# FIXTURES ARE VERBATIM REAL EVENTS from the fleet log (project/pi-devbox), not
|
||||
# invented shapes — including the exact incident that motivated isWithdrawn:
|
||||
# mbp-m1-2020 withdrew a v1.8.13 rollout ask to tor-ms22 at seq 119 and the
|
||||
# withdrawal had no effect, so tor-ms22 was still being told it owed a reply 41h
|
||||
# later for a release it never installed.
|
||||
#
|
||||
# Usage: scripts/test-owed-withdrawal.sh [source.ts]
|
||||
# exit 0 = all rules behave 1 = a rule broke
|
||||
# exit 2 = source unreadable or does not parse 3 = the gate itself cannot run
|
||||
#
|
||||
# The optional argument exists so the suite can be pointed at a deliberately
|
||||
# MUTATED copy of the source to prove it is sensitive — a suite that has never
|
||||
# been observed to fail is not evidence. See the mutation check in the CHANGELOG
|
||||
# entry that introduced isWithdrawn.
|
||||
set -euo pipefail
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
SRC="${1:-$REPO_ROOT/extensions/pi/mempalace.ts}"
|
||||
WORK="$(mktemp -d)"
|
||||
trap 'rm -rf "$WORK"' EXIT
|
||||
|
||||
[ -r "$SRC" ] || { echo "FAIL: cannot read $SRC" >&2; exit 2; }
|
||||
|
||||
# A gate that cannot run must not pass — the standing rule in this repo.
|
||||
#
|
||||
# NOT `node --check`: it does not type-strip, so it rejects ANY TypeScript —
|
||||
# `const x: number = 1` included, not just the inline type-import at the top of
|
||||
# mempalace.ts. It passed on node 22.x and stopped passing on node 24.x
|
||||
# (measured on pi-devbox v1.9.1, node v24.21.0: this gate exited 2 on an
|
||||
# unmodified mempalace.ts, so all 17 assertions below refused to run). Strip
|
||||
# first, then syntax-check the emitted JS. `mode: "strip"` blanks type syntax
|
||||
# without moving anything, so byte offsets and line numbers survive and a
|
||||
# reported error line still points at the right line of the ORIGINAL .ts.
|
||||
#
|
||||
# The two failure modes are reported separately and on purpose. "Cannot strip"
|
||||
# is a fact about the toolchain; "does not parse" is a fact about the source.
|
||||
# Collapsing them is what made this very defect present itself as
|
||||
# "mempalace.ts does not parse" when mempalace.ts was fine.
|
||||
#
|
||||
# SC2016 is disabled deliberately: the single quotes are the point. What follows
|
||||
# is JavaScript, and `${process.version}` must reach node, not be expanded by the
|
||||
# shell first.
|
||||
# shellcheck disable=SC2016
|
||||
node --no-warnings -e '
|
||||
const { readFileSync, writeFileSync } = require("node:fs");
|
||||
const { stripTypeScriptTypes } = require("node:module");
|
||||
if (typeof stripTypeScriptTypes !== "function") {
|
||||
console.error(`FAIL: node ${process.version} cannot strip TypeScript ` +
|
||||
`(module.stripTypeScriptTypes needs >= 22.13) — the gate is unavailable, ` +
|
||||
`which is NOT a statement about the source`);
|
||||
process.exit(3);
|
||||
}
|
||||
let js;
|
||||
try {
|
||||
js = stripTypeScriptTypes(readFileSync(process.argv[1], "utf8"), { mode: "strip" });
|
||||
} catch (err) {
|
||||
// The stripper is itself a parser, so a syntax error lands HERE, not in the
|
||||
// --check below. This is a statement about the source: exit 2, not 3.
|
||||
console.error(`FAIL: ${process.argv[1]} does not parse: ${err.message}`);
|
||||
process.exit(2);
|
||||
}
|
||||
writeFileSync(process.argv[2], js);
|
||||
' "$SRC" "$WORK/stripped.mjs" || exit $?
|
||||
node --check "$WORK/stripped.mjs" \
|
||||
|| { echo "FAIL: $SRC does not parse (stripped output rejected)" >&2; exit 2; }
|
||||
|
||||
# ---------------------------------------------------------------- extractor ---
|
||||
cat >"$WORK/extract.mjs" <<'EXTRACT'
|
||||
import { readFileSync, writeFileSync } from "node:fs";
|
||||
|
||||
const src = readFileSync(process.argv[2], "utf8");
|
||||
|
||||
/**
|
||||
* Cut `const <name> = ...;` out of the source by matching delimiters, so the
|
||||
* test runs the shipped text. Returns the declaration verbatim.
|
||||
*/
|
||||
function decl(name) {
|
||||
const start = src.indexOf(`const ${name} =`);
|
||||
if (start < 0) throw new Error(`declaration not found in source: ${name}`);
|
||||
let depth = 0;
|
||||
let inStr = null;
|
||||
for (let i = start; i < src.length; i++) {
|
||||
const c = src[i];
|
||||
const prev = src[i - 1];
|
||||
if (inStr) {
|
||||
if (c === inStr && prev !== "\\") inStr = null;
|
||||
continue;
|
||||
}
|
||||
if (c === '"' || c === "'" || c === "`") { inStr = c; continue; }
|
||||
if (c === "/" && src[i + 1] === "/") { i = src.indexOf("\n", i); if (i < 0) break; continue; }
|
||||
if (c === "{" || c === "(" || c === "[") depth++;
|
||||
else if (c === "}" || c === ")" || c === "]") depth--;
|
||||
else if (c === ";" && depth === 0) return src.slice(start, i + 1);
|
||||
}
|
||||
throw new Error(`unterminated declaration: ${name}`);
|
||||
}
|
||||
|
||||
const parts = ["TERMINAL_STATUS", "isStrictlyAfter", "isAnswered", "isWithdrawn"].map(decl);
|
||||
|
||||
// Sanity: the extractor must have found real bodies, not empty matches. A
|
||||
// silently-empty extraction would make every assertion below pass vacuously.
|
||||
for (const [i, p] of parts.entries()) {
|
||||
if (p.length < 40) throw new Error(`extracted declaration ${i} is implausibly short: ${p}`);
|
||||
}
|
||||
if (!parts[3].includes("withdraws")) throw new Error("isWithdrawn does not mention its marker key");
|
||||
|
||||
writeFileSync(process.argv[3], parts.join("\n\n"));
|
||||
EXTRACT
|
||||
|
||||
node "$WORK/extract.mjs" "$SRC" "$WORK/extracted.ts"
|
||||
echo "[extract] pulled 4 declarations from $(basename "$SRC") ($(wc -c <"$WORK/extracted.ts") bytes)"
|
||||
|
||||
# ------------------------------------------------------------------ harness ---
|
||||
{
|
||||
cat <<'HEAD'
|
||||
type LogEvent = {
|
||||
id?: string;
|
||||
seq?: number;
|
||||
hlc?: string;
|
||||
type?: string;
|
||||
status?: string;
|
||||
from_agent?: string;
|
||||
to_agent?: string;
|
||||
correlation_id?: string | null;
|
||||
created_at?: string;
|
||||
body?: string;
|
||||
metadata?: Record<string, unknown> | null;
|
||||
};
|
||||
|
||||
const mailboxAddress = "pi@tor-ms22";
|
||||
|
||||
HEAD
|
||||
cat "$WORK/extracted.ts"
|
||||
cat <<'TAIL'
|
||||
|
||||
// ---- VERBATIM fixtures from the real fleet log, stream project/pi-devbox ----
|
||||
const REP = "rep_d344e349ba276d6fc11997cd552f6937";
|
||||
|
||||
/** seq 112 — mbp's v1.8.13 rollout ask to tor-ms22. The obligation. */
|
||||
const ask112: LogEvent = {
|
||||
id: "evt_20260907T125110_2bef6f9b07a8",
|
||||
seq: 112,
|
||||
hlc: `1788785470242-000000-${REP}`,
|
||||
type: "task.request",
|
||||
status: "open",
|
||||
from_agent: "pi@mbp-m1-2020",
|
||||
to_agent: "pi@tor-ms22",
|
||||
correlation_id: "v1813-client-rollout-tor-ms22",
|
||||
};
|
||||
|
||||
/** seq 119 — mbp's withdrawal of its OWN ask. Terminal, directed, marker present. */
|
||||
const withdraw119: LogEvent = {
|
||||
id: "evt_20260908T223949_8f5b9bec7c92",
|
||||
seq: 119,
|
||||
hlc: `1788907189286-000000-${REP}`,
|
||||
type: "task.reply",
|
||||
status: "superseded",
|
||||
from_agent: "pi@mbp-m1-2020",
|
||||
to_agent: "pi@tor-ms22",
|
||||
correlation_id: "v1813-client-rollout-tor-ms22",
|
||||
metadata: {
|
||||
closes: "v1813-client-rollout-tor-ms22",
|
||||
nothing_owed: "no reply required to this event",
|
||||
replacement: "v1814-client-rollout-tor-ms22",
|
||||
},
|
||||
};
|
||||
|
||||
/** seq 120 — the LIVE v1.8.14 ask. Must survive the v1813 withdrawal. */
|
||||
const ask120: LogEvent = {
|
||||
id: "evt_20260908T224118_808de43d6982",
|
||||
seq: 120,
|
||||
hlc: `1788907278075-000000-${REP}`,
|
||||
type: "task.request",
|
||||
status: "open",
|
||||
from_agent: "pi@mbp-m1-2020",
|
||||
to_agent: "pi@tor-ms22",
|
||||
correlation_id: "v1814-client-rollout-tor-ms22",
|
||||
};
|
||||
|
||||
/** seq 122 — tor-ms22's OWN terminal reply, which is what actually cleared 112. */
|
||||
const myReply122: LogEvent = {
|
||||
id: "evt_20260909T061851_cf9bd18800db",
|
||||
seq: 122,
|
||||
hlc: `1788934731146-000000-${REP}`,
|
||||
type: "task.reply",
|
||||
status: "superseded",
|
||||
from_agent: "pi@tor-ms22",
|
||||
to_agent: "pi@mbp-m1-2020",
|
||||
correlation_id: "v1813-client-rollout-tor-ms22",
|
||||
metadata: { closes: "v1813-client-rollout-tor-ms22 — from the RECIPIENT side, which is the only side that can" },
|
||||
};
|
||||
|
||||
const clone = (e: LogEvent, over: Partial<LogEvent>): LogEvent => ({ ...e, ...over });
|
||||
|
||||
let failed = 0;
|
||||
const check = (name: string, expected: boolean, actual: boolean): void => {
|
||||
const ok = expected === actual;
|
||||
if (!ok) failed++;
|
||||
console.log(`${ok ? " ok " : " FAIL"} ${name} (expected ${expected}, got ${actual})`);
|
||||
};
|
||||
|
||||
console.log("\nisWithdrawn — the requester retracting its own ask");
|
||||
// THE INCIDENT. Before this rule existed the answer was false and tor-ms22 was
|
||||
// told it owed a reply for a release that no longer existed.
|
||||
check("real seq 119 withdraws real seq 112", true, isWithdrawn(ask112, [withdraw119]));
|
||||
check("withdrawal does NOT touch the live v1814 ask", false, isWithdrawn(ask120, [withdraw119]));
|
||||
|
||||
console.log("\nisWithdrawn — the ways a mailbox must NOT be silently emptied");
|
||||
check(
|
||||
"no marker: a bare terminal event from the requester is not a withdrawal",
|
||||
false,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { metadata: { nothing_owed: "no reply required" } })]),
|
||||
);
|
||||
check(
|
||||
"marker naming a DIFFERENT thread does not clear this one",
|
||||
false,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { metadata: { closes: "v1814-client-rollout-tor-ms22" } })]),
|
||||
);
|
||||
check(
|
||||
"marker carrying PROSE does not count (tor-ms22's own seq 122 shape)",
|
||||
false,
|
||||
isWithdrawn(ask112, [
|
||||
clone(withdraw119, {
|
||||
metadata: { closes: "v1813-client-rollout-tor-ms22 — from the RECIPIENT side, which is the only side that can" },
|
||||
}),
|
||||
]),
|
||||
);
|
||||
check(
|
||||
"a THIRD PARTY cannot retract someone else's ask",
|
||||
false,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { from_agent: "pi@emb-7kj4vr4g" })]),
|
||||
);
|
||||
check(
|
||||
"a BROADCAST cannot empty every machine's mailbox at once",
|
||||
false,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { to_agent: "*" })]),
|
||||
);
|
||||
check(
|
||||
"a NON-TERMINAL status is not a withdrawal",
|
||||
false,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { status: "open" })]),
|
||||
);
|
||||
check(
|
||||
"an event addressed to a DIFFERENT device does not clear my ask",
|
||||
false,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { to_agent: "pi@emb-7kj4vr4g" })]),
|
||||
);
|
||||
check(
|
||||
"ORDERING: a withdrawal cannot retire an ask the requester sent LATER",
|
||||
false,
|
||||
isWithdrawn(clone(ask112, { seq: 121, hlc: `1788907300000-000000-${REP}` }), [withdraw119]),
|
||||
);
|
||||
|
||||
console.log("\nisWithdrawn — accepted marker spellings");
|
||||
check(
|
||||
"canonical `withdraws` naming the correlation",
|
||||
true,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { metadata: { withdraws: "v1813-client-rollout-tor-ms22" } })]),
|
||||
);
|
||||
check(
|
||||
"marker naming the ask's EVENT ID instead of its correlation",
|
||||
true,
|
||||
isWithdrawn(ask112, [clone(withdraw119, { metadata: { withdraws: ask112.id } })]),
|
||||
);
|
||||
|
||||
console.log("\nisAnswered — regression guard, unchanged behaviour");
|
||||
check("my own terminal reply still closes my own ask", true, isAnswered(ask112, [myReply122]));
|
||||
check("my reply on one thread does not close another", false, isAnswered(ask120, [myReply122]));
|
||||
check(
|
||||
"ORDERING still load-bearing: my reply cannot pre-close a LATER ask",
|
||||
false,
|
||||
isAnswered(clone(ask112, { seq: 130, hlc: `1788999999999-000000-${REP}` }), [myReply122]),
|
||||
);
|
||||
|
||||
console.log("\nEND TO END — the derived owed set for tor-ms22 on 2026-09-09");
|
||||
// The behaviour change, stated as the derivation states it. `mine` deliberately
|
||||
// EXCLUDES seq 122 so this measures the new rule rather than the reply that
|
||||
// happened to be written first.
|
||||
const candidates = [ask112, ask120];
|
||||
const mine: LogEvent[] = [];
|
||||
const inbound = [withdraw119];
|
||||
const owed = candidates.filter((c) => !isAnswered(c, mine) && !isWithdrawn(c, inbound));
|
||||
const owedIds = owed.map((e) => e.correlation_id).join(", ");
|
||||
check("exactly one ask remains owed", true, owed.length === 1);
|
||||
check("and it is the v1814 one, not the withdrawn v1813", true, owedIds === "v1814-client-rollout-tor-ms22");
|
||||
console.log(` derived owed set: [${owedIds}]`);
|
||||
|
||||
console.log(failed === 0 ? "\n=== PASSED ===" : `\n=== FAILED: ${failed} ===`);
|
||||
process.exit(failed === 0 ? 0 : 1);
|
||||
TAIL
|
||||
} >"$WORK/harness.ts"
|
||||
|
||||
node --experimental-strip-types "$WORK/harness.ts"
|
||||
Executable
+88
@@ -0,0 +1,88 @@
|
||||
#!/usr/bin/env bash
|
||||
# test-rsync-ship-idempotency.sh — regression test for the ship-step bug fixed
|
||||
# 2026-08-27 (bin/mempalace-pi-session: rsync --update -> --checksum).
|
||||
#
|
||||
# THE BUG: the stage file's mtime is deliberately the SOURCE transcript's mtime
|
||||
# (os.utime() at :903, "preserve session mtime for dedup stability"), so a
|
||||
# re-export of a session that has not been appended to since the last ship
|
||||
# carries a mtime that is NOT newer than the receiver copy. rsync --update
|
||||
# skips a file whose mtime is not strictly newer than the destination's, so a
|
||||
# scrubbed re-export of a DORMANT session was silently dropped: content
|
||||
# differed, mtime did not, "success" was reported, and unscrubbed bytes stayed
|
||||
# in the palace host's inbox indefinitely. Measured on mbp-m1-2020 2026-08-27.
|
||||
#
|
||||
# THE ASSERTION mirrors that exactly and needs no ssh, no palace, no secret:
|
||||
# stage a file, ship it, rewrite the content while RESTORING the original
|
||||
# mtime (the same thing os.utime() does), ship again, assert the receiver's
|
||||
# sha256 changed. On the pre-fix flag (--update) this fails; with --checksum
|
||||
# (content comparison, mtime-independent) it passes.
|
||||
#
|
||||
# Deliberately does NOT invoke bin/mempalace-pi-session itself: that needs a
|
||||
# live ssh target, a palace, MEMPALACE_* env — disproportionate scaffolding for
|
||||
# what is, at its core, one rsync flag. Pulls the flag list out of the script
|
||||
# instead of hand-copying it, so a future change to the ship command either
|
||||
# updates this test's expectation or fails it loudly rather than drifting
|
||||
# silently out of sync with what actually ships.
|
||||
#
|
||||
# Usage: scripts/test-rsync-ship-idempotency.sh (exit 0 = fix still holds)
|
||||
set -euo pipefail
|
||||
|
||||
REPO_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
SHIP_SCRIPT="$REPO_ROOT/bin/mempalace-pi-session"
|
||||
SRC="$(mktemp -d)"; DST="$(mktemp -d)"
|
||||
trap 'rm -rf "$SRC" "$DST"' EXIT
|
||||
|
||||
EXPECT='rsync -a --checksum --no-owner --no-group'
|
||||
if ! grep -qF "$EXPECT" "$SHIP_SCRIPT"; then
|
||||
echo "FAIL: $SHIP_SCRIPT no longer ships with '$EXPECT'." >&2
|
||||
echo " Either the fix regressed, or the flags changed -- update this" >&2
|
||||
echo " test's EXPECT string to match, don't just delete the test." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
ship() {
|
||||
rsync -a --checksum --no-owner --no-group \
|
||||
--include='*.jsonl' --exclude='*' \
|
||||
"$SRC/" "$DST/"
|
||||
}
|
||||
|
||||
# Same BYTE LENGTH on purpose, mirroring mbp-m1-2020's own reasoning for why
|
||||
# dropping --update outright would not be enough: rsync's default quick check
|
||||
# (no --update, no --checksum) already transfers when SIZE differs, so a test
|
||||
# with mismatched lengths would pass by accident and prove nothing about
|
||||
# --checksum specifically. A real redaction whose placeholder happens to match
|
||||
# the secret's length is exactly the case that stays silently undetected unless
|
||||
# content itself, not size or mtime, is compared.
|
||||
UNSCRUBBED='live: MEMPALACE_REMOTE_TOKEN=AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA'
|
||||
SCRUBBED='live: MEMPALACE_REMOTE_TOKEN=<redacted:MEMPALACE_REMOTE_TOKEN---------->'
|
||||
if [[ ${#UNSCRUBBED} -ne ${#SCRUBBED} ]]; then
|
||||
echo "FAIL: test fixture bug -- the two fixture strings must be equal length (${#UNSCRUBBED} vs ${#SCRUBBED}); this test would not exercise --checksum at all otherwise." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# 1. Baseline: ship a session, receiver now matches sender.
|
||||
printf '%s\n' "$UNSCRUBBED" > "$SRC/session.jsonl"
|
||||
touch -d '2026-08-23T23:53:18' "$SRC/session.jsonl"
|
||||
ship
|
||||
before="$(sha256sum "$DST/session.jsonl" | cut -d' ' -f1)"
|
||||
|
||||
# 2. The bug's exact shape: rewrite content (same length!) as a scrubbed
|
||||
# re-export would, then restore the ORIGINAL mtime, as the exporter's
|
||||
# os.utime() does.
|
||||
printf '%s\n' "$SCRUBBED" > "$SRC/session.jsonl"
|
||||
touch -d '2026-08-23T23:53:18' "$SRC/session.jsonl"
|
||||
|
||||
ship
|
||||
after="$(sha256sum "$DST/session.jsonl" | cut -d' ' -f1)"
|
||||
|
||||
if [[ "$before" == "$after" ]]; then
|
||||
echo "FAIL: receiver sha256 unchanged after a content-only re-export at a held-constant mtime." >&2
|
||||
echo " This is the exact defect --checksum was added to fix." >&2
|
||||
exit 1
|
||||
fi
|
||||
if ! grep -q '<redacted:MEMPALACE_REMOTE_TOKEN' "$DST/session.jsonl"; then
|
||||
echo "FAIL: receiver does not contain the scrubbed content." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "PASS: ship step transfers changed content even when mtime is deliberately held constant (before=$before after=$after)"
|
||||
Reference in New Issue
Block a user