309980b62c2ad6487ab5786f690c64a6a1930db8
81 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
309980b62c |
fix(pi-ext): stop the feed tick from launching overlapping mines
"[mempalace ext] feed (tick) failed: mine timed out after 30000ms" was parked as
cosmetic on 2026-08-27. It is not cosmetic: the tight deadline was the trigger,
but the defect is a positive feedback loop that puts multiple writers on a
single-writer palace.
lastFeedAt was assigned only AFTER a successful await. The Promise.race abandons
our WAIT and cannot cancel the server's work, so on a mine that takes longer
than the deadline -- measured at 30-60s in normal operation, against a 30s
deadline -- the catch ran with lastFeedAt UNCHANGED. That left both guards in
the agent_settled handler open at once: the debounce test
(`Date.now() - lastFeedAt < feedDebounceMs`) passed because lastFeedAt was still
stale, and feedInFlight was already null because `run` had settled. Every
subsequent settled turn therefore launched another mine on top of the one still
running, each making the next slower and the next timeout likelier -- which is
why operators saw the message many times per session instead of at most once per
10-minute debounce window.
Fix, three lines:
- move `lastFeedAt = Date.now()` to before the await, so a timeout still starts
the debounce clock. A timeout is not a "did not happen": the mine is running
server-side and `mine --mode convos` dedups by source_file and is idempotent.
- raise MEMPALACE_FEED_MINE_TIMEOUT_MS from 30_000 to 300_000. The mine is the
slowest thing this extension does yet carried the tightest deadline: 4x
tighter than the prepare step before it (120_000) and 10x tighter than the
init handshake (300_000), a fast call. All three were introduced together in
|
||
|
|
e2b060a940 |
feat(mailbox): let a requester withdraw its own ask, with an explicit marker
RFC 003 §3.3 clause 3 clears an ask only on "no event OF YOURS", and the skill states the consequence outright: "there is nothing anyone can do about it from the other end". That asymmetry is deliberate and mostly right — owed-ness is a statement about the RECIPIENT's accountability. It is wrong in exactly one case: the requester retracting its own ask. MEASURED COST. pi@mbp-m1-2020 withdrew a v1.8.13 rollout ask to pi@tor-ms22 at seq 119 — terminal `superseded`, same correlation_id, metadata.closes naming the thread, body "DO NOT SPEND A MINUTE ON v1.8.13" — and recorded it as done. It had no effect: seq 119's from_agent is mbp, so it could never satisfy a join that only inspects tor-ms22's own events. tor-ms22's next wake-up, 8h later and on the first boot of the image shipping this very derivation, still listed the ask as owed, 41h old, for a release it never installed. The failure is invisible from the sender's side, which is why it went unnoticed for 41h. WHY AN EXPLICIT MARKER RATHER THAN "ANY TERMINAL EVENT FROM THE REQUESTER". This file's standing rule is that every failure stays on the noisy-but-visible side: a resurfacing item costs one turn of human correction, a suppressed unanswered ask is silent and permanent. Under the naive rule a requester appending `applied` for its own bookkeeping — on the correlation, addressed to me, before I ever replied — would silently delete a real obligation. So release must be STATED: metadata.withdraws (canonical) or metadata.closes (already this fleet's de-facto marker), whose value must NAME the ask — its correlation_id or its event id. Prose does not count. Five further guards, each pinned by a mutation test: only the original requester; addressed to this device exactly, never '*'; terminal status; strictly after the ask by hlc; and joined by ack_of or correlation_id. This does NOT break the fixed point in §9.2. That section rejects letting terminal directed events into the owed set, because then every closure mints a fresh obligation. That is about CANDIDATES; this adds a CLEARER. A withdrawal is terminal, so it can never be a candidate, and the asserting shape (open) and the clearing shape (terminal) stay disjoint. Also fixed here, because the new rule depends on the same windowing: the `mine` query used event_list's DEFAULT `asc` order with limit 100, i.e. the OLDEST 100 events this device ever wrote. Once a device passes 100 authored events its most RECENT replies fall out of the join window and every ask it just answered resurfaces as owed. Latent, not theoretical — tor-ms22 was at ~20. Both windows are now anchored at the newest end with order: "desc". Verification: scripts/test-owed-withdrawal.sh, 17 assertions over VERBATIM fixtures from the real log (seq 112/119/120/122). It extracts the predicates from mempalace.ts by brace matching and runs the SHIPPED text rather than a pasted copy — this repo has already paid for a divergent second copy. Sensitivity proven by four mutations: removing the marker requirement flips exactly the 3 marker assertions, and disabling the third-party / broadcast / ordering guards each flip exactly their own. Control passes. |
||
|
|
e45f6b4301 |
feat(mailbox): surface replies that CLOSE this device's own asks
The mailbox could report what this device OWES, and structurally nothing else.
deriveOwed() queries the log with status:"open", and a reply that closes an ask is
by definition not open — so a peer answering my delegation was invisible to it at
every poll, forever, not just late.
Measured 2026-09-07: emb-7kj4vr4g closed correlation v1813-client-rollout-emb with
a task.reply at status=applied. Nothing was announced. The feature was CORRECT by
its own definition ("owed" = "you must reply", and nothing was owed) and wrong by
the operator's, who asked why no notification arrived. The single most useful thing
a fleet can say to a human is "the thing you asked for is done" — and that was the
one category it could not say.
Independent corroboration that this was a real gap and not a preference: emb's own
reply ends "The amd64 narrowing is filed as a drawer as well, SINCE A TERMINAL
EVENT REACHES NO MAILBOX." A peer had already diagnosed the hole and was routing
around it by hand.
deriveClosed(): my task.requests (correlation + directed) joined against inbound
events with NO status filter, keeping the newest TERMINAL_STATUS reply per
correlation. Verified by computing the exact predicate over the two real events:
the new path yields evt_20260907T183143 (applied); the old status:"open" query
yields 0. Announced once each (never resurfaced — a finished ask is not a nag),
and the copy warns that a peer who did the work is the likeliest party to have
found your premise wrong. Here both premises were wrong: mine that emb was amd64,
and emb's that tor-ms22 was therefore the last candidate.
Deliberate asymmetry, documented at the call site: a broadcast is excluded from the
owed set (it owes nobody) but allowed to CLOSE, since a peer answering on my
correlation_id is news however widely it was addressed.
|
||
|
|
21023e7aa0 |
rfc-003: an event addressed to a nobody is write-only
§7.13 — the owed set keys on to_agent, and from_agent is whatever string the writer puts there, so nothing stops authoring under an identity no live session runs as. When the reply comes back addressed to that string, it is stored and delivered to no one. Distinct from §7.12: the defect is the address, not the status. Measured 2026-08-27: two directed task.requests planted under a synthetic from_agent (a provenance label, not a run identity) drew two correct replies, addressed back to the label as the protocol requires. Neither reply was ever delivered. One carried a live-credential exposure finding; it sat unread for ~2h20m and was found only because a human asked whether mail had arrived. General rule stated, not just the instance: this fleet has already produced the same failure shape three ways (synthetic sender above; an event addressed to a decommissioned device; a device-identity case mismatch). Permitted exception carried over from existing fleet practice: a synthetic sender is fine for a controlled experiment, but the body must then name the real reply-to identity. Proposed (not implemented): warn at event_append when from_agent differs from the writer's own session identity, using a DISTINCT from_agent lookup to say whether anyone has ever authored under it — stated honestly as a heuristic that catches 'never authored, certainly unread' but not a decommissioned device that once did. Decision 9 in the open-decisions list gets a numbered sibling, 10, pointing at this section — related to decision 7 (agent-name collision between two devices) but a different failure: not two devices sharing a name, but one writer using a name nobody runs. |
||
|
|
a361b71c40 |
ship: don't trust the mtime the exporter deliberately backdates
rsync --update skips a file whose mtime is not strictly newer than the receiver's. The stage file's mtime IS the source transcript's mtime (os.utime() at :903, "preserve session mtime for dedup stability"), so re-exporting a session that has not been appended to since its last ship produces a mtime that is not newer than what's already at the receiver — exactly the case a redactor upgrade needs to ship, because content differs while mtime does not. --update reports success and sends nothing. Reported and patched by pi@mbp-m1-2020 (evt_20260827T211925_9674a31da0b4, artifact art_20260827T211839_fa52af563105, sha256 f217e47e…), measured live: a scrubbed re-export of a dormant session (pi_01a03022-…f542) sat unshipped in the palace host's inbox while every local signal reported a clean stage, saved only because a host-side sweep happened to rewrite the remote copy independently that same day. --checksum compares content and ignores size/mtime entirely. Dropping --update outright was considered and rejected: rsync's default quick check already transfers on a SIZE difference alone, which is why the observed case (33-byte placeholder vs a 43-byte token) would have been masked as "fixed" by a change that only works until a redaction whose placeholder happens to match the secret's length. os.utime() at :903 is untouched — its backdating is a separate, load-bearing design call for dedup stability, out of scope for this fix. Added scripts/test-rsync-ship-idempotency.sh: ships a file, rewrites its content to an EQUAL-LENGTH string while restoring the original mtime (what os.utime() does), ships again, asserts the receiver's sha256 changed. Equal length is deliberate, not cosmetic — mismatched lengths would pass via the quick check alone and prove nothing about --checksum specifically; this is the same reasoning that ruled out dropping --update. Verified the test discriminates: passes against today's --checksum, fails against --update (checked by temporarily substituting the flag in a copy, not committed). Runs offline — a local rsync destination path exercises the same size/mtime/checksum comparison as the ssh transfer, no palace or network needed. |
||
|
|
b2b50afcc1 |
rfc-003: propose a news surface that needs no cursor, and say why widening the owed set cannot work
Decision 9 (news vs obligations) gets a proposed direction in a new §9.2, following how §9.1 promoted the retention decision. The load-bearing part is the negative result. The tempting fix — let terminal directed events into the owed set so a report addressed to a device reaches it — breaks the derivation's fixed point. Owed-ness is "directed at me, not mine, and not joined by a later terminal event of mine", so the asserting shape (open) and the clearing shape (terminal) have to be disjoint. Make terminal events owed and a reply becomes owed by its requester, whose closure is itself directed + terminal and therefore owed by the original author: every closure mints a fresh obligation and the loop never terminates. status="open" is not editorial taste about tone, it is what makes owed-ness terminate. Worth writing down before someone "fixes" it. The proposal itself avoids the cost that sinks the naive version. "Since your last session" implies a per-device read cursor, and container-local state is exactly what --force-recreate erases. No new state is needed: the device's own last authored event is already a cursor, it lives in the shared log, and it is comparable across replicas for the same reason the owed-set join now uses hlc (§7.3). Named the non-obvious constraint too — the anchor must be computed PER STREAM, because a global one lets a chatty stream advance past unread news in a quiet one, and that failure is silent. Also tightened §7.12: candidacy requires exactly `open`, not merely "non-terminal", so a `claimed` announcement is as undelivered as a finished report. Claiming still earns its keep on handoff-prone work (it is what distinguishes "nobody started" from "someone started and the container died"), but its audience is a log reader, not the requester — and it does not quiet the claimer's own mailbox either, which is what the state machine's "open --> claimed does NOT clear" already implies. |
||
|
|
de59571966 |
docs: unclip fleet-memory's diagrams, and measure the host instead of guessing
Seven labels in fleet-memory.md were losing their last line for readers, in the
document held up as the model for user-facing docs. Found by running a new
mermaid checker against a repo it was not written for, which is the strongest
evidence available that this is a systemic defect class and not one author's slip.
Measured on the real Gitea render rather than a simulation. Gitea serves each
mermaid block from a sandboxed same-origin <iframe srcdoc>, so its
.markup{line-height:1.5!important} never reaches the labels and the page is
directly measurable. On the published version: 31 labels, 6 of them rendering
MORE lines than were authored, worst overflow +0.6px past the clip edge. Mermaid
measures a label with its own metrics, commits to a box, then clips whatever does
not fit; per-line rounding accumulates, so every cut label observed anywhere in
this fleet had four or five rendered lines.
Fixes, all preserving the information rather than deleting it:
- The five store nodes drop their third line; the content shapes they carried
("verbatim text, embedded", "typed facts, time-bounded", "addressed events and
their artifacts", "links between rooms", "read by recency") move into the prose
under the diagram, which is where a qualifier belongs anyway.
- "ask by ADDRESS + ORDER" soft-wrapped its own first line because a caps line
nearly as wide as the box overflows first; shortened to "ask by ADDRESS", with
"read in append order" added to the prose so the ordering property survives.
The table already lists "address, correlation, append order".
- The decision diamond loses "machine or agent" to the sentence above it, which
now asks whether "a specific machine or agent" must act — same words, no clip.
- "to_agent = the specific agent" and the BOTH node shortened; both are spelled
out in full in the worked-examples table directly below.
Nothing was cut for brevity's sake: every phrase removed from a node reappears in
the prose or was already in the table beneath it.
|
||
|
|
f0bffd1b93 |
feed: chase the symlink, or fail-closed takes the whole fleet's feed down
The image installs /usr/local/bin/mempalace-pi-session as a symlink into
/opt/mempalace-toolkit/bin, and ${BASH_SOURCE[0]} reports the path the script was
INVOKED as, not the resolved target. So the sibling-module lookup pointed at
/usr/local/bin, the redactor was not there, and the fail-closed import did exactly
what it was told: refused to stage.
MEASURED, not theorised: /usr/local/bin/mempalace-pi-session --dry-run exited
"[FATAL] secret scrubber unavailable ... refusing to stage" while
/opt/mempalace-toolkit/bin/mempalace-pi-session --dry-run scrubbed 40 findings on
the same input. The symlink is how every device invokes it, so at the next image
bake every feeder tick on every machine would have stopped staging — a silent,
fleet-wide memory outage, which is a worse outcome than the leak the scrubber
exists to prevent. My own tests missed it by calling bin/... directly from the
checkout, i.e. the one invocation path the fleet never uses.
FIX: chase the symlink chain in portable shell and offer fallbacks instead of
betting on a single answer. MEMPALACE_REDACT_DIR is now colon-separated —
resolved dir, invoked dir, then the image's known install path /opt/... — and the
Python side inserts the first existing candidate. readlink -f is deliberately NOT
used: it is GNU/newer-BSD only and this script also runs directly on macOS hosts,
so the chase is a plain while [ -L ] loop handling relative link targets.
Verified on all three invocation shapes: via the /usr/local/bin symlink, via the
direct /opt path, and via a second-hop symlink in an unrelated directory. All
three now report the same 40 redactions.
LESSON worth keeping: fail-closed is correct for a secret scrubber, but it
converts "module not found" into an outage, so the module lookup becomes
load-bearing infrastructure and must be tested through the real invocation path,
not the convenient one.
|
||
|
|
836e35b320 |
redact: name the tiers where they are used, and stop the docstring lying about tier 3
Review caught that the tier vocabulary was used in the report, the commit message and the docs without being defined anywhere the reader would land, and inspecting that turned up two real defects rather than just a wording gap. 1. STALE DOCSTRING. The module still described tier 3 as if it redacts, which stopped being true when the 403-hit measurement demoted it to report-only. It also credited tier 3 with resolving the 40-hex-PAT-vs-commit-sha collision — false by default, since a reporting rule resolves nothing. Corrected, with the consequence stated plainly: in the default configuration a sha-shaped PAT is caught if and only if it belongs to THIS machine, because only a known value (tier 1) or a naming key (tier 3, reporting) can separate it from a commit sha. That is an accepted gap; the alternative is redacting every sha in the palace. 2. THE VOCABULARY NEVER REACHED THE OUTPUT. The tool prints rule names (github-pat, env-value, url-credentials) and nothing printed a tier, so the docs' tier language was unconnected to what an operator actually sees. Added RULE_TIERS as the authoritative rule -> tier mapping, tier_of(), and Finding.tier; the feeder now prints "T2:github-pat=6" so what matched and how much to trust it are both visible on one line. A self-test asserts every rule that can appear in a Finding maps to a tier, so adding a rule without classifying it fails the tests instead of printing "T?". Tiers, for the record, are three kinds of EVIDENCE (not three severities): T1 known value from this process's env — near-certain, zero FP by construction; T2 known vendor shape — strong, the prefix is meaningful; T3 key name says secret — candidate only, measured FP-heavy, reported. T0 is reserved for suspicions(), which is a measured NON-detection. Docs gain worked one-line examples per tier and a "which tier fired?" section showing real output. 46 self-test cases pass. |
||
|
|
3d47937d06 |
feed: scrub secrets before staging a transcript, name-anchored not entropy-anchored
A palace is mined from transcripts, and transcripts contain whatever the terminal
printed. Measured on this fleet: one leaked bearer token had reached 3 drawers,
13 feeder inbox files across all three devices, and 10 local files spanning 10
days — from an agent inspecting an env var while debugging. That frequency is the
premise: this is a pipeline problem, not a discipline problem.
WHERE. bin/mempalace_redact.py, called from mempalace-pi-session at the point the
staged transcript is written. That single hook covers both transports, because
local mode mines the staged file and remote mode rsyncs that same file
byte-for-byte. Scrubbing operates on the parsed objects rather than the
serialized text, so string values are rewritten while keys, ids and structure are
untouched — scanning raw JSONL instead invents keys like "tapiKey" out of the \t
escape preceding a field name (observed, not theorised).
WHY NOT ENTROPY. The obvious "redact long random-looking strings" is actively
destructive here: drawer ids, chunk ids, event ids, replica ids, HLCs and commit
SHAs are all high-entropy, are the majority of random-looking text in a palace,
and redacting them is silent and permanent. Detection is anchored on meaning
instead: tier 1 literal values from this process's env whose NAME says secret
(zero false positives by construction, and the only tier that can tell a 40-hex
Gitea PAT from a git commit sha); tier 2 vendor-prefixed shapes (ghp_, glpat-,
xox*-, sk-, AKIA, AIza, hf_, JWT, PEM blocks, URL credentials, Authorization
headers); tier 3 name=value assignments.
THE MEASUREMENT THAT CHANGED THE DESIGN. Tier 3 was going to redact. Against
52 MB of real fleet transcripts it produced 403 hits, and inspection with values
masked showed most were ${VAR} interpolation in compose files, TypeScript
identifiers, a TYPE ANNOTATION (credentials: Credentials), an IPA attribute
holding a date (krbPasswordExpiration), AAAK diary shorthand, and terminal output
following a "Password:" prompt. Redacting those corrupts code and docs held as
memory to catch what tier 1 already catches by value. After adding guards for
interpolation, code context, non-secret key suffixes and all-digit values, the
enforced count fell 403 -> 29 on the same corpus. So tier 3 REPORTS and does not
rewrite unless MEMPALACE_REDACT_STRICT=1.
HONESTY ABOUT MISSES. Known false negatives are documented rather than papered
over: novel formats in bare prose, another machine's secrets, base64-of-a-secret,
line-split secrets. Every run prints a count including "0 redaction(s)", because
silence is indistinguishable from a scrubber that never ran, and suspicions()
reports high-entropy strings it did NOT redact as (length, fingerprint) so the
miss rate is measurable. Findings never carry the value — rule, label, length and
sha256[:8], enough to recognise a recurrence, not enough to recover the secret.
FAIL CLOSED: no scrubber, no staging (exit 3), overridable with
MEMPALACE_FEED_ALLOW_UNSCRUBBED=1 for a machine older than this file.
Tested: 42-case corpus in --self-test, including every palace id shape as a
must-not-redact case, idempotency, and a compound case the corpus caught where a
vendor placeholder was then re-matched by the URL rule (nested placeholder — the
secret was hidden either way, so only an exact-output assertion catches it).
End-to-end on 45 real sessions: 29 enforced, fail-closed verified at rc=3,
override verified loud. Server-side layer specified in docs/secret-hygiene.md §5
but NOT implemented — it is tier 2 only there, since the hub cannot see a
client's environment.
|
||
|
|
ecc2a9c574 |
mailbox notify: record that the terminal path is unverified through tmux
Measured on tor-ms22: the layering here is kitty -> tmux (on the host) -> docker exec -> pi (in the container), with two clients attached to one tmux session. Test OSC sequences written straight to pi's tty produced no notification on the remote client; the local client is still unobserved. Most likely tmux is dropping OSC types it does not implement — reaching the outer terminal needs tmux's DCS passthrough (ESC P tmux; ... ESC \, inner ESC doubled) plus allow-passthrough on, which is not the default and is not implemented here. The point worth keeping is structural, not incidental: the container can see NEITHER layer. KITTY_WINDOW_ID is absent because docker exec does not forward it, and TMUX is absent because tmux runs one level further out on the host. So a containerised client cannot detect the terminal it is speaking to or the multiplexer it must speak through, and autodetection is not merely unreliable here — it is blind. Explicit configuration is the only route, which is why the previous commit added forced kitty/osc777 modes. No behaviour or defaults changed: in-TUI notify remains the default and the terminal path stays opt-in. Also flagged the unsettled routing question — a pane's output reaches every attached client, so a work laptop and a home machine would both ping from one arriving ask. |
||
|
|
e91766286e |
mailbox notify: name the protocol, because docker exec hides the terminal
Auto-detection is useless in the deployment this ships in, measured on tor-ms22. Terminal identity lives in env vars set by the emulator (KITTY_WINDOW_ID, TERM_PROGRAM) and docker exec does not forward them — pi inside the container sees only TERM=xterm-256color no matter what is rendering it. So the =desktop detection can never see Kitty from in there and always falls through to OSC 777, which Kitty does not implement. Result: on a containerised client the ping silently does nothing, which is the worst available failure for a feature whose entire purpose is to break a silence. Adds MEMPALACE_MAILBOX_NOTIFY=kitty and =osc777 to force one protocol and skip detection. =desktop keeps the notify.ts-style autodetect for the non-container case, unset still means in-TUI only, 0/off still silent. The Windows toast branch is dropped from the comment rather than the code path it never had here: it needs powershell.exe on PATH, which is a WSL fact, not a Linux-container one. |
||
|
|
a92c75d070 |
mailbox: say it is queued, and ping the human who is not looking (RFC 003 §7.11)
Two changes, neither touching the no-triggerTurn decision, which stands.
B — a delivery note in the message itself. The mid-session text explained how to
CLOSE an ask and never said when it would be SEEN, so the only reader who needed
that fact — the human watching an idle session — was the one not told. It now
says: this is a queued message, nothing woke the agent, any message starts the
turn that handles it, and the agent is not ignoring the ask, it is not running.
Costs nothing and changes no behaviour; it converts "why is it ignoring me?" into
"right, I nudge it". Deliberately NOT added to the wake-up injection, where a
turn is already starting and the note would be false.
A — a notification at poll time, because B only helps someone already looking and
the case that loses an ask is nobody looking. MEMPALACE_MAILBOX_NOTIFY: unset →
in-TUI ctx.ui.notify, the surface session_start already uses; =desktop →
additionally a terminal-native notification (Kitty OSC 99, else OSC 777), reusing
the detection the fleet's own notify.ts already proves in this harness; =0/off →
silent. The desktop path is how a ping escapes a container with no notify-send,
no DBus and no host access: the escape sequence is written to stdout and
interpreted by the terminal emulator on the human's own machine. Opt-in because
writing raw escapes is a behaviour change on a shared machine, not because it is
unreliable — say the word and the default flips.
Placement and firing conditions are deliberate: the notify call sits AFTER
sendMessage so a ping can never be the only thing that happened, and it fires
only when something is due — the same condition as delivery. A notification on
an empty poll would train its reader to ignore it, which is the failure this
whole feature exists to reverse. The wording names the nudge ("send any message
to handle") because "you have mail" without "press a key" reproduces exactly the
confusion B fixes.
Tested: 12 cases. Mode parsing (unset/empty/desktop/DESKTOP-with-space/0/off/1),
escape hygiene, and the OSC invariant that matters — a hostile title or body
containing ";" or a BEL cannot forge an OSC field or terminate the sequence
early, which is worth asserting because event bodies arrive from other machines.
ctx.hasUI is checked and the notify call is wrapped, since the UI can be gone by
the time an unawaited poll resolves. Syntax checked with node --strip-types.
|
||
|
|
982b001001 |
docs: RFC 003 §7.11/§7.12 — delivered is not read, and the mailbox carries obligations not news
Both measured 2026-08-26 by peer devices, and the second one measured on me: the operator had to point at an event id before I read a report that had been addressed to me for two hours. The mailbox was working correctly the whole time. §7.12 is the structural one. Candidates are drawn with status="open", so an event carrying a TERMINAL status is not a mailbox candidate at all, whoever it is addressed to. A task.reply written to a named device to share a finding is therefore delivered to nobody, ever — and neither is any event.ack. So the most natural inter-device message, "here is something you should know", is exactly the shape that gets no delivery. Two things that do work: address it as a directed status="open" ask (correct when a response is actually wanted), or accept it as pull-only and pair it with a drawer, which the peer's search will surface. A terminal report plus an expectation of attention does not. §7.11 is the peer's finding, and it explains the other half of why that report sat unread: delivery is deliverAs "steer" with deliberately no triggerTurn, and the poll fires on agent_settled — i.e. when the agent is IDLE, with no inference running. The text is queued for the next turn, so the human is the trigger. Measured on another device: a delivery landed at ~20:50Z and sat visibly unreacted-to until the operator asked "do I have to nudge you?". Not a defect; the no-triggerTurn decision was deliberate and stands. But the delivered text explains how to CLOSE an ask and never says when it will be SEEN, so the one reader who needs that fact — the human watching the window — is the one not told. Recorded with the three fixes that do not wake a model, as decision 8. §7.4 is upgraded from inferred to measured, on two devices independently, by a two-arm control: a to_agent='*' event with status='open' survives the raw candidate filter and therefore reaches the exclusion branch, which no previously written broadcast (all non-open) ever did. Raw query returned it, derived owed set did not, delivery named only the directed arm. Two arms differing only in to_agent is what makes it evidence rather than an absence. And it carries a correction of my own record, in place and dated: I had filed that this branch COULD NOT be exercised, because every broadcast in the log is status='ready' and the protocol forbids writing an open one. Wrong in an instructive way — the protocol forbids it as PRODUCTION TRAFFIC, which does not forbid planting one as a labelled, self-closed control. Where a rule appears to block a measurement, check whether it blocks the use or the test. |
||
|
|
bfe9c5cd4f |
mailbox: join the owed set on hlc, not seq, before a second replica exists
deriveOwed decides "is this ask still owed?" by asking whether one of my own terminal replies is LATER than the ask. It compared `seq` — this database's arrival rowid. On a single hub that is global order, so it was correct; the comment above it already said hlc was the durable key "once mesh_peers reports actual peers". Making the switch now, while one replica means the two orderings agree, costs nothing; making it later means changing the rule while two machines already disagree about order. The bug being pre-empted is specific: with a second replica the same event gets a different `seq` in each database, because arrival order is not authorship order. A reply authored after its ask can arrive first and take the lower seq; the join then concludes "no later reply exists" and an already-answered ask reappears as owed — permanently, on that machine. isStrictlyAfter() prefers `hlc` when both events carry one and falls back to `seq` otherwise (a server predating the field, or an un-backfilled row). hlc is rendered fixed-width, <unix_ms:13 digits>-<counter:6 hex>-<replica_id>, so a plain string comparison IS the causal comparison, with the replica id as final tiebreak. Verified before writing the code that the field is actually on the wire — event_list returns it per event (logstream.py:593) — because a fallback that never fires would have made this a no-op dressed as a fix. `created_at` stays rejected, and the reason is now written down where the decision is: it is server-generated at second precision, so ties are routine, and a tie can suppress an UNANSWERED ask. Both remaining failure modes are on the noisy-but-visible side — an answered item resurfacing is annoying, an unanswered ask going silent defeats the mailbox. Tested: 18 cases against the extracted comparator — hlc later/earlier/equal, same-ms counter ties in hex (0x10 vs 0x9, which is where a non-padded format would break), cross-replica tiebreak, both mesh reorderings, every seq-fallback path, and degenerate input (null seq, numeric hlc, empty strings, no keys) which must never claim "answered". All pass. Real-data check on the positive-control pair: seq 26/27 carry hlc ...3071857/...3574085, so the orderings agree today and the switch is a no-op now and correct later. The cursor keeps the opposite ordering ON PURPOSE — since_event_id is local-arrival ordered so a tail consumer still sees late-arriving remote ops whose hlc is older (hlc.py:19-21). RFC 003 §9.3 now says so explicitly, because that asymmetry looks like a bug worth "fixing" and is not. |
||
|
|
d2764bf78e |
docs: take the Phase 1 exposure record private, leave a moved-note
463 lines with 64 mentions of specific hosts, in a public repo: the primary and tunnel hosts by name, the registrar/DNS step, the tunnel resource wiring, shared token custody, per-machine flip dates, and a palace lineage naming three work machines. Now in the private fleet repo (fleet-ops 093fb65); this file becomes a moved-note in the shape docs/synlig-primary-runbook.md already established. Git history keeps the old text, so this limits future exposure rather than undoing it. Unlike the primary-host runbook, this file was MIXED — and the stub says so instead of quietly implying the toolkit still documents HTTP exposure. §1.1-§1.3 (why one shared fleet token rather than per-device proxy users, and the one place per-device identity does exist), §2 (the bind trap and the Host/Origin pin) and §3.7 (a client is flipped by three env vars that travel as a set) are reusable mechanism now published nowhere else. Named in the stub so the extraction is tracked debt rather than a silent loss, and named in the private copy too so whoever extracts it can delete the duplicate. Inbound references fixed rather than left pointing at content that moved: two §3.8 pointers in extensions/pi/README.md are replaced by the instruction they were pointing at (the palace path reported by mempalace_status must be the remote host's — a half-flipped client looks healthy while reporting a local path), and the bind-trap reference is replaced by the Host/Origin sentence itself, so the extension README no longer depends on the moved file. contrib also loses a hostname and a seeding date it did not need to make its point. |
||
|
|
d4d8bb6109 |
docs: retention direction for the coordination log, and drop the operator's name from RFC 002 §5
RFC 003 §9.1 — retention is settled in direction (rotate old traffic out of
the way, logrotate-style, moved aside rather than destroyed) and the sketch
records the parts that are not obvious:
- Tier artifacts before events. Events are a few KB; artifacts are capped at
4 MiB and stored in-row, so moving artifact CONTENT cold while keeping the
kind/sha256/size/created_by stub reclaims nearly all the space and keeps the
audit trail ("what was handed over, by whom, verified how") intact.
- If events rotate at all, the unit is the correlation_id THREAD whose latest
event is terminal — never the row. Archiving an ask while leaving its reply
(or the reverse) breaks the owed-set join, and both failure modes are bad:
an ask that can never be cleared resurfaces as owed forever, or a reply is
orphaned from what it answered.
- Never rotate an event that is still owed. Owed-ness is DERIVED at read time,
so an unanswered ask is indistinguishable from a stale one except by that
derivation — a purely time-based sweep would discard the live obligations of
a machine that has merely been offline for a month, which is precisely the
case this log exists to serve.
- Rotation invalidates held cursors: since_event_id RAISES on an unknown id
(§7.5), so archiving an event a watcher holds as its resume point turns its
next poll into an error. Either announce rotation ahead of live cursors, or
teach the anchor lookup to fall back to created_at/hlc.
- Archive → verify (row counts, artifact sha256) → only then DELETE + VACUUM,
with a --dry-run that reports in thread units.
RFC 002 §5: "Open decisions for ALC" → "Open decisions". RFC 001 never names
the operator anywhere; impersonal is the mature precedent and ALC was never a
real identifier in the first place (it is the AAAK spec's illustrative code for
"Alice", copy-forwarded into ~700 diary entries without verification).
Same edit also removes two device names and a hostname from §5.3, which is
host inventory and belongs in the private fleet repo, not a public one. The
mechanism it teaches — a palace in a Docker named volume dies on the next
container recreate, so census before flipping — is unchanged and is the part
that mattered.
|
||
|
|
e1cc7592a1 |
docs: write the RFC 003 that the code has been citing all along, plus an operator-facing fleet-memory guide
`logstream.py`'s module docstring is headed "Agent coordination event log for MemPalace (RFC 003)" and enumerates five "Design constraints (RFC 003)". Comments cite "RFC 003 phase 5", "RFC 003 suggested defaults" and "the first RFC 003 dogfood". Every event/artifact tool description cites RFC 003. The document has never existed — confirmed by searching this repo and the primary host. So this is a retrospective spec: it transcribes what the implementation already believes, and records what it does NOT do. docs/rfc-003-coordination-log.md, verified line-by-line against mempalace 3.8.0 (every claim cites file.py:LINE, indexed in §10 for re-verification). The parts that are not visible from the tool descriptions: - No idempotency guard on event_append or put_artifact (§7.1). The replication path checks `id OR (origin_replica, origin_seq)` before applying; the client path checks nothing. So peer replay is safe and CLIENT RETRY IS NOT — a retried append forks a coordination thread into two ids. This makes the fleet's "a timeout is not a failure, verify before retrying" rule load-bearing rather than advisory. - Coordination traffic is deliberately exempt from BOTH palace locks (§4) — _HTTP_LOCK_FREE_TOOLS and _PEER_WRITER_EXEMPT_TOOLS, each with its own rationale in-source. "One large mine blocks every client" is true of drawer writes and false of coordination writes. - `mempalace sync` never touches the log, and no DELETE FROM events exists anywhere (§2, §9.1) — answering for the logstream a question RFC 001 §7.2 left open, and making the log permanent and unbounded. - from_agent is shape-validated and never authenticated; there is no read scoping at all (§6). The log authenticates the fleet, not the agent. That is simultaneously the security limitation and the only way to positive-control the mailbox. - Three distinct orderings — seq (local arrival rowid), origin_seq (author's counter), hlc (fleet-wide, lexicographically sortable) — and origin_replica identifies the PALACE, not the writer, which is why from_agent/to_agent carry the whole distinction between machines (§3.2). - GET /logstream/events does not exist and never did (§7.8). An earlier measurement saw it 404 and blamed the reverse proxy; that inference was right for /sync/* and /logstream/stream and wrong for this one. Corrected in extensions/pi/README.md §3 too, in place, dated. docs/fleet-memory.md is the operator-facing companion the repo lacked entirely: the front-door README had zero mentions of coordination, so the channel was undiscoverable unless a message happened to arrive. It covers the five stores and what each is for (drawers/wings/rooms, diaries, KG, palace graph, coordination log), what a central palace buys a fleet — awareness, non-repetition of expensive work, and retractions that travel — and a decision flow for drawer vs KG vs event. Six mermaid diagrams, all rendered and inspected as images, not merely parsed: the first pass produced a truncated state label from an HTML entity and a self-loop that drew a meaningless dotted lasso. Validation says "no syntax error"; only looking says "correct". Also: a Documentation table in the top-level README so all of the above is reachable from the front door. Deliberately host-agnostic, per synlig-primary-runbook.md's precedent — no device names, hostnames or operator names in either new document. |
||
|
|
5b8d78f946 |
pi bridge: the log gets read, not just written
Adds the auto-delivered mailbox. Until now the bridge stamped events on the way out and never read the log, so a directed ask reached an agent only if that agent happened to run event_list itself — which in practice meant ALC telling it to. The channel had real cross-machine traffic since 2026-08-18 and no reader. DELIVERY, two points, both fail-silent and both additive (223 insertions, 0 deletions; the feed's agent_settled handler is byte-identical): - session start: one more sections.push() in the existing before_agent_start wake-up injection, beside mempalace_status and diary_read. - mid-session: a second agent_settled handler, poll floored at MEMPALACE_MAILBOX_POLL_MS (default 300000 = 5 min), delivered with pi.sendMessage(deliverAs: "steer"). The cadence is chosen from measured arrival, not taste: 22 events since 2026-08-18, of which ELEVEN landed inside one 5h38m window today. Arrival is bursty and correlates with the agent's own activity, because events arrive when another machine is working the same thread — so agent_settled (activity-coupled) is the right trigger and a wall-clock timer is the wrong one. Tightest observed gap was 2m12s, so a 5-minute floor coalesces a burst into one message instead of delivering five. POLLING IS DECOUPLED FROM DELIVERY, which is the part that keeps this from becoming noise: polling is cheap and frequent, but an item is only announced if it has not been surfaced this session, or was surfaced more than MEMPALACE_MAILBOX_RESURFACE_MS ago (default 1 h). Re-announcing the same ask every five minutes would train the reader to ignore it — the exact failure the status filter was introduced to prevent. The dedup map is in memory on purpose: after a restart it may re-show something already seen, and that is the SAFE failure direction (a resurfacing item is visible noise; a suppressed unanswered ask is silent and permanent). Owed-ness is DERIVED, never read off a field. event_ack appends and status is written once, so a directed `open` matches the mailbox query forever, answered or not — measured on this device, where the raw filter returned 3 asks of which 2 were already answered. A candidate is answered only when one of this device's own events has a strictly higher seq, joins via metadata.ack_of or a shared correlation_id, and carries a terminal status. The seq test is load-bearing: without it one terminal reply suppresses every later ask on that correlation forever, verified against the live thread where a seq-16 reply precedes the seq-17 request it cannot have answered. `*` broadcasts are excluded even though to_agent=<me> matches them, because the protocol says a broadcast owes nobody a reply. Leaving them in would have made this code contradict the skill documenting it, and would have made every machine think it personally owed the same answer. It also gives "don't broadcast an ask" teeth: broadcasting one now demonstrably reaches no owed set. GATE: on when MEMPALACE_PI_DEVICE and MEMPALACE_REMOTE_URL are both set (the same pair as the stamper — an unstamped client has no address to be reached at), off with MEMPALACE_MAILBOX=0. Default-on is deliberate and ALC's call: the problem being fixed is that nobody reads the inbox, and an opt-in fix for a nobody-does-it problem only relocates the forgetting. Inert on a solitary palace. README §2 rewritten in the same commit — it asserted "the bridge is write-only today: there is no mailbox, no poll, no delivery", which this commit falsifies. Shipping the code without the doc edit would have left a record asserting something untrue in the very file documenting the fix for that class of defect. VERIFIED: tsc 5.9.3 --strict, 0 errors, against the real pi types, with the harness mutation-tested first (an injected error on a new line was caught, then restored clean) because this repo has no package.json, no tsconfig and no tsc on PATH — nothing in-repo will re-run this. Owed-set logic extracted verbatim from the implementation and run against the live fixture: 3 candidates -> owed [seq 21] only; demoting the seq-19 reply to seq 5 makes seq 17 owed again (proves the ordering guard is live, not dead code); an open `*` broadcast and a self-authored open ask are both excluded; null correlation_id does NOT join itself (a plain === would have had null === null clear every uncorrelated ask). NOT verified: no live palace call from the implementation, and the queue-vs- interrupt semantics of "steer" are read from docs/extensions.md, not observed. |
||
|
|
e70bef2b5d |
docs: the edge stamper and the coordination log the fleet actually uses
Two gaps, both found by using the thing rather than reading it.
PROVENANCE WAS SHIPPED UNDOCUMENTED.
|
||
|
|
553d86570c |
provenance: stamp device+harness at the edge, not in the agent's head
RFC 001 §7.3.2 ranks "agent stamps provenance via a skill instruction" as the ❌ worst possible place — per-call boilerplate, forgettable, improvisable. It was right, and we had shipped exactly that: the mempalace skill told the agent to pass added_by="<harness>@<device>" by hand. Measured on the shared palace, 199 rows had reached it unresolvable, 10 of them filed by the very agent that wrote the instruction, in a drawer about host provenance. The trigger was a cross-host misattribution: a session on tor-ms22 read its own diary, could not tell that the entries were written on EMB-7KJ4VR4G, and reported another machine's verification as this one's. Move the same convention into the ⚠️ edge row, where it is uniform and unforgettable (§7.3.5): * extensions/pi/mempalace.ts defaults the writer field on every tool that has one — added_by (add_drawer, checkpoint), agent (mine), from_agent (event_append), created_by (artifact_put) — from $MEMPALACE_PI_DEVICE. An explicit value always wins, so filing for another device stays possible. The allowlist is per tool, never blanket: 3.8.0's dispatcher hard-rejects undeclared args with -32602, so injecting added_by into diary_write or kg_add (which have no such property) would break the call outright. * mine gets miner@<device> when the caller invokes it, but <harness>@<device> for the bridge's own transcript feed — bulk extraction is not agent-authored memory, and that keeps the pi/opencode/miner taxonomy honest. * diary_write has no metadata slot at all, and the device must never go in agent_name (wing = f"wing_{agent_name}" would splinter the diary per host). So the entry TEXT carries an AAAK field, HOST:<device>|SESSION:… — which is also the only channel a READER sees: search projects a fixed key set and diary_read returns content, so no metadata fix, not even a server-authoritative one, would have prevented the misattribution. * The wake-up block now states the device and warns that diary_read interleaves every machine's diary. * R1: doubly gated on MEMPALACE_PI_DEVICE and MEMPALACE_REMOTE_URL, so a solitary devbox stamps nothing and behaves exactly as before — which is also the correct semantics per §7.3.3. Version the reconciler that was living only on synlig (bin/ + contrib/systemd/), add --dry-run, and teach it two new rules: diary_host_marker reads the HOST: field, and sibling_chunk propagates a resolved origin across a drawer's chunks (a text marker lands in chunk 0 only, so a 5-chunk diary entry would otherwise stamp 1 and leave 4 blank). --dry-run against the real palace before deploying earned its keep twice, and scripts/test-device-stamp.sh pins both findings with the strings it found: HOST: was ALREADY in use with a composite grammar (HOST:emb-7kj4vr4g.f1d3c3f89e3e.v1.8.3.pi0.84.2) and for bare container ids, so an unvalidated rule invented devices like "f1d3c3f89e3e.pi0.84.2"; and HOST: also carries a different SENSE elsewhere (HOST:exec.via.ssh-controlmaster->…, meaning where I was executing). Validating against the known-device set both refuses those and recovers the composite entries correctly. A marker convention inherits every prior meaning of its own name. Deployed and verified on synlig: device 14,217 → 14,317, integrity ok, idempotent on immediate re-run, no invented device values. RFC updates: §7.3.1 corrected (the arg whitelist is a hard -32602 in 3.8.0, not a silent drop; get_drawer DOES return metadata, search structurally cannot; triples and logstream live in separate databases the stamper cannot reach), §7.3.5 added (what is deployed, including the divergence from §7.3.4's opaque origin_device — tor-ms22 vs tor-ms22-native is that cost already visible), and Phase 4 now carries per-device tokens motivated FIRST by revocation, with the finding that tokens are the cheap half: core holds one scalar auth_token and has zero device concept, so authoritative stamping needs a component we own. |
||
|
|
0fe64c480e |
docs: MEMPALACE_PI_DEVICE now also attributes drawers
Follow-up to
|
||
|
|
c64ffa1d93 |
feeder: default agent to pi@<device> so palace writes carry provenance
mempalace core 3.7.1 records neither which machine nor which harness produced a drawer, and the single shared bearer token means the server cannot distinguish clients. Today both facts survive only incidentally -- device in the per-device inbox path, harness in the pi_*.jsonl filename -- so attribution for anything filed outside the feeder has to be inferred after the fact. Defaulting --agent to pi@$MEMPALACE_PI_DEVICE records both explicitly in a field that already flows through to drawer metadata (added_by), for every enrolled device, with no core change. Falls back to $USER, which is what pre-existing drawers carry (added_by=joakim on fed transcripts, hence the ambiguity). Not pushed: lands on machines at the next image build. |
||
|
|
fd8b15f570 |
docs: the convos miner does check mtime — finish a correction that stopped half-way
ARCHITECTURE.md and README.md still carried the claim from |
||
|
|
947604b25d |
docs: backup and recovery, plus units; move host runbook to a private repo
Adds bin/mempalace-backup and docs/backup-and-recovery.md — the mechanism a palace actually needs, none of it site-specific. Why a palace cannot be backed up with cp: it is chroma.sqlite3 (authoritative), knowledge_graph.sqlite3 (usually WAL, so -wal/-shm make a plain copy a same-instant gamble), derived HNSW segment dirs, hallways.json, the embedder descriptor, and a HIDDEN .mempalace/origin.json. Both SQLite files are therefore copied through the online-backup API. Two bugs are documented because both produce a backup that looks fine: "$PALACE"/*/ silently skips the hidden dir, and per-directory rsync collides the identically named data_level0.bin in every HNSW segment. Treating the palace as one tree fixes both and makes a backup a faithful palace IMAGE, so restore is a copy rather than a procedure. Two modes: hot (default, zero downtime, ~4 s, index may lag but SQLite is authoritative and repair --mode from-sqlite rebuilds) and cold (--cold, ~5 s downtime, byte-consistent, restart trapped so a failed run still brings the server back). Verification runs on the COPY — quick_check plus row counts — and the backup is committed by mv only after it passes, with retention pruned only after a verified commit, so a broken new backup cannot delete the last good one. Documented because they are easy to get wrong: the sqlite3 CLI is often absent where the Python module is present; mempalace_embedder.json must be restored with the drawers or search silently degrades; a tested restore means running status AND search against the restored copy, since search is what actually exercises the index; mempalace-serve is a USER unit, so root systemctl reports "not found"; Persistent=true is what makes a missed window run after boot; and installing against the system Python couples the palace's availability to distribution upgrades, with the uv-managed-interpreter fix plus the two PATH traps that bite scripted upgrades. Also moves docs/synlig-primary-runbook.md out to a private fleet repository, leaving a stub that explains the split, since a host inventory is operator data for one deployment rather than part of a public toolkit. The path stays valid so existing links do not break. Remaining host references in the README, RFCs and ARCHITECTURE are left alone deliberately: they are load-bearing prose, contain no secrets, and are best generalised as they are next edited rather than in one churn-heavy pass. |
||
|
|
b609cf5a69 |
docs(synlig): the transcript inbox is the primary's third moving part; and survive an unset HOME
Runbook gaps found while fixing the 2026-08-15 feed failure: - §2.5 still titled "written but not installed" and still asserting "Not installed, not enabled" — false since 2026-08-12. A reader landing there got a flat contradiction of §4 item 3. Retitled, with the verified-2026-08-16 process line, and it now states the fact §2.6 depends on: the server is a NATIVE process (no mempalace container on synlig), so it sees host paths. - New §2.6 documents ~/mempalace-feed/<device>/: why transcripts cannot travel over the HTTPS leg at all, the three client variables, and why MEMPALACE_PI_REMOTE_PATH is the trap (its /data/feed default assumes a containerized server; here it must equal the ssh-target path, and a mismatch fails with rsync succeeding and only the mine failing). Plus the operational notes that cost time: dedup keys on the absolute path so the inbox path is load-bearing, grown sessions are purged+refiled by mtime, and a client-side MCP timeout is NOT a failed mine. - Header status: counts refreshed with an explicit "treat counts as timestamps". Also, bin/mempalace-pi-session: default HOME from the passwd database when it is unset. `docker run --entrypoint="" <image>` inherits no HOME when the image config declares none, and every default is HOME-anchored under `set -u`, so the script — including the palace-free --self-test — died with "HOME: unbound variable" in exactly the environment pi-devbox's smoke suite uses. pi-devbox v1.8.0 lost a release to the same assumption from the other side. |
||
|
|
6e1f4f30fc |
fix(pi-session): a failed remote mine reported success — MCP escapes the payload
The remote-mine leg decided success with `'"error"' in body`. MCP answers a
hard tool failure with HTTP 200 and a JSON-RPC *result* whose content[].text
carries the tool's own JSON as an ESCAPED string, so those bytes are
\"error\" and the substring never matches. On 2026-08-15 (EMB-7KJ4VR4G, first
boot of the fresh pi-devbox image) a mine that failed with
{"success": false, "error": "source directory not found: '/data/feed/emb-7kj4vr4g'"}
printed "Done. Wing 'wing_conversations' updated." directly under that error
and exited 0. Transcripts had been rsynced for the whole session and filed
nowhere; the only artifact anyone would check said it worked.
- classify(): parse the envelope instead of grepping it. Catches JSON-RPC
errors, MCP isError, and inner success=false/error, and distinguishes
"verified ok" from "unverified: no JSON tool payload" rather than assuming.
- --self-test: six recorded MCP responses (fixture 1 is the real 2026-08-15
body) plus a regression guard asserting the old substring check is blind to
it. Needs no palace, no network, no sessions dir.
- Preflight warning when the rsync destination path and
MEMPALACE_PI_REMOTE_PATH disagree. The /data/feed default assumes a
CONTAINERIZED palace server; a native one (systemd unit / uv tool) sees host
paths, and then the two must match. Warned in preflight so --dry-run and
--prepare surface it too.
- Remote mode no longer previews NEW/SKIP from the LOCAL palace: dedup happens
on the palace host keyed on the remote inbox path, so this machine cannot
answer it. Tags become [?] and the summary says who decides. It had been
reporting "6 already filed" about a palace it was not feeding.
|
||
|
|
f60cf9c732 |
feat(census): RFC 002 Phase A — read-only join census, and three RFC corrections it found
bin/mempalace-census classifies a palace on disk into the RFC 002 §2 classes
(mined / diary / agent-authored) and emits a human report or a --json manifest
that feeds Phases B/C. Read-only: every sqlite handle is opened mode=ro, no
-wal/-shm is created, safe against a live mempalace-serve. Reads LOCAL DISK only
and warns if MEMPALACE_REMOTE_URL is set, so a local census can't be mistaken
for a remote one.
The design point is that it SELF-VERIFIES instead of trusting ids.py's
docstrings: for every replayable drawer it reassembles content from chunks,
recomputes the upstream id and compares to the stored id. That single check
covers the id recipe, the chunk reassembly order and the classifier at once --
176/176 accounted for on the reference palace -- and it falsified three things
the RFC previously asserted from a docs-only reading (now RFC 002 §2.1):
(a) The hash input is LENGTH-PREFIXED, not "|"-joined. ids.py:31 defines
_DELIM = "|" and the make_* docstrings describe f"{wing}|{room}|{content}",
but _DELIM is dead code and _delimited_sha256 builds
"".join(f"{len(part)}:{part}"). Measured: length-prefixed reproduces real
ids 5/5, pipe-joined 0/5. Diary ids differ again -- a PLAIN sha256.
(b) id_recipe is NOT a mined-only marker. It looked like a clean
discriminator (same 14,586 count as source_file) but the server stamps
'v3' on content ids too, so classifying on `source_file OR id_recipe`
swallowed all 60 agent-authored drawers into MINED -- the dangerous
direction, since Phase C would try to re-mine drawers that have no source
file and silently drop them. Discriminator is a TRUTHY source_file (the
writer stores "" rather than omitting the key), cross-checked against the
miner-only keys source_mtime / normalize_version; disagreement is now a
first-class warning.
(c) Content ids DRIFT: 9 of 60 agent-authored drawers no longer reproduce
their own id, because update_drawer preserves the id while rewriting and
re-chunking. So "recompute the content id and skip if present" -- the
strategy this RFC specified for regime A -- misses every drifted drawer
and duplicates it. Phase C must key on the STORED id. Flagged as
edited_since_filing in the manifest so Phase C can be tested on them.
Also corrects §4.1's headline number: 15,949 mined / 98.9% was reconstructible
exactly as 16,338 (all embeddings rows) - 192 (diary rows) - 197 (agent rows),
i.e. it counted rows rather than parent drawers AND spanned both collections,
absorbing all 1,560 non-joinable mempalace_closets rows into the mined total.
Correct figures: 14,389 mined / 116 diary / 60 agent-authored = 14,565 parents,
replay surface 176 (1.2%). Two rules now enforced in the tool: always filter by
collection (one sqlite file holds both), and always say whether a count is rows
or parent drawers -- a chunked drawer contributes N rows and no parent row.
One implementation trap worth recording: Chroma splits metadata across
string_value and int_value, so reading only string_value nulls every numeric key
(source_mtime, chunk_index, line_start) -- which made the miner-marker
cross-check report 100% conflict until the loader coalesced the two columns.
|
||
|
|
2f9170428c |
docs(rfc-002): run the census for real — the replay surface is 176 records, and closets were missing
Ran Phase A's classifier against the frozen EMB-7KJ4VR4G archive. Since the fleet shares one devbox image this is a reasonable prior for tor-ms22 and MBP-M1-2020: mined (source_file set) 15949 98.9% re-mine, never replay diary entries 116 0.7% replay + §7.6 suffix skip agent-authored drawers 60 0.4% replay, idempotent KG open facts 34 -- server guard dedupes KG closed facts 0 -- nothing to do Two consequences that shrink this project: the replay-only surface is 176 records, not thousands, so the writer is a small job and the §7.6 diary guard treated as the blocker governs 116 records; and there are ZERO closed KG facts, so the unguarded-closed-fact gap is real in the code but empty in the data. filed_at spread in the same archive -- 12 in May, 52 in June, 15174 in July, 887 in August -- is the concrete argument for the history-preserving regime. Two corrections to the first draft: 1. CLOSETS were omitted entirely. mempalace_closets is a second Chroma collection (1560 rows, ~10% of the palace) and there is NO MCP tool that writes one -- mcp_server.py exposes only _purge_source_closets, and closet_llm.py says outright that regex closets are always created by the miner. So closets cannot be replayed even in principle; they return only by re-mining. Same category as hallways/known_entities/palace-graph, and it resolves itself for the ~99% that gets re-mined anyway. 2. Census gotcha: chroma.sqlite3 holds BOTH collections, so an embeddings-wide query over-counts by ~10%. Must join segments->collections and keep mempalace_drawers. Doing so reconciles exactly: 14,778 = the 14,777 seeded to the primary + the one known post-snapshot chunk. Also drops the earlier "438 of 14,829" figure, which conflated central's no_source count (which includes its own later agent writes) with the archive's actual replay surface. |
||
|
|
a25e22922b |
docs: scope the joiner (RFC 002) — and the chronology loss nobody had costed
RFC 001 §4.4 designs a join as "idempotent replay of local history", but no
replay tool exists and the two joins done so far were whole-palace file copies
that cannot merge. tor-ms22 and MBP-M1-2020 each hold a palace that needs to
reach the primary, so this scopes the tool.
The finding that drives the design: MCP replay CANNOT preserve filed_at.
add_drawer stamps it server-side (mcp_server.py:2580) with no override, and
diary_write builds its own now()-based id. A pure MCP replay would therefore
collapse months of history into the join instant -- on a palace whose value IS
its chronology, and where list_drawers filters on filed_at. RFC 001 does not
mention this. kg_add is the exception: it takes valid_from/valid_to, so fact
windows survive.
Hence two regimes, and a recommendation to build the history-preserving one
first: direct disk write on synlig (preserves ids + filed_at, bypasses the
server guards, needs the service stopped) vs MCP replay (guards work, timestamps
flatten). migrate.py is already a working model for the direct path -- it reads
drawers straight from the palace sqlite and re-adds them preserving ids,
documents and metadata.
Also concretised: every dedup key verified against mempalace 3.6.0 source rather
than assumed --
- agent-authored drawers: sha256(wing|room|content)[:24], fully deterministic
- mined drawers: sha256(source_file|chunk_index)[:24], PATH-dependent, so
replay duplicates instead of deduping -> re-mine on synlig, do not replay
- diaries: full ids never repeat (wall-clock component); match the
sha256(entry)[:12] suffix only -- this is §7.6
- KG closed facts: no server guard at all (it is scoped to valid_to IS NULL)
- hallways/known_entities/palace-graph are built at MINE time, so replayed
drawers arrive with no co-occurrence edges and traverse under-reports
Scope is phased so the useful half lands first: Phase A is a read-only census
that sizes the job and cannot break anything, and is the piece to run the moment
tor-ms22 is reachable. The replay-only surface is small -- only 438 of 14,829
drawers on the primary have no source_file.
Method note recorded in the doc: an attempt to gather these mechanisms via a
delegated subagent returned confident, fabricated code for a package path that
does not exist on this machine. Everything here is cited to file:line and was
re-read directly.
|
||
|
|
c349d007e1 |
docs(§7.2): measure the shared-palace sync blast radius — it is 0 today, and the feeder is what changes that
Unscoped `mempalace_sync` dry-run against the live primary (14,829 drawers): out_of_scope 14391, no_source 438, missing 0, kept 0. Zero drawers deletable. The mechanism is the finding: an entirely-absent source root yields `out_of_scope`, NOT `missing`, and only missing/gitignored drawers get removed. So a wipe needs the root to EXIST while files under it do not -- not merely a host that lacks the repos. §7.2's "from a laptop that lacks the repos, it is a fleet-wide wipe" therefore overstates today's risk. It also understates tomorrow's, which is the part worth acting on. Two guards currently prevent the laptop scenario and neither was designed to: the CLI has no remote support, so a client's `mempalace sync` cannot reach central at all; and the MCP tool runs server-side on synlig, where clients' /workspace roots do not exist. That is safety by coincidence of layout. The feeder's remote mode dissolves it: it rsyncs staged transcripts into per-device inboxes ON synlig, so those sources begin existing on the palace host and become in-scope for the first time. A later stage rotation or inbox cleanup then marks that device's conversation drawers `missing` -- prunable, and prunable from a different device. mempalace-pi-session already documents the single-machine form of this; a shared palace makes it cross-device. So stage retention on synlig plus the sync guard belong to enabling the feeder fleet-wide, not to a later cleanup pass. |
||
|
|
6e8172d93a |
docs: record the primary's real lineage — EMB-X1JY06WJ -> EMB-7KJ4VR4G -> synlig
Three docs said the primary was "seeded from EMB-7KJ4VR4G's palace", which is
true but stops one hop short. That palace was itself carried over from
EMB-X1JY06WJ, the previous work computer, when it was replaced around
2026-07-06.
Evidence, read read-only out of the frozen archive on EMB-7KJ4VR4G: 76 drawers
carry `source_machine=EMB-X1JY06WJ` and they are the oldest in the store --
earliest filed_at 2026-05-04, two months before that machine's stack existed --
with no other source_machine value present. The remaining ~16k were filed
locally afterwards (15,189 in July, 1,073 in August). Last write is
2026-08-14T15:07:06, the seed instant.
Two consequences, recorded in rfc-001 S4.4:
- It was the *second* whole-palace file copy, not the first, and both were
single-source clones. So the method has two successes behind it and still
zero exercises of merge semantics -- S7.6 is *less* tested than "we've done
this twice" would suggest, not more.
- The primary now holds records from a machine that no longer exists, and
`source_machine` is the only thing marking them. Not noise; don't prune it.
Also corrects
|
||
|
|
a94eb7fdd0 |
docs: native pi has no .env to flip, and EMB-7KJ4VR4G has no native pi
Follow-up to
|
||
|
|
ec436ed3ad |
docs: reconcile the RFC-001 docs with what is actually deployed
Audit of every doc touching the global-palace rollout against the running
fleet. Each correction below was verified against the filesystem or the host,
not against another doc:
- synlig-primary-runbook: the decommission `rm -rf ~/.mempalace` now carries a
STOP block. That tree holds the fleet palace *and* the only copy of the
bearer token every client authenticates with; the old "empty today" comment
stopped being true when the palace was seeded on 2026-08-14. Adds an ordered
safe decommission, and drops count-based join verification.
- phase-1-exposure-runbook: new S3.8, how to verify a flip actually took --
the procedure that until now existed only in an untracked handover file.
Three claims that fail independently (env var / curl / the palace-path
discriminator) plus an explicit list of checks that produce FALSE POSITIVES:
drawer counts (both sides were seeded from the same palace, and `status`
counts chunks not drawers), write-then-read through the same transport, and
the `mempalace` CLI -- which has no remote support at all, so post-flip it
reads the dead local archive and reports success.
- rfc-001: status Draft -> Phases 0-1 implemented. Records that the join was a
file-level copy, which SIDESTEPPED the S7.6 diary-dedup question rather than
answering it -- so S7.6 remains a hard blocker for the second machine, which
is the one that will actually exercise merge semantics.
- ARCHITECTURE, SKILL, contrib/README, extensions/pi/README all claimed pi
feeds the palace automatically, unconditionally. That is gated on
mempalace-toolkit >=
|
||
|
|
2293f1c89b |
docs: the silent-skip fix is base-image-gated, so today's fleet still skips silently
pi-devbox cbd7cf5 makes the remote-palace-without-inbox skip announce itself, but entrypoint-user.sh is COPY'd in Dockerfile.base, so the fix only reaches a container after a base rebuild. Every image running today still skips silently -- recording both behaviours with the cut-off, so the table stays true for whichever image a container is actually on, rather than describing a fix that has not shipped yet. |
||
|
|
2e73a9eae8 |
docs: REMOTE_URL includes /mcp, and the entrypoint skips silently where the feeder exits 1
Two details needed before the first client flip. 1. MEMPALACE_REMOTE_URL is the full endpoint including /mcp with no trailing slash. The feeder POSTs to it verbatim (bin/mempalace-pi-session:718), and the server matches `path != "/mcp"` exactly, so a base URL or a trailing slash both 404 -- and a 404 here looks like a routing/proxy fault, not a config typo, which is a bad hour to spend. Matches the existing pi-devbox examples. For this fleet: https://mempalace.jordbo.se/mcp 2. The 3.7 trap has TWO symptoms, not one, and I had only documented the loud one. Direct runs / session-end / cron exit 1 with a clear error (:298-300). But pi-devbox's entrypoint-user.sh:134 checks the same condition and skips *quietly* -- and the skip happens before the subshell that writes ~/.pi/agent/mempalace-catchup.log, so there is not even an empty log to find. A container flipped with REMOTE_URL but no SSH_TARGET therefore contributes nothing to the palace and leaves no artifact explaining why. The skip is correct in itself (there is genuinely no inbox to ship to) but it is indistinguishable from a healthy run with nothing to do. Documented with two commands that tell those apart after a flip. |
||
|
|
4cb70ce3e1 |
docs: record the 302 auth-redirect fingerprint and the server's exact HTTP surface
Phase 1 exposure now verified end to end (2026-08-12): /healthz -> ok and an unauthenticated POST /mcp -> 401, both from a client container and from the primary itself. 3.4 and 3.6 marked passed. Two findings from the failure in between, worth more than a checkbox: 1. Leaving Pangolin's resource authentication on does NOT surface as 401 or 403. It is a 302 with `location: https://pangolin.jordbo.se/auth/resource/<uuid> ?redirect=...`, sent *before* the request is proxied -- so the palace never sees it and its journal stays silent, `curl -s` prints an empty body, and an MCP client gets non-JSON. That is 1.1's "Pangolin's HTTP auth is browser-shaped; the clients are not" arriving as a concrete symptom rather than an argument. Now recorded with the exact header shape and the two curl flags that reveal it (-D-, -w '%{redirect_url}'), plus the reading: a 302 is good news, because DNS, TLS and routing all worked and only auth intervened. 2. Read the server's routing to settle whether path-scoped proxy rules are sufficient. The entire HTTP surface is two exact paths (mcp_server.py:5299- 5318): GET /healthz (no token, Host/Origin gated) and POST /mcp (Bearer, compare_digest on the exact string). Everything else is a 404 from the palace itself. So path-scoped rules are tighter than a host-wide proxy and lose nothing. Two client-facing consequences: /mcp is matched exactly, so a trailing slash 404s -- configure clients without one; and there is no GET /mcp, no SSE, no session id, no DELETE, so this is plain JSON-RPC over POST, not MCP streamable-HTTP. A strict client opening with a GET handshake sees 404. Also means the verify step needs no initialize and no Accept: text/event-stream, which the old snippet left ambiguous. Also: run the 200 and 401 from two different networks, not one -- passing from only the primary leaves split-horizon DNS untested. |
||
|
|
08e344b047 |
docs: proxy target is http:// not https://; loopback probe refuses, it does not 403
Both corrections come from the first real Phase 1 start on synlig (2026-08-12).
1. The Pangolin resource target was documented as bare `172.17.0.1:8765` with no
scheme, and the obvious guess from that is `https://` -- which cannot work.
`contrib/systemd/mempalace-serve.service` runs `serve --host 172.17.0.1
--port 8765` with no --tls-cert, so the primary speaks plaintext HTTP; TLS
terminates at Pangolin, which is the whole point of the RFC 6.2 decision.
Point a proxy at https:// and it attempts a TLS handshake against a plaintext
listener: 502 from outside, while `curl 172.17.0.1:8765/healthz` on the box
still says ok -- a confusing pair of symptoms. Now spelled `http://` with the
failure mode named, in the runbook and in the unit's comments.
2. `curl -s 127.0.0.1:8765/healthz` was documented as "expect 403". Wrong: the
real run returns empty. With a docker0-only bind nothing is listening on
loopback, so the connection is refused at TCP level before any header is sent
(%{http_code} -> 000, exit 7). The 403 is the *loopback-bind* case verified
2026-08-10 -- server on 127.0.0.1 answering a proxy-forwarded foreign Host.
Two distinct behaviours had been collapsed into one expectation in three
places (both runbooks and the unit).
Worth stating why the correction matters rather than just fixing it: refusal
is the *stronger* signal. A 403 proves only that a request was rejected; a
refused connection proves the loopback and LAN surface is not listening at
all. Someone who expected 403, saw silence, and "fixed" it by rebinding to
0.0.0.0 would have converted a correct configuration into an exposed one.
The docs now also say what to do if it hangs, or if ss shows 0.0.0.0:8765.
|
||
|
|
00a95d1a2f |
docs: why the tunnel and the feeder's SSH path are not redundant; newt done
Asked "why do we need Pangolin if you proposed rsync/ssh?", and the runbook did
not actually answer it -- it stated both were needed without saying why neither
substitutes. New section 1.3:
- Pangolin/HTTPS carries the MCP tool surface (search, add_drawer,
diary_write, kg_*) -- every live tool call, from any MCP client.
- SSH/rsync carries transcript *files* only, because mempalace_mine expands
its source path server-side, so the server can only mine its own disk.
HTTPS alone is a palace you can query but cannot feed; SSH alone is files with
no query API. The rsync is not a transport preference, it is a workaround for
where `mine` resolves paths.
Records honestly that `ssh -L 8765:172.17.0.1:8765 synlig` *would* replace the
tunnel for MCP, and why we don't: synlig dials out (reaching for a dial-out
tunnel is itself the evidence inbound was unavailable), MCP clients want a
durable URL rather than a per-session forward, and the forward must be up on
every device before every session.
And the design's weak point, stated instead of glossed: the rsync runs
client -> synlig, so mining needs synlig's SSH reachable *from the client*. Were
that true everywhere, no tunnel would be needed for MCP either. Honest
expectation after Phase 1 is therefore: query/write from anywhere, mine only
from devices that can reach synlig's SSH. Section 4 now carries the upstream ask
that would close the gap -- have the feeder send content over MCP (add_drawer /
diary_write, which it already calls) instead of asking the server to mine a path
it must first rsync there.
New section 3.7, a live trap for the imminent client flip:
MEMPALACE_REMOTE_URL on its own does not degrade to local feeding, it *stops*
feeding. auto mode switches to remote as soon as the URL is set (:286) and
remote mode then exits 1 without MEMPALACE_PI_SSH_TARGET (:298-300), before
anything is staged or filed -- so a cron feeder just starts failing, and the
loudest symptom is silence. Two safe orders given: set all three variables in
one edit, or set URL+token and pin --mode local until the SSH target exists.
newt is installed on synlig and connected to Pangolin (done 2026-08-12), marked
here and in the synlig runbook's item 2; the blocker is now item 3, the one
sudo. Added the follow-up that "connected to Pangolin" only proves newt reached
nyvaken -- reaching the *palace* is a separate claim that fails independently,
so probe 172.17.0.1:8765/healthz from inside newt's namespace.
All seven code citations verified against the source at commit time.
|
||
|
|
29e660e18f |
feeders: stage beside the palace, not in ~/.cache; document Phase 1 exposure
Staging default moves out of ~/.cache to <palace-root>/pi-stage (pi) and <palace-root>/opencode-stage (opencode), resolved with mempalace's own palace-path precedence ($MEMPALACE_PALACE_PATH -> $MEMPAL_PALACE_PATH -> ~/.mempalace/config.json -> ~/.mempalace/palace), then dirname. Why: the convos miner keys dedup on the *staged* path, so a wiped stage plus a sync scoped to include it prunes the drawers mined from those sources -- deleting memories, not a cache. Under ~/.cache that state was reachable by anything treating a cache as disposable. Staging inside the palace makes the coupling structural: the stage cannot be wiped without touching the palace itself. Overrides ($MEMPALACE_PI_STAGE / $MEMPALACE_SESSION_STAGE, --stage) are unchanged. Note the old default had never been created on any host, so this closed a latent hazard, not a live one. Measured, and the docs now claim only this much: sync prunes only within the scope it is given -- wing-only, 1299 scanned / 1299 out of scope / 0 removed; scoped at the palace root, 651 kept / 648 out of scope. The previous blanket "sync prunes every drawer" wording overstated it, which is a liability: the next reader disproves the overstatement and discards the real constraint with it. Also in this change: - cron log dir ~/.cache/mempalace-session -> ~/.cache/mempalace-logs. The stage left that namespace, so the old name now read as "the stage". - AGENTS.md: the convos miner *does* check mtime (verified against upstream convo_miner.py); the previous "no mtime check" claim was wrong. - smoke-test assertions use `mktemp -d` for --sessions-dir. One pointed at /tmp, which still held earlier synthetic transcripts, so a --dry-run exported a fake session into the real stage: --dry-run skips the mine, not the export. docs/phase-1-exposure-runbook.md -- the newt/DNS/auth step that RFC 001 and the synlig runbook leave open (runbook section 4, items 2 and 5). Port 8765 at /mcp, newt targets 172.17.0.1, and the authentication is the single shared bearer token (RFC 6.2, decided 2026-08-09) rather than per-device proxy users. The latter cannot work today: mempalace validates exactly one token, and Pangolin's SSO/PIN/password are browser-shaped while every client here is a headless JSON-RPC POST -- enabling that protection breaks the clients it protects. The per-device axis that *does* exist is the feeder's SSH key + per-device inbox. New finding recorded there: a loopback bind does not merely 403 behind a tunnel (already known, runbook 2.4) -- it also silently starts the server with no token at all, because auto-minting is gated on the bind being non-loopback. extensions/pi/README.md: the HTTP transport IS authenticated as of mempalace 3.6.0; the "sessionless and unauthenticated" note dated from the v1.3.0 era. Closes the RFC section 8 Phase-0 hygiene item. |
||
|
|
3626946013 |
Phase 0 on synlig: install, palace layout, verified Host/Origin policy
synlig is greenfield — no mempalace, no ~/.mempalace at all — so Phase 0 became "provision correctly from birth" rather than "migrate carefully". Nothing is serving; no client config was touched. Done: - mempalace 3.6.0 installed via uv, pinned to the fleet version (the id recipes and idempotency probes this RFC leans on are version-specific). - Embedder pre-warmed. This was the real unknown: the first embed pulls a 79.3 MB ONNX model from the chroma CDN, and an egress-filtered work VM would have failed at the worst moment — the first client write. Pulled at ~20 MB/s, no proxy interference. Done in a throwaway palace so the real one never saw it. - Palace at the stock default ~/.mempalace/palace, so no config file and no MEMPALACE_PALACE_PATH is needed on synlig at all. Two corrections to the RFC, both from provisioning rather than reading: - §7.1 was understated. DEFAULT_PALACE_PATH (~/.mempalace/palace) and DEFAULT_KG_PATH (~/.mempalace/knowledge_graph.sqlite3) already differ with stock defaults, so the KG split is out-of-the-box behaviour, not a consequence of a custom --palace, and it is permanent rather than one-time: serve always passes --palace, any CLI call without it uses the HOME path. A one-time mv does not fix that, it only picks which of the two files gets populated. Fixed instead by converging both rules on one inode via relative symlinks, and verified the load-bearing assumption: a dangling symlink is created on connect, -wal/-shm land next to the target (so the palace dir stays a self-contained backup unit, which is the part that mattered), cross-path read works, same inode. hallways.json deliberately left alone — already palace-derived, HOME path is a warning-only probe. - §6.2 upgraded from "test early" to verified: 11/11 as predicted. The headline is that the safe-sounding reflex is the failure mode — loopback bind + proxy forwarding a public Host is 403, non-loopback is 200. Also confirmed Origin is never relaxed on either bind, and /healthz is Host/Origin-gated but token-free, so it works as the tunnel probe. Recommends binding the docker0 gateway over 0.0.0.0: non-loopback so the pin relaxes, but reachable only from the host and its containers. New docs/synlig-primary-runbook.md carries the discovered facts about the box, the evidence tables, an explicit "deliberately not done" list, and tomorrow's ordered steps. New contrib/systemd/mempalace-serve.service carries the bind rationale inline so nobody "fixes" it back to loopback; staged on synlig with a .staged suffix so systemd cannot pick it up by accident. Flagged for tomorrow: synlig has no newt/tunnel client (docker ps shows only the Gitea runner and digikam), so Pangolin on nyvaken cannot reach it until one is added — easy to miss, because Pangolin will look healthy from its own side. |
||
|
|
7e51055c96 |
rfc-001: join protocol, diary non-idempotency, deployment decisions
Second planning round. Three findings from verification, and the decisions that were blocking phasing. Verified in mempalace 3.6.0 and written up: - NEW §7.6 — diary_write has no idempotency guard at all. The id embeds datetime.now() at microsecond resolution and the write is a bare col.add (mcp_server.py:3546) with no col.get probe, in direct contrast to add_drawer's probe 900 lines earlier (:2593-2600). So §3's "log the intent, replay the intent" is false for diaries — replay duplicates. This lands on the critical path because diaries are `replicated` and are exactly what a join replays. - NEW §4.4 — join/bootstrap protocol, which the RFC simply lacked. Not "seed the primary from one palace": every container joins the same way, repeatedly, so a join is idempotent replay and the only question per record type is what dedupes it. Drawers and open KG facts need zero client bookkeeping; closed KG facts and diaries need client-side keys. Join state belongs in the palace dir, not the container (~/.mempalace is not preserved by default), which also makes two containers sharing one host's palace the easy case rather than a double-upload hazard. - §3.2 — the add_triple guard is scoped to `valid_to IS NULL`, so closed historical facts are unguarded. §3 read as unconditionally idempotent. - §7.3 — drawer/diary metadata is exhaustive: no session, PID or conversation field. Two concurrent pi sessions are indistinguishable, and pi-vs-opencode is only accidentally distinguishable because the pi extension never sets added_by (it sets identity for diaries only, mempalace.ts:758). Corrects the §7.3.3 bullet that claimed multi-harness attribution was already solved — the field is the right home, but nothing populates it. Design rule stated: device + agent, never session; provenance granularity equals token granularity. - §6.2 — Host-pinning is coupled to the bind address (enforce_host_pin = _http_is_loopback(host), :5367). The trap is the RFC's own "bind loopback" reflex: behind a proxy that forwards a public Host, a loopback bind 403s, while a non-loopback bind deliberately relaxes the pin. The Origin check is never relaxed. This replaces the warning I first wrote, which had it backwards. - §9.3 — largely resolved: the last-chunk-only probe is deliberate (batched upsert is all-or-nothing, :2586-2592). Only a crash mid-upsert remains untested. - §4.1 — fix shape for opencode's stale-config asymmetry: make the mcp.mempalace subtree env-authoritative, with a fingerprint so hand-edits still win. pi-devbox/entrypoint-user.sh:131-162 already has the jq deep-merge pattern to port; no OPENCODE_CONFIG* layer exists upstream, so the real file must be written. Now Phase 1.5. Decisions (new §8.1): primary = synlig; TLS at Pangolin; single shared token for now (so origin_device stays advisory); diaries replicated. Work/personal is per-wing, not per-device — two primaries split along machine lines is rejected, because pi-devbox/opencode-devbox are simultaneously work and home projects and the device where work happened cannot classify the project. MEMPALACE_REMOTE_URL stays scalar so multi-store remains additive later (new §9.7 keeps the placement question open). |
||
|
|
fdcd5871de |
rfc-001: resolve open questions 1 and 6 — both permissive
Recon before planning the implementation, and both blockers dissolved. Q1, does opencode support remote MCP: yes. Its published schema defines McpRemoteConfig (type/url/headers/oauth) as a sibling of McpLocalConfig, with headers as a free string→string map. The RFC had been treating "every sampled config in myconfigs is type:local" as evidence about the schema when it was only evidence about deployments. Consequence: the edge proxy is not the only option for opencode, so nothing in the phasing hangs on this. What opencode still lacks is offline/local-first, which is the honest Phase 2 argument. Q6, how opencode-devbox learns the opt-in: it already does. generate-config.py, run from entrypoint-user.sh:117, registers mempalace as remote+bearer when MEMPALACE_REMOTE_URL is set and local stdio otherwise, and its comment says it deliberately mirrors the mempalace.ts env contract. The image also ships mempalace by default, so R5's premise was wrong and is corrected. Phase 2 there is a third branch in an existing script, not a new mechanism. One new constraint found while confirming this, and it is the sharpest edge in the whole opt-in story: generate-config.py never overwrites an existing config and ~/.config/opencode is a named volume, so flipping the .env on a container that already has a config is a no-op — it only writes an opencode.jsonc.proposed sidecar. pi re-reads env every start; opencode does not. Documented in §4.1 and §9.6, and it applies symmetrically to R6 reversibility. |
||
|
|
052dbb8038 |
docs(rfc-001): provenance belongs to the sync boundary, not the agent
Reverses the previous §7.3 on review. It said "stamp added_by everywhere, now, because it cannot be backfilled" and was about to become a skill instruction telling agents to do it. Both halves were wrong. Wrong on ownership: provenance answers "which device asserted this?", so only a party that can verify the answer should write it. An agent must shell out to read env, can forget, and will improvise when the values are absent — the worst possible stamper, and its claim is unverifiable by anyone. Under the §4 design every write reaches the primary over an authenticated channel, including offline ones at outbox-flush time, so the primary can stamp the complete set with no client cooperation. Provenance is a property of the sync channel, not of the record's author. §7.3.2 adds the trust ladder; edge-side stamping is demoted to an advisory interim, because 3.6.0's serve takes a single shared bearer token (mcp_server.py:5291-5293) and the package has zero device/origin concept, so authoritative stamping needs the per-device credentials of §6 — Phase 4, not Phase 0. Wrong on backfill: a solitary devbox is a single-origin store by definition, so origin is a property of the whole palace and can be assigned wholesale at import (one --origin-device flag) at the moment it stops being solitary. Bulk attribution is strictly more reliable than per-record stamping since it cannot be partially applied. Per-record provenance is only needed where origins interleave, which is only the primary. Solitary containers therefore stamp nothing and lose nothing — more R1-compliant than the previous draft, which quietly asked users who had opted out to carry metadata for the feature. Multi-harness-on-one-host stays solved by added_by = agent name. Keeps the verified mechanics (fixed metadata schema, argument whitelisting at mcp_server.py:4777 silently dropping unknown fields, the diary agent_name → wing_pi@host trap, kg_add having no provenance slot, added_by absent from search results) and the identity findings (a container cannot discover its host's identity; hostnames are neither unique nor stable; rename splits one device's history in two). §7.3.5 keeps the fail-closed rule for whichever component does stamp. Phase 0 drops its provenance item accordingly, and §6.2 now specifies server-side stamping from the authenticated credential. Also flags in-document that these notes are a poisoning vector: a future agent reading them out of the palace must not conclude it should hand-stamp. |
||
|
|
661ee20b39 |
docs(rfc-001): require solitary-first operation, opt-in centralization
Adds §1.1 (R1–R6) and §1.2 as a hard constraint on the design rather than a preference. A shared palace is valuable only for the multi-machine / multi-container pattern; for most users of the published pi-devbox and opencode-devbox images it is useless overhead. Evidence: all three sampled opencode-devbox deployments contain zero mempalace references. Requirements: solitary operation stays the default and stays byte-identical (no extra process, no outbox, no network calls); opt-in via docker-compose.yml + .env only, never an image rebuild; credentials only in .env, never in compose, the image, a command line, or a log; no new required services (the primary stays a separate standalone compose project); degrade-not-fail where mempalace is absent; fully reversible. §1.2 extends the convention pi-devbox already ships (.env.example:12-23, "local by default" with commented MEMPALACE_REMOTE_URL/_TOKEN) into a three-state ladder — local stdio (default, unchanged) / direct remote (unchanged) / edge (new, MEMPALACE_EDGE=1) — instead of inventing a new mechanism. Notes that the devbox-palace volume coupling reverses under edge mode: the local palace holds the outbox, so persisting it becomes required rather than irrelevant. Consequence recorded in §4.1: mempalace-edge must be selected at registration time, not left always-in-path to decide by env at runtime, since that would insert a process and a failure mode into every solitary user's setup. pi branches in createClient(); opencode needs its static MCP JSON templated at container start, which is new open question §9.6. Also: adds an R1 acceptance test, marks pi-devbox/.env.example:21 as stale (advertises mempalace-mcp --transport http rather than mempalace serve --token/--tls-cert), and extends the evidence index. |
||
|
|
35b1e3d81d |
docs: add RFC 001 — global palace with local fallback (mempalace-edge)
Design for moving from one MemPalace per machine per harness to a single primary palace with per-machine offline fallback, plus the source-verified archaeology behind it. Key decisions recorded: - Sync operations (MCP tool calls), not databases. Embeddings are computed client-side and are not portable across architectures; the KG `triples` table has no UNIQUE(subject,predicate,object,valid_from) and its ids embed datetime.now(), so row copies duplicate facts. Replaying tool calls is idempotent where it matters. - Implement as `mempalace-edge`, a local stdio MCP proxy (child mempalace-mcp + HTTPS to the primary + outbox.sqlite), not as per-harness patches. Needs zero mempalace internals, so it serves pi, opencode and the CLI alike and survives mempalace upgrades. - Merged reads (query both, re-sort, dedupe by drawer id) give fleet-wide recall without replication — which is why Phase 3 (pull replication) is deferred: "own writes plus whatever it can reach" is good enough. - Per-wing replication policy: curated/content-addressed wings replicate, mined code/docs stay local (derived, re-mineable, path-dependent ids). Also documents two silently destructive footguns to avoid during rollout: `--palace` vs MEMPALACE_PALACE_PATH (the KG follows the flag only, so `mempalace serve` can start with a silently empty knowledge graph), and `mempalace sync`, which is gitignore-aware drawer deletion rather than replication and would wipe fleet memory when run from a host lacking the repos. Notes that mempalace 3.6.0's `serve` already ships token auth + TLS, making the "unauthenticated, front it with a proxy" notes elsewhere stale. Includes an evidence index mapping each claim to file:line in mempalace 3.6.0, and four upstream candidates (origin_host provenance, per-wing ACL, mempalace_kg_supersede missing from service.py WRITE_TOOLS, and a sync --refuse-shared guard). |
||
|
|
96699f2a17 |
feat(pi-bridge): external MemPalace transport via MEMPALACE_REMOTE_URL
Let the pi<->mempalace bridge connect to a shared MemPalace over HTTP instead of always spawning a local mempalace-mcp: - Extract IMcpClient; rename McpClient -> StdioMcpClient (ctor command, arg-less start()). - Add RemoteMcpClient (vendored from pi-extensions/mcp-loader.ts): streamable-HTTP with AbortController timeouts, protocolVersion pinned 2024-11-05, alive/ensureAlive. mempalace-mcp --transport http is sessionless JSON-RPC today; session/SSE/404 branches retained for a future streamable-HTTP server. - createClient() selects transport from MEMPALACE_REMOTE_URL; MEMPALACE_REMOTE_TOKEN -> Authorization: Bearer. Lifecycle automation (wake-up, /mempalace-diary) unchanged. - scripts/check-mcp-client-sync.sh: drift guard vs canonical mcp-loader.ts. - README: document local-vs-external transport. Typechecks clean (strict); both transports smoke-tested against live mempalace-mcp. |
||
|
|
e12b624cf7 |
feat(pi-ext): self-healing respawn + scoped init timeout for mempalace-mcp
A stall-kill (or any crash) of mempalace-mcp was a permanent latch: available flipped off and stayed off until pi restart. Now the next tool call transparently respawns the server and retries. - ensureAlive(): bounded respawn with capped exponential backoff (MEMPALACE_MCP_MAX_RESPAWNS, default 2; MEMPALACE_MCP_RESPAWN_BACKOFF_MS, default 1000). Respawn budget resets on any successful JSON-RPC response, so a recovered server regains full patience while a persistently-broken one hits the cap and stays down (no hot-loop). - Init timeout default raised 120000 -> 300000 (scoped to init only): a genuine virtiofs cold-open shouldn't be killed mid-progress only to respawn and re-pay the same cost. Per-call timeout stays 60000. - Concurrency hardening: generation counter so a late exit from a killed old process can't tear down a fresh respawn; explicit healthy flag replaces racy proc!=null liveness check. - README: document self-heal, new env vars, and why generous-init + bounded-respawn compose rather than overlap. |
||
|
|
a3b8829991 |
feat(pi-ext): per-request timeout + stall-kill for mempalace-mcp
A wedged mempalace-mcp (classically an OrbStack virtiofs cold-open of a large chroma.sqlite3 / HNSW load) left the awaiting JSON-RPC promise pending forever, freezing the pi TUI uninterruptibly: ESC cancels the LLM stream, not a pending tool execute(). The JSON-RPC client now arms a per-request timer. On expiry it rejects the request AND kills the stalled child (SIGTERM->SIGKILL), so pi gets an error instead of hanging; the extension flips available=false so later calls fail fast (restart pi to retry). Per-REQUEST, not per-process: the long-lived server only dies on a genuine stall. Knobs: MEMPALACE_MCP_TIMEOUT_MS (default 60000), MEMPALACE_MCP_INIT_TIMEOUT_MS (default 120000), 0 = disable. This supersedes the planned standalone stdio-watchdog shim: the extension already owns request/response correlation, so a separate framing-reparsing shim is unnecessary. |
||
|
|
ce09d25c97 |
Rename to @earendil-works/pi-coding-agent + earendil-works/pi URL
Pi moved to its new home at earendil-works on 2026-05-07 (https://pi.dev/news/2026/5/7/pi-has-a-new-home). Sweep: - extensions/pi/mempalace.ts: 'import type { ExtensionAPI } from "@mariozechner/pi-coding-agent"' -> @earendil-works/pi-coding-agent. - README and extensions/pi/README: github.com/mariozechner/pi-coding-agent URL refs -> github.com/earendil-works/pi. - install.sh: same URL substitution in the user-facing pointer line. Brew install references (`brew install pi-coding-agent`) left as-is: formula still works at 0.73.1, tap update tracked upstream at earendil-works/pi#2755. |