Compare commits

..

7 Commits

Author SHA1 Message Date
joakimp b615571913 changelog: release v1.8.11 — the mail nobody could deliver, and the ship that skipped a file
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 19s
Publish Docker Image / base-decide (push) Successful in 10s
Publish Docker Image / build-base (push) Successful in 53m20s
Publish Docker Image / smoke (push) Successful in 4m48s
Publish Docker Image / smoke-studio (push) Successful in 4m58s
Publish Docker Image / build-variant-studio (push) Successful in 17m0s
Publish Docker Image / build-variant (push) Successful in 26m43s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 24s
Names three things riding in via floating refs, since the rule this
fleet adopted (v1.8.9) is to name a behaviour change before tagging,
not after:

- mempalace-toolkit b2b50af -> 21023e7: the rsync --checksum ship fix
  (a real behaviour change to every remote-mode client) and RFC 003
  §7.13 (docs only).
- skillset 6eb20af -> a12fe5e: mermaid-diagrams cutU normalisation and
  the mempalace skill's from_agent identity rule, plus the vendored
  fallback snapshot re-pinned to match (this pi-devbox commit).

Dependency audit run by direct command against every component, not
assumed: pi/mempalace/pi-atelier/pi-studio/pi-toolkit/pi-extensions/
pi-fork/pi-observational-memory all confirmed unchanged since v1.8.10.
pi-studio's apparent SHA drift on ls-remote turned out to be an
annotated-tag-object hash vs its underlying commit hash, not real
drift -- checked with ^{commit} before writing it down as 'none'.
2026-08-27 23:39:39 +02:00
joakimp 495b7e3859 vendor: resync mempalace skill snapshot to skillset a12fe5e
scripts/vendor-mempalace-skill.sh, real refresh not --check: skillset
moved 6eb20af -> a12fe5e (mermaid-diagrams cutU normalisation, and the
from_agent identity rule this same release ships in RFC 003). --check
reported stale-but-truthful (exit 0, the sanctioned skip) but this
release's point is getting today's fixes live fleet-wide, and the base
rebuild is already forced by the entrypoint change and the floating
mempalace-toolkit ref moving -- so the incremental cost of also
bumping this pin is zero. Phrase canary in scripts/smoke-test.sh
unaffected: neither pinned phrase ('Provenance is stamped for you'
present, 'Attribute what you file yourself' absent) is in the section
that changed; both verified still correct in the new snapshot.
2026-08-27 23:36:03 +02:00
joakimp 45850bc973 entrypoint: put back the shell state a recreate eats
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 25s
cli_utils' install.sh reaches PATH by symlinking bin/ into ~/.local/bin.
That is persistent on a host and ephemeral in a container, so the same
installer produced opposite durability and every --force-recreate sent the
human back to typing /workspace/cli_utils/bin/git-status-all. Re-link at
start instead, and add a per-device boot hook so the next question of this
shape needs no image change at all.

Symlinks rather than a PATH edit in an rc file, deliberately: ~/.local/bin
is already ahead of /usr/local/bin in ENV PATH, so links resolve in
NON-interactive shells too (docker exec, agent tool shells, scripts). An
rc-file PATH edit cannot reach those because ~/.bashrc returns early when
not interactive — measured, that asymmetry is exactly why `command -v
git-status-all` failed in one shell and worked in another on the same box.
Shell FUNCTIONS remain the sourced file's job; a symlink cannot carry them.

Guards, because ~/.local/bin is shared: a real file is never clobbered, a
symlink pointing elsewhere is never stolen, ours are refreshed, and links
into a cli_utils/bin whose target vanished are pruned — a dangling link on
PATH reads as a broken container rather than a removed script.

The hook (~/.config/devbox-shell/init.sh) adds no trust boundary: that dir
is already sourced into every interactive shell by the baked bash_aliases,
so it is already arbitrary code from the same owner. Only WHEN it runs is
new. bash <file>, never sourced, exit status ignored, output to a log.

Caught before commit, and the reason the loops use `if` bodies instead of
`&&` chains: under `set -euo pipefail` a for-loop whose last command is a
false test exits non-zero, and with no match the /workspace/*/cli_utils
glob stays literal — so the first draft would have failed to START a
container on every machine that does not have this repo, rather than merely
skipping the links. Re-tested with set -e in place: no-cli_utils/empty-HOME
no-op, guards, idempotence, CLI_UTILS_LINK=0, a hook that exits 7, and a
hook that tries to mutate CLI_UTILS_BIN — all exit 0 with intact state.

Not covered by CI: docker-publish.yml runs only on tags and lint.yml lints
workflow run: steps, so neither executes this file. Validated by extracting
both sections and running them against fixtures, then for real in a live
v1.8.10 container. No smoke assertion added on purpose — the positive path
needs a /workspace mount smoke does not have, and asserting it there would
repeat the v1.8.0 mistake of a smoke check written against a stage that
does not exist at run time.

Moves the base hash (base-decide folds `cat entrypoint.sh
entrypoint-user.sh`), so this rides along with the next tag's ~40-minute
base rebuild rather than justifying a tag of its own.
2026-08-27 21:52:51 +02:00
joakimp 6891dc32b8 changelog: release v1.8.10 — deploy the scrubber, and name the check that must not be skipped
Publish Docker Image / build-variant-studio (push) Successful in 17m4s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 14s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-variant (push) Successful in 18m45s
Publish Docker Image / build-base (push) Successful in 41m59s
Publish Docker Image / update-description (push) Successful in 6s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / smoke (push) Successful in 4m55s
Lint / hadolint (push) Successful in 8s
Publish Docker Image / smoke-studio (push) Successful in 5m9s
Audited every component against upstream rather than assuming. Only two moved
since v1.8.9:

- mempalace-toolkit 5b8d78f -> b2b50af (ours): feeder scrubber, the symlink fix
  that stops it refusing to stage, hlc owed-set join, queued-delivery note,
  explicit MEMPALACE_MAILBOX_NOTIFY protocol modes, RFC 003 + fleet-memory docs.
- pi-studio v0.9.48 -> v0.9.52 (upstream, studio variant only): 22 commits, all
  additive — PDF previews, hideable header, contextual side questions. No removals
  or renames. Its pi floor is >=0.84.3 against our exact 0.84.3 pin: satisfied,
  zero headroom, named as a watch item because the next floor bump breaks the
  studio job only, after core has already published.

Verified unchanged, with dates proving they predate the v1.8.9 bake: pi 0.84.3 (=
npm latest), mempalace 3.8.0 (= PyPI latest), pi-atelier v0.8.2, playwright 1.62.1
(no drift this cycle despite floating on latest), pi-fork bf702b4, pi-obsmem
ce9fc98, pi-toolkit 0e1369e, pi-extensions 2022887. No breaking changes anywhere,
so everything can ship in one tag.

The reason for tagging now is not features. The scrubber has been live on exactly
ONE device since this morning; every other device has kept staging unscrubbed
transcripts into a shared palace, and cleanup after the fact is manual redaction
(done twice today, with freed-page residue left behind by choice).

The entry leads with the acceptance check instead of burying it, because this
release's worst failure is silent: the feeder is fail-closed, so a packaging or
path mistake stops the fleet's memory feed and nothing complains — refusing to
stage looks exactly like a quiet session. f0bffd1 exists because that very bug was
real (BASH_SOURCE reports the symlink path, not the target). First client on the
new image must confirm a [scrub] summary line appears, the drawer count moves, and
exit 3 did not fire. Silence is the failure signal, not success.
2026-08-27 17:51:46 +02:00
Joakim Persson 8a673ec143 docs: unclip the diagrams, and answer what compaction leaves behind
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 15s
Two problems reported against docs/observational-memory.md, one cosmetic and one
substantive. Both turned out to be worth more than the fix.

Clipping. Several boxes lost their bottom line of text in the viewer, and neither
the source nor my own render showed it. Cause: Mermaid measures a node label with
its own font metrics, commits to a box size, then renders that label as real HTML
in a <foreignObject> — so any host stylesheet touching line-height or font-size
inflates the text past a box that is already fixed, and the overflow is clipped.
The error accumulates per line, which is why it always eats the last line of the
tallest labels.

Rejected the obvious fix after testing it rather than assuming it:
%%{init: {'flowchart': {'htmlLabels': false}}}%% *is* honoured (labels switch from
16 foreignObject to 7 tspan) and clips identically, because inflated font-size
inherits into SVG text too. The fix that works is a hard limit of two short lines
per node, with the detail moved into prose under each diagram — one- and two-line
boxes have the vertical slack to absorb inflation, three- and four-line boxes do
not. The nodes were carrying paragraph-sized text; the diagrams are better for
losing it.

Regression harness: render every block with line-height: 1.7 !important forced
onto the label HTML, screenshot, read it. That caught two survivors of the rewrite
that looked fine in the clean render — a long unbreakable /opt path wrapping to a
third line, and a cylinder shape whose curved bottom leaves less room than a
rectangle for the same two lines.

New §4, because the document explained that compaction folds the ledger and never
said what that leaves in the context. Asked directly: is the session back to
knowing nothing? No. A verbatim tail survives, sized by keepRecentTokens (20k) and
cut only at turn boundaries; the system prompt and AGENTS.md were never in the
compacted region because they are rebuilt from disk each request; nothing is
deleted from disk, since compaction appends a compaction entry rather than
rewriting lines; and recall keeps resolving ids whose sources left the context
because it reads the full branch via sessionManager.getBranch() and never consults
the context window. Repeated compaction renders from live records, not from the
previous summary's prose, so there is no generation-loss spiral.

Also corrects this repo's own "compaction calls no model" to the steady-state
claim it actually is: with an empty ledger the hook returns nothing and explicitly
declines ownership, and pi's native model-based summariser runs. Snippet quoted in
the doc.

And a bug shipped in the first version: §9 said the ledger entries are
custom_message and specifically not custom. Exactly backwards, so the one grep
that section existed to get right was the one it got wrong. Verified empirically
against the live session file — 11 om.observations.recorded and 6
om.reflections.recorded, all "type":"custom", next to "type":"custom_message"
entries whose customType is mempalace-mailbox and mempalace-wakeup. That is where
the confusion came from, and the distinction is load-bearing rather than
cosmetic: the mailbox uses the context-visible append API, om's ledger uses the
invisible one. Which makes it a feature the doc now advertises — the ledger costs
zero context until it is folded.
2026-08-27 14:24:19 +02:00
joakimp cdb6fc0950 changelog: name what the floating toolkit ref will pull into the next tag
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 16s
MEMPALACE_TOOLKIT_REF=main floats and docker-publish.yml resolves it to a SHA at
build time, so toolkit main moving 5b8d78f -> f0bffd1 (10 commits) ships in the
next tagged image whether or not this repo has a commit. v1.8.9 adopted the rule
after that shape bit twice (553d865, 5b8d78f): name the behaviour change BEFORE
tagging. This is that rule obeyed rather than re-learned — the work was pushed to
toolkit main earlier today and this entry was missing, which is exactly the gap
that caused a cross-host misattribution in v1.8.7.

Contents: the feeder-side secret scrubber (three tiers, T3 report-only after a
measured 403 -> 29 false-positive calibration on 52 MB of real transcripts, fail
closed); the symlink near-miss that fail-closed would have turned into a
fleet-wide silent memory outage at bake time, caught before tagging; the mailbox
work (hlc owed-set join, queued-until-next-turn note, explicit notify protocol
modes, tmux path documented unverified); and the documentation set (RFC 003,
fleet-memory.md, secret-hygiene.md). Also states what remains unscrubbed: the
opencode bridge write path and the unbuilt server-side layer.

No tag pushed — per the release protocol, no tag means no build.
2026-08-27 14:20:35 +02:00
Joakim Persson 14371e2da6 docs: explain the memory that runs itself, and fix a claim the mailbox falsified
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 16s
`pi-observational-memory` is baked, registered in the seeded settings.json, and
handed a cheaper model than the session it serves — and the only description in
this repo was five words in a feature list. Someone meeting `/om:status` or a
"compacted memory" block had nothing to read that said whether to leave any of it
on. New docs/observational-memory.md, 272 lines and five diagrams.

Scoped by who is authoritative, so there is one copy of each claim:

- Upstream (/opt/pi-observational-memory/docs/) already documents the mechanism
  well — concepts.md, how-it-works.md, configuration.md, including a v3 lifecycle
  diagram that matches the deployed code. Linked, not re-derived.
- This document takes the four facts pi-devbox owns and can change: the pinned
  commit it bakes (v3.0.4 ce9fc98, matching build-manifest.json), the packages[]
  entry that decides which copy loads, the Haiku-workers-vs-Opus-session split,
  and the devbox-pi-config volume that makes the ledger outlive the container.
- Plus the confusion this image creates by shipping two things called memory: a
  section contrasting it with MemPalace, on the line "observational memory keeps
  a session coherent, the palace keeps the fleet coherent".

Placement follows the split fleet-ops states for itself — reusable mechanism is
not deployment data — so a "why is this in my container" document belongs in the
repo that pins and wires the component. Linked twice from the README, because
until this commit the README referenced docs/ zero times and the file already
sitting there was reachable only by listing the directory.

Numbers were read out of the live container and the baked tree rather than out of
release notes, which caught one thing the pi-extensions skill still has wrong:
the dropper is gated on a successful same-turn reflection, not on a token
threshold of its own.

Also fixes a README sentence that v1.8.9 made false. § Cross-machine agent
coordination ended with "Nothing in this image polls the log on the agent's
behalf"; the mailbox has shipped since aac4a1c. Replaced with the three knobs and
their defaults, derived-not-read owed-ness, and the queued-into-the-next-turn
delivery measured on two devices — and dated to "as baked in v1.8.9
(mempalace-toolkit 5b8d78f)", pointing at RFC 003 §7.11–§7.12 for the mechanism,
because toolkit main is already ahead (a92c75d pings the human who is not
looking) and describing that here would trade a stale-behind claim for a
stale-ahead one.

The cause is worth more than the fix: the behaviour arrived through the floating
MEMPALACE_TOOLKIT_REF, so no diff in this repo ever touched the paragraph making
the claim. v1.8.9's own rule fired for the CHANGELOG and nobody swept the README.
The CHANGELOG records what changed, the README asserts what is true, and only the
first is reviewed at release time — so the rule now extends to grepping the README
for absolute claims (nothing, never, does not, only) about a component whose SHA
moved.

Diagrams were verified by rendering, not by parsing. Both comparison diagrams
parsed clean and rendered with their meaning reversed: Mermaid laid the second
declared subgraph out first, putting "with observational memory" before "without"
and MemPalace before observational memory in the diagram whose entire job was
that contrast. Rebuilt as declaration-ordered chains and re-rendered at mermaid@11
— the version pi-studio pins — in the baked headless browser, then read back as an
image. mermaid.parse() proves syntax and says nothing about layout.
2026-08-27 13:55:55 +02:00
7 changed files with 1126 additions and 4 deletions
+8
View File
@@ -146,6 +146,14 @@ GIT_USER_EMAIL=
# Detection is automatic if the skillset lives at WORKSPACE_PATH/skillset.
# SKILLSET_CONTAINER_PATH=
# ── cli_utils (standalone commands from a mounted checkout) ──────────
# If a cli_utils repo is mounted, the entrypoint symlinks its bin/ commands
# into ~/.local/bin on every start, so they survive container recreate and
# resolve in non-interactive shells too (docker exec, agent tool shells).
# Detection is automatic at WORKSPACE_PATH/cli_utils (or one level below).
# CLI_UTILS_CONTAINER_PATH=
# CLI_UTILS_LINK=0 # disable the linking entirely
# ── Locale ───────────────────────────────────────────────────────────
# LANG=sv_SE.UTF-8
# LANGUAGE=sv_SE:sv
+505
View File
@@ -11,6 +11,511 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
---
## v1.8.11 — 2026-08-27
**Shell state that the writable layer eats on every recreate now gets rebuilt at
start.** Two additions to `entrypoint-user.sh`, both idempotent, both silent
no-ops when the thing they wire up is absent.
**`cli_utils` commands are linked onto `PATH`.** If a `cli_utils` checkout is
mounted, every executable in its `bin/` is symlinked into `~/.local/bin` at
container start — `git-status-all`, `git-pull-all`, `devbox-sanity`,
`pi-devbox-sanity`, `pi-session-repair`, `docker-clean`, `vpn-status`. Detection:
`CLI_UTILS_CONTAINER_PATH` → `/workspace/cli_utils` → `$HOME/cli_utils` →
`/workspace/*/cli_utils`; `CLI_UTILS_LINK=0` disables it.
The reason this is an *image* concern and not the user's problem to re-solve: on a
host, `cli_utils/install.sh` puts those commands on `PATH` by symlinking them into
`~/.local/bin`, which is persistent there — and **ephemeral here**. Same installer,
same repo, opposite durability, so the fix died on every `--force-recreate` and the
next session was back to typing `/workspace/cli_utils/bin/git-status-all`. Running
`install.sh` *inside* a container is the trap rather than the fix: it re-creates
the same disposable state.
**Symlinks rather than a `PATH` edit in an rc file, deliberately.** `~/.local/bin`
is already ahead of `/usr/local/bin` in `ENV PATH`, so links resolve in
**non-interactive** shells too — `docker exec <c> git-status-all`, agent tool
shells, scripts. An rc-file `PATH` edit cannot reach those: `~/.bashrc` returns
early when the shell is not interactive. Measured on tor-ms22 2026-08-27,
`command -v git-status-all` failed in a non-interactive shell while succeeding in
an interactive one, from exactly that asymmetry. Guards, because `~/.local/bin` is
shared with other tooling: a real file is never clobbered, a symlink pointing
somewhere else is never stolen, our own links are refreshed, and links into a
`cli_utils/bin` whose target vanished are pruned — a dangling link on `PATH`
reports "No such file or directory" and reads as a broken container rather than a
removed script.
**A per-device boot hook: `~/.config/devbox-shell/init.sh`.** If the host provides
one, it runs once at start with output to `~/.pi/agent/devbox-init.log`. That
directory is the host-owned bind-mount already sourced into every interactive
shell by `/etc/skel-devbox/.bash_aliases`, so this is its boot-time twin — the
same ownership and the same persistence, but running *before any shell*, which is
what non-interactive fixups (symlinks, directories, one-off migrations) need. **It
introduces no new trust boundary**: that path is already arbitrary code from the
same owner; only *when* it runs is new. Invoked as `bash <file>`, never sourced,
and its exit status is ignored — a hook must not be able to mutate the
entrypoint's own shell state or stop a container from starting.
With the hook in place, the next "can this run on every recreate?" question needs
no image change at all — which is the point, given what the next paragraph costs.
**This moves the base hash.** `base-decide` folds `cat entrypoint.sh
entrypoint-user.sh` into it, so this change forces the ~40-minute base rebuild at
the next tag whether or not anything else in the base moved. It is a rider, not a
reason to tag.
**How it was validated, since CI cannot.** `docker-publish.yml` runs only on
`push: tags: v*`, and `lint.yml` runs `actionlint` over workflow `run:` steps —
neither one executes `entrypoint-user.sh`. So both sections were extracted and run
against fixtures in a throwaway `$HOME` before commit: real file not clobbered,
foreign symlink respected, stale link pruned, new command picked up, second run
byte-identical, `CLI_UTILS_LINK=0` honoured, and "no `cli_utils` anywhere" a silent
`exit 0`. Then run for real in a live v1.8.10 container, after which
`command -v git-status-all` resolved in a *non-interactive* shell. No
`smoke-test.sh` assertion was added on purpose: the positive path needs a
`/workspace` mount that smoke does not have, and asserting it there would repeat
the v1.8.0 mistake of a smoke assertion written against a stage that does not
exist at run time. `workflow_dispatch` with `smoke_only` remains the way to
exercise this against `HEAD` before a tag.
**Also carried by the floating `mempalace-toolkit` main ref** (resolved at build
time, not by a pi-devbox commit — `MEMPALACE_TOOLKIT_REF=main`):
**A scrubbed re-export of a dormant session could silently never reach the
palace host.** `bin/mempalace-pi-session` ships to the palace with
`rsync -a --update`, and the stage file's mtime is deliberately the SOURCE
transcript's mtime (`os.utime()`, "preserve session mtime for dedup
stability"). Re-exporting a session that has not been appended to since its
last ship therefore produces a mtime that is *not newer* than the receiver's —
exactly the case a redactor upgrade needs to ship, since content differs while
mtime does not. `--update` reported success and sent nothing. Found and
patched by `pi@mbp-m1-2020` (mempalace-toolkit `a361b71`): `--update` →
`--checksum`, which compares content and ignores size/mtime entirely.
Dropping `--update` outright was considered and rejected — rsync's default
quick check already transfers on a size difference alone, which would have
masked the *next* instance of this (a redaction whose placeholder happens to
match the secret's length) as fixed. `os.utime()` is untouched; its backdating
is a separate, load-bearing design call for dedup stability. New regression
test, `scripts/test-rsync-ship-idempotency.sh`, runs fully offline (a local
rsync destination exercises the same size/mtime/checksum comparison as the ssh
transfer) and is built to *discriminate*: it must fail against `--update` and
pass against `--checksum,` not merely exercise the code path — the first draft
of the test used fixture strings of different lengths and passed for the wrong
reason (rsync's quick check transfers on size difference alone regardless of
`--update`), which is the same trap the patch itself was written to avoid.
**Acceptance line for this class of change going forward:** "receiver sha256
matches sender for every staged file", not "local stage is clean" — a clean
local stage says nothing about what a dormant session already sent.
**An event addressed to an identity no session runs as is delivered to
nobody, and this fleet has now hit it three separate ways.** RFC 003 gains
§7.13 and open-decision 10 (mempalace-toolkit `21023e7`, docs only, no image
behaviour change): the owed-set derivation — the log's only push channel — is
keyed on `to_agent`, and a reply is always addressed back to whatever string
the *original writer* put in `from_agent`. Nothing validates that string
against a live session identity, so authoring under a synthetic or foreign
name makes every reply to that event write-only. Measured cost this cycle: a
directed ask planted under a synthetic sender drew a correct reply containing
an urgent security finding, and it sat unread for ~2h20m, found only because a
human asked whether mail had arrived. Permitted exception, unchanged: a
synthetic sender is fine for a deliberate control experiment, provided the
body names the real identity to reply to.
**Also carried by the live `skillset` mount** (each device's own clone, not
baked — except the `mempalace` skill's fallback snapshot, re-vendored below):
**The mermaid-diagrams checker's cut gate moved from client pixels to a
per-SVG user-space unit.** `CUT_PX` was calibrated against one live page at
one render scale; sweeping `--viewport` 500→1600 on an *unchanged* document
moved the worst overflow −1.0px → −3.0px, near-proportional to the viewport,
i.e. a constant geometric overflow viewed through a changing scale. `cutU =
cutPx / scale` (scale taken per-SVG, never a page average — one page mixes
scales 0.643–0.988) recovers that invariant: the sweep now collapses to
exactly −3.0u at every width. Re-deriving the threshold against the live host
surfaced a real false negative the old pixel gate had: a label at
`cutPx=0.4, scale=0.678` read as healthy under `CUT_PX=0.5` but is `0.59u` —
a genuine cut hiding behind a compressed render scale. `CUT_U` stays `0.5`;
`cutPx` and `scale` are still printed on every issue so a devtools ruler still
confirms the number on the actual page. A new, explicitly-deferred finding
from the same review: `cut` only measures vertically, so an unbreakable token
wider than its box (a long URL, a `snake_case` identifier) is invisible to
soft-wrap, tall, *and* cut simultaneously — filed as a backlog item, not
implemented, pending a fifth acceptance control.
**The `from_agent`-identity finding above is also now in the `mempalace`
skill itself** ("Writing to another machine", and Anti-Patterns), and the
baked fallback snapshot of that skill was refreshed to match
(`vendor-mempalace-skill.sh`, `6eb20af` → `a12fe5e`) — sanctioned to skip on
its own (`--check` reported stale-but-truthful), done anyway because this
release's point is getting today's fixes live, and the base rebuild below was
already forced regardless.
### Dependency audit (2026-08-27)
Every component checked against upstream by direct command, not assumed:
| Component | Baked in v1.8.10 | Upstream now | Action |
|---|---|---|---|
| **mempalace-toolkit** | `b2b50af` | **`21023e7`** | ships the rsync ship-fix + RFC 003 §7.13 (both above) |
| **skillset** (mempalace fallback snapshot) | `6eb20af` | **`a12fe5e`** | re-vendored (above); live-mounted devices already had it |
| pi | `0.84.3` (pinned) | `0.84.3` is npm latest | none |
| mempalace | `3.8.0` (pinned) | `3.8.0` is PyPI latest | none |
| pi-atelier | `v0.8.2` (pinned) | `v0.8.2` highest tag | none |
| pi-studio (studio variant) | `v0.9.52` | `v0.9.52` — `main`'s commit and the tag's commit are identical (0 either direction) | none |
| pi-toolkit | `0e1369e` | `0e1369e` (local clone HEAD == `origin/main`) | none |
| pi-extensions | `2022887` | `2022887` (local clone HEAD == `origin/main`) | none |
| pi-fork | `bf702b4` | `bf702b4` | none |
| pi-observational-memory | `ce9fc98` | `ce9fc98` | none |
pi-toolkit / pi-extensions checked against their actual Gitea origin (the
Dockerfile's `PI_TOOLKIT_REPO` / `PI_EXTENSIONS_REPO`), not a GitHub mirror —
querying `api.github.com` for those two returned nothing (rate-limited or
blocked; not investigated, the local clones are the source of truth anyway).
No SHA above is fork-supplied; a fork inventing plausible-looking upstream SHAs
is a recorded failure mode (v1.8.9), so every value here came from
`git ls-remote`, a local clone's own `origin/HEAD`, `npm view`/registry JSON,
or the PyPI JSON API, run directly.
---
## v1.8.10 — 2026-08-27
**This tag exists to deploy a fix and a safety net that are currently running on
exactly one machine.** The feeder scrubber has been hand-copied to `/opt` on one
device since this morning; every other device has kept staging unscrubbed
transcripts into the shared palace. Nothing here is a new capability for its own
sake.
`MEMPALACE_TOOLKIT_REF=main` floats: `docker-publish.yml` resolves it to a
concrete SHA at build time, so whatever is on toolkit `main` when the tag is
pushed ships in that image whether or not this repo has a commit. That is the
rule v1.8.9 adopted after `553d865`/`5b8d78f` shipped undocumented twice — *name
the behaviour change before tagging, not after* — and this entry is that rule
being obeyed rather than re-learned.
**`mempalace-toolkit` main moves `5b8d78f` → `b2b50af`** (13 commits, ~2100
insertions / ~520 deletions). No pi-devbox commit implements any of it.
### ⚠️ THE ONE CHECK THIS RELEASE MUST NOT SKIP
**On the first client that runs this image, verify the memory feed still stages.**
The feeder is *fail-closed* by design: no redactor module, no staging (`exit 3`).
That is correct behaviour and it is also the failure mode with no alarm — a
packaging or path mistake stops the fleet's entire transcript feed and nothing
complains loudly, because refusing to stage looks exactly like a quiet session.
This is not hypothetical. `f0bffd1` exists because the feeder is installed as a
symlink (`/usr/local/bin/mempalace-pi-session` → `/opt/mempalace-toolkit/bin/…`)
and `${BASH_SOURCE[0]}` reports the *symlink* path, so the module lookup landed in
a directory where it does not exist. Had that shipped, every device would have
refused to stage on first boot. It was caught by execution, not by review.
Acceptance, in order, on the first recreated client:
1. Run a session, then confirm the feeder logged a scrub summary — a
`[scrub]` line with tier-tagged counts (`T1:env-value=…`, `T2:github-pat=…`),
or an explicit "zero redactions". **Silence is the failure signal**, not success.
2. Confirm the palace drawer count *moved* for that session (the feed reached the
server, not just the stager).
3. Confirm `exit 3` did **not** fire: `mempalace-pi-session` invoked through the
`/usr/local/bin` symlink must find `mempalace_redact.py`.
4. Only then trust the rest of this release.
If step 1 or 3 fails, the memory feed is down fleet-wide until it is fixed, and
sessions that ran in the meantime are not recoverable from the palace — they were
never staged. `MEMPALACE_FEED_ALLOW_UNSCRUBBED=1` is the loud escape hatch, and
using it means accepting unscrubbed transcripts until the packaging is repaired.
### Dependency audit (2026-08-27)
Every component checked against upstream, not assumed:
| Component | Baked in v1.8.9 | Upstream now | Action |
|---|---|---|---|
| **mempalace-toolkit** | `5b8d78f` | **`b2b50af`** | ships the scrubber + symlink fix + `hlc` join |
| **pi-studio** (studio variant) | `v0.9.48` | **`v0.9.52`** | 22 commits, additive only — see below |
| pi | `0.84.3` (pinned) | `0.84.3` is npm latest | none |
| mempalace | `3.8.0` (pinned) | `3.8.0` is PyPI latest | none |
| pi-atelier | `v0.8.2` (pinned) | `v0.8.2` highest tag | none |
| playwright | `1.62.1` (floats `latest`) | `1.62.1` | none — no drift this cycle |
| pi-fork | `bf702b4` | `bf702b4` (2026-08-24) | none |
| pi-observational-memory | `ce9fc98` (v3.0.4) | `ce9fc98` | none |
| pi-toolkit | `0e1369e` | `0e1369e` (2026-08-07) | none |
| pi-extensions | `2022887` | `2022887` (2026-08-17) | none |
**`pi-studio` `v0.9.48` → `v0.9.52`** — four releases, 22 commits, all additive:
PDFs open directly in Studio with watched previews, the header can hide, and
contextual *side questions* arrive (selected-tool use, frozen git context, export,
keyboard shortcuts). No removals or renames in the diff; the changes are
concentrated in `client/studio-client.js`, `index.ts` and three new `shared/`
helpers.
**Its `pi` floor is `>=0.84.3` and we pin exactly `0.84.3` — satisfied with zero
headroom.** Worth naming as a watch item rather than a problem: the next studio
release that raises the floor breaks the studio variant until `PI_VERSION` moves,
and that failure surfaces at build time in the studio job only, after the core
variant has already published.
### Also pulled in by the floating toolkit ref (documentation only)
RFC 003 gains **§9.2**, a proposed direction for the one open decision this
fleet keeps tripping over — that a report addressed to a device is never
delivered, because mailbox candidacy requires exactly `status="open"`. It records
a negative result worth keeping: widening the owed set to include terminal events
cannot work, since the asserting shape and the clearing shape must be disjoint or
every closure mints a fresh obligation. No code implements §9.2 in this release.
The skillset snapshot also moves, so this image bakes the mermaid-diagrams skill's
Playwright driver and the honest note that a `claimed` ack notifies nobody.
### Transcripts get scrubbed before they are staged (`3d47937`, `836e35b`, `f0bffd1`)
`bin/mempalace_redact.py`, called from `mempalace-pi-session` at the moment the
staged transcript is written — one hook covering both transports, because local
mode mines that file and remote mode rsyncs the same bytes.
- **Why it exists, measured rather than argued.** One leaked bearer token had
reached 3 drawers, 13 feeder inbox files across all three devices, and 10 local
files spanning 10 days, from an agent printing an env var while debugging. A
second sweep then found `GITEA_ACCESS_TOKEN` in 2 more drawers and
`GITEA_EGL_ACCESS_TOKEN` in 3. This is routine agent behaviour, so the fix
belongs in the pipeline, not in discipline.
- **Detection is name-anchored, never entropy-anchored.** A palace's own primary
keys — drawer ids, chunk ids, event ids, replica ids, HLCs, commit SHAs — *are*
its high-entropy strings, so an entropy detector eats the memory it protects,
silently and unrecoverably. Three tiers instead: T1 literal values from this
process's env whose name says secret (zero false positives by construction);
T2 vendor shapes (`ghp_`, `glpat-`, `xox*-`, `sk-`, `AKIA`, JWT, PEM, URL
credentials, `Authorization:`); T3 key-name-says-secret.
- **T3 is report-only, because the false-positive rate was measured.** On 52 MB
of real fleet transcripts T3 fired 403 times, mostly `${VAR}` interpolation in
compose files, TypeScript identifiers, a *type annotation*
(`credentials: Credentials`), an IPA attribute holding a date
(`krbPasswordExpiration`), AAAK diary shorthand, and terminal output following
an ssh `Password:` prompt. With interpolation/code-context/key-suffix guards the
enforced count fell **403 → 29** on the same corpus. `MEMPALACE_REDACT_STRICT=1`
makes T3 enforce.
- **Operational shape.** Fail closed — no redactor, no staging (`exit 3`),
overridable with `MEMPALACE_FEED_ALLOW_UNSCRUBBED=1`. Every run prints a count
*including* `0 redaction(s)`, because silence is indistinguishable from a
scrubber that never ran. Findings carry rule, label, length and `sha256[:8]` —
never the value.
**Near-miss this image would have shipped, caught before tagging (`f0bffd1`).**
The image installs `/usr/local/bin/mempalace-pi-session` as a **symlink** into
`/opt/mempalace-toolkit/bin`, and `${BASH_SOURCE[0]}` reports the invoked path,
not the target — so the sibling-module lookup resolved to `/usr/local/bin`, the
redactor was absent, and fail-closed did as instructed: `[FATAL] ... refusing to
stage`. Measured side by side, the symlinked invocation FATALed while the direct
one scrubbed 40 findings. **At the next bake that would have stopped every
feeder tick on every device — a silent fleet-wide memory outage, worse than the
leak the scrubber prevents.** Fixed by chasing the symlink chain in portable
shell (`readlink -f` avoided: GNU/newer-BSD only, and this script also runs
directly on macOS hosts) with colon-separated fallback candidates. Verified via
the symlink, the direct path, and a second-hop symlink. General lesson: fail-closed
converts "module not found" into an outage, which makes the module lookup
load-bearing infrastructure that must be tested through the invocation path the
fleet actually uses — not the convenient one from a checkout.
### The mailbox becomes explainable and mesh-safe (`bfe9c5c`, `a92c75d`, `e917662`, `ecc2a9c`)
- **Owed-set derivation joins on `hlc`, not `seq`** (`bfe9c5c`). `seq` is a
replica-local arrival counter — the same event is `#7` in one database and `#12`
in another — so a second replica would let already-answered asks resurrect.
`hlc` is immutable and replicated, fixed-width, so string comparison *is* causal
comparison. A safe no-op on today's single replica (verified: the positive-control
pair orders identically under both keys), correct once a mesh exists.
- **Delivered text now says it is queued** (`a92c75d`). Delivery uses `steer`
with no `triggerTurn`, and the poll fires on `agent_settled`, so nothing wakes
the model — a delivered ask sits until a human starts the next turn. Measured
case: a directed report sat unread for 2.5 hours. The note explains the agent is
not ignoring the ask, it is not running.
- **`MEMPALACE_MAILBOX_NOTIFY` gains explicit `=kitty` / `=osc777` modes**
(`e917662`). Terminal autodetection inside a container is not unreliable, it is
*blind*: `docker exec` forwards neither `KITTY_WINDOW_ID` nor `TERM_PROGRAM`, and
`TMUX` is unset because tmux runs on the host. Verified on a live process:
`TERM=xterm-256color` and nothing else.
- **The terminal path through tmux is documented as UNVERIFIED** (`ecc2a9c`).
Test sequences written to the pty produced no notification on a remote client;
tmux likely drops unknown OSC types without `allow-passthrough`, and multi-client
routing (one ask pinging every attached client) is an open question.
### Documentation (`e1cc759`, `982b001`, `d4d8bb6`, `d2764bf`)
- **RFC 003, the coordination-log spec the code had been citing all along** — it
did not exist anywhere. 11 sections, retrospective against mempalace 3.8.0
`logstream.py`, incl. owed-set derivation, ten dogfooded landmines and seven open
decisions. Non-obvious findings: `event_append` has **no** idempotency guard on
the write path (verify-before-retry; the replication path *is* guarded),
coordination tools are exempt from both palace locks by design,
`GET /logstream/events` **never existed** in 3.8.0 (not proxy-blocked), and
`mempalace sync` never touches the logstream — the log is permanent and unbounded.
- **`docs/fleet-memory.md`**, operator-facing: five storage types, a decision tree,
latency expectations (~2–5 min live session; next session while offline),
broadcast exclusion by design, fan-out, and the search-before-answer /
diary-at-session-end / verify-don't-retry habits.
- **`docs/secret-hygiene.md`**, incl. the tier definitions, the measured FP data,
stated false negatives, and the three server-side call sites (specified, not
built — tier 2 only there, since the hub cannot see a client's env).
- Phase 1 exposure record moved to the private fleet repo with a moved-note stub;
retention direction for the unbounded log (logrotate-style: never rotate
still-owed events, rotation invalidates held cursors, archive-verify-delete).
### The other memory system finally gets explained — `docs/observational-memory.md`
`pi-observational-memory` has been baked for several releases and described in
one line of the feature list (*"the `recall` tool for session compaction"*),
which is enough to name it and not nearly enough to use it. New 272-line
explainer with five diagrams, aimed at someone who has seen `/om:status` or a
"compacted memory" block and wondered whether to leave any of it switched on.
**Scoped to what this repo is authoritative for, because upstream already
documents the mechanism well.** `/opt/pi-observational-memory/docs/` ships
`concepts.md`, `how-it-works.md` and `configuration.md`, including a correct v3
lifecycle diagram — so the new document links those for depth and spends its own
words on the four facts pi-devbox owns and can change: the pinned commit it bakes
(v3.0.4 `ce9fc98`, the value in `build-manifest.json`), the `packages[]` entry
that decides which copy loads, the Haiku-workers-vs-Opus-session split seeded
into `~/.pi/agent/settings.json`, and the `devbox-pi-config` volume that makes
the ledger survive `--force-recreate`. Plus the confusion this image creates by
shipping two things called memory: a section contrasting it with MemPalace, on
the line *observational memory keeps a session coherent, the palace keeps the
fleet coherent*.
Every stated number was read out of the live container or the baked tree rather
than copied from release notes — including the correction that the dropper is
gated on a **successful same-turn reflection** and not on a token threshold of
its own, which is the one detail `pi-extensions/SKILL.md` still gets wrong.
Placement follows the audience split fleet-ops states for itself: reusable
mechanism is not deployment data, so a "why is this in my container" document
belongs in the repo that **pins and wires** the component, pointing upstream for
depth. Linked twice from the README, because before this commit the README
referenced `docs/` zero times and the one file already there
(`mempalace-broker-design.md`) was reachable only by listing the directory.
### A README claim that v1.8.9 made false, and how it got there
**§ Cross-machine agent coordination ended with "Nothing in this image polls the
log on the agent's behalf." The mailbox shipped in v1.8.9, so that sentence has
been wrong since `aac4a1c`.** Replaced with the three knobs and their defaults
(`MEMPALACE_MAILBOX`, `MEMPALACE_MAILBOX_POLL_MS` 300000,
`MEMPALACE_MAILBOX_RESURFACE_MS` 3600000), the fact that owed-ness is *derived*
rather than read off `status`, and the queued-into-the-next-turn delivery
semantics measured on two devices.
The cause is the one v1.8.9 wrote a rule about: the behaviour arrived through the
floating `MEMPALACE_TOOLKIT_REF`, so **no diff in this repo ever touched the
paragraph that made the claim**. v1.8.9's rule ("name a floating-ref behaviour
change in the CHANGELOG before tagging") fired for the CHANGELOG and nobody swept
the README. The CHANGELOG records what *changed*; the README asserts what is
*true*, and only the first is reviewed at release time. Extending the rule
accordingly: grep the README for absolute claims — *nothing*, *never*, *does
not*, *only* — about any component whose SHA moved.
**The replacement is dated on purpose.** It says it describes the bridge *as baked
in v1.8.9* (`mempalace-toolkit` `5b8d78f`) and points at that repo's
`docs/rfc-003-coordination-log.md` §7.11–§7.12 for the mechanism, because toolkit
main is already ahead of the baked copy (`a92c75d` makes delivery say it is queued
and ping the human who is not looking; `e917662` and `ecc2a9c` refine that notify
path) and none of it reaches a container until a base rebuild. Documenting those
here would have swapped a stale-behind claim for a stale-ahead one — the same
defect with the sign flipped.
### Diagrams verified by rendering, not by parsing
Both comparison diagrams **parsed clean and rendered with their meaning
reversed**: Mermaid laid the second declared `subgraph` out first, so "with
observational memory" appeared before "without", and MemPalace before
observational memory in the diagram whose entire job was that contrast. A third
was legible only at 1280px. Rebuilt as declaration-ordered node chains, then
re-rendered at mermaid@11 — the version `pi-studio` pins — in the baked headless
browser and read back as an image. Recorded because it generalises:
`mermaid.parse()` proves syntax and says nothing about layout, so a diagram is
unverified until someone has looked at it.
### … and rendering it in *my* browser was still not enough
Reported from a real viewer: several boxes had their bottom line of text sliced
off. Reproduced and root-caused rather than nudged — **Mermaid measures a node
label with its own font metrics, computes the box, then renders the label as real
HTML inside a `<foreignObject>`.** Any host stylesheet that touches the
`line-height` or `font-size` of that HTML makes the text taller than the box
already committed to, and the overflow is clipped at the box edge. Error
accumulates per line, so the loss always lands on the last line of the tallest
labels — which is exactly what was reported.
Two fixes were tried and only the second works:
- `%%{init: {'flowchart': {'htmlLabels': false}}}%%` — **rejected, and verified
ineffective rather than assumed so.** The directive *is* honoured (label
elements switch from 16 `foreignObject` to 7 `tspan`), and the clipping is
identical, because the inflated font-size still inherits into SVG text.
- **A hard limit of two short lines per node, with the detail moved into the prose
under each diagram.** One- and two-line boxes have enough vertical slack to
absorb the inflation; three- and four-line boxes do not. This is also better
documentation — the old nodes were carrying paragraph-sized text.
The regression harness is now the interesting artefact: render every block with a
deliberately inflated `line-height: 1.7 !important` on the label HTML, screenshot,
and read it. Two survivors of the rewrite were caught only by that harness — a
long unbreakable `/opt/pi-observational-memory` path silently wrapping to a third
line, and a cylinder (`[( )]`) shape, whose curved bottom leaves less room than a
rectangle for the same two lines.
### §4 answers the question the document left hanging: what compaction does to your context
Asked directly and worth writing down: *if the old conversation is folded away, is
the session back to knowing nothing?* No — and the specifics are all checkable
against pi 0.84.3's own `docs/compaction.md` and the extension's source:
- **A verbatim tail survives, sized by a token budget rather than a message
count.** Pi walks back from the newest entry until `keepRecentTokens` [20000],
and everything from that `firstKeptEntryId` onward is kept **unchanged**. Cut
points land on turn boundaries, never mid-tool-call.
- **The system prompt and `AGENTS.md` are not in the compacted region at all** —
they are rebuilt from disk on every request, so compaction cannot lose them.
- **Nothing is deleted from disk.** Compaction *appends* a `compaction` entry
carrying the summary and the cut pointer; no session line is rewritten in place.
- **`recall` therefore still resolves ids whose sources left the context**, because
it reads the full branch via `sessionManager.getBranch()` and never consults the
context window.
- **Repeated compaction does not summarise the summary.** The text is always
rendered from live observation/reflection records, so there is no
generation-loss spiral; the projection is incremental against the last full-fold
boundary and escalates to a true re-fold from the branch root at
`observationsPoolMaxTokens` [20000].
And one correction to this repo's own earlier claim: **"compaction calls no model"
is a steady-state property, not an absolute.** If the ledger is empty — compaction
firing before the observer has ever run — the hook returns nothing and explicitly
declines ownership (`// Decline ownership so Pi's native summarizer preserves the
pre-cut context.`), and pi's own model-based summariser runs. The doc now says so,
with the snippet.
### A shipped doc bug: the ledger entry type was stated exactly backwards
§9 told readers the entries are `custom_message` and specifically *not* `custom`.
It is the other way round, so the one grep the section existed to get right was
the one it got wrong. Corrected against the live session file — 11
`om.observations.recorded` and 6 `om.reflections.recorded` entries, all
`"type":"custom"`, alongside `"type":"custom_message"` entries whose `customType`
is `mempalace-mailbox` and `mempalace-wakeup`, which is precisely where the
confusion came from: **the mailbox uses the context-visible API, om's ledger uses
the invisible one.**
That is not a typo but a load-bearing distinction, and the fix turns it into a
feature the doc now advertises: `custom` entries *"do not participate in LLM
context"* (pi `docs/session-format.md`), so **the ledger costs zero context until
it is folded** — now a row in the cost table.
### Not covered by any of this
The opencode bridge is a separate write path the feeder hook never sees, and the
server-side layer is unbuilt — so a secret typed straight into `add_drawer`, or
staged by a non-pi client, still lands unscrubbed.
---
## v1.8.9 — 2026-08-26
The coordination log gets a reader, and the release checklist's last gate stops
+1 -1
View File
@@ -305,7 +305,7 @@ ARG MEMPALACE_TOOLKIT_REF=main
# no ~67-minute base rebuild. (scripts/check-base-hash.sh scans only
# Dockerfile.base, so no folding into the base hash is required — nor would
# it be correct, since this ARG changes nothing about the base's contents.)
ARG SKILLSET_SNAPSHOT_REF=6eb20af181f0147cb8c1377f6e36a6a47a68e8e5
ARG SKILLSET_SNAPSHOT_REF=a12fe5ecc71e60feb24791e3e33571105f1afba7
# Dockerfile.base sets description="pi-devbox — base image (variant-independent)"
# and every variant INHERITS it, so both published images used to advertise
+76 -2
View File
@@ -20,7 +20,9 @@ on the host.
- `pi-extensions` — TypeScript extensions for pi (preview, MCP bridges,
mempalace integration, etc.)
- `pi-fork` — the `fork` tool for spawning sub-agents
- `pi-observational-memory` — the `recall` tool for session compaction
- `pi-observational-memory` — durable session memory: the ledger that makes
compaction cheap, plus the `recall` tool. See
[`docs/observational-memory.md`](docs/observational-memory.md)
- `pi-atelier` — TUI sidebar: ordered panels, split-pane, themes. Pinned to an
audited tag; see [Version pins](#version-pins-pi-pi-atelier-mempalace)
@@ -536,6 +538,35 @@ to refresh.
Anything not on a volume is on the writable layer and is lost on
container recreate.
### Rebuilding ephemeral shell state at start
Two entrypoint steps put back the kind of state that the writable layer eats, so a
recreate does not cost you a manual re-install:
- **`cli_utils` commands.** If a `cli_utils` checkout is mounted, every
executable in its `bin/` is symlinked into `~/.local/bin` on start, so
`git-status-all` and friends are on `PATH` without a path prefix. Detection:
`CLI_UTILS_CONTAINER_PATH` → `/workspace/cli_utils` → `$HOME/cli_utils` →
`/workspace/*/cli_utils`. Set `CLI_UTILS_LINK=0` to disable. Existing real files
in `~/.local/bin` and symlinks pointing elsewhere are left alone, so a
deliberate override still wins; links whose target disappeared are pruned.
Do **not** run a host installer's `install.sh` inside the container to achieve
this — it writes to the ephemeral home and dies on the next recreate.
- **A per-device boot hook.** If `~/.config/devbox-shell/init.sh` exists it is run
once at start (`bash`, never sourced, exit status ignored), with output in
`~/.pi/agent/devbox-init.log`. `~/.config/devbox-shell/` is the host-owned
bind-mount whose `bash_aliases` is already sourced into every interactive shell,
so a hook there persists across recreates with no image change. Use it for
fixups that must exist *before any shell* — symlinks, directories, one-off
migrations.
The distinction that decides which mechanism you want: `~/.local/bin` is on `ENV
PATH`, so symlinks there work in **non-interactive** shells too (`docker exec <c>
<cmd>`, agent tool shells, scripts). A `PATH` edit in `bash_aliases` reaches only
*interactive* shells, because `~/.bashrc` returns early when non-interactive —
which is also why shell **functions** (fzf helpers and the like) can only come
from the sourced file, never from a symlink.
## MemPalace integration
MemPalace is installed in the base image and pre-warmed with the
@@ -579,7 +610,50 @@ convention that a directed event with `status="open"` is a request owed a reply
while a `*` broadcast owes nothing. The mechanism side (what the bridge stamps,
and why live SSE push depends on the palace deployment's reverse proxy rather
than on this image) is documented in the toolkit's `extensions/pi/README.md`.
Nothing in this image polls the log on the agent's behalf.
**Since v1.8.9 the bridge reads the log for you.** Earlier images were write-only
— they stamped provenance on the way out and never read back, so a directed ask
reached an agent only if that agent happened to run `mempalace_event_list`
itself. The mailbox is gated on the same two variables as the stamper, is on by
default, and derives what is *owed* rather than trusting `status` (an acked event
keeps matching a `status="open"` query forever, because the log is append-only):
| Variable | Default | Effect |
|---|---|---|
| `MEMPALACE_MAILBOX` | unset (on) | `0` disables mailbox reads entirely |
| `MEMPALACE_MAILBOX_POLL_MS` | `300000` | minimum gap between mid-session polls |
| `MEMPALACE_MAILBOX_RESURFACE_MS` | `3600000` | re-announce a still-owed ask after this long |
Delivery **queues, it never interrupts**: the poll runs when pi goes idle and the
message is steered into the *next* turn, so nothing wakes the model on inbound
fleet traffic. The practical consequence, measured on two devices: the message
appears in your session window and the agent acts on it when the next turn
starts — you are the trigger. (That describes the bridge **as baked in v1.8.9**,
`mempalace-toolkit` `5b8d78f`; the mailbox's own mechanism and landmines live in
the toolkit's `docs/rfc-003-coordination-log.md` §7.11–§7.12, which moves ahead of
whatever this image has baked.)
## Observational memory (in-session memory)
The image also bakes [pi-observational-memory](https://github.com/elpapi42/pi-observational-memory),
which is memory of a *different kind* from the palace and is easy to confuse with
it. It keeps a small branch-local ledger of observations and reflections while a
session runs, so when pi compacts the conversation the summary is a
**deterministic fold of that ledger rather than a model call**, and every item
keeps a 12-character id that `recall(<id>)` resolves back to the exact source.
In one line: **observational memory keeps a session coherent; the palace keeps
the fleet coherent.**
It is on by default, needs no habit from you, and sends its background work to a
cheaper model than your session (Haiku while the session runs Opus, in the seeded
`~/.pi/agent/settings.json`). Inspect it from inside pi with `/om:status` and
`/om:view`; turn all proactive work off for one run with
`PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi`.
What it is for, how the lifecycle works, what it costs, every setting and its
default, and how it differs from MemPalace:
[`docs/observational-memory.md`](docs/observational-memory.md).
## Agent skills
+371
View File
@@ -0,0 +1,371 @@
# Observational memory — why this image has it, and what it does for you
**Audience:** anyone using this container for long pi sessions who has wondered
what `recall`, `/om:status` and "compacted memory" are, or whether they should
leave any of it switched on.
**Companion documents:** the extension ships its own reference docs at
`/opt/pi-observational-memory/docs/` —
[`concepts.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/concepts.md)
(the model),
[`how-it-works.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/how-it-works.md)
(hooks and internals) and
[`configuration.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/configuration.md)
(every setting). Pi's own compaction mechanics are in
`/usr/lib/node_modules/@earendil-works/pi-coding-agent/docs/compaction.md`.
Those are normative; this document is the **deployment** view — what is pinned
here, how it is wired, what it costs, and how it differs from MemPalace. For the
palace, see
[`mempalace-toolkit/docs/fleet-memory.md`](https://gitea.jordbo.se/joakimp/mempalace-toolkit/src/branch/main/docs/fleet-memory.md).
> Verified on pi-devbox **v1.8.9** (`release_tag v1.8.9`, source `aac4a1c`),
> which bakes pi-observational-memory **v3.0.4** at commit `ce9fc98` — the value
> in `/etc/pi-devbox/build-manifest.json` → `components.pi-observational-memory`.
> Every number below was read from that tree, from pi 0.84.3's own docs, or from
> the live container.
---
## 1. The problem it solves
A long pi session outgrows the model's context window. Pi's answer is
**compaction**: fold the older part of the conversation into a summary and keep
recent messages verbatim. That is unavoidable, and it is where sessions go
wrong — the summary is produced *at the moment of pressure*, by a model, about a
transcript that is about to leave the context.
Observational memory changes *when* the remembering happens. Instead of
summarising in a panic at the end, it keeps a small **ledger** up to date while
the session runs, and compaction then just folds that ledger.
```mermaid
flowchart LR
A0["plain compaction"] --> A1["context fills"]
A1 --> A2["a model summarises<br/>under pressure"]
A2 --> A3["prose summary,<br/>no way back"]
B0["with observational<br/>memory"] --> B1["context fills"]
B1 --> B2["ledger written<br/>as you work"]
B2 --> B3["compaction folds<br/>the ledger"]
B3 --> B4["ids you can<br/>recall"]
```
Top row is pi on its own: one model call at the worst possible moment, detail
chosen in a hurry, and the original wording gone from view. Bottom row is this
image's default: the thinking happened earlier on a cheap model, the fold is
deterministic, and every line in the result carries an id that resolves back to
the exact source.
## 2. The mental model: three layers and a ledger
| Layer | What it is | Example |
|---|---|---|
| **Observation** | a timestamped, source-backed event from the conversation | "user rejected option B because it needs a base rebuild" |
| **Reflection** | a durable conclusion *backed by* observations | "the user optimises for avoiding 67-minute rebuilds" |
| **Drop** | a tombstone retiring an observation from active memory | the superseded detail of a bug that is now fixed |
These are appended to the session as silent ledger entries
(`om.observations.recorded`, `om.reflections.recorded`,
`om.observations.dropped`) and **folded** — replayed in order — to produce the
memory state. The ledger is the source of truth; what you see in a compacted
session is a rendering of it.
Two properties follow, and both matter later:
- **The ledger itself costs no context.** Those entries are pi `custom` entries,
which *"do not participate in LLM context"* (pi `docs/session-format.md`). They
sit in the session file and reach the model only via the fold at compaction.
- **Memory is branch-local.** A pi session is a tree (resume, fork), and the fold
follows the current branch only, so a forked branch does not inherit another
branch's view.
## 3. The lifecycle
Three background workers and one compaction hook, driven by *raw token
progress* rather than wall-clock time. Defaults in brackets.
```mermaid
flowchart TD
T(["turn_end"]) --> O{"10k raw tokens<br/>since observing?"}
O -- yes --> OBS["<b>observer</b> runs"]
O -- "no" --> R{"20k tokens<br/>since reflecting?"}
R -- yes --> REF["<b>reflector</b> runs"]
REF -- "if pool over 10k" --> DR["<b>dropper</b> prunes"]
S(["agent_settled"]) --> C{"81k tokens<br/>since compacting?"}
C -- yes --> CP["ctx.compact()"]
CP --> H(["session_before_compact"])
H --> F["fold the ledger<br/>no model call"]
F --> VIS["compacted memory"]
```
- **observer** — `observeAfterTokens` [10000]: writes observations for the
conversation it has not covered yet.
- **reflector** — `reflectAfterTokens` [20000]: promotes patterns across
observations into durable reflections.
- **dropper** — no clock of its own. It is post-reflection maintenance, gated on
a *successful same-turn* reflection **and** an active pool above
`observationsPoolTargetTokens` [10000]. Not a third worker on a third
threshold.
- **compaction** — `compactAfterTokens` [81000], checked when pi goes idle, so it
never interrupts a turn. Pi will also compact on its own when the context is
nearly full (`contextTokens > contextWindow - reserveTokens`, `reserveTokens`
[16384]).
## 4. What compaction actually does to your context
This is the question the rest of the document used to leave hanging: if the old
conversation is folded away, is the session back to knowing nothing?
**No.** Compaction replaces *part* of the context, not all of it, and it deletes
nothing at all from disk.
```mermaid
flowchart LR
SYS["system prompt<br/>+ AGENTS.md"] --> CTX["what the model sees<br/>on the next turn"]
SUM["folded memory:<br/>reflections + observations"] --> CTX
TAIL["recent turns,<br/>verbatim"] --> CTX
DISK[("session .jsonl: all of it")] -. "recall(id)" .-> CTX
```
Where each piece comes from:
- **System prompt and `AGENTS.md` — never compacted, because they were never
conversation.** Pi rebuilds them from disk on every request
(`loadContextFileFromDir`), so they cannot be lost by compaction.
- **The verbatim tail — sized by a token budget, not a message count.** Pi walks
backwards from the newest entry accumulating token estimates until
`keepRecentTokens` [20000] is reached; that entry becomes `firstKeptEntryId`,
and *everything from there on is kept unchanged*. Cut points land on turn
boundaries, never mid-tool-call. So the most recent ~20k tokens of real work —
your last instructions, the diffs, the test output — survive word for word.
- **The folded memory — replaces only what came before that cut.** Rendered from
the ledger's records: reflections and observations, each with its 12-hex id.
- **The session file — untouched.** Compaction *appends* a `compaction` entry
(`{"type":"compaction", summary, firstKeptEntryId, tokensBefore, …}`) and
rebuilds context from it on later turns. Nothing is rewritten in place; the
only documented way to remove session content is deleting the whole `.jsonl`.
That last point is what makes the answer to "is the detail gone?" *no* rather
than *mostly*: `recall` does not read the context window at all. It calls
`sessionManager.getBranch()` — the full branch from the root — and resolves an
observation id back to the original entries. Detail that left the model's view
an hour ago is still one `recall` away.
**Repeated compaction does not summarise the summary.** The rendered text is
always built from live observation/reflection *records*, never from the previous
compaction's prose, so there is no generation-loss spiral. (Mechanically the
projection is incremental — it re-derives back to the last full-fold boundary and
carries the rest forward, escalating to a genuine re-fold from the branch root
when the observation pool reaches `observationsPoolMaxTokens` [20000].)
So the honest summary of the state after compaction: **the model keeps its
instructions, keeps recent work verbatim, trades older turns for a dense
id-carrying digest of them, and can pull any of it back on demand.** Not a fresh
start — a smaller, cheaper, still-navigable one.
### One caveat about "no model call"
If the ledger is empty — compaction fires before the observer has ever run — the
hook returns nothing and *declines ownership*, and pi's own model-based
summariser runs instead:
```ts
const summary = renderSummary(projection.reflections, projection.observations);
if (summary.length === 0) {
// Decline ownership so Pi's native summarizer preserves the pre-cut context.
return;
}
```
In steady state (any session old enough to have produced one observation) om's
hook wins and compaction is model-free. "Never calls a model" is true in practice
and false in principle; the fallback is deliberate, so an empty ledger degrades
to normal pi rather than to no summary at all.
## 5. What you actually get
- **Compaction stops being a stall.** In steady state the latency path is
deterministic work over ledger entries, not a summarisation call.
- **Nothing important vanishes silently.** Compaction is lossy by design, but
every item keeps a 12-character id, and `recall(<id>)` returns the exact
evidence — original wording, reasoning, file path, error text.
- **The bookkeeping runs on a cheaper model than your session.** In this image
that is deliberate and visible (§7): background workers on Haiku, session on
Opus.
- **It is automatic.** No habit to maintain, unlike the palace protocol — which is
exactly why the two complement each other (§11).
- **Forks stay clean.** Branch-local memory means a `fork` sub-agent's noise does
not leak into the parent's folded memory.
## 6. `recall` is not a search tool
`recall` takes **one specific 12-hex id** that already appears in compacted
memory or in `/om:view`. It cannot be given a topic. It can return an observation
(marked `active` or `dropped`), or a reflection together with the observations
supporting it.
```mermaid
sequenceDiagram
participant M as compacted memory
participant A as agent
participant L as ledger
M->>A: "[high] user rejected option B (a1b2c3d4e5f6)"
A->>L: recall("a1b2c3d4e5f6")
L-->>A: exact observation + source ids
Note over A: acts on the original wording
```
The rule of thumb the agent skill uses: recall **before a load-bearing action**
that rests on a compressed memory — shipping a change, asserting a fact,
answering "why do you believe that". One recall is cheap; redoing finished work
is not.
## 7. How it is wired in this image
```mermaid
flowchart TB
IMG["baked in the image:<br/>v3.0.4 @ ce9fc98"] --> REG["settings.json<br/>packages[]"]
REG --> SESS["your pi session"]
SESS -- "your turns" --> SM["session model:<br/>Opus"]
SESS -- "observer, reflector,<br/>dropper" --> WM["memory model:<br/>Haiku"]
SESS -- "ledger entries" --> JL["session .jsonl"]
JL --> VOL[("devbox-pi-config<br/>volume")]
```
Four consequences of that wiring:
1. **There is no separate database.** Memory *is* entries inside the ordinary pi
session file (`~/.pi/agent/sessions/<project>/<timestamp>_<uuid>.jsonl`).
Nothing extra to back up, nothing to migrate.
2. **It survives container recreate**, because `~/.pi` is the `devbox-pi-config`
named volume (`docker-compose.yml`) — the same one holding your pi config and
session history.
3. **`packages[]` is the only source of truth for which copy is loaded.** A clone
at `/workspace/pi-observational-memory` may exist (and today matches `/opt`
byte-for-byte at `ce9fc98`) — its presence proves nothing. To run a patched
build you point `packages[]` at it explicitly and start a new session.
4. **The worker model is a deliberate choice, and it is yours to change.** The
seeded config sends background work to Haiku while your session runs Opus:
```json
"observational-memory": {
"model": { "provider": "amazon-bedrock", "id": "eu.anthropic.claude-haiku-4-5-20251001-v1:0" },
"debugLog": false
}
```
## 8. What it costs
| Resource | Cost |
|---|---|
| Model calls | up to **three** background calls per consolidation pass (observer, reflector, dropper), each capped at `agentMaxTurns` [16], on the configured memory model — not your session model |
| Latency in your turns | none by construction: workers run from `turn_end`, compaction runs when pi is idle, and the fold itself does no model work |
| Disk | negligible — JSON lines inside a session file that would exist anyway (measured here: `~/.pi/agent/sessions` = 30 MB total, tens of `om.*` entries per session) |
| Context window | **zero until compaction.** `custom` entries do not enter LLM context; only the folded summary does |
| Attention | none once configured; there is no protocol for you or the agent to remember |
If that is still more than you want on a given run, §9's `passive` switch turns
off all proactive work while keeping `recall` and `/om:*` usable.
## 9. Configuration
Global: `~/.pi/agent/settings.json` (persisted in the volume). Per project:
`<project>/.pi/settings.json`, which overrides global. Precedence is
project → global → environment, and the environment can only override `passive`.
| Key | Default | What it changes |
|---|---|---|
| `observeAfterTokens` | `10000` | observer cadence — lower means smaller chunks and more calls |
| `reflectAfterTokens` | `20000` | reflector cadence (and thereby dropper opportunities) |
| `observerChunkMaxTokens` | 20% of the memory model's context window, else `60000` | cap on one observer run's input |
| `compactAfterTokens` | `81000` | when proactive auto-compaction fires |
| `observationsPoolMaxTokens` | `20000` | pool size at which compaction does a full re-fold from the branch root |
| `observationsPoolTargetTokens` | half of max (`10000`) | what the dropper aims back down to |
| `agentMaxTurns` | `16` | shared turn cap for the three workers |
| `model` | unset → session model | send background work to a cheaper/faster model |
| `showWorkerNotifications` | `true` | routine "observer ran" notices |
| `passive` | `false` | **kill switch** for all proactive background work; `recall` and `/om:*` still work |
| `debugLog` | `false` | per-session NDJSON trace at `~/.pi/agent/observational-memory/debug/<session-id>.ndjson` |
Pi's own compaction knobs live under a separate `compaction` key —
`keepRecentTokens` [20000] sets the verbatim tail from §4, `reserveTokens`
[16384] the headroom that triggers pi's own compaction.
One-off passive run, no config edit:
```bash
PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi
```
Invalid values are ignored rather than fatal, so a typo degrades to the default
instead of breaking your session — which also means a typo is silent. Check with
`/om:status`.
## 10. Confirming it is actually working
Do not infer health from the absence of a warning; look:
```bash
# 1. inside pi — the authoritative view
/om:status # visible-vs-full drift, thresholds, worker state
/om:view # what the agent currently sees
/om:view full # full ledger truth at the branch tip
# 2. from a shell — are ledger entries being written, and has it compacted?
grep -o '"customType":"om\.[a-z.]*"' \
"$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)" | sort | uniq -c
grep -c '"type":"compaction"' "$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)"
# 3. which copy is loaded, and at what commit
python3 -c "import json;print(json.load(open('$HOME/.pi/agent/settings.json'))['packages'])"
git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory rev-parse HEAD
```
Ledger entries are `"type":"custom"` with `"customType":"om.…"`. Do not grep for
`custom_message` — that is a *different* pi API for entries that **do** enter LLM
context, used here by the MemPalace mailbox (`customType: "mempalace-mailbox"`),
not by om.
## 11. It is not the same thing as MemPalace
Both are called "memory" and they solve different problems. Nothing is wrong with
running both — this image does, and they cover each other's failure modes.
```mermaid
flowchart LR
O0["observational<br/>memory"] --> O1["horizon:<br/>this session"]
O1 --> O2["scope: one branch,<br/>one machine"]
O2 --> O3["automatic"]
O3 --> O4["retrieval:<br/>recall(id)"]
P0["MemPalace"] --> P1["horizon: months,<br/>machines"]
P1 --> P2["scope:<br/>the fleet"]
P2 --> P3["protocol-driven"]
P3 --> P4["retrieval:<br/>search, KG, mailbox"]
```
| Question | Answer |
|---|---|
| "What did we decide 200 turns ago in *this* session?" | observational memory (and `recall` for the exact wording) |
| "What did we decide last month, or on another machine?" | MemPalace (`mempalace_search`, diaries) |
| "What is true *right now* about version X?" | MemPalace knowledge graph |
| "Does another machine need something from me?" | MemPalace coordination log — see [Cross-machine agent coordination](../README.md#cross-machine-agent-coordination) |
| "Why is compaction not losing my session?" | observational memory |
The crisp version: **observational memory keeps a session coherent; the palace
keeps the fleet coherent.** A container recreate wipes neither — but only because
`~/.pi` and the palace both live outside the container filesystem.
## 12. Gotchas
- **Branch-local means branch-local.** Resuming or forking changes which ledger
is folded. Memory that "disappeared" is usually on another branch.
- **`recall` needs an id, not a topic.** If you only have a topic, that is a
palace search, not a recall.
- **A `/workspace` clone is not evidence of what is loaded** — see §7.3.
- **`showWorkerNotifications: true` is not proof of work**; it reports runs, and
an observer that deliberately emits nothing writes no ledger entry and simply
retries after another `observeAfterTokens`.
- **A turn bigger than `keepRecentTokens` splits.** The cut then lands mid-turn at
an assistant message and pi merges two summaries — rare, but it is why a very
large single turn can lose more verbatim detail than you would expect.
- **`git log` in the baked tree needs `safe.directory`** (`/opt` is root-owned):
`git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory log`.
+105
View File
@@ -188,6 +188,111 @@ if [ "${MEMPALACE_FEED:-1}" != "0" ] && [ -n "$MEMPALACE_FEEDER" ]; then
fi
fi
# ── cli_utils: link workspace bin/ commands onto PATH ────────────────
# Standalone commands from a mounted cli_utils checkout (git-status-all,
# git-pull-all, devbox-sanity, pi-session-repair, ...) live in <repo>/bin. On a
# host they reach PATH via cli_utils' own install.sh, whose install_bin step
# symlinks them into ~/.local/bin — but that home is on the container's WRITABLE
# LAYER, so every recreate loses them and the human is back to typing
# /workspace/cli_utils/bin/git-status-all. This is the container equivalent of
# that install step, re-run at every start.
#
# WHY SYMLINKS RATHER THAN A PATH EDIT IN AN rc FILE: ~/.local/bin is already
# ahead of /usr/local/bin in ENV PATH (Dockerfile.base), so links here resolve in
# NON-interactive shells too — `docker exec <c> git-status-all`, agent tool
# shells, scripts. An rc-file PATH edit cannot reach those, because ~/.bashrc
# returns early when the shell is not interactive. Measured 2026-08-27 on
# tor-ms22: `command -v git-status-all` failed in a non-interactive shell while
# working in an interactive one, from exactly that asymmetry.
#
# Detection order (first hit wins):
# 1. CLI_UTILS_CONTAINER_PATH explicit, for non-standard layouts
# 2. /workspace/cli_utils repo directly in the workspace root
# 3. $HOME/cli_utils dedicated mount
# 4. /workspace/*/cli_utils workspace root holds several repo groups
# CLI_UTILS_LINK=0 disables. Absent repo = silent no-op, which is the common
# case for anyone who does not use cli_utils.
if [ "${CLI_UTILS_LINK:-1}" != "0" ]; then
CLI_UTILS_BIN=""
if [ -n "${CLI_UTILS_CONTAINER_PATH:-}" ] && [ -d "${CLI_UTILS_CONTAINER_PATH}/bin" ]; then
CLI_UTILS_BIN="${CLI_UTILS_CONTAINER_PATH}/bin"
elif [ -d /workspace/cli_utils/bin ]; then
CLI_UTILS_BIN=/workspace/cli_utils/bin
elif [ -d "$HOME/cli_utils/bin" ]; then
CLI_UTILS_BIN="$HOME/cli_utils/bin"
else
# `if` bodies, not `&&` chains: under `set -e` a loop whose LAST command is a
# false test exits non-zero and would abort the entrypoint. With no match the
# glob stays literal, so that is the normal case on any machine without this
# repo — i.e. the bug would have been "container will not start", not "links
# missing".
for _cu in /workspace/*/cli_utils/bin; do
if [ -d "$_cu" ]; then
CLI_UTILS_BIN="$_cu"
break
fi
done
unset _cu
fi
if [ -n "$CLI_UTILS_BIN" ]; then
mkdir -p "$HOME/.local/bin" 2>/dev/null || true
# Never clobber a real file, and never steal a link that points elsewhere: a
# deliberate user override in ~/.local/bin must win, and silently shadowing
# an image-provided command is worse than the missing command.
for _f in "$CLI_UTILS_BIN"/*; do
if [ ! -f "$_f" ] || [ ! -x "$_f" ]; then
continue
fi
_link="$HOME/.local/bin/$(basename "$_f")"
if [ -e "$_link" ] && [ ! -L "$_link" ]; then
continue
fi
if [ -L "$_link" ]; then
case "$(readlink "$_link")" in
"$CLI_UTILS_BIN"/*) ;;
*) continue ;;
esac
fi
ln -sf "$_f" "$_link" 2>/dev/null || true
done
# Prune links we own whose target vanished (command renamed, repo moved),
# mirroring the skillset deploy's --prune-stale. A dangling link on PATH
# reports "No such file or directory" for a command that simply no longer
# exists, which reads as a broken container rather than a removed script.
for _link in "$HOME/.local/bin"/*; do
[ -L "$_link" ] || continue
case "$(readlink "$_link")" in
*/cli_utils/bin/*) [ -e "$_link" ] || rm -f "$_link" ;;
esac
done
unset _f _link
fi
unset CLI_UTILS_BIN
fi
# ── Per-device boot hook ─────────────────────────────────────────────
# Runs ~/.config/devbox-shell/init.sh if the host provides one. That directory is
# the host-owned, bind-mounted shell-sharing dir (see "Volumes and persistence"),
# so a hook placed there survives every recreate WITHOUT an image change — the
# boot-time twin of the interactive bridge in /etc/skel-devbox/.bash_aliases,
# which sources ~/.config/devbox-shell/bash_aliases for every interactive shell.
#
# NO NEW TRUST BOUNDARY: that same directory is already sourced into every
# interactive shell, i.e. it is already arbitrary code from the same owner. What
# is new is only WHEN it runs — once at start, before any shell — which is what
# non-interactive fixups (symlinks, dirs, one-off migrations) need.
#
# Deliberately `bash <file>`, not `.` — a hook must not be able to mutate this
# entrypoint's own shell state, and its exit status must not matter. Output goes
# to a log rather than the container's start output, so a chatty hook cannot
# masquerade as a startup error.
if [ -r "$HOME/.config/devbox-shell/init.sh" ]; then
mkdir -p "$HOME/.pi/agent" 2>/dev/null || true
bash "$HOME/.config/devbox-shell/init.sh" \
>"$HOME/.pi/agent/devbox-init.log" 2>&1 || true
fi
# ── Git config defaults ──────────────────────────────────────────────
if [ -n "${GIT_USER_NAME:-}" ] && ! git config --global user.name &>/dev/null; then
git config --global user.name "$GIT_USER_NAME"
@@ -428,10 +428,56 @@ An obligation you never agreed to is noise, so the sender states it:
| `to_agent="*"` (any status) | broadcast FYI | nothing |
| any other status (`ready`, `applied`, `blocked`, …) | a statement of fact | nothing |
**That table says what you *owe*. Delivery is stricter, and the difference bites:
the mailbox is an obligation channel, not a news channel.** Mailbox candidates are
drawn with `status="open"`, so an event carrying any **terminal** status
(`applied`, `superseded`, `failed`, `blocked`) is never a candidate — *whoever it
is addressed to*. A `task.reply` written to a named machine to share a finding is
delivered to nobody, ever, and neither is any `event_ack`. It sits in the log
until somebody reads the log.
So the most natural inter-machine message — *"here is something you should
know"* — is exactly the shape that gets no delivery. Pick deliberately:
| You want the peer to… | Write |
|---|---|
| **do something**, and you need it tracked until done | directed `status="open"` ask, with a `correlation_id` |
| **know something**, no response needed | terminal-status event **plus a drawer** — the drawer is what actually reaches them, via search |
What does **not** work is a terminal report plus an expectation of attention.
Measured 2026-08-26: a detailed report addressed to `pi@<peer>` with
`status="applied"` went unread for two and a half hours until the operator quoted
the event id by hand, with the mailbox working correctly the whole time. Full
mechanism in the toolkit's `docs/rfc-003-coordination-log.md` §7.12.
One more timing fact, because it looks like negligence and is not: a delivered
ask is queued into the agent's **next turn** (`deliverAs: "steer"`, deliberately
no `triggerTurn`), and the poll fires when the agent is *idle*. Between delivery
and the next turn no inference runs, so **a human starting a turn is the
trigger** (§7.11). An agent that "has not reacted" has usually not been running.
Ack with `mempalace_event_ack(event_id=…, from_agent="<you>", status=…)`. It
**appends a new event** and never mutates the original; the correlation id is
copied for you, and `metadata.ack_of` is set to the event you answered.
**Claiming, and what it does not do.** `status="claimed"` announces that you have
picked work up. Nothing requires it — a directed open ask owes "an ack *or* a
reply", and finishing the work is a complete answer. Do it anyway when the work is
long or the machine is unreliable, because it is the only thing that later
distinguishes *nobody started this* from *someone started and their container
died mid-task*. Be clear about its limits, both of which follow from candidacy
requiring exactly `status="open"`:
- **It does not notify the requester.** `claimed` is not `open`, so a claim is no
more deliverable than a finished report is (see the delivery table above). Its
reader is whoever pulls the log.
- **It does not quiet your own mailbox.** The ask stays owed until a *terminal*
event of yours joins it, so a claimed-then-silent thread keeps resurfacing —
correctly.
Prefer a prompt terminal reply over a claim plus a long silence; claim *in
addition*, when the gap between pickup and finish is where a machine might die.
#### What you actually owe — derive it, do not read it off `status`
The log is append-only and `status` is written **once**, so it is an honest
@@ -502,6 +548,17 @@ Two consequences worth internalising:
- **Address the stamped name you actually saw** in a `from_agent` field, e.g.
`pi@tor-ms22`. A bare `pi` reaches nobody's mailbox once stamping is live, and
older events in the log still carry bare names — do not copy them.
- **The rule runs in reverse too: what you put in YOUR OWN `from_agent` decides
where every reply to your event goes.** Nothing stops you writing a synthetic
or borrowed identity there, and a reply is always addressed back to exactly
that string — so if no live session ever runs as it, the reply is stored,
searchable, and delivered to no one. Measured cost: a directed ask sent under
a synthetic sender got two correct replies, one of them an urgent security
finding, and both sat unread for ~2h20m because nobody's mailbox was that
identity (RFC 003 §7.13). Authoring under a synthetic name is fine for a
deliberate control experiment — this fleet does it on purpose — but then
**name the real identity to reply to inside the body**, because the address
line is not a safe place to also carry provenance.
- **Use `status="open"` only when you truly need an answer.** It places an
obligation on another machine.
- **Never broadcast an ask.** `to_agent="*"` + `status="open"` obliges everyone
@@ -513,7 +570,8 @@ Two consequences worth internalising:
be matched to it at all.
- **Corrections are new events, never edits.** Say explicitly what you retract
and name the id — drawer or event — that carried the withdrawn claim.
- **Put a retraction where the reader will look.** An event reaches a live agent;
- **Put a retraction where the reader will look.** A *directed open ask* reaches a
live agent's mailbox; a **terminal-status event reaches no mailbox at all**, and
a *drawer* is what a future semantic search finds. If you filed advice as a
drawer and later withdraw it, file the withdrawal as a drawer too — otherwise
the next agent finds your original confident advice and no trace of the
@@ -600,4 +658,5 @@ Entity-relationship triples with temporal validity. Query with `mempalace_kg_que
- **Don't treat the palace as a task list.** It's for knowledge and context, not todos.
- **Don't broadcast an ask, and don't leave one unanswered.** On a shared palace, `to_agent="*"` + `status="open"` obliges every machine and therefore none of them. And don't expect acking to tidy your mailbox: `status` is immutable, so the event keeps matching either way — what a terminal reply buys you is that the *derived* owed set (see *What you actually owe*) stops counting it. Leave asks unanswered and that set only grows, until everyone learns to stop looking. "Seen, not doing it" is a complete answer — silence is not.
- **Don't assume you would have heard.** Nothing pushes another machine's message into your session. If you did not run the mailbox query at wake-up, a correction addressed to you by name can sit unread while you confidently rebuild the thing it warned you about.
- **Don't author an ask under an identity nobody runs as, including your own throwaway labels.** The failure is symmetric to the one above: it is not that you missed a message, it is that nothing could ever have delivered the reply to you, because you addressed it at a name instead of an agent. If you must use a synthetic sender for a control or an experiment, say inside the body who should actually receive the reply.
- **Don't invent provenance metadata, and don't hand-stamp it either.** An earlier version of this list told you to set `added_by="<harness>@<device>"` by hand; that instruction has been withdrawn, because RFC 001 §7.3.2 places provenance at the client/server boundary and the pi bridge now does it uniformly (see *Provenance is stamped for you* above) — but the withdrawal only holds where the bridge is live, so run the one-line check in that bullet first; on an older image hand-stamping is still the only signal a hand-filed drawer gets. DO NOT invent values for the palace's own metadata fields (`device`, `agent_kind`, `origin_device`): those are stamped by infrastructure that also records *how* each was determined, and a fabricated value is worse than none because it silently corrupts a future merge. DO pass `source_drawer_id` on `kg_add`. And never put a machine name in a diary's `agent_name` — it becomes the wing name and hides your entries from `diary_read`.