docs: explain the memory that runs itself, and fix a claim the mailbox falsified
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 16s

`pi-observational-memory` is baked, registered in the seeded settings.json, and
handed a cheaper model than the session it serves — and the only description in
this repo was five words in a feature list. Someone meeting `/om:status` or a
"compacted memory" block had nothing to read that said whether to leave any of it
on. New docs/observational-memory.md, 272 lines and five diagrams.

Scoped by who is authoritative, so there is one copy of each claim:

- Upstream (/opt/pi-observational-memory/docs/) already documents the mechanism
  well — concepts.md, how-it-works.md, configuration.md, including a v3 lifecycle
  diagram that matches the deployed code. Linked, not re-derived.
- This document takes the four facts pi-devbox owns and can change: the pinned
  commit it bakes (v3.0.4 ce9fc98, matching build-manifest.json), the packages[]
  entry that decides which copy loads, the Haiku-workers-vs-Opus-session split,
  and the devbox-pi-config volume that makes the ledger outlive the container.
- Plus the confusion this image creates by shipping two things called memory: a
  section contrasting it with MemPalace, on the line "observational memory keeps
  a session coherent, the palace keeps the fleet coherent".

Placement follows the split fleet-ops states for itself — reusable mechanism is
not deployment data — so a "why is this in my container" document belongs in the
repo that pins and wires the component. Linked twice from the README, because
until this commit the README referenced docs/ zero times and the file already
sitting there was reachable only by listing the directory.

Numbers were read out of the live container and the baked tree rather than out of
release notes, which caught one thing the pi-extensions skill still has wrong:
the dropper is gated on a successful same-turn reflection, not on a token
threshold of its own.

Also fixes a README sentence that v1.8.9 made false. § Cross-machine agent
coordination ended with "Nothing in this image polls the log on the agent's
behalf"; the mailbox has shipped since aac4a1c. Replaced with the three knobs and
their defaults, derived-not-read owed-ness, and the queued-into-the-next-turn
delivery measured on two devices — and dated to "as baked in v1.8.9
(mempalace-toolkit 5b8d78f)", pointing at RFC 003 §7.11–§7.12 for the mechanism,
because toolkit main is already ahead (a92c75d pings the human who is not
looking) and describing that here would trade a stale-behind claim for a
stale-ahead one.

The cause is worth more than the fix: the behaviour arrived through the floating
MEMPALACE_TOOLKIT_REF, so no diff in this repo ever touched the paragraph making
the claim. v1.8.9's own rule fired for the CHANGELOG and nobody swept the README.
The CHANGELOG records what changed, the README asserts what is true, and only the
first is reviewed at release time — so the rule now extends to grepping the README
for absolute claims (nothing, never, does not, only) about a component whose SHA
moved.

Diagrams were verified by rendering, not by parsing. Both comparison diagrams
parsed clean and rendered with their meaning reversed: Mermaid laid the second
declared subgraph out first, putting "with observational memory" before "without"
and MemPalace before observational memory in the diagram whose entire job was
that contrast. Rebuilt as declaration-ordered chains and re-rendered at mermaid@11
— the version pi-studio pins — in the baked headless browser, then read back as an
image. mermaid.parse() proves syntax and says nothing about layout.
This commit is contained in:
Joakim Persson
2026-08-27 13:55:55 +02:00
parent aac4a1c323
commit 14371e2da6
3 changed files with 396 additions and 2 deletions
+77
View File
@@ -11,6 +11,83 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
---
## Unreleased
### The other memory system finally gets explained — `docs/observational-memory.md`
`pi-observational-memory` has been baked for several releases and described in
one line of the feature list (*"the `recall` tool for session compaction"*),
which is enough to name it and not nearly enough to use it. New 272-line
explainer with five diagrams, aimed at someone who has seen `/om:status` or a
"compacted memory" block and wondered whether to leave any of it switched on.
**Scoped to what this repo is authoritative for, because upstream already
documents the mechanism well.** `/opt/pi-observational-memory/docs/` ships
`concepts.md`, `how-it-works.md` and `configuration.md`, including a correct v3
lifecycle diagram — so the new document links those for depth and spends its own
words on the four facts pi-devbox owns and can change: the pinned commit it bakes
(v3.0.4 `ce9fc98`, the value in `build-manifest.json`), the `packages[]` entry
that decides which copy loads, the Haiku-workers-vs-Opus-session split seeded
into `~/.pi/agent/settings.json`, and the `devbox-pi-config` volume that makes
the ledger survive `--force-recreate`. Plus the confusion this image creates by
shipping two things called memory: a section contrasting it with MemPalace, on
the line *observational memory keeps a session coherent, the palace keeps the
fleet coherent*.
Every stated number was read out of the live container or the baked tree rather
than copied from release notes — including the correction that the dropper is
gated on a **successful same-turn reflection** and not on a token threshold of
its own, which is the one detail `pi-extensions/SKILL.md` still gets wrong.
Placement follows the audience split fleet-ops states for itself: reusable
mechanism is not deployment data, so a "why is this in my container" document
belongs in the repo that **pins and wires** the component, pointing upstream for
depth. Linked twice from the README, because before this commit the README
referenced `docs/` zero times and the one file already there
(`mempalace-broker-design.md`) was reachable only by listing the directory.
### A README claim that v1.8.9 made false, and how it got there
**§ Cross-machine agent coordination ended with "Nothing in this image polls the
log on the agent's behalf." The mailbox shipped in v1.8.9, so that sentence has
been wrong since `aac4a1c`.** Replaced with the three knobs and their defaults
(`MEMPALACE_MAILBOX`, `MEMPALACE_MAILBOX_POLL_MS` 300000,
`MEMPALACE_MAILBOX_RESURFACE_MS` 3600000), the fact that owed-ness is *derived*
rather than read off `status`, and the queued-into-the-next-turn delivery
semantics measured on two devices.
The cause is the one v1.8.9 wrote a rule about: the behaviour arrived through the
floating `MEMPALACE_TOOLKIT_REF`, so **no diff in this repo ever touched the
paragraph that made the claim**. v1.8.9's rule ("name a floating-ref behaviour
change in the CHANGELOG before tagging") fired for the CHANGELOG and nobody swept
the README. The CHANGELOG records what *changed*; the README asserts what is
*true*, and only the first is reviewed at release time. Extending the rule
accordingly: grep the README for absolute claims — *nothing*, *never*, *does
not*, *only* — about any component whose SHA moved.
**The replacement is dated on purpose.** It says it describes the bridge *as baked
in v1.8.9* (`mempalace-toolkit` `5b8d78f`) and points at that repo's
`docs/rfc-003-coordination-log.md` §7.11–§7.12 for the mechanism, because toolkit
main is already ahead of the baked copy (`a92c75d` makes delivery say it is queued
and ping the human who is not looking; `e917662` and `ecc2a9c` refine that notify
path) and none of it reaches a container until a base rebuild. Documenting those
here would have swapped a stale-behind claim for a stale-ahead one — the same
defect with the sign flipped.
### Diagrams verified by rendering, not by parsing
Both comparison diagrams **parsed clean and rendered with their meaning
reversed**: Mermaid laid the second declared `subgraph` out first, so "with
observational memory" appeared before "without", and MemPalace before
observational memory in the diagram whose entire job was that contrast. A third
was legible only at 1280px. Rebuilt as declaration-ordered node chains, then
re-rendered at mermaid@11 — the version `pi-studio` pins — in the baked headless
browser and read back as an image. Recorded because it generalises:
`mermaid.parse()` proves syntax and says nothing about layout, so a diagram is
unverified until someone has looked at it.
---
## v1.8.9 — 2026-08-26
The coordination log gets a reader, and the release checklist's last gate stops
+47 -2
View File
@@ -20,7 +20,9 @@ on the host.
- `pi-extensions` — TypeScript extensions for pi (preview, MCP bridges,
mempalace integration, etc.)
- `pi-fork` — the `fork` tool for spawning sub-agents
- `pi-observational-memory` — the `recall` tool for session compaction
- `pi-observational-memory` — durable session memory: the ledger that makes
compaction cheap, plus the `recall` tool. See
[`docs/observational-memory.md`](docs/observational-memory.md)
- `pi-atelier` — TUI sidebar: ordered panels, split-pane, themes. Pinned to an
audited tag; see [Version pins](#version-pins-pi-pi-atelier-mempalace)
@@ -579,7 +581,50 @@ convention that a directed event with `status="open"` is a request owed a reply
while a `*` broadcast owes nothing. The mechanism side (what the bridge stamps,
and why live SSE push depends on the palace deployment's reverse proxy rather
than on this image) is documented in the toolkit's `extensions/pi/README.md`.
Nothing in this image polls the log on the agent's behalf.
**Since v1.8.9 the bridge reads the log for you.** Earlier images were write-only
— they stamped provenance on the way out and never read back, so a directed ask
reached an agent only if that agent happened to run `mempalace_event_list`
itself. The mailbox is gated on the same two variables as the stamper, is on by
default, and derives what is *owed* rather than trusting `status` (an acked event
keeps matching a `status="open"` query forever, because the log is append-only):
| Variable | Default | Effect |
|---|---|---|
| `MEMPALACE_MAILBOX` | unset (on) | `0` disables mailbox reads entirely |
| `MEMPALACE_MAILBOX_POLL_MS` | `300000` | minimum gap between mid-session polls |
| `MEMPALACE_MAILBOX_RESURFACE_MS` | `3600000` | re-announce a still-owed ask after this long |
Delivery **queues, it never interrupts**: the poll runs when pi goes idle and the
message is steered into the *next* turn, so nothing wakes the model on inbound
fleet traffic. The practical consequence, measured on two devices: the message
appears in your session window and the agent acts on it when the next turn
starts — you are the trigger. (That describes the bridge **as baked in v1.8.9**,
`mempalace-toolkit` `5b8d78f`; the mailbox's own mechanism and landmines live in
the toolkit's `docs/rfc-003-coordination-log.md` §7.11–§7.12, which moves ahead of
whatever this image has baked.)
## Observational memory (in-session memory)
The image also bakes [pi-observational-memory](https://github.com/elpapi42/pi-observational-memory),
which is memory of a *different kind* from the palace and is easy to confuse with
it. It keeps a small branch-local ledger of observations and reflections while a
session runs, so when pi compacts the conversation the summary is a
**deterministic fold of that ledger rather than a model call**, and every item
keeps a 12-character id that `recall(<id>)` resolves back to the exact source.
In one line: **observational memory keeps a session coherent; the palace keeps
the fleet coherent.**
It is on by default, needs no habit from you, and sends its background work to a
cheaper model than your session (Haiku while the session runs Opus, in the seeded
`~/.pi/agent/settings.json`). Inspect it from inside pi with `/om:status` and
`/om:view`; turn all proactive work off for one run with
`PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi`.
What it is for, how the lifecycle works, what it costs, every setting and its
default, and how it differs from MemPalace:
[`docs/observational-memory.md`](docs/observational-memory.md).
## Agent skills
+272
View File
@@ -0,0 +1,272 @@
# Observational memory — why this image has it, and what it does for you
**Audience:** anyone using this container for long pi sessions who has wondered
what `recall`, `/om:status` and "compacted memory" are, or whether they should
leave any of it switched on.
**Companion documents:** the extension ships its own reference docs at
`/opt/pi-observational-memory/docs/` —
[`concepts.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/concepts.md)
(the model),
[`how-it-works.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/how-it-works.md)
(hooks and internals) and
[`configuration.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/configuration.md)
(every setting). Those are normative; this document is the **deployment** view —
what is pinned here, how it is wired, what it costs, and how it differs from
MemPalace. For the palace, see
[`mempalace-toolkit/docs/fleet-memory.md`](https://gitea.jordbo.se/joakimp/mempalace-toolkit/src/branch/main/docs/fleet-memory.md).
> Verified on pi-devbox **v1.8.9** (`release_tag v1.8.9`, source `aac4a1c`),
> which bakes pi-observational-memory **v3.0.4** at commit `ce9fc98` — the value
> in `/etc/pi-devbox/build-manifest.json` → `components.pi-observational-memory`.
> Every number below was read from that tree or from the live container, not
> from release notes.
---
## 1. The problem it solves
A long pi session outgrows the model's context window. Pi's answer is
**compaction**: fold the older part of the conversation into a summary and keep
recent messages verbatim. That is unavoidable, and it is where sessions go
wrong — the summary is produced *at the moment of pressure*, by a model, about a
transcript that is about to be discarded.
Observational memory changes *when* the remembering happens. Instead of
summarising in a panic at the end, it keeps a small **ledger** up to date while
the session runs, and compaction then just folds that ledger.
```mermaid
flowchart LR
A0["<b>plain compaction</b>"] --> A1["long session,<br/>context filling"]
A1 --> A2["a model summarises<br/>the transcript<br/><i>at the moment of pressure</i>"]
A2 --> A3["one prose summary:<br/>detail chosen in a hurry,<br/>no route back to the original"]
B0["<b>with observational memory</b>"] --> B1["long session,<br/>context filling"]
B1 --> B2["<i>while you work:</i><br/>a cheap background model appends<br/>observations + reflections<br/>to a ledger"]
B2 --> B3["compaction folds the ledger<br/><b>no model call</b>"]
B3 --> B4["every item keeps a 12-hex id —<br/><code>recall(id)</code> returns<br/>the exact source"]
```
## 2. The mental model: three layers and a ledger
| Layer | What it is | Example |
|---|---|---|
| **Observation** | a timestamped, source-backed event from the conversation | "user rejected option B because it needs a base rebuild" |
| **Reflection** | a durable conclusion *backed by* observations | "the user optimises for avoiding 67-minute rebuilds" |
| **Drop** | a tombstone retiring an observation from active memory | the superseded detail of a bug that is now fixed |
These are appended as silent ledger entries (`om.observations.recorded`,
`om.reflections.recorded`, `om.observations.dropped`) and **folded** — replayed
from the branch root — to produce the memory state. The ledger is the source of
truth; the compaction summary is a rendering of it.
Memory is **branch-local**. A pi session is a tree (resume, fork), and the fold
follows the current branch only, so a forked branch does not inherit another
branch's view.
## 3. The lifecycle
Three background workers and one compaction hook, driven by *raw token
progress* rather than wall-clock time. Defaults in brackets.
```mermaid
flowchart TD
T(["turn_end"]) --> O{"tokens since last<br/>observation coverage<br/>≥ observeAfterTokens<br/>[10000]?"}
O -- yes --> OBS["<b>observer</b> runs:<br/>appends om.observations.recorded"]
O -- "no (observer not due)" --> R{"tokens since last<br/>reflection<br/>≥ reflectAfterTokens<br/>[20000]?"}
R -- yes --> REF["<b>reflector</b> runs:<br/>appends om.reflections.recorded"]
REF -- "non-empty reflection AND active pool<br/>&gt; observationsPoolTargetTokens [10000]" --> DR["<b>dropper</b> runs:<br/>appends om.observations.dropped"]
S(["agent_settled"]) --> C{"tokens since last compaction<br/>≥ compactAfterTokens [81000]<br/>and pi is idle?"}
C -- yes --> CP["extension calls ctx.compact()"]
CP --> H(["session_before_compact"])
H --> F["<b>deterministic fold</b> of the ledger:<br/>no model call, no waiting on workers"]
F --> VIS["compacted memory the agent reads"]
```
Two details worth carrying, because they are easy to get wrong:
- **The dropper has no clock of its own.** It is post-reflection maintenance,
gated on a *successful same-turn* reflection — not a third worker on a third
threshold.
- **Compaction itself calls no model.** That is the whole design: the expensive
thinking happened earlier, in the background, on a cheap model.
## 4. What you actually get
- **Compaction stops being a stall.** The latency path is deterministic work on
ledger entries, not a summarisation call.
- **Nothing important vanishes silently.** Compaction is lossy by design, but
every observation and reflection keeps a 12-character id, and `recall(<id>)`
returns the exact evidence behind it — the original wording, the reasoning,
the file path, the error text.
- **The bookkeeping runs on a cheaper model than your session.** In this image
that is deliberate and visible (§6): background workers on Haiku, the session
on Opus.
- **It is automatic.** There is no habit to maintain, unlike the palace protocol
— which is exactly why the two systems complement each other (§10).
- **Forks stay clean.** Branch-local memory means a `fork` sub-agent's noise does
not leak into the parent's folded memory.
## 5. `recall` is not a search tool
`recall` takes **one specific 12-hex id** that already appears in compacted
memory or in `/om:view`, and looks it up in the ledger on the current branch. It
cannot be given a topic. It can return an observation (marked `active` or
`dropped`), or a reflection together with the observations supporting it.
```mermaid
sequenceDiagram
participant M as compacted memory
participant A as agent
participant L as ledger (current branch)
M->>A: "[high] user rejected option B (a1b2c3d4e5f6)"
A->>L: recall("a1b2c3d4e5f6")
L-->>A: exact observation + its source entry ids
Note over A: acts on the original wording,<br/>not on the compressed paraphrase
```
The rule of thumb the agent skill uses: recall **before a load-bearing action**
that rests on a compressed memory — shipping a change, asserting a fact,
answering "why do you believe that". One recall is cheap; redoing finished work
is not.
## 6. How it is wired in this image
```mermaid
flowchart TB
IMG["<b>image layer</b><br/>/opt/pi-observational-memory<br/>v3.0.4 @ ce9fc98 (pinned)"] --> REG["<b>~/.pi/agent/settings.json</b><br/>packages[]: ../../../../opt/pi-observational-memory<br/><i>the only source of truth for which copy loads</i>"]
REG --> SESS["<b>your pi session</b>"]
SESS -->|"your turns"| SM["session model<br/>claude-opus-5"]
SESS -->|"observer / reflector / dropper"| WM["memory model<br/>claude-haiku-4-5 (Bedrock)"]
SESS -->|"ledger entries (custom_message)"| JL["~/.pi/agent/sessions/&lt;project&gt;/&lt;timestamp&gt;_&lt;uuid&gt;.jsonl"]
JL --> VOL[("devbox-pi-config volume — docker-compose.yml:59<br/>survives --force-recreate")]
```
Four consequences of that wiring:
1. **There is no separate database.** Memory *is* entries inside the ordinary pi
session `.jsonl`. Nothing extra to back up, and nothing to migrate.
2. **It survives container recreate**, because `~/.pi` is the
`devbox-pi-config` named volume — the same one that keeps your pi config and
sessions.
3. **`packages[]` is the only source of truth for which copy is loaded.** A
clone at `/workspace/pi-observational-memory` may exist (and today matches
`/opt` byte-for-byte at `ce9fc98`) — its presence proves nothing. If you ever
run a patched build, you point `packages[]` at it explicitly and start a new
session.
4. **The worker model is a deliberate choice, and it is yours to change.** The
seeded config sends background work to Haiku while your session runs Opus:
```json
"observational-memory": {
"model": { "provider": "amazon-bedrock", "id": "eu.anthropic.claude-haiku-4-5-20251001-v1:0" },
"debugLog": false
}
```
## 7. What it costs
| Resource | Cost |
|---|---|
| Model calls | up to **three** background calls per consolidation pass (observer, reflector, dropper), each capped at `agentMaxTurns` [16], on the configured memory model — not your session model |
| Latency in your turns | none by construction: workers run from `turn_end`, and the compaction hook does no model work |
| Disk | negligible — JSON lines inside a session file that would exist anyway (measured here: `~/.pi/agent/sessions` = 30 MB total, tens of `om.*` entries per session) |
| Attention | none once configured; there is no protocol for you or the agent to remember |
If that is still more than you want on a given run, §8's `passive` switch turns
off all proactive work while keeping `recall` and `/om:*` usable.
## 8. Configuration
Global: `~/.pi/agent/settings.json` (persisted in the volume). Per project:
`<project>/.pi/settings.json`, which overrides global. Precedence is
project → global → environment, and the environment can only override `passive`.
| Key | Default | What it changes |
|---|---|---|
| `observeAfterTokens` | `10000` | observer cadence — lower means smaller chunks and more calls |
| `reflectAfterTokens` | `20000` | reflector cadence (and thereby dropper opportunities) |
| `observerChunkMaxTokens` | 20% of the memory model's context window, else `60000` | cap on one observer run's input |
| `compactAfterTokens` | `81000` | when proactive auto-compaction fires |
| `observationsPoolMaxTokens` | `20000` | pressure at which compaction does a full fold |
| `observationsPoolTargetTokens` | half of max (`10000`) | what the dropper aims back down to |
| `agentMaxTurns` | `16` | shared turn cap for the three workers |
| `model` | unset → session model | send background work to a cheaper/faster model |
| `showWorkerNotifications` | `true` | routine "observer ran" notices |
| `passive` | `false` | **kill switch** for all proactive background work; `recall` and `/om:*` still work |
| `debugLog` | `false` | per-session NDJSON trace at `~/.pi/agent/observational-memory/debug/<session-id>.ndjson` |
One-off passive run, no config edit:
```bash
PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi
```
Invalid values are ignored rather than fatal, so a typo degrades to the default
instead of breaking your session — which also means a typo is silent. Check with
`/om:status`.
## 9. Confirming it is actually working
Do not infer health from the absence of a warning; look:
```bash
# 1. inside pi — the authoritative view
/om:status # visible-vs-full drift, thresholds, worker state
/om:view # what the agent currently sees
/om:view full # full ledger truth at the branch tip
# 2. from a shell — are ledger entries being written?
grep -o '"customType":"om\.[a-z.]*"' \
"$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)" | sort | uniq -c
# 3. which copy is loaded, and at what commit
python3 -c "import json;print(json.load(open('$HOME/.pi/agent/settings.json'))['packages'])"
git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory rev-parse HEAD
```
Note the entry type in the transcript is `custom_message` (with
`customType: "om.…"`), not `custom` — an easy grep to get wrong and conclude
nothing is happening.
## 10. It is not the same thing as MemPalace
Both are called "memory" and they solve different problems. Nothing is wrong
with running both — this image does, and they cover each other's failure modes.
```mermaid
flowchart LR
O0["<b>observational memory</b>"] --> O1["horizon:<br/><b>this session</b>"]
O1 --> O2["scope:<br/>one branch, one machine"]
O2 --> O3["automatic —<br/>no habit required"]
O3 --> O4["retrieval:<br/>an id you already hold<br/>→ <code>recall</code>"]
P0["<b>MemPalace</b>"] --> P1["horizon:<br/><b>sessions, months, machines</b>"]
P1 --> P2["scope:<br/>the whole fleet, every harness"]
P2 --> P3["protocol-driven —<br/>search first, diary at the end"]
P3 --> P4["retrieval:<br/>semantic search, entity+time,<br/>mailbox"]
```
| Question | Answer |
|---|---|
| "What did we decide 200 turns ago in *this* session?" | observational memory (and `recall` for the exact wording) |
| "What did we decide last month, or on another machine?" | MemPalace (`mempalace_search`, diaries) |
| "What is true *right now* about version X?" | MemPalace knowledge graph |
| "Does another machine need something from me?" | MemPalace coordination log — see [Cross-machine agent coordination](../README.md#cross-machine-agent-coordination) |
| "Why is compaction not losing my session?" | observational memory |
The crisp version: **observational memory keeps a session coherent; the palace
keeps the fleet coherent.** A container recreate wipes neither — but only
because `~/.pi` and the palace both live outside the container filesystem.
## 11. Gotchas
- **Branch-local means branch-local.** Resuming or forking changes which ledger
is folded. Memory that "disappeared" is usually on another branch.
- **`recall` needs an id, not a topic.** If you only have a topic, that is a
palace search, not a recall.
- **A `/workspace` clone is not evidence of what is loaded** — see §6.3.
- **`showWorkerNotifications: true` is not proof of work**; it reports runs, and
an observer that deliberately emits nothing writes no ledger entry and simply
retries after another `observeAfterTokens`.
- **`git log` in the baked tree needs `safe.directory`** (`/opt` is
root-owned): `git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory log`.