docs: explain the memory that runs itself, and fix a claim the mailbox falsified
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 16s

`pi-observational-memory` is baked, registered in the seeded settings.json, and
handed a cheaper model than the session it serves — and the only description in
this repo was five words in a feature list. Someone meeting `/om:status` or a
"compacted memory" block had nothing to read that said whether to leave any of it
on. New docs/observational-memory.md, 272 lines and five diagrams.

Scoped by who is authoritative, so there is one copy of each claim:

- Upstream (/opt/pi-observational-memory/docs/) already documents the mechanism
  well — concepts.md, how-it-works.md, configuration.md, including a v3 lifecycle
  diagram that matches the deployed code. Linked, not re-derived.
- This document takes the four facts pi-devbox owns and can change: the pinned
  commit it bakes (v3.0.4 ce9fc98, matching build-manifest.json), the packages[]
  entry that decides which copy loads, the Haiku-workers-vs-Opus-session split,
  and the devbox-pi-config volume that makes the ledger outlive the container.
- Plus the confusion this image creates by shipping two things called memory: a
  section contrasting it with MemPalace, on the line "observational memory keeps
  a session coherent, the palace keeps the fleet coherent".

Placement follows the split fleet-ops states for itself — reusable mechanism is
not deployment data — so a "why is this in my container" document belongs in the
repo that pins and wires the component. Linked twice from the README, because
until this commit the README referenced docs/ zero times and the file already
sitting there was reachable only by listing the directory.

Numbers were read out of the live container and the baked tree rather than out of
release notes, which caught one thing the pi-extensions skill still has wrong:
the dropper is gated on a successful same-turn reflection, not on a token
threshold of its own.

Also fixes a README sentence that v1.8.9 made false. § Cross-machine agent
coordination ended with "Nothing in this image polls the log on the agent's
behalf"; the mailbox has shipped since aac4a1c. Replaced with the three knobs and
their defaults, derived-not-read owed-ness, and the queued-into-the-next-turn
delivery measured on two devices — and dated to "as baked in v1.8.9
(mempalace-toolkit 5b8d78f)", pointing at RFC 003 §7.11–§7.12 for the mechanism,
because toolkit main is already ahead (a92c75d pings the human who is not
looking) and describing that here would trade a stale-behind claim for a
stale-ahead one.

The cause is worth more than the fix: the behaviour arrived through the floating
MEMPALACE_TOOLKIT_REF, so no diff in this repo ever touched the paragraph making
the claim. v1.8.9's own rule fired for the CHANGELOG and nobody swept the README.
The CHANGELOG records what changed, the README asserts what is true, and only the
first is reviewed at release time — so the rule now extends to grepping the README
for absolute claims (nothing, never, does not, only) about a component whose SHA
moved.

Diagrams were verified by rendering, not by parsing. Both comparison diagrams
parsed clean and rendered with their meaning reversed: Mermaid laid the second
declared subgraph out first, putting "with observational memory" before "without"
and MemPalace before observational memory in the diagram whose entire job was
that contrast. Rebuilt as declaration-ordered chains and re-rendered at mermaid@11
— the version pi-studio pins — in the baked headless browser, then read back as an
image. mermaid.parse() proves syntax and says nothing about layout.
This commit is contained in:
Joakim Persson
2026-08-27 13:55:55 +02:00
parent aac4a1c323
commit 14371e2da6
3 changed files with 396 additions and 2 deletions
+272
View File
@@ -0,0 +1,272 @@
# Observational memory — why this image has it, and what it does for you
**Audience:** anyone using this container for long pi sessions who has wondered
what `recall`, `/om:status` and "compacted memory" are, or whether they should
leave any of it switched on.
**Companion documents:** the extension ships its own reference docs at
`/opt/pi-observational-memory/docs/` —
[`concepts.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/concepts.md)
(the model),
[`how-it-works.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/how-it-works.md)
(hooks and internals) and
[`configuration.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/configuration.md)
(every setting). Those are normative; this document is the **deployment** view —
what is pinned here, how it is wired, what it costs, and how it differs from
MemPalace. For the palace, see
[`mempalace-toolkit/docs/fleet-memory.md`](https://gitea.jordbo.se/joakimp/mempalace-toolkit/src/branch/main/docs/fleet-memory.md).
> Verified on pi-devbox **v1.8.9** (`release_tag v1.8.9`, source `aac4a1c`),
> which bakes pi-observational-memory **v3.0.4** at commit `ce9fc98` — the value
> in `/etc/pi-devbox/build-manifest.json` → `components.pi-observational-memory`.
> Every number below was read from that tree or from the live container, not
> from release notes.
---
## 1. The problem it solves
A long pi session outgrows the model's context window. Pi's answer is
**compaction**: fold the older part of the conversation into a summary and keep
recent messages verbatim. That is unavoidable, and it is where sessions go
wrong — the summary is produced *at the moment of pressure*, by a model, about a
transcript that is about to be discarded.
Observational memory changes *when* the remembering happens. Instead of
summarising in a panic at the end, it keeps a small **ledger** up to date while
the session runs, and compaction then just folds that ledger.
```mermaid
flowchart LR
A0["<b>plain compaction</b>"] --> A1["long session,<br/>context filling"]
A1 --> A2["a model summarises<br/>the transcript<br/><i>at the moment of pressure</i>"]
A2 --> A3["one prose summary:<br/>detail chosen in a hurry,<br/>no route back to the original"]
B0["<b>with observational memory</b>"] --> B1["long session,<br/>context filling"]
B1 --> B2["<i>while you work:</i><br/>a cheap background model appends<br/>observations + reflections<br/>to a ledger"]
B2 --> B3["compaction folds the ledger<br/><b>no model call</b>"]
B3 --> B4["every item keeps a 12-hex id —<br/><code>recall(id)</code> returns<br/>the exact source"]
```
## 2. The mental model: three layers and a ledger
| Layer | What it is | Example |
|---|---|---|
| **Observation** | a timestamped, source-backed event from the conversation | "user rejected option B because it needs a base rebuild" |
| **Reflection** | a durable conclusion *backed by* observations | "the user optimises for avoiding 67-minute rebuilds" |
| **Drop** | a tombstone retiring an observation from active memory | the superseded detail of a bug that is now fixed |
These are appended as silent ledger entries (`om.observations.recorded`,
`om.reflections.recorded`, `om.observations.dropped`) and **folded** — replayed
from the branch root — to produce the memory state. The ledger is the source of
truth; the compaction summary is a rendering of it.
Memory is **branch-local**. A pi session is a tree (resume, fork), and the fold
follows the current branch only, so a forked branch does not inherit another
branch's view.
## 3. The lifecycle
Three background workers and one compaction hook, driven by *raw token
progress* rather than wall-clock time. Defaults in brackets.
```mermaid
flowchart TD
T(["turn_end"]) --> O{"tokens since last<br/>observation coverage<br/>≥ observeAfterTokens<br/>[10000]?"}
O -- yes --> OBS["<b>observer</b> runs:<br/>appends om.observations.recorded"]
O -- "no (observer not due)" --> R{"tokens since last<br/>reflection<br/>≥ reflectAfterTokens<br/>[20000]?"}
R -- yes --> REF["<b>reflector</b> runs:<br/>appends om.reflections.recorded"]
REF -- "non-empty reflection AND active pool<br/>&gt; observationsPoolTargetTokens [10000]" --> DR["<b>dropper</b> runs:<br/>appends om.observations.dropped"]
S(["agent_settled"]) --> C{"tokens since last compaction<br/>≥ compactAfterTokens [81000]<br/>and pi is idle?"}
C -- yes --> CP["extension calls ctx.compact()"]
CP --> H(["session_before_compact"])
H --> F["<b>deterministic fold</b> of the ledger:<br/>no model call, no waiting on workers"]
F --> VIS["compacted memory the agent reads"]
```
Two details worth carrying, because they are easy to get wrong:
- **The dropper has no clock of its own.** It is post-reflection maintenance,
gated on a *successful same-turn* reflection — not a third worker on a third
threshold.
- **Compaction itself calls no model.** That is the whole design: the expensive
thinking happened earlier, in the background, on a cheap model.
## 4. What you actually get
- **Compaction stops being a stall.** The latency path is deterministic work on
ledger entries, not a summarisation call.
- **Nothing important vanishes silently.** Compaction is lossy by design, but
every observation and reflection keeps a 12-character id, and `recall(<id>)`
returns the exact evidence behind it — the original wording, the reasoning,
the file path, the error text.
- **The bookkeeping runs on a cheaper model than your session.** In this image
that is deliberate and visible (§6): background workers on Haiku, the session
on Opus.
- **It is automatic.** There is no habit to maintain, unlike the palace protocol
— which is exactly why the two systems complement each other (§10).
- **Forks stay clean.** Branch-local memory means a `fork` sub-agent's noise does
not leak into the parent's folded memory.
## 5. `recall` is not a search tool
`recall` takes **one specific 12-hex id** that already appears in compacted
memory or in `/om:view`, and looks it up in the ledger on the current branch. It
cannot be given a topic. It can return an observation (marked `active` or
`dropped`), or a reflection together with the observations supporting it.
```mermaid
sequenceDiagram
participant M as compacted memory
participant A as agent
participant L as ledger (current branch)
M->>A: "[high] user rejected option B (a1b2c3d4e5f6)"
A->>L: recall("a1b2c3d4e5f6")
L-->>A: exact observation + its source entry ids
Note over A: acts on the original wording,<br/>not on the compressed paraphrase
```
The rule of thumb the agent skill uses: recall **before a load-bearing action**
that rests on a compressed memory — shipping a change, asserting a fact,
answering "why do you believe that". One recall is cheap; redoing finished work
is not.
## 6. How it is wired in this image
```mermaid
flowchart TB
IMG["<b>image layer</b><br/>/opt/pi-observational-memory<br/>v3.0.4 @ ce9fc98 (pinned)"] --> REG["<b>~/.pi/agent/settings.json</b><br/>packages[]: ../../../../opt/pi-observational-memory<br/><i>the only source of truth for which copy loads</i>"]
REG --> SESS["<b>your pi session</b>"]
SESS -->|"your turns"| SM["session model<br/>claude-opus-5"]
SESS -->|"observer / reflector / dropper"| WM["memory model<br/>claude-haiku-4-5 (Bedrock)"]
SESS -->|"ledger entries (custom_message)"| JL["~/.pi/agent/sessions/&lt;project&gt;/&lt;timestamp&gt;_&lt;uuid&gt;.jsonl"]
JL --> VOL[("devbox-pi-config volume — docker-compose.yml:59<br/>survives --force-recreate")]
```
Four consequences of that wiring:
1. **There is no separate database.** Memory *is* entries inside the ordinary pi
session `.jsonl`. Nothing extra to back up, and nothing to migrate.
2. **It survives container recreate**, because `~/.pi` is the
`devbox-pi-config` named volume — the same one that keeps your pi config and
sessions.
3. **`packages[]` is the only source of truth for which copy is loaded.** A
clone at `/workspace/pi-observational-memory` may exist (and today matches
`/opt` byte-for-byte at `ce9fc98`) — its presence proves nothing. If you ever
run a patched build, you point `packages[]` at it explicitly and start a new
session.
4. **The worker model is a deliberate choice, and it is yours to change.** The
seeded config sends background work to Haiku while your session runs Opus:
```json
"observational-memory": {
"model": { "provider": "amazon-bedrock", "id": "eu.anthropic.claude-haiku-4-5-20251001-v1:0" },
"debugLog": false
}
```
## 7. What it costs
| Resource | Cost |
|---|---|
| Model calls | up to **three** background calls per consolidation pass (observer, reflector, dropper), each capped at `agentMaxTurns` [16], on the configured memory model — not your session model |
| Latency in your turns | none by construction: workers run from `turn_end`, and the compaction hook does no model work |
| Disk | negligible — JSON lines inside a session file that would exist anyway (measured here: `~/.pi/agent/sessions` = 30 MB total, tens of `om.*` entries per session) |
| Attention | none once configured; there is no protocol for you or the agent to remember |
If that is still more than you want on a given run, §8's `passive` switch turns
off all proactive work while keeping `recall` and `/om:*` usable.
## 8. Configuration
Global: `~/.pi/agent/settings.json` (persisted in the volume). Per project:
`<project>/.pi/settings.json`, which overrides global. Precedence is
project → global → environment, and the environment can only override `passive`.
| Key | Default | What it changes |
|---|---|---|
| `observeAfterTokens` | `10000` | observer cadence — lower means smaller chunks and more calls |
| `reflectAfterTokens` | `20000` | reflector cadence (and thereby dropper opportunities) |
| `observerChunkMaxTokens` | 20% of the memory model's context window, else `60000` | cap on one observer run's input |
| `compactAfterTokens` | `81000` | when proactive auto-compaction fires |
| `observationsPoolMaxTokens` | `20000` | pressure at which compaction does a full fold |
| `observationsPoolTargetTokens` | half of max (`10000`) | what the dropper aims back down to |
| `agentMaxTurns` | `16` | shared turn cap for the three workers |
| `model` | unset → session model | send background work to a cheaper/faster model |
| `showWorkerNotifications` | `true` | routine "observer ran" notices |
| `passive` | `false` | **kill switch** for all proactive background work; `recall` and `/om:*` still work |
| `debugLog` | `false` | per-session NDJSON trace at `~/.pi/agent/observational-memory/debug/<session-id>.ndjson` |
One-off passive run, no config edit:
```bash
PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi
```
Invalid values are ignored rather than fatal, so a typo degrades to the default
instead of breaking your session — which also means a typo is silent. Check with
`/om:status`.
## 9. Confirming it is actually working
Do not infer health from the absence of a warning; look:
```bash
# 1. inside pi — the authoritative view
/om:status # visible-vs-full drift, thresholds, worker state
/om:view # what the agent currently sees
/om:view full # full ledger truth at the branch tip
# 2. from a shell — are ledger entries being written?
grep -o '"customType":"om\.[a-z.]*"' \
"$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)" | sort | uniq -c
# 3. which copy is loaded, and at what commit
python3 -c "import json;print(json.load(open('$HOME/.pi/agent/settings.json'))['packages'])"
git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory rev-parse HEAD
```
Note the entry type in the transcript is `custom_message` (with
`customType: "om.…"`), not `custom` — an easy grep to get wrong and conclude
nothing is happening.
## 10. It is not the same thing as MemPalace
Both are called "memory" and they solve different problems. Nothing is wrong
with running both — this image does, and they cover each other's failure modes.
```mermaid
flowchart LR
O0["<b>observational memory</b>"] --> O1["horizon:<br/><b>this session</b>"]
O1 --> O2["scope:<br/>one branch, one machine"]
O2 --> O3["automatic —<br/>no habit required"]
O3 --> O4["retrieval:<br/>an id you already hold<br/>→ <code>recall</code>"]
P0["<b>MemPalace</b>"] --> P1["horizon:<br/><b>sessions, months, machines</b>"]
P1 --> P2["scope:<br/>the whole fleet, every harness"]
P2 --> P3["protocol-driven —<br/>search first, diary at the end"]
P3 --> P4["retrieval:<br/>semantic search, entity+time,<br/>mailbox"]
```
| Question | Answer |
|---|---|
| "What did we decide 200 turns ago in *this* session?" | observational memory (and `recall` for the exact wording) |
| "What did we decide last month, or on another machine?" | MemPalace (`mempalace_search`, diaries) |
| "What is true *right now* about version X?" | MemPalace knowledge graph |
| "Does another machine need something from me?" | MemPalace coordination log — see [Cross-machine agent coordination](../README.md#cross-machine-agent-coordination) |
| "Why is compaction not losing my session?" | observational memory |
The crisp version: **observational memory keeps a session coherent; the palace
keeps the fleet coherent.** A container recreate wipes neither — but only
because `~/.pi` and the palace both live outside the container filesystem.
## 11. Gotchas
- **Branch-local means branch-local.** Resuming or forking changes which ledger
is folded. Memory that "disappeared" is usually on another branch.
- **`recall` needs an id, not a topic.** If you only have a topic, that is a
palace search, not a recall.
- **A `/workspace` clone is not evidence of what is loaded** — see §6.3.
- **`showWorkerNotifications: true` is not proof of work**; it reports runs, and
an observer that deliberately emits nothing writes no ledger entry and simply
retries after another `observeAfterTokens`.
- **`git log` in the baked tree needs `safe.directory`** (`/opt` is
root-owned): `git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory log`.