# Observational memory — why this image has it, and what it does for you **Audience:** anyone using this container for long pi sessions who has wondered what `recall`, `/om:status` and "compacted memory" are, or whether they should leave any of it switched on. **Companion documents:** the extension ships its own reference docs at `/opt/pi-observational-memory/docs/` — [`concepts.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/concepts.md) (the model), [`how-it-works.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/how-it-works.md) (hooks and internals) and [`configuration.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/configuration.md) (every setting). Those are normative; this document is the **deployment** view — what is pinned here, how it is wired, what it costs, and how it differs from MemPalace. For the palace, see [`mempalace-toolkit/docs/fleet-memory.md`](https://gitea.jordbo.se/joakimp/mempalace-toolkit/src/branch/main/docs/fleet-memory.md). > Verified on pi-devbox **v1.8.9** (`release_tag v1.8.9`, source `aac4a1c`), > which bakes pi-observational-memory **v3.0.4** at commit `ce9fc98` — the value > in `/etc/pi-devbox/build-manifest.json` → `components.pi-observational-memory`. > Every number below was read from that tree or from the live container, not > from release notes. --- ## 1. The problem it solves A long pi session outgrows the model's context window. Pi's answer is **compaction**: fold the older part of the conversation into a summary and keep recent messages verbatim. That is unavoidable, and it is where sessions go wrong — the summary is produced *at the moment of pressure*, by a model, about a transcript that is about to be discarded. Observational memory changes *when* the remembering happens. Instead of summarising in a panic at the end, it keeps a small **ledger** up to date while the session runs, and compaction then just folds that ledger. ```mermaid flowchart LR A0["plain compaction"] --> A1["long session,
context filling"] A1 --> A2["a model summarises
the transcript
at the moment of pressure"] A2 --> A3["one prose summary:
detail chosen in a hurry,
no route back to the original"] B0["with observational memory"] --> B1["long session,
context filling"] B1 --> B2["while you work:
a cheap background model appends
observations + reflections
to a ledger"] B2 --> B3["compaction folds the ledger
no model call"] B3 --> B4["every item keeps a 12-hex id —
recall(id) returns
the exact source"] ``` ## 2. The mental model: three layers and a ledger | Layer | What it is | Example | |---|---|---| | **Observation** | a timestamped, source-backed event from the conversation | "user rejected option B because it needs a base rebuild" | | **Reflection** | a durable conclusion *backed by* observations | "the user optimises for avoiding 67-minute rebuilds" | | **Drop** | a tombstone retiring an observation from active memory | the superseded detail of a bug that is now fixed | These are appended as silent ledger entries (`om.observations.recorded`, `om.reflections.recorded`, `om.observations.dropped`) and **folded** — replayed from the branch root — to produce the memory state. The ledger is the source of truth; the compaction summary is a rendering of it. Memory is **branch-local**. A pi session is a tree (resume, fork), and the fold follows the current branch only, so a forked branch does not inherit another branch's view. ## 3. The lifecycle Three background workers and one compaction hook, driven by *raw token progress* rather than wall-clock time. Defaults in brackets. ```mermaid flowchart TD T(["turn_end"]) --> O{"tokens since last
observation coverage
≥ observeAfterTokens
[10000]?"} O -- yes --> OBS["observer runs:
appends om.observations.recorded"] O -- "no (observer not due)" --> R{"tokens since last
reflection
≥ reflectAfterTokens
[20000]?"} R -- yes --> REF["reflector runs:
appends om.reflections.recorded"] REF -- "non-empty reflection AND active pool
> observationsPoolTargetTokens [10000]" --> DR["dropper runs:
appends om.observations.dropped"] S(["agent_settled"]) --> C{"tokens since last compaction
≥ compactAfterTokens [81000]
and pi is idle?"} C -- yes --> CP["extension calls ctx.compact()"] CP --> H(["session_before_compact"]) H --> F["deterministic fold of the ledger:
no model call, no waiting on workers"] F --> VIS["compacted memory the agent reads"] ``` Two details worth carrying, because they are easy to get wrong: - **The dropper has no clock of its own.** It is post-reflection maintenance, gated on a *successful same-turn* reflection — not a third worker on a third threshold. - **Compaction itself calls no model.** That is the whole design: the expensive thinking happened earlier, in the background, on a cheap model. ## 4. What you actually get - **Compaction stops being a stall.** The latency path is deterministic work on ledger entries, not a summarisation call. - **Nothing important vanishes silently.** Compaction is lossy by design, but every observation and reflection keeps a 12-character id, and `recall()` returns the exact evidence behind it — the original wording, the reasoning, the file path, the error text. - **The bookkeeping runs on a cheaper model than your session.** In this image that is deliberate and visible (§6): background workers on Haiku, the session on Opus. - **It is automatic.** There is no habit to maintain, unlike the palace protocol — which is exactly why the two systems complement each other (§10). - **Forks stay clean.** Branch-local memory means a `fork` sub-agent's noise does not leak into the parent's folded memory. ## 5. `recall` is not a search tool `recall` takes **one specific 12-hex id** that already appears in compacted memory or in `/om:view`, and looks it up in the ledger on the current branch. It cannot be given a topic. It can return an observation (marked `active` or `dropped`), or a reflection together with the observations supporting it. ```mermaid sequenceDiagram participant M as compacted memory participant A as agent participant L as ledger (current branch) M->>A: "[high] user rejected option B (a1b2c3d4e5f6)" A->>L: recall("a1b2c3d4e5f6") L-->>A: exact observation + its source entry ids Note over A: acts on the original wording,
not on the compressed paraphrase ``` The rule of thumb the agent skill uses: recall **before a load-bearing action** that rests on a compressed memory — shipping a change, asserting a fact, answering "why do you believe that". One recall is cheap; redoing finished work is not. ## 6. How it is wired in this image ```mermaid flowchart TB IMG["image layer
/opt/pi-observational-memory
v3.0.4 @ ce9fc98 (pinned)"] --> REG["~/.pi/agent/settings.json
packages[]: ../../../../opt/pi-observational-memory
the only source of truth for which copy loads"] REG --> SESS["your pi session"] SESS -->|"your turns"| SM["session model
claude-opus-5"] SESS -->|"observer / reflector / dropper"| WM["memory model
claude-haiku-4-5 (Bedrock)"] SESS -->|"ledger entries (custom_message)"| JL["~/.pi/agent/sessions/<project>/<timestamp>_<uuid>.jsonl"] JL --> VOL[("devbox-pi-config volume — docker-compose.yml:59
survives --force-recreate")] ``` Four consequences of that wiring: 1. **There is no separate database.** Memory *is* entries inside the ordinary pi session `.jsonl`. Nothing extra to back up, and nothing to migrate. 2. **It survives container recreate**, because `~/.pi` is the `devbox-pi-config` named volume — the same one that keeps your pi config and sessions. 3. **`packages[]` is the only source of truth for which copy is loaded.** A clone at `/workspace/pi-observational-memory` may exist (and today matches `/opt` byte-for-byte at `ce9fc98`) — its presence proves nothing. If you ever run a patched build, you point `packages[]` at it explicitly and start a new session. 4. **The worker model is a deliberate choice, and it is yours to change.** The seeded config sends background work to Haiku while your session runs Opus: ```json "observational-memory": { "model": { "provider": "amazon-bedrock", "id": "eu.anthropic.claude-haiku-4-5-20251001-v1:0" }, "debugLog": false } ``` ## 7. What it costs | Resource | Cost | |---|---| | Model calls | up to **three** background calls per consolidation pass (observer, reflector, dropper), each capped at `agentMaxTurns` [16], on the configured memory model — not your session model | | Latency in your turns | none by construction: workers run from `turn_end`, and the compaction hook does no model work | | Disk | negligible — JSON lines inside a session file that would exist anyway (measured here: `~/.pi/agent/sessions` = 30 MB total, tens of `om.*` entries per session) | | Attention | none once configured; there is no protocol for you or the agent to remember | If that is still more than you want on a given run, §8's `passive` switch turns off all proactive work while keeping `recall` and `/om:*` usable. ## 8. Configuration Global: `~/.pi/agent/settings.json` (persisted in the volume). Per project: `/.pi/settings.json`, which overrides global. Precedence is project → global → environment, and the environment can only override `passive`. | Key | Default | What it changes | |---|---|---| | `observeAfterTokens` | `10000` | observer cadence — lower means smaller chunks and more calls | | `reflectAfterTokens` | `20000` | reflector cadence (and thereby dropper opportunities) | | `observerChunkMaxTokens` | 20% of the memory model's context window, else `60000` | cap on one observer run's input | | `compactAfterTokens` | `81000` | when proactive auto-compaction fires | | `observationsPoolMaxTokens` | `20000` | pressure at which compaction does a full fold | | `observationsPoolTargetTokens` | half of max (`10000`) | what the dropper aims back down to | | `agentMaxTurns` | `16` | shared turn cap for the three workers | | `model` | unset → session model | send background work to a cheaper/faster model | | `showWorkerNotifications` | `true` | routine "observer ran" notices | | `passive` | `false` | **kill switch** for all proactive background work; `recall` and `/om:*` still work | | `debugLog` | `false` | per-session NDJSON trace at `~/.pi/agent/observational-memory/debug/.ndjson` | One-off passive run, no config edit: ```bash PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi ``` Invalid values are ignored rather than fatal, so a typo degrades to the default instead of breaking your session — which also means a typo is silent. Check with `/om:status`. ## 9. Confirming it is actually working Do not infer health from the absence of a warning; look: ```bash # 1. inside pi — the authoritative view /om:status # visible-vs-full drift, thresholds, worker state /om:view # what the agent currently sees /om:view full # full ledger truth at the branch tip # 2. from a shell — are ledger entries being written? grep -o '"customType":"om\.[a-z.]*"' \ "$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)" | sort | uniq -c # 3. which copy is loaded, and at what commit python3 -c "import json;print(json.load(open('$HOME/.pi/agent/settings.json'))['packages'])" git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory rev-parse HEAD ``` Note the entry type in the transcript is `custom_message` (with `customType: "om.…"`), not `custom` — an easy grep to get wrong and conclude nothing is happening. ## 10. It is not the same thing as MemPalace Both are called "memory" and they solve different problems. Nothing is wrong with running both — this image does, and they cover each other's failure modes. ```mermaid flowchart LR O0["observational memory"] --> O1["horizon:
this session"] O1 --> O2["scope:
one branch, one machine"] O2 --> O3["automatic —
no habit required"] O3 --> O4["retrieval:
an id you already hold
→ recall"] P0["MemPalace"] --> P1["horizon:
sessions, months, machines"] P1 --> P2["scope:
the whole fleet, every harness"] P2 --> P3["protocol-driven —
search first, diary at the end"] P3 --> P4["retrieval:
semantic search, entity+time,
mailbox"] ``` | Question | Answer | |---|---| | "What did we decide 200 turns ago in *this* session?" | observational memory (and `recall` for the exact wording) | | "What did we decide last month, or on another machine?" | MemPalace (`mempalace_search`, diaries) | | "What is true *right now* about version X?" | MemPalace knowledge graph | | "Does another machine need something from me?" | MemPalace coordination log — see [Cross-machine agent coordination](../README.md#cross-machine-agent-coordination) | | "Why is compaction not losing my session?" | observational memory | The crisp version: **observational memory keeps a session coherent; the palace keeps the fleet coherent.** A container recreate wipes neither — but only because `~/.pi` and the palace both live outside the container filesystem. ## 11. Gotchas - **Branch-local means branch-local.** Resuming or forking changes which ledger is folded. Memory that "disappeared" is usually on another branch. - **`recall` needs an id, not a topic.** If you only have a topic, that is a palace search, not a recall. - **A `/workspace` clone is not evidence of what is loaded** — see §6.3. - **`showWorkerNotifications: true` is not proof of work**; it reports runs, and an observer that deliberately emits nothing writes no ledger entry and simply retries after another `observeAfterTokens`. - **`git log` in the baked tree needs `safe.directory`** (`/opt` is root-owned): `git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory log`.