# Observational memory — why this image has it, and what it does for you **Audience:** anyone using this container for long pi sessions who has wondered what `recall`, `/om:status` and "compacted memory" are, or whether they should leave any of it switched on. **Companion documents:** the extension ships its own reference docs at `/opt/pi-observational-memory/docs/` — [`concepts.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/concepts.md) (the model), [`how-it-works.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/how-it-works.md) (hooks and internals) and [`configuration.md`](https://github.com/elpapi42/pi-observational-memory/blob/main/docs/configuration.md) (every setting). Pi's own compaction mechanics are in `/usr/lib/node_modules/@earendil-works/pi-coding-agent/docs/compaction.md`. Those are normative; this document is the **deployment** view — what is pinned here, how it is wired, what it costs, and how it differs from MemPalace. For the palace, see [`mempalace-toolkit/docs/fleet-memory.md`](https://gitea.jordbo.se/joakimp/mempalace-toolkit/src/branch/main/docs/fleet-memory.md). > Verified on pi-devbox **v1.8.9** (`release_tag v1.8.9`, source `aac4a1c`), > which bakes pi-observational-memory **v3.0.4** at commit `ce9fc98` — the value > in `/etc/pi-devbox/build-manifest.json` → `components.pi-observational-memory`. > Every number below was read from that tree, from pi 0.84.3's own docs, or from > the live container. --- ## 1. The problem it solves A long pi session outgrows the model's context window. Pi's answer is **compaction**: fold the older part of the conversation into a summary and keep recent messages verbatim. That is unavoidable, and it is where sessions go wrong — the summary is produced *at the moment of pressure*, by a model, about a transcript that is about to leave the context. Observational memory changes *when* the remembering happens. Instead of summarising in a panic at the end, it keeps a small **ledger** up to date while the session runs, and compaction then just folds that ledger. ```mermaid flowchart LR A0["plain compaction"] --> A1["context fills"] A1 --> A2["a model summarises
under pressure"] A2 --> A3["prose summary,
no way back"] B0["with observational
memory"] --> B1["context fills"] B1 --> B2["ledger written
as you work"] B2 --> B3["compaction folds
the ledger"] B3 --> B4["ids you can
recall"] ``` Top row is pi on its own: one model call at the worst possible moment, detail chosen in a hurry, and the original wording gone from view. Bottom row is this image's default: the thinking happened earlier on a cheap model, the fold is deterministic, and every line in the result carries an id that resolves back to the exact source. ## 2. The mental model: three layers and a ledger | Layer | What it is | Example | |---|---|---| | **Observation** | a timestamped, source-backed event from the conversation | "user rejected option B because it needs a base rebuild" | | **Reflection** | a durable conclusion *backed by* observations | "the user optimises for avoiding 67-minute rebuilds" | | **Drop** | a tombstone retiring an observation from active memory | the superseded detail of a bug that is now fixed | These are appended to the session as silent ledger entries (`om.observations.recorded`, `om.reflections.recorded`, `om.observations.dropped`) and **folded** — replayed in order — to produce the memory state. The ledger is the source of truth; what you see in a compacted session is a rendering of it. Two properties follow, and both matter later: - **The ledger itself costs no context.** Those entries are pi `custom` entries, which *"do not participate in LLM context"* (pi `docs/session-format.md`). They sit in the session file and reach the model only via the fold at compaction. - **Memory is branch-local.** A pi session is a tree (resume, fork), and the fold follows the current branch only, so a forked branch does not inherit another branch's view. ## 3. The lifecycle Three background workers and one compaction hook, driven by *raw token progress* rather than wall-clock time. Defaults in brackets. ```mermaid flowchart TD T(["turn_end"]) --> O{"10k raw tokens
since observing?"} O -- yes --> OBS["observer runs"] O -- "no" --> R{"20k tokens
since reflecting?"} R -- yes --> REF["reflector runs"] REF -- "if pool over 10k" --> DR["dropper prunes"] S(["agent_settled"]) --> C{"81k tokens
since compacting?"} C -- yes --> CP["ctx.compact()"] CP --> H(["session_before_compact"]) H --> F["fold the ledger
no model call"] F --> VIS["compacted memory"] ``` - **observer** — `observeAfterTokens` [10000]: writes observations for the conversation it has not covered yet. - **reflector** — `reflectAfterTokens` [20000]: promotes patterns across observations into durable reflections. - **dropper** — no clock of its own. It is post-reflection maintenance, gated on a *successful same-turn* reflection **and** an active pool above `observationsPoolTargetTokens` [10000]. Not a third worker on a third threshold. - **compaction** — `compactAfterTokens` [81000], checked when pi goes idle, so it never interrupts a turn. Pi will also compact on its own when the context is nearly full (`contextTokens > contextWindow - reserveTokens`, `reserveTokens` [16384]). ## 4. What compaction actually does to your context This is the question the rest of the document used to leave hanging: if the old conversation is folded away, is the session back to knowing nothing? **No.** Compaction replaces *part* of the context, not all of it, and it deletes nothing at all from disk. ```mermaid flowchart LR SYS["system prompt
+ AGENTS.md"] --> CTX["what the model sees
on the next turn"] SUM["folded memory:
reflections + observations"] --> CTX TAIL["recent turns,
verbatim"] --> CTX DISK[("session .jsonl: all of it")] -. "recall(id)" .-> CTX ``` Where each piece comes from: - **System prompt and `AGENTS.md` — never compacted, because they were never conversation.** Pi rebuilds them from disk on every request (`loadContextFileFromDir`), so they cannot be lost by compaction. - **The verbatim tail — sized by a token budget, not a message count.** Pi walks backwards from the newest entry accumulating token estimates until `keepRecentTokens` [20000] is reached; that entry becomes `firstKeptEntryId`, and *everything from there on is kept unchanged*. Cut points land on turn boundaries, never mid-tool-call. So the most recent ~20k tokens of real work — your last instructions, the diffs, the test output — survive word for word. - **The folded memory — replaces only what came before that cut.** Rendered from the ledger's records: reflections and observations, each with its 12-hex id. - **The session file — untouched.** Compaction *appends* a `compaction` entry (`{"type":"compaction", summary, firstKeptEntryId, tokensBefore, …}`) and rebuilds context from it on later turns. Nothing is rewritten in place; the only documented way to remove session content is deleting the whole `.jsonl`. That last point is what makes the answer to "is the detail gone?" *no* rather than *mostly*: `recall` does not read the context window at all. It calls `sessionManager.getBranch()` — the full branch from the root — and resolves an observation id back to the original entries. Detail that left the model's view an hour ago is still one `recall` away. **Repeated compaction does not summarise the summary.** The rendered text is always built from live observation/reflection *records*, never from the previous compaction's prose, so there is no generation-loss spiral. (Mechanically the projection is incremental — it re-derives back to the last full-fold boundary and carries the rest forward, escalating to a genuine re-fold from the branch root when the observation pool reaches `observationsPoolMaxTokens` [20000].) So the honest summary of the state after compaction: **the model keeps its instructions, keeps recent work verbatim, trades older turns for a dense id-carrying digest of them, and can pull any of it back on demand.** Not a fresh start — a smaller, cheaper, still-navigable one. ### One caveat about "no model call" If the ledger is empty — compaction fires before the observer has ever run — the hook returns nothing and *declines ownership*, and pi's own model-based summariser runs instead: ```ts const summary = renderSummary(projection.reflections, projection.observations); if (summary.length === 0) { // Decline ownership so Pi's native summarizer preserves the pre-cut context. return; } ``` In steady state (any session old enough to have produced one observation) om's hook wins and compaction is model-free. "Never calls a model" is true in practice and false in principle; the fallback is deliberate, so an empty ledger degrades to normal pi rather than to no summary at all. ## 5. What you actually get - **Compaction stops being a stall.** In steady state the latency path is deterministic work over ledger entries, not a summarisation call. - **Nothing important vanishes silently.** Compaction is lossy by design, but every item keeps a 12-character id, and `recall()` returns the exact evidence — original wording, reasoning, file path, error text. - **The bookkeeping runs on a cheaper model than your session.** In this image that is deliberate and visible (§7): background workers on Haiku, session on Opus. - **It is automatic.** No habit to maintain, unlike the palace protocol — which is exactly why the two complement each other (§11). - **Forks stay clean.** Branch-local memory means a `fork` sub-agent's noise does not leak into the parent's folded memory. ## 6. `recall` is not a search tool `recall` takes **one specific 12-hex id** that already appears in compacted memory or in `/om:view`. It cannot be given a topic. It can return an observation (marked `active` or `dropped`), or a reflection together with the observations supporting it. ```mermaid sequenceDiagram participant M as compacted memory participant A as agent participant L as ledger M->>A: "[high] user rejected option B (a1b2c3d4e5f6)" A->>L: recall("a1b2c3d4e5f6") L-->>A: exact observation + source ids Note over A: acts on the original wording ``` The rule of thumb the agent skill uses: recall **before a load-bearing action** that rests on a compressed memory — shipping a change, asserting a fact, answering "why do you believe that". One recall is cheap; redoing finished work is not. ## 7. How it is wired in this image ```mermaid flowchart TB IMG["baked in the image:
v3.0.4 @ ce9fc98"] --> REG["settings.json
packages[]"] REG --> SESS["your pi session"] SESS -- "your turns" --> SM["session model:
Opus"] SESS -- "observer, reflector,
dropper" --> WM["memory model:
Haiku"] SESS -- "ledger entries" --> JL["session .jsonl"] JL --> VOL[("devbox-pi-config
volume")] ``` Four consequences of that wiring: 1. **There is no separate database.** Memory *is* entries inside the ordinary pi session file (`~/.pi/agent/sessions//_.jsonl`). Nothing extra to back up, nothing to migrate. 2. **It survives container recreate**, because `~/.pi` is the `devbox-pi-config` named volume (`docker-compose.yml`) — the same one holding your pi config and session history. 3. **`packages[]` is the only source of truth for which copy is loaded.** A clone at `/workspace/pi-observational-memory` may exist (and today matches `/opt` byte-for-byte at `ce9fc98`) — its presence proves nothing. To run a patched build you point `packages[]` at it explicitly and start a new session. 4. **The worker model is a deliberate choice, and it is yours to change.** The seeded config sends background work to Haiku while your session runs Opus: ```json "observational-memory": { "model": { "provider": "amazon-bedrock", "id": "eu.anthropic.claude-haiku-4-5-20251001-v1:0" }, "debugLog": false } ``` ## 8. What it costs | Resource | Cost | |---|---| | Model calls | up to **three** background calls per consolidation pass (observer, reflector, dropper), each capped at `agentMaxTurns` [16], on the configured memory model — not your session model | | Latency in your turns | none by construction: workers run from `turn_end`, compaction runs when pi is idle, and the fold itself does no model work | | Disk | negligible — JSON lines inside a session file that would exist anyway (measured here: `~/.pi/agent/sessions` = 30 MB total, tens of `om.*` entries per session) | | Context window | **zero until compaction.** `custom` entries do not enter LLM context; only the folded summary does | | Attention | none once configured; there is no protocol for you or the agent to remember | If that is still more than you want on a given run, §9's `passive` switch turns off all proactive work while keeping `recall` and `/om:*` usable. ## 9. Configuration Global: `~/.pi/agent/settings.json` (persisted in the volume). Per project: `/.pi/settings.json`, which overrides global. Precedence is project → global → environment, and the environment can only override `passive`. | Key | Default | What it changes | |---|---|---| | `observeAfterTokens` | `10000` | observer cadence — lower means smaller chunks and more calls | | `reflectAfterTokens` | `20000` | reflector cadence (and thereby dropper opportunities) | | `observerChunkMaxTokens` | 20% of the memory model's context window, else `60000` | cap on one observer run's input | | `compactAfterTokens` | `81000` | when proactive auto-compaction fires | | `observationsPoolMaxTokens` | `20000` | pool size at which compaction does a full re-fold from the branch root | | `observationsPoolTargetTokens` | half of max (`10000`) | what the dropper aims back down to | | `agentMaxTurns` | `16` | shared turn cap for the three workers | | `model` | unset → session model | send background work to a cheaper/faster model | | `showWorkerNotifications` | `true` | routine "observer ran" notices | | `passive` | `false` | **kill switch** for all proactive background work; `recall` and `/om:*` still work | | `debugLog` | `false` | per-session NDJSON trace at `~/.pi/agent/observational-memory/debug/.ndjson` | Pi's own compaction knobs live under a separate `compaction` key — `keepRecentTokens` [20000] sets the verbatim tail from §4, `reserveTokens` [16384] the headroom that triggers pi's own compaction. One-off passive run, no config edit: ```bash PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi ``` Invalid values are ignored rather than fatal, so a typo degrades to the default instead of breaking your session — which also means a typo is silent. Check with `/om:status`. ## 10. Confirming it is actually working Do not infer health from the absence of a warning; look: ```bash # 1. inside pi — the authoritative view /om:status # visible-vs-full drift, thresholds, worker state /om:view # what the agent currently sees /om:view full # full ledger truth at the branch tip # 2. from a shell — are ledger entries being written, and has it compacted? grep -o '"customType":"om\.[a-z.]*"' \ "$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)" | sort | uniq -c grep -c '"type":"compaction"' "$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)" # 3. which copy is loaded, and at what commit python3 -c "import json;print(json.load(open('$HOME/.pi/agent/settings.json'))['packages'])" git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory rev-parse HEAD ``` Ledger entries are `"type":"custom"` with `"customType":"om.…"`. Do not grep for `custom_message` — that is a *different* pi API for entries that **do** enter LLM context, used here by the MemPalace mailbox (`customType: "mempalace-mailbox"`), not by om. ## 11. It is not the same thing as MemPalace Both are called "memory" and they solve different problems. Nothing is wrong with running both — this image does, and they cover each other's failure modes. ```mermaid flowchart LR O0["observational
memory"] --> O1["horizon:
this session"] O1 --> O2["scope: one branch,
one machine"] O2 --> O3["automatic"] O3 --> O4["retrieval:
recall(id)"] P0["MemPalace"] --> P1["horizon: months,
machines"] P1 --> P2["scope:
the fleet"] P2 --> P3["protocol-driven"] P3 --> P4["retrieval:
search, KG, mailbox"] ``` | Question | Answer | |---|---| | "What did we decide 200 turns ago in *this* session?" | observational memory (and `recall` for the exact wording) | | "What did we decide last month, or on another machine?" | MemPalace (`mempalace_search`, diaries) | | "What is true *right now* about version X?" | MemPalace knowledge graph | | "Does another machine need something from me?" | MemPalace coordination log — see [Cross-machine agent coordination](../README.md#cross-machine-agent-coordination) | | "Why is compaction not losing my session?" | observational memory | The crisp version: **observational memory keeps a session coherent; the palace keeps the fleet coherent.** A container recreate wipes neither — but only because `~/.pi` and the palace both live outside the container filesystem. ## 12. Gotchas - **Branch-local means branch-local.** Resuming or forking changes which ledger is folded. Memory that "disappeared" is usually on another branch. - **`recall` needs an id, not a topic.** If you only have a topic, that is a palace search, not a recall. - **A `/workspace` clone is not evidence of what is loaded** — see §7.3. - **`showWorkerNotifications: true` is not proof of work**; it reports runs, and an observer that deliberately emits nothing writes no ledger entry and simply retries after another `observeAfterTokens`. - **A turn bigger than `keepRecentTokens` splits.** The cut then lands mid-turn at an assistant message and pi merges two summaries — rare, but it is why a very large single turn can lose more verbatim detail than you would expect. - **`git log` in the baked tree needs `safe.directory`** (`/opt` is root-owned): `git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory log`.