Files
pi-devbox/docs/observational-memory.md
T
Joakim Persson 14371e2da6
Lint / hadolint (push) Successful in 12s
Lint / actionlint (push) Successful in 16s
docs: explain the memory that runs itself, and fix a claim the mailbox falsified
`pi-observational-memory` is baked, registered in the seeded settings.json, and
handed a cheaper model than the session it serves — and the only description in
this repo was five words in a feature list. Someone meeting `/om:status` or a
"compacted memory" block had nothing to read that said whether to leave any of it
on. New docs/observational-memory.md, 272 lines and five diagrams.

Scoped by who is authoritative, so there is one copy of each claim:

- Upstream (/opt/pi-observational-memory/docs/) already documents the mechanism
  well — concepts.md, how-it-works.md, configuration.md, including a v3 lifecycle
  diagram that matches the deployed code. Linked, not re-derived.
- This document takes the four facts pi-devbox owns and can change: the pinned
  commit it bakes (v3.0.4 ce9fc98, matching build-manifest.json), the packages[]
  entry that decides which copy loads, the Haiku-workers-vs-Opus-session split,
  and the devbox-pi-config volume that makes the ledger outlive the container.
- Plus the confusion this image creates by shipping two things called memory: a
  section contrasting it with MemPalace, on the line "observational memory keeps
  a session coherent, the palace keeps the fleet coherent".

Placement follows the split fleet-ops states for itself — reusable mechanism is
not deployment data — so a "why is this in my container" document belongs in the
repo that pins and wires the component. Linked twice from the README, because
until this commit the README referenced docs/ zero times and the file already
sitting there was reachable only by listing the directory.

Numbers were read out of the live container and the baked tree rather than out of
release notes, which caught one thing the pi-extensions skill still has wrong:
the dropper is gated on a successful same-turn reflection, not on a token
threshold of its own.

Also fixes a README sentence that v1.8.9 made false. § Cross-machine agent
coordination ended with "Nothing in this image polls the log on the agent's
behalf"; the mailbox has shipped since aac4a1c. Replaced with the three knobs and
their defaults, derived-not-read owed-ness, and the queued-into-the-next-turn
delivery measured on two devices — and dated to "as baked in v1.8.9
(mempalace-toolkit 5b8d78f)", pointing at RFC 003 §7.11–§7.12 for the mechanism,
because toolkit main is already ahead (a92c75d pings the human who is not
looking) and describing that here would trade a stale-behind claim for a
stale-ahead one.

The cause is worth more than the fix: the behaviour arrived through the floating
MEMPALACE_TOOLKIT_REF, so no diff in this repo ever touched the paragraph making
the claim. v1.8.9's own rule fired for the CHANGELOG and nobody swept the README.
The CHANGELOG records what changed, the README asserts what is true, and only the
first is reviewed at release time — so the rule now extends to grepping the README
for absolute claims (nothing, never, does not, only) about a component whose SHA
moved.

Diagrams were verified by rendering, not by parsing. Both comparison diagrams
parsed clean and rendered with their meaning reversed: Mermaid laid the second
declared subgraph out first, putting "with observational memory" before "without"
and MemPalace before observational memory in the diagram whose entire job was
that contrast. Rebuilt as declaration-ordered chains and re-rendered at mermaid@11
— the version pi-studio pins — in the baked headless browser, then read back as an
image. mermaid.parse() proves syntax and says nothing about layout.
2026-08-27 13:55:55 +02:00

14 KiB

Observational memory — why this image has it, and what it does for you

Audience: anyone using this container for long pi sessions who has wondered what recall, /om:status and "compacted memory" are, or whether they should leave any of it switched on.

Companion documents: the extension ships its own reference docs at /opt/pi-observational-memory/docs/ — concepts.md (the model), how-it-works.md (hooks and internals) and configuration.md (every setting). Those are normative; this document is the deployment view — what is pinned here, how it is wired, what it costs, and how it differs from MemPalace. For the palace, see mempalace-toolkit/docs/fleet-memory.md.

Verified on pi-devbox v1.8.9 (release_tag v1.8.9, source aac4a1c), which bakes pi-observational-memory v3.0.4 at commit ce9fc98 — the value in /etc/pi-devbox/build-manifest.json → components.pi-observational-memory. Every number below was read from that tree or from the live container, not from release notes.


1. The problem it solves

A long pi session outgrows the model's context window. Pi's answer is compaction: fold the older part of the conversation into a summary and keep recent messages verbatim. That is unavoidable, and it is where sessions go wrong — the summary is produced at the moment of pressure, by a model, about a transcript that is about to be discarded.

Observational memory changes when the remembering happens. Instead of summarising in a panic at the end, it keeps a small ledger up to date while the session runs, and compaction then just folds that ledger.

flowchart LR
    A0["<b>plain compaction</b>"] --> A1["long session,<br/>context filling"]
    A1 --> A2["a model summarises<br/>the transcript<br/><i>at the moment of pressure</i>"]
    A2 --> A3["one prose summary:<br/>detail chosen in a hurry,<br/>no route back to the original"]
    B0["<b>with observational memory</b>"] --> B1["long session,<br/>context filling"]
    B1 --> B2["<i>while you work:</i><br/>a cheap background model appends<br/>observations + reflections<br/>to a ledger"]
    B2 --> B3["compaction folds the ledger<br/><b>no model call</b>"]
    B3 --> B4["every item keeps a 12-hex id —<br/><code>recall(id)</code> returns<br/>the exact source"]

2. The mental model: three layers and a ledger

Layer What it is Example
Observation a timestamped, source-backed event from the conversation "user rejected option B because it needs a base rebuild"
Reflection a durable conclusion backed by observations "the user optimises for avoiding 67-minute rebuilds"
Drop a tombstone retiring an observation from active memory the superseded detail of a bug that is now fixed

These are appended as silent ledger entries (om.observations.recorded, om.reflections.recorded, om.observations.dropped) and folded — replayed from the branch root — to produce the memory state. The ledger is the source of truth; the compaction summary is a rendering of it.

Memory is branch-local. A pi session is a tree (resume, fork), and the fold follows the current branch only, so a forked branch does not inherit another branch's view.

3. The lifecycle

Three background workers and one compaction hook, driven by raw token progress rather than wall-clock time. Defaults in brackets.

flowchart TD
    T(["turn_end"]) --> O{"tokens since last<br/>observation coverage<br/>≥ observeAfterTokens<br/>[10000]?"}
    O -- yes --> OBS["<b>observer</b> runs:<br/>appends om.observations.recorded"]
    O -- "no (observer not due)" --> R{"tokens since last<br/>reflection<br/>≥ reflectAfterTokens<br/>[20000]?"}
    R -- yes --> REF["<b>reflector</b> runs:<br/>appends om.reflections.recorded"]
    REF -- "non-empty reflection AND active pool<br/>&gt; observationsPoolTargetTokens [10000]" --> DR["<b>dropper</b> runs:<br/>appends om.observations.dropped"]
    S(["agent_settled"]) --> C{"tokens since last compaction<br/>≥ compactAfterTokens [81000]<br/>and pi is idle?"}
    C -- yes --> CP["extension calls ctx.compact()"]
    CP --> H(["session_before_compact"])
    H --> F["<b>deterministic fold</b> of the ledger:<br/>no model call, no waiting on workers"]
    F --> VIS["compacted memory the agent reads"]

Two details worth carrying, because they are easy to get wrong:

  • The dropper has no clock of its own. It is post-reflection maintenance, gated on a successful same-turn reflection — not a third worker on a third threshold.
  • Compaction itself calls no model. That is the whole design: the expensive thinking happened earlier, in the background, on a cheap model.

4. What you actually get

  • Compaction stops being a stall. The latency path is deterministic work on ledger entries, not a summarisation call.
  • Nothing important vanishes silently. Compaction is lossy by design, but every observation and reflection keeps a 12-character id, and recall(<id>) returns the exact evidence behind it — the original wording, the reasoning, the file path, the error text.
  • The bookkeeping runs on a cheaper model than your session. In this image that is deliberate and visible (§6): background workers on Haiku, the session on Opus.
  • It is automatic. There is no habit to maintain, unlike the palace protocol — which is exactly why the two systems complement each other (§10).
  • Forks stay clean. Branch-local memory means a fork sub-agent's noise does not leak into the parent's folded memory.

5. recall is not a search tool

recall takes one specific 12-hex id that already appears in compacted memory or in /om:view, and looks it up in the ledger on the current branch. It cannot be given a topic. It can return an observation (marked active or dropped), or a reflection together with the observations supporting it.

sequenceDiagram
    participant M as compacted memory
    participant A as agent
    participant L as ledger (current branch)
    M->>A: "[high] user rejected option B (a1b2c3d4e5f6)"
    A->>L: recall("a1b2c3d4e5f6")
    L-->>A: exact observation + its source entry ids
    Note over A: acts on the original wording,<br/>not on the compressed paraphrase

The rule of thumb the agent skill uses: recall before a load-bearing action that rests on a compressed memory — shipping a change, asserting a fact, answering "why do you believe that". One recall is cheap; redoing finished work is not.

6. How it is wired in this image

flowchart TB
    IMG["<b>image layer</b><br/>/opt/pi-observational-memory<br/>v3.0.4 @ ce9fc98 (pinned)"] --> REG["<b>~/.pi/agent/settings.json</b><br/>packages[]: ../../../../opt/pi-observational-memory<br/><i>the only source of truth for which copy loads</i>"]
    REG --> SESS["<b>your pi session</b>"]
    SESS -->|"your turns"| SM["session model<br/>claude-opus-5"]
    SESS -->|"observer / reflector / dropper"| WM["memory model<br/>claude-haiku-4-5 (Bedrock)"]
    SESS -->|"ledger entries (custom_message)"| JL["~/.pi/agent/sessions/&lt;project&gt;/&lt;timestamp&gt;_&lt;uuid&gt;.jsonl"]
    JL --> VOL[("devbox-pi-config volume — docker-compose.yml:59<br/>survives --force-recreate")]

Four consequences of that wiring:

  1. There is no separate database. Memory is entries inside the ordinary pi session .jsonl. Nothing extra to back up, and nothing to migrate.
  2. It survives container recreate, because ~/.pi is the devbox-pi-config named volume — the same one that keeps your pi config and sessions.
  3. packages[] is the only source of truth for which copy is loaded. A clone at /workspace/pi-observational-memory may exist (and today matches /opt byte-for-byte at ce9fc98) — its presence proves nothing. If you ever run a patched build, you point packages[] at it explicitly and start a new session.
  4. The worker model is a deliberate choice, and it is yours to change. The seeded config sends background work to Haiku while your session runs Opus:
"observational-memory": {
  "model": { "provider": "amazon-bedrock", "id": "eu.anthropic.claude-haiku-4-5-20251001-v1:0" },
  "debugLog": false
}

7. What it costs

Resource Cost
Model calls up to three background calls per consolidation pass (observer, reflector, dropper), each capped at agentMaxTurns [16], on the configured memory model — not your session model
Latency in your turns none by construction: workers run from turn_end, and the compaction hook does no model work
Disk negligible — JSON lines inside a session file that would exist anyway (measured here: ~/.pi/agent/sessions = 30 MB total, tens of om.* entries per session)
Attention none once configured; there is no protocol for you or the agent to remember

If that is still more than you want on a given run, §8's passive switch turns off all proactive work while keeping recall and /om:* usable.

8. Configuration

Global: ~/.pi/agent/settings.json (persisted in the volume). Per project: <project>/.pi/settings.json, which overrides global. Precedence is project → global → environment, and the environment can only override passive.

Key Default What it changes
observeAfterTokens 10000 observer cadence — lower means smaller chunks and more calls
reflectAfterTokens 20000 reflector cadence (and thereby dropper opportunities)
observerChunkMaxTokens 20% of the memory model's context window, else 60000 cap on one observer run's input
compactAfterTokens 81000 when proactive auto-compaction fires
observationsPoolMaxTokens 20000 pressure at which compaction does a full fold
observationsPoolTargetTokens half of max (10000) what the dropper aims back down to
agentMaxTurns 16 shared turn cap for the three workers
model unset → session model send background work to a cheaper/faster model
showWorkerNotifications true routine "observer ran" notices
passive false kill switch for all proactive background work; recall and /om:* still work
debugLog false per-session NDJSON trace at ~/.pi/agent/observational-memory/debug/<session-id>.ndjson

One-off passive run, no config edit:

PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi

Invalid values are ignored rather than fatal, so a typo degrades to the default instead of breaking your session — which also means a typo is silent. Check with /om:status.

9. Confirming it is actually working

Do not infer health from the absence of a warning; look:

# 1. inside pi — the authoritative view
/om:status          # visible-vs-full drift, thresholds, worker state
/om:view            # what the agent currently sees
/om:view full       # full ledger truth at the branch tip

# 2. from a shell — are ledger entries being written?
grep -o '"customType":"om\.[a-z.]*"' \
  "$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)" | sort | uniq -c

# 3. which copy is loaded, and at what commit
python3 -c "import json;print(json.load(open('$HOME/.pi/agent/settings.json'))['packages'])"
git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory rev-parse HEAD

Note the entry type in the transcript is custom_message (with customType: "om.…"), not custom — an easy grep to get wrong and conclude nothing is happening.

10. It is not the same thing as MemPalace

Both are called "memory" and they solve different problems. Nothing is wrong with running both — this image does, and they cover each other's failure modes.

flowchart LR
    O0["<b>observational memory</b>"] --> O1["horizon:<br/><b>this session</b>"]
    O1 --> O2["scope:<br/>one branch, one machine"]
    O2 --> O3["automatic —<br/>no habit required"]
    O3 --> O4["retrieval:<br/>an id you already hold<br/>→ <code>recall</code>"]
    P0["<b>MemPalace</b>"] --> P1["horizon:<br/><b>sessions, months, machines</b>"]
    P1 --> P2["scope:<br/>the whole fleet, every harness"]
    P2 --> P3["protocol-driven —<br/>search first, diary at the end"]
    P3 --> P4["retrieval:<br/>semantic search, entity+time,<br/>mailbox"]
Question Answer
"What did we decide 200 turns ago in this session?" observational memory (and recall for the exact wording)
"What did we decide last month, or on another machine?" MemPalace (mempalace_search, diaries)
"What is true right now about version X?" MemPalace knowledge graph
"Does another machine need something from me?" MemPalace coordination log — see Cross-machine agent coordination
"Why is compaction not losing my session?" observational memory

The crisp version: observational memory keeps a session coherent; the palace keeps the fleet coherent. A container recreate wipes neither — but only because ~/.pi and the palace both live outside the container filesystem.

11. Gotchas

  • Branch-local means branch-local. Resuming or forking changes which ledger is folded. Memory that "disappeared" is usually on another branch.
  • recall needs an id, not a topic. If you only have a topic, that is a palace search, not a recall.
  • A /workspace clone is not evidence of what is loaded — see §6.3.
  • showWorkerNotifications: true is not proof of work; it reports runs, and an observer that deliberately emits nothing writes no ledger entry and simply retries after another observeAfterTokens.
  • git log in the baked tree needs safe.directory (/opt is root-owned): git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory log.