Files
pi-devbox/docs/observational-memory.md
T
Joakim Persson 8a673ec143
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 15s
docs: unclip the diagrams, and answer what compaction leaves behind
Two problems reported against docs/observational-memory.md, one cosmetic and one
substantive. Both turned out to be worth more than the fix.

Clipping. Several boxes lost their bottom line of text in the viewer, and neither
the source nor my own render showed it. Cause: Mermaid measures a node label with
its own font metrics, commits to a box size, then renders that label as real HTML
in a <foreignObject> — so any host stylesheet touching line-height or font-size
inflates the text past a box that is already fixed, and the overflow is clipped.
The error accumulates per line, which is why it always eats the last line of the
tallest labels.

Rejected the obvious fix after testing it rather than assuming it:
%%{init: {'flowchart': {'htmlLabels': false}}}%% *is* honoured (labels switch from
16 foreignObject to 7 tspan) and clips identically, because inflated font-size
inherits into SVG text too. The fix that works is a hard limit of two short lines
per node, with the detail moved into prose under each diagram — one- and two-line
boxes have the vertical slack to absorb inflation, three- and four-line boxes do
not. The nodes were carrying paragraph-sized text; the diagrams are better for
losing it.

Regression harness: render every block with line-height: 1.7 !important forced
onto the label HTML, screenshot, read it. That caught two survivors of the rewrite
that looked fine in the clean render — a long unbreakable /opt path wrapping to a
third line, and a cylinder shape whose curved bottom leaves less room than a
rectangle for the same two lines.

New §4, because the document explained that compaction folds the ledger and never
said what that leaves in the context. Asked directly: is the session back to
knowing nothing? No. A verbatim tail survives, sized by keepRecentTokens (20k) and
cut only at turn boundaries; the system prompt and AGENTS.md were never in the
compacted region because they are rebuilt from disk each request; nothing is
deleted from disk, since compaction appends a compaction entry rather than
rewriting lines; and recall keeps resolving ids whose sources left the context
because it reads the full branch via sessionManager.getBranch() and never consults
the context window. Repeated compaction renders from live records, not from the
previous summary's prose, so there is no generation-loss spiral.

Also corrects this repo's own "compaction calls no model" to the steady-state
claim it actually is: with an empty ledger the hook returns nothing and explicitly
declines ownership, and pi's native model-based summariser runs. Snippet quoted in
the doc.

And a bug shipped in the first version: §9 said the ledger entries are
custom_message and specifically not custom. Exactly backwards, so the one grep
that section existed to get right was the one it got wrong. Verified empirically
against the live session file — 11 om.observations.recorded and 6
om.reflections.recorded, all "type":"custom", next to "type":"custom_message"
entries whose customType is mempalace-mailbox and mempalace-wakeup. That is where
the confusion came from, and the distinction is load-bearing rather than
cosmetic: the mailbox uses the context-visible append API, om's ledger uses the
invisible one. Which makes it a feature the doc now advertises — the ledger costs
zero context until it is folded.
2026-08-27 14:24:19 +02:00

18 KiB

Observational memory — why this image has it, and what it does for you

Audience: anyone using this container for long pi sessions who has wondered what recall, /om:status and "compacted memory" are, or whether they should leave any of it switched on.

Companion documents: the extension ships its own reference docs at /opt/pi-observational-memory/docs/ — concepts.md (the model), how-it-works.md (hooks and internals) and configuration.md (every setting). Pi's own compaction mechanics are in /usr/lib/node_modules/@earendil-works/pi-coding-agent/docs/compaction.md. Those are normative; this document is the deployment view — what is pinned here, how it is wired, what it costs, and how it differs from MemPalace. For the palace, see mempalace-toolkit/docs/fleet-memory.md.

Verified on pi-devbox v1.8.9 (release_tag v1.8.9, source aac4a1c), which bakes pi-observational-memory v3.0.4 at commit ce9fc98 — the value in /etc/pi-devbox/build-manifest.json → components.pi-observational-memory. Every number below was read from that tree, from pi 0.84.3's own docs, or from the live container.


1. The problem it solves

A long pi session outgrows the model's context window. Pi's answer is compaction: fold the older part of the conversation into a summary and keep recent messages verbatim. That is unavoidable, and it is where sessions go wrong — the summary is produced at the moment of pressure, by a model, about a transcript that is about to leave the context.

Observational memory changes when the remembering happens. Instead of summarising in a panic at the end, it keeps a small ledger up to date while the session runs, and compaction then just folds that ledger.

flowchart LR
    A0["plain compaction"] --> A1["context fills"]
    A1 --> A2["a model summarises<br/>under pressure"]
    A2 --> A3["prose summary,<br/>no way back"]
    B0["with observational<br/>memory"] --> B1["context fills"]
    B1 --> B2["ledger written<br/>as you work"]
    B2 --> B3["compaction folds<br/>the ledger"]
    B3 --> B4["ids you can<br/>recall"]

Top row is pi on its own: one model call at the worst possible moment, detail chosen in a hurry, and the original wording gone from view. Bottom row is this image's default: the thinking happened earlier on a cheap model, the fold is deterministic, and every line in the result carries an id that resolves back to the exact source.

2. The mental model: three layers and a ledger

Layer What it is Example
Observation a timestamped, source-backed event from the conversation "user rejected option B because it needs a base rebuild"
Reflection a durable conclusion backed by observations "the user optimises for avoiding 67-minute rebuilds"
Drop a tombstone retiring an observation from active memory the superseded detail of a bug that is now fixed

These are appended to the session as silent ledger entries (om.observations.recorded, om.reflections.recorded, om.observations.dropped) and folded — replayed in order — to produce the memory state. The ledger is the source of truth; what you see in a compacted session is a rendering of it.

Two properties follow, and both matter later:

  • The ledger itself costs no context. Those entries are pi custom entries, which "do not participate in LLM context" (pi docs/session-format.md). They sit in the session file and reach the model only via the fold at compaction.
  • Memory is branch-local. A pi session is a tree (resume, fork), and the fold follows the current branch only, so a forked branch does not inherit another branch's view.

3. The lifecycle

Three background workers and one compaction hook, driven by raw token progress rather than wall-clock time. Defaults in brackets.

flowchart TD
    T(["turn_end"]) --> O{"10k raw tokens<br/>since observing?"}
    O -- yes --> OBS["<b>observer</b> runs"]
    O -- "no" --> R{"20k tokens<br/>since reflecting?"}
    R -- yes --> REF["<b>reflector</b> runs"]
    REF -- "if pool over 10k" --> DR["<b>dropper</b> prunes"]
    S(["agent_settled"]) --> C{"81k tokens<br/>since compacting?"}
    C -- yes --> CP["ctx.compact()"]
    CP --> H(["session_before_compact"])
    H --> F["fold the ledger<br/>no model call"]
    F --> VIS["compacted memory"]
  • observer — observeAfterTokens [10000]: writes observations for the conversation it has not covered yet.
  • reflector — reflectAfterTokens [20000]: promotes patterns across observations into durable reflections.
  • dropper — no clock of its own. It is post-reflection maintenance, gated on a successful same-turn reflection and an active pool above observationsPoolTargetTokens [10000]. Not a third worker on a third threshold.
  • compaction — compactAfterTokens [81000], checked when pi goes idle, so it never interrupts a turn. Pi will also compact on its own when the context is nearly full (contextTokens > contextWindow - reserveTokens, reserveTokens [16384]).

4. What compaction actually does to your context

This is the question the rest of the document used to leave hanging: if the old conversation is folded away, is the session back to knowing nothing?

No. Compaction replaces part of the context, not all of it, and it deletes nothing at all from disk.

flowchart LR
    SYS["system prompt<br/>+ AGENTS.md"] --> CTX["what the model sees<br/>on the next turn"]
    SUM["folded memory:<br/>reflections + observations"] --> CTX
    TAIL["recent turns,<br/>verbatim"] --> CTX
    DISK[("session .jsonl: all of it")] -. "recall(id)" .-> CTX

Where each piece comes from:

  • System prompt and AGENTS.md — never compacted, because they were never conversation. Pi rebuilds them from disk on every request (loadContextFileFromDir), so they cannot be lost by compaction.
  • The verbatim tail — sized by a token budget, not a message count. Pi walks backwards from the newest entry accumulating token estimates until keepRecentTokens [20000] is reached; that entry becomes firstKeptEntryId, and everything from there on is kept unchanged. Cut points land on turn boundaries, never mid-tool-call. So the most recent ~20k tokens of real work — your last instructions, the diffs, the test output — survive word for word.
  • The folded memory — replaces only what came before that cut. Rendered from the ledger's records: reflections and observations, each with its 12-hex id.
  • The session file — untouched. Compaction appends a compaction entry ({"type":"compaction", summary, firstKeptEntryId, tokensBefore, …}) and rebuilds context from it on later turns. Nothing is rewritten in place; the only documented way to remove session content is deleting the whole .jsonl.

That last point is what makes the answer to "is the detail gone?" no rather than mostly: recall does not read the context window at all. It calls sessionManager.getBranch() — the full branch from the root — and resolves an observation id back to the original entries. Detail that left the model's view an hour ago is still one recall away.

Repeated compaction does not summarise the summary. The rendered text is always built from live observation/reflection records, never from the previous compaction's prose, so there is no generation-loss spiral. (Mechanically the projection is incremental — it re-derives back to the last full-fold boundary and carries the rest forward, escalating to a genuine re-fold from the branch root when the observation pool reaches observationsPoolMaxTokens [20000].)

So the honest summary of the state after compaction: the model keeps its instructions, keeps recent work verbatim, trades older turns for a dense id-carrying digest of them, and can pull any of it back on demand. Not a fresh start — a smaller, cheaper, still-navigable one.

One caveat about "no model call"

If the ledger is empty — compaction fires before the observer has ever run — the hook returns nothing and declines ownership, and pi's own model-based summariser runs instead:

const summary = renderSummary(projection.reflections, projection.observations);
if (summary.length === 0) {
  // Decline ownership so Pi's native summarizer preserves the pre-cut context.
  return;
}

In steady state (any session old enough to have produced one observation) om's hook wins and compaction is model-free. "Never calls a model" is true in practice and false in principle; the fallback is deliberate, so an empty ledger degrades to normal pi rather than to no summary at all.

5. What you actually get

  • Compaction stops being a stall. In steady state the latency path is deterministic work over ledger entries, not a summarisation call.
  • Nothing important vanishes silently. Compaction is lossy by design, but every item keeps a 12-character id, and recall(<id>) returns the exact evidence — original wording, reasoning, file path, error text.
  • The bookkeeping runs on a cheaper model than your session. In this image that is deliberate and visible (§7): background workers on Haiku, session on Opus.
  • It is automatic. No habit to maintain, unlike the palace protocol — which is exactly why the two complement each other (§11).
  • Forks stay clean. Branch-local memory means a fork sub-agent's noise does not leak into the parent's folded memory.

6. recall is not a search tool

recall takes one specific 12-hex id that already appears in compacted memory or in /om:view. It cannot be given a topic. It can return an observation (marked active or dropped), or a reflection together with the observations supporting it.

sequenceDiagram
    participant M as compacted memory
    participant A as agent
    participant L as ledger
    M->>A: "[high] user rejected option B (a1b2c3d4e5f6)"
    A->>L: recall("a1b2c3d4e5f6")
    L-->>A: exact observation + source ids
    Note over A: acts on the original wording

The rule of thumb the agent skill uses: recall before a load-bearing action that rests on a compressed memory — shipping a change, asserting a fact, answering "why do you believe that". One recall is cheap; redoing finished work is not.

7. How it is wired in this image

flowchart TB
    IMG["baked in the image:<br/>v3.0.4 @ ce9fc98"] --> REG["settings.json<br/>packages[]"]
    REG --> SESS["your pi session"]
    SESS -- "your turns" --> SM["session model:<br/>Opus"]
    SESS -- "observer, reflector,<br/>dropper" --> WM["memory model:<br/>Haiku"]
    SESS -- "ledger entries" --> JL["session .jsonl"]
    JL --> VOL[("devbox-pi-config<br/>volume")]

Four consequences of that wiring:

  1. There is no separate database. Memory is entries inside the ordinary pi session file (~/.pi/agent/sessions/<project>/<timestamp>_<uuid>.jsonl). Nothing extra to back up, nothing to migrate.
  2. It survives container recreate, because ~/.pi is the devbox-pi-config named volume (docker-compose.yml) — the same one holding your pi config and session history.
  3. packages[] is the only source of truth for which copy is loaded. A clone at /workspace/pi-observational-memory may exist (and today matches /opt byte-for-byte at ce9fc98) — its presence proves nothing. To run a patched build you point packages[] at it explicitly and start a new session.
  4. The worker model is a deliberate choice, and it is yours to change. The seeded config sends background work to Haiku while your session runs Opus:
"observational-memory": {
  "model": { "provider": "amazon-bedrock", "id": "eu.anthropic.claude-haiku-4-5-20251001-v1:0" },
  "debugLog": false
}

8. What it costs

Resource Cost
Model calls up to three background calls per consolidation pass (observer, reflector, dropper), each capped at agentMaxTurns [16], on the configured memory model — not your session model
Latency in your turns none by construction: workers run from turn_end, compaction runs when pi is idle, and the fold itself does no model work
Disk negligible — JSON lines inside a session file that would exist anyway (measured here: ~/.pi/agent/sessions = 30 MB total, tens of om.* entries per session)
Context window zero until compaction. custom entries do not enter LLM context; only the folded summary does
Attention none once configured; there is no protocol for you or the agent to remember

If that is still more than you want on a given run, §9's passive switch turns off all proactive work while keeping recall and /om:* usable.

9. Configuration

Global: ~/.pi/agent/settings.json (persisted in the volume). Per project: <project>/.pi/settings.json, which overrides global. Precedence is project → global → environment, and the environment can only override passive.

Key Default What it changes
observeAfterTokens 10000 observer cadence — lower means smaller chunks and more calls
reflectAfterTokens 20000 reflector cadence (and thereby dropper opportunities)
observerChunkMaxTokens 20% of the memory model's context window, else 60000 cap on one observer run's input
compactAfterTokens 81000 when proactive auto-compaction fires
observationsPoolMaxTokens 20000 pool size at which compaction does a full re-fold from the branch root
observationsPoolTargetTokens half of max (10000) what the dropper aims back down to
agentMaxTurns 16 shared turn cap for the three workers
model unset → session model send background work to a cheaper/faster model
showWorkerNotifications true routine "observer ran" notices
passive false kill switch for all proactive background work; recall and /om:* still work
debugLog false per-session NDJSON trace at ~/.pi/agent/observational-memory/debug/<session-id>.ndjson

Pi's own compaction knobs live under a separate compaction key — keepRecentTokens [20000] sets the verbatim tail from §4, reserveTokens [16384] the headroom that triggers pi's own compaction.

One-off passive run, no config edit:

PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi

Invalid values are ignored rather than fatal, so a typo degrades to the default instead of breaking your session — which also means a typo is silent. Check with /om:status.

10. Confirming it is actually working

Do not infer health from the absence of a warning; look:

# 1. inside pi — the authoritative view
/om:status          # visible-vs-full drift, thresholds, worker state
/om:view            # what the agent currently sees
/om:view full       # full ledger truth at the branch tip

# 2. from a shell — are ledger entries being written, and has it compacted?
grep -o '"customType":"om\.[a-z.]*"' \
  "$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)" | sort | uniq -c
grep -c '"type":"compaction"' "$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)"

# 3. which copy is loaded, and at what commit
python3 -c "import json;print(json.load(open('$HOME/.pi/agent/settings.json'))['packages'])"
git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory rev-parse HEAD

Ledger entries are "type":"custom" with "customType":"om.…". Do not grep for custom_message — that is a different pi API for entries that do enter LLM context, used here by the MemPalace mailbox (customType: "mempalace-mailbox"), not by om.

11. It is not the same thing as MemPalace

Both are called "memory" and they solve different problems. Nothing is wrong with running both — this image does, and they cover each other's failure modes.

flowchart LR
    O0["observational<br/>memory"] --> O1["horizon:<br/>this session"]
    O1 --> O2["scope: one branch,<br/>one machine"]
    O2 --> O3["automatic"]
    O3 --> O4["retrieval:<br/>recall(id)"]
    P0["MemPalace"] --> P1["horizon: months,<br/>machines"]
    P1 --> P2["scope:<br/>the fleet"]
    P2 --> P3["protocol-driven"]
    P3 --> P4["retrieval:<br/>search, KG, mailbox"]
Question Answer
"What did we decide 200 turns ago in this session?" observational memory (and recall for the exact wording)
"What did we decide last month, or on another machine?" MemPalace (mempalace_search, diaries)
"What is true right now about version X?" MemPalace knowledge graph
"Does another machine need something from me?" MemPalace coordination log — see Cross-machine agent coordination
"Why is compaction not losing my session?" observational memory

The crisp version: observational memory keeps a session coherent; the palace keeps the fleet coherent. A container recreate wipes neither — but only because ~/.pi and the palace both live outside the container filesystem.

12. Gotchas

  • Branch-local means branch-local. Resuming or forking changes which ledger is folded. Memory that "disappeared" is usually on another branch.
  • recall needs an id, not a topic. If you only have a topic, that is a palace search, not a recall.
  • A /workspace clone is not evidence of what is loaded — see §7.3.
  • showWorkerNotifications: true is not proof of work; it reports runs, and an observer that deliberately emits nothing writes no ledger entry and simply retries after another observeAfterTokens.
  • A turn bigger than keepRecentTokens splits. The cut then lands mid-turn at an assistant message and pi merges two summaries — rare, but it is why a very large single turn can lose more verbatim detail than you would expect.
  • git log in the baked tree needs safe.directory (/opt is root-owned): git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory log.