Files
pi-devbox/docs/observational-memory.md
T
joakimp a2846a5f7e
Lint / hadolint (push) Successful in 11s
Lint / actionlint (push) Successful in 21s
Publish Docker Image / resolve-versions (push) Successful in 19s
Publish Docker Image / base-decide (push) Successful in 16s
Publish Docker Image / build-base (push) Successful in 59m39s
Publish Docker Image / smoke-studio (push) Successful in 5m22s
Publish Docker Image / smoke (push) Successful in 17m53s
Publish Docker Image / build-variant-studio (push) Successful in 17m14s
Publish Docker Image / build-variant (push) Successful in 18m22s
Publish Docker Image / promote-base-latest (push) Successful in 12s
Publish Docker Image / update-description (push) Successful in 20s
release: adopt pi 0.84.4 + pi-atelier v0.10.0, and fix the doc claim the pi bump invalidates
pi 0.84.3 -> 0.84.4 (no Breaking Changes / Removed heading in that section,
grepped). Adopted for three fixes that land on machinery this fleet runs:
#6879 (large tool results crossing the auto-compaction threshold were sent to
the provider before compacting), #8345 (a resumed session corrupted its next
appended entry when the JSONL lacked a trailing newline -- that file is the
memory feeder's input; measured 49/49 clean here beforehand), and #8537
(triggerTurn:false messages sent mid-run were inserted between a tool call and
its result). The mempalace mailbox is outside #8537's precondition: it delivers
at agent_settled with deliverAs:"steer" and no triggerTurn, and 0.84.4 leaves
the documented steer semantics unchanged.

pi-atelier v0.8.2 -> v0.10.0: two minor releases, both UI-only, no BREAKING
notice. v0.9.0 raises its minimum pi to 0.84.0 and, unlike the
0.7.1-under-pi-0.84 startup-hang precedent, encodes it in peerDependencies
(>=0.84.0). Satisfied by PI_VERSION=0.84.4. Both executable floors compare with
sort -V, so 0.10.0 >= 0.7.1 evaluates correctly.

docs/observational-memory.md: pi's own compaction.md gained one paragraph in
0.84.4 -- autoCompact is now also checked mid-run, after a tool batch's results
are appended. Our text said compaction is checked only when pi goes idle and so
"never interrupts a turn"; that was only ever true of the OM trigger. The
section now states both entry points into session_before_compact and the
diagram carries the second edge (mermaid checker re-run: 6 blocks, 44 labels,
0 soft-wrapped, no cut glyphs at 1280px and 800px).

README: the version-pin table had been wrong since v1.8.6 -- 93f986e moved
ARG PI_VERSION to 0.84.3 and MEMPALACE_VERSION to 3.8.0 and neither table row,
so it advertised pi 0.84.2 / mempalace 3.7.1. Corrected, plus the
--expected-version example that would now fail against a 0.84.4 image.

CHANGELOG: Unreleased retitled v1.8.12 (2026-08-31) with the audits above and a
dependency-audit table -- every other component measured SAME (skillset
snapshot --check OK at a12fe5e, 0 commits since baked).
2026-08-31 07:08:41 +02:00

19 KiB

Observational memory — why this image has it, and what it does for you

Audience: anyone using this container for long pi sessions who has wondered what recall, /om:status and "compacted memory" are, or whether they should leave any of it switched on.

Companion documents: the extension ships its own reference docs at /opt/pi-observational-memory/docs/ — concepts.md (the model), how-it-works.md (hooks and internals) and configuration.md (every setting). Pi's own compaction mechanics are in /usr/lib/node_modules/@earendil-works/pi-coding-agent/docs/compaction.md. Those are normative; this document is the deployment view — what is pinned here, how it is wired, what it costs, and how it differs from MemPalace. For the palace, see mempalace-toolkit/docs/fleet-memory.md.

Verified on pi-devbox v1.8.9 (release_tag v1.8.9, source aac4a1c), which bakes pi-observational-memory v3.0.4 at commit ce9fc98 — the value in /etc/pi-devbox/build-manifest.json → components.pi-observational-memory. Every number below was read from that tree, from pi's own docs, or from the live container. The pi-side mechanics were first read at pi 0.84.3 and re-checked at 0.84.4 (v1.8.12), which moved one of them — see §3.


1. The problem it solves

A long pi session outgrows the model's context window. Pi's answer is compaction: fold the older part of the conversation into a summary and keep recent messages verbatim. That is unavoidable, and it is where sessions go wrong — the summary is produced at the moment of pressure, by a model, about a transcript that is about to leave the context.

Observational memory changes when the remembering happens. Instead of summarising in a panic at the end, it keeps a small ledger up to date while the session runs, and compaction then just folds that ledger.

flowchart LR
    A0["plain compaction"] --> A1["context fills"]
    A1 --> A2["a model summarises<br/>under pressure"]
    A2 --> A3["prose summary,<br/>no way back"]
    B0["with observational<br/>memory"] --> B1["context fills"]
    B1 --> B2["ledger written<br/>as you work"]
    B2 --> B3["compaction folds<br/>the ledger"]
    B3 --> B4["ids you can<br/>recall"]

Top row is pi on its own: one model call at the worst possible moment, detail chosen in a hurry, and the original wording gone from view. Bottom row is this image's default: the thinking happened earlier on a cheap model, the fold is deterministic, and every line in the result carries an id that resolves back to the exact source.

2. The mental model: three layers and a ledger

Layer What it is Example
Observation a timestamped, source-backed event from the conversation "user rejected option B because it needs a base rebuild"
Reflection a durable conclusion backed by observations "the user optimises for avoiding 67-minute rebuilds"
Drop a tombstone retiring an observation from active memory the superseded detail of a bug that is now fixed

These are appended to the session as silent ledger entries (om.observations.recorded, om.reflections.recorded, om.observations.dropped) and folded — replayed in order — to produce the memory state. The ledger is the source of truth; what you see in a compacted session is a rendering of it.

Two properties follow, and both matter later:

  • The ledger itself costs no context. Those entries are pi custom entries, which "do not participate in LLM context" (pi docs/session-format.md). They sit in the session file and reach the model only via the fold at compaction.
  • Memory is branch-local. A pi session is a tree (resume, fork), and the fold follows the current branch only, so a forked branch does not inherit another branch's view.

3. The lifecycle

Three background workers and one compaction hook, driven by raw token progress rather than wall-clock time. Defaults in brackets.

flowchart TD
    T(["turn_end"]) --> O{"10k raw tokens<br/>since observing?"}
    O -- yes --> OBS["<b>observer</b> runs"]
    O -- "no" --> R{"20k tokens<br/>since reflecting?"}
    R -- yes --> REF["<b>reflector</b> runs"]
    REF -- "if pool over 10k" --> DR["<b>dropper</b> prunes"]
    S(["agent_settled"]) --> C{"81k tokens<br/>since compacting?"}
    C -- yes --> CP["ctx.compact()"]
    CP --> H(["session_before_compact"])
    A(["pi autoCompact<br/>idle, or mid-run<br/>after a tool batch"]) --> H
    H --> F["fold the ledger<br/>no model call"]
    F --> VIS["compacted memory"]
  • observer — observeAfterTokens [10000]: writes observations for the conversation it has not covered yet.
  • reflector — reflectAfterTokens [20000]: promotes patterns across observations into durable reflections.
  • dropper — no clock of its own. It is post-reflection maintenance, gated on a successful same-turn reflection and an active pool above observationsPoolTargetTokens [10000]. Not a third worker on a third threshold.
  • compaction — compactAfterTokens [81000], checked at agent_settled, so this trigger never interrupts a turn. Pi will also compact on its own when the context is nearly full (contextTokens > contextWindow - reserveTokens, reserveTokens [16384]), and from pi 0.84.4 that check also runs mid-run — after a tool batch's results are appended, before the next assistant response, skipped only when the batch ends the run and no queued message needs another response. So session_before_compact has two entry points and the second one can fire inside a turn. Harmless for the fold itself, which makes no model call, but worth stating plainly: "never interrupts a turn" was only ever true of the observational-memory trigger, and reads as a promise about pi's.

4. What compaction actually does to your context

This is the question the rest of the document used to leave hanging: if the old conversation is folded away, is the session back to knowing nothing?

No. Compaction replaces part of the context, not all of it, and it deletes nothing at all from disk.

flowchart LR
    SYS["system prompt<br/>+ AGENTS.md"] --> CTX["what the model sees<br/>on the next turn"]
    SUM["folded memory:<br/>reflections + observations"] --> CTX
    TAIL["recent turns,<br/>verbatim"] --> CTX
    DISK[("session .jsonl: all of it")] -. "recall(id)" .-> CTX

Where each piece comes from:

  • System prompt and AGENTS.md — never compacted, because they were never conversation. Pi rebuilds them from disk on every request (loadContextFileFromDir), so they cannot be lost by compaction.
  • The verbatim tail — sized by a token budget, not a message count. Pi walks backwards from the newest entry accumulating token estimates until keepRecentTokens [20000] is reached; that entry becomes firstKeptEntryId, and everything from there on is kept unchanged. Cut points land on turn boundaries, never mid-tool-call. So the most recent ~20k tokens of real work — your last instructions, the diffs, the test output — survive word for word.
  • The folded memory — replaces only what came before that cut. Rendered from the ledger's records: reflections and observations, each with its 12-hex id.
  • The session file — untouched. Compaction appends a compaction entry ({"type":"compaction", summary, firstKeptEntryId, tokensBefore, …}) and rebuilds context from it on later turns. Nothing is rewritten in place; the only documented way to remove session content is deleting the whole .jsonl.

That last point is what makes the answer to "is the detail gone?" no rather than mostly: recall does not read the context window at all. It calls sessionManager.getBranch() — the full branch from the root — and resolves an observation id back to the original entries. Detail that left the model's view an hour ago is still one recall away.

Repeated compaction does not summarise the summary. The rendered text is always built from live observation/reflection records, never from the previous compaction's prose, so there is no generation-loss spiral. (Mechanically the projection is incremental — it re-derives back to the last full-fold boundary and carries the rest forward, escalating to a genuine re-fold from the branch root when the observation pool reaches observationsPoolMaxTokens [20000].)

So the honest summary of the state after compaction: the model keeps its instructions, keeps recent work verbatim, trades older turns for a dense id-carrying digest of them, and can pull any of it back on demand. Not a fresh start — a smaller, cheaper, still-navigable one.

One caveat about "no model call"

If the ledger is empty — compaction fires before the observer has ever run — the hook returns nothing and declines ownership, and pi's own model-based summariser runs instead:

const summary = renderSummary(projection.reflections, projection.observations);
if (summary.length === 0) {
  // Decline ownership so Pi's native summarizer preserves the pre-cut context.
  return;
}

In steady state (any session old enough to have produced one observation) om's hook wins and compaction is model-free. "Never calls a model" is true in practice and false in principle; the fallback is deliberate, so an empty ledger degrades to normal pi rather than to no summary at all.

5. What you actually get

  • Compaction stops being a stall. In steady state the latency path is deterministic work over ledger entries, not a summarisation call.
  • Nothing important vanishes silently. Compaction is lossy by design, but every item keeps a 12-character id, and recall(<id>) returns the exact evidence — original wording, reasoning, file path, error text.
  • The bookkeeping runs on a cheaper model than your session. In this image that is deliberate and visible (§7): background workers on Haiku, session on Opus.
  • It is automatic. No habit to maintain, unlike the palace protocol — which is exactly why the two complement each other (§11).
  • Forks stay clean. Branch-local memory means a fork sub-agent's noise does not leak into the parent's folded memory.

6. recall is not a search tool

recall takes one specific 12-hex id that already appears in compacted memory or in /om:view. It cannot be given a topic. It can return an observation (marked active or dropped), or a reflection together with the observations supporting it.

sequenceDiagram
    participant M as compacted memory
    participant A as agent
    participant L as ledger
    M->>A: "[high] user rejected option B (a1b2c3d4e5f6)"
    A->>L: recall("a1b2c3d4e5f6")
    L-->>A: exact observation + source ids
    Note over A: acts on the original wording

The rule of thumb the agent skill uses: recall before a load-bearing action that rests on a compressed memory — shipping a change, asserting a fact, answering "why do you believe that". One recall is cheap; redoing finished work is not.

7. How it is wired in this image

flowchart TB
    IMG["baked in the image:<br/>v3.0.4 @ ce9fc98"] --> REG["settings.json<br/>packages[]"]
    REG --> SESS["your pi session"]
    SESS -- "your turns" --> SM["session model:<br/>Opus"]
    SESS -- "observer, reflector,<br/>dropper" --> WM["memory model:<br/>Haiku"]
    SESS -- "ledger entries" --> JL["session .jsonl"]
    JL --> VOL[("devbox-pi-config<br/>volume")]

Four consequences of that wiring:

  1. There is no separate database. Memory is entries inside the ordinary pi session file (~/.pi/agent/sessions/<project>/<timestamp>_<uuid>.jsonl). Nothing extra to back up, nothing to migrate.
  2. It survives container recreate, because ~/.pi is the devbox-pi-config named volume (docker-compose.yml) — the same one holding your pi config and session history.
  3. packages[] is the only source of truth for which copy is loaded. A clone at /workspace/pi-observational-memory may exist (and today matches /opt byte-for-byte at ce9fc98) — its presence proves nothing. To run a patched build you point packages[] at it explicitly and start a new session.
  4. The worker model is a deliberate choice, and it is yours to change. The seeded config sends background work to Haiku while your session runs Opus:
"observational-memory": {
  "model": { "provider": "amazon-bedrock", "id": "eu.anthropic.claude-haiku-4-5-20251001-v1:0" },
  "debugLog": false
}

8. What it costs

Resource Cost
Model calls up to three background calls per consolidation pass (observer, reflector, dropper), each capped at agentMaxTurns [16], on the configured memory model — not your session model
Latency in your turns none by construction: workers run from turn_end, compaction runs when pi is idle, and the fold itself does no model work
Disk negligible — JSON lines inside a session file that would exist anyway (measured here: ~/.pi/agent/sessions = 30 MB total, tens of om.* entries per session)
Context window zero until compaction. custom entries do not enter LLM context; only the folded summary does
Attention none once configured; there is no protocol for you or the agent to remember

If that is still more than you want on a given run, §9's passive switch turns off all proactive work while keeping recall and /om:* usable.

9. Configuration

Global: ~/.pi/agent/settings.json (persisted in the volume). Per project: <project>/.pi/settings.json, which overrides global. Precedence is project → global → environment, and the environment can only override passive.

Key Default What it changes
observeAfterTokens 10000 observer cadence — lower means smaller chunks and more calls
reflectAfterTokens 20000 reflector cadence (and thereby dropper opportunities)
observerChunkMaxTokens 20% of the memory model's context window, else 60000 cap on one observer run's input
compactAfterTokens 81000 when proactive auto-compaction fires
observationsPoolMaxTokens 20000 pool size at which compaction does a full re-fold from the branch root
observationsPoolTargetTokens half of max (10000) what the dropper aims back down to
agentMaxTurns 16 shared turn cap for the three workers
model unset → session model send background work to a cheaper/faster model
showWorkerNotifications true routine "observer ran" notices
passive false kill switch for all proactive background work; recall and /om:* still work
debugLog false per-session NDJSON trace at ~/.pi/agent/observational-memory/debug/<session-id>.ndjson

Pi's own compaction knobs live under a separate compaction key — keepRecentTokens [20000] sets the verbatim tail from §4, reserveTokens [16384] the headroom that triggers pi's own compaction.

One-off passive run, no config edit:

PI_OBSERVATIONAL_MEMORY_PASSIVE=1 pi

Invalid values are ignored rather than fatal, so a typo degrades to the default instead of breaking your session — which also means a typo is silent. Check with /om:status.

10. Confirming it is actually working

Do not infer health from the absence of a warning; look:

# 1. inside pi — the authoritative view
/om:status          # visible-vs-full drift, thresholds, worker state
/om:view            # what the agent currently sees
/om:view full       # full ledger truth at the branch tip

# 2. from a shell — are ledger entries being written, and has it compacted?
grep -o '"customType":"om\.[a-z.]*"' \
  "$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)" | sort | uniq -c
grep -c '"type":"compaction"' "$(ls -t ~/.pi/agent/sessions/*/*.jsonl | head -1)"

# 3. which copy is loaded, and at what commit
python3 -c "import json;print(json.load(open('$HOME/.pi/agent/settings.json'))['packages'])"
git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory rev-parse HEAD

Ledger entries are "type":"custom" with "customType":"om.…". Do not grep for custom_message — that is a different pi API for entries that do enter LLM context, used here by the MemPalace mailbox (customType: "mempalace-mailbox"), not by om.

11. It is not the same thing as MemPalace

Both are called "memory" and they solve different problems. Nothing is wrong with running both — this image does, and they cover each other's failure modes.

flowchart LR
    O0["observational<br/>memory"] --> O1["horizon:<br/>this session"]
    O1 --> O2["scope: one branch,<br/>one machine"]
    O2 --> O3["automatic"]
    O3 --> O4["retrieval:<br/>recall(id)"]
    P0["MemPalace"] --> P1["horizon: months,<br/>machines"]
    P1 --> P2["scope:<br/>the fleet"]
    P2 --> P3["protocol-driven"]
    P3 --> P4["retrieval:<br/>search, KG, mailbox"]
Question Answer
"What did we decide 200 turns ago in this session?" observational memory (and recall for the exact wording)
"What did we decide last month, or on another machine?" MemPalace (mempalace_search, diaries)
"What is true right now about version X?" MemPalace knowledge graph
"Does another machine need something from me?" MemPalace coordination log — see Cross-machine agent coordination
"Why is compaction not losing my session?" observational memory

The crisp version: observational memory keeps a session coherent; the palace keeps the fleet coherent. A container recreate wipes neither — but only because ~/.pi and the palace both live outside the container filesystem.

12. Gotchas

  • Branch-local means branch-local. Resuming or forking changes which ledger is folded. Memory that "disappeared" is usually on another branch.
  • recall needs an id, not a topic. If you only have a topic, that is a palace search, not a recall.
  • A /workspace clone is not evidence of what is loaded — see §7.3.
  • showWorkerNotifications: true is not proof of work; it reports runs, and an observer that deliberately emits nothing writes no ledger entry and simply retries after another observeAfterTokens.
  • A turn bigger than keepRecentTokens splits. The cut then lands mid-turn at an assistant message and pi merges two summaries — rare, but it is why a very large single turn can lose more verbatim detail than you would expect.
  • git log in the baked tree needs safe.directory (/opt is root-owned): git -c safe.directory=/opt/pi-observational-memory -C /opt/pi-observational-memory log.