docs(rfc-002): run the census for real — the replay surface is 176 records, and closets were missing
Ran Phase A's classifier against the frozen EMB-7KJ4VR4G archive. Since the fleet shares one devbox image this is a reasonable prior for tor-ms22 and MBP-M1-2020: mined (source_file set) 15949 98.9% re-mine, never replay diary entries 116 0.7% replay + §7.6 suffix skip agent-authored drawers 60 0.4% replay, idempotent KG open facts 34 -- server guard dedupes KG closed facts 0 -- nothing to do Two consequences that shrink this project: the replay-only surface is 176 records, not thousands, so the writer is a small job and the §7.6 diary guard treated as the blocker governs 116 records; and there are ZERO closed KG facts, so the unguarded-closed-fact gap is real in the code but empty in the data. filed_at spread in the same archive -- 12 in May, 52 in June, 15174 in July, 887 in August -- is the concrete argument for the history-preserving regime. Two corrections to the first draft: 1. CLOSETS were omitted entirely. mempalace_closets is a second Chroma collection (1560 rows, ~10% of the palace) and there is NO MCP tool that writes one -- mcp_server.py exposes only _purge_source_closets, and closet_llm.py says outright that regex closets are always created by the miner. So closets cannot be replayed even in principle; they return only by re-mining. Same category as hallways/known_entities/palace-graph, and it resolves itself for the ~99% that gets re-mined anyway. 2. Census gotcha: chroma.sqlite3 holds BOTH collections, so an embeddings-wide query over-counts by ~10%. Must join segments->collections and keep mempalace_drawers. Doing so reconciles exactly: 14,778 = the 14,777 seeded to the primary + the one known post-snapshot chunk. Also drops the earlier "438 of 14,829" figure, which conflated central's no_source count (which includes its own later agent writes) with the archive's actual replay surface.
This commit is contained in:
+43
-8
@@ -64,11 +64,24 @@ Chunking is deterministic and replay-safe: `DEFAULT_CHUNK_SIZE = 800` (`config.p
|
||||
|
||||
### Not carried by either regime
|
||||
|
||||
**Hallways, `known_entities.json`, and the palace graph are built at mine time**, not by `add_drawer`.
|
||||
Replayed drawers therefore arrive with no hallway/co-occurrence edges, so `list_hallways` and `traverse`
|
||||
will under-report for joined content. These are derived artifacts — rebuildable by re-mining, or
|
||||
acceptably degraded. Decide which; do not discover it later. (`hallways.py`, `entities.py`,
|
||||
`palace_graph.py` — none referenced from the `add_drawer` write path.)
|
||||
**Hallways, `known_entities.json`, the palace graph, and closets are built at mine time**, not by
|
||||
`add_drawer`. Replayed drawers therefore arrive with no hallway/co-occurrence edges and no closets, so
|
||||
`list_hallways` and `traverse` will under-report for joined content, and joined drawers get no
|
||||
`closet_boost` in search ranking. All are derived artifacts — rebuilt by re-mining, or acceptably
|
||||
degraded. Decide which; do not discover it later.
|
||||
|
||||
**Closets specifically:** `mempalace_closets` is a second Chroma collection (1,560 rows in the archive,
|
||||
~10% of the palace), and **there is no MCP tool that writes a closet** — `mcp_server.py` only exposes
|
||||
`_purge_source_closets`. `closet_llm.py` states it plainly: *"Regex closets are always created by the
|
||||
miner"*, with the LLM path an opt-in regeneration afterwards. So closets cannot be replayed even in
|
||||
principle; they come back only by re-mining. Since ~99% of a devbox palace is mined content that gets
|
||||
re-mined on synlig anyway (§4), this resolves itself for the bulk.
|
||||
|
||||
> **Census gotcha — filter by collection.** `chroma.sqlite3` holds *both* collections. A naive
|
||||
> `embeddings`-wide query over-counts: the archive yields 16,338 rows total, which is
|
||||
> `mempalace_drawers` 14,778 + `mempalace_closets` 1,560. Join through
|
||||
> `segments`→`collections` and keep `mempalace_drawers`, or the census inflates by ~10%. (14,778 also
|
||||
> reconciles exactly with the 14,777 seeded to the primary plus the one known post-snapshot chunk.)
|
||||
|
||||
---
|
||||
|
||||
@@ -101,13 +114,33 @@ contract. Two backends behind one interface:
|
||||
## 4. Explicitly out of scope
|
||||
|
||||
- **Mined wings.** RFC 001 §5 makes them `local`; their ids are path-dependent, so replay produces
|
||||
duplicates rather than dedup. Re-mine on synlig from sources present there. This is also the bulk by
|
||||
count — on the primary, only **438 of 14,829** drawers have no `source_file`, so the genuinely
|
||||
replay-only surface is hundreds, not tens of thousands.
|
||||
duplicates rather than dedup. Re-mine on synlig from sources present there.
|
||||
- **Merging two palaces into a third.** Every join targets the existing primary.
|
||||
- **`mempalace sync` interaction.** See RFC 001 §7.2 — settle the guard separately; a joiner must never
|
||||
call it.
|
||||
|
||||
### 4.1 The census, run for real — the replay surface is tiny
|
||||
|
||||
Phase A's classifier was run against the frozen EMB-7KJ4VR4G archive (2026-08-15). Since the fleet
|
||||
shares one devbox image, this is a reasonable prior for what tor-ms22 and MBP-M1-2020 hold:
|
||||
|
||||
| Class | Count | Share | Joiner action |
|
||||
| --- | --- | --- | --- |
|
||||
| Mined (`source_file` set) | 15,949 | **98.9%** | re-mine on synlig; **never replay** |
|
||||
| Diary entries | 116 | 0.7% | replay + §7.6 suffix skip |
|
||||
| Agent-authored drawers | 60 | 0.4% | replay, idempotent by content id |
|
||||
| KG open facts | 34 | — | replay, server guard dedupes |
|
||||
| KG **closed** facts | **0** | — | nothing to do — see below |
|
||||
|
||||
**Two consequences that shrink this project sharply.** First, the genuinely replay-only surface is
|
||||
**176 records**, not thousands — so Phase C's writer is a small job, and the §7.6 diary guard that has
|
||||
been treated as the blocker governs *116 records*. Second, **there are zero closed KG facts**, so the
|
||||
unguarded-closed-fact gap (§2) is real in the code but currently empty in the data; it needs handling for
|
||||
correctness, not for this join.
|
||||
|
||||
`filed_at` spread in the same archive — 12 records in May, 52 in June, 15,174 in July, 887 in August — is
|
||||
the concrete case for regime A: an MCP replay would restamp all of it to the join date.
|
||||
|
||||
## 5. Open decisions for ALC
|
||||
|
||||
1. **Chronology: keep it or flatten it?** Regime A keeps it and costs a service stop plus an rsync of the
|
||||
@@ -145,3 +178,5 @@ mempalace 3.6.0, `/opt/uv-tools/mempalace/lib/python3.13/site-packages/mempalace
|
||||
| Chunk size 800 | `config.py:274` |
|
||||
| Direct-write precedent preserving ids+metadata | `migrate.py:326-330` |
|
||||
| No join/replay/import tooling exists | `cli.py` subcommands; `exporter.py` emits markdown (lossy, not a join primitive); `diary_ingest.py` ingests *daily-summary files*, unrelated to agent diaries; `migrate.py` is single-palace chromadb recovery; `dedup.py` is cosine-similarity pruning within one `source_file` |
|
||||
| Closets are miner-derived, no MCP write path | `closet_llm.py:8-10` ("Regex closets are always created by the miner"); `mcp_server.py:2913` exposes only `_purge_source_closets` |
|
||||
| Two collections in one sqlite file | `mempalace_drawers` (14,778 rows) and `mempalace_closets` (1,560) in the archive, via `segments`→`collections` |
|
||||
|
||||
Reference in New Issue
Block a user