diff --git a/CHANGELOG.md b/CHANGELOG.md index b344ae0..74d3563 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -13,6 +13,94 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`). ## Unreleased +**The rule "use `pi-task`, not `fork`, for a brief that carries a prohibition" was +written in the global `AGENTS.md` and in the pi-extensions skill, and it lost to +the `fork` tool's own description anyway.** Measured on tor-ms22, 2026-09-17: five +of five fork briefs in one session carried "do not"; one returned verbatim quotes +that did not exist in the source, and four had disjoint write boundaries that +fork cannot enforce — re-run as `pi-task`, all four passed their envelope. The +image now ships the rule where the decision is *made*, not where it is read about: + +- **`fork-gate.ts`** (pi-extensions `25c1265`): a `tool_call` hook that blocks a + `fork` whose brief contains a prohibition (*do not / never / only …*), a write + boundary (*only touch / read-only / stay within …*) or a clause-initial + file-changing imperative (*Edit …, Commit …, Fix …*). The block reason the model + reads **is** the `task(...)` call to make instead. It matches wording, not + intent, and says so. `PI_FORK_GATE=off` logs instead of blocking. +- **`task.ts`** (same commit): `pi-task` registered as the `task` tool, with the + decision rule in its description and in `promptGuidelines` — which pi appends + to the **system prompt**, the one place compaction cannot remove it from. + Before spending a model run it rejects the two spec errors that make a boundary + violation certain (`write_allowed` not an *exact subset* of `roots` — pi-task + keys deltas by root string; a writable root nested inside a watched-only root, + whose porcelain would change every time) and serialises sibling tasks whose + roots overlap (parallel siblings saw each other's writes as violations, + 2026-09-17). A FAIL verdict is a *result*; only the CLI refusing to run is an + error. +- **`pi-global-AGENTS.md`** (pi-toolkit `9c87ee8`): the delegation section now + opens "`task` first, `fork` second" — the one discriminator (what the child + sees), a three-question pre-flight before any `fork(...)`, a copy-paste minimal + call with the roots contract. This file is in the system prompt; the skill is + not, which is why the rule lives here. +- **Skill floor refreshed** to `25c1265` (`check-skill-floor.sh` OK, tree + `9b85a633…`); mirror in skillset `debc8f6`. + +Why prose failed, as mechanisms: `fork` is a **tool** — its self-recommending +description ("exploration, implementation, testing, review…") is in the tool +list on every turn and survives compaction; the skill is gone after the first +compaction; `pi-task` was a CLI to be remembered and reached through `bash` with a +hand-written JSON spec. The asymmetry *widens* in exactly the long sessions where +fork is worst, and fork deletes its temp dir on exit, so its failures were found +only by re-verifying the narrative. Third time on this fleet that a rule held in +prose and violated in practice was fixed by moving it into a hook +(`check-secrets`, `check-egl-only`, now delegation). + +Evidence, two-sided: `test/fork-gate.test.mjs` pins 15 must-block briefs +(including the real shapes) and 10 must-pass (including *"Write a summary…"*, +*"Report which files were modified…"*, *"Give me an update…"* — verbs a naive +list misfires on); the shipped classifier over the five real briefs from the +motivating session redirects 5/5. Live in `pi -p`: the fork was intercepted +before any child spawned (no `/tmp/pi-fork-*` directory), `task` returned PASS in +4 s / $0.014 with an evidence pointer and audit dir, a `usd=0.000001` budget came +back as a FAIL *result* (`isError=false`), and a nested-root spec was rejected +with no audit dir created. + +Nothing to wire in this repo: `PI_EXTENSIONS_REF=main` floats and +`pi-extensions/install.sh` symlinks every `extensions/*.ts` on container start, so +the next build ships both. A container running today has neither — +`~/.pi/agent/extensions/` links into `/opt/pi-extensions`, which is the baked ref. + +--- + +**`[mempalace ext] feed (tick) failed: mempalace remote request 'tools/call' +failed: timed out after 60000ms` is the same event as the `mine timed out after +30000ms` message the 2026-09 toolkit fix addressed, one deadline further down — +and that fix was incomplete.** mempalace-toolkit `817b3a8` (2026-09-18; v1.9.2 +baked `dab989b`; the floating `MEMPALACE_TOOLKIT_REF=main` picks it up on the next +build). `MEMPALACE_FEED_MINE_TIMEOUT_MS` had been raised to 300 000 but was only +*raced* against `client.callTool("mempalace_mine")`; `callTool()` had no way to +carry a deadline, so every mine went out under the transport's generic +per-request timeout — `MEMPALACE_MCP_TIMEOUT_MS`, 60 000, the value the "Stall +protection" comment in `Dockerfile.base` documents — which fired first on every +honest 60 s+ mine on the shared single-writer palace. The 300 s was unreachable. +Over HTTP nothing is lost (the mine continues server-side and is idempotent); +over stdio it was worse than noise — that transport **kills the child** on +timeout, so there the mine really was aborted at 60 s. + +Fix: `callTool(name, args, { timeoutMs })` on both transports, the feed passes +its own deadline down, plain calls keep 60 s (a *query* that slow is wedged; the +race stays as the liveness guard for a transport with its timeout disabled). +`scripts/test-mcp-call-timeout.sh` cuts `RemoteMcpClient` out of the shipped +file, drives it against a local JSON-RPC server that delays `tools/call`, and +asserts three things — a plain call rejects at the generic deadline, the override +outlives it, the override is itself a deadline: 2 of 6 fail on `dab989b`, 6 of 6 +pass on `817b3a8`. `test-owed-withdrawal.sh` (17) and `check-mcp-client-sync.sh` +stay green; sync token `v1` untouched, since nothing in the protocol changed. +Reading for the fleet: on a fixed build that message means a mine exceeded +*five* minutes — look at palace size or a competing writer, not at the timeout. + +--- + **`scripts/recreate-sanity-check.sh` asserted that `/tmp/sshcm` exists while every `ssh` in the container was dying `rc=255`, and it was right to — it was checking the directory the *image* creates, and the breakage was in the directory a diff --git a/Dockerfile.base b/Dockerfile.base index 071939b..e732c9a 100644 --- a/Dockerfile.base +++ b/Dockerfile.base @@ -452,7 +452,10 @@ RUN ARCH=$(case "${TARGETARCH}" in amd64) echo "x86_64" ;; arm64) echo "aarch64" # uninterruptibly. A stall-kill is no longer a permanent latch either: the # next tool call respawns the server with capped exponential backoff (the # budget resets on any successful response). Tunables: -# MEMPALACE_MCP_TIMEOUT_MS (default 60000), MEMPALACE_MCP_INIT_TIMEOUT_MS +# MEMPALACE_MCP_TIMEOUT_MS (default 60000; the feed's `mempalace_mine` carries +# its own longer MEMPALACE_FEED_MINE_TIMEOUT_MS, default 300000, since toolkit +# 817b3a8 — before that the 60 s deadline cut every honest mine off), +# MEMPALACE_MCP_INIT_TIMEOUT_MS # (default 300000 — generous so a genuine first cold-open isn't killed), # MEMPALACE_MCP_MAX_RESPAWNS (default 2; 0 disables self-heal), # MEMPALACE_MCP_RESPAWN_BACKOFF_MS (default 1000); timeouts of 0 disable.