From adfb553f5c8a21e3705ff7902fd2f1c3b5a52808 Mon Sep 17 00:00:00 2001 From: "pi@mbp-m1-2020" Date: Tue, 8 Sep 2026 22:08:03 +0200 Subject: [PATCH] docs: which subtask mechanism, for the operator who will not read a skill MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The pi-task section already explained the tool in depth. What was missing was the decision: given a subtask, which mechanism, and why. Adds an end-user section built on the L0-L4 ladder — the child's context volume is the axis that explains nearly every observed good and bad behaviour — plus when to use neither. Carries the measured cost so the choice is priced, not guessed: 9 runs, $0.669 total, $0.0027 (fast, refused an over-budget spec in 2.5s) to $0.165 (balanced, 87s). Includes the jq one-liner, and the warning that a crashed run leaves no result.json — one of the ten here is exactly that, so a rollup must tolerate missing files rather than assume runs == directories. Records that fork spend is NOT in that tree (pi-fork aggregates from the parent's own toolResult entries into the status bar), so the two mechanisms report spend in two different places and nothing adds them up today. pi-global-AGENTS.md gets one cheat-sheet bullet so the choice is visible without loading a skill, and names pi-task as a CLI rather than a tool. Both files also carry the inverted capability-floor trap ([] = floor on, null = floor off). --- README.md | 89 +++++++++++++++++++++++++++++++++++++++++++++ pi-global-AGENTS.md | 7 ++++ 2 files changed, 96 insertions(+) diff --git a/README.md b/README.md index 0f61ba7..4748d63 100644 --- a/README.md +++ b/README.md @@ -161,6 +161,95 @@ Removes the keybindings and `AGENTS.md` symlinks, plus the shell-loader and `pi- --- +## Which mechanism for a subtask? The context ladder (L0–L4) + +Both `fork` and `pi-task` run a *second* pi as a child process. The difference that +matters is **how much of your session the child can see** — and that one choice +explains most of the good and bad behaviour observed so far. Five rungs, from +nothing to everything: + +| rung | what the child sees | how you get it | built? | +|---|---|---|---| +| **L0** | nothing but the goal | `pi-task` default — fresh `--session-id`, prompt is goal + deliverable | yes | +| **L1** | goal + the **names** of files it should read itself | `pi-task` spec `context.files` / `context.commands` | yes | +| **L2** | goal + an **excerpt you curated** | `pi-task` spec `context.facts`, pasted verbatim into the prompt | yes | +| **L3** | a **truncated tail** of your session | *not built* — needs a new spec key plus `--session ` | **no** | +| **L4** | your **entire** session branch | `fork(task=…)` — this is the only thing fork does | yes | + +**Pick the lowest rung that can still do the job.** Context is not free in either +direction: too little and the child re-derives what you already know; too much and +it starts finishing *your* pending work instead of its own. The 2026-07-29 case is +the cautionary one — a 4645-character brief with four explicit prohibitions was +overridden because the inherited transcript showed work in flight, and the child +resolved the conflict toward "finish the obvious thing". + +### When to use which + +Use **`fork`** (L4) when: + +- the subtask only makes sense against the current conversation ("does this fit + what we just decided?"); +- you want several **independent** opinions in parallel from a single message; +- it is read-only exploration whose detail you do not want to keep; and +- you will verify every load-bearing claim it returns anyway. + +Use **`pi-task`** (L0–L2) when: + +- the child must **not** inherit your intentions — in particular anything whose + brief contains a prohibition; +- you want a **pass/fail** answer rather than prose (the envelope either parses or + the run FAILED, however fluent the report); +- you need an **audit trail** afterwards — spec, exact prompt, argv, raw NDJSON, + before/after boundary snapshots; +- the task touches files and writes outside an authorised set must be **caught**; or +- you will run it again later and want the same spec to produce a comparable run. + +Use **neither** when the task is trivial (under ~30 seconds yourself), iterative +(both mechanisms are one-shot), or when the judgement needs context only you have. + +Neither rung buys you honesty. A fresh L0 context removes the *narrative* failures +— answering in your voice, inventing continuity — but it does not stop a child from +filling the `deliverable` slot when the task itself is under-specified. That was +measured directly: an adversarial spec built on a false premise still returned a +confident shape, and only the `unsure` field exposed it. Verify decisive claims +from the filesystem either way. + +### What it costs + +Measured here, 9 runs, 2026-09-07: **$0.67 total**, from $0.0027 (`fast`, correctly +refused an over-budget spec in 2.5 s) to $0.165 (`balanced`, 87 s, 5 tool calls). +`fast` runs land near $0.02, `balanced` near $0.11–0.16. Every run records its own +figure: + +```bash +jq -s 'map(.metrics.cost_usd)|{runs:length,total:add}' ~/.pi/agent/pi-task/*/result.json +``` + +A crashed run leaves its directory **without** `result.json` — one of the ten here +is exactly that — so a rollup must tolerate missing files instead of assuming +`runs == directories`. + +`fork` spend is *not* in that tree. pi-fork aggregates it live into pi's status bar +from the parent transcript's own `toolResult` entries, so fork children never exist +as separate session files. The two mechanisms therefore report spend in two +different places for two different reasons, and **no single view adds them up +today**. + +### One settings trap worth knowing + +`~/.pi/agent/settings.json` → `pi-fork.extensions: []` is what removes the +mempalace bridge from fork children, making palace writes impossible by +construction. The check in `pi-fork/src/runner.ts` is `if (extensions !== null)`, +so the semantics are inverted from intuition: + +- `[]` → `--no-extensions` is passed → floor **on** (what you want) +- `null` → nothing passed → floor **off**, forks can write to the palace again + +Because `null` is documented as the way to "restore normal extension loading", +changing `[]` to `null` as a tidy-up silently re-arms the thing that was +deliberately disarmed. `pi-task` hardcodes `--no-extensions`, so it cannot drift +this way. + ## Deploying pi on a new machine Full recipe from a clean macOS or Linux box to a working pi install. Follow in order. diff --git a/pi-global-AGENTS.md b/pi-global-AGENTS.md index b0aa97c..5a64984 100644 --- a/pi-global-AGENTS.md +++ b/pi-global-AGENTS.md @@ -17,6 +17,13 @@ Core triggers (cheat-sheet — the skill has the full guidance): `deep`=opus (architecture, security, ambiguous debugging). Always state decision authority, pass verified context, specify the deliverable, ask for an "unsure about" section. Don't fork trivial or iterative work. +- **Task** (`pi-task`, a CLI at `/opt/pi-toolkit/bin/pi-task` — *not* a tool in + your list, you invoke it with bash) when the child must NOT inherit your + session: any brief containing a prohibition, anything needing a machine-checked + pass/fail envelope, a durable audit trail, or write-boundary enforcement over + named roots. Same ladder, different rung: `fork` is **L4** (child sees your + whole branch), `pi-task` is **L0–L2** (nothing / named files / curated facts). + Pick the lowest rung that can do the job. **L3** (truncated branch) is not built. - **Recall** (`recall(<12-char-hex-id>)`) before a load-bearing action (edit code, ship a change, assert a fact) that rests on a `[high]`/`[critical]` observation or a reflection you did not produce this turn. The compaction