docs: which subtask mechanism, for the operator who will not read a skill

The pi-task section already explained the tool in depth. What was missing was
the decision: given a subtask, which mechanism, and why. Adds an end-user
section built on the L0-L4 ladder — the child's context volume is the axis that
explains nearly every observed good and bad behaviour — plus when to use
neither.

Carries the measured cost so the choice is priced, not guessed: 9 runs,
$0.669 total, $0.0027 (fast, refused an over-budget spec in 2.5s) to $0.165
(balanced, 87s). Includes the jq one-liner, and the warning that a crashed run
leaves no result.json — one of the ten here is exactly that, so a rollup must
tolerate missing files rather than assume runs == directories.

Records that fork spend is NOT in that tree (pi-fork aggregates from the
parent's own toolResult entries into the status bar), so the two mechanisms
report spend in two different places and nothing adds them up today.

pi-global-AGENTS.md gets one cheat-sheet bullet so the choice is visible
without loading a skill, and names pi-task as a CLI rather than a tool.

Both files also carry the inverted capability-floor trap ([] = floor on,
null = floor off).
This commit is contained in:
2026-09-08 22:08:03 +02:00
parent 02af927f26
commit adfb553f5c
2 changed files with 96 additions and 0 deletions
+89
View File
@@ -161,6 +161,95 @@ Removes the keybindings and `AGENTS.md` symlinks, plus the shell-loader and `pi-
---
## Which mechanism for a subtask? The context ladder (L0–L4)
Both `fork` and `pi-task` run a *second* pi as a child process. The difference that
matters is **how much of your session the child can see** — and that one choice
explains most of the good and bad behaviour observed so far. Five rungs, from
nothing to everything:
| rung | what the child sees | how you get it | built? |
|---|---|---|---|
| **L0** | nothing but the goal | `pi-task` default — fresh `--session-id`, prompt is goal + deliverable | yes |
| **L1** | goal + the **names** of files it should read itself | `pi-task` spec `context.files` / `context.commands` | yes |
| **L2** | goal + an **excerpt you curated** | `pi-task` spec `context.facts`, pasted verbatim into the prompt | yes |
| **L3** | a **truncated tail** of your session | *not built* — needs a new spec key plus `--session <trimmed snapshot>` | **no** |
| **L4** | your **entire** session branch | `fork(task=…)` — this is the only thing fork does | yes |
**Pick the lowest rung that can still do the job.** Context is not free in either
direction: too little and the child re-derives what you already know; too much and
it starts finishing *your* pending work instead of its own. The 2026-07-29 case is
the cautionary one — a 4645-character brief with four explicit prohibitions was
overridden because the inherited transcript showed work in flight, and the child
resolved the conflict toward "finish the obvious thing".
### When to use which
Use **`fork`** (L4) when:
- the subtask only makes sense against the current conversation ("does this fit
what we just decided?");
- you want several **independent** opinions in parallel from a single message;
- it is read-only exploration whose detail you do not want to keep; and
- you will verify every load-bearing claim it returns anyway.
Use **`pi-task`** (L0–L2) when:
- the child must **not** inherit your intentions — in particular anything whose
brief contains a prohibition;
- you want a **pass/fail** answer rather than prose (the envelope either parses or
the run FAILED, however fluent the report);
- you need an **audit trail** afterwards — spec, exact prompt, argv, raw NDJSON,
before/after boundary snapshots;
- the task touches files and writes outside an authorised set must be **caught**; or
- you will run it again later and want the same spec to produce a comparable run.
Use **neither** when the task is trivial (under ~30 seconds yourself), iterative
(both mechanisms are one-shot), or when the judgement needs context only you have.
Neither rung buys you honesty. A fresh L0 context removes the *narrative* failures
— answering in your voice, inventing continuity — but it does not stop a child from
filling the `deliverable` slot when the task itself is under-specified. That was
measured directly: an adversarial spec built on a false premise still returned a
confident shape, and only the `unsure` field exposed it. Verify decisive claims
from the filesystem either way.
### What it costs
Measured here, 9 runs, 2026-09-07: **$0.67 total**, from $0.0027 (`fast`, correctly
refused an over-budget spec in 2.5 s) to $0.165 (`balanced`, 87 s, 5 tool calls).
`fast` runs land near $0.02, `balanced` near $0.11–0.16. Every run records its own
figure:
```bash
jq -s 'map(.metrics.cost_usd)|{runs:length,total:add}' ~/.pi/agent/pi-task/*/result.json
```
A crashed run leaves its directory **without** `result.json` — one of the ten here
is exactly that — so a rollup must tolerate missing files instead of assuming
`runs == directories`.
`fork` spend is *not* in that tree. pi-fork aggregates it live into pi's status bar
from the parent transcript's own `toolResult` entries, so fork children never exist
as separate session files. The two mechanisms therefore report spend in two
different places for two different reasons, and **no single view adds them up
today**.
### One settings trap worth knowing
`~/.pi/agent/settings.json` → `pi-fork.extensions: []` is what removes the
mempalace bridge from fork children, making palace writes impossible by
construction. The check in `pi-fork/src/runner.ts` is `if (extensions !== null)`,
so the semantics are inverted from intuition:
- `[]` → `--no-extensions` is passed → floor **on** (what you want)
- `null` → nothing passed → floor **off**, forks can write to the palace again
Because `null` is documented as the way to "restore normal extension loading",
changing `[]` to `null` as a tidy-up silently re-arms the thing that was
deliberately disarmed. `pi-task` hardcodes `--no-extensions`, so it cannot drift
this way.
## Deploying pi on a new machine
Full recipe from a clean macOS or Linux box to a working pi install. Follow in order.
+7
View File
@@ -17,6 +17,13 @@ Core triggers (cheat-sheet — the skill has the full guidance):
`deep`=opus (architecture, security, ambiguous debugging). Always state
decision authority, pass verified context, specify the deliverable, ask for
an "unsure about" section. Don't fork trivial or iterative work.
- **Task** (`pi-task`, a CLI at `/opt/pi-toolkit/bin/pi-task` — *not* a tool in
your list, you invoke it with bash) when the child must NOT inherit your
session: any brief containing a prohibition, anything needing a machine-checked
pass/fail envelope, a durable audit trail, or write-boundary enforcement over
named roots. Same ladder, different rung: `fork` is **L4** (child sees your
whole branch), `pi-task` is **L0–L2** (nothing / named files / curated facts).
Pick the lowest rung that can do the job. **L3** (truncated branch) is not built.
- **Recall** (`recall(<12-char-hex-id>)`) before a load-bearing action (edit
code, ship a change, assert a fact) that rests on a `[high]`/`[critical]`
observation or a reflection you did not produce this turn. The compaction