docs: which subtask mechanism, for the operator who will not read a skill

The pi-task section already explained the tool in depth. What was missing was
the decision: given a subtask, which mechanism, and why. Adds an end-user
section built on the L0-L4 ladder — the child's context volume is the axis that
explains nearly every observed good and bad behaviour — plus when to use
neither.

Carries the measured cost so the choice is priced, not guessed: 9 runs,
$0.669 total, $0.0027 (fast, refused an over-budget spec in 2.5s) to $0.165
(balanced, 87s). Includes the jq one-liner, and the warning that a crashed run
leaves no result.json — one of the ten here is exactly that, so a rollup must
tolerate missing files rather than assume runs == directories.

Records that fork spend is NOT in that tree (pi-fork aggregates from the
parent's own toolResult entries into the status bar), so the two mechanisms
report spend in two different places and nothing adds them up today.

pi-global-AGENTS.md gets one cheat-sheet bullet so the choice is visible
without loading a skill, and names pi-task as a CLI rather than a tool.

Both files also carry the inverted capability-floor trap ([] = floor on,
null = floor off).
This commit is contained in:
2026-09-08 22:08:03 +02:00
parent 02af927f26
commit adfb553f5c
2 changed files with 96 additions and 0 deletions
+89
View File
@@ -161,6 +161,95 @@ Removes the keybindings and `AGENTS.md` symlinks, plus the shell-loader and `pi-
--- ---
## Which mechanism for a subtask? The context ladder (L0–L4)
Both `fork` and `pi-task` run a *second* pi as a child process. The difference that
matters is **how much of your session the child can see** — and that one choice
explains most of the good and bad behaviour observed so far. Five rungs, from
nothing to everything:
| rung | what the child sees | how you get it | built? |
|---|---|---|---|
| **L0** | nothing but the goal | `pi-task` default — fresh `--session-id`, prompt is goal + deliverable | yes |
| **L1** | goal + the **names** of files it should read itself | `pi-task` spec `context.files` / `context.commands` | yes |
| **L2** | goal + an **excerpt you curated** | `pi-task` spec `context.facts`, pasted verbatim into the prompt | yes |
| **L3** | a **truncated tail** of your session | *not built* — needs a new spec key plus `--session <trimmed snapshot>` | **no** |
| **L4** | your **entire** session branch | `fork(task=…)` — this is the only thing fork does | yes |
**Pick the lowest rung that can still do the job.** Context is not free in either
direction: too little and the child re-derives what you already know; too much and
it starts finishing *your* pending work instead of its own. The 2026-07-29 case is
the cautionary one — a 4645-character brief with four explicit prohibitions was
overridden because the inherited transcript showed work in flight, and the child
resolved the conflict toward "finish the obvious thing".
### When to use which
Use **`fork`** (L4) when:
- the subtask only makes sense against the current conversation ("does this fit
what we just decided?");
- you want several **independent** opinions in parallel from a single message;
- it is read-only exploration whose detail you do not want to keep; and
- you will verify every load-bearing claim it returns anyway.
Use **`pi-task`** (L0–L2) when:
- the child must **not** inherit your intentions — in particular anything whose
brief contains a prohibition;
- you want a **pass/fail** answer rather than prose (the envelope either parses or
the run FAILED, however fluent the report);
- you need an **audit trail** afterwards — spec, exact prompt, argv, raw NDJSON,
before/after boundary snapshots;
- the task touches files and writes outside an authorised set must be **caught**; or
- you will run it again later and want the same spec to produce a comparable run.
Use **neither** when the task is trivial (under ~30 seconds yourself), iterative
(both mechanisms are one-shot), or when the judgement needs context only you have.
Neither rung buys you honesty. A fresh L0 context removes the *narrative* failures
— answering in your voice, inventing continuity — but it does not stop a child from
filling the `deliverable` slot when the task itself is under-specified. That was
measured directly: an adversarial spec built on a false premise still returned a
confident shape, and only the `unsure` field exposed it. Verify decisive claims
from the filesystem either way.
### What it costs
Measured here, 9 runs, 2026-09-07: **$0.67 total**, from $0.0027 (`fast`, correctly
refused an over-budget spec in 2.5 s) to $0.165 (`balanced`, 87 s, 5 tool calls).
`fast` runs land near $0.02, `balanced` near $0.11–0.16. Every run records its own
figure:
```bash
jq -s 'map(.metrics.cost_usd)|{runs:length,total:add}' ~/.pi/agent/pi-task/*/result.json
```
A crashed run leaves its directory **without** `result.json` — one of the ten here
is exactly that — so a rollup must tolerate missing files instead of assuming
`runs == directories`.
`fork` spend is *not* in that tree. pi-fork aggregates it live into pi's status bar
from the parent transcript's own `toolResult` entries, so fork children never exist
as separate session files. The two mechanisms therefore report spend in two
different places for two different reasons, and **no single view adds them up
today**.
### One settings trap worth knowing
`~/.pi/agent/settings.json` → `pi-fork.extensions: []` is what removes the
mempalace bridge from fork children, making palace writes impossible by
construction. The check in `pi-fork/src/runner.ts` is `if (extensions !== null)`,
so the semantics are inverted from intuition:
- `[]` → `--no-extensions` is passed → floor **on** (what you want)
- `null` → nothing passed → floor **off**, forks can write to the palace again
Because `null` is documented as the way to "restore normal extension loading",
changing `[]` to `null` as a tidy-up silently re-arms the thing that was
deliberately disarmed. `pi-task` hardcodes `--no-extensions`, so it cannot drift
this way.
## Deploying pi on a new machine ## Deploying pi on a new machine
Full recipe from a clean macOS or Linux box to a working pi install. Follow in order. Full recipe from a clean macOS or Linux box to a working pi install. Follow in order.
+7
View File
@@ -17,6 +17,13 @@ Core triggers (cheat-sheet — the skill has the full guidance):
`deep`=opus (architecture, security, ambiguous debugging). Always state `deep`=opus (architecture, security, ambiguous debugging). Always state
decision authority, pass verified context, specify the deliverable, ask for decision authority, pass verified context, specify the deliverable, ask for
an "unsure about" section. Don't fork trivial or iterative work. an "unsure about" section. Don't fork trivial or iterative work.
- **Task** (`pi-task`, a CLI at `/opt/pi-toolkit/bin/pi-task` — *not* a tool in
your list, you invoke it with bash) when the child must NOT inherit your
session: any brief containing a prohibition, anything needing a machine-checked
pass/fail envelope, a durable audit trail, or write-boundary enforcement over
named roots. Same ladder, different rung: `fork` is **L4** (child sees your
whole branch), `pi-task` is **L0–L2** (nothing / named files / curated facts).
Pick the lowest rung that can do the job. **L3** (truncated branch) is not built.
- **Recall** (`recall(<12-char-hex-id>)`) before a load-bearing action (edit - **Recall** (`recall(<12-char-hex-id>)`) before a load-bearing action (edit
code, ship a change, assert a fact) that rests on a `[high]`/`[critical]` code, ship a change, assert a fact) that rests on a `[high]`/`[critical]`
observation or a reflection you did not produce this turn. The compaction observation or a reflection you did not produce this turn. The compaction