AGENTS.md: delegating work — task first, fork second; the rule that survives compaction

The cheat-sheet listed Fork first (with "balanced … default: exploration/
impl/test") and Task as the exception, and it lost to the fork tool's
self-recommending description every time it mattered (five of five fork
briefs in one 2026-09-17 session carried "do not"). This file is the one
copy of the rule in the SYSTEM PROMPT — the skill is gone after the first
compaction — so it now carries: the one discriminator (what the child
sees), task as the default for anything that writes or must obey a rule,
fork only for read-only exploration against this conversation or parallel
opinions, a three-question pre-flight before any fork(...), a copy-paste
minimal task(...) call with the roots contract (write_allowed ⊆ roots, never
nested in a watched root), and a note that fork-gate now enforces the split.

README: pi-task is now also the `task` tool (pi-extensions task.ts) with
fork-gate alongside; where the wrapper looks for this script.
This commit is contained in:
2026-09-19 16:42:03 +02:00
parent adfb553f5c
commit 9c87ee843f
2 changed files with 62 additions and 19 deletions
+15
View File
@@ -92,6 +92,21 @@ wait out.
./bin/pi-task run examples/task-*.json # exit 0 = PASS, 1 = FAIL ./bin/pi-task run examples/task-*.json # exit 0 = PASS, 1 = FAIL
``` ```
**As a tool, not just a CLI (2026-09-19).** `pi-extensions/extensions/task.ts`
registers this runner as the `task` tool, and `fork-gate.ts` blocks a `fork`
whose brief carries a prohibition, a write boundary or a file-changing
imperative, handing back the `task(...)` call to make instead. Reason: the
correct prose rule ("pi-task for briefs with prohibitions") lost to the tool
list for months — `fork` was a tool with a self-recommending description, this
was a CLI to be remembered — and the skill that carried the rule is gone after
the first compaction. The tool wrapper also rejects, before spending a model
run, the two spec errors that make a violation certain (`write_allowed` not an
exact subset of `roots`; a writable root nested in a watched-only root), and
serialises tasks whose roots overlap. The wrapper looks for this script at
`$PI_TASK_BIN`, `/opt/pi-toolkit/bin/pi-task`, `~/src/pi-toolkit/bin/pi-task`,
`~/src/src_local/pi-toolkit/bin/pi-task`, `/workspace/pi-toolkit/bin/pi-task`,
then `$PATH`.
`selftest` is not decoration: it feeds the validator one known-good and six `selftest` is not decoration: it feeds the validator one known-good and six
known-bad envelopes plus a two-sided boundary check, and **aborts** if any pair known-bad envelopes plus a two-sided boundary check, and **aborts** if any pair
fails to discriminate. A validator that has only ever returned PASS has not been fails to discriminate. A validator that has only ever returned PASS has not been
+47 -19
View File
@@ -2,33 +2,61 @@
## Session start: load the pi-extensions skill ## Session start: load the pi-extensions skill
If the `fork` and/or `recall` tools are present in your tool list (you are If the `fork`, `task` and/or `recall` tools are present in your tool list (you
running inside the **pi** harness with the pi-fork / pi-observational-memory are running inside the **pi** harness with the pi-fork / pi-extensions /
packages), **read `~/.agents/skills/pi-extensions/SKILL.md` before doing any pi-observational-memory packages), **read `~/.agents/skills/pi-extensions/SKILL.md` before doing any
non-trivial work.** These extensions are routinely under-utilised when left to non-trivial work.** These extensions are routinely under-utilised when left to
on-demand description matching; reading the skill up front fixes that. on-demand description matching; reading the skill up front fixes that.
Core triggers (cheat-sheet — the skill has the full guidance): ## Delegating work: `task` first, `fork` second
Two ways to run a child agent. They differ in ONE thing, and it decides the
quality of the result: **what the child sees.**
- **`task`** (tool; wraps the `pi-task` CLI) — the child sees **only your spec**
(L0–L2: goal + named files + curated facts). Immutable spec, machine-checked
PASS/FAIL envelope, every root diffed before/after (a write outside
`write_allowed` FAILS), audit dir under `~/.pi/agent/pi-task/`. **Default for
any delegated work that changes files or must obey a rule.**
- **`fork`** (tool) — the child inherits your **entire branch** (L4) with the
brief as its last message. Measured on this fleet: in a long session the
parent narrative outweighs the brief — forks ignored prohibitions, invented
quotes, answered in the user's voice, and left no audit trail (their temp dir
is deleted on exit). **Use only for read-only exploration that needs this
conversation's context, or N parallel opinions on one question.**
Pre-flight before **any** `fork(...)` — one *yes* makes it a `task`:
1. Does the brief say *do not / only / never / must not*?
2. Will the child write, edit, commit or push anything?
3. Do I want a PASS/FAIL I can check, rather than prose?
The `fork` tool's own description recommends itself for "implementation"; it is
wrong about that, and `fork-gate` (a `tool_call` hook) will block a fork whose
brief trips the three questions and hand back the `task(...)` to make instead.
This paragraph is in the system prompt and survives compaction; the skill does
not — that is why the rule lives here.
Minimal `task` call (`pi-task schema` lists every field):
task(id="slug", goal="…verbatim, the child has NO other context…",
deliverable="…exact shape wanted…", effort="fast|balanced|deep",
read_only=false,
roots=["/abs/repo/docs", "/abs/repo/src"], # WATCHED, each diffed alone
write_allowed=["/abs/repo/docs"], # exact subset of roots; never nested in a watched root
facts=["verified fact"], files=["/abs/path/to/read"])
Result = the CLI report (verdict, problems, deliverable, evidence pointers,
audit dir). Spot-check the pointers: the envelope proves shape, not truth.
Tasks with overlapping roots run one at a time (the tool serialises them).
Tiers: `fast`=haiku (mechanical/lookups), `balanced`=sonnet (default),
`deep`=opus (architecture, security, ambiguous debugging).
CLI fallback if the tool is missing: `/opt/pi-toolkit/bin/pi-task run <spec.json>`.
- **Fork** (`fork(task=..., effort=fast|balanced|deep)`) when a subtask needs
reading many files you won't keep, runs in parallel, or is a well-scoped
one-shot whose detail would pollute the main thread. Tiers: `fast`=haiku
(mechanical/lookups), `balanced`=sonnet (default: exploration/impl/test),
`deep`=opus (architecture, security, ambiguous debugging). Always state
decision authority, pass verified context, specify the deliverable, ask for
an "unsure about" section. Don't fork trivial or iterative work.
- **Task** (`pi-task`, a CLI at `/opt/pi-toolkit/bin/pi-task` — *not* a tool in
your list, you invoke it with bash) when the child must NOT inherit your
session: any brief containing a prohibition, anything needing a machine-checked
pass/fail envelope, a durable audit trail, or write-boundary enforcement over
named roots. Same ladder, different rung: `fork` is **L4** (child sees your
whole branch), `pi-task` is **L0–L2** (nothing / named files / curated facts).
Pick the lowest rung that can do the job. **L3** (truncated branch) is not built.
- **Recall** (`recall(<12-char-hex-id>)`) before a load-bearing action (edit - **Recall** (`recall(<12-char-hex-id>)`) before a load-bearing action (edit
code, ship a change, assert a fact) that rests on a `[high]`/`[critical]` code, ship a change, assert a fact) that rests on a `[high]`/`[critical]`
observation or a reflection you did not produce this turn. The compaction observation or a reflection you did not produce this turn. The compaction
summary is lossy by design; one recall is cheap, redoing finished work is not. summary is lossy by design; one recall is cheap, redoing finished work is not.
Not a search tool — you must already have the ID. Not a search tool — you must already have the ID.
For depth on tier selection, fork-brief design, boundary discipline, and the For depth on the context ladder, brief design, boundary discipline, and the
observational-memory model, read the full skill. observational-memory model, read the full skill.