docs: write down the coordination channel the fleet already runs on
The logstream has carried cross-machine work since 2026-08-18 — patch handoff, review, a v1->v2 supersede — and nothing in this repo said it existed. That gap had a measurable cost this morning: another host addressed a retraction to pi@tor-ms22 by name and it was read only because the human said "read the logstream", while the agent was actively rebuilding the thing it warned about. Split by what each document is authoritative for, so there is one copy of each claim rather than three that drift: - README § Cross-machine agent coordination — what the CONTAINER needs. MEMPALACE_REMOTE_URL selects the shared palace; MEMPALACE_PI_DEVICE is what makes this machine reachable, because where every host is a thin client of one palace the stamped agent name is the only thing that distinguishes them. Stated as a rule with teeth: set both or neither, since a container missing the device var can read the log but is addressable by nobody. - AGENTS.md release checklist step 2 — the vendored-snapshot refresh, as a MECHANISM in the document a releasing agent actually reads, not a comment hoping to be noticed. It says the refresh costs a base rebuild, that skipping it is legitimate (every enrolled host reads its live clone), and that skipping it silently is not. - CHANGELOG — the three-way split itself, plus the measurement that shaped the ack contract: unfiltered, the mailbox returned 5 events, 4 of them finished broadcasts from eight days earlier; with status="open", exactly the 1 that needed an answer. Norms live in the skillset skill (82a8d3c, already live on every host that mounts the skillset — no rebuild) and mechanism in mempalace-toolkit's extensions/pi/README.md (e70bef2, which also documents the edge stamper that 553d8657 shipped undocumented). Deliberately NOT duplicated here. Consequence recorded rather than hidden: the skill edit lands in the skillset, so this repo's SKILLSET_SNAPSHOT_REF now honestly reports itself behind, and --check exits 1 with "has moved to 82a8d3c; the snapshot describes the older c04cd15". That message is also fixed in this commit — it previously blamed "the working tree" even when the tree was clean and only the ref had moved, which is the same defect class as a canary pinned to a phrase the release deleted: a message that names the wrong cause. Now distinguishes moved-HEAD from dirty-tree, verified against both plus the in-sync case.
This commit is contained in:
@@ -64,8 +64,22 @@ re-brand of opencode-devbox's `pi-only` variant.
|
||||
(`curl -sf 'https://registry.npmjs.org/@earendil-works%2Fpi-coding-agent/latest' | jq -r .version`).
|
||||
Check release notes at https://github.com/earendil-works/pi/releases for
|
||||
the upstream changelog to include in `CHANGELOG.md`.
|
||||
2. Update `CHANGELOG.md` Unreleased → vX.Y.Z section.
|
||||
3. Verify `docker compose up` works locally with the current `latest` image
|
||||
2. **Refresh the vendored mempalace skill snapshot if the skillset moved:**
|
||||
`scripts/vendor-mempalace-skill.sh --check` (reads a real skillset clone,
|
||||
writes nothing). Exit 1 means either the recorded `SKILLSET_SNAPSHOT_REF`
|
||||
does not describe the shipped bytes, or upstream has moved past it — the
|
||||
message distinguishes the two. Refresh with
|
||||
`scripts/vendor-mempalace-skill.sh`, which rewrites the file **and** the ARG
|
||||
together so they cannot drift apart.
|
||||
Two consequences to accept deliberately: the snapshot is hashed into
|
||||
`base_tag`, so refreshing costs a base rebuild (~67 min); and if the section
|
||||
the phrase canary names has changed, re-pin it in `scripts/smoke-test.sh`.
|
||||
**Skipping this is legitimate** — every enrolled host reads its own live
|
||||
skillset clone, so the baked copy is a no-mount fallback. What is *not*
|
||||
legitimate is skipping it silently: the drift is visible in
|
||||
`pi-devbox-version` and in the manifest, so decide rather than forget.
|
||||
3. Update `CHANGELOG.md` Unreleased → vX.Y.Z section.
|
||||
4. Verify `docker compose up` works locally with the current `latest` image
|
||||
if you're upgrading users from a previous version. Then run the
|
||||
**post-recreate sanity check** inside the running container to confirm
|
||||
persisted volumes survived and the pi runtime wiring re-deployed (not just
|
||||
@@ -73,8 +87,8 @@ re-brand of opencode-devbox's `pi-only` variant.
|
||||
`docker compose exec devbox bash scripts/recreate-sanity-check.sh --expected-version X.Y.Z`
|
||||
(or just `pi-devbox-sanity --expected-version X.Y.Z` if `cli_utils/bin` is
|
||||
on PATH). This is the runtime peer of the build-time `smoke-test.sh` gate.
|
||||
4. Push tag: `git tag vX.Y.Z && git push origin vX.Y.Z`.
|
||||
5. Watch CI: smoke job builds amd64 only and asserts size + extensions +
|
||||
5. Push tag: `git tag vX.Y.Z && git push origin vX.Y.Z`.
|
||||
6. Watch CI: smoke job builds amd64 only and asserts size + extensions +
|
||||
pi version + new-base-tooling presence. Variant build is multi-arch
|
||||
(amd64 + arm64) only after smoke passes. A tag push fires **only**
|
||||
`docker-publish.yml` — `lint.yml` is scoped to `branches: ['**']`, which
|
||||
@@ -85,9 +99,9 @@ re-brand of opencode-devbox's `pi-only` variant.
|
||||
discovery on `head_sha` **and** the workflow `path` — see *Gitea API access*
|
||||
below — because that guard costs nothing and a future workflow added on `v*`
|
||||
would silently reintroduce the ambiguity.
|
||||
6. Verify the Hub tags appear (latest + vX.Y.Z, the `-studio` pair, plus
|
||||
7. Verify the Hub tags appear (latest + vX.Y.Z, the `-studio` pair, plus
|
||||
base-latest if the base was rebuilt this run).
|
||||
7. **Revoke any short-lived Gitea PAT** used during the release at
|
||||
8. **Revoke any short-lived Gitea PAT** used during the release at
|
||||
`gitea.jordbo.se/user/settings/applications`. N/A if you used the
|
||||
`GITEA_ACCESS_TOKEN` env var instead (see *Gitea API access* below) —
|
||||
its lifecycle is managed host-side, nothing to revoke.
|
||||
|
||||
@@ -187,6 +187,45 @@ the skillset actually is — a maintainer's clone, or any running container.
|
||||
`pi-devbox-version` degrades quietly on such an image: no fingerprint, no
|
||||
annotation, verified against the real v1.8.7 manifest.
|
||||
|
||||
### Documented
|
||||
|
||||
- **The fleet's cross-machine coordination, which was working and unwritten.**
|
||||
The RFC 003 logstream has carried real work between hosts since 2026-08-18 —
|
||||
patch handoff, design review, a v1→v2 supersede — and no document in this repo
|
||||
or the toolkit said so. Written up in three places, split by what each is
|
||||
authoritative for:
|
||||
- `README.md` § *Cross-machine agent coordination* — what the **container**
|
||||
needs: `MEMPALACE_REMOTE_URL` selects the shared palace, and
|
||||
`MEMPALACE_PI_DEVICE` is what makes this machine *reachable* on the log,
|
||||
because when every host is a thin client of one palace the stamped agent name
|
||||
is the only thing distinguishing them. Set both or neither: a container
|
||||
without the device var can read the log but is addressable by nobody.
|
||||
- the skillset's `mempalace` skill (`82a8d3c`, live on every host that mounts
|
||||
the skillset, no rebuild needed) — the **norms**: a mailbox query at wake-up,
|
||||
and the sender-declared ack contract, where a *directed* event with
|
||||
`status="open"` is owed a reply and a `*` broadcast owes nothing. Measured
|
||||
while designing it: an unfiltered mailbox returned 5 events, 4 of them
|
||||
finished broadcasts from eight days earlier, where `status="open"` returned
|
||||
exactly the 1 that needed an answer — an unfiltered mailbox trains you to
|
||||
ignore it, so the filter is the feature.
|
||||
- mempalace-toolkit `extensions/pi/README.md` (`e70bef2`) — the **mechanism**,
|
||||
including that the bridge is *write-only* today (it stamps events going out
|
||||
and never reads the log, so nothing in this image polls on the agent's
|
||||
behalf), and that live SSE push is a palace-deployment question: the server
|
||||
implements `GET /logstream/stream`, but a reverse proxy exposing only `/mcp`
|
||||
makes it unreachable — verified by 404s against the real endpoint.
|
||||
|
||||
⚠️ **This makes the vendored snapshot stale on purpose.** The skill edit is in
|
||||
the skillset (`82a8d3c`), so `SKILLSET_SNAPSHOT_REF` still records `c04cd15`
|
||||
and `scripts/vendor-mempalace-skill.sh --check` now exits 1 with
|
||||
*"has moved to 82a8d3c; the snapshot describes the older c04cd15"*. Refreshing
|
||||
it is a deliberate release-day decision, not an oversight — hence the new
|
||||
step 2 in `AGENTS.md` § *Release-day checklist*, which states both that the
|
||||
refresh costs a base rebuild and that skipping it is legitimate because every
|
||||
enrolled host reads its live clone. What is not legitimate is skipping it
|
||||
*silently*, which is precisely what the new manifest fields and
|
||||
`pi-devbox-version` output make impossible.
|
||||
|
||||
---
|
||||
|
||||
## v1.8.7 — 2026-08-25
|
||||
|
||||
@@ -555,6 +555,32 @@ session/docs mining; the 29 MCP tools (search, kg-query, drawer-add,
|
||||
diary-write, etc.) are wired into pi automatically by the pi-extensions
|
||||
mempalace bridge.
|
||||
|
||||
### Cross-machine agent coordination
|
||||
|
||||
When `MEMPALACE_REMOTE_URL` points at a *shared* palace, the container gets more
|
||||
than shared search: it joins an append-only coordination log (RFC 003) that other
|
||||
machines' agents can address it on — used here for design review, patch handoff
|
||||
and retraction between hosts.
|
||||
|
||||
Two container-side settings make it work:
|
||||
|
||||
| Variable | Why it matters |
|
||||
|---|---|
|
||||
| `MEMPALACE_REMOTE_URL` | selects the shared palace; unset means a purely local palace, and the log then contains only this machine's own events |
|
||||
| `MEMPALACE_PI_DEVICE` | the bridge stamps `pi@<device>` as the writer, which is the **only** way the log can tell two machines apart when both are thin clients of one palace |
|
||||
|
||||
So a container with no `MEMPALACE_PI_DEVICE` can read the log but is not
|
||||
reachable *on* it: messages addressed to a bare `pi` match nobody. Set both, or
|
||||
neither.
|
||||
|
||||
What the agent is expected to *do* with this lives in the mempalace skill
|
||||
(`~/.agents/skills/mempalace/SKILL.md`) — the mailbox query at wake-up, and the
|
||||
convention that a directed event with `status="open"` is a request owed a reply
|
||||
while a `*` broadcast owes nothing. The mechanism side (what the bridge stamps,
|
||||
and why live SSE push depends on the palace deployment's reverse proxy rather
|
||||
than on this image) is documented in the toolkit's `extensions/pi/README.md`.
|
||||
Nothing in this image polls the log on the agent's behalf.
|
||||
|
||||
## Agent skills
|
||||
|
||||
pi discovers skills under `~/.agents/skills/`. Two delivery paths feed that
|
||||
|
||||
@@ -118,8 +118,16 @@ if [ "$MODE" = "check" ]; then
|
||||
printf 'OK: the vendored snapshot is exactly skillset@%s:%s\n' "${recorded:0:7}" "$REL_PATH"
|
||||
fi
|
||||
if [ "$vendored_sha" != "$upstream_sha" ]; then
|
||||
printf 'STALE: the working tree of %s differs from the snapshot (HEAD %s)\n' \
|
||||
# Name the ACTUAL cause. "working tree differs" is wrong when the tree is
|
||||
# clean and the ref simply moved on — a message that names the wrong cause is
|
||||
# the same defect class as a canary pinned to a deleted phrase.
|
||||
if [ "$recorded" != "$head_sha" ] && [ "$blob_sha" = "$upstream_sha" ]; then
|
||||
printf 'STALE: %s has moved to %s; the snapshot describes the older %s\n' \
|
||||
"$ROOT" "${head_sha:0:7}" "${recorded:0:7}" >&2
|
||||
else
|
||||
printf 'STALE: the working tree of %s differs from the snapshot (HEAD %s)\n' \
|
||||
"$ROOT" "${head_sha:0:7}" >&2
|
||||
fi
|
||||
rc=1
|
||||
fi
|
||||
exit "$rc"
|
||||
|
||||
Reference in New Issue
Block a user