Compare commits

...

2 Commits

Author SHA1 Message Date
joakimp b615571913 changelog: release v1.8.11 — the mail nobody could deliver, and the ship that skipped a file
Lint / hadolint (push) Successful in 13s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 19s
Publish Docker Image / base-decide (push) Successful in 10s
Publish Docker Image / build-base (push) Successful in 53m20s
Publish Docker Image / smoke (push) Successful in 4m48s
Publish Docker Image / smoke-studio (push) Successful in 4m58s
Publish Docker Image / build-variant-studio (push) Successful in 17m0s
Publish Docker Image / build-variant (push) Successful in 26m43s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 24s
Names three things riding in via floating refs, since the rule this
fleet adopted (v1.8.9) is to name a behaviour change before tagging,
not after:

- mempalace-toolkit b2b50af -> 21023e7: the rsync --checksum ship fix
  (a real behaviour change to every remote-mode client) and RFC 003
  §7.13 (docs only).
- skillset 6eb20af -> a12fe5e: mermaid-diagrams cutU normalisation and
  the mempalace skill's from_agent identity rule, plus the vendored
  fallback snapshot re-pinned to match (this pi-devbox commit).

Dependency audit run by direct command against every component, not
assumed: pi/mempalace/pi-atelier/pi-studio/pi-toolkit/pi-extensions/
pi-fork/pi-observational-memory all confirmed unchanged since v1.8.10.
pi-studio's apparent SHA drift on ls-remote turned out to be an
annotated-tag-object hash vs its underlying commit hash, not real
drift -- checked with ^{commit} before writing it down as 'none'.
2026-08-27 23:39:39 +02:00
joakimp 495b7e3859 vendor: resync mempalace skill snapshot to skillset a12fe5e
scripts/vendor-mempalace-skill.sh, real refresh not --check: skillset
moved 6eb20af -> a12fe5e (mermaid-diagrams cutU normalisation, and the
from_agent identity rule this same release ships in RFC 003). --check
reported stale-but-truthful (exit 0, the sanctioned skip) but this
release's point is getting today's fixes live fleet-wide, and the base
rebuild is already forced by the entrypoint change and the floating
mempalace-toolkit ref moving -- so the incremental cost of also
bumping this pin is zero. Phrase canary in scripts/smoke-test.sh
unaffected: neither pinned phrase ('Provenance is stamped for you'
present, 'Attribute what you file yourself' absent) is in the section
that changed; both verified still correct in the new snapshot.
2026-08-27 23:36:03 +02:00
3 changed files with 160 additions and 3 deletions
+99 -1
View File
@@ -11,7 +11,7 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
---
## Unreleased
## v1.8.11 — 2026-08-27
**Shell state that the writable layer eats on every recreate now gets rebuilt at
start.** Two additions to `entrypoint-user.sh`, both idempotent, both silent
@@ -78,6 +78,104 @@ the v1.8.0 mistake of a smoke assertion written against a stage that does not
exist at run time. `workflow_dispatch` with `smoke_only` remains the way to
exercise this against `HEAD` before a tag.
**Also carried by the floating `mempalace-toolkit` main ref** (resolved at build
time, not by a pi-devbox commit — `MEMPALACE_TOOLKIT_REF=main`):
**A scrubbed re-export of a dormant session could silently never reach the
palace host.** `bin/mempalace-pi-session` ships to the palace with
`rsync -a --update`, and the stage file's mtime is deliberately the SOURCE
transcript's mtime (`os.utime()`, "preserve session mtime for dedup
stability"). Re-exporting a session that has not been appended to since its
last ship therefore produces a mtime that is *not newer* than the receiver's —
exactly the case a redactor upgrade needs to ship, since content differs while
mtime does not. `--update` reported success and sent nothing. Found and
patched by `pi@mbp-m1-2020` (mempalace-toolkit `a361b71`): `--update` →
`--checksum`, which compares content and ignores size/mtime entirely.
Dropping `--update` outright was considered and rejected — rsync's default
quick check already transfers on a size difference alone, which would have
masked the *next* instance of this (a redaction whose placeholder happens to
match the secret's length) as fixed. `os.utime()` is untouched; its backdating
is a separate, load-bearing design call for dedup stability. New regression
test, `scripts/test-rsync-ship-idempotency.sh`, runs fully offline (a local
rsync destination exercises the same size/mtime/checksum comparison as the ssh
transfer) and is built to *discriminate*: it must fail against `--update` and
pass against `--checksum,` not merely exercise the code path — the first draft
of the test used fixture strings of different lengths and passed for the wrong
reason (rsync's quick check transfers on size difference alone regardless of
`--update`), which is the same trap the patch itself was written to avoid.
**Acceptance line for this class of change going forward:** "receiver sha256
matches sender for every staged file", not "local stage is clean" — a clean
local stage says nothing about what a dormant session already sent.
**An event addressed to an identity no session runs as is delivered to
nobody, and this fleet has now hit it three separate ways.** RFC 003 gains
§7.13 and open-decision 10 (mempalace-toolkit `21023e7`, docs only, no image
behaviour change): the owed-set derivation — the log's only push channel — is
keyed on `to_agent`, and a reply is always addressed back to whatever string
the *original writer* put in `from_agent`. Nothing validates that string
against a live session identity, so authoring under a synthetic or foreign
name makes every reply to that event write-only. Measured cost this cycle: a
directed ask planted under a synthetic sender drew a correct reply containing
an urgent security finding, and it sat unread for ~2h20m, found only because a
human asked whether mail had arrived. Permitted exception, unchanged: a
synthetic sender is fine for a deliberate control experiment, provided the
body names the real identity to reply to.
**Also carried by the live `skillset` mount** (each device's own clone, not
baked — except the `mempalace` skill's fallback snapshot, re-vendored below):
**The mermaid-diagrams checker's cut gate moved from client pixels to a
per-SVG user-space unit.** `CUT_PX` was calibrated against one live page at
one render scale; sweeping `--viewport` 500→1600 on an *unchanged* document
moved the worst overflow −1.0px → −3.0px, near-proportional to the viewport,
i.e. a constant geometric overflow viewed through a changing scale. `cutU =
cutPx / scale` (scale taken per-SVG, never a page average — one page mixes
scales 0.643–0.988) recovers that invariant: the sweep now collapses to
exactly −3.0u at every width. Re-deriving the threshold against the live host
surfaced a real false negative the old pixel gate had: a label at
`cutPx=0.4, scale=0.678` read as healthy under `CUT_PX=0.5` but is `0.59u` —
a genuine cut hiding behind a compressed render scale. `CUT_U` stays `0.5`;
`cutPx` and `scale` are still printed on every issue so a devtools ruler still
confirms the number on the actual page. A new, explicitly-deferred finding
from the same review: `cut` only measures vertically, so an unbreakable token
wider than its box (a long URL, a `snake_case` identifier) is invisible to
soft-wrap, tall, *and* cut simultaneously — filed as a backlog item, not
implemented, pending a fifth acceptance control.
**The `from_agent`-identity finding above is also now in the `mempalace`
skill itself** ("Writing to another machine", and Anti-Patterns), and the
baked fallback snapshot of that skill was refreshed to match
(`vendor-mempalace-skill.sh`, `6eb20af` → `a12fe5e`) — sanctioned to skip on
its own (`--check` reported stale-but-truthful), done anyway because this
release's point is getting today's fixes live, and the base rebuild below was
already forced regardless.
### Dependency audit (2026-08-27)
Every component checked against upstream by direct command, not assumed:
| Component | Baked in v1.8.10 | Upstream now | Action |
|---|---|---|---|
| **mempalace-toolkit** | `b2b50af` | **`21023e7`** | ships the rsync ship-fix + RFC 003 §7.13 (both above) |
| **skillset** (mempalace fallback snapshot) | `6eb20af` | **`a12fe5e`** | re-vendored (above); live-mounted devices already had it |
| pi | `0.84.3` (pinned) | `0.84.3` is npm latest | none |
| mempalace | `3.8.0` (pinned) | `3.8.0` is PyPI latest | none |
| pi-atelier | `v0.8.2` (pinned) | `v0.8.2` highest tag | none |
| pi-studio (studio variant) | `v0.9.52` | `v0.9.52` — `main`'s commit and the tag's commit are identical (0 either direction) | none |
| pi-toolkit | `0e1369e` | `0e1369e` (local clone HEAD == `origin/main`) | none |
| pi-extensions | `2022887` | `2022887` (local clone HEAD == `origin/main`) | none |
| pi-fork | `bf702b4` | `bf702b4` | none |
| pi-observational-memory | `ce9fc98` | `ce9fc98` | none |
pi-toolkit / pi-extensions checked against their actual Gitea origin (the
Dockerfile's `PI_TOOLKIT_REPO` / `PI_EXTENSIONS_REPO`), not a GitHub mirror —
querying `api.github.com` for those two returned nothing (rate-limited or
blocked; not investigated, the local clones are the source of truth anyway).
No SHA above is fork-supplied; a fork inventing plausible-looking upstream SHAs
is a recorded failure mode (v1.8.9), so every value here came from
`git ls-remote`, a local clone's own `origin/HEAD`, `npm view`/registry JSON,
or the PyPI JSON API, run directly.
---
## v1.8.10 — 2026-08-27
+1 -1
View File
@@ -305,7 +305,7 @@ ARG MEMPALACE_TOOLKIT_REF=main
# no ~67-minute base rebuild. (scripts/check-base-hash.sh scans only
# Dockerfile.base, so no folding into the base hash is required — nor would
# it be correct, since this ARG changes nothing about the base's contents.)
ARG SKILLSET_SNAPSHOT_REF=6eb20af181f0147cb8c1377f6e36a6a47a68e8e5
ARG SKILLSET_SNAPSHOT_REF=a12fe5ecc71e60feb24791e3e33571105f1afba7
# Dockerfile.base sets description="pi-devbox — base image (variant-independent)"
# and every variant INHERITS it, so both published images used to advertise
@@ -428,10 +428,56 @@ An obligation you never agreed to is noise, so the sender states it:
| `to_agent="*"` (any status) | broadcast FYI | nothing |
| any other status (`ready`, `applied`, `blocked`, …) | a statement of fact | nothing |
**That table says what you *owe*. Delivery is stricter, and the difference bites:
the mailbox is an obligation channel, not a news channel.** Mailbox candidates are
drawn with `status="open"`, so an event carrying any **terminal** status
(`applied`, `superseded`, `failed`, `blocked`) is never a candidate — *whoever it
is addressed to*. A `task.reply` written to a named machine to share a finding is
delivered to nobody, ever, and neither is any `event_ack`. It sits in the log
until somebody reads the log.
So the most natural inter-machine message — *"here is something you should
know"* — is exactly the shape that gets no delivery. Pick deliberately:
| You want the peer to… | Write |
|---|---|
| **do something**, and you need it tracked until done | directed `status="open"` ask, with a `correlation_id` |
| **know something**, no response needed | terminal-status event **plus a drawer** — the drawer is what actually reaches them, via search |
What does **not** work is a terminal report plus an expectation of attention.
Measured 2026-08-26: a detailed report addressed to `pi@<peer>` with
`status="applied"` went unread for two and a half hours until the operator quoted
the event id by hand, with the mailbox working correctly the whole time. Full
mechanism in the toolkit's `docs/rfc-003-coordination-log.md` §7.12.
One more timing fact, because it looks like negligence and is not: a delivered
ask is queued into the agent's **next turn** (`deliverAs: "steer"`, deliberately
no `triggerTurn`), and the poll fires when the agent is *idle*. Between delivery
and the next turn no inference runs, so **a human starting a turn is the
trigger** (§7.11). An agent that "has not reacted" has usually not been running.
Ack with `mempalace_event_ack(event_id=…, from_agent="<you>", status=…)`. It
**appends a new event** and never mutates the original; the correlation id is
copied for you, and `metadata.ack_of` is set to the event you answered.
**Claiming, and what it does not do.** `status="claimed"` announces that you have
picked work up. Nothing requires it — a directed open ask owes "an ack *or* a
reply", and finishing the work is a complete answer. Do it anyway when the work is
long or the machine is unreliable, because it is the only thing that later
distinguishes *nobody started this* from *someone started and their container
died mid-task*. Be clear about its limits, both of which follow from candidacy
requiring exactly `status="open"`:
- **It does not notify the requester.** `claimed` is not `open`, so a claim is no
more deliverable than a finished report is (see the delivery table above). Its
reader is whoever pulls the log.
- **It does not quiet your own mailbox.** The ask stays owed until a *terminal*
event of yours joins it, so a claimed-then-silent thread keeps resurfacing —
correctly.
Prefer a prompt terminal reply over a claim plus a long silence; claim *in
addition*, when the gap between pickup and finish is where a machine might die.
#### What you actually owe — derive it, do not read it off `status`
The log is append-only and `status` is written **once**, so it is an honest
@@ -502,6 +548,17 @@ Two consequences worth internalising:
- **Address the stamped name you actually saw** in a `from_agent` field, e.g.
`pi@tor-ms22`. A bare `pi` reaches nobody's mailbox once stamping is live, and
older events in the log still carry bare names — do not copy them.
- **The rule runs in reverse too: what you put in YOUR OWN `from_agent` decides
where every reply to your event goes.** Nothing stops you writing a synthetic
or borrowed identity there, and a reply is always addressed back to exactly
that string — so if no live session ever runs as it, the reply is stored,
searchable, and delivered to no one. Measured cost: a directed ask sent under
a synthetic sender got two correct replies, one of them an urgent security
finding, and both sat unread for ~2h20m because nobody's mailbox was that
identity (RFC 003 §7.13). Authoring under a synthetic name is fine for a
deliberate control experiment — this fleet does it on purpose — but then
**name the real identity to reply to inside the body**, because the address
line is not a safe place to also carry provenance.
- **Use `status="open"` only when you truly need an answer.** It places an
obligation on another machine.
- **Never broadcast an ask.** `to_agent="*"` + `status="open"` obliges everyone
@@ -513,7 +570,8 @@ Two consequences worth internalising:
be matched to it at all.
- **Corrections are new events, never edits.** Say explicitly what you retract
and name the id — drawer or event — that carried the withdrawn claim.
- **Put a retraction where the reader will look.** An event reaches a live agent;
- **Put a retraction where the reader will look.** A *directed open ask* reaches a
live agent's mailbox; a **terminal-status event reaches no mailbox at all**, and
a *drawer* is what a future semantic search finds. If you filed advice as a
drawer and later withdraw it, file the withdrawal as a drawer too — otherwise
the next agent finds your original confident advice and no trace of the
@@ -600,4 +658,5 @@ Entity-relationship triples with temporal validity. Query with `mempalace_kg_que
- **Don't treat the palace as a task list.** It's for knowledge and context, not todos.
- **Don't broadcast an ask, and don't leave one unanswered.** On a shared palace, `to_agent="*"` + `status="open"` obliges every machine and therefore none of them. And don't expect acking to tidy your mailbox: `status` is immutable, so the event keeps matching either way — what a terminal reply buys you is that the *derived* owed set (see *What you actually owe*) stops counting it. Leave asks unanswered and that set only grows, until everyone learns to stop looking. "Seen, not doing it" is a complete answer — silence is not.
- **Don't assume you would have heard.** Nothing pushes another machine's message into your session. If you did not run the mailbox query at wake-up, a correction addressed to you by name can sit unread while you confidently rebuild the thing it warned you about.
- **Don't author an ask under an identity nobody runs as, including your own throwaway labels.** The failure is symmetric to the one above: it is not that you missed a message, it is that nothing could ever have delivered the reply to you, because you addressed it at a name instead of an agent. If you must use a synthetic sender for a control or an experiment, say inside the body who should actually receive the reply.
- **Don't invent provenance metadata, and don't hand-stamp it either.** An earlier version of this list told you to set `added_by="<harness>@<device>"` by hand; that instruction has been withdrawn, because RFC 001 §7.3.2 places provenance at the client/server boundary and the pi bridge now does it uniformly (see *Provenance is stamped for you* above) — but the withdrawal only holds where the bridge is live, so run the one-line check in that bullet first; on an older image hand-stamping is still the only signal a hand-filed drawer gets. DO NOT invent values for the palace's own metadata fields (`device`, `agent_kind`, `origin_device`): those are stamped by infrastructure that also records *how* each was determined, and a fabricated value is worse than none because it silently corrupts a future merge. DO pass `source_drawer_id` on `kg_add`. And never put a machine name in a diary's `agent_name` — it becomes the wing name and hides your entries from `diary_read`.