Compare commits

...

9 Commits

Author SHA1 Message Date
pi f645e6654f smoke: fix the snapshot canary that blocked v1.8.7, and make it bidirectional
Lint / hadolint (push) Successful in 15s
Lint / actionlint (push) Successful in 27s
Publish Docker Image / resolve-versions (push) Successful in 1m5s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-base (push) Successful in 41m8s
Publish Docker Image / smoke (push) Successful in 4m49s
Publish Docker Image / smoke-studio (push) Successful in 18m28s
Publish Docker Image / build-variant (push) Successful in 15m46s
Publish Docker Image / promote-base-latest (push) Successful in 11s
Publish Docker Image / update-description (push) Successful in 20s
Publish Docker Image / build-variant-studio (push) Successful in 16m52s
Run 589 built the base cleanly and then failed both smoke jobs 81-passed/1-failed
on 'mempalace skill snapshot is current'. That canary greps a phrase from the
vendored mempalace skill to detect a stale snapshot, and the phrase it pinned was
'Attribute what you file yourself' — the heading of the hand-stamping instruction
that THIS release withdraws. So it fired correctly: the snapshot changed and the
expectation did not. Every publish job was skipped, so nothing reached the
registry and v1.8.7 was never consumed.

Rather than bump the string:

* the assertion is now BIDIRECTIONAL — the new phrase must be present AND the
  withdrawn one absent. A one-way canary only catches half the drift: it cannot
  notice a re-vendored stale snapshot that happens to contain the pinned phrase.
  Verified against v1.8.6's snapshot, which now correctly fails.
* the comment records the structural limit rather than just the fix: a phrase
  canary can only ever detect 'older than what I remembered to pin', never
  'older than skillset main'. Only a diff against the skillset repo can do that,
  which is now a Still-open item — it needs a CI clone credential for a private
  repo, i.e. a policy decision, not a code change.

Changelog consolidated: the SSH sidecar multiplexing default moves from
Unreleased into v1.8.7, since the retag will sit on a commit that contains it,
and the v1.8.7 summary now records the failed first attempt rather than quietly
presenting the second one as the whole story.
2026-08-26 08:09:09 +02:00
pi 657b1ad856 ssh sidecar: default to multiplexing, as a default and not an override
Lint / actionlint (push) Successful in 15s
Lint / hadolint (push) Successful in 16s
A target whose ~/.ssh/config entry never mentioned ControlMaster got no
multiplexing from the sidecar (only ControlPath was supplied), so every ssh call
opened a fresh TCP connection. On 2026-08-25 that produced ~12 connections to
one host in 15 min and a fail2ban block that looked like an outage — the tell
being that HTTPS to the same estate stayed healthy.

The correctness of this depends entirely on WHERE the block goes. ssh_config is
first-value-wins:

  ControlPath   before the Include -> override (the user's value points at
                read-only ~/.ssh and cannot work in the container)
  ControlMaster after  the Include -> default  (an explicit per-host
                'ControlMaster no' must keep winning)

Force what is broken, default what is merely absent. The first draft put both in
the leading block and would have silently overridden an explicit 'no'.

Verified with ssh -G rather than from the man page, including the counterfactual:
under the shipped layout an explicit 'no' resolves to controlmaster false while a
silent host resolves to auto; under the rejected layout the 'no' host flips to
auto. So the test discriminates position, not presence. Plus a sandbox render of
the real script, bash -n, and shellcheck -S error (the v1.8.7 gate) clean.

Effect measured on 41 real host aliases: 22 silent entries gain auto+10m, 0
overridden. Note the fleet's one deliberate opt-out is written as absence plus a
comment ('# No ControlMaster — VPN means direct route'), which ssh cannot
distinguish from no opinion; that host now multiplexes, which its own comment
says is unnecessary rather than harmful.

Skill documents the sidecar-vs-~/.ssh trap (the failure misleads: read-only
ControlPath makes multiplexing look impossible rather than misconfigured) and
the stale-master recovery, ssh -O check / -O exit.
2026-08-25 23:09:46 +02:00
pi ebd0de0be2 changelog: v1.8.7 — device provenance reaches the fleet
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 17s
Publish Docker Image / resolve-versions (push) Successful in 10s
Publish Docker Image / base-decide (push) Successful in 12s
Publish Docker Image / build-base (push) Successful in 42m4s
Publish Docker Image / smoke (push) Failing after 4m44s
Publish Docker Image / build-variant (push) Has been skipped
Publish Docker Image / promote-base-latest (push) Has been skipped
Publish Docker Image / update-description (push) Has been skipped
Publish Docker Image / smoke-studio (push) Failing after 7m46s
Publish Docker Image / build-variant-studio (push) Has been skipped
The provenance fix's client half lives in mempalace-toolkit, which the image
clones at build time, so it only reaches the fleet through a tag. Records both
routes (extension via MEMPALACE_TOOLKIT_REF folded into base_tag; vendored skill
via rootfs), the three design points (stamp in the client not the agent; diary
marker in TEXT because metadata is invisible to readers; solitary devbox stamps
nothing), and why the allowlist is per tool (3.8.0 hard-fails -32602 on
undeclared args). Carries the CI-hardening work already sitting in Unreleased,
and a Still open block for the three known bounds.
2026-08-25 22:47:53 +02:00
pi 4f1aa0d0dd skills: refresh the vendored mempalace snapshot (withdrawn hand-stamping)
Lint / hadolint (push) Successful in 9s
Lint / actionlint (push) Successful in 19s
VENDORED.md's freshness model for `mempalace` is "Option 2 only — refreshed
manually per release", and it had drifted since 2026-08-23. The stale snapshot
still carried the instruction to hand-stamp added_by="<harness>@<device>", which
skillset 73c7c8e withdrew: the pi bridge now stamps at the edge
(mempalace-toolkit 553d865), and RFC 001 §7.3.2 ranks agent-side stamping ❌
worst-possible.

That matters specifically for the fallback case this snapshot exists to serve — a
container started WITHOUT the private skillset mounted would otherwise be the
only kind of container still being taught to do it by hand.
2026-08-25 22:27:04 +02:00
joakimp 9e744d701f lint: shellcheck the repo's own shell scripts, not just workflow run: steps
Lint / actionlint (push) Successful in 15s
Lint / hadolint (push) Successful in 15s
lint.yml has shellchecked every workflow `run:` step since the dash-vs-bash
incidents, but nothing had ever pointed shellcheck at entrypoint.sh, scripts/*.sh
or the extensionless tools under rootfs/usr/local/bin/. That gap is not
hypothetical: the skillset repo's ci-release-watcher template shipped
`echo "$json" | python3 <<'EOF' ... json.load(sys.stdin)` for two months, where
the heredoc IS python's stdin (no script arg) so the load hit EOF and the
function silently returned nothing. shellcheck names exactly that at severity
ERROR — SC2259, "This redirection overrides piped input" — and could have named
it the whole time.

New step in the existing actionlint job, so no second container pull: shellcheck
-S error plus bash -n over every shell file, discovered as *.sh UNION a shebang
scan (the glob alone misses pi-devbox-version, devbox-skill-reconcile, dot-watch
and studio-expose; a shebang scan alone would miss a sourced fragment without
one). Fails loudly on a zero-file match, because a green tick over an empty set
is not a check.

Severity chosen by measurement, not taste: -S error is 0 findings across all 11
shell files today, so the gate is green on arrival with no cleanup, while
-S warning is NOT free (19x SC2088 tilde-in-quotes in recreate-sanity-check.sh
plus assorted SC2016, all intentional) and would train everyone to ignore the
job — the same reasoning as the SHELLCHECK_OPTS exclusions already on the
actionlint step.
2026-08-25 21:11:55 +02:00
joakimp 2b8c3a4db4 ci: audit MEMPALACE_VERSION the way PI_VERSION is audited
Lint / hadolint (push) Successful in 8s
Lint / actionlint (push) Successful in 17s
Closes the item v1.8.6 (and v1.8.5 before it) listed as "Still open": the
palace pin was a literal string in Dockerfile.base with zero references in
docker-publish.yml, while PI_VERSION had a concreteness gate, a
published-on-registry check and a never-silently-adopt drift warning.

resolve-versions now applies all of those to MEMPALACE_VERSION, read from
Dockerfile.base so a local `docker build` and CI install the same version by
construction, plus one gate pi does not need: a YANKED release is refused,
because an exact pin installs one silently under PEP 592 and would have
shipped a withdrawn palace client to the whole fleet.

smoke gains `installed mempalace matches CI's audited pin` via a new
EXPECTED_MEMPALACE_VERSION threaded into both smoke jobs. It is not redundant
with `manifest mempalace_version matches the installed core`: that compares two
properties of one image and cannot notice that both are the wrong version. The
case this covers is a variant built FROM a cached base carrying an older pin —
internally consistent, silently stale.

Mutation-tested by extracting the shipped block out of the YAML and stubbing
curl: 9 cases covering every gate, then once end-to-end against live PyPI. That
found a real defect in the first draft — the yank message inlined a jq program
inside $(...) inside a double-quoted string, where the escaping broke the
filter (jq compile error) while the surrounding `exit 1` still fired: a gate
that looked correct and reported garbage.

Note: correcting Dockerfile.base's now-false "known gap, carried forward"
comment forces a base rebuild (~67 min) on the next tag. Leaving a comment
asserting the audit does not exist was the worse option.
2026-08-25 20:33:19 +02:00
joakimp cb7b8ad2ae smoke: assert manifest VALUES, not the presence of field names
Lint / actionlint (push) Successful in 16s
Lint / hadolint (push) Successful in 17s
Four build-provenance assertions grepped the manifest for a field name and
never looked at the value:

    run_expect "manifest records pi_version" "cat …manifest.json" '"pi_version"'

which passes on {"pi_version": ""} and on {"pi_version": null}. The tell was
in its own passing output the whole time — `✅ manifest records pi_version (got
"pi_version")` echoes the key back as the thing it claims to have found. Found
while reading run 579's smoke log to confirm v1.8.6's new assertions had really
executed rather than merely gone green.

Now checked against values, and against ground truth where it exists:

- every required component key present, naming the one that vanished
- every component value a full 40-hex SHA (null allowed for pi-studio alone,
  which is legitimately absent in the non-studio variant)
- pi_version equal to `pi --version`, mirroring the mempalace ground-truth check
- release_tag non-empty; source_revision 40-hex and build_date ISO-8601 *when
  populated*, since both default empty on a plain local `docker build` and
  demanding them would fail honest local smoke runs
- --json compared byte-for-byte with the file, which is assertable because that
  mode is a verbatim cat; the old form grepped its output for "release_tag"

Key presence and value shape are deliberately SEPARATE assertions: a single
"all values are valid SHAs" loop passes vacuously on components:{}, because
jq's all() over an empty list is true. Combining them would reproduce the same
shape of hole as the three false greens already recorded in CHANGELOG.md.

Dropped `manifest has no unresolved ('unknown') components`: the 40-hex check
strictly subsumes it ("unknown" is not 40-hex, and only rev() emits it, feeding
components{} exclusively). Removed rather than kept, because a check that can
no longer fail independently is one more green tick that means nothing.

Mutation-tested twice rather than reasoned about: nine fabricated manifests
through the raw jq filters, then twelve through the shipped assertions using
the real `run` helper's `sh -c` quoting path — the quoting is load-bearing,
since a jq filter dying on a quoting error exits non-zero and looks exactly
like a caught defect. Measured on the same twelve defects: old caught 3, missed
9; new catches 12. Three legitimate variations stay green (empty
source_revision, empty build_date, null pi-studio).

Also corrects a factually wrong "Still open" bullet in the released v1.8.6
entry, which claimed pi-devbox-version's human output does not show
mempalace_version and that only --json surfaces it. Both halves are false: it
prints a `palace:` line with live-vs-baked drift annotation, verified against
fabricated manifests (match, skew, and pre-v1.8.6 absent-field cases). Left as
a struck-through correction rather than deleted, since v1.8.6 is published.
2026-08-25 17:15:39 +02:00
joakimp 93f986e90e v1.8.6: adopt pi 0.84.3 + mempalace 3.8.0, close the v1.8.5 doc/observability gaps
Lint / hadolint (push) Successful in 9s
Publish Docker Image / resolve-versions (push) Successful in 15s
Publish Docker Image / base-decide (push) Successful in 8s
Lint / actionlint (push) Successful in 1m9s
Publish Docker Image / build-base (push) Successful in 41m23s
Publish Docker Image / smoke (push) Successful in 4m50s
Publish Docker Image / smoke-studio (push) Successful in 5m5s
Publish Docker Image / build-variant (push) Successful in 15m46s
Publish Docker Image / update-description (push) Successful in 7s
Publish Docker Image / promote-base-latest (push) Successful in 15s
Publish Docker Image / build-variant-studio (push) Successful in 19m59s
Three coupled pieces of work, all of which ride on the base rebuild that the
mempalace bump forces anyway.

DRIFT ADOPTED
- pi 0.84.2 -> 0.84.3. Its release notes carry a "Breaking Changes" line
  (GoogleThinkingLevel -> GoogleApiThinkingLevel). Audited before adopting:
  zero references across all four vendored companions (pi-fork,
  pi-observational-memory, pi-atelier, pi-studio), so it is inert for us. The
  reason to adopt is two skill-discovery fixes that land directly on v1.8.5's
  vendored-skill work: nested Markdown skills inside grouping directories were
  not discovered, and root README.md/AGENTS.md in skill dirs were reported as
  broken skills.
- mempalace core 3.7.1 -> 3.8.0. Additive/reliability only. Its sync fix
  (#2320/#2322) stops sync --apply deleting drawers whose source_file was
  unreachable *at that moment* -- which does NOT relax the standing landmine
  against sync on the shared palace, because that landmine is about paths
  permanently absent from whichever host runs the sync. Different failure
  shape; the caution stands.

DOCS -- three defects, one of them public
- DOCKER_HUB.md advertised "neovim (LazyVim defaults)". Nothing in the image
  installs LazyVim; the only nvim config is a 19-line sysinit.vim. CI PATCHes
  this file into the Docker Hub description on every release, so this was a
  false claim published to the world. Removed.
- agent-browser + Playwright + Chromium is the single largest addition in the
  image (~625 MB) and had zero mentions in README, DOCKER_HUB or THIRD_PARTY --
  it was documented only to agents, in the AGENTS.md managed block. Now
  documented to humans, including the Chromium licence dimension.
- typst and socat appeared in README prose but not in the "What's inside"
  inventory. Added.

OBSERVABILITY -- the three gaps v1.8.5 listed as still open
- build-manifest.json now records mempalace core, read from the live binary
  (ground truth, not the build ARG). Placed as a sibling of pi_version rather
  than inside components{}, because pi-devbox-version renders that map through
  [0:12] and would truncate a version string.
- smoke asserts the pi-observational-memory clone actually CONTAINS the ce9fc98
  auth fix, pinned to src/runtime.ts. Deliberately not a repo-wide grep: two of
  the three markers also live under tests/, so the repo-wide form stays green
  with the fix site reverted. That is the third false-green of this exact family
  in this repo (canary phrase in both snapshots; reconciler fixture using a
  non-owned name; now this) -- pin containment checks to the fix site.
- smoke asserts the feeder's pi@<device> agent default behaviourally. The
  earlier audit concluded this needed a --print-config added upstream; it does
  not. AGENT is assigned before arg parsing, so `bash -x mempalace-pi-session
  --help` observes the real resolution with no toolkit change. Two-sided:
  device set => pi@<device>, unset => must not be pi@*.
- pi-devbox-version now prints a palace: line with the same live-vs-baked drift
  detection pi already had. This matters more than it looks: mempalace is the
  one component that is both client (here) and server (synlig), so skew between
  them is a real failure mode. Degrades quietly on pre-v1.8.6 images.

Deferred deliberately: a native arm64 act_runner on tor-ms22 (the current
runner is on synlig, x86_64, so every arm64 layer ships QEMU-emulated).
Analysis and caveats filed to the palace rather than actioned here.
2026-08-25 15:29:22 +02:00
joakimp 26f223568d .env.example: document MEMPALACE_PALACE_PATH and why nothing exports it
Lint / hadolint (push) Successful in 10s
Lint / actionlint (push) Successful in 50s
The only MemPalace variable the template never mentioned, and the one that
moves the feeders' stage as a side effect: the palace root resolves as
$MEMPALACE_PALACE_PATH -> $MEMPAL_PALACE_PATH -> ~/.mempalace/config.json ->
~/.mempalace/palace, and the stage is derived from it (<palace-root>/pi-stage).

The comment states that precedence, records why neither the image nor the
entrypoint exports it (pinning the palace without carrying the stage along
re-creates the split a shared root removed, v1.8.2), warns that a stage whose
persistence differs from the palace makes a scoped `mempalace sync` prune
conversation drawers whose dedup key is the staged path, and notes it is a
path INSIDE the container unlike the host-side WORKSPACE_PATH/SSH_KEY_PATH
above it.

Found while auditing a live host whose .env sets it redundantly to the
default value.
2026-08-23 23:27:42 +02:00
14 changed files with 1069 additions and 23 deletions
+13
View File
@@ -18,6 +18,19 @@ SSH_KEY_PATH=~/.ssh
# the staged files and the palace dedup keys pointing at them cannot be
# separated.
#
# That palace root is resolved with mempalace's own precedence
# ($MEMPALACE_PALACE_PATH -> $MEMPAL_PALACE_PATH -> ~/.mempalace/config.json ->
# ~/.mempalace/palace), and the feeders derive their stage FROM it
# (<palace-root>/pi-stage). Neither the image nor the entrypoint exports it, by
# design: pinning the palace without carrying the stage along re-creates the
# very split that a shared root removed. Override it only to move the palace off
# the default -- e.g. onto a different mount -- and only to a path with the SAME
# persistence as the palace itself. A stage that outlives its palace (or dies
# first) makes a scoped `mempalace sync` prune conversation drawers, because
# their dedup key is the staged path. Setting it to the default buys nothing.
# Unlike WORKSPACE_PATH/SSH_KEY_PATH above, this is a path INSIDE the container.
# MEMPALACE_PALACE_PATH=/home/developer/.mempalace/palace
#
# To instead share ONE MemPalace across containers/harnesses (pi + opencode
# + native), set the URL below. When set, the extension connects over HTTP
# and NO local mempalace-mcp is spawned; the devbox-palace volume is then
+52
View File
@@ -142,6 +142,7 @@ jobs:
image: catthehacker/ubuntu:act-latest
outputs:
pi_version: ${{ steps.resolve.outputs.pi_version }}
mempalace_version: ${{ steps.resolve.outputs.mempalace_version }}
fork_ref: ${{ steps.resolve.outputs.fork_ref }}
obsmem_ref: ${{ steps.resolve.outputs.obsmem_ref }}
toolkit_ref: ${{ steps.resolve.outputs.toolkit_ref }}
@@ -242,6 +243,54 @@ jobs:
fi
echo "pi_version=${PI_VERSION}" >> "$GITHUB_OUTPUT"
# ── mempalace core: same audit as pi, from Dockerfile.base ────
# Until now this pin had NO CI-side audit at all — a literal string
# in Dockerfile.base with zero references in this workflow, while
# PI_VERSION got a concreteness gate, a published-on-registry check
# and a drift warning. It is the same class of risk: the palace's MCP
# tool schema is the agent-facing contract, and a client/server skew
# against the shared central palace is a fleet-wide, not local,
# problem. Read from Dockerfile.base (not duplicated here) so a local
# `docker build` and CI install the same version by construction.
MEMPALACE_VERSION=$(sed -n 's/^ARG MEMPALACE_VERSION=\([^[:space:]]*\).*/\1/p' Dockerfile.base | head -n1)
if ! printf '%s' "${MEMPALACE_VERSION:-}" | grep -qE '^[0-9]+\.[0-9]+\.[0-9]+$'; then
echo "::error::ARG MEMPALACE_VERSION in Dockerfile.base is not a concrete version (got '${MEMPALACE_VERSION:-<empty>}'). CI refuses to build from a floating palace version — see the pin policy comment above that ARG."
exit 1
fi
# One fetch, two gates. `curl -sf` exits non-zero and prints nothing
# on 404 (PyPI's answer for an unpublished version), so an empty body
# lands in the "not published" branch with its own message.
MEMPALACE_PYPI=$(curl -sf "https://pypi.org/pypi/mempalace/${MEMPALACE_VERSION}/json" || true)
MEMPALACE_PUBLISHED=$(printf '%s' "$MEMPALACE_PYPI" | jq -r '.info.version // empty' 2>/dev/null || true)
if [ "${MEMPALACE_PUBLISHED:-}" != "${MEMPALACE_VERSION}" ]; then
echo "::error::Pinned mempalace version ${MEMPALACE_VERSION} is not published on PyPI (registry returned '${MEMPALACE_PUBLISHED:-<empty>}'). Fix ARG MEMPALACE_VERSION in Dockerfile.base."
exit 1
fi
# A yanked release still installs when pinned exactly (PEP 592), so
# `uv tool install mempalace==X` would succeed silently and ship a
# version upstream has withdrawn to the whole fleet. The escape hatch
# is the same one-line bump that got us here.
MEMPALACE_YANKED=$(printf '%s' "$MEMPALACE_PYPI" | jq -r '.info.yanked // false' 2>/dev/null || true)
if [ "${MEMPALACE_YANKED:-false}" = "true" ]; then
# Reason hoisted into its own variable rather than inlined as a
# $(...) inside the message: a jq program nested in a substitution
# inside a double-quoted string needs escaping that silently breaks
# the FILTER (jq compile error) while the surrounding `exit 1` still
# fires, so the gate looks correct and reports garbage. Caught by
# the mutation test, not by review.
MEMPALACE_YANK_REASON=$(printf '%s' "$MEMPALACE_PYPI" | jq -r '.info.yanked_reason // "no reason given"' 2>/dev/null || true)
echo "::error::Pinned mempalace version ${MEMPALACE_VERSION} is YANKED on PyPI (${MEMPALACE_YANK_REASON:-no reason given}). An exact pin installs a yanked release without complaint — bump ARG MEMPALACE_VERSION in Dockerfile.base."
exit 1
fi
# Informational only, exactly like pi's npm drift warning: a newer
# palace must never be adopted implicitly. `|| true` so a transient
# PyPI failure cannot fail a release whose pin is already verified.
MEMPALACE_PYPI_LATEST=$(curl -sf "https://pypi.org/pypi/mempalace/json" | jq -r '.info.version // empty' 2>/dev/null || true)
if [ -n "${MEMPALACE_PYPI_LATEST:-}" ] && [ "${MEMPALACE_PYPI_LATEST}" != "${MEMPALACE_VERSION}" ]; then
echo "::warning::mempalace ${MEMPALACE_PYPI_LATEST} is published; this build ships the audited pin ${MEMPALACE_VERSION}. To adopt it: read the upstream CHANGELOG for MCP tool-schema changes (the agent-facing contract) and for sync/delete semantics, check the skew it introduces against the central palace host's server version, then bump ARG MEMPALACE_VERSION in Dockerfile.base and note the audit in CHANGELOG.md."
fi
echo "mempalace_version=${MEMPALACE_VERSION}" >> "$GITHUB_OUTPUT"
# pi-fork / pi-observational-memory (GitHub) → commit SHAs.
FORK_REF=$(curl -sf -H "Accept: application/vnd.github.sha" \
"https://api.github.com/repos/elpapi42/pi-fork/commits/master" || true)
@@ -337,6 +386,7 @@ jobs:
echo "studio_tag=${STUDIO_TAG}" >> "$GITHUB_OUTPUT"
echo "Resolved PI_VERSION=${PI_VERSION} (pinned in Dockerfile.variant; npm latest is ${PI_NPM_LATEST:-unknown})"
echo "Resolved MEMPALACE_VERSION=${MEMPALACE_VERSION} (pinned in Dockerfile.base; PyPI latest is ${MEMPALACE_PYPI_LATEST:-unknown})"
echo "Resolved PI_ATELIER_REF=${ATELIER_REF} (pi-atelier ${ATELIER_TAG}, pinned)"
echo "Resolved PI_FORK_REF=${FORK_REF}, PI_OBSMEM_REF=${OBSMEM_REF}"
echo "Resolved PI_TOOLKIT_REF=${TOOLKIT_REF}, PI_EXTENSIONS_REF=${EXTENSIONS_REF}"
@@ -471,6 +521,7 @@ jobs:
- name: Smoke test (amd64)
env:
EXPECTED_PI_VERSION: ${{ needs.resolve-versions.outputs.pi_version }}
EXPECTED_MEMPALACE_VERSION: ${{ needs.resolve-versions.outputs.mempalace_version }}
run: bash scripts/smoke-test.sh pi-devbox:smoke
# ── Phase 3b: amd64 smoke for the studio variant ────────────────────
@@ -533,6 +584,7 @@ jobs:
- name: Smoke test studio (amd64)
env:
EXPECTED_PI_VERSION: ${{ needs.resolve-versions.outputs.pi_version }}
EXPECTED_MEMPALACE_VERSION: ${{ needs.resolve-versions.outputs.mempalace_version }}
run: bash scripts/smoke-test.sh pi-devbox:smoke-studio
# ── Phase 4: multi-arch publish ─────────────────────────────────────
+49
View File
@@ -52,6 +52,55 @@ jobs:
apt-get update
apt-get install -y --no-install-recommends shellcheck python3-yaml
- name: "Shellcheck + syntax-check repository scripts (severity: error)"
# Gap being closed: everything else in this job shellchecks workflow
# `run:` steps ONLY, via actionlint. The repo's own shell scripts —
# entrypoint.sh, scripts/*.sh, and the extensionless tools under
# rootfs/usr/local/bin/ — have never been shellchecked. That exact gap
# (a sibling repo with no shell-script lint at all) is how a defect
# shipped invisibly for two months: `echo "$json" | python3 <<'EOF'
# ... json.load(sys.stdin)` cannot work — with no script argument
# python reads its SCRIPT from stdin, so the heredoc IS stdin and the
# json.load call hits EOF. shellcheck flags exactly this at severity
# ERROR (SC2259, "This redirection overrides piped input"); nothing
# ever ran it. Measured before adding this gate: `-S error` is 0
# findings across every shell file in THIS repo today, so it is free
# to add. `-S warning` is NOT free here (19x SC2088 tilde-in-quotes in
# scripts/recreate-sanity-check.sh, plus assorted SC2016 — both
# intentional), so warning-level would train people to ignore the job;
# hence error-only, matching the SHELLCHECK_OPTS philosophy below.
#
# Discovery is *.sh UNION a shebang scan, because rootfs/usr/local/
# bin/{pi-devbox-version,devbox-skill-reconcile,dot-watch,studio-expose}
# are shell scripts with no extension. -print0/mapfile -d '' so a path
# with a space cannot silently split, and the file count is asserted
# non-zero — a green tick over an empty file set is not a check.
run: |
# Union of two signals, because either alone misses a real case:
# a shebang scan misses a sourced fragment with no shebang, and a
# *.sh glob misses the extensionless tools in rootfs/usr/local/bin/.
# Silent skipping is precisely the failure mode this gate exists to
# prevent, so err toward over-collecting.
mapfile -d '' -t all_files < <(find . -not -path './.git/*' -type f -print0)
sh_files=()
for f in "${all_files[@]}"; do
case "$f" in *.sh) sh_files+=("$f"); continue;; esac
if head -n1 "$f" 2>/dev/null | grep -qE '^#!.*\b(bash|sh)\b'; then
sh_files+=("$f")
fi
done
echo "Checking ${#sh_files[@]} shell file(s)"
if [ "${#sh_files[@]}" -eq 0 ]; then
echo "::error::no shell files found — the shebang scan or the checkout is wrong"
exit 1
fi
shellcheck -S error -f gcc "${sh_files[@]}"
rc=0
for f in "${sh_files[@]}"; do
bash -n "$f" || { echo "::error file=$f::bash -n failed"; rc=1; }
done
exit "$rc"
- name: Gitea shell guard (catches the actionlint blind spot)
# actionlint models GitHub Actions, where the default run shell is
# bash, so it does NOT flag bash syntax in a step that merely OMITS
+492
View File
@@ -11,6 +11,498 @@ Pre-v1.0.0 tags followed the pi npm version (`v{pi_version}[letter]`).
---
## v1.8.7 — 2026-08-25
Patch release, and the fastest turnaround in the series (~9 h after v1.8.6) for
one reason: **v1.8.6 shipped a container that cannot tell you which machine it
is running on**, and that anonymity produced a real misattribution the same
evening — a session on tor-ms22 read *another host's* diary out of the shared
palace, reported its verification as its own, and built a causal inference on
top of the coincidence. The client-side half of the fix lives in
`mempalace-toolkit`, which the image clones **at build time**, so it can only
reach the fleet through a tag. The CI-hardening work that had accumulated since
v1.8.6 rides along.
Gitea-hosted refs re-resolved immediately before tagging (2026-08-25T22:35Z):
pi-toolkit `0e1369e6` and pi-extensions `20228878` **unchanged** since v1.8.6;
mempalace-toolkit `0fe64c48` → `553d8657` (the provenance change below). CI
re-resolves pi-fork / pi-observational-memory / pi-atelier / pi-studio at build
time as usual. **Base rebuild is forced twice over** — `Dockerfile.base` changed
(the `MEMPALACE_VERSION` audit) *and* `base_tag` deliberately folds in the
mempalace-toolkit SHA ("otherwise a toolkit-only fix never lands") — so expect
~67 min, and note that either cause alone would have sufficed.
⚠️ **The first tag of this version did not publish.** Run 589 built the base
fine, then **both** smoke jobs failed 81-passed/1-failed on a single assertion —
`mempalace skill snapshot is current`, a canary pinning a phrase from the
vendored skill. The phrase it pinned was the heading of the very instruction this
release *withdraws*, so refreshing the snapshot without re-pinning the canary
made it fire correctly on a healthy image. Every publish job was skipped, so
nothing reached the registry and the version was never consumed; the tag was
moved to include the fix below. Fixing the canary is what this release is
*for*, in miniature: the gate was right and the expectation was stale.
### Added
- **Palace writes now carry the device that made them, and diary entries say so
in text.** The container is host-anonymous by construction — `hostname` is a
Docker hash, `$DEVBOX_HOST_ALIAS` is generic, the virtiofs source tag is
generic, and two hosts in this fleet are both `aarch64` — so nothing inside it
distinguished tor-ms22 from EMB-7KJ4VR4G. In a *local* palace that costs
nothing (one origin, so origin is a property of the whole store). In the
**shared** palace it means every drawer and all 621 diary entries read as
though written here, which is exactly how a v1.8.6 verification performed on
EMB was reported as tor-ms22's own.
Two halves, arriving by different routes:
| Half | Where it lives | How it gets into this image |
|---|---|---|
| the writer — stamps `<harness>@<device>` on `add_drawer`/`checkpoint`/`mine`/`event_append`/`artifact_put`, and prefixes diary entries with `HOST:<device>\|` | `mempalace-toolkit` `extensions/pi/mempalace.ts` (553d8657) | cloned in `Dockerfile.base` at `MEMPALACE_TOOLKIT_REF`, whose SHA is folded into `base_tag` |
| the consumer skill — stops telling the agent to do it by hand, adds the read-side warning | vendored `rootfs/…/skills/mempalace/SKILL.md`, refreshed from skillset `73c7c8e6` | `rootfs/*` is hashed into `base_tag` too |
Three design points worth recording, because each was arrived at the hard way:
- **The stamp goes in the *client*, not the agent.** RFC 001 §7.3.2 ranks
"agent stamps it via a skill instruction" as the ❌ *worst possible* place,
and the skill had carried exactly that instruction since 2026-08-23. It
failed as predicted: the agent that wrote the instruction then filed its own
provenance drawer without it. 199 rows reached the palace unresolvable.
One `execute()` wrapper cannot forget.
- **The diary marker is in the entry TEXT on purpose.** `diary_write` has no
metadata parameter, but the deeper reason is that mempalace's `search`
projects a fixed key set and `diary_read` returns content — **metadata is
invisible to the agent who will later read the entry**, so no metadata-only
fix, not even a server-authoritative one, would have prevented the
misattribution. The marker is an AAAK field, so it is machine-parseable
*and* the first thing a reader sees. The wake-up preamble now also names the
device and warns that `diary_read` interleaves every machine's diary.
- **A solitary devbox stamps nothing.** Gated on `MEMPALACE_PI_DEVICE` **and**
`MEMPALACE_REMOTE_URL` — both set only when the palace is actually shared
(RFC 001 R1). Unset either and behaviour is byte-identical to v1.8.6.
Never injected into `diary_write` or `kg_add`: mempalace 3.8.0 hard-rejects
undeclared arguments with JSON-RPC `-32602` rather than dropping them (the
behaviour changed since the RFC's 2026-08-09 note, now corrected), so a
blanket injection would **break** those two calls instead of being ignored.
The allowlist is per tool for that reason.
- **CI now shellchecks the repo's own shell scripts, not just workflow `run:` steps.** `.gitea/workflows/lint.yml`'s `actionlint` job already shellchecks every workflow step, but nothing had ever pointed shellcheck at `entrypoint.sh`, `scripts/*.sh`, or the extensionless tools under `rootfs/usr/local/bin/` (`pi-devbox-version`, `devbox-skill-reconcile`, `dot-watch`, `studio-expose`). The gap is not hypothetical: a sibling repo (skillset's `ci-release-watcher` templates) shipped `echo "$json" | python3 <<'EOF' ... json.load(sys.stdin)` for two months without anyone noticing it silently returned nothing — with no script argument python reads its *script* from stdin, so the heredoc is stdin and the JSON load hits EOF. shellcheck flags exactly this at severity **error** (`SC2259`, "This redirection overrides piped input"); it had been available to catch it the whole time, just never run.
New step in the `actionlint` job, `Shellcheck + syntax-check repository scripts`, runs `shellcheck -S error` plus `bash -n` over every shell file in the repo, **discovered by `*.sh` union a shebang scan** (neither alone suffices) so the extensionless `rootfs/usr/local/bin/*` tools are covered too. Measured before adding it: `-S error` is 0 findings across all 11 shell files today, so the gate is green on arrival with no cleanup. `-S warning` is *not* free (19× `SC2088` tilde-in-quotes in `scripts/recreate-sanity-check.sh`, plus assorted `SC2016`, both intentional here) — a warning-level gate would train people to ignore it, so it stays error-only, same reasoning as the existing `SHELLCHECK_OPTS` exclusions on the actionlint step. File-count guard included: the step fails loudly if the shebang scan matches zero files, since a green check over an empty set is not a check.
- **`MEMPALACE_VERSION` now gets the same CI audit as `PI_VERSION`** — closing
the item v1.8.6 (and v1.8.5 before it) listed as "Still open". The pin was a
literal string in `Dockerfile.base` with **zero** references anywhere in
`.gitea/workflows/docker-publish.yml`, while `PI_VERSION` had ~20: a
concreteness gate, a published-on-registry check, and a never-silently-adopt
drift warning. `resolve-versions` now applies all of them to the palace pin,
read from `Dockerfile.base` (not duplicated in the workflow, so a local
`docker build` and CI install the same version by construction):
| Gate | Behaviour |
|---|---|
| not a concrete `X.Y.Z` | **error** — no floating palace version, same policy as pi |
| not published on PyPI | **error** at resolve time, instead of a `uv tool install` failure mid-build |
| **yanked** on PyPI | **error** — an exact pin installs a yanked release silently under PEP 592, so `mempalace==X` would have shipped a withdrawn client to the whole fleet |
| newer release exists | **warning** naming what to audit before adopting (MCP tool-schema = the agent-facing contract; client/server skew against the central palace) |
Plus one smoke assertion, `installed mempalace matches CI's audited pin`,
gated on a new `EXPECTED_MEMPALACE_VERSION` env threaded into both the `smoke`
and `smoke-studio` jobs. It is **not** redundant with the existing `manifest
mempalace_version matches the installed core`: that one compares two
properties of a single image and therefore cannot notice that *both* are the
wrong version. The failure mode this one covers is a variant built `FROM` a
cached base whose `MEMPALACE_VERSION` pin was older — internally consistent,
silently stale, invisible to every other assertion (the risk
`scripts/check-base-hash.sh` exists to reduce but cannot eliminate).
**Mutation-tested rather than reasoned about**, by extracting the shipped
block out of the YAML and running it with a stubbed `curl`: 9 cases —
`latest` / `3.8` / absent ARG refused; 404 and a registry echoing a different
version refused; a yanked release refused *with its reason*; a newer release
warning without failing; a transient PyPI outage not failing a build whose pin
is already verified; happy path silent and emitting the job output. Then once
more end-to-end against live PyPI with the real `Dockerfile.base`. **This
found a genuine defect in the first draft**: the yank message inlined a jq
program inside a `$(...)` inside a double-quoted string, where the escaping
broke the *filter* (jq compile error) while the surrounding `exit 1` still
fired — a gate that looked correct and reported garbage. The reason is now
hoisted into its own variable. The new smoke assertion was checked the same
way, through the real `run` helper's `sh -c` quoting path: passes on `3.8.0`,
fails on `3.7.1` *and* on `3.8.01` (exact equality, not the substring match
the pi assertion uses), and skips cleanly when the env is unset so a local
`smoke-test.sh` run is unaffected.
⚠️ **Costs a base rebuild on the next tag**: `Dockerfile.base` is hashed
wholesale into `base_tag`, and its now-false "Known gap, carried forward"
comment had to be corrected in place (leaving a comment that says the audit
does not exist would repeat the shipped-false-claim mistake corrected below).
Expect ~67 min, as for v1.8.5/v1.8.6.
- **The SSH sidecar now defaults to connection multiplexing, without overriding
anyone's explicit choice.** `~/.ssh-local/config` already forced `ControlPath`
into the writable sidecar dir, but nothing supplied `ControlMaster` for targets
coming from the user's own bind-mounted `~/.ssh/config`. An entry that never
mentioned it therefore opened a **fresh TCP connection per `ssh` call** — and
an agent doing a dozen calls in a few minutes is exactly the traffic shape that
trips fail2ban or a CGNAT flow-table cap. Observed 2026-08-25 on this fleet:
~12 connections to one host in 15 minutes, after which port 22 stopped
answering while HTTPS to the same estate stayed healthy in 0.44 s (that
asymmetry is the tell for rate-limiting rather than an outage).
**The fix is where the block sits, not what it says.** `ssh_config` is
first-value-wins, so position encodes intent, and the two settings need
opposite treatment:
| Setting | Position | Meaning | Why |
|---|---|---|---|
| `ControlPath` | **before** `Include ~/.ssh/config` | override | the user's value points at read-only `~/.ssh`; it cannot work here, so it must lose |
| `ControlMaster auto` + `ControlPersist 10m` | **after** the `Include` | default | an explicit per-host `ControlMaster no` must keep winning; we only supply an opinion where the user expressed none |
*Force what is broken, default what is merely absent.* The first draft of this
put both in the leading block, which would have silently overridden an explicit
`ControlMaster no` — the counterfactual is in the test below.
Verified with `ssh -G` (the resolved-config oracle) rather than by reading the
man page, against a fixture with one host set to `no`, one silent, one set to
`auto`: the explicit `no` resolves to `controlmaster false` **and** still gets
the writable `ControlPath`, the silent host resolves to `auto`, and the same
fixture under the rejected layout flips the `no` host to `auto` — so the test
discriminates the *position*, not merely the presence of the block. Then
end-to-end: the real script rendered in a sandbox `HOME`, block last, `bash -n`
clean, `shellcheck -S error` clean (the gate added in v1.8.7).
Measured effect on the author's own config (41 host aliases): **22 were silent
about `ControlMaster` and gain `auto` + 10 m persist; 0 are overridden**, since
the fleet contains no explicit `no`. Worth noting *how* the one deliberate
exception is written — `proxmox002-vpn` carries `# No ControlMaster — VPN means
direct route, no CGNAT flow cap`, i.e. the intent is expressed as **absence
plus a comment**, which `ssh` cannot distinguish from "no opinion". That host
does now get multiplexing; its comment says multiplexing is *unnecessary*
there, not harmful. Anything that must stay unmultiplexed needs a literal
`ControlMaster no`.
Why the ordering matters beyond this one config: `~/.ssh/config` is
**per-machine**, differs across the fleet, and future machines' versions do not
exist yet to be audited. A default-not-override design is correct without
needing to inspect any of them.
`ControlPersist` is deliberately short (10 m idle, and each new session resets
the idle timer — long enough to collapse an agent's burst, short enough that an
abandoned socket ages out). A per-host entry that sets its own value keeps it:
hosts already specifying `ControlPersist 4h` still resolve to 4 h. The known
cost of multiplexing is the **stale master** — socket present, daemon gone,
after a suspend or network change — which makes every later `ssh` to that host
hang; recovery is `ssh -F ~/.ssh-local/config -O exit <host>`, now documented
in the `pi-devbox-environment` skill along with `-O check`.
### Fixed
- **The vendored-snapshot canary was one-way, and pinned a phrase the same
release deleted.** `mempalace skill snapshot is current` grepped for
*"Attribute what you file yourself"* — the heading of the hand-stamping
instruction withdrawn above. It therefore did its job (snapshot changed,
expectation did not) and blocked an otherwise-green build. Two changes rather
than a string bump: the assertion is now **bidirectional** (the new phrase must
be present **and** the withdrawn one absent, so a re-vendored *stale* snapshot
fails as loudly as a forgotten bump — verified by running it against v1.8.6's
snapshot, which correctly fails), and the comment now states the structural
limit: a phrase canary can only detect *"older than what I remembered to pin"*,
never *"older than skillset main"*.
- **Four build-provenance smoke assertions verified the presence of a manifest
field name and never looked at its value.** The originals were literally:
```sh
run_expect "manifest records pi_version" "cat …build-manifest.json" '"pi_version"'
```
which passes on `{"pi_version": ""}` and on `{"pi_version": null}`. The tell
was sitting in the passing output all along — `✅ manifest records pi_version
(got "pi_version")` echoes the *key* back as the thing it claims to have
found — and it was spotted while reading run 579's smoke log to confirm the
new v1.8.6 assertions had actually executed.
Replaced with checks against the values, and against ground truth where
ground truth exists:
| Assertion | What it now enforces |
|---|---|
| `manifest declares every required component key` | all seven components present by name, failing with *which* key vanished |
| `manifest component values are resolved 40-hex commits` | each value is a full 40-hex SHA; `null` allowed for `pi-studio` alone (absent in the non-studio variant) |
| `manifest pi_version matches the installed pi` | manifest value equals `pi --version`, same ground-truth shape as the mempalace check |
| `manifest top-level fields are well-formed, not merely present` | `release_tag` non-empty; `source_revision` 40-hex *when populated*; `build_date` ISO-8601 *when populated* |
| `pi-devbox-version --json round-trips the manifest byte-for-byte` | actual string equality with the file, since `--json` is a verbatim `cat` |
**Why five value-checks replace four name-checks (total assertion count
unchanged at 61), and specifically why key presence and value shape are kept
apart:** an "every component value is a valid SHA" loop passes **vacuously** on `components:{}`, because jq's `all()`
over an empty list is true. A single combined check would therefore go green
on a manifest that had lost every component — which is the same shape of hole
as the three false greens already recorded in this file. They are separate on
purpose.
Mutation-tested rather than reasoned about, twice: nine fabricated manifests
through the raw jq filters, then twelve through the shipped assertions using
the real `run` helper's `sh -c` quoting path (the quoting is load-bearing here
— a jq filter that dies on a quoting error exits non-zero and *looks* like a
caught defect). Measured against the old assertions on the same twelve
defects: **old caught 3, missed 9; new catches 12.** The three the old set
caught were key *disappearance* (grepping for a key name does fail when the
key is gone) and the literal string `"unknown"`; every value-level defect —
empty string, `null`, a 12-hex truncation, a wrong-but-plausible version, a
malformed `source_revision` — was invisible. Three legitimate variations are
correctly *not* flagged: empty `source_revision` and empty `build_date` (both
default empty on a plain local `docker build`, so demanding them would fail
honest local smoke runs) and `pi-studio: null`.
- **Dropped the now-redundant `manifest has no unresolved ('unknown')
components` assertion.** The 40-hex value check strictly subsumes it:
`"unknown"` is not 40-hex, and only `rev()` in Dockerfile.variant ever emits
that string, feeding `components{}` exclusively. Removed rather than left in
place, because a redundant check that can never fail independently is one more
green tick that means nothing.
- **Corrected a factually wrong "Still open" bullet in the v1.8.6 entry below**
(see the strikethrough there). It claimed `pi-devbox-version`'s human output
does not display `mempalace_version` and that only `--json` surfaces it. Both
halves are false — v1.8.6 shipped a `palace:` line with the same live-vs-baked
drift annotation `pi` already had. Verified by running the shipped script
against fabricated manifests: matching versions print `palace: 3.7.1`, a skew
prints `palace: 3.7.1 (baked as 3.8.0 — drift detected)`, and a pre-v1.8.6
manifest with no baked field prints the live value un-annotated. The bullet
appears to describe an intermediate state of the working tree and was never
re-checked before tagging. Left visible as a struck-through correction rather
than deleted, since v1.8.6 is already published and someone may have read it.
---
### Still open
- **Make the vendored-snapshot check automatic instead of a remembered string.**
Tonight's failure is the third iteration of the same maintenance burden (v1.8.4:
phrase present in both copies; v1.8.7: phrase deleted by the release that
refreshed the snapshot). A phrase canary structurally cannot answer *"is this
snapshot older than skillset main?"* — only a diff can. Proposed: a lint job
that clones the skillset repo and compares
`rootfs/usr/local/share/pi-devbox/skills/mempalace/SKILL.md` against it,
failing with the diff when they drift. Open question first: the skillset repo is
**private**, so this needs a CI clone credential, which is a policy decision
rather than a code change.
- **Provenance stops at Chroma's metadata.** The hourly reconciler on the palace
host stamps `device`/`agent_kind` in `chroma.sqlite3`, but knowledge-graph
triples and coordination events live in *separate* SQLite files
(`knowledge_graph.sqlite3`, `logstream.sqlite3`) it cannot reach. 156 triples
carry no origin field at all; `logstream`'s `from_agent` is free-form and
already inconsistent (`pi@tor-ms22`, `pi@emb-7kj4vr4g`, and bare `pi` in the
same table). Tracked in RFC 001 §7.3.1.
- **The stamp is self-asserted, and cannot be otherwise yet.** mempalace 3.8.0
authenticates with a *single scalar* bearer token and has zero device concept,
so a verified stamp needs per-device credentials plus an origin field in six
write paths across three databases. Deferred to RFC 001 Phase 4, where it is
now motivated primarily by **revocation** (one shared token covers every
device, so cutting off one laptop means rotating the fleet) rather than by
provenance. Forward-compatible by design: every stamp records *how* it was
determined, so an authoritative pass overwrites with `device_source='token'`
and nothing has to be undone.
- **`tor-ms22` and `tor-ms22-native` are one machine with two device values**
(4,680 and 3,826 rows). That is the hostname-as-identity cost RFC 001 §7.3.4
warned about, now visible in data: a rename splits one device's history
silently. Repairing it means a device-identity mapping, not a relabel.
## v1.8.6 — 2026-08-25
Patch release. Adopts the drift that accumulated in the ~2 days since v1.8.5
(pi `0.84.3`, mempalace core `3.8.0`), then closes the documentation and
observability gaps that v1.8.5 itself listed as "Still open". No component
was adopted without an audit note recording *why* it is safe.
All moving refs re-resolved immediately before tagging (2026-08-25T13:28Z):
pi-toolkit `0e1369e6`, pi-extensions `20228878`, mempalace-toolkit `0fe64c48`
and pi-observational-memory `ce9fc982` all unchanged since v1.8.5;
pi-fork `f1ff8087` → `bf702b4c`; pi-atelier holds at `v0.8.2` (floor for
pi ≥0.84 satisfied); pi-studio's CI-resolved newest tag has moved again to
`v0.9.51`. Base rebuild is forced (Dockerfile.base changed), so the 16
floating base-tooling ARGs re-roll — expect ~67 min as for v1.8.5.
### Changed
- **`mempalace` core `3.7.1` → `3.8.0`.** Released 2026-08-23T21:19Z, hours
after this project's own v1.8.5 tag the same day. Additive/reliability only
— reviewed for MCP tool-schema changes before bumping, as always: none.
`sync --apply` (PR #2320/#2322) no longer deletes a drawer solely because
its `source_file` was unreachable *at that moment* — it asks for
corroboration first. **This does not relax the standing landmine** against
running `mempalace_sync` / `mempalace_delete_by_source` beyond dry-run on
the shared central palace: that failure mode is paths *permanently* absent
from whichever host runs the sync, not transient unavailability, and 3.8.0
doesn't touch it. Server-side perf fix PR #2307 (long-running Chroma servers
no longer invalidate their own HNSW cache on their own writes) likewise does
not make `mempalace_reconnect` unnecessary — that tool covers *external*
writes bypassing the in-process client, a different scenario. Full reasoning
lives in the `Dockerfile.base` comment above `ARG MEMPALACE_VERSION`.
**Deployment note:** synlig's central palace currently serves `3.7.1`
server-side via `docker-compose.mempalace.yml` (which reuses this image) —
this client bump introduces version skew until that stack is separately
redeployed; sequence accordingly.
- **`pi` `0.84.2` → `0.84.3`.** Published 2026-08-24T11:09Z. Release notes
carry one "Breaking Changes" line — `GoogleThinkingLevel` renamed to
`GoogleApiThinkingLevel` — checked against all four vendored packages
(`pi-fork`, `pi-observational-memory`, `pi-atelier`, `pi-studio`): zero
references, inert here. 0.84.3 also fixes two skill-discovery bugs that
land directly on this repo's own vendored-skill work: nested Markdown
skills inside `.agents/skills/` grouping directories not being discovered,
and root Markdown files (`README.md`/`AGENTS.md`) in skill directories being
wrongly reported as broken skills.
### Added
- **Browser automation is now documented to humans, not just to agents.**
`agent-browser` + Playwright + a headless Chromium (~625 MB — the single
largest addition in the image) previously had zero mentions in `README.md`,
`DOCKER_HUB.md` or `THIRD_PARTY.md`; it existed only in the agent-facing
`AGENTS.md` managed block. Added a `README.md` "Browser automation"
subsection, a `DOCKER_HUB.md` feature entry, and `THIRD_PARTY.md` license
rows for `agent-browser` (Apache-2.0), Playwright (Apache-2.0), and Chromium
(BSD-3-Clause for Chromium's own code plus a large set of bundled
third-party components under their own licenses; the binary here is not
compiled by this repo — it's Playwright's own "Chrome for Testing" download
via `playwright install --with-deps chromium`).
- **`THIRD_PARTY.md` gains rows for `pi-atelier` (MIT) and `mempalace` core
(MIT per the GitHub repo; noted that the PyPI package's own metadata omits
a license classifier, so verify against the repo's `LICENSE` rather than
sdist/wheel metadata if clearance is needed from the artifact alone).**
- **`typst` and `socat` added to `README.md`'s tooling inventory.** Both were
already used in prose (typst as pandoc's `--pdf-engine`, socat by
`studio-expose`) but missing from the "What's inside" lists, so the
inventory didn't match what the image actually ships.
- **`mempalace` core version recorded in `/etc/pi-devbox/build-manifest.json`.**
Previously absent — a published image couldn't answer "which palace version
shipped?", and a palace bug couldn't be correlated to an image version.
Derived from the live installed binary (matching the manifest's existing
ground-truth-not-build-args philosophy), degrading to `null` rather than
failing the build if the binary is missing or its output format changes.
Verified landed: new top-level `"mempalace_version"` key, sibling to
`pi_version` rather than a member of `components{}` (that map is rendered
truncated to 12 chars by `pi-devbox-version`, which would mangle a longer
version string).
- **New smoke assertions**, all landed in `scripts/smoke-test.sh`: (1) the
`pi-observational-memory` clone is checked for the actual `ce9fc98`
auth-fix markers pinned to their fix site, `src/runtime.ts`
(`availability_recheck`, `providerCredentialConfigured`,
`hasConfiguredAuth`) — not merely clone existence, and deliberately not a
repo-wide grep: all three identifiers also appear under `tests/`, so a
repo-wide search would stay green even with the fix reverted in
`src/runtime.ts` alone; (2) the manifest's new `mempalace_version` field is
asserted present, non-null, and equal to what `mempalace --version` reports
live, so the manifest can't silently drift from the installed package —
expected to fail against any pre-v1.8.6 image, by design; (3) a
behavioural check for the mempalace-toolkit feeder's `--agent` default
(see below — this one turned out to be possible after all).
### Fixed
- **A false claim was being published to Docker Hub on every release.**
`DOCKER_HUB.md` advertised "neovim (LazyVim defaults)". Nothing in this
repo installs LazyVim — the only nvim configuration is a 19-line
`sysinit.vim` that sets `termguicolors`. `update-description` pushes this
file verbatim (with `{{PI_VERSION}}` substituted) to the Hub description, so
the error was public, not internal. Corrected to describe what's actually
there.
### Component audit for this release
Checked against upstream 2026-08-25 (two days after v1.8.5's own audit):
`mempalace` core moved `3.7.1` → `3.8.0` (see Changed, above — timing is
notable: released *hours after* v1.8.5 tagged, so v1.8.5 could not have caught
it no matter how carefully it was audited). `pi` moved `0.84.2` → `0.84.3`
(see Changed). `pi-toolkit` `0e1369e6`, `pi-extensions` `20228878`,
`pi-observational-memory` `ce9fc982`, and `pi-atelier` `v0.8.2` are all
**unchanged** from v1.8.5 — in particular `pi-observational-memory` still sits
exactly at the auth-fix commit with nothing landed upstream since, and
`pi-atelier` is still the newest tag with the `≥0.7.1` floor for `pi ≥ 0.84`
trivially satisfied. `pi-fork` has one upstream commit not adopted this
release: `f1ff8087` → `bf702b4c`, a text-only rewording of the fork task
preamble (no code-path change) — **left un-pulled** for this release since it
is a moving ref CI resolves fresh at every build anyway; it will be adopted
automatically on the next build regardless of this entry. `pi-studio` (studio
variant) has drifted two tags upstream, `v0.9.48` (pinned at build time via
CI's newest-semver-tag resolution) → `v0.9.51` at tag time, purely additive
(watched PDF previews, opening PDFs directly in Studio, Studio header
hide) — nothing to bump in this repo since studio-tag resolution happens in
CI, not the Dockerfile, but note it **will** auto-adopt `v0.9.51` on the next
studio-variant build. `mempalace-toolkit` unchanged — this release's manifest
and pi-bump work in `Dockerfile.variant` stayed within that file's ownership
and did not require a toolkit-side change.
### Still open
- **`MEMPALACE_VERSION` has no CI-side audit equivalent to `PI_VERSION`'s.**
`PI_VERSION` is verified published-on-npm and warns (never silently adopts)
on drift; `MEMPALACE_VERSION` is a literal Dockerfile string with zero
references in `.gitea/workflows/docker-publish.yml`. Flagged in v1.8.5's
audit as a gap; still a gap.
- ~~**`pi-devbox-version`'s human-readable output does not display
`mempalace_version`.** Its render path is a fixed sequence
(`release_tag`, `build_date`, `source_revision`, `pi`, then `components{}`)
and the new top-level field isn't in it — only `--json` mode (which `cat`s
the manifest directly) surfaces it today. One line in
`rootfs/usr/local/bin/pi-devbox-version` would fix this; deferred since the
field's stated purpose (correlating a palace bug to an image) is already
served by `--json`, but worth doing in a follow-up if this becomes a
routine manual check.~~
**CORRECTION (2026-08-25, post-tag):** this bullet is wrong and was never
true of the tagged tree. `pi-devbox-version` *does* print a `palace:` line
in human mode, with live-vs-baked drift detection, degrading quietly on
pre-v1.8.6 manifests. Nothing is open here. See the Unreleased entry above.
**Resolved during this release, not left open:** the feeder `--agent`
default behavioural hook initially looked like it might need a
mempalace-toolkit change (a `--print-config` flag that doesn't exist). It
didn't — `mempalace-pi-session` assigns `AGENT` before argument parsing and
`--help` exits 0 with no side effects, so `bash -x mempalace-pi-session
--help` observes the real resolution (env interpolation and fallback)
without needing a source change. The new smoke assertion exploits exactly
that, checked both ways: with `MEMPALACE_PI_DEVICE` set it must resolve to
`pi@<device>`; with it unset it must NOT be `pi@*` (catches a regression to
the old unconditional `$USER`/`mempalace` default).
`mempalace-toolkit` commit `c64ffa1` changed the feeder's `--agent` default
from `$USER` to `pi@<device>`, but there is still no way for smoke to assert
this default is actually in effect from this repo alone, since
`mempalace-toolkit` is a separate repo this release does not modify. If the
concurrent smoke-test work could not find an honest assertion from the
existing `/opt/mempalace-toolkit` surface (help text, `--self-test`), this
remains open pending a toolkit-side `--print-config`-style hook — a
toolkit-repo change, not a pi-devbox one.
- **16 base-tooling `ARG *_VERSION=latest` pins remain unrecorded.** (Corrected
count — v1.8.5's entry said "~14"; the actual count from `Dockerfile.base`
is 16, plus 5 more that float with no ARG at all: `rustup-init`, AWS CLI v2,
Chromium-via-Playwright, Node's minor version via `setup_22.x`, and
`DEBIAN_VERSION=trixie-slim` itself.) None of these are recorded anywhere
once the build completes — not in the manifest, not in a label — so a
published image cannot answer "which nvim/uv/chromium shipped?" without
exec-ing in and asking the binary.
### Documentation
- **`.env.example` documents `MEMPALACE_PALACE_PATH`.** It was the only MemPalace
variable the template never mentioned, while being the one that silently moves
the feeders' stage: the palace root resolves as `$MEMPALACE_PALACE_PATH` →
`$MEMPAL_PALACE_PATH` → `~/.mempalace/config.json` → `~/.mempalace/palace`, and
the stage is derived from it (`<palace-root>/pi-stage`). The comment states the
precedence, says why neither the image nor the entrypoint exports it (pinning
the palace without carrying the stage re-creates the split a shared root
removed — see v1.8.2), warns that a stage whose persistence differs from the
palace makes a scoped `mempalace sync` prune conversation drawers whose dedup
key is the staged path, and notes it is a *container* path unlike the
host-side `WORKSPACE_PATH`/`SSH_KEY_PATH` above it. Found while auditing a live
host whose `.env` sets the variable redundantly to the default.
---
## v1.8.5 — 2026-08-23
Patch release with two fixes in the container's skill wiring — one behavioural,
+8 -1
View File
@@ -65,12 +65,19 @@ The entrypoint deploys/registers all of these on first container start. Re-runni
### Document and image tooling
- **pandoc** — universal Markdown↔HTML/Org/RST/etc. conversion. Useful well beyond pi: agent-driven doc exports, format conversion, etc.
- **Typst** — markup-based typesetting, used as pandoc's `--pdf-engine`
- **graphviz** (`dot`) — diagram rendering pipelines
- **imagemagick** (`magick`) — image conversion / resizing
### Browser automation
- **agent-browser** — CLI for driving a real browser (open pages, click/fill/`eval`, snapshot the DOM, screenshots) so agents can verify front-end work instead of guessing
- **Playwright** + a headless **Chromium** are pre-installed and pinned together; `AGENT_BROWSER_EXECUTABLE_PATH` is preset to the baked browser, so `agent-browser open <url>` works out of the box with no setup
- **socat** — TCP bridge used to expose the pi-studio server outside the container's loopback
### Modern CLI tooling
- **Editor**: neovim (LazyVim defaults), tmux (configured for 0-indexed sessions)
- **Editor**: neovim (system-wide `termguicolors` default; bring your own config/plugins), tmux (configured for 0-indexed sessions)
- **Search/nav**: ripgrep, fd, fzf, zoxide
- **Display**: bat, eza, htop, tree
- **Data**: jq, yq
+42 -1
View File
@@ -396,7 +396,48 @@ ARG INSTALL_MEMPALACE=true
# (refuse the write) rather than fail open. Neither affects the container's
# normal MCP-server-plus-CLI-feeder pattern, which already serialised on the
# same lock under 3.6.0.
ARG MEMPALACE_VERSION=3.7.1
#
# 3.8.0 (2026-08-23, PyPI, released hours after this project's own v1.8.5 tag
# the same day) is additive/reliability only — reviewed for MCP tool-schema
# changes before bumping, as always: there are NONE. Two PRs matter:
# - PR #2320/#2322: `sync --apply` no longer deletes a drawer solely because
# its source_file was unreachable AT THAT MOMENT — it now asks for
# corroboration first. This fixes losing a whole mined project to one
# `sync --apply` while its volume happened to be unmounted.
# IMPORTANT — do not over-read this fix: it addresses TRANSIENT
# unreachability, not the standing landmine (documented in the operator's
# global AGENTS.md) against running `mempalace_sync` / `mempalace_delete_by_source`
# beyond dry-run on the SHARED central palace. On that palace most
# source_file paths are PERMANENTLY absent from whichever host runs the
# sync — a different machine's paths simply do not exist here, ever, not
# merely "right now". That is a different failure shape than #2320/#2322
# fixes. The landmine still stands; this bump does not relax it.
# - PR #2307: long-running Chroma servers no longer invalidate their own
# HNSW cache on their own writes (server-side perf fix). This does NOT
# make `mempalace_reconnect` unnecessary — that tool exists for EXTERNAL
# writes bypassing the in-process client (e.g. direct sqlite backfills,
# CLI commands against a running server), a different scenario #2307
# does not touch.
#
# CI-side audit (added after v1.8.6, closing that release's "Still open" item):
# resolve-versions now treats this pin exactly as it treats PI_VERSION — it
# reads the ARG from THIS file, refuses a non-concrete value, verifies the
# version is published on PyPI, refuses a YANKED release (an exact pin installs
# one silently under PEP 592), and WARNS — never silently adopts — when PyPI has
# a newer release. smoke-test.sh then asserts the installed core equals that
# audited pin, which catches a stale cached base layer that no manifest-internal
# check can see. So a bump here is now gated end to end; what remains manual is
# the JUDGEMENT above (MCP schema review, server/client sequencing), which is
# the part that should stay manual.
#
# Deployment sequencing note for whoever ships this bump: synlig (the shared
# central palace host) currently serves mempalace 3.7.1 SERVER-SIDE via
# docker-compose.mempalace.yml, which reuses this same devbox image. Bumping
# this ARG changes only the CLIENT version baked into pi-devbox images: it
# introduces client/server skew until synlig's compose stack is separately
# rebuilt/redeployed with the new pin. Not something to code around here —
# just sequence the redeploy.
ARG MEMPALACE_VERSION=3.8.0
ENV UV_TOOL_DIR=/opt/uv-tools
ENV UV_TOOL_BIN_DIR=/usr/local/bin
RUN if [ "${INSTALL_MEMPALACE}" = "true" ]; then \
+33 -1
View File
@@ -56,7 +56,22 @@ ARG USER_NAME=developer
# current when it was first populated (shipped the same bytes for pi-devbox
# v0.74.0..v0.75.5; discovered + fixed in v0.75.5b, 2026-05-23). The `latest`
# branch below is kept only for a deliberate local `docker build` override.
ARG PI_VERSION=0.84.2
#
# AUDITED AT 0.84.3 (2026-08-25, was 0.84.2): upstream's notes carry a
# "Breaking Changes" heading — `GoogleThinkingLevel` renamed to
# `GoogleApiThinkingLevel`. INERT FOR THIS IMAGE: all four vendored companions
# (/opt/pi-fork, /opt/pi-observational-memory, /opt/pi-atelier, /opt/pi-studio)
# were grepped for that symbol and reference it ZERO times, so nothing here
# couples to the renamed type. Recorded because the heading will look alarming
# to the next reader doing step 1 above — the audit is done, don't redo it.
# Adopted for two fixes that land squarely on this repo's own vendored-skill
# wiring (see devbox-skill-reconcile, v1.8.5): nested Markdown skills inside
# `.agents/skills/<group>/` directories were not discovered, and root Markdown
# files such as README.md / AGENTS.md inside a skill dir were reported as
# broken skills unless they declared valid skill frontmatter.
# pi-atelier needs no companion bump: v0.8.2 clears the >=0.7.1 floor that
# pi >= 0.84 requires (see PI_ATELIER_REF below).
ARG PI_VERSION=0.84.3
ARG PI_TOOLKIT_REF=main
ARG PI_EXTENSIONS_REF=main
# Repo URLs default to the canonical gitea origin but are overridable so a
@@ -297,6 +312,19 @@ RUN set -e; \
mkdir -p /etc/pi-devbox; \
rev() { git -C "$1" rev-parse HEAD 2>/dev/null || echo "unknown"; }; \
PI_V="$(pi --version 2>/dev/null | head -n1 | tr -d '\r\n')"; \
# mempalace CORE (the PyPI package behind the MCP tools) is installed in
# Dockerfile.base via `uv tool install`, so no /opt clone reveals it and
# until v1.8.6 the manifest could not answer "which palace shipped here?" —
# a palace bug could not be correlated to an image, which is precisely the
# correlation this file exists to provide. Read from the INSTALLED BINARY,
# not from ARG MEMPALACE_VERSION, per the ground-truth rule above: that is
# what catches an install which resolved to something other than the pin.
# `mempalace --version` prints "MemPalace 3.7.1" — NAME-PREFIXED, unlike
# pi's bare "0.84.2" — hence the $NF pick rather than a straight read. The
# leading-digit test then rejects usage/error text (a renamed flag prints a
# usage block) and degrades to JSON null, so this can never fail the build.
MP_V="$(mempalace --version 2>/dev/null | head -n1 | tr -d '\r' | awk '{print $NF}')"; \
case "$MP_V" in [0-9]*) MP_CORE="\"${MP_V}\"" ;; *) MP_CORE='null' ;; esac; \
STUDIO_REV='null'; \
if [ -d /opt/pi-studio/.git ]; then STUDIO_REV="\"$(rev /opt/pi-studio)\""; fi; \
{ \
@@ -305,6 +333,10 @@ RUN set -e; \
echo " \"build_date\": \"${BUILD_DATE}\","; \
echo " \"source_revision\": \"${SOURCE_REVISION}\","; \
echo " \"pi_version\": \"${PI_V}\","; \
# Sibling of pi_version, NOT a member of components{}: that map holds git
# SHAs and `pi-devbox-version` renders it with .value[0:12], which would
# silently truncate a longer version string.
echo " \"mempalace_version\": ${MP_CORE},"; \
echo " \"components\": {"; \
echo " \"pi-toolkit\": \"$(rev /opt/pi-toolkit)\","; \
echo " \"pi-extensions\": \"$(rev /opt/pi-extensions)\","; \
+21 -2
View File
@@ -70,9 +70,27 @@ so `TERM=xterm-kitty` is understood. Override either in your own
### Document and image tooling
- `pandoc` — universal Markdown↔HTML/Org/RST/etc. converter
- `typst` — markup-based typesetting, wired up as pandoc's `--pdf-engine` (see
[Generating a PDF with pandoc + typst](#generating-a-pdf-with-pandoc--typst))
- `graphviz` — `dot` rendering for diagram pipelines
- `imagemagick` — image conversion / resizing (invoked as `magick`)
### Browser automation
- `agent-browser` — CLI for driving a real headless browser: open pages,
click/fill/`eval`, snapshot the DOM, take screenshots. Useful whenever a task
involves a web UI or verifying how a page actually renders (live DOM, WebGL,
layout, popup positioning) instead of guessing from source.
- `playwright` + a pre-installed headless **Chromium** back it.
`AGENT_BROWSER_EXECUTABLE_PATH` is preset to the baked browser via a stable
`/usr/local/bin/agent-chrome` symlink (insulated from Playwright's
per-version/arch install directory), so `agent-browser open <url>` works
out of the box with no setup. Run `agent-browser skills get core --full`
for the command set and workflow patterns.
- `socat` — TCP bridge used by `studio-expose` to reach pi-studio's
loopback-bound server from outside the container (see
[Using pi-studio](#using-pi-studio--studio-variant))
### Language toolchains
- `python3` + `python3-venv` + `python3-pip` (system Python)
@@ -787,8 +805,9 @@ docker inspect --format '{{json .Config.Labels}}' joakimp/pi-devbox:latest | jq
`org.opencontainers.image.{version,revision,created}` plus
`se.jordbo.pi-devbox.*-ref` record the intended pi version and companion
refs. The on-disk `/etc/pi-devbox/build-manifest.json` records **ground
truth** — the actual checked-out commit of each `/opt` clone and the live
`pi --version` — so a tag is reconstructable after CI logs rotate:
truth** — the actual checked-out commit of each `/opt` clone, the live
`pi --version`, and (from v1.8.6) the live `mempalace --version` of the
installed palace core — so a tag is reconstructable after CI logs rotate:
```bash
docker run --rm --entrypoint= joakimp/pi-devbox:latest cat /etc/pi-devbox/build-manifest.json
+15
View File
@@ -18,8 +18,23 @@ for OS packages, the per-package copyright files inside the image at
| pi-fork | github.com/elpapi42/pi-fork | MIT |
| pi-observational-memory | github.com/elpapi42/pi-observational-memory | MIT |
| pi-studio *(`-studio` variant only)* | github.com/omaclaren/pi-studio | MIT |
| pi-atelier | github.com/michaelmjhhhh/pi-atelier | MIT |
| pi-toolkit, pi-extensions, mempalace-toolkit | authored by the maintainer (Joakim Persson) | MIT |
## MemPalace (AI memory)
| Component | Upstream | License |
| --- | --- | --- |
| mempalace (core, MCP server) | github.com/MemPalace/mempalace (PyPI: `mempalace`) | MIT — the GitHub repo declares MIT; the PyPI package's own metadata omits a license classifier, so if you need clearance from the package artifact alone, verify against the repo's `LICENSE` file rather than the sdist/wheel metadata |
## Browser automation
| Component | Upstream | License |
| --- | --- | --- |
| agent-browser | github.com/vercel-labs/agent-browser (npm: `agent-browser`) | Apache-2.0 |
| Playwright | github.com/microsoft/playwright (npm: `playwright`) | Apache-2.0 |
| Chromium | chromium.googlesource.com/chromium/src | BSD-3-Clause for Chromium's own code, plus a large set of bundled third-party components each under their own license (see Chromium's own `LICENSE`/`about:credits`). The binary in this image is **not compiled here** — it is the build Playwright downloads for its pinned version ("Chrome for Testing"), installed via `playwright install --with-deps chromium` at `/usr/local/share/ms-playwright/`. Treat Playwright's own distribution terms for that build as authoritative over any summary here. |
## Tooling baked into the base image
| Component | Upstream | License (best effort) |
+24
View File
@@ -55,6 +55,10 @@ release_tag=$(jq -r '.release_tag' "$MANIFEST")
build_date=$(jq -r '.build_date' "$MANIFEST")
source_rev=$(jq -r '.source_revision' "$MANIFEST")
pi_version_baked=$(jq -r '.pi_version' "$MANIFEST")
# `// empty` matters: images built before v1.8.6 have no such field, and
# `jq -r` renders a JSON null as the 4-char string "null" — which would
# print as a bogus version rather than being treated as absent.
mp_version_baked=$(jq -r '.mempalace_version // empty' "$MANIFEST")
if [ "$MODE" = "quiet" ]; then
printf '%s (%s)\n' "$release_tag" "${source_rev:0:7}"
@@ -71,6 +75,16 @@ if command -v pi >/dev/null 2>&1; then
pi_version_live=$(pi --version 2>/dev/null | head -n1 | tr -d '\r\n')
fi
# Same check for the palace, which matters more than it looks: mempalace is
# the one component that is BOTH client (here) and server (synlig runs this
# same image), so a skew between the two is a real failure mode rather than
# cosmetic. `mempalace --version` prints "MemPalace 3.8.0" — name-prefixed,
# unlike pi's bare "0.84.3" — hence $NF rather than reading the whole line.
mp_version_live=""
if command -v mempalace >/dev/null 2>&1; then
mp_version_live=$(mempalace --version 2>/dev/null | head -n1 | awk '{print $NF}' | tr -d '\r\n')
fi
printf 'pi-devbox %s\n' "$release_tag"
printf ' built: %s (source %s)\n' "$build_date" "${source_rev:0:12}"
if [ -n "$pi_version_live" ] && [ "$pi_version_live" != "$pi_version_baked" ]; then
@@ -79,5 +93,15 @@ else
printf ' pi: %s\n' "${pi_version_live:-$pi_version_baked}"
fi
# Printed only when known, so this degrades quietly on pre-v1.8.6 images
# instead of showing an empty or "null" palace line.
if [ -n "$mp_version_live" ] || [ -n "$mp_version_baked" ]; then
if [ -n "$mp_version_live" ] && [ -n "$mp_version_baked" ] && [ "$mp_version_live" != "$mp_version_baked" ]; then
printf ' palace: %s \033[33m(baked as %s — drift detected)\033[0m\n' "$mp_version_live" "$mp_version_baked"
else
printf ' palace: %s\n' "${mp_version_live:-$mp_version_baked}"
fi
fi
printf ' components:\n'
jq -r '.components | to_entries[] | select(.value != null) | " \(.key): \(.value[0:12])"' "$MANIFEST"
@@ -197,6 +197,45 @@ EOF
)
fi
# ── Multiplexing default, deliberately LAST ───────────────────────────
# Why this block exists: ControlPath above is forced, but ControlMaster is not
# set anywhere for targets that come from the user's own ~/.ssh/config. A target
# whose entry omits ControlMaster therefore opens a NEW TCP connection per ssh
# call, and an agent doing a dozen calls in a few minutes can trip fail2ban or a
# CGNAT flow-table cap on the far end — observed 2026-08-25: ~12 connections in
# 15 min and port 22 stopped answering while HTTPS to the same estate stayed fine.
#
# WHY IT IS AT THE BOTTOM, and ControlPath is at the top. ssh_config is
# first-value-wins, so position encodes intent:
# * BEFORE the Include = an OVERRIDE. Correct for ControlPath, whose value in
# the user's config points at read-only ~/.ssh and simply cannot work here.
# * AFTER the Include = a DEFAULT. Correct for ControlMaster, because an
# explicit per-host 'ControlMaster no' (or 'auto', or any value) in the
# user's own config must keep winning. We are supplying an opinion only
# where the user expressed none.
# That asymmetry is the whole design: force what is broken, default what is
# merely absent. It also means this needs no audit of anyone's ~/.ssh/config —
# which matters because that file is per-machine, differs across the fleet, and
# future machines' versions do not exist yet to be audited.
#
# Caveat worth knowing (and documented in the pi-devbox-environment skill): a
# stale master socket — file present, daemon gone, e.g. after the host suspends
# or changes network — makes every later ssh to that host hang. Recovery is
# 'ssh -F ~/.ssh-local/config -O exit <host>'. ControlPersist is deliberately
# short (10m idle, and each new session resets the idle timer) so an abandoned
# socket ages out on its own rather than lingering for hours.
MULTIPLEX_DEFAULT_BLOCK=$(cat <<'EOF'
# Multiplexing DEFAULT — intentionally after the Include above, so any explicit
# per-host ControlMaster in your own ~/.ssh/config still wins (first-value-wins).
# Applies only to targets that never mentioned ControlMaster at all.
# Stale socket after a suspend/network change? ssh -O exit <host>.
Host *
ControlMaster auto
ControlPersist 10m
EOF
)
cat > "$CONFIG" <<EOF
# AUTO-GENERATED by setup-lan-access.sh on every container start. Do not edit
# by hand — edits are overwritten. Used via: ssh -F ~/.ssh-local/config <host>
@@ -216,6 +255,7 @@ ${JUMP_BLOCK}
${LAN_CONF_BLOCK}
${AUTOJUMP_BLOCK}
${INCLUDE_BLOCK}
${MULTIPLEX_DEFAULT_BLOCK}
EOF
chmod 600 "$CONFIG" 2>/dev/null || true
@@ -79,6 +79,79 @@ mempalace_search(query="<keywords>", wing="<project>")
**Never guess about facts that might be in the palace.** Wrong is worse than slow. Say "let me check" and query.
#### Search Before You *Probe*
The rule above covers **questions**. This one covers **actions** — and it is the one
that actually gets skipped, because mid-task the impulse is to go and *look* rather
than to remember. The palace is a **fleet** record: another machine's agent has
usually already paid the cost of discovering how this environment is wired, and its
notes include the corrections that came afterwards, which a fresh probe cannot show
you.
**Before you SSH somewhere to find out how it is set up, enumerate infrastructure,
or derive a deployment — search.** Concrete triggers, all meaning *search first*:
- about to run `ssh <host> …`, `docker ps`, `systemctl list-units`, `ip addr` to
discover how something is deployed or connected
- about to establish topology: which hosts/runners/services exist, where they live,
which of them can reach which
- about to conclude "this isn't documented anywhere" or "there's no way to know"
- about to assert an environment fact you learned **earlier in this same session**
**That last trigger is the sharp edge.** A compacted session summary is lossy by
design, and a belief you formed 40 turns ago may already be *retracted* in the
palace by another machine. Trusting your own context over the shared record is how a
withdrawn claim gets re-published as fact.
Search broadly before narrowing — fleet knowledge often sits in another machine's
wing, or inside a mined conversation, not where you would file it yourself:
```
mempalace_search(query="<topic> <host> <mechanism>") # no wing filter first
mempalace_search(query="…", wing="<likely-wing>") # then narrow
```
Two or three searches cost seconds. Re-deriving infrastructure costs minutes **and
can be wrong**: a probe shows one host's present state, while the palace records
intent, history, and what was already disproved.
> **Worked example (real, 2026-08-25).** An agent evaluating whether to add an ARM
> CI runner probed hosts directly instead of searching. It concluded "the runner
> lives on synlig" — there are **four** — and that "synlig is on the home LAN" —
> it is an OpenStack VM with a public floating IP that cannot reach the home LAN at
> all. Both facts were already in the palace, the second one as an **explicit
> retraction of the very same mistake** made weeks earlier. The palace also held
> the runner labels and the deliberate `capacity: 1` setting, which the probe never
> revealed. Cost: a wrong recommendation written into the palace twice, then
> corrected twice.
**A search that comes back empty is not an answer — least of all about recent work.**
Semantic search is weakest exactly where the fleet record is freshest: a drawer filed
minutes ago is unranked against a keyword-shaped query, and the drawer you most need
is *by construction* the newest one, because the other machine files its release,
handoff and correction drawers at the **end** of its session. So a single miss proves
nothing. **If the work is 0-2 days old and the first search looks stale or empty,
enumerate before concluding:**
```
mempalace_list_drawers(wing="<wing>", since="<today>") # or room=, or no filter
mempalace_diary_read(agent_name="<you>", wing="<wing>") # the other machine's handoff
```
Enumeration is exact where embeddings are probabilistic. Treat "I searched and found
nothing" as a hypothesis you have not yet tested, and never as licence to go probing.
> **Worked example (real, 2026-08-25, same fleet as above).** An agent asked to
> orient on an in-flight release *did* search first — `"v1.8.6 release run 579
> Docker Hub verification"` — and got back only v1.6.4 / v0.78.0 era hits, because
> the release drawer it needed was **58 seconds old**. It accepted the miss and went
> off to probe Docker Hub and the Gitea API. The user had to prompt "maybe there is a
> note in mempalace"; `list_drawers(wing="pi-devbox", since=<today>)` then returned
> the drawer immediately, along with the diary entry naming the exact open item. The
> rule above was present and correct in this very file at the time — the failure was
> not knowing to *retry differently* after a bad first hit.
#### Mine New Projects
When working on a new codebase for the first time:
@@ -290,13 +363,14 @@ Zechner's pi-coding-agent). Implications:
- **Session feeders run on different schedules.** Pi sessions are fed Tue 03:00, opencode sessions Mon 03:00 (launchd `Weekday`: `0`/`7`=Sunday, `1`=Monday, `2`=Tuesday — misreading this by one day is easy). Recent sessions from either harness can lag the palace by up to a week, so absence-of-evidence in `wing_conversations` is not evidence-of-absence for recent work.
- **Reading another harness's diary is useful.** When orienting after a gap, `mempalace_diary_read agent_name=pi` (or whichever sibling agent has been active) often gives a fresher picture than waiting for the conversations feeder to catch up.
When the palace is **central** (shared across machines), five more things apply:
When the palace is **central** (shared across machines), these further things apply:
- **Check which machine a conversation came from.** Transcripts are fed per device, so `source_path` reads `…/mempalace-feed/<device>/pi_<uuid>.jsonl` while the displayed `source_file` is only the basename. One search can legitimately return hits from several machines at once — look at the device segment before attributing a decision to *this* project.
- **Attribute what you file yourself.** Drawers now carry `device` and `agent_kind` metadata (plus `device_source`/`agent_kind_source` recording *how* each was determined, so an inference is never mistaken for a fact). Mined content gets these for free — the inbox path gives the device, the filename shape gives the harness — and a timer on the palace host re-stamps hourly, because live re-mining replaces metadata rows and silently drops earlier stamps. But for anything **you** file by hand, the only signal is what you pass: set `added_by="<harness>@<device>"` (e.g. `pi@emb-7kj4vr4g`, from `$MEMPALACE_PI_DEVICE`) on `add_drawer`/`checkpoint`/`mine`. Skip it and your drawer joins the ~16k historic `/workspace` project mines that are permanently unattributable, because `/workspace` exists identically on every devbox. Note the palace preserves `source_file` in full (see `source_path`) but *displays* only the basename — so a device prefix there survives storage even though it looks stripped.
- **Provenance is stamped for you — leave it alone.** Drawers carry `device` and `agent_kind` metadata (plus `device_source`/`agent_kind_source` recording *how* each was determined, so an inference is never mistaken for a fact). You do **not** set these, and you no longer set `added_by` either: the pi bridge defaults the writer field to `<harness>@<device>` on `add_drawer`/`checkpoint`/`mine`/`event_append`/`artifact_put`, and prefixes diary entries with `HOST:<device>|`, from host-supplied `$MEMPALACE_PI_DEVICE`. RFC 001 §7.3.2 ranks "agent stamps it via a skill instruction" as the *worst possible* place for exactly the reason you would expect — it is per-call boilerplate that gets forgotten, and it did: the agent who wrote the previous version of this bullet then filed its own provenance drawer as `added_by=checkpoint`. Two things remain yours: pass `source_drawer_id` on `kg_add` (triples have no provenance field, so that pointer is the only path back to a device), and pass an explicit `added_by` **only** when deliberately filing on behalf of another device. Never invent values for `device`/`agent_kind`/`origin_device` — a fabricated value is worse than a blank, because it silently corrupts a future merge.
- **Metadata is invisible to search — so check the text, not the fields.** `search` results are built from a fixed key list and `diary_read` returns content, so neither ever shows `device`/`added_by`. Only `mempalace_get_drawer` reveals them. This is why diary entries carry an in-text `HOST:<device>` marker: it is the only attribution a reader actually sees. **A diary entry with no `HOST:` marker predates the convention and may be from any machine — do not assume it is this one's history.**
- **Mined drawers carry the MINE date, not the session date.** When history is imported, or re-mined on the palace host, `filed_at`/`created_at` is the *import* time — so sorting by them does not give chronological order. Real session time is recoverable from the UUIDv7 in `pi_<uuid>.jsonl`: the first 12 hex digits are milliseconds since the epoch (and UUIDv7 sorts lexicographically in time order, so a plain filename sort is already chronological). Agent-authored drawers and diaries have no such backdoor — for those `filed_at` is the only chronology, which is why it must never be restamped.
- **Beware the timezone mismatch when you combine those.** Palace `filed_at`/`created_at` are naive timestamps in the palace host's local time, while a UUIDv7 decodes to UTC. Comparing them directly introduces a silent offset (2 h for a CEST host). Normalise before drawing conclusions about ordering.
- **`agent_name` is not device-scoped.** `mempalace_diary_read(agent_name="pi")` returns *every* machine's `pi` diary, interleaved. Read the entry before assuming it is your own history.
- **`agent_name` is not device-scoped.** `mempalace_diary_read(agent_name="pi")` returns *every* machine's `pi` diary, interleaved. Read the entry before assuming it is your own history — and note that a container cannot tell you which machine it is on (`hostname` is a docker hash, `$DEVBOX_HOST_ALIAS` is generic). `$MEMPALACE_PI_DEVICE` is the cheap answer; `ssh -F ~/.ssh-local/config host hostname` is the independent one.
- **One writer, no queue.** A concurrent mine returns a structured `already-running` error rather than waiting its turn, and one large mine can make the palace unresponsive to every client for minutes. After another client's mine, call `mempalace_reconnect` to see the new drawers. A client-side timeout is not evidence of failure — verify before retrying, or you file a duplicate.
### Rooms
@@ -330,10 +404,13 @@ Entity-relationship triples with temporal validity. Query with `mempalace_kg_que
## Anti-Patterns
- **Don't guess when you can search.** If a question touches past work, search first.
- **Don't probe what the fleet already knows.** Before SSH-ing into a host, enumerating infrastructure, or deriving how something is deployed, search the palace. A probe reveals one host's present state; the palace holds intent, history and prior corrections — including the ones that contradict what you are about to conclude.
- **Don't trust this session's context over the palace.** A compacted summary is lossy, and another machine may have corrected the fact since. Verify load-bearing environment claims against the shared record before acting on them.
- **Don't take one empty search as proof the palace is silent.** Fresh drawers rank worst, and the drawer that matters is usually the newest one. For anything 0-2 days old, enumerate with `mempalace_list_drawers(since=…)` and read the other machine's diary before you go and probe.
- **Don't infer elapsed time from session or container boundaries.** A restart isn't a new day. Compare the actual timestamp (`timestamp` / `created_at`) against the current date/time before saying "yesterday", "last week", etc.
- **Don't skip the diary.** A session without a diary entry is a session forgotten.
- **Don't summarize drawer content.** File verbatim — the embedding model needs the original words.
- **Don't mine .git directories or node_modules.** The CLI miner respects .gitignore by default.
- **Don't create duplicate drawers.** Use `mempalace_check_duplicate` before adding manually.
- **Don't treat the palace as a task list.** It's for knowledge and context, not todos.
- **Don't hand-craft provenance.** Leave `added_by` alone (and never put a machine name in a diary's `agent_name` — it becomes the wing name and hides your entries from `diary_read`). Recording *which device* wrote a record is client/server infrastructure, not your job: a hostname or container ID is not a stable identity, and an invented value is worse than none because it silently corrupts any future palace merge. If you find notes in the palace describing an `origin_device` scheme, that is a design for the client to implement — not an instruction for you to start stamping.
- **Don't invent provenance metadata, and don't hand-stamp it either.** An earlier version of this list told you to set `added_by="<harness>@<device>"` by hand; that instruction has been withdrawn, because RFC 001 §7.3.2 places provenance at the client/server boundary and the pi bridge now does it uniformly (see *Provenance is stamped for you* above). DO NOT invent values for the palace's own metadata fields (`device`, `agent_kind`, `origin_device`): those are stamped by infrastructure that also records *how* each was determined, and a fabricated value is worse than none because it silently corrupts a future merge. DO pass `source_drawer_id` on `kg_add`. And never put a machine name in a diary's `agent_name` — it becomes the wing name and hides your entries from `diary_read`.
@@ -185,6 +185,16 @@ entrypoint's `setup-lan-access.sh` writes a **writable SSH sidecar** at
- A `Host *` block redirecting `ControlPath` into the writable `~/.ssh-local/cm`
(because `~/.ssh` is typically bind-mounted **read-only**, so a master socket
can't be created under it), plus `Include ~/.ssh/config`.
- A **trailing** `Host *` block supplying `ControlMaster auto` + `ControlPersist
10m` as a *default*. Position is the design: `ControlPath` sits **before** the
`Include` (an override — the value in your own config points at read-only
`~/.ssh` and cannot work here), while `ControlMaster` sits **after** it (a
default — an explicit per-host `ControlMaster no`/`auto` in your own config
still wins, because ssh_config is first-value-wins). **Force what is broken,
default what is merely absent.** Without this, a target whose entry never
mentioned `ControlMaster` opens a fresh TCP connection per `ssh` call, and an
agent making a dozen calls in a few minutes can trip fail2ban or a CGNAT
flow-table cap on the far end.
- Aliases **`host` / `mac`** → `host.docker.internal` (user comes from
`HOST_SSH_USER`) — i.e. SSH back into the Docker host.
- On VM-backed hosts only: an **SSH-jump-via-host** block so the container can
@@ -199,6 +209,25 @@ ssh -F "$HOME/.ssh-local/config" mac 'hostname; whoami' # reach the host
ssh -F "$HOME/.ssh-local/config" <lan-peer> '…' # reach a LAN peer (if configured)
```
**Always go through the sidecar, never `-F ~/.ssh/config`.** This is the single
easiest way to break SSH from inside the container, and the failure actively
misleads: the read-only path makes the master socket uncreatable, so
multiplexing appears *impossible* rather than misconfigured. What follows is a
burst of fresh connections and, on a rate-limiting peer, a block that looks like
an outage. The tell that it is rate-limiting and not an outage: HTTPS to the same
estate keeps working while port 22 stops answering. (Recorded 2026-08-25 — an
agent hit exactly this, concluded "ControlMaster is impossible here", disabled
multiplexing, and filed that as a lesson. The sidecar had solved it since v1.4.)
If every `ssh` to one host suddenly hangs, suspect a **stale master** — socket
file present, daemon gone, typically after the host suspended or changed
network. Check and clear it:
```sh
ssh -F "$HOME/.ssh-local/config" -O check <host> # "Master running (pid=…)" or no master
ssh -F "$HOME/.ssh-local/config" -O exit <host> # tear down a stale one
```
Two related mechanisms (don't reinvent them):
- **ControlMaster multiplexing** is preconfigured (`/tmp/sshcm/`) to survive
+170 -14
View File
@@ -5,6 +5,7 @@
#
# Verifies:
# - pi binary present and (if EXPECTED_PI_VERSION set) matches CI's resolved version
# - mempalace core matches the audited pin (if EXPECTED_MEMPALACE_VERSION set)
# - new v1.0.0 base additions (pandoc, graphviz, imagemagick, yq, tealdeer)
# - typst PDF engine for pandoc (Unreleased) — `pandoc --pdf-engine=typst`
# - non-modal editors nano + micro (alongside nvim)
@@ -151,6 +152,33 @@ run "pi stage follows MEMPALACE_PALACE_PATH" '
mempalace-pi-session --dry-run --reason smoke --sessions-dir "$(mktemp -d)" 2>&1) || true
echo "$out" | grep -q "stage=/tmp/alt/.mempalace/pi-stage/"
'
# The feeder's --agent default is WHO a drawer is attributed to. mempalace core
# records neither the machine nor the harness on a write, and one shared bearer
# token means the server cannot tell clients apart, so toolkit c64ffa1 changed
# this default from $USER to pi@$MEMPALACE_PI_DEVICE — the one string that makes
# a write attributable to both. Nothing ever PRINTED the resolved value (the
# banner shows mode= and stage= only), so an image built from a pre-c64ffa1
# toolkit ref would ship unattributed writes with every check still green.
#
# `--help` assigns AGENT (script top) before it parses args, then exits 0 with
# no side effects — so `bash -x` observes the REAL resolution, env interpolation
# and fallback included, rather than grepping the source for a literal line that
# any reformat would break. Two-sided on purpose: device set => pi@<device>;
# device UNSET => must not be pi@anything. The second half is what fails against
# the old unconditional $USER default, which ignored the device entirely.
#
# Probes the PATH entry (a symlink into the /opt clone) rather than that clone
# path directly: this is the invocation the systemd/launchd timers and
# entrypoint-user.sh actually use, so it is the default that reaches the palace.
run "feeder resolves --agent to pi@<device> (drawer attribution)" '
f=$(command -v mempalace-pi-session) || { echo "feeder not on PATH" >&2; exit 1; }
with=$(MEMPALACE_PI_DEVICE=smoke-device bash -x $f --help 2>&1 | sed -n "s/^+* *AGENT=//p" | tail -n1)
without=$(env -u MEMPALACE_PI_DEVICE bash -x $f --help 2>&1 | sed -n "s/^+* *AGENT=//p" | tail -n1)
echo "resolved with-device=[$with] without-device=[$without]" >&2
[ "$with" = "pi@smoke-device" ] || exit 1
case "$without" in pi@*) exit 1 ;; esac
echo ok
'
# Regression guard for the pi transcript exporter. If pi ever changes its
# session JSONL shape, the exporter stops recognising sessions and the palace
# silently gets nothing (or, worse, raw JSON chunked as prose). Feed it a
@@ -263,6 +291,26 @@ run "pi-fork clone + node_modules" \
"test -f /opt/pi-fork/package.json && test -d /opt/pi-fork/node_modules"
run "pi-observational-memory clone + node_modules" \
"test -f /opt/pi-observational-memory/package.json && test -d /opt/pi-observational-memory/node_modules"
# ...and that the clone carries the AUTH FIX, not merely that it exists. om's
# pre-flight hasUsableAuth() check silently disabled `recall` for ~8 weeks once
# pi moved to request-time SigV4 signing and stopped exposing a static Bedrock
# key; upstream fixed it in ce9fc98, adopted in v1.8.4. PI_OBSMEM_REF tracks
# master, so an upstream revert or force-push would ship a dead `recall` with
# the clone assertion above still green — the exact gap flagged as open in the
# v1.8.5 changelog.
#
# Pin the markers to src/runtime.ts, the fix SITE, rather than grepping the
# repo: two of these three strings also appear under tests/, so a repo-wide
# grep stays green with runtime.ts itself reverted. That is a false green of the
# same family as the old skill-snapshot canary.
run "pi-observational-memory carries the ce9fc98 auth fix (recall stays alive)" '
f=/opt/pi-observational-memory/src/runtime.ts
test -f "$f" || { echo "fix site missing: $f" >&2; exit 1; }
for m in availability_recheck providerCredentialConfigured hasConfiguredAuth; do
grep -q "$m" "$f" || { echo "marker absent from runtime.ts: $m" >&2; exit 1; }
done
echo ok
'
# pi-atelier: deliberately NO node_modules assertion, unlike its siblings —
# it declares zero runtime dependencies (only peerDeps, satisfied by the baked
# pi) and has no build step, so Dockerfile.variant skips `npm install` for it.
@@ -301,24 +349,118 @@ echo ""
echo "── Build provenance ──"
run "/etc/pi-devbox/build-manifest.json present" \
"test -f /etc/pi-devbox/build-manifest.json"
run_expect "manifest records pi-extensions component" \
"cat /etc/pi-devbox/build-manifest.json" '"pi-extensions"'
run_expect "manifest records pi-atelier" \
"cat /etc/pi-devbox/build-manifest.json" '"pi-atelier"'
run_expect "manifest records pi_version" \
"cat /etc/pi-devbox/build-manifest.json" '"pi_version"'
# These next checks replace three that grepped the manifest for the FIELD NAME
# and never looked at the value:
#
# run_expect "manifest records pi_version" "cat …manifest.json" '"pi_version"'
#
# which passes on {"pi_version": ""} and on {"pi_version": null}. The tell was
# visible in its own passing output — `✅ manifest records pi_version (got
# "pi_version")` echoes the key back as the thing it claims to have found.
# Two failure modes were therefore invisible: a key that survives with an empty
# or garbage value, and a key that vanishes from the manifest while every
# remaining value still looks fine.
#
# Those two need SEPARATE assertions, and the reason is a trap worth keeping in
# writing: an "every component value is a valid SHA" loop passes VACUOUSLY on
# components:{} — jq's all() over an empty list is true — so the value check
# alone would go green on a manifest that lost every component. Mutation-tested
# 2026-08-25 across nine fabricated manifests (empty map, deleted key, "",
# null, "unknown", 12-hex truncation, 40 non-hex chars, legit null pi-studio).
run "manifest declares every required component key" '
req="pi-toolkit pi-extensions pi-fork pi-observational-memory pi-atelier mempalace-toolkit pi-studio"
for k in $req; do
jq -e --arg k "$k" "(.components|has(\$k))" /etc/pi-devbox/build-manifest.json >/dev/null \
|| { echo "manifest lost component key: $k" >&2; exit 1; }
done
'
# Subsumes the old `! grep -q \"unknown\"` check ("unknown" is not 40-hex), and
# also catches "", null and truncated SHAs, which that grep let through. null is
# legitimate for pi-studio alone: the non-studio variant has no such clone.
run "manifest component values are resolved 40-hex commits" '
jq -e "
.components
| to_entries
| all(if .key == \"pi-studio\" and .value == null then true
else (.value|type) == \"string\" and (.value|test(\"^[0-9a-f]{40}\$\")) end)
" /etc/pi-devbox/build-manifest.json >/dev/null
'
# pi_version against ground truth, same shape as the mempalace check below.
# Chains with the "pi version matches build arg" assertion earlier in this file:
# together they tie build arg -> installed binary -> recorded manifest, so a
# manifest written from a stale variable cannot pass by agreeing with itself.
run "manifest pi_version matches the installed pi" '
m=$(jq -r ".pi_version // empty" /etc/pi-devbox/build-manifest.json)
b=$(pi --version 2>/dev/null | head -n1 | tr -d "\r")
echo "manifest=[$m] installed=[$b]" >&2
[ -n "$m" ] && [ "$m" = "$b" ]
'
# Top-level provenance fields: assert the SHAPE of each value, and only when the
# field is populated. source_revision and build_date legitimately default to
# empty (Dockerfile.variant ARGs) on a plain local `docker build`, so demanding
# them would fail honest local smoke runs; a populated-but-malformed value is
# the actual defect. release_tag defaults to "dev", so empty means a broken write.
run "manifest top-level fields are well-formed, not merely present" '
j=/etc/pi-devbox/build-manifest.json
t=$(jq -r ".release_tag // empty" $j)
r=$(jq -r ".source_revision // empty" $j)
d=$(jq -r ".build_date // empty" $j)
echo "release_tag=[$t] source_revision=[$r] build_date=[$d]" >&2
[ -n "$t" ] || { echo "release_tag empty (ARG default is dev)" >&2; exit 1; }
if [ -n "$r" ]; then
printf "%s" "$r" | grep -qxE "[0-9a-f]{40}" || { echo "source_revision not a 40-hex commit" >&2; exit 1; }
fi
if [ -n "$d" ]; then
printf "%s" "$d" | grep -qE "^[0-9]{4}-[0-9]{2}-[0-9]{2}T" || { echo "build_date not ISO-8601" >&2; exit 1; }
fi
'
# mempalace CORE was absent from the manifest through v1.8.5: the toolkit SHA
# was recorded but the palace version behind the MCP tools was not, so a palace
# bug could not be correlated to an image version. Assert the field exists AND
# equals the installed binary — recording it from ARG MEMPALACE_VERSION instead
# would look identical here yet drift silently the first time an install
# resolved to something other than the pin, which is the whole reason this file
# is built from ground truth. `// empty` matters: jq -r prints the 4-char
# string "null" for a JSON null, which would satisfy a naive -n test.
run "manifest mempalace_version matches the installed core" '
m=$(jq -r ".mempalace_version // empty" /etc/pi-devbox/build-manifest.json)
b=$(mempalace --version 2>/dev/null | head -n1 | tr -d "\r"); b=${b##* }
echo "manifest=[$m] installed=[$b]" >&2
[ -n "$m" ] && [ "$m" = "$b" ]
'
# ... and, when CI supplies it, that the installed core is the version CI
# actually AUDITED (published + not yanked on PyPI, in resolve-versions). This
# does NOT duplicate the check above, which compares two properties of one
# image and so cannot notice that BOTH are the wrong version. The live failure
# mode it covers: the variant builds `FROM` a base tag chosen by base-decide's
# content hash, so a bug in that hashing (the reason scripts/check-base-hash.sh
# exists) could reuse a cached base built from an OLDER MEMPALACE_VERSION pin —
# internally consistent, silently stale, invisible to every other assertion.
if [ -n "${EXPECTED_MEMPALACE_VERSION:-}" ]; then
run "installed mempalace matches CI's audited pin (${EXPECTED_MEMPALACE_VERSION})" "
b=\$(mempalace --version 2>/dev/null | head -n1 | tr -d '\r'); b=\${b##* }
echo \"installed=[\$b] audited_pin=[${EXPECTED_MEMPALACE_VERSION}]\" >&2
[ \"\$b\" = \"${EXPECTED_MEMPALACE_VERSION}\" ]
"
fi
# Every component must be a resolved commit (or null for pi-studio in the
# non-studio variant) — 'unknown' means a clone silently failed to resolve.
run "manifest has no unresolved ('unknown') components" \
"! grep -q '\"unknown\"' /etc/pi-devbox/build-manifest.json"
# pi-devbox-version wraps the manifest into a human-first command (this
# PR); verify the binary is present, executable, and both output modes work.
# non-studio variant) — now enforced by the 40-hex value check above, which
# strictly subsumes the old whole-file grep for '"unknown"'. Only rev() ever
# emits "unknown" and rev() feeds components only, so nothing is lost.
# pi-devbox-version wraps the manifest into a human-first command; verify the
# binary is present, executable, and that all three output modes work.
run "pi-devbox-version binary present + executable" \
"test -x /usr/local/bin/pi-devbox-version"
run_expect "pi-devbox-version human output shows release tag" \
"pi-devbox-version" "pi-devbox "
run_expect "pi-devbox-version --json round-trips the manifest" \
"pi-devbox-version --json" '"release_tag"'
# --json is a verbatim `cat` of the manifest, so "round-trips" is assertable
# literally. The old form grepped the output for the string "release_tag" — the
# key name again — which would pass on a truncated or re-serialised dump.
run "pi-devbox-version --json round-trips the manifest byte-for-byte" '
a=$(cat /etc/pi-devbox/build-manifest.json)
b=$(pi-devbox-version --json)
[ "$a" = "$b" ] || { echo "--json output differs from the manifest on disk" >&2; exit 1; }
'
run_expect "pi-devbox-version --quiet is a compact one-liner" \
"pi-devbox-version --quiet | wc -l" "1"
# OCI labels live in the image config, not the container fs — inspect them
@@ -385,7 +527,21 @@ exec_test "mempalace skill linked (fallback)" 'test -L $HOME/.agents/skills
# multiple harnesses", a phrase present in BOTH the stale and the fresh copy.
# A snapshot canary must pin the NEWEST section, so update this string whenever
# the snapshot is refreshed — that is the point of it.
exec_test "mempalace skill snapshot is current" 'grep -q "Attribute what you file yourself" $HOME/.agents/skills/mempalace/SKILL.md && echo ok'
#
# v1.8.7: this fired for real, and on the release that changed the snapshot. The
# pinned phrase was "Attribute what you file yourself", the heading of the
# instruction telling agents to hand-stamp added_by — which that same release
# WITHDREW (RFC 001 §7.3.2 ranks agent-side stamping worst-possible; the bridge
# now does it). So the canary correctly reported "snapshot changed, expectation
# did not", and blocked publication of an otherwise-green build (81 passed, 1
# failed, twice). Two lessons kept in the assertion itself:
# * it is now BIDIRECTIONAL — the new phrase must be present AND the withdrawn
# one absent, so a re-vendored stale snapshot fails just as loudly as a
# forgotten bump. A one-way canary only catches half the drift.
# * a phrase canary can only ever detect "older than what I remembered to pin",
# never "older than skillset main". The real fix is a CI job diffing this
# file against the skillset repo — see the Unreleased changelog note.
exec_test "mempalace skill snapshot is current" 'f=$HOME/.agents/skills/mempalace/SKILL.md; grep -q "Provenance is stamped for you" "$f" && ! grep -q "Attribute what you file yourself" "$f" && echo ok'
# Link TARGETS, not just link existence: with no skillset mounted (as here) the
# baked tree must be what resolves, for all three vendored skills.
exec_test "vendored skills resolve to the baked tree (no skillset mounted)" \